Speed will win the AI computing battle with Tuhin Srivastava from Baseten

21 Mar 2024 · 39 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: No Priors: Speed Will Win the AI Computing Battle with Tuhin Srivastava from Baseten

Episode Overview In this episode, co-hosts Elad Gil and Sarah Guo speak with Tuhin Srivastava, CEO and co-founder of Baseten, about the significant challenges and advancements in AI computing. The conversation centers around the importance of speed in AI products and how Baseten provides scalable AI infrastructures, particularly focusing on inference.

Key Participants

  • Sarah Guo: Co-host, startup investor, and founder of Conviction.
  • Elad Gil: Co-host, serial entrepreneur, and author of the High Growth Handbook.
  • Tuhin Srivastava: Guest, CEO and co-founder of Baseten.

Episode Highlights

Introduction to Baseten

  • What is Baseten?
  • Baseten is an infrastructure product offering fast, scalable AI solutions, especially in inference.
  • Launched four and a half years ago, it addresses challenges in AI development by focusing on efficient, scalable code rather than the no-code approach.

Key Discussions

The Importance of Efficient Code

  • Efficient Code vs. No Code: Tuhin emphasizes that while no-code solutions are user-friendly, they lack the flexibility and control that efficient code provides.
  • Focus on Empowering Engineers: Baseten aims to create intuitive abstractions that make powerful coding accessible without sacrificing the ability to customize.

Applications and Use Cases

  • Baseten serves a variety of clients from small startups to larger companies like Descript and Patreon, enabling rapid AI feature deployment.
  • Notable example: Planned AI, which builds SDKs for call centers, and Picnic Health, which uses AI for medical data extraction.

AI Workloads

Training vs. Inference

  • Differences in Workloads:
  • Training focuses on large-scale, less time-sensitive computations, while inference demands low latency and high reliability.
  • Inference workflows are often more repeatable and customer-driven.

Market Dynamics

  • Propensity to Buy vs. Build: A shift is observed where companies are more willing to purchase infrastructure rather than develop their own due to speed advantages.
  • Accelerating Market Demand: The increased focus on AI has led to surprising acceleration in market needs, with many enterprises rushing to adopt AI technologies.

Optimization and Performance

  • Driving Factors of Performance in Inference:
  • Tuhin discusses recent advancements in performance, including speculative decoding and better GPU utilization.
  • Baseten has been successful in achieving high throughput and low latency, crucial for real-time AI applications.

Future of AI Infrastructure

  • Mass Scale Adoption of AI: Tuhin speculates on the timeline for enterprise adoption of AI, suggesting that while immediate growth may be limited, a significant uptick is expected in the coming years.
  • Changing Hardware Landscape: The conversation touches on GPU availability and the ongoing challenges and improvements in accessing necessary computing power.

Defensibility in AI

  • Discussion on how companies can build defensible positions in a rapidly evolving market where numerous startups are emerging simultaneously.
  • Emphasis on understanding the infrastructure as a competitive advantage and recognizing that speed and efficiency in deployment are crucial.

Key Takeaways

  • Speed is Critical: Companies that can deploy AI solutions quickly will have a significant advantage in the market.
  • Adoption Trends: The trend is shifting towards purchasing AI infrastructures rather than building them, highlighting the importance of speed and scalability.
  • Performance Optimization: Continual improvements in software and hardware are vital to reducing latency and enhancing AI application efficiency.
  • Market Evolution: The AI infrastructure landscape is evolving rapidly, opening doors for both opportunities and challenges.

Conclusion Tuhin Srivastava's insights into the importance of speed and efficiency in AI computing provide a comprehensive overview of the current landscape and the future direction of AI infrastructure. The episode emphasizes the strategic decisions companies must make regarding their AI capabilities to remain competitive in an increasingly digital world.

---

Show Links

  • [Baseten](https://www.baseten.co/)
  • [Benchmarking fast Mistral 7B inference](https://www.baseten.co/benchmarking)

Contact

  • Feedback: show@no-priors.com
  • Follow on Twitter: [@NoPriorsPod](https://twitter.com/NoPriorsPod), [@Saranormous](https://twitter.com/Saranormous), [@EladGil](https://twitter.com/EladGil), [@tuhinone](https://twitter.com/tuhinone)

Subscribe

  • Follow the show on [Apple Podcasts](https://podcasts.apple.com/us/podcast/no-priors/id1685863667), [Spotify](https://open.spotify.com/show/7zwnc0J2dUkIy9f8jUO3Bx), or your favorite podcast platform.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Hi, listeners. Welcome to another episode of No Priors. Today, Alad and I are catching up with Tuhan Srivastava, the CEO and co-founder of Base10, which gives teams fast scalable AI infrastructure, starting with inference. They're one of the players at the center of the battle heating up around AI computing. Welcome, Tuhan. Hi, thanks for having me. Good to see you guys. Let's start at the beginning. For any listeners who don't know, what is Base 10 and how did you start working on it? Base 10 is an infrastructure product, so we provide fast scalable AI infrastructure for engineering teams working with large models.

0:36Currently, we're focused on inference and we want to do a lot more after that. But for the past, I'd say, four and a half years, actually, oh, that's a long time. For the last four and a half years, we've been cutting our teeth and trying to build this thing. I think it's been pretty rewarding over the last 12 months seeing the market kind of show up and everyone get equally excited about AI infrastructure. We started this honestly because, firstly, we thought ML was pretty cool in 2019. We thought it was going somewhere. And we wanted to build a picks and shovel business and kind of solve the problems that we were running into.

1:11I think the side note here is that I wanted to start a company with my friends. You often say that base 10 isn't no code, it's efficient code. Like, why does that difference matter? That wasn't always the case, I'd say. I'd say like, you know, those times when we had elements which were definitely a bit no codey. I think what we've learned over the last three or four years is, you know, code is just incredibly powerful and engineers want to write code. Even in its best form, you know, you want to build really, really tight abstractions. But I think the ability to turn the knobs under the hood is very, very important.

1:45I think no code kind of makes that a lot harder. I don't think it removes it, but it makes it a lot harder. So what we do is just build very strong intuitive abstractions that try to make the easy things super easy and still make the hard things possible. So, you know, you can get a lot of value really quickly. But I say, unlike a lot of other infrastructure products that have been built over the last 10 years, we're trying to solve against the graduation problem, which is that, you know, that we're able to support teams as they grow on scale. And just to sort of make it a little bit more visceral for our listeners, like what are the types of applications that run on Base 10?

2:22Like what's the scale of the platform? Do you have a favorite application? Everything from, you know, tiny side projects on weekends, all the way to companies that are pretty AI native. We've supported foundation model companies. We work with companies like Descript, where AI is very, very core to the product experience. We power a lot of AI features that Patreon has shipped. But I'd say some of the more interesting use cases from our perspective, actually, from my perspective at least, are either the really small teams that we're giving a lot of leverage to so that they can ship things very quickly.

2:55So a really good example of that might be a company like Planned AI, which is basically building an SDK for call centers is how I describe it. But, you know, they're able to ship models and, you know, co-locate workloads so that they can get, you know, sub 300 millisecond or sub 200 millisecond responses without, you know, months and months of infrastructure effort. But I think, on the other hand, it's really exciting to see companies become AI-enabled. That's where we see a lot of the value is going to be over the next decade. If I look at a company like Picnic Health, which has actually been around for a decade, and it's starting to do a very, very interesting thing with this corpus of data that they've gathered over the last 10 years, and supporting those use cases, I think their model is called Picnic GPT, which extracts information from medical records.

3:44And to me, those are the really exciting use cases where you're giving leverage to companies that are good at the domain that they are working in. Their model might be proprietary, their data might be proprietary, but the infrastructure doesn't necessarily need to be proprietary. And we can give them just an easy way to deploy that stuff without many, many people months. It's become like in vogue to compare the size of your GPU cluster. Like people are spending a lot of money on GPUs. We hear about 600 ,000 H100 equivalents and lots of venture rounds being raised often to train, you know, large models in some domain or another or even, you know, more and more expensive post-training.

4:29Are training and inference workloads different? Yeah, I think so. I think they just have very different, almost like SLAs for the customer. Like, you know, things that matter for inference are things like, you know, your cluster is somewhat co-located with where you're doing your work, whereas training, you know, stuff like that matters a bit less. It doesn't really matter where that training is happening as long as all your GPUs are somewhat together. You know, even the GPU clusters themselves, like, you know, the full training networking is a very, very important piece to have networking on the racks themselves, where Inference matters a little less because you're doing a little bit more on individual GPUs and less so across GPUs.

5:13I think from a user perspective, there's a lot more workflow, I'd say, in Inference that's repeated across customers as opposed to, you know, you guys work with a bunch of companies training models. You know, the state of the art really is, give me some SSH keys and let me go at it. Whereas with Inference, there's definitely repeated workflow where people are trying to get similar things out of the inference infrastructure, whether that be version management, whether that be the way they deploy it, hooking it up into CICD, cold starts, and so on and so forth. So I'd say it seems a bit more repeatable today how the problem is being solved by customers.

5:53I'd say the hardware requirements are quite different. They're probably a little less to some degree. But I think resiliency and reliability matters a lot more. you know, downtime is unacceptable from an input perspective. Nodes get terminated all the time from a training perspective. You were quite early to this market, and I think you folks pioneered a lot of the sort of early ML sort of infrastructure for these sorts of use cases and applications. What has been the most surprising thing, or what did you least expect relative to how things evolved? Yeah, I think I can answer that question to two different abstinence.

6:28I think you can answer that question from like a market perspective, which is that, I think, Eli, you have some old writing, which is basically markets or all that matter. I think we've felt that very viscerally in some sense, which is that you can build all this cool stuff. And then when markets show up, you feel that. And that really pushes the customer forward and the needs for your product forward. So that's one thing, which is the acceleration we saw through the end of 2022 and in 2023. definitely took us by surprise. I can be really honest and say that from 2019 to 2022, it was pretty quiet.

7:10We had happy customers, but the demands weren't necessarily there. I think from a practitioner perspective, how fast some of these teams move has really shocked me, which is I think what is really clear in AI and early-stage AI in general, and I think the enterprises are waking up to realize this right now, is that speed is actually your number one advantage. Things are moving so fast that if you're not competing on speed, you're going to be left behind. And so there's actually a lot of propensity to buy versus build, I'd say, which is people are happy to buy technology. where I'd say in the past people were pretty hesitant to buy infrastructure.

8:01We talk with companies all the time where we think that, oh, they probably have something built out where they're a lot less sophisticated than you think and they're handling a lot more scale. So they need to be able to have the infrastructure to support that. So I think that's probably one of the larger things that we've been surprised by, which is how fast people need to move to be relevant. I think the other thing, which is just like how GPU needs have evolved going from, you know, at the end of 2022, most of our customers were using T4s and AT &Gs. You know, after that, it changed to A100s.

8:37Now it's gone to H100s. Like the compute needs aren't necessarily going down. They're only going up, especially as these services scale up. You guys were just highlighted for leading on the independent artificial analysis.ai benchmarks for highest throughput and lowest latency serving. Congrats. Can you give us some intuition for what is driving that? Like, you know, what makes inference hard to run fast? Inference is, you know, I think there's like multiple things which are quite difficult about inference. You know, I think there's a workflow headaches, which we've talked about a bit. I can talk more about this, like the scalability and reliability bottlenecks.

9:15I think the stuff you're talking about is performance optimization, which is really like how long does one generation take? I think, you know, there's a lot of work and research that's being done to run generations as fast as possible to get the maximal throughput and the minimal latency. I think historically, historically, well, I mean, it's funny when I say historically because I mean over the last six months. But over the last six months, a lot of work has been done in the research community to get basically these things to move faster, stuff like speculative decoding, which came out sometime in the last six months.

9:52It started really being used. For us, what it means is how well you can use the GPU, how you can scale across multiple GPUs, and how you can honestly be really, really up-to-date with the latest things that are happening in open source and research. We've partnered with a company, with NVIDIA, with a company. We've partnered with NVIDIA and really worked really closely with their LLM engine called TRT-LM. And that's actually driven a lot of the performance gains that we've worked with and we've contributed to that, we've forked that. But the hard thing there, a lot of the optimization you're doing is pretty low level and there's no real abstraction.

10:42So you either have to learn how to use OpenSource or rewrite some of these kernels by yourself. If you look at something like OpenAI, what do people complain about? A lot of the time it's speed. It's speed. And that's probably one of the core performance advantages of open source is that you can get these smaller models to run faster. And I think that will continue to be a massive focus for us going forward as well. On the benchmarks, I think it's pretty crazy how that's evolved as well. I think we've gone from state-of-the-art being 90 tokens a second, then it got over 100, now it's over 200.

11:22Now we're talking for some people up to 300, 400. And I think that's going to continue to be a very, very important place to innovate. And we think over time, it will get somewhat commoditized, the performance, especially for language models, to be honest. I think more and more of that stuff should run locally to some degree, I think. But being on top of it and making sure that we're kind of attached to the state of the art is, if we're not, existential risk to the business. And so we have to do it. How much optimization have you been seeing for other types of models? So diffusion models, some of the Texas speech models, other areas like that.

12:08I'm just sort of curious. There seems like there's different types of optimizations happening across different foundation model types as well. So I was curious, you know, what's state of the art there and how you're thinking about it. I don't have the metrics on hand, but we're seeing and, you know, also pushing limits there as well. So just yesterday, for example, you know, we were able to get Whisper running. So there's Whisper, there's faster Whisper, and there's Whisper on TRT, which again is an NVIDIA thing. I think what we are seeing is that there's more and more focus on bringing these experiences to real-time as possible.

12:41And so one of our customers is a company called Gamma, AI-powered storytelling software. They use stable diffusion image models to generate images. Getting that, not waiting four or five seconds there and still having high-quality images is, again, core to their business and making it very, very fast and easy to use. And I think, you know, we are seeing like definitely from a customer requirement perspective, we're seeing that. I think, you know, have we been able to juice as much there? Not yet, but I think we're getting there. Speaking of the applications, you know, driven by these models still being generally startlingly slow, including the like really amazing capable ones from ChatGPT to Cognition to, you know, things like Pika and MidJourney.

13:33I mean, like the, you know, in a way that consumers have not seen in many years, we are waiting, you know, seconds for interactions. Is your view of that like it'll change because the models are getting smaller? It'll change because more smaller models will get more powerful. People do distillation. People just get better at running these things. Like we'll get better hardware. What's the path to like not waiting 10 seconds a generation, two minutes a generation? If you look at so many of the gains of running this trial fast, let's take that as an example, the step function gains come from a few things.

14:14The first one is running on H100 and not running on A100. So I think as hardware gets better and better, you get this almost like leg up. And so as hopefully those prices go down and that gets more available, that'll be one thing. Hopefully the H200 comes up next and we do that. I think the next piece is around software optimization. So stuff like continuous batching and dynamic batching and speculative decoding, which basically makes it easier to either parallelize or batch process a bunch of things or makes each individual generation faster or offsets it, offloads it to another model. There's lots of stuff in that place.

14:52I think these models are also just going to get smaller, to be honest. I think that's the really, really powerful small model that does one thing, that's pretty exciting. I think stuff like Ollama is very interesting. But I think if you saw, Sourcegraph has this where they basically announced that Cody is now running locally for a lot of customers. I think that's a pretty exciting proposition. We have to figure out where we sit in the world like that. It's not necessarily amazing for cloud providers like ourselves because we want everything to run on the cloud. but I think things do get smaller.

15:31Things do get more efficient and we have more distilled, more powerful models, sharper models to do. Yeah, it's the unbundling of models. How and when do customers choose to deploy their own models or open source models or fine-tuned open source models on their own infra versus use public model endpoints? Or like what guidance would you give people? This is a general trend, right? Which is that you got to open out, you got to Anthropic, you have an API that works. Really quickly, you're like, that's too slow. or that's too expensive or I don't need something that powerful. So we often see, if customers have a lot of money, they end up going, paying for private deployments in Azure, they still have the sloniness issue, then they go to open source models.

16:13When you're starting out with open source models, you're going to go to a shared endpoint. There's lots of great shared endpoint providers. But there are things that might matter to you which that shared endpoint providers can't give you. Firstly, you want your own SLAs. So you don't want kind of this noisy neighbors problem. If me and Elad have two competing apps and Elad apps get slammed, I don't want my app to slow down. My model calls still need to be fast. So in that case, you might want dedicated compute. But there's also like data and privacy stuff. Maybe you don't want your data running on the same infrastructure, data going through the same infrastructure that other folks' infrastructure is going to.

16:56And honestly, maybe even at some scale, it's cheaper to run it yourself. And they're the three things. I think then when you go to larger companies, because of no brain, like shared endpoints don't really work for large companies. They're definitely not going to work for enterprises. And a lot of times you're going to want a, you might want Mistral 7B not coming off the shared endpoint provider, maybe not even on a dedicated endpoint. You might want it self-hosted within your own AWS and GCP. And that's kind of actually where we see a lot of the world going as well, which is that especially larger customers, they're going to have actually pretty good compute deals.

17:32And they're going to have their own credit system or spend commit with the marketplaces. And actually running it on the infrastructure solves a lot of problems and has a massive cost advantage to them as well. And so I think there's like three stages to it, which is like you started shared inference endpoint providers, you go to dedicated in the cloud. But I think for some customers, that's not enough either. And you want to go into your own cloud. One prediction on the enterprise side would be that, you know, if ChatGPT only launched 15, 16 months ago and GPT-4 just came out a year ago, then most enterprises are still in a planning cycle and they haven't really adopted AI at any real scale.

18:12Which means for infrastructure providers like Base10, it's like a huge opportunity that's about to come, right? Already you're cresting this giant wave and the wave is about to get 10 times bigger potentially. Ashley, what do you view as the timeline for really mass scale enterprise adoption of AI? And where do you think things will be in terms of order of magnitude usage overall a year from now, two years from now? I'm just sort of curious, like what your view of that future is. I think it's a good question. I think we've been so wrong with time, with time scales here. But I do think like what we see it right now is that when we go to talk to enterprises, I'd say a lot of them, you know, like for honestly, what we're seeing now is that like co-pilots, especially co-gen stuff, that actually has already made its way into the enterprise.

18:59Like most enterprises we talk to, like when you say, when you ask them how advanced are you on the AI strategy, they'll tell you, well, everyone uses co-pilot. That's like the first big foray. I think the next piece then is like using like OpenAI or Anthropic in some way. Well, I think that is going on right now. And I think people are starting to experiment with that. I think the fear, I would say, someone was just telling me that I think it was Pfizer, Ian Mark, tens of millions of dollars for ML investment or AI investment over the next 12 to 18 months. That's kind of frightening to me in some degree.

19:36I think it's great. It's great for us. I love to hear it as a business builder here. But it's also like, to me, that means that the pressure is actually coming from the above. And, you know, that's kind of like, I'd say, the ML trap we fell into in 2018 to 2020 when CIOs were buying software. And it wasn't really attached to real user value or product value. And so, like, I would actually say that we're probably overestimating how big enterprise will get in the next 12 to 18 months. But we're underestimating where we'll be in, you know, three to five years from now. And I think 10x is like a pretty massive underinvestment.

20:12underinvestment. I'd say what we see, we're working with a customer right now that has four engineers and has, by the end of this year, will have mid-hundreds of thousands of dollars annual spent. That's for one use case with sub-thousand users for them. And they're already cashflow positive as a business, which is insane to me in general. But I can't even imagine what these workloads are going to look like when we get to the enterprise. When you start to think about, you know, take a customer service and chatbots is probably like the number one place where people think about efficiency. The volume that some of these customers, you know, my brother is the head of AI at Sunrun, which is like a public company that does solar panels.

21:07And I think that's like another really good example of a company where there's so much opportunity for so much opportunity for AI just to eat away at processes. And the volume is just so much higher than what we're thinking from like a traditional business perspective with harder requirements, which even drive spending higher. Yeah, I think one reset in stance that people should have on spending here is that traditionally, people were looking at software companies. you got really concerned as a software investor if your cost of goods sold was affected by a lot of like data processing, basically.

21:49And so, you know, you have this like expectation that your average SaaS business at maturity might have 80 percent gross margins. And I think like, you know, now people understand that the training businesses have a big upfront CapEx investment. that, you know, may or may not pay off. But I think one of the things that you're pointing out is that you can actually spend a lot on the inference and the core intelligence and that actually, you know, end up with a very valuable business on the other end with perhaps fewer people. And so I think people have talked about that shape of company, but they don't really think of it as a norm yet.

22:29And, you know, at least in my portfolio, we are seeing more efficiency on headcount and like a lot more compute spend. And I know for some of the base 10 customers, that compute spend for inference is actually like the, you know, one of the largest items on the P &L. We were working with this customer and we basically, and this was like a challenge for us. So we asked for an upfront payment for the year of compute. And the CTO came back to us and said, hey, I appreciate what you're doing, but no, this is, after payroll, this is our second biggest expense for the year. We're not going to do that.

23:16And I think that's somewhat indicative of how much spend there is here. And it's probably also somewhat indicative of how big I personally think the market can be once we start seeing mass scale adoption. But I do think, I think that's a good point, Sarah, which is I think it's somewhat of a reset in terms of, I don't know if this looks like normal SaaS business. Actually, I know, sure, it does not. And I think even like traditional multiples, it's really, really hard to think about. And what's crazy about it though is I think the most efficient businesses through markups and through software optimization can actually drive pretty healthy margins and still have these really aggressive consumption contracts.

24:07And I think that's, I don't know, I think that's right. I can't think of, I feel like you guys see more businesses than I do, I hope. And then, but I hope, but you guys will be able to chime in on that more. Is that unique to this industry? Where else have you seen that? It's been a while since I've seen so many companies ramping so quickly. And sometimes they were fake ramps. So like, you know, in the internet wave of the 90s, it was kind of startups selling to each other and kind of bootstrapping off of venture capital. and then there was giant telecom buildouts on like a five-year cycle that caused huge revenue uplift and then suddenly there was a glut and things dropped dramatically here it feels like things are ramping really fast off of products that are a couple months old which sometimes suggests that there's not defensibility and so then the question starts to become okay how do you build defensibility and what does that mean and how do things get commoditized and do they and you know so there's a couple different markets where suddenly you see three companies all go from zero to five or zero to 10 million of revenue in a year.

25:16And then you're like, okay, there's three of these companies and they all ramped at the same time, the same amount. And so there's enormous demand. But what does that mean in terms of, do they cannibalize each other? Can three more entrants come in and do the same thing? Like what, what is the basis for competition in that market? And so I think there's a lot of that happening too, which is at least for me, pretty unexpected. And I think it's just because we have such a big technology capability shift that suddenly you can do things you literally couldn't do a year ago. You know, it's kind of amazing.

25:44I think it's particularly exciting when you go and apply that to, I think, Eli, you've cut your teeth on a bunch of different healthcare initiatives. Healthcare is like, you know, a really interesting place where you look at, you know, like Nuance, if you remember Nuance Technologies, like, you know, they had the stranglehold over this market for years. And honestly, it always looked like they were kind of struggling all along the way as well. and then Whisper comes along and you see that market now of note-taking for the medical thing. It's insane how fast it's going and there clearly is real value there.

26:22And then I think the question actually goes, maybe this does look like a SaaS business again when you're like, okay, what's the workflow? And is the power and defensibility of the workflow you're powering? Yeah, it's a really good point. And then the other thing I think that people often forget is that many markets are not monopolies. Many of them are oligopolies. You know, that's payments with Stripe and AdGen and PayPal and all these things, right? And so it's also possible some of these market structures are oligopoly markets. And then it's possible it actually ends up being win or take all.

26:50And there's some network effect or data effect. But if you look at some of these types of companies like healthcare, to your point, is a great example where it's deal driven, right? You have large deployments with big customers and you lock them in for multi-year deals. and if you're actually able to lock down customer bases and effectively you you can fragment in an oligopoly market more easily than if you have renewables every year right so i think also part of it is just like what's the contractual structure of a market and people really don't talk about that kind of stuff but i think it's really fascinating to think about through the lens of what actually is a sustainable business in each each one of these categories you know beyond health care it's like where are these businesses going to get disrupted um i think like financial services is an obvious one.

Read the full transcript

27:29Funnily enough, I think they've been at the cusp of this stuff in the past. I don't actually think they're on the cusp of it as much as you'd think. If you think about a lot of the big data stuff like 10 years ago, the hedge funds were all over that. They're like, hey, there's alpha here. And I know that some of them are science and look at large models and language models and whatnot. But I do feel like they're actually being a bit laggard in terms of their adoption of these things. It might actually be because they were so deep in the other sphere, in the old ML world, that's hard to kind of really quickly turn things around.

28:07It's just such a different capability set that I think like old school machine learning or where you're just effectively doing regressions and just pulling out patterns and data is kind of different from some of the generative stuff in terms of what it does and what it can do for you. And one of the things I've been thinking about recently related to what you just said is, what are the companies that just don't care about this? And that would be a very good thing because they're defensible, right? In the era of AI eating everything, like what can't be eaten? And therefore maybe those are really good things to get involved with or to work on because you're not threatened by a dozen different new startups.

28:40That becomes really hard when you start to think about some of the demos we've seen over the last couple of weeks. I thought my job was safe until yesterday. Yes. One of my partners asked me how long I thought venture capital was going to last in terms of like, you know, an agent-based automation taking over because he was like all excited that he got out of software engineering at exactly the right time before his skills became useless. But I tried to give him, I tried to give him a real answer, which is I think on the early stage, at the early stages, a lot of the data doesn't exist, right? Like you'd have to capture real world data.

29:18You have increasingly meetings over Zoom, but you would want to capture a lot of information about people. So much of it is access and the information about who is like leaving and like a 100x engineer and entrepreneurial and product oriented. And it works with velocity. Like a lot of that is not collected today. Right. So you have this big inputs and there's no digital trail for it. And so you have this big inputs problem. I think the decisioning, like if you think about like what is actually structurally predictable, maybe if you have all that data, the like people are the most identifiable piece.

29:57but the um maybe you can maybe you have a model that is doing continuous learning and can and learn like metastructures like a lot is talking about like oh this is a market that operates as an oligopoly where these the core um uh core drivers of you know differentiation and these the dimensions of competition and such but i think that that feels quite hard when you're investing in a technology landscape that is always changing. Right. And so like you're kind of always out of distribution and you don't have the data on the people. And like, I don't know how you make decisions on whether products are any good because you'd have to have all the customer point of view or you'd have to have taste.

30:37Maybe models will have taste. So I think Sarah is saying that her job is defensible, which I think is what everybody says. No, no, no, not my job. I am I'm just saying this morning I gave it a good think and I was like should we hire people to go work on this and I was like nah that doesn't feel like a tenable problem this year but I I commit to if it is feasible like we're gonna be first but I just I feel like it'll be like you know another six months or so so that means it'll be next month um you give them all actually in about 20 minutes it was a company launching wait I gotta ask you one more question because like you you know, you're working closely with NVIDIA, you work with the hardware providers.

31:17People are really interested in this topic now of like, I mean, generally, like, do you believe in hardware heterogeneity, right? There are some strong opinions on this from, you know, Databricks and others here. And do you like, you know, do you still see the same supply demand dynamics around GPU shortage from your customers that you did maybe beginning of last year? I think the chips have just changed. So I think there is like, before there was a shortage, it just felt like of everything. Like if you wanted anything except a T4, you could not get it. That was hard. I think right now, there's two things we're seeing is that one, And it is now possible for us to acquire compute pretty quickly.

32:09We have big spend, though. And so everything should be kind of conditioned on that. We're making long-term commitments with providers. So that gives us negotiating power. I think customers are still struggling with availability for the most premium chips. And I think, you know, whether that's H100 or A100, I think even when there is availability, you're oftentimes looking for like three to six weeks of negotiating with cloud providers. And then your rep calling in favors in exchange for something or the other. Like the amount of times that with the cloud providers we've escalated conversations just to get people moving faster is unreal.

32:53And so I think customers are still running into it. But I do think it's getting better, and I do think that it will, you know, it seems like it will go away. I think the heterogeneity argument around different things, I mean, it's probably a good thing, right? Like, if there is more than one provider of chips. But that being said, I personally think that it's pretty overstated how easy it is to run something that looks like CUDA or CUDA in some form on an AMD chip. seems like a challenge to me. I know a lot of people, there are people who believe that they've got it done. I think the amount of time we spend debugging bad nodes we get in a place where that you have a lot of information about existing infrastructure, that's challenging as it is.

33:44I can't imagine what they'll be on these chips that are untested. And so like, I think over time, yes, I hope so. That'd be great. I think short term, you know, it's really hard for me to see how we make investments beyond NVIDIA, especially when there's a customer crunch on the other side from customers. We're like, hey, we need this now. And I don't think we want to... The other thing we don't want to do is that, you know, what Base 10 doesn't give you, it doesn't give you real access to GPUs. It's conditioned on this inference problem today. That, you know, like, if I gave you a GPU to use on Base 10, there's not that much you could do on it, except inference, just in terms of the access control.

34:24that we give to you. That being said, our customers do end up fiddling with NVIDIA drivers. They do end up installing new versions of PyTorch and have custom Docker images. And I think running those things on things that aren't NVIDIA, and especially doing them in an abstracted way, I think it'd be easier if I was building a service on top of these AMD chips and saying, take this service. But our customers do interface with the GPs in some way. And I think building an abstracted service where there's heterogeneity, that just sounds very, very challenging. I'm sure there are people much smarter than me who could figure it out.

35:05But I think for us, I think that would just add a lot of complexity and really slow down how fast we can move. But I hope, I think it is a good world when there is more than one option. How do you see customers thinking about build versus buy? And how do you think that's going to be evolving over time? Speed is the only thing that matters in this market. And what this means is that, you know, if you spend time fiddling around with your infrastructure and your service goes down when you launch it, I think that actually hurts the end user experience a lot. And it's something you just don't want to mess around with.

35:39And we see this from customers, you know, even customers like where AI is their core core thing is that they are understanding that what is proprietary to them is models, data, and workflow. What is repeated for them is infrastructure. And I think the amount of times that we've seen in the last 12 months that we're going to build this ourselves, which is very much how infrastructure engineers were thinking a decade ago, only to come back three months later, is like I have a Docker dumpster fire somewhere. It's like, it's our super qualifier. Have you built it yourself is how we know someone's going to be a great base 10 customer because they empathize with the pain and they know that this is going to allow them to move a lot faster.

36:27We had a company with a four-person AI infrastructure team that had been building this for two years, migrate all their workloads over the base 10 in 36 hours. And I think that is a pretty amazing case study for them, which is like, holy crap, we can now take these four engineers and focus on what is actually our competitive differentiated advantage. And the way we think about our business is not, we don't need to scoop everything, something off every single customer either. Like we offer options where you can run this in your own environment and pay us a license fee. And I think it is very, very cost-effective the way that us and honestly other providers are doing this.

37:12I think it's crazy, to be honest, to try to build this yourself, especially at the scale that some of these customers are operating at. I was looking at one of the customers that we chatted with this morning who was tinkering around and they said they were doing a billion tokens a day. This is a six-person chatbot company that has a billion tokens a day going through them. Like to build the infrastructure that supports that with the elasticity, reliability and the performance and then build the product experience around that, that's impossible for a six person team. And like, I think you should try to take away things that, you know, other people can do just as well, if not better.

37:54And that's my take. And I think that's kind of what I see the market coming around to as well, which is that speed is a competitive advantage. let's spend out like we can spend our way we can buy that competitive advantage without a long build cycle. It was an awesome conversation. Thanks for doing it, Tuhan. Thanks for joining me. Thanks, Sarah. Thanks a lot. Find us on Twitter at NoPriorsPod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-priors.com.

38:31Lastly, put it on the exceeding of 2020. Thank you.

From the publisher

At a time when users are being asked to wait unthinkable seconds for AI products to generate art and answers, speed is what will win the battle heating up in AI computing. At least according to today’s guest, Tuhin Srivastava, the CEO and co-founder of Baseten which gives customers scalable AI infrastructures starting with interference. In this episode of No Priors, Sarah, Elad, and Tuhin discuss why efficient code solutions are more desirable than no code, the most surprising use cases for Baseten, and why all of their jobs are very defensible from AI. 

Show Links:

Baseten

Benchmarking fast Mistral 7B inference

Sign up for new podcasts every week. Email feedback to show@no-priors.com
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @tuhinone

Show Notes: 
(0:00) Introduction
(1:19) Capabilities of efficient code enabled development
(4:11) Difference in training inference workloads
(6:12) AI product acceleration
(8:48) Leading on inference benchmarks at Baseten
(12:08) Optimizations for different types of models
(16:11) Internal vs open source models
(19:01) timeline for enterprise scale
(21:53) Rethinking investment in compute spend
(27:50) Defensibility in AI industries
(31:30) Hardware and the chip shortage
(35:47) Speed is the way to win in this industry
(38:26) Wrap

More from No Priors: Artificial Intelligence | Technology | Startups

All 169 episodes
Speed will win the AI computing battle with Tuhin Srivastava from BasetenNo Priors: Artificial Intelligence | Technology | Startups · 39 min
Listen in VO