Inferact: Building the Infrastructure That Runs Modern AI

22 Jan 2026 · 44 min · 21 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

a16z Podcast Summary: Inferact: Building the Infrastructure That Runs Modern AI

Podcast Overview

  • Title: a16z Podcast
  • Description: The a16z Podcast covers tech and culture trends, news, and future insights, featuring industry experts and leaders. Produced by Andreessen Horowitz, a Silicon Valley-based venture capital firm.

Episode Details

  • Title: Inferact: Building the Infrastructure That Runs Modern AI
  • Description: This episode discusses Inferact, an AI infrastructure company founded by core maintainers of vLLM. They aim to create a universal, open-source inference layer that enhances the efficiency of running large AI models across various hardware and architectures. The conversation covers the complexities of AI model inference and the evolution of vLLM into Inferact.

Key Participants

  • Matt Bornstein: General Partner at Andreessen Horowitz
  • Simon Mo: Co-founder of Inferact and contributor to vLLM
  • Woosuk Kwon: Co-founder of Inferact and contributor to vLLM

---

Key Concepts and Discussions

  1. The Challenge of Inference
  2. Inference Complexity: Inference, the execution of pre-trained AI models, has become one of the complex problems in AI infrastructure.
  3. Unlike traditional programming, AI models face unpredictable workload demands, requiring real-time adaptations.
  4. Large language models (LLMs) differ from earlier models because every input and output can vary significantly.
  1. Emergence of vLLM
  2. Origin and Development:
  3. vLLM started as a prototype project by Woosuk Kwon during his PhD at UC Berkeley, later evolving into a robust open-source project.
  4. Initial work was motivated by the need to optimize running large models in production, beginning with the OPT model from Meta.
  1. Open Source as a Foundation
  2. Open Source Importance: The founders believe that open-source collaboration is essential for advancing AI infrastructure.
  3. Open source allows diverse contributions, fostering innovation across different models and hardware.
  4. Successful AI infrastructure requires a cooperative ecosystem, where model and hardware developers work together.
  1. Scaling Challenges
  2. Model and Hardware Diversity: The increasing size and variety of models (e.g., models with trillions of parameters) make inference more complex.
  3. Issues like load balancing, communication overhead, and memory management arise when deploying large models across clusters of GPUs.
  1. Insights from Community Engagement
  2. Community Building: The vLLM project has grown to over 2,000 contributors, with a culture that encourages collaboration and input from various stakeholders.
  3. Regular meetups and open discussions help maintain engagement and encourage innovation.
  1. Technical Architecture of Inference Engines
  2. Components of an Inference Engine:
  3. API server for handling requests.
  4. Tokenizer for transforming inputs into a consumable format for models.
  5. Scheduler and memory manager to optimize performance under diverse workloads.
  1. Future of AI Inference
  2. Agentic AI Models: The transition from simple request-response models to more interactive, agent-like systems introduces new challenges.
  3. The need for maintaining state across long interactions complicates the design of inference engines.

---

Conclusion

  • Inferact's Vision: Inferact aims to be a pivotal player in AI infrastructure by developing a universal inference engine that supports any model on any hardware efficiently.
  • The company prioritizes open-source development and views collaboration as a strategic advantage in the rapidly evolving AI landscape.

Follow the Guests

  • [Matt Bornstein on X](https://twitter.com/BornsteinMatt)
  • [Simon Mo on X](https://twitter.com/simon_mo_)
  • [Woosuk Kwon on X](https://twitter.com/woosuk_k)
  • [vLLM Project on X](https://twitter.com/vllm_project)

Stay Connected

  • [a16z on X](https://x.com/a16z)
  • [a16z on LinkedIn](https://www.linkedin.com/company/a16z)
  • Listen to the podcast on [Spotify](https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711).

---

This summary captures the essence and key discussions from the a16z Podcast episode on Inferact, highlighting the challenges and innovations in AI inference infrastructure.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Complexity of AI Inference

0:45 to 2:31

Discussion on the challenges of maintaining AI systems in real-time and the emergence of complex inference problems.

“Then our company, in a sense, has the right meaning and to be able to support everybody around it.”

Introducing the Guests

2:31 to 3:06

Matt Bornstein introduces Simon Mo and Wusak Kwon, co-founders of Infract and contributors to VLLM.

“Matt Bornstein, General Partner at Andreessen Horowitz, is joined by Simon Mo and Rusa Kwan, co-founders of Infract and creators of the open source inference engine VLLM.”

The Origins of VLLM

3:06 to 3:24

Exploration of the inception of the VLLM project and its significance in AI.

“We're going to talk a little bit about VLLM, the Open Source Project.”

Challenges in Optimizing AI Models

3:24 to 4:28

Wusak discusses the challenges faced when optimizing the VLLM model and its evolution from a prototype.

“VLLM project started from actually Wussak's prototype project at UC Berkeley doing his PhD and grow into today's OpenSource project on GitHub for inference runtime for everybody.”

The Curiosity Behind VLLM's Development

4:28 to 4:52

Wusak shares his motivation and curiosity in tackling inference problems during VLLM's development.

“because this autoregressive language model is pretty different.”

Differences in Computing Workloads

4:52 to 7:19

Discussion on how autoregressive language models differ from traditional workloads in computing.

“pretty well-defined open source project as yeah, more and more people got interested in it.”

Dynamic Inference Challenges

7:19 to 8:47

Examination of the dynamic nature of language model inputs and how it complicates inference.

“like compute-heavy workload versus like deep learning workload, I would say.”

Innovations in Scheduling and Memory Management

8:47 to 10:12

Exploration of the scheduling and memory management innovations needed for effective inference.

“And yeah, back in the day, that was not like people didn't have a clear idea about how to handle it.”

Community and Growth of VLLM

10:12 to 11:38

Discussion on the growing community around VLLM and the diversity of its contributors.

“That's why scheduling is the first problem to solve.”

Motivations Behind Open Source Contributions

11:38 to 14:00

Simon and Wusak discuss the various motivations driving contributions to the VLLM project.

“So this was the very first VLLM meetup, right?”
Show all 21 chapters

Managing Contributor Dynamics in Open Source

14:00 to 18:03

Explore how to effectively manage a large pool of contributors in an open-source project.

“So everyone in VIA, using VIA has ability to choose about different silicons for accelerated computing.”

The Role of Funding in Open Source Development

18:04 to 19:38

Learn about the importance of funding and grants in sustaining open source projects.

“However, I did hear a rumor that at the time that we made the grant funding, that you guys put a portion of the money into NVIDIA stock.”

Understanding Inference Engines

19:39 to 23:25

Get a detailed overview of what inference engines are and how they work.

“Yeah, I mean, it makes sense for us and for other corporate sponsors of YLLM.”

The Challenges of Scaling AI Inference

23:26 to 28:00

Discuss the increasing complexities of running inference for large AI models.

“And with larger models, presumably you need more nodes working concurrently.”

The Evolving Challenges of AI Inference

28:00 to 30:27

Explore the complexities of managing state and caching in AI inference layers.

“So now this has been diverging quite a bit.”

Open Source vs Closed Source in AI

30:27 to 33:59

Understand the importance of open source in building diverse AI models and architectures.

“the patterns got pretty disrupted by the new paradigm.”

Real-world Deployments of VLOM

33:59 to 36:40

Learn about the deployment of VLOM in major companies like Amazon and LinkedIn.

“And this is kind of the first sort of magical experience in a way.”

Introducing Infraact: The Future of AI Infrastructure

36:40 to 39:04

Discover the vision behind Infraact and its goals in the AI ecosystem.

“This is a thing we heard over and over again that people just tell us, we just cannot keep up with VLM.”

The Role of Yang Stojka and Learning from Mentorship

39:04 to 41:35

Insights into the mentorship of Yang Stojka and its impact on the company's direction.

“So on that topic, what are some of the big problems you need to solve now?”

Building the Universal Inference Layer

41:35 to 42:01

Explore the concept of a universal inference layer and its significance in AI.

Building the Universal Inference Layer for AI

42:01 to 42:38

Learn about the concept of a universal inference layer and its role in AI systems.

“got innovated over time, databases got innovated over time, with the new information we have ahead.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Our goal is to make VLM the world's inference engine, really push the capabilities on the open source front, and then build a universal inference layer. That means we'll have the runtime to power any new model on new hardware for new application, be able to tailor that to extreme efficiency and support all the AI workload going forward. I fundamentally believe that open source, especially how VR itself is structured, is critical to the AI infrastructure in the world. And what we want to do with Infraq is to support, maintain, steward, and push forward the open source ecosystem. It is only that VLM wins, VLM becomes a standard, and VLM helps everybody to achieve what they need to do.

0:48Then our company, in a sense, has the right meaning and to be able to support everybody around it.

1:01What if the hardest problem in artificial intelligence isn't training smarter models, but simply keeping them running? For most of the history of computing, once a system was built, the hard part was over. You wrote the program, pressed run, and the machine behaved predictably. Even early machine learning followed that pattern. Inputs were standardized, workloads were regular, the computer did its job and stopped. Large language models quietly broke that assumption. Every request is different. Prompts can be a sentence or an entire archive. Outputs can end instantly or stretch on indefinitely.

1:34Thousands of users can arrive at once, each making incompatible demands on the same hardware. And all of this has to happen in real time on GPUs that were never designed for this kind of unpredictability. Over the last few years, this problem has moved from obscure to essential. As models have grown larger, more diverse, and more deeply embedded into products, the challenge of running AI systems has started to rival the challenge of building them. That's where the tension lies. A public story of AI progress is about better models and bigger breakthroughs. But underneath it is a quieter systems problem.

2:06How do you schedule chaotic requests efficiently? How do you manage memory when you don't know when a conversation is actually finished? And what changes when AI systems stop behaving like single-turn tools and start acting like agents that think, pause, and interact with the world over time. This episode focuses on the hidden layer. We examine why inference, the act of running trained AI models, has become one of the most complex and important problems in modern computing and why open source infrastructure is increasingly central to solving it. Matt Bornstein, General Partner at Andreessen Horowitz, is joined by Simon Mo and Rusa Kwan, co-founders of Infract and creators of the open source inference engine VLLM.

2:44This is a conversation about the infrastructure beneath AI and why it may matter more than the models themselves.

2:55We are here today with Simon Mo and Wusak Kwon, lead contributors on the VLLM Open Source Project, and co-founders of Infraact, a new AI inference company. Super excited to have you guys on the show today. Thank you. Thank you so much for coming. We're going to talk a little bit about VLLM, the Open Source Project. We're going to talk a lot about inference and what inference technology really is. And then we'll talk a little bit about Infract, the new company. So to start, can you talk a little bit about where VLLM came from? What is it? How did you start it? And why is it such an exciting project?

3:25Thank you for having us. VLLM project started from actually Wussak's prototype project at UC Berkeley doing his PhD and grow into today's OpenSource project on GitHub for inference runtime for everybody. Maybe Wussak can talk a little bit about the page attention paper. Oh, yeah. So basically, I think it kind of started in 2022 when Meta released the OPT model as open source. I'm not actually sure how many people actually remember the model nowadays, but it was kind of one of the first open-weight large language models to reproduce GPT-3. And our lab created a demo service to run the model and to demonstrate it for the broader audience.

4:10And yeah, it was working, but super slow. So I started a small side project to optimize that demo service. That was kind of at the beginning. And initially, I was thinking that it may only take a couple of weeks to optimize the service end to end. But it turns out that it actually has a lot of open problems inside of it because this autoregressive language model is pretty different. Actually, it was pretty different from other traditional ML workloads. and it wasn't actually, it was kind of like a brand new at least like outside this Frontier Labs back in the day. I started to work on it and it became a research project and then we wrote a paper and it even became like an open source project, pretty well-defined open source project as yeah, more and more people got interested in it.

4:57So 2022, this is pre-GPT4, obviously. Yeah, pre-ChatGPT. Yeah, pre-ChatGPT. Yeah. And you're thinking like, oh, I'll just like work on this inference server. this should be a fairly straightforward problem. Like four years later, actually, you're like doing more work instead of less. Exactly, exactly. Yeah. Why did you think this is a meaningful problem to work on at the time? Because like, I would say most people in the world at that time saw GPT-3 as a curiosity in some sense. And OPT was kind of like a curiosity attached to a curiosity in a way. Like what made you and your lab mates sort of excited to work on this back then?

5:32I think I also started from curiosity. I didn't really think it's the most important problem in the world back in the day. I just wanted to have a hands-on experience on how this actually works. I mean, I think I'm also impressed by the size of the model. The OPT largest model has 175 billion parameters, and that was the largest model available. So it's kind of like a meaningful for me, like it was kind of pretty rewarding to work on such a large model. This reminds me of when like, when I was like growing up, you know, we would build like computers. That was like the cool thing to do. And each step change in like memory capacity was such a big, I was like, oh my my God, this one has four megabytes of RAM.

6:08Oh my God, this one has 512 megabytes of RAM. Looking back, it's silly. But at the time, that was actually like, maybe it's because we're like nerds, but it like it gets, you get like emotionally excited about like the numbers getting bigger on these systems. Right, right, right. Yeah. I think that was like one of the main motivation, clearly. And so you started to say the sort of technical problem is different for autoregressive transformers compared to traditional machine learning. Do you mind explaining a little bit, you know, how that is and even compare just to normal kind of computing workloads for, you know, listeners who, you know, are engineers who may not be familiar with AI workload.

6:43So basically compared to the traditional workload, you know, the clear difference is definitely like GPUs, right? Now all the computes are kind of most of the computer happening on GPU and we have to optimize for the, which like presumably have less memory than CPU, at least back in the day. Now, like GPUs are much, has much larger memory, but typically it has much smaller memory than CPU, maybe still. And like, you know, like all the computations happens on GPU. So you have to write program in a different language and a different type of parallelism in mind. Yeah, so that's kind of like a fundamental difference from the traditional, like compute-heavy workload versus like deep learning workload, I would say.

7:25And, but within the deep learning workload, there's actually still a huge difference between the kind of traditional deep learning workload versus like lower language model inference. So for traditional workload, I think the biggest kind of characteristic is that it is pretty like static, which means, for example, like for image models, like back in the day, like CNNs, like what people do is, you know, we may have a different, several images with different sizes. Then what we do is we like, we resize them or crop them into the same size and then we batch them and then we put it to the model to run the inference at once.

8:02And this is basically, yeah, because of this resizing and cropping, like they all kind of, at the end, they're kind of compressing to the same size tensor. And that actually makes things much simpler for the GPU to handle, right? All the shapes are pretty regular, static, and it's kind of like well-defined. But for a large language model, if you think about it, they're pretty dynamic. You know, your prompt can be either like, hello, like a single word, or your prompt can be like a bunch of documents spanning like hundreds of pages. And this kind of like dynamism exists inherently in the language model.

8:41And this makes things a whole like kind of in a different world. Like we have to handle this dynamism as a first class citizen. And yeah, back in the day, that was not like people didn't have a clear idea about how to handle it. And fortunately, we were one of the first to see the problem. That's very interesting. So kind of regularizing a batch of inputs, it sounds like one of the first problems you had to solve. It's actually more about scheduling and memory management as well. Yeah, so the problem we're solving before in all the serving system is about just what we call micro-batching. to leverage first CPU's fundamental vectorization in the early days before LMs, and then early GPU for vision models like ResNet.

9:30It's all about micro-batching. You put together four requests together that arrive around the same time. But the change in the LM world is you always have requests that are continuously filling and coming in, and then each request looks differently. You just cannot really normalize them. So that's why you have to have a notion of a step within the LM engine to process one token across all the requests at the same time, regardless of each request having different kinds of input lengths and output lengths. Now, it's also non-deterministic. The language model itself will decide when does it stop.

10:04Instead of in the traditional sense of other machine learning servings, it's very much like work like a clockwork. And here it is very stochastic. It's always flowing. It is always continuous. That's why scheduling is the first problem to solve. And then memory management, that's where page attention comes about. is a second problem to solve. So when did you get involved in the project, Simon? Well, I got involved in around 2023. I first issued a call in the Skylab Slack channel to say, hey, we need someone to work with us on this page attention paper and kernel. Actually, surprisingly, I was on spring break and I was like, look, someone else can do this.

10:43Let me just play with GBT for the entire week. So I just ended up just playing with prompt engineering. So I actually didn't end up joining with Usook. And so this is what a vacation looks like in Jan Stoika's lab, playing with models for a week. And he's playing with kernels. Yeah, exactly. So he's playing with kernels. I was trying to build more prompt engineering and explore different kinds of early agentic workflow. And then over the summer, and especially this is when around August and September, and we really get to work together. Actually, this is where you come in. We get to work together on our very first VLM meetup, ACQZ.

11:20And where I had the experience of managing open source project before, as well as deeply interested in actually building a serving platform and into a fully open source project. And this is where I start to get involved right through my first lines of code and like sort of build up the CS system, build up the performance benchmarking systems, and then really much work with Wusuk ever since. I had forgotten about that. So this was the very first VLLM meetup, right? Yeah. I was in this office. In this office. Yeah, in this office. On the exact floor. I think we are previously anticipating just 10, like 10, 20, maybe 50 people showed up.

11:57And then the registration was like exactly over the anticipated capacity. People are extremely interested in this technology. I remember that very well because we run events here for ourselves. And it's always very hard to get people to show up. We're always scrambling. And instead, I got a call from our security team saying, too many people have been approved for this. We all have meet up. We need to scale it back. This isn't safe. I'm like, oh, okay. Probably don't tell. I don't think we ever scaled it back. So don't tell the security. It was quite crowded. The piece that ran out like the first 10 minutes.

12:28But this is a big deal, right? Because this is not like a consumer app, right? That you're building. This is pulling from systems engineers, right? For the most part, who want to learn about how to serve LLMs and contribute to it. So it's actually a big deal to get, I think, so much interest from such a kind of narrow, sophisticated group of people who don't like meeting other humans in real life that often either. you know, at least speaking for myself. So can you talk a little bit more about the community behind VLLM? Like, how big is it now? How did it come together? And like, how do you guys manage it as it's gotten big?

12:59Yeah, so in the beginning, of course, it's just a few grad students working on it. And then so, but over time, we started to having this very much open-minded and developing the open kind of mindset. So as of now, we're looking at 50 or more regular full-time contributors who open up GitHub every single day to work on VLM. We cross 2 ,000 contributor bars on GitHub, one of the fastest growing top open source projects ranked by GitHub itself. And then this is really a diverse community. So there is folks like Usook and I are sort of the team from UC Berkeley from grad student days, and as well as Meta and Red Hat pulling their way behind this open source project.

13:44And then as well as, of course, people who are not just people who are making the model, Mestral and Quen team. And of course, like anyone who's making open way model are participating in our community. And then on the model side, NVIDIA, AMD, Google, AWS, Intel, they're all having their own participation and be able to support the ecosystem. So everyone in VIA, using VIA has ability to choose about different silicons for accelerated computing. Oh, that's very interesting though, which I think is a property that many successful open source projects have, which is that people aren't all contributing for the same reason, right?

14:21Some people I'm sure just love the technology, but it sounds like you're saying the model providers actually have incentives to contribute to the project because they want their models to run well. The silicon providers want it to run well in their silicon. The infra providers want to have first divs on running it so they can sell infra, that kind of thing. Yeah, this is kind of a classic worth solving the M times M problem so that as a model provider, you don't need to talk to everybody. And as a hardware provider, you can just go into this one system and then magically you'll work for all the models out there in the world.

14:51And then for applications who are using VLM as well as infrastructure, building with VLM, like having a common ground where everybody can participate in and then innovate together is way easier and cheaper, in fact, in the end to deploy. What's your philosophy for managing a pool of contributors this large? Do you tell them what to do? Do they choose themselves? How do you maintain high code quality? It's a constant sort of iteration, month over month, year after year. So for this, I have to go back to my previous open source project, which I was working on a project called Ray, and then later AnyScale, where we have this kind of, where I learned this community-driven approach in a way that have a clear requirement, have a clear roadmap, have a clear sort of milestone being set.

15:39So we kind of tried to borrow that, but also really study this really successful open source project out there. I went all the way back to Linux and then to, and then study Kubernetes, study Postgres. How are these community operating and together? So in VLM, we had kind of a special model that we do like any normal engineering organization, set clear team scope, but also clear objective and result and milestones with different kinds of technical features we want to push forward and build. So this is where we have set forward our vision every quarter. And then, but also invite the community to contribute.

16:17So we're saying, great, we're working on these. We also need help on these items that we don't have anyone actively working on. If you are brand new and want to engage with us or engage with the community, here's what you can work on. And additionally, we keep an extremely open mind to all the GitHub pull requests that people just opened up that we're seeing, oh, is this a good request? Is this a good feature? And then as well as a request for common processes. So it kind of is a blend of all the lessons learned from previously other open source projects. And then code quality wise, code reviews, but also a lot of constant refactoring and iterations.

16:54Yeah, yeah. I do a lot of refactoring like every six months. And actually one thing to add is, we do in-person meetups every two months and we are kind of expanding to globally, actually, like sometimes in Europe, sometimes in some other places in Asia. And yeah, like we actually from the first meetup in AS &Z, we learned that it's actually super, super useful to meet, you know, those like collaborators and, you know, users in person. And yeah, we are continuing doing that. It's funny. It's another one of these lessons that like, you know, Silicon Valley engineers, like we've gotten so kind of like, you know, high up the abstraction stack that were like relearning, you know, lessons from a thousand years ago saying, oh, it turns out in person communication is high bandwidth and doesn't suffer from consistency problems.

17:41So around the time you guys did that first meetup, we also made grant funding to the project through the academic lab. I think it was a small amount of money, but it was actually the very first open source grant that we made. So it's super, you know, just like fun and kind of gratifying for us to see like the money was actually put to good use and the project grew massively. And then we even had a chance to invest in a related company later. However, I did hear a rumor that at the time that we made the grant funding, that you guys put a portion of the money into NVIDIA stock. Can you confirm or deny?

18:13Not him. So someone else in the recipient list. So you probably turned our tiny grant into 10 times as much money before. Oh, in terms of the funding for VLM. A lot of these funding for VLM is that we set aside for project development and sort of project development, testing, and everything around operating this project. And one thing we're actually super grateful for the first grant is actually kicked off a culture, and nowadays you can get even a tradition, for people really opened up to sponsor open source projects in a quite significant way. Because running via our CI bill, for example, is more than 100K a month.

18:56that could be tiny for some folks. And it's like overgrowing over time. This is where at a burn of million dollar amounts and sorry, a year, a million dollar a year. And for an academic project, it's actually very... Yeah, because we want to make sure every single commit is well tested. And then this is something that people are going to deploy at not thousands, but potentially millions of GPUs across the world in different environments. So we want to make sure it's well tested, it is reliable. And then this requirement, this infrastructure, right now all comes from contribution and sponsorship and from everybody are chipping in to help on this project.

19:35And of course, we also run meetups and sometimes expenses associated with meetups are directly leveraging the grants that you're providing. Yeah, I mean, it makes sense for us and for other corporate sponsors of YLLM. It benefits the whole ecosystem, right? So I think it makes a lot of sense. Let's talk more about the technical aspects of the problem, if that's okay with you guys. Do you mind to start just defining exactly what an inference server, an inference engine is? Sure. So an inference engine turns, it takes a already trained model. So this can be a very small model like Q1B. It could be a very big model on DeepSQL, Kimi K2, run it on a accelerated computing device.

20:20And its job is to fully utilize the computing device to be able to generate text and images and videos, essentially. But this all got tokenized into individual tokens. So the goal of Inference Engine is to produce... The goal of Inference Engine is to run the model at highly efficient speed to make sure that we can produce maximum output at the highest efficiency. And just from a high level, can you explain some of the architecture, how sort of a typical inference engine works? What are just the few most important components that people would be interested to learn more about? Maybe one goes through a life of a request.

21:00Like if I say hello, what will happen to VLM? Yeah, so basically there's a kind of traditional API server. Definitely, you know, gets the retest. And once the model generates the output, it streambacks the tokens one by one. Yeah, so there's definitely a traditional API server layer. and inside in it, we have kind of typically something called tokenizer, right? Like to transform this like input to like the tokens, basically some integers, the least of integers that the language model can consume. And inside in it, we have basically an engine, what we call like engine. And that includes a scheduler to, you know, which decides how to batch the recast to incoming recast.

21:42and we have a memory manager to manage something called KV cache, which is a kind of the core part of the transformer for idle limbs. And we definitely have some kind of a worker, this is a very generic term, but which basically actually initialize the model and run the model and get the output and do all the pre-processing for the input and post-processing for the model output. Yeah, so yeah, that's basically, I mean, in a sense, it's not like a crazy new architecture, but each one basically highly optimized and specialized for this LM inference workflow. Do you think it's getting easier or harder over time running inference?

22:24Yeah, definitely. I think it is definitely getting much more difficult over time. Actually, honestly, maybe one and a half years ago, I wasn't thinking inference as a hard problem at all, to be very honest. But now things have changed. The trend has changed so far. So I think there are kind of three factors. One is scale. Another is diversity. And the last one is kind of agents. So for scale, you know, like the models are definitely getting larger. And, you know, right now we have Kimi K2 with like more than 100, more than a trillion parameters. but I think we believe we will see like multi-trillion parameter open source model this year and I think that's still clearly a trend that people will be training a larger model and definitely it's much more challenging to deal with such a model compared to the early days of LLMs where we just only deal with small llama models.

23:28And with larger models, presumably you need more nodes working concurrently. you need, you have more memory to manage that may or may not fit in each, you know, chip's available memory. Can you describe some of the challenges from scale? Yeah. For these kind of large models, we definitely need to shard, you know, distribute the model into multiple like GPUs, multiple nodes. Right. And then, and yeah, then there's like definitely like a problem of how to shard, how to distribute this model. Right. There are actually many dimensions we can use to shard the model and they have a different like trade-offs.

24:01and trade-offs, for example, in terms of how much communication we should pay to share the model in this way. And also there's a trade-off in terms of load balancing. If I share this in this dimension, then how significant is the load imbalance? So these all need to take into account for the final performance estimation and to get the best performance. And yeah, it's becoming more and more a bigger problem as the models get larger. And what about just cluster scale? I mean, I think, Simon, how many nodes is VLLM running on at any given time? Right now, we're looking at, this is through our sort of like a very small subsample of our usage statistics that's used for us to figure out what feature to deprecate.

24:51Just literally from this one signal, we're looking at 400K to 500K GPUs, 24-7 running VLLM. And there's quite a big scale thinking about the global deployment of GPU footprints. And we definitely believe there's a lot more out there. And of course, this is a wide diversity of different kinds of GPUs, GPU architecture, as well as model architecture being deployed. We're not seeing like a one size fit all. People are using it for just one singular use case. I see. And this is sort of your point. Your second point was about diversity sort of making it a harder problem over time. Yeah, the chip diversity, harder diversity is definitely one factor.

Read the full transcript

25:27And also models are getting also diverse. You know, if you think about the, like, for example, like for NVIDIA, like a year ago, I think they only released a few series of open source models. But now they're releasing many open source models, like every month in different domains, right? Some are on the video, some are on the robotics, some are on the language. And yeah, this kind of like open sourcing trend is getting expanding. and that people are training many different kinds of models in many different domains and releasing them every month. So there's model diversity. And even just for text models, they're all transformers in that, but their detailed architecture still are very diverse.

26:10And we see they're even diverging. Like say for DeepCity 3.2 was using sparse attention, something called sparse attention. But say for Q1 and Kimi, they're kind of exploring like linear attention, which is kind of different attention mechanism. And they have different ways to manage the memory. So yeah, this model architecture divergence is also getting more significant. And so is it up to you as, you know, meaning VLM to implement all of these, like all the two, you know, implement sparse attention, for instance, so that it's available for the models to use? Yeah, definitely. we basically leverage open source community definitely.

26:51Like we, you know, because we collaborate with these model vendors, like we often get help from these model vendors. They basically provide some kernels or at least like reference implementations of, you know, of these new kind of like operations. And yeah, we like, our job is often like basically leverage this collaboration and making more mature and also available for more diverse environments. I remember early on in open source models, there was some standardization, like everyone was kind of using LAMA. I think everyone's using sort of like the same tokenizer and the same like input format and, you know, and like end of stream token and stuff like that.

27:31Is that still the case or is it like, is it different for each provider now? Yeah, it is. Yeah, it diverged quite a bit over the last few years, maybe last couple of years. Yeah. One thing is that many, yeah, like the model architecture itself has changed a lot, you know, especially on the attention side. And also even for like input output processing, because like different labs have different kind of their own ways to form, you know, how to form the conversation and how to form the tool codes, for example, for their own models. So now this has been diverging quite a bit. And now, yeah, this has been diverging quite a bit for the last couple of years.

28:07I see. Okay, so scale of models, diversity of models, and hardware deployment scenarios, and then agents were the third thing you mentioned, sort of getting hard over it. Yeah, yeah. You know, like for agents, we need a, definitely we need a kind of different, I mean, beyond, just beyond the inference engine, we also need to set up the whole new, like environment, actually whole new infrastructure to support all the tool callings and to support all the multi-agent things. Yeah, that's probably becoming kind of a new emerging challenge for inference as well. Do you think this means there will be more state managed in the inference layer over time?

28:45As before, the paradigm has been text in, text out, and then just single request response. But as we evolve into the year and the decade of agents, we're seeing multi-term conversation turning into hundreds and thousands of terms. And then these terms also involves external tool use, like interacting with sandbox, performing web searches, running Python script or any programming languages, and be able to have this kind of long iterative process where LOM is involved, but also external environment interaction is involved. And this really kicked off a huge wave of co-optimizing agentic architecture with inference architecture.

29:30So just to give an example, that when, because just to give an example, it is very important for VLM to understand whether or not the conversation is still happening. If the conversation is no longer happening, we can remove the KV cache. That is the persistent state associated with each text completion streams. But in agentic use cases, you actually don't know whether or not the agent will think it finishes or also wait the interaction previously. And the interaction previously was just a human typing in the text box. But now it becomes external environment interaction. It could be one second just for a single script to finish.

30:13It could be 10 seconds for a search or like a complex analysis to finish. And then it could also be minutes, hours, if there's humans in the loop. Now, with that uncertainty, we actually don't even know when is the request going to come back. And then the uniformity of cache access pattern and eviction pattern got kind of, the patterns got pretty disrupted by the new paradigm. I see. I see. And so you have to be much smarter about how you manage the cache as one. As one example of that. Yeah. Gotcha, gotcha. Which is one of the unsolvable problems in computer science. Caching validation. Yeah, exactly.

30:49So I can see how that would get harder over time. I think I know the answer to this, but are you guys big believers in open source AI compared to closed source? And can you just explain how you think about that? We're definitely big believers in open source. What we believe is diversity will triumph, that sort of single of anything at all. So that means we believe in diversities in models, diversity in chip architecture. Fundamentally, this is because the world is complex. For any application, you're going to need to find and tailor the right sort of model architecture to the right chip architecture for that right exact use cases.

31:28And the best way to promote diversity and improve that is through open source. Because open source, everybody knows where everybody else is up to and be able to make their opinion take based off the common ground. And finally, if you look at the history of computer science, operating system, cluster managers, databases, every single system field gets better when they're starting to have a common standard and everybody that deviate a little bit, innovate on top of each other versus following a single line of trend that is proprietary and single source control. I see. That's very interesting. So you're almost saying OpenAI will tune their stack very tightly for their use case, which is ChatGPT or whatever other apps they're running.

32:13For an enterprise or another tech company, if I want that same level of tuning, I can't just use off-the-shelf closed-source models because I don't sort of control the whole stack and the different participants in the stack kind of aren't paying attention. Yeah, of course, one part is data. One part is the model architecture itself, which will impact the performance. And then just on the model architecture itself, right? How smart do you want the model to be? Do you want the model to be able to handle millions of contacts, token contacts, or just shorter contacts is totally fine, right? And then you also need to specialize that model to your exact compute architecture.

32:49What chip are you using? For example, for NVIDIA, the model you design for a H100 chip is very different from a B200 chip. And then it is very different for a GB200 MVL72 system. And then compared to, for example, the model architecture you design for TPU, then again, that is also drastically different. And then using it for vision model, video generation, and for reasoning, mass coding. In the end, we'll all look at the vertical stack integration. and we're like, wow, they're so much different from each other. That makes sense. Can you just share any stories about live VLOM deployments that you thought were particularly interesting or important?

33:31I have a few. One is, I think around 2024, we learned that Amazon is running VLOM to power their Rufus Assistant bot, which was really surprising to all of us because one, at the point, of course, Like we believe VLM can be deployed at scale, but seeing this as a massive scale, like kind of global e-commerce deploying this as like front page feature. That means when everybody, when they're opening Amazon app and clicking the bot's suggestion or even entering a search query is going through VLM. And this is kind of the first sort of magical experience in a way. One of the first experiences was, wow, my purchase is going through VLM right now.

34:15It's kind of exciting, but also scary. You're like PhD students at the time. And also across not just Amazon, LinkedIn, and every major deployment of VLM, we're surprised to find out they're always the first adopter of cutting-edge features. So I've seen one of the examples of deployment of VLM within Character AI was when we first make the N-Grain speculation for SpectreCode available as just a single PR, pull request in VLM, not even merged. and then while we're still iterating on that feature and I heard someone from Character AI saying, oh, actually, we already wrote it out. You have hundreds of GPUs at scale given just your first iteration of this feature.

34:57So it's really much everybody is staying on the cutting edge of VLM and we're quite excited about that, yeah. Okay, should we talk about the company then, Infraact? What is Infraact and why did you guys decide to start the company? So Infraq, created by the creators and maintainers of the VLM project, our goal is to make VLM the world's inference engine, really push the capabilities on the open source front, and then build a universal inference layer. That means we'll have the runtime to power any new model on new hardware for new application, be able to tailor that to extreme efficiency and support all the AI workload going forward.

35:41And implicit in what you just said is that you're devoting a lot of resources, I think, to the open source project. Could you, I guess, is that right? And can you expand on that? Yeah, one thing I believe is, I fundamentally believe that open source, especially how VRM itself is structured, is critical to the AI infrastructure. And what we want to do with Infrared is to support, maintain, steward, and push forward the open source ecosystem. It is only that VLM, when VLM becomes a standard and VLM helps everybody to achieve what they need to do, then our company, in a sense, have the right meaning and to be able to support everybody around it.

36:23So open source is definitely number one and, in fact, sometimes the only priority of our company right now. You're not supposed to tell your investors, by the way. bit of that. We do believe that open source project is also kind of a secret weapon in a sense that having this community all work together for this open source, we have the execution beyond any single entity can have. This is a thing we heard over and over again that people just tell us, we just cannot keep up with VLM. So that's why we're using VLM. We have our internal team, we'll have our internal fork, we'll have our internal inference engine.

37:00But open source moves so fast that the only way to stay ahead is adopting. And that's why we want to make happen. And in fact, this is exactly why we're staying all in on open source. That's awesome. We mentioned Jan Stojka before, obviously one of the founders of Databricks. He was your, I think, both of your PhD advisors at Berkeley. And he's going to be involved in Infract too. Can you talk about maybe a little bit how he's going to be involved in this company? And even more importantly, what have you guys learned from him as his students and about startups and distributed systems and all this stuff?

37:34Sure. Yeah. Yeah, you're exactly right. Yang is both of our advisors. I have actually worked with Yang since 2017, since I was an undergrad working on my first open source project for serving. And then work with him at any scale for my second open source project for serving. You're just addicted to like Berkeley-based open source AI serving companies. Yeah. So at this company, I am young is quite involved as so as a company, he will be a co-founder. And then as an open source project, he has been advising this project for since its inception. Young knows open source project, academic project, industry research trend, you know.

38:13So from what we're working together on, Yang really helps us with both clearly understanding all the lessons learned about bringing open source through the final miles of adoption in companies' enterprises, as well as what is actually happening on the research world. A sky computing lab over the last few years has produced amazing infrastructure and new research ideas. And Yang continued to explore a new frontier on that front. And then we're quite excited to hear that and also innovate on the open source together. Yeah, and he also helps recruiting a lot. And he's involved in all of our hiring process.

38:55He basically teaches us how to tell talents, where to find talents. These are all amazingly helpful. Cool. So on that topic, what are some of the big problems you need to solve now? And what type of people are you hiring to help you solve? Definitely, you know, the inference at scale is kind of one of the biggest challenge, I think, in the field, not only for us, but in the field overall. So we are trying to hire more like a very experienced ML infra engineers overall to make, But for example, what would be the best way to utilize the GB200, GB300, NBL72 rack entirely for the giant open source model?

39:39Still, I think it's an open problem. There are definitely some endeavors in academia and industry, but I think there are some room for improvements. So yeah, that's some of our focus at the moment. Here's my pitch from a computer science point. Pretty rare if people ask me this question. And that is, if you're working at a vertically integrated company that have an end product for, let's say, for chatbots, for assistant, you are working on the vertical slice of the problem. In Infraact, you will be working on an abstraction of horizontal layer. And this is similar to operating system, databases, and different kinds of abstraction that people have built over the years.

40:26operating system, abstracted CPU and memory, databases and file system, abstracted storage devices and networking. For accelerated computing, there's a brand new physical device that's inference and VL abstracted a large part of it for inference-specific work. Of course, it's training, but we are a singular focus is on inference. And this necessitates a layer, a software layer that abstracts away GPUs and assertive computing device for models. And this is as important from my point of view as abstraction you need to build for OS, for databases, which I both feel we're really passionate about when we're PhD students too.

41:10So that's why ML System is fundamentally a new system research and system deployment. So you here at Infrared will be working on this layer that is not a vertical slice, but a fundamental runtime and impacting all the future generation of software that will run on a cellular computing device. And your work will stand from both working with different models and then working with different applications, as well as understanding the pros and cons of different chips, as well as their whole integrated data center systems, to be able to figure out, oh, actually, for these, we should build the abstraction this way.

41:52And we'll constantly remove abstraction, break abstraction, and build it over and over again, just like how operating system got innovated over time, databases got innovated over time, with the new information we have ahead. So you will come here to have the constant exercise of building an actual widely deployed production system that will be at the frontier of interest. And this is what you call universal inference layer? Yeah. It's purposely vague in a way, but what we really focus on is going from page attention, from going from the serving system to the whole runtime you need for intelligence.

42:38Usuk, Simon, thank you so much for being here today. Thrilled to have you on the podcast, of course. And we're thrilled to be working together in the company. It feels like it's been a few years. We've already been working together, But yeah, great to have you here. And congratulations on getting off to a great start. Thank you for having us. Thank you.

43:18or security and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments, please see A16Z.com forward slash disclosures.

From the publisher

Inferact is a new AI infrastructure company founded by the creators and core maintainers of vLLM. Its mission is to build a universal, open-source inference layer that makes large AI models faster, cheaper, and more reliable to run across any hardware, model architecture, or deployment environment. Together, they broke down how modern AI models are actually run in production, why “inference” has quietly become one of the hardest problems in AI infrastructure, and how the open-source project vLLM emerged to solve it. The conversation also looked at why the vLLM team started Inferact and their vision for a universal inference layer that can run any model, on any chip, efficiently.

Follow Matt Bornstein on X: https://twitter.com/BornsteinMatt

Follow Simon Mo on X: https://twitter.com/simon_mo_

Follow Woosuk Kwon on X: https://twitter.com/woosuk_k

Follow vLLM on X: https://twitter.com/vllm_project

Stay Updated:

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Show on Spotify

Listen to the a16z Show on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

 

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from The a16z Show

All 489 episodes
Inferact: Building the Infrastructure That Runs Modern AIThe a16z Show · 44 min
Listen in VO