Dataflow Computing for AI Inference with Kunle Olukotun - #751

14 Oct 2025 · 58 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Notes on Podcast Episode: Dataflow Computing for AI Inference with Kunle Olukotun - #751

Podcast Overview Podcast Title: The TWIML AI Podcast Host: Sam Charrington Guest: Kunle Olukotun, Professor of Electrical Engineering and Computer Science at Stanford University, Co-founder and Chief Technologist at SambaNova Systems. Episode Summary: The episode explores reconfigurable dataflow architectures for AI inference, their advantages over traditional architectures, and how they cater to large language models (LLMs) and agentic workflows.

---

Key Concepts Discussed

  1. Reconfigurable Dataflow Architecture
  2. Definition: A computational paradigm that allows hardware to be dynamically configured to match the dataflow graph of an AI model instead of relying on a traditional instruction-fetching mechanism of CPUs/GPUs.
  3. Origin: Emerges from the need to effectively implement machine learning algorithms in hardware.
  1. Advantages Over Traditional Architectures
  2. Dynamic Configuration: Instead of fetching instructions, reconfigurable architectures configure hardware according to the dataflow graph, enhancing processing efficiency.
  3. Elimination of Bottlenecks: Reduces memory bandwidth bottlenecks encountered in traditional architectures, crucial for LLM inference.
  1. Memory Handling
  2. SambaNova's architecture integrates large, tiered memory that enhances model-switching capabilities and allows efficient multi-model serving.
  3. Memory Architecture:
  4. On-chip memory: 0.5 GB
  5. HBM (High Bandwidth Memory): 64 GB
  6. High-capacity DDR memory: 1.5 TB
  1. Efficient Inference
  2. Inference is constrained by memory bandwidth; thus, the system optimizes data flow to minimize unnecessary data movements.
  3. Example Usage: For an 8 billion parameter model (like LAMA 3.1), the architecture maps decoders onto multiple reconfigurable data units (RDUs), allowing efficient parallel processing.
  1. Multi-Model Serving
  2. The architecture can host multiple models (up to 5 trillion parameters) and switch between them rapidly (in about 1 ms), making it suitable for applications needing real-time responses.
  1. Agentic Workflows
  2. Agentic systems often require multiple models to operate simultaneously. The architecture supports such scenarios by allowing quick model switches and efficient resource utilization.
  3. Collaborations with frameworks like Crew AI and Autogen AI enable the orchestration of these workflows.

---

Research Directions and Future Prospects

Dynamic Reconfigurable Architectures

  • Kunle is exploring architectures that allow for dynamic reconfigurability with low overhead, adjusting to the changing needs of AI models and workloads efficiently.

AI Agent-Driven Compiler Development

  • Research is ongoing to leverage AI agents for building compilers for new architectures. This could drastically reduce the complexity of optimizing software for specialized hardware.

Improvements and Opportunities

  • Predictions suggest significant performance improvements (5 to 10x) in both processing speed and energy efficiency as architectures become more specialized for AI tasks.

---

Conclusion

  • Future of AI Inference: The discussion indicates a shift towards more adaptable and efficient computing environments, significantly impacting how AI and machine learning models are developed and deployed, particularly in real-time scenarios and complex workflows.
  • Call to Action: Kunle encourages exploration of new architectures and emphasizes the importance of continuous innovation in both hardware and software to meet the demands of evolving AI technologies.

Related Resources

  • Complete show notes and additional resources can be found at [TWIML AI Podcast Show Notes](https://twimlai.com/go/751).

---

Episode Transcript

  • The complete transcript can be accessed through the podcast's official page.

---

This structured summary aims to provide a clear understanding of the episode's content while highlighting the advancements and innovations discussed by Kunle Olukotun in the realm of AI inference computing.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I'd like to thank our friends at Capital One for sponsoring today's episode. Capital One's tech team isn't just talking about multi-agenticentic AI. they already deployed one. It's called Chat Concierge and it's simplifying car shopping. Using self-reflection and layered reasoning with live API checks, it doesn't just help buyers find a car they love. It helps schedule a test drive, get pre-approved for financing, and estimate trade in value. Advanced, intuitive, and deployed. That's how they stack. That's technology at Capital One.

0:34Hey folks, Stephen Johnson here, co-founder of Notebook LM. As an author, I've always been obsessed with how software could help organize ideas and make connections. So we built Notebook LM as an AI-first tool for anyone trying to make sense of complex information. Upload your documents and Notebook LM instantly becomes your personal expert, uncovering insights and helping you brainstorm. Try it at notebooklm.google.com.

1:41all right everyone welcome to another episode of the twimloidai podcast i am your host sam sharrington today i'm joined by kunle olukotin kunle is a professor of electrical engineering and computer science at stanford university and co-founder and chief technologist at SambaNova Systems. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Kunle, welcome back to the podcast. It has been a minute. Good to be back and glad to be here. We were just chatting. It was, I think, at a NURPS when we first spoke back in 2018. clearly a very different world in AI since then and it is going to be really good to catch up with you and what's been new in your research and at SambaNova.

2:34Let's get started by having you share a little bit about your background and what you've been working on recently. Well thank you for having me on again and it's uh good to be uh back on on the uh the cast podcast and uh you know my background is is is i'm a computer architect uh well known uh for doing some of the uh pioneering multi-core work uh back in the in the uh mid-90s and uh done a lot of work on parallel programming environments using domain-specific languages. And this work has led to the work we're doing now in research and at Sandanova about how to take reconfigurable data flow architectures and efficiently execute these domain-specific language models to create environments that are better for doing AI computation, right?

3:35And so our focus is on how you do very efficient, fast inference on big models as energy efficiently as possible, right? And so the focus has been on very large models, you know, models with trillions of parameters that can fit in a single rack, on models that are good for agentic solutions, where you've got hundreds of agents that potentially need to work together and switch between these agents with very low latency, and then sort of very fast inference because you have interactive environments that you want to support multiple LLM model calls sequentially for a chain of thought kinds of reasoning or because you've got a workflow that requires lots of LLM calls.

4:35You mentioned in there the idea of reconfigurable data flow architecture. Let's kind of step back for a second and talk a little bit about what that means. What is a reconfigurable data flow architecture and where does that idea come from? Well, it comes basically from thinking about how you take the computational paradigm that comes from machine learning and implement it in hardware, right? So you think about designing a AI or ML algorithm, right? And you do it in your favorite domain-specific language framework. And these days, everybody's favorite is PyTorch, right? You know, back in the day, we did some early work on a language that you've probably never heard of called Optimal, but it predated PyTorch, right?

5:36But it was the same idea, right? And then, of course, you know, Google had TensorFlow and for a while there it was in contention, but I think, you know, it's fallen by the wayside. So today, it's all PyTorch. And what'd you get out of a PyTorch uh you know once you write your your algorithm in pytorch what you get is what computer scientists would call a data flow graph right and so the data flow graph is a graph which has nodes and edges right because all of us have nodes nodes and edges and the nodes in this case are kernels of computation, like matrix multiply or convolution or softmax or maybe your favorite kernel that you came up with to support your specialized model.

6:31And then the edges are basically tensor data, right? Matrices that flow between these nodes, right? Right. So there's a famous computer architect named Jim Smith, and he designed Cray computers back in the day. And his favorite quote that I often use is, you know, if you have a vector problem, and most of the problems were vector problems, then you should build a vector computer. Right. So. All right. So we have a data flow problem, right? So let's build a data flow computer, right? So let's build a computer that matches the data flows as represented by the algorithm in the hardware. In other words, these graphs.

7:27These graphs, right? So why does it need to be reconfigurable? Well, because you don't have one graph, right? Right. The graph is your program. It's your model. It's how you change what the hardware does. You reconfigure it, right? Instead of thinking about executing instructions, you think about configuring your hardware to match the graph that you've developed in this domain-specific language. That is the fundamental idea of reconfigurable data flow. So compare and contrast this to kind of traditional computer architectures, right? You know, there are data flows that I guess by, you know, you can squint and also call them reconfigurable data flows.

8:18You've got registers and you've got memory buses and data flows in different directions. And you've got. Yeah. So I think the key difference is whether you have instructions, right? You know. So meaning instructions versus kernels or? Instructions versus configuration, right? Versus do you fetch an instruction every cycle and decide what to do? Or do you configure your hardware and leave it for that way for the whole program? Okay. That's the difference, right? That's one difference. The other difference is a kind of technical issue, but also is very important, which is how do you synchronize the computation, right?

9:09You have parallel computation, and the way that parallel computation typically works is the way that you communicate is through a single address space called shared memory, right? And whenever you do this, you need some sort of synchronization to make sure that, you know, the data that one thread put in the memory is ready for the next thread to go pick up, right? So locks, barriers, all this stuff. Reconfigurable data flow has none of that. You don't have any of that, right? So this is a huge benefit because it means... shared memory or because there's an alternative to locks for ensuring synchronization and accessing shared memory?

9:57Both. I thought those were mutually exclusive propositions.

10:07There's no shared memory. Okay. But you can access, there's no explicit shared memory. That's true. And the way that you do the synchronization is by using data flow tags, right? So you've got hardware mechanisms that essentially tokens, which determine when data is ready, right? So I send you a stream of data and I've got some tokens there, which you can look at and decide when the data is ready. Okay, so that is the way that you do communication, and it matches very nicely with this model of the graph where one kernel does a computation, and it generates data, and it sends that to another kernel, which is executing on a different piece of the chip.

11:10One of the pictures that I'm getting of this is kind of analogous to like stream processing versus traditional database, like centralized processing. Is that? Streaming is definitely a mechanism that we use. Like at the hardware, like chip level. At the chip level, you're doing streaming and you've got mechanisms that make sure that the streaming has a zero overhead. Okay. Yeah. That's the way to think about it. Interesting.

11:46And with regards to reconfigurability, I'm trying to get a picture for, on the one side, and this is not the right spectrum, I don't think, but on the one side, you're fetching instructions every cycle on another side like you've got like uh um like uh like a reconfigurable chip like a xilinx or like a uh i'm the blanking the words blanking for me but like fpga yeah exactly yeah so it has some similarities to fpga but it's it's it's much coarser grained right you're not thinking doing things at the at the logic gate level you're doing things at the coarser tensor It's a core level, right? And then you've got this very explicit mechanism for doing this synchronization based on tokens.

12:43But if it's FPGA-oriented, like, I would... So don't think FPGA because it's not, you know, you don't program it with a hardware description. That's what I'm thinking of, like, the timescale of utility of configuration is, like, what I'm trying to get at. Yeah, yeah, yeah. So it's much, much quicker, right? So you're on the order of microseconds to reconfigure. Oh, okay. Got it. So you can push a PyTorch graph and reconfigure it. Maybe you wouldn't want to do that per transaction or per inference, but you could? You might. You might. You might. Because, you know, so think of it, right? Eventually, you're going to get a graph that is bigger than the system that you're trying to map it on.

13:29So now you have to reuse that hardware for different parts of the graph. Okay. Interesting. So you have to be able to reconfigure. Taking a step back here, we're talking about this idea of reconfigurable data flow architectures. And at the point we were talking, this kind of was, I think, a little bit closer to the cusp of research to practice. Now this is something that you've been doing, you know, particularly since like the, I think the, you know, the chat GPT moment, you know, shortly after there we saw a shift. I remember seeing a shift from Salmonova from like this theoretical computer that is scalable that you go figure out what you want to do with it to like, we're going to get really good at this LLM thing and, you know, make it, you know, available to folks.

14:23Talk a little bit about how you got from this idea of reconfigurable data flow architectures to we've got these LLMs, this is the right thing for it, and what we need to do to either accommodate the LLMs on the architecture or accommodate the architecture for the LLMs. Yeah. So it turns out the architecture was very well suited for LLMs. is fundamentally, especially for inference, right? And so inference was something that we saw was really going to be a bottleneck. And if you think about inference, if you think about doing inference, it's fundamentally constrained by the memory bandwidth, right?

15:17So I should back up and say, okay, You know, when I talked to you about 2018, in 2018, this idea was pretty theoretical. Now we have five generations of chip. The chip that we currently have in the market is called the SN40L, which is, you know, 100 billion transistors, five nanometer chip. And it has three tiers of memory, right? So it's got half a gigabyte of memory on chip. It's got HBM, 64 gigabytes of HBM connected to the accelerator. And then it also has a high capacity memory, 1.5 terabytes of DDR, right? So this is unique in that most accelerators do not have that high capacity memory.

16:10And so when you think about doing inference on a large language model, It's fundamentally constrained by the bandwidth to fetch the parameters and to fetch the KV cache from the HBM. That's the limit. So what you want to do then is make sure that you don't send any extraneous data that should should not be necessary over that highly constrained, that critical resource, that HBM bandwidth, right? And so reconfigurable data flow does this by, you know, you can go back to your streaming idea, right? So if you take a LLM model, right? So let's take a simple one like LAMA 3.1 8B, right? 8 billion parameter model, it's got these decoders, and it's got 32 decoders, right, which are essentially the same operation.

17:23What you can do is that you can take that whole decoder and map it onto a set of 16 RDU chips in space. So that whole... An RDU chip is... Is a reconfigurable data unit. The SN40, yeah. Yeah, yeah. So we call reconfigurable data flow architecture and the chip we call a reconfigurable data flow unit, RDU. So take 16 of these RDUs and you take the one decoder and you map it in space across all those chips. Okay, so now it's resonant on your system. And then to generate the tokens, then you just loop through that 16 times. And no data ever goes between the different components of the decoder ever crosses the HBM accelerator boundary.

18:31And the only data that comes... so you don't need like InfiniBand or, you know, some super low latency interconnects? Well, you do have interconnect between the chips that is lower latency, you know, some sort of custom chip implementation, custom interchip protocol. The 16 RDUs, like, should we think of these as cores that are on one die, or are they like chips that are on one board, or are they systems that are in one rack? So typically you've got a couple of chips per board, and so you can think of these as across eight boards. So you think of an RDU as an Argus to a GPU in some sense. So now, of course, this idea in the compiler land is known as fusion, right?

19:34So you essentially have created a single kernel, a single fused kernel that encompasses the whole decoder, right? So, you know, you probably have heard of this idea of flash attention, right? In fact, I was just going to ask you, you talked about the memory bandwidth limitation being a significant issue and the way we usually do inference, like with GPUs. And we've got all these tricks like flash attention and speculative decoding and other things that are trying to help us overcome that limitation. and the question that I was going to ask was like, do you not have to do these kinds of things or do you do them differently?

20:22You can do them on steroids, right? Oh, okay. The same tricks work, except you get to do them much better. Okay, okay. Talk more about that. So if you think about this idea of flash attention, right? What you're doing is you're taking some component of the decoder that does the attention mechanism, and you're fusing things together such that you can reduce the memory bandwidth requirements. We can go even further and fuse the whole decoder, right? So you get flash attention across benefits across the whole decoder, right? Which is a huge reduction in memory bandwidth requirements, okay? Okay? So then that's one thing.

21:11The other thing that we do, it comes back to this idea of synchronization, right? And making sure that the synchronization is very low overhead. As you know, in any computer system, you have to make sure that the critical resource is used 100 % of the time. Because that's going to limit your performance. And the performance limiter for inference is, as we've just discussed, HBM bandwidth for the parameters and the KV cache. And so what you want to do is you want to make sure that you never have that interface idle and you overlap the computation happening in the CPU, with the memory access on the HBM and the communication across chip to the other RDUs.

22:12You keep all of this, you know, it's like everything everywhere all at once, right? You know, you want everything to be busy all at the same time. And you want to make sure that you use that interface to the HBM all the time. And what we do is we can achieve the fraction of busy time that we can achieve on the RDU is two to three times higher than GPUs can achieve. because we have this data flow idea of asynchronous execution of all the components, right? So there's no sequential instruction access, right? So one of the other, you alluding to your earlier question of why is this different? Well, the problem with sequential instruction execution is, you know, has a benefit is that humans can understand what's going on.

23:23The big downside is it keeps things sequential, right? Sequential, right? So what you want is you want asynchronous, right? So GPUs are starting to get into this idea, but we've taken it to the extreme in data flow. As I said, it's everything all at once, all the time, right? So it's all happening at the same time, much like you would a hardware system. And so there is nothing keeping us from keeping things busy all the time. And so is that HBM utilization, is that like just a simple following from the parallelization? Or are you also layering tricks like imagining like speculative prefetching or, you know, various things you might want to...

24:23It comes naturally. It comes naturally from once you kind of design things in this data flow manner, you don't have to prefetch, right? Because you naturally have prefetching going on because that unit accessing the HBM is working independently at the rate at which the HBM can go, right? You don't have to prefetch. I'm imagining then that like the, you're thinking a lot about, you're either thinking about InDesign or you're limited by like some ratio of RDUs to, to like interconnect or in order to keep that HBM utilized. Like you can't go beyond a certain scale of compute because that determines like that utilization level.

25:13Yeah, so part of the trick is sort of how you're going to map the model to your architecture, but you're going to map it in such a way that you make sure that the HBM bandwidth is the limiter, right, for the data you need to move. And so the idea is you don't move any data that you shouldn't have to across to using the HBM bandwidth. So you only use the HBM bandwidth for the absolutely necessary data, right? So a GPU puts intermediate data between the kernels that flows back and forth the GPU. That's wasted bandwidth, right? So you don't do any of that because you've got this data flow mechanisms that move the data between the kernels on chip and between chips using chip to chip interconnect.

26:09And then you also make sure that that critical resource is used at the highest utilization that you can muster, which is upwards of 90%, which GPUs cannot touch for multiple reasons. So talk a little bit about that process of mapping a model to the architecture. You talked about one of the key constraints there is utilizing that HPM bandwidth, but what goes into it when Meta comes out with a new model or Quinn or whoever comes out with a new model? How do you then ensure that that's optimized for this approach? Yeah, yeah, that's a good question. So, you know, we start with PyTorch models, right?

27:00And so PyTorch gives us a graph of operators, right? And then we have a implementation of these operators in our environment. And so we've got a way of, you know, so if you think about these operators, you can describe them as they're either gems or there may be some custom operation, right? And this custom operation could be described using a map reduce paradigm. So we've got a Python environment that allows you to describe specialized operators using MapReduce. And so now you've got a set of kernels. And now the question is, how should you fuse these kernels together to create a representation that can be mapped?

28:08And the way that you do things is that you need to paralyze within the chip by splitting up the tensors. You might want to reduce the size of things by tiling, just by taking the tile of computation, thinking back to what you had to do manually in something like flash attention. And then you also want to think about how you shard the computation across multiple chips. And so our compiler has mechanisms for you to specify what you want based on the tensors. This is all about the tensors. You take a tensor, you say, am I going to tile the tensor? Am I going to paralyze the tensor? Am I going to shard the tensor?

28:55And then that's the way you get the mapping. And then now you have something that can map to your machine, which is a collection of RDUs. And the whole goal is there will be some performance tweaking. But at the end of the day, what you're trying to do is to make sure that you are using the critical resource. If it's inference, it's HBM bandwidth to the maximum extent. extent. Yeah, I remember thinking that in the early days of Salmonova, like having to have, you know, enterprise engineers that have built their own custom models, like manage that, you know, mapping process, you know, sounded like a difficult problem just from like a go-to-market perspective.

29:47But it's a different world when you've got, even at the pace that we're seeing new models now, you know, that's work that your engineers can do. They know how to do it. And, um, right. And we've gotten very adept at doing this, this sort of thing. Right. And, and fundamentally they're all transformer based models, even variants. In fact, you know, once, once deep seek came out, which has a slightly different, uh, multi-headed latent attention, which is slightly different. It took us about a week to implement that and get it, get that running on our system. So, you know, as long as they're basically transformer-based models, then we can do this fairly quickly.

30:32But, you know, even if you have to create a custom kernel, you can do it at the Python level. You don't ever have to write kernel, I have the right kudo, I should say. And so we can make this transition fairly easily. You know, in thinking about kind of optimizing the system for inference, like are you primarily optimizing around tokens per second or are there, how do you think about the different metrics that are important for inference and how you tune your system to deliver those? We are optimized for both tokens per second because we have this kind of very low latency capability when you're doing a matrix vector style computation, right?

31:31So we do that very quickly and we're, you know, sort of 10x or faster than GPUs at that point. But even when you increase the batch size in order to get higher throughput, right, as you know, there's this latency throughput tradeoff that you see in inference where as you increase the batch size, you get higher throughput. And so from the point of view of the server, you know, this looks better because it costs you less power. You know, you can get a higher throughput in terms of the requests that come in. But the latency, of course, goes up, right? And if you look at GPUs, the latency really goes up once you get to sort of, you know, 1 ,024 or above.

32:26And what you see on a SN40L-based system is, yes, our latency increases, but it doesn't increase nearly as much. We're still kind of 5x better in terms of latency, even at these very, very high throughputs. So that's what cause you're not batching? or? We are batching, but fundamentally we still retain the advantages that we had, uh, you know, with the, the lower batch, uh, uh, implementations, right? Yeah. Why is that? Like I, I, when I think about the, you know, that throughput latency trade-off, I think that, you know, the reason why the latency grows up is because you're waiting to start your computation because of the batch, right?

33:14And so how do you not have to deal with that? Well, you still have to deal with the latency, the queuing latency, but the fundamental computation that we can do is more efficient because we can do tensor parallelism in ways that GPUs cannot, right? GPUs aren't efficient with tensor parallelism because they can't overlap communication very efficiently. And we can use tensor parallelism to decrease the time that the overall computation takes. And because we can overlap and hide the latency of communication, this is still effective for us, whereas it might not be, and typically it's not effective if you try and do that on GPUs.

34:05so yeah you you don't you don't get a you don't you the queuing latency you're going to pay but the question so there's two components of the latency there's one is the queuing latency uh and then in the gpu you know is it fair to think about as like a serialization latency um because the batch is processed serially and then in your case the batch is processed the batch is is processed in parallel but the thing is that that fundamentally yeah yeah fundamentally you can't take advantage of the kinds of things that we can do on the RDU, right? So yeah, you can think of it as some serialization that happens on a GPU that we can paralyze away on the RDU.

34:51That's one way to think about it. I assume you mentioned one other thing about inference while we're on that subject, which is one of the things that we have because we've got this high-capacity memory is we can hold multiple models at the same time, up to 5 trillion parameters in total. And then because you've got, you know, it's not nearly HBM bandwidth, but you've got much higher bandwidth to the accelerator than you would have if you were trying to communicate over PCIe to the host, Now you can switch models in about a millisecond, right? And so you can think about if you've got multiple custom models potentially, right?

Read the full transcript

35:43Because you fine-tuned the models. you can then so you fine-tune some model for a particular user then you can then switch between these models with very low latency. That presumably helps like at a system level if you've got multiple models you can deploy them in parallel. Yeah, so if you've got a multi-tenancy environment or you've got an environment where you're serving all these models and you could So one of the benefits of serving a custom model is you can charge more for it, right? And the question is now how do you make that efficient, right? So you need to be able to support. If you dedicate a GPU to a custom model, then the utilization of that GPU might be below.

36:38But if you can switch between these models with very low latency, then you can keep the overall utilization of your system high, even though you're supporting a bunch of custom models. Got it. Along the lines of customization, you know, there's, you know, inference is clearly important. And, you know, in the big picture, you know, we'll see a lot more of it than training. But tuning and post-training is also increasingly important for, you know, supporting enterprise use cases, does or how does the architecture support post-training? Yeah, it supports post-training just fine, right? And in fact, you can post-train very effectively, again, using this big memory, because it turns out that the benefit of training or even pre-fill is it's basically compute bound, right?

37:43So we switched from inference, which is memory bound, to training, which is compute bound. And so now you're not bandwidth constrained anymore. So now you can actually use the big memory to hold lots of parameters, right? And so you can do training very effectively, right? Because you don't need to hold all of the data in the HBM, you can put part of it in the DDR. I'm curious if you're seeing any particular use cases emerge based on the capabilities of the system, or are they in line with what everyone else is doing, or what we're seeing, a better way to put it, are they in line with what we're seeing across the broader market?

38:31Well, I mean, what we're seeing is we're seeing a lot of use cases that are very latency sensitive, right? So, you know, real-time voice is one that seems to be something that, you know, we've been working with 11 Labs and a couple of other companies. And they really care about low latency. And so that's one use case that we see that's certainly something that we're leaning into. Then the other use case is sort of these agentic systems that require large numbers of models to coexist at the same time. right and then using this this model switching capability we can we can support that with far fewer uh you know resources than you would be required if you had to uh put each of the these models on a separate gpu based system so these are the two use cases that that seem to be uh coming to the fore.

39:45Talk a little bit more about the agentic side of things. How does, I guess you mentioned that, you know, part of it is that, you know, the agentic system is composed of different models and the architecture supports that. Are there other things that you're doing to enable these agentic use cases? Yeah, I mean, the agentic use cases are, you know, you need to work with the companies that are kind of thinking about how to create the environments for using agentic systems, right? So the ecosystem is bigger. It's not just supporting the models. It's now supporting the frameworks that orchestrate the models.

40:39Yeah, so frameworks like Crew AI and Autogen AI. And so we're starting to work with these partners to support the ability to develop these agentic workflows on our system. And then, you know, we're using our capability to support multiple models to make it easier to rapidly switch between these models and support the agentic workflows that get defined using these building and orchestration frameworks. looks. Yeah, kind of going back to the beginning of the conversation, we talked a little bit about how your research, you know, from when we spoke in 2018 kind of led to, you know, Sam Manova, and your research continues to evolve.

41:35I'm really curious, like, how your research has evolved, and what it says to you about, like, the future of LLMs and agentic systems and how we build computing environments to support them? Yeah, that's a really interesting question, and one which, you know, my research group has continued to pursue, you know. So, I mean, I can start with, you know, the great thing about sort of being in computer systems, being a computer architect is that you get influenced by both sides, both by where the applications and the use cases are going, but also by the constraints of the hardware. And so in thinking about where the models are going, the models are fundamentally becoming more dynamic, right?

42:40So you're getting these mixture of experts, styles of models. You're getting environments that, you know, which you've got multiple users with different context lengths that need to be served. You are getting models like graph-based neural nets that have a kind of dynamic data access patterns. People are even starting to talk about fine-grained, sparse tensor computations, right? And so if you think about our current architecture as expressed, as I described, as sort of taking a graph and mapping it to the machine as a static process. And then, you know, if you think about, you know, re-changing that, you know, then it might take you a microsecond or more, right?

43:50Maybe a couple microseconds. So how about if you fundamentally think about something that's much more dynamic, right, that can change with much lower latency or can change just a portion of it, right? So you could potentially dynamically change a portion of the system to execute a different kind of graph, or maybe you are going to change the size of your tensor, right? So the tensor computation was 1024 by 1024 for the current set of token generations. And now you want to change it to 512 or maybe increase it to 2048. And so how can you dynamically enable these sorts of things? And so we're working on a architecture that we call a kind of a, we have a representation for a dynamic reconfigurable data flow architecture, right?

45:08So, you know, if you follow the history of CPU architectures, we started with statically scheduled architectures that we call VIW, very long instruction word, which were fixed. and then we eventually they found out that hey those weren't flexible enough and what you needed was a dynamically scheduled machine which is where what we have today right where you you you the scheduling is all done uh at runtime and so the question is can we make things more dynamic more flexible, maybe easier to change. And, you know, you still want to execute the static case very efficiently, but you want to allow dynamic cases to be implemented also with low overhead.

46:07So that's where we're going from an architecture point of view. Can you take these fundamental benefits of reconfigurable data flow architectures and put the D in front of it? Can you make them dynamic, right? With low overhead, without, you know, you can make anything dynamic, but of course, if you throw away all the advantages, then you haven't solved the problem, right? So that's one thing we're doing. The other thing we're doing is we're thinking, okay, hey, these agentic AI systems are going to become more prevalent. How do we create... And before we get to the agentic side, when you think about the...

46:48On this dynamic side, are there specific papers or publications or presentations that folks can look for that kind of illustrate the way you're thinking about incorporating that dynamic nature? So these are brand new ideas. We've just submitted a paper, and it's on our intermediate representation we call STEP streaming tensor programs. And, you know, I think we're going to put it on archive in the next, or it might already soon be on archive. And so that's the work we're doing there. And then, of course, this is always a sequence of research that involves both how to think about the software and compiler portions of the system, and then how to think about what you should do in terms of the hardware implementation.

47:54So I've got a couple of really bright PhD students thinking about the hardware. I've got a couple who are working on the software piece. So the early software piece has sort of come to fruition, but we haven't said anything about how to do the hardware yet because we don't know precisely how that's going to work. got it got it so still early days uh on that front yeah nice okay so you're about to talk a little bit about how you're thinking about supporting agentic use cases like from an architecture systems and software level from a systems point of view i mean you know there what we uh have been doing if you think about agentic systems right you've got kind of two components right which is uh you're given some high level tasks that you want to perform and you give that task to a fairly capable reasoning uh based uh lm like you know um uh deep seek uh r1 right 671 billion parameter and it comes up with a plan right and then that plan gets executed by a bunch of agents and may access the web, it may invoke a bunch of different specialized LLMs.

49:27And so one of the things we want to do is sort of how well could we use caching to cache that plan, right? So that it could be reused by subsequent users. And so it turned out to be a pretty good idea. And we got significant improvements, right? So caching is, of course, an idea that gets used all over the place. But you can't just cache the result. You've got to cache something that's got more semantic content to it than that. And it ends up being different from context caching, which has been pretty broadly. Yeah, yeah. If you just try to apply context caching, it doesn't work, right? So there's more to it than that, right?

50:14So that's what we're doing there. And then we were kind of expanding these ideas. I'm just very curious about the caching approach here. Like if you're not caching it as context, like that, you know, my next thought is like, are you like trying to like create some kind of symbolic representation of like the orchestration flow or something like that? Yeah, you have some representation of sort of what the plan looks like. And then sort of how similar is it to, you know, someone says, give me the financial results for company X, you know, you can do the same sort of thing for company Y, right? So that's the level at which you're trying to reuse plans.

50:58Okay. So that's one idea. And there's a whole bunch of research that's focused on how you make more efficient agentic systems. So caching is just one of the approaches, and we have other ideas there. And then there's the fundamental problem of accelerators are proliferating for both ML and for other kinds of problems because, as you know, everything is energy constrained, right? And so the only way to get around this is to become more specialized. and of course now that you've got this more specialized architecture you need some sort of programming environment some sort of compiler uh to map to it and uh this is typically the the long pole in any uh you know systems design you know even from sam vanova's point of view the the most difficult part of the whole endeavor has been delivering uh the compiler infrastructure for our systems so the question then is is could you use ai to make this problem easier and uh you know so lms are very capable but they're not as cape not as capable as you might think right in particular they can't really do any search of of of the space right and so if you want to make them do search then you need to add some agentic workflow around them and the space that you want them to be searching is like the space of compilers you want so in particular uh the the space that we that we want is we want to uh uh the problem that we that we focused on was sort of how could you take a new architecture as represented by some low-level representation that we call architectural-specific programming language and create a ML library.

53:14So, of course, NVIDIA has QD and N, right? So what if you wanted to create a new library, a new library for, for example, that could be targeted by PyTorch for your new architecture. Right. And the problem is, of course, the LLM is not going to be very good at doing this because it doesn't have any training examples. Right. This is a brand new architecture. So that was the idea. And so you want to create your new ML library, and you've got your new domain-specific architecture with specialized functions, and you want to create your new ML library. And so we created a agentic system for doing this and an adaptive self-improving loop around that in order to implement the solution.

54:22and it turned out to be a lot, lot better, like three and a half to four times better than just using an LLM out of the box to try and solve this problem. Imagining the agentic piece is like decomposing new things into oldish things or something like that, that the LLM can know something about? Well, just using the LLM to create samples in the space, right? And then having, you know, using the examples it comes up with as information that you feed back to it to inform it, to make it better. Right. And so when you think about like, when you kind of project forward, we've like created this new like approach to building software, Identic AI, and we've like, we're running it on the stuff that we've, you know, the old stuff, like it does it, yeah, does it like, you know, 10x, 100x, like, what do you see as like the broad opportunity when you think about like, in some future in which we're running agents on like custom you know on a hardware that is you know better suited for agents like what's yeah what's like is it incremental i'm trying to get a sense for is it incremental or is it like um i think the opportunities for five to 10x improvement in both uh performance and efficiency right so uh so i think there's a lot you know there's a lot of improvement that you can get by kind of using this uh very fast inference technique right so that and you've got you need to distinct for ultra low latency inference because you've got multiple llm calls that you need to make uh and i think there's a further benefit you get from fast model switching so you can support these multiple models uh and And you combine these capabilities with the fundamental underlying circuit techniques that we have pioneered at SambaNova to increase energy efficiency and the performance per watt benefits of the whole capability, I think will give you a factor of 5 to 10 improvement in performance per watt.

56:56Awesome. Well, Kunle, it has been wonderful to catch up with you. We've got to make sure it's not another seven plus years until the next time. Yeah, absolutely. He's definitely doing interesting things, you know, both on the research side and at Sam and Oven. It's been, you know, great to get reconnected. Okay. Yeah, it's been a lot of fun. And yeah, let's do it again much sooner. All righty. Take care. Thank you. Thanks. Bye.

57:26Thank you.

From the publisher

In this episode, we're joined by Kunle Olukotun, professor of electrical engineering and computer science at Stanford University and co-founder and chief technologist at Sambanova Systems, to discuss reconfigurable dataflow architectures for AI inference. Kunle explains the core idea of building computers that are dynamically configured to match the dataflow graph of an AI model, moving beyond the traditional instruction-fetch paradigm of CPUs and GPUs. We explore how this architecture is well-suited for LLM inference, reducing memory bandwidth bottlenecks and improving performance. Kunle reviews how this system also enables efficient multi-model serving and agentic workflows through its large, tiered memory and fast model-switching capabilities. Finally, we discuss his research into future dynamic reconfigurable architectures, and the use of AI agents to build compilers for new hardware.

The complete show notes for this episode can be found at https://twimlai.com/go/751.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Dataflow Computing for AI Inference with Kunle Olukotun - #751The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 58 min
Listen in VO