Mamba, Mamba-2 and Post-Transformer Architectures for Generative AI with Albert Gu - #693

17 Jul 2024 · 58 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Mamba, Mamba-2 and Post-Transformer Architectures for Generative AI with Albert Gu - #693

Episode Overview In this episode of *The TWIML AI Podcast*, host Sam Charrington interviews Albert Gu, an assistant professor at Carnegie Mellon University. They discuss Gu's research on post-transformer architectures, specifically focusing on his recent papers, Mamba and Mamba-2, which explore state-space models in generative AI. The conversation delves into the efficiency of attention mechanisms, the limitations of transformer architectures, and the adaptability of new state-space models.

Key Points Discussed

Introduction to Albert Gu

  • Background in theoretical computer science and machine learning.
  • Transitioned focus during his PhD at Stanford, emphasizing efficiency in neural networks.
  • Development of structured matrices for enhancing model efficiency.

Post-Transformer Architectures

  • State-Space Models: A significant focus of Gu's work, particularly Mamba and Mamba-2, which attempt to overcome some limitations of traditional transformers.
  • The efficiency tradeoff between performance and memory usage is highlighted, especially in handling high-resolution perceptual data.

Limitations of the Transformer Architecture

  • Transformers excel in tasks with clear tokenization (e.g., language modeling) but struggle with high-dimensional raw data (e.g., images, audio).
  • Attention mechanisms in transformers are inefficient for certain data types, leading to the exploration of alternatives like state-space models.

Importance of Tokenization

  • The discussion emphasizes the role of effective tokenization for model performance.
  • Tokens serve as abstract representations of data, and their meaningfulness directly impacts model efficiency.

Hybrid Models

  • There is an emerging trend in hybrid models that combine state-space models and attention mechanisms.
  • Gu notes that even with pure transformer architectures, researchers are increasingly finding that fewer attention layers, with substantial model compression, can yield significant performance.

Future Directions

  • Continued research in selectivity mechanisms and state update processes is crucial for enhancing model efficiency.
  • The potential for state-space models in various applications beyond language, including multimodal data processing, is acknowledged.

Current Adoption and Applications

  • Gu asserts that while Mamba is still heavily researched, there are efforts in industry to implement these models in practical applications.
  • Companies like AI21 and Zifra are already developing hybrid models based on Gu's work, indicating a growing interest in state-space models.

Philosophical Considerations

  • The discussion touches on the importance of moving away from handcrafted architectures to more adaptable, learnable systems that can process data more efficiently and effectively.

Conclusion Albert Gu provides insightful perspectives on the evolving landscape of generative AI, particularly regarding the advancements in state-space models and the limitations of traditional transformer architectures. The episode highlights the importance of efficient tokenization, the emergence of hybrid models, and the potential for future innovations in this field.

For further details, the complete show notes are available at [TWIML AI Podcast Episode 693](https://twimlai.com/go/693).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00This is like one of the fundamental concepts of all of these alternate sequence models or post-transformer models. The entire axis that is the most relevant is basically the performance to efficiency tradeoff. And basically the predominant factor there is what is the model remembering in between time steps? And so we just talked about two broad classes of things, one of which is attention-like things where you still kind of preserve the idea of storing a cache of things. And the question is just which things do you store? And then the other family is recurrent or stateful models have like a fuzzier state that is.

0:42All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington, and today I'm joined by Albert Gu. Albert is an assistant professor at Carnegie Mellon University. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Albert, welcome to the podcast. Thanks for having me, Sam. Excited to be here. I'm excited to have you on the show. We are going to be talking about some of your research into what we'll call maybe post-transformer architectures for language models and generative AI. In particular, your research on Mamba and Mamba 2.

1:19To get us going, I'd love to have you share a little bit about your background and how you came to work in the field. Fairly recently, I finished at Stanford my PhD. Before that, I kind of went through a standard university. My focus was in both math and CS, and that has kind of informed the way that my research has shaped up. I went to Stanford actually starting in theoretical computer science, but I started shifting over to machine learning because although I enjoy the type of work that went into more pure math and like TCS problems, I cared more about the problems that the machine learning field cared about more, their ultimate goals and stuff.

1:54And I worked on a whole bunch of things during my PhD, including a lot of stuff on general efficiency, lots of different compressed structures for neural networks and how to make them theoretically and actually practically efficient. What's an example of a compressed structure for neural networks? Yeah, so I spent many years basically first working in a purely theoretical context on a concept called structured matrices. The idea is basically all like matrices are everywhere in machine learning. And these kind of like use up the bulk of the parameters and computation in neural networks. And structured matrices are basically any type of matrix that can be represented in fewer parameters.

2:32And so there's many, many different types of these from things such as sparsity to low rank to many, many more types. And one example that I and my lab mate and close collaborator, Tri Dao, have worked on are things called butterfly matrices, which later led to what's called monarch matrices. And these are all just different representations of matrices that can be represented in fewer parameters. They're less expressive and more structured, which means they could have better bias or fit for certain types of problems. And overall, they can be made a lot more efficient. So I spent quite a while kind of working on this, developing different classes of structured matrices and developing algorithms for them, such as matrix multiplication that can be done faster than the bound for unstructured dense matrices.

3:13And so this was pretty detached from applications, but later on I spent a while trying to apply them in neural networks. And the obvious way is just that everywhere you see giant matrix multiplications in neural networks, you can instead replace them with structured matrices that can be worked with much more efficiently. But this really kind of informed more is just that I was always kind of interested in efficiency and these type of problems throughout my PhD, but the type of problems I was working on were fairly different. But this also helped me learn and develop like a pretty large toolbox of these connections to like structure matrices that most machine learning practitioners aren't really aware of.

3:49And so this did directly influence some of my later work on state-space models. They've basically been a recurring theme in a lot of the work I've done from the early state-space models like S4. I needed to use these structured matrices to figure out how to make things fast. And then in Mamba 2, a major theme of that paper was about structured matrices again and just different ways that they appear in sequence modeling more directly. All the work I did in sequence modeling happened in the last two or three years. And I just really thought that the problem was super interesting. And I really liked thinking about conceptual capability questions, such as how can you develop models to handle super long context or potentially infinite context?

4:25Was that the core question for you, handling very large context? Yeah, so actually, the very first time I started working on sequence models, that was actually the main motivation. And I kind of liked this approach of using recurrent models from the very beginning. Of course, I was aware of what was happening in transformer space. But there was something that was nice about the recurrent approach that I always liked. And so I always kind of spent a lot of time and effort trying to understand that more deeply and making it work. But the very first question I was trying to address was actually in the context of RL.

4:59And it was about long-term credit assignment. And kind of the realization there was just that if you could fit your RL trajectory in the context of the sequence model, that's kind of enough. And then you just need to figure out how to make your model able to actually learn the dynamics across that extended length. Before Transformers, people were really working on this a lot in RNN land. And then Transformers came along and they have particularly good ways of doing this, but kind of at the expense of efficiency. And so then a lot of my work has been like trying to kind of still target some of these kind of interesting conceptual questions like, are there these new capabilities that we can get while also preserving efficiency and trying to do it in the right way?

5:36Can we take a step back and talk a little bit about the landscape of kind of post-transformer approaches that you're seeing evolve out there? State space models is a big part of that, but I'm assuming there are other efforts that we can talk a little bit about. And then you just mentioned that many of these approaches are related to one another. If you can talk a little bit about those relationships and the way to think about them. I think there was a really interesting note in the Mamba paper where you kind of talked about sequence modeling as being fundamentally about compressing context into a smaller state.

6:13And that is one of the things that ties all these different approaches together. If you can kind of riff on that and explain that idea, that would be great. So this is an interesting question. So I think actually before we start talking about post-transformer architectures, one aspect that is useful to draw some nuance on is about what sort of tasks do we care about or what sort of capabilities do we care about. I say that because I think there's kind of a mindset these days that the architecture of models has kind of been solved and attention is all you need. People are throwing transformers pretty much everywhere and for the most part, it does work really well in lots of places.

6:51But actually, I think that's not quite accurate and there's more nuance there. And what's happening more is that different types of models are good on different types of data. And people are finding ways to transform their data in ways that are amenable for transformers. And so when people are kind of using transformers everywhere, that's ignoring like the other processing aspects and so on of what else is happening in the data. Are there some examples that jump out at you? An example that I talk about a lot is that for working on modalities that are more close to underlying signals. So for example, audio, video, and images are kind of captured from some like underlying physical process.

7:27and it's like a continuous map you can view. And then our image is kind of a discretized version of that. But the closer you get to like a true perceptual modality, the worse attention is basically. That's where you actually need other types of models and architectures that are better suited for working with that type of data. And why do you think attention is worse as you get closer to these underlying modalities? It's a little tricky to describe, but it is basically related to a lot of the other concepts that, for example, you just asked the question about compression and stuff, and these are all sort of related.

8:00Perhaps one way I can try to explain it is that attention, the way we use it in machine learning, it's called softmax attention. And one can view that as an approximation to hard attention, which would be in the extreme case, you can imagine that you have like a big sequence or set of things and hard attention would be like an operation where I need to pick out one of these things and use that for my next layer. And softmax attention is basically an approximation where it can theoretically kind of attend to one thing, but also it's, you know, taking a soft average over things. But really, that's kind of, in some sense, that's the spirit of attention.

8:33And now, so when you're dealing with things like language, that's very chunked in into abstract concepts, sometimes you will need to pay attention to like this one word in my sentence is the thing I need to pay attention to, to understand, you know, whatever my task is or to pass into the next layer. But when you're dealing with a really, really high-resolution image, then it doesn't really make sense to say that what I want is some attention mechanism that's going to pick out one pixel in this image and pay attention to that one. This is also related to compressibility. So attention is very non-compressible.

9:03What it does is it stores a cache or a map of all possible things and allows you to pick and choose any one of them out of it. But instead, other models, what they do is they try to compress their data into smaller latent representations. And that's really what you need when you're dealing with super high-dimensional raw data, like a super high-resolution image, for example. That high-resolution image is extremely compressible. That's why we have such good compression algorithms that have been handcrafted in the past, but you can view other types of models as also doing some form of compression and extracting features out of them that attention is really not doing because attention is about storing the raw data and leaving it as kind of like a database or a book or retrieval thing that you can pick and choose things out of or you can attend to that, right?

9:51So these are all kind of like different facets of it. You know, it's a little hard to explain a lot of intuition, but that's kind of maybe one way of getting at it. And so the point is that the way that many like end-to-end pipelines are set up is that there's kind of some different model that works on the raw data and it could be handcrafted or it could be other types of neural networks such as convolutions are still very, very good for working with high-resolution data, such as raw audio waveforms or high-resolution images or videos. And some flavors of state-space models are also very good for that sort of data.

10:25And then the way that N-PyPlan's work is that you often first have a different model, such as a convolutional model, learn some sort of compressed representation of the high-resolution data, and then that can kind of tokenize it or discretize it in some way. So people talk about like patrifying the data, stuff like that. That's all done by basically like some form of like a convolution or some other type of model. And only when you've kind of tokenized things in a sense, then attention is really good for working on tokens. Can you dig into that last point? If you've got additional thoughts there, you mentioned that attention is really good on working on tokens.

11:03Tokens obviously comes up all the time in LLM. So we know that that's true from that perspective. But riff on that a little bit more. Is the core idea that tokens are fundamentally a more abstracted representation than some underlying thing? And that's why transformers work on them? Or is there something about the vocabulary nature of tokenization that is important? Yeah, so that's kind of exactly right. Like tokens, what they are, some sort of compressed, abstracted out, and hopefully like meaningful representation, like higher level representation of the data. And then if you can get your data to be processed into a state where you're working over these units of tokens, where each token really has an intrinsic meaning, maybe they can be composed.

11:49if you've gotten your data into the form where that is the actual unit you want to reason over, then that's where attention kind of shines. And so kind of using this analogy of like, on what level would you want to use hard attention? Like on what level of abstraction of your data would you actually want to pay attention to one individual unit? That is then the unit that like attention really shines at, like transformers really shine at. And so for example, we kind of talked about say in... perceptual modalities, it doesn't really make sense to want to attend to like one particular pixel in an image most of the time, especially if you have high resolution.

12:27So that's kind of the wrong unit of tokenization in a sense. I mean, in a sense, like everything is tokens, right? So the question really is like, have you kind of processed your data into the right abstract unit that is a meaningful unit of representation? And then for example, for language modeling, in the past, people would care about all these different regimes or like settings for language modeling. So there would be like word level language modeling, or there could be character level language modeling. And these would be treated as like different regimes. Nowadays, we kind of throw away that distinction.

12:58And the tokenizer is kind of part, it's just like another hyper parameter almost or like part of your model. And you don't really classify things in a different settings like, oh, I care about character level modeling or word level modeling, and so on. People used to care about that. And that does have different behaviors. And for example, it's been observed that tokenizer matters a lot for modern language models. And I think maybe that's kind of related to the idea that certain types of models, such as transformers, really work best when your data has been processed into the cleanest possible level of representation or intermediate language in terms of which tokens are the most meaningful.

13:34And if you had to constrain yourself to work purely on character level, then they start working worse. And later on, when we start talking about stasis models more in depth, I can say a little bit more about that. But kind of one intuition there is that it's like the wrong unit of representation, the wrong type of tokenization. And you would never really imagine yourself as a human if you're reading a block of text that you would ever pay hard attention to like one individual character in a word. Like, I mean, sometimes you do. And of course, like for example, spaces are important characters. But for the most part, we don't think like that.

14:05And we've already chunked things into different representations, higher level ones that we then reason over. And similarly, having that correct representation, I think is extremely important for transformers. And that's one of the things that alternate models are better at and one of the things that they try to address. The implication is that if you're not working on text or however you would define the sweet spot of transformers, then you're to some degree or another kind of force fitting your data so that it can be processed by the transformer. And by extension, perhaps there may be better architectures that natively work with your data.

14:46Right. Right. Yeah. But now once we are kind of like, have acknowledged that we are in the regime where like we have very nicely tokenized data. And this is really the most interesting regime. That's why people care about it. And language, of course, is a very natural and important example. then we can start talking about like what are the strengths and weaknesses of transformers there and what are the potential like post-transformer alternatives when we're talking about working on like you know language and stuff like that yeah here the landscape is rapidly evolving and there's a lot of interesting questions here so I think to sort of start talking about this once again I think it's kind of useful to think about what is the role of the attention and like what is the issue with it right And so the idea sort of is that, so attention is this operation where you kind of look at your sequence of data, like your sentence or your corpus of your document.

15:44And it will kind of do, for every single word or token in your sequence, it's going to look at that and compare it to like every single other token in the sequence and kind of do some pattern matching or something and use that to figure out what it should attend to elsewhere and use that to transform your representation for the next layer. and like this is in some sense like a very very flexible powerful um mechanism which is why it just works so well um and it's kind of been like the the thing that everyone uses um and and it's for a good reason because there's not really any other like even all these alternate models don't really uh capture that idea of like storing every single other word or token that you've seen and kind of using it as a database so that you can inform your next thing.

16:36So let's maybe focus to the setting of auto-aggressive language modeling, which is how GPT works and the most important setting perhaps right now. And so the idea is that in auto-aggressive modeling, you're trying to use your model to predict the next word step-by-step. And the way I like to abstract things now is that any auto-aggressive model has a state, which is its representation of all the previous contexts it's seen. And so as you're trying to predict or generate the next word, you have some state that's representing everything that you've generated so far or everything you've seen so far.

17:11And now what the transformer does is that it will store every single thing. Its state is basically a cache, what's called the KV cache, which is basically, abstractly, it's a cache of everything it's seen before in the sequence. and then the transformer can then use, it can attend to or pick and choose which previous words has it seen that it wants to use in a way to process or to predict the next word. And for some types of applications, you really just can't get away with anything other than this mechanism. For example, if you wanted to just copy a large block of text, you can't do anything other than just like storing that entire block of text, right?

17:50And so that's why kind of attention is indispensable even now and why it's necessary. The issue is basically that for the majority of cases, it is extremely wasteful because you're storing every single thing you've seen. And it's just not a good use of your representation power. And so all of these so-called post-transformer architectures are aiming to try to preserve the best parts of the transformer while addressing this limitation of efficiency. and in particular, trying to see in what cases, can you not literally cache everything you've seen? So that's kind of a useful way to kind of think about what all of the alternatives do.

18:32And most of them have sort of moved toward this like state for a current model that has a state that has a compressed representation of everything that the model has seen. And then finding ways to define kind of the state update mechanism that can compress the state and then leverage it in the future. So obviously, even before Transformers, people were using recurrent models like RNNs, and I'll touch on them again later. In the middle, people started using models like convolutions. And there were two works that kind of simultaneously started trying to use convolutions for language modeling. One of them was the CK-conv or continuous kernel convolution by David Romero.

19:13And the other one was the line of work that I worked on. Normal convolutions don't really work well for language because they have local context. And what we were trying to do was use global convolutions that kind of have an infinitely long sliding window that work over the entire sequence's past. And this did improve over some of the earlier attempts. But I think kind of quickly, I realized that these models suffered from the issue that they were not flexible enough, I guess. And what I mean is that, like we were talking about how attention, it kind of stores a cache of every word you've seen in the past and then it can pick and choose it can attend over any one of these words every single time step it's very very flexible in terms of which words it's attending over and now the issue with convolutions is that they're doing this kind of very clumsy thing where a convolution basically says that I'm always going to take the combination of my last say like k words with the same linear weights You're taking a linear sum of your last words.

20:18Those weights don't change. And this is just very, very different from the way attention can flexibly choose any one of them. It doesn't have the selectivity that an attention model would have. Exactly, yeah. And so that is kind of quickly realized that it was not really suited for language modeling. People kind of kept trying for a while. There was a bunch of other follow-ups. so the first convolution one was called S4 the safest model paper people tried to use it adjusting the architecture for language modeling in this paper called H3 other forms of convolutions were proposed which was this model called hyena and many many other variants were proposed but kind of like I think it was pretty clear to me that fundamentally these would never work for language modeling for the reason I just said and so I'll refer folks to my prior conversations with Dan Fu and Eric Nguyen, which talked about some of these other models.

21:17Yeah. And so they're, of course, very good for other things, particularly convolutions, like I mentioned before, very good for kind of more continuous high resolution data and maybe for other language-like modalities, but are not quite as like, maybe, you know, the tokenization matters in all of these concepts. But for normal language modeling, for that intuition I just talked about, they're not really that good. And the truth was, when I was working on this, I actually moved toward convolutions as kind of a crutch for efficiency. I had actually wanted to work with different types of recurrent models, but I couldn't make them fast enough.

21:54And I found a way to reframe some of them as convolutions that was pretty elegant and was good at other problems, but not for language. So after this is when I kind of try to move back to the fundamentals and figure out how I could make the recurrent models work better. And so the idea of the recurrent models, like even older things like LSTMs, is that it's back to this idea of how can you take your context and compress it into a state, perhaps selectively where you can pick and choose the right information you want into your state. And then how can you use your state to selectively produce the next output or use it for whatever task you care about.

22:32So this was something that I had been working on even from the very first time I started sequence modeling. Back when I started working on it in 2019, what I was working on was something called gating mechanism, which is something that people were using in RNNs back from a long, long time ago, like from the LSTM in the early 90s. And the idea there is that you have a recurrent state that's kind of evolving through time. And you want to kind of, they call it gate, something that controls the flow of information into your state. So you have like a recurrent state and then you have a new input. And then you ask yourself, how do I combine my input with my state to produce my next state for a future time step?

23:15And then the gate is something that controls how much of the input do I kind of let into the state. So I was working on this quite early on and my work on state space models can be seen as kind of an evolution of those ideas that slowly refine them into this pure and pure concept of the SSM. and then ever since the beginning i had wanted to try to figure out like the cleanest mathematical way of defining gates and so this is what led to this concept of discretization in space models which is a little bit in the weeds but turns out that there's like very nice mathematical foundations for where do gates and arnons come about and um it basically turns out to be exactly equivalent to this concept of a discretization which is kind of saying like how long do i wait in between how much time has elapsed in between inputs.

24:03And so you can imagine if too much time has elapsed, then you've kind of forgotten your old state and you're kind of resetting your state and focusing on a new input. Or if not enough time has elapsed, then that means that each input you're looking at is like individually not that important. So you don't pay attention to it as much. And so long story short, there was a lot of kind of principles that I was kind of thinking of how to incorporate selectivity. And then after S4, all these convolutional models, I went back and tried to figure out how can I incorporate this idea of selectivity again, but also make it efficient.

24:36And so that's kind of what led to Mamba. And at the same time that the SSM line of work was developing, another group was working on a model called RWKB, which is another recurrent model that was also trying to incorporate some early ideas of like, is there a way that we can incorporate some sort of selection there. And it turns out that they were also more or less using kind of convolutional models though. So kind of suffered from the same fundamental issues, but had some architectural changes that made it better in some settings. And so the selectivity was kind of one of the big breakthroughs, which was exactly what Mabo was about.

25:14And so the idea there was basically about the idea that it's very closely related to the idea of gating I just mentioned. and the idea that you need to choose how much of your inputs to pay attention to and how to select what you need to care about. And so to do that, basically, in the language of state-sense models, we basically just had to make a small change to... Basically, these are current models. They have some parameter that controls how much of my state do I decay and how much of my input do I add into my state. And all the previous models basically had fixed dynamics, which means that these parameters that control how much do I decay, how much do I pay attention to new input.

26:03So a static hyperparameter? Exactly, yeah. And these are learnable parameters, but kind of they're static along the sequence, right? No matter what data you see. No matter what data you see, they're always going to be constant.

Read the full transcript

26:19and actually I originally didn't want to do this but I introduced this for the first SDSM model S4 because it was the only way to make it efficient enough but then after realizing that we were just never going to be able to address language modeling properly I tried to go back and see how can we reintroduce this concept but make it faster And so after the first SSM paper, there was a bunch of other simplifications that people did, including Ankit Gupta. And I kind of found a way to make diagonal SSMs work and basically like simplify the structure of the model in a bunch of ways while still retaining its performance.

27:00And then finally, with all these simplifications, I was able to finally kind of like reintroduce this concept of gates or selectivity, but in a way that was still efficient enough. and I collaborated with my friend Tree who famously is the creator of Flash Attention and his expertise on how to make things really efficient on hardware. We were kind of able to do this, create this model that was a recurrent model that still had selectivity, still had all the other advantages of state-space models in terms of the flexibility of their state, having like a large expressive state that could try to compress the context by just doing it in a little bit better, more data-dependent way of incorporating what goes into your state and doing that in a way that could still be run on like a GPU efficiently.

27:47So that was kind of a bunch of things tied together there at the same time. So including how do you define your model in the right way that is as simple as possible while still having all of the capabilities that we were targeting, such as like this, making the transition parameters that we were just talking about instead of being constant, the main change was making them data dependent but when doing this you then lose the convolutional equivalents, you lose the efficiency and we found an alternate way to do it by kind of developing custom kernels targeted for GPU and like specific for GPUs and kind of materializing these states in the right part of the GPU to make it efficient enough and so this was, a lot of this was kind of Tree's work on the efficiency and the hardware algorithms to make this implementable.

28:37You talk about MAMA having this large expressive state. How do you contrast that to the KV store and you not wanting that being large? So the type of states that these SSM use and attention uses are kind of almost fundamentally different where the idea behind a stasis model is finding a state that's kind of the exact right size that can is large enough to remember all the information you care about or is large enough to store a compression of your history like a minimum compression is kind of what we're after we want a state that's large enough that for our task we can compress all the data exactly into the right size but And so you can't have it too small or else it's going to be too lossy.

29:29But if it's too big, then you're less efficient. And so the state is exactly this. Whereas with attention, you are remembering things that you saw as opposed to a compression of it. Yeah, you don't really have the control over what you're storing in your attention. And so that's one of the big differences with these state-space models and other recurrent models in general is just this idea that you have this controllable knob that kind of controls the tradeoff between efficiency versus how much can you remember. And so that's one of the main things, actually, I think that space-based models introduced over other types of recurrent models.

30:03So especially compared to older RNNs, pre-transformer, people didn't really think about this concept of how large the state was. And even a lot of the alternate models around this time weren't really paying attention to how large was the state and this idea that the state is a very important part of the model in terms of what it's remembering. Now, it is interesting, though, that there are ways to compress the transformer state a lot as well. And they kind of just work differently. So the way that SSMs compress state is that they don't really remember a notion of keeping track of which timestep or which token it has seen in the past.

30:41It kind of is just squishing them all together into some abstract state that hopefully the model has learned to encode in the way that is as good for the model as possible. but it's a bit of a black box. We try to give the model a flexible mechanism and guide the model toward learning what we want it to learn. But at the end of the day, it's not like in its state, we have no idea whether it is disentangled, actually remembering that this thing is this word in the past and so on. Whereas attention kind of does that by design and there's many efficient attention variants that in some ways kind of basically are still, the core idea of attention is still that it is doing some cache of the history.

31:26And so a lot of the efficient variants will just find, for example, the controllable knob is going to be like, if I want to make it more efficient, the controllable knob is, which tokens in the past am I storing? So sliding window attention is a variant that's pretty popular right now that's saying, I'm just going to remember my last, say, 1000 tokens. And then you can get into other types of attention structures that maybe remember some sort of strided version of this where you remember like things farther back in the past but forget but like in a coarser grain resolution or you can have like other sorts of maybe adaptive attention mechanisms that try to be more clever in picking and choosing which things they're remembering in the past and so this this becomes a little bit more close to a state-state model where the idea is like they're trying to figure out how to compress the past into a smaller state and having that state be as meaningful as possible and remembering the right information.

32:23This is like one of the fundamental concepts of all of these alternate sequence models or post-transformer models. The entire axis that is the most relevant is basically the performance to efficiency trade-off, especially say inference time. And basically the predominant factor there is what is the model remembering in between time steps. And so we kind of just talked about two broad classes of things, one of which is attention-like things where you still kind of preserve the idea of storing a cache of things. And the question is just, which things do you store? And then the other family is recurrent or stateful models kind of have like a fuzzier state that is a fixed size.

33:00And we don't quite know exactly what the model is learning in it, but we're trying to give the model an expressive mechanism to do its state updates. And the whole question is just It's basically, given my current state, given the next input, how do I combine them in a way to the next state? And so that's the major axis of variation. So I think especially after Mamba, there's been lots of other developments in this space. And a lot of people have since adopted this idea of state expansion. So like larger states, as well as the idea of selectivity became clear that it's extremely important. And then the major axis of variation then is basically like, even with this, there's still many, many ways to perform a state update.

33:39And so the type of state-space models that I work on have kind of stuck to this very simple equation that's been inspired by classical state-space models. And it has basically a very simple linear update in the state. And that seems to have been good enough or as good as any of the other post-transformer models. But there have been many other types of variants and attempts there. And this is definitely something that's still very underexplored. These improvements are primarily changing the way that the state is evolving as the model is being trained, as the data is being processed. Yeah, exactly.

34:18People have also realized that there is a fundamental limitation to any of these models in that the paradigm here is basically how do you do a form of optimal compression? The idea I just mentioned, how do you remember exactly what you need to remember? which I think is the core question or a core question that's needed to really optimize your efficiency to performance trade-off, right? That's the most salient question to me. Yeah, there are definitely some types of capabilities applications where the optimal part of that trade-off is that you just have to remember everything. So there is basically some niche applications or capabilities, such as I gave an example of what if you just need to literally memorize your whole document so that you can regurgitate it in the future.

35:18And then you have no choice but to just memorize that entire document. And it's these types of abilities. Or another issue, for example, is that... So this is one thing that people bring up, is this an issue with physics models, which is once you've decided what's in your state, then going forward, you can't go back and retrieve the thing that needs to be in your state again. As a human, I can think of my brain as sort of like a stateful model. And once I've forgotten something, there's no way to recover that memory, basically. That's just by definition the way a state works. And so this is kind of one limitation of the stasis models as well.

36:00The hope is that when it's trained on enough data and the right data, it will basically know what is important to encode or not. Just the same way humans, they can't remember everything, but they've learned to encode what's important and will kind of filter out or forget irrelevant things. And that is probably important to intelligence. Like I'm saying, the fundamental question is in terms of the efficiency to performance trade-off, you do need to have compression is fundamental to that. Yeah, at the same time, like there are settings where you can't get around needing to remember your whole history.

36:35And so that's why basically lots of people have started moving toward these hybrid models where the bulk of the work is done by a stateful model. And at the same time, kind of incorporating a small amount of vanilla attention, which is basically you can think of as a perfect kind of like a memorization of the context of the model. What's an example of one of those hybrid models? People have been trying to do these hybrid models from early on. So this model called H3 that came out of my old PhD lab was one of the first ones that really tried to explore the combination of a non-transformer model with a few sparse layers of attention.

37:13attention. And then people have been using these hybrid models since then. Some of the most recent ones that built on top of Mamba were, there's one called Jamba by AI21, I think. And then another one called Zamba by a company called Zifra. They kind of all had the same idea of basically like making a model that was mostly Mamba layers and then kind of a few attention layers. And many people have independently kind of found that around 10 % attention seems optimal. I also recently had a collaboration with NVIDIA where they were trying to explore this more systematically and kind of found the same thing again.

37:49And more interestingly, kind of like started making progress into figuring out what exactly those attention layers are useful for. And I think my current kind of best understanding is that basically this idea that like having some cache of your models of the entire history is indispensable for doing some sort of like tasks, for example, maybe even like, maybe it can help, for example, your model's state, the bulk of the state is doing its compression and like forgetting things, filtering things out. But maybe later on, if it realizes it needs to be reminded of something, that's where the KV cache comes in.

38:25And some like the small amount of attention maybe can help restore information in the state. I don't actually know, you know, this is all kind of just high level conceptual hypotheses. And I would be really interested if the research community kind of digs into these potential mechanisms further. But one thing that is a trend that's become apparent is that even with pure transformer models, it seems like people have been more and more finding that basically using hybrids where there's not very many attention layers and also finding new ways of reducing the size of the KV cache. For example, Character AI just had a blog post that talked a bit about their model and had similar ideas.

39:03And a bunch of academic papers just concurrently all did a similar idea, which was like, let's make a hybrid model where there's only a few like full attention layers. And in between them, you can have lots of layers of like a compressed model, whether it's a state space model, or whether it's like a sliding window attention. And then for your global attention layers, not only can you get away with like very few of them, but you can also share their KV caches. So normally, each one of these layers has a separate cache of the model's history. But people realize that you can actually share, if I have every layer share the same one and still preserve most of the performance.

39:36And this all kind of points to the idea that you just need some storage of the entire, every single thing the model has seen. And the size of that thing just doesn't have to be too big. It's just some small representation of every token you've seen in the past might be enough just to recover this one capability of needing to remember everything. And then the rest of the model can and maybe should be different types of layers. And of course, I'm very bullish on different variants of state-space models. But it's still, like I said, a very active area where all these different mechanisms are, there's a lot of room to continue developing them.

40:15So you mentioned the state-space evolution as one area that folks are iterating on to experiment with. And you mentioned also incorporating aspects of attention and these hybrid approaches, is there work being done around kind of tweaking the selectivity mechanism itself or those selectivity parameters? Like, do you think that's an important aspect of this as well? Yeah, I think that kind of just relates to the general question of what is the best state update mechanism? How do you design a mechanism that is like how your state evolves over time, and how is it incorporating new information. All of that is, you can say, are forms of selectivity.

41:09And the question is, what is the best one that has the best ability to process data while also being efficiently implementable? My current feeling is that once people started incorporating the selectivity and the state expansion that Mamba started doing, most of them seem like quite similar performance. I don't know that there's like qualitatively different performance. So for example, like in the gap between older RNNs like LSTMs versus Transformers, there was kind of a qualitative shift in the ability of these models. And now with some of the more recent recurrent stateful models, there's also kind of a like a qualitatively different or new type of mechanism that has been explored.

42:00And within that space, the work so far seems to be pretty, most of them seem pretty similar. So it kind of just seems like, it almost seems like any form of selectivity will be pretty reasonably similar in performance. It remains to be seen whether there's like a completely different type that will be like, you know, say like have different scaling laws or whatever. But for the most part, kind of all these papers I've seen have very, very similar scaling laws and performance and similar strengths and weaknesses. So yeah, I think the jury's still out. It's a very, very interesting area of research there.

42:34You started our conversation talking a little bit about the areas where transformers excel versus the areas where we kind of force fit our data to help transformers along. Now that we've talked about state-space models and Mamba, relate that back to this idea of data and in terms of performance and efficiency, have you found then that these models perform better for multimodal data? Talk a little bit about the way you look at performance broadly for these projects. Yeah, absolutely. I was actually just about to start talking about that myself. I think one of the areas where these models are most interesting is in new modalities and things that have co-evolved less with transformers.

43:28So for example, one simple example, even in the context of language, is I alluded to this earlier about how when we disentangle the tokenizer from the main model, or we force ourselves to use less optimized tokenizers, then the alternative models seem to start doing way better than transformers especially if you're trying to match on a compute-controlled basis because like we've been saying the main issue with transformers is efficiency because it's doing this wasteful storage of all the things it's seen and now when your data hasn't been tokenized into a nice representation. The whole point is that it only works if your data has already been compressed into a way where each unit you're storing is meaningful.

44:22And that's kind of this fuzzy notion of semantics or meaningfulness I've been talking about because if you're not storing one meaningful unit of information, then you're just wasting space. And so Sasha Rush and one of his students, Junshung at Cornell, had this paper called MambaBite, which investigated what happens if you're modeling raw bytes. And basically, and they basically found it seemed like fairly large gaps in performance where alternate models like Mamba were significantly better than transformers, especially when compute matched. And hopefully, this kind of makes more intuitive sense because this data is just uncompressed and then the transformer is just doing, it's just extremely wasteful.

45:06and so kind of seems like the less processed the data is the better some of these alternate models are at dealing with it in an elegant way and like compressing it the way it needs to be done now that particular phenomenon I actually don't know if I'm not aware of other groups that have reproduced it I would actually love to see it be reproduced and just like you know it was like an academic thing and seeing it independently verified especially at larger scale would be very enlightening. But I have observed similar things in other modalities. So in the original Mamba paper, we worked on DNA, which is a language-like thing that has a small vocabulary, but it's been tokenized in a way where one token is just one DNA-based pair.

45:56And certainly there has not been any co-evolution of tokenizers with the architecture. And so here, as a drop-in model, what we found was that Mamba was way better than Transformers there as well and I would kind of expect this to be true with lots of alternate models other than attention like any of these recurrent stateful models that have selectivity I think would probably be better and even like vanilla convolutions like S4 and Hyena seem to do reasonable there and when you say with modalities that haven't co-evolved with Transformers do you mean modalities for which that end-to-end architecture that includes that pre-processing into tokens hasn't been created yet?

46:39Yes, yeah, that's more what I meant. That's a, yeah, I guess what I said was kind of a... Let's say you've done that already. Will you find any advantage in a Mamba-type architecture since you've already put in that work? Are they roughly equivalent in terms of performance or are those architectures better because they're kind of, by definition, and fine-tune for a specific modality or use case? Yeah, so this touches on kind of the age-old debate of handcraftedness in machine learning. And so from one perspective, it's like if you already have something that's like a really nice tuned end-to-end pipeline, it seems like maybe you should just go with it.

47:22But also like we were talking about previously, even in that regime, kind of, I do think that the transformer attention is indispensable, but it seems like it's actually, it actually should be like a small part of the model instead of the bulk of it. And other things are better even there. But especially as you were saying, the less processed the data is, the more important it is to use other types of models. And now the question is, you asked a really good question, which was like, is that better or is it good enough? Like, should we just use, like when we have really strong priors on how to process our data, should we just use that?

47:58I think it's a bit of a stopgap because I think ultimately the way that things should be done is everything should be as kind of elegant and end-to-end as possible. That's kind of the whole premise and promise of deep learning. And the field has evolved to just slowly and slowly strip away all of these hard-coded, handcrafted definitions and processing and so on. And it will always evolve in that direction. And hopefully one day, you know, we'll be at the point where there's like something that's as generic and flexible as possible. And so kind of a, from a philosophical standpoint, I think it's important to continue developing things that are just more flexible and end to end and particularly can do this form of like, it can hopefully like we can learn the right compression or even for a tokenizer, right?

48:43Like hopefully we can learn the right tokenizer instead of having kind of these predefined tokenization schemes using an, like learn when using an end-to-end model. And so that's kind of a philosophical answer. But even without that, I think the reason why this is important is because when you have the handcrafted pipelines, you'll inevitably run into issues that maybe your handcrafted pipeline is going to address most of the issues, but then you'll run into other things. And that's just why learning everything from data is the most important to capture the way interactions really should be. What are some examples of the kinds of issues that you're thinking of?

49:23Yeah, so I think even in language and tokenizers, which has been our predominant example of a data modality and its processing schemes, these are still pretty ingrained, but it seems that the community is definitely observing issues that tokenizers cause. for example andre kipathy has a pretty viral tweet about all the issues that tokenizers cause including for example i think it's been speculated that one of the reasons why modern llms have a lot of trouble doing like arithmetic or paying attention to like even basic things like i think if you know ask it to generate words starting with a certain letter i think it has trouble with anything related to spelling seems like a big struggle for these models and that is probably it's been hypothesized that it's related to the tokenizer.

50:14Because spelling is at the level of characters, right? And the tokenizers operate at a higher level than that. And so then the model loses the ability to reason over characters. And so that's just one example of why doing things from raw data end to end as much as possible, I believe, is the right approach in general for machine learning. Reducing structure, putting it more compute. The tokenization is this artifact that's introduced to kind of pre-process a modality for the transformer. And what you're advocating for is getting rid of those types of pre-processing steps and using modalities that can directly or using rather models that can be directly trained on modalities and, you know, therefore will get rid of these artifacts that come from the pre-processing steps.

51:05Right, yeah, no matter what those models are. And I think we're probably, there's still a little bit of work to be done to figure out how to do that because the tokenization is there for a reason. It drastically improves efficiency and whatnot. So it's there for a reason, but it's just the way the whole field has always evolved is structure is there for a reason. Maybe because of compute constraints, for example, or we don't have the models for it yet. But as we get more tools for it, we always want to reduce structure, use better model suited for like more general unstructured data and uh and put in more compute and have basically have things learned as end to end and as automatically as possible it's just always um the the way the the way things should be done to really like learn properly right and what's your sense of mamba adoption in the wild is it you know exclusively an academic exercise or Are you aware of teams that have implemented it and are in some stage of putting it into production to solve some problem?

52:08And what can you tell us about who or what kinds of problems or any of that stuff? Yeah, so there's a lot of academics have certainly been very interested in it. there is also I mean I think that there is a bunch of teams that have been really interested in it for like even like language for example like some of the models we talked about like Jamba and Zamba these have all been done by industry labs also there's another one I forgot called Zamba which is another hybrid and so I can't speak to like whether they're using it in production somewhere but I I do expect that that is going to happen.

52:53And I do expect that there's going to be certain types of like modalities or applications where these state-space models really are very well suited for and they're going to be more and more productionized. Aside from that, all I can say is that at my company, Cartesia, we are also using state-space models to no one's surprise and using them in real products as well for generative AI where they work very well for lots of modalities. So I do expect this to become kind of an increasing trend. And of course, it might not just be Mamba, but any of these other models in this large family of stateful recurrent models, or lots of them you can also say are different other variants of state-based models aside from Mamba.

53:42Yeah, I think this kind of whole family will have a place to stay in the AI landscape. Beyond continuing to iterate on some of the areas that we talked about previously, how do you see the future for these types of models? There's a lot of ongoing things. So one thing I mentioned is continuing to improve the core mechanisms. Then there's a lot of work to be done in understanding them still. So for example, Mamba 2, it was an improvement to Mamba for efficiency. But actually, a lot of it was basically developing new theoretical frameworks for understanding sequence models. And as I mentioned, for example, relating them to structured matrices and structured matrix transformations on sequences.

54:28And I think it's just getting started in terms of the developments that occur downstream of that. so even just kind of my students at Carnegie Mellon have been working on lots of different extensions and applications of these things for example from sequences which can be viewed as one example of a simple graph structure it's like when your data is structured in a sequential order how can you generalize this to other forms of graphical structures with different types of dependencies so one of the things we've worked on is kind of extending it to other types of graphical structures another kind of related thing is how do you extend these to bidirectional sequence modeling in the most natural way and that's something that one of my students is going to release very very soon, hopefully this week actually there's a whole class of problems that are kind of how can we leverage the existing work done for Transformers to improve all of these other models and so this both includes in terms of model design of course which is one thing that Amatu tried to do but also in terms of how do you actually use the pre-trained models which have been scaled to very large, very powerful models how can we use them to bootstrap the space of post-Transformer models in a way that we don't have to reinvent the wheel entirely.

56:02And so I think this is a very important and interesting direction as well. And for example, my group has also started working on distillation techniques for how can you take a pre-trained transformer model and convert it into a state-state model with a much smaller amount of training. So these are kind of examples of big general problems. But yeah, there's... I'm sure there's like so many that I'm, I don't even, I don't have the ability to follow up with or keep up with. They're just like, you know, every day there's a new, there's a new like a Mamba follow-up, which is very exciting to see. And so, yeah, I'm very excited for where the future of the field continues to take us.

56:48Awesome. Awesome. Well, Albert, thanks so much for jumping on and sharing a bit about what you've been working on and how you see this particular space evolving. Yeah, thanks so much for having me, Sam. These were some really great questions. Thanks so much.

From the publisher

Today, we're joined by Albert Gu, assistant professor at Carnegie Mellon University, to discuss his research on post-transformer architectures for multi-modal foundation models, with a focus on state-space models in general and Albert’s recent Mamba and Mamba-2 papers in particular. We dig into the efficiency of the attention mechanism and its limitations in handling high-resolution perceptual modalities, and the strengths and weaknesses of transformer architectures relative to alternatives for various tasks. We dig into the role of tokenization and patching in transformer pipelines, emphasizing how abstraction and semantic relationships between tokens underpin the model's effectiveness, and explore how this relates to the debate between handcrafted pipelines versus end-to-end architectures in machine learning. Additionally, we touch on the evolving landscape of hybrid models which incorporate elements of attention and state, the significance of state update mechanisms in model adaptability and learning efficiency, and the contribution and adoption of state-space models like Mamba and Mamba-2 in academia and industry. Lastly, Albert shares his vision for advancing foundation models across diverse modalities and applications.

The complete show notes for this episode can be found at https://twimlai.com/go/693.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Mamba, Mamba-2 and Post-Transformer Architectures for Generative AI with Albert Gu - #693The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 58 min
Listen in VO