In short
Practical AI Podcast Episode Summary
Episode Title
Mamba & Jamba Date: [Insert Date] Guests: Yoav Shoham (Co-founder and Co-CEO of AI21), Chris Benson, Daniel Whitenack
---
Episode Overview In this episode, the hosts discuss the evolution and significance of AI21's latest model, Jamba, which integrates non-transformer architecture with attention layers. The conversation explores Jamba's design, its implications for enterprises, and the broader landscape of large language models (LLMs).
Key Themes
- Introduction of Jamba and its architectural significance.
- The evolution of AI21 as a company.
- The importance of language processing in AI.
- Practical applications of LLMs in various industries.
- The shift towards open-source models and community collaboration.
---
Key Discussions
- AI21's Mission and Background
- Founding Philosophy: AI21 was established to merge deep learning with reasoning, enhancing AI capabilities beyond just statistical models.
- Focus on Language: Language is seen as a deeper expression of human thought, necessitating sophisticated AI models to handle its complexities.
- Introduction of Jamba
- Model Architecture: Jamba combines structured space state models (SSSM) with transformer elements for improved efficiency and performance.
- Performance Metrics: Jamba boasts a context length of 250K tokens and is designed to operate on a single 80GB GPU, enhancing accessibility for enterprises.
- Enterprise Applications
- Unlocking Text Data: Enterprises often have large amounts of unstructured text data that can provide insights when leveraged correctly.
- Use Cases:
- Contextual answers from manuals and technical documentation.
- Summarization of lengthy reports.
- Automation of product description generation for e-commerce.
- Challenges in AI Development
- Reliability and Trust: Emphasis on creating reliable models to avoid 'hallucinations' (incorrect outputs) that can jeopardize enterprise applications.
- Task-Specific Models: AI21's approach focuses on developing models tailored for specific use cases, improving efficiency and accuracy.
- Open-source Collaboration
- Shift to Open-source: The decision to open-source Jamba was driven by the belief that community involvement would accelerate model improvement and innovation.
- Community Engagement: Encouragement for developers to experiment with and improve the model, enhancing its utility for various applications.
- Future Directions
- Focus on Understanding in AI: The goal is to develop systems that genuinely understand language and context, moving past the limitations of current models.
- Innovation in AI Systems: Anticipation of more sophisticated AI systems that integrate multiple functionalities rather than functioning in isolation.
---
Conclusion The episode highlights the innovative steps AI21 is taking with Jamba and the broader implications for AI in enterprise environments. The commitment to open-source development and community collaboration is viewed as a crucial factor in advancing the capabilities and reliability of AI systems.
---
Key Takeaways
- Jamba represents a significant advancement in the integration of different AI architectures.
- The enterprise has untapped potential in leveraging AI for processing text data.
- Open-source initiatives can foster rapid innovation and improvement in AI models.
- Trust and reliability are paramount for the successful application of AI technologies in real-world scenarios.
---
Additional Resources
- [AI21 Labs Official Website](https://www.ai21.com/)
- [Technical White Paper on Jamba](https://www.ai21.com/jamba)
- [Discussion on Practical AI](https://changelog.zulipchat.com/#narrow/stream/456003-practicalai)
---
Sponsors
- Fly.io: A platform for deploying apps and databases globally without operations.
- Changelog News: A podcast and newsletter combo for tech updates.
---
*Thank you for tuning into this episode of Practical AI. Stay updated with the latest in AI technology!*
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related tech is changing the world, this is the show for you. Thank you to our partners at Fly.io, the home of changelog.com. Fly transforms containers into micro VMs that run on their hardware in 30 plus regions on six continents. So you can launch your app near your users. Learn more at Fly.io.
0:42Welcome to another episode of the Practical AI Podcast. My name is Daniel Whitenack. I am CEO and founder at Prediction Guard. And I'm joined as always by my co-host, Chris Benson, who is a principal AI research engineer at Lockheed Martin. How are you doing, Chris? Doing great today, Daniel. How's it going? It's going great. The sun is out and summer is upon us, along with lots of new AI models and excitement going on in the space. And on that note, specifically as related to large language models, we're really excited to have with us today Yoav, who is the co-founder and co-CEO of AI21 and professor emeritus at Stanford.
1:28Welcome, Yoav. How are you doing? I'm doing good. Really a pleasure to be with you guys. Yeah, yeah. We're so excited to have you on. It's a show we've been wanting to have for some time now. I'm wondering if you could kind of give us a little bit of the background of AI21 and specifically maybe how you view AI21 as fitting into this wider landscape of LLM companies and technology. So maybe a good starting point will be to say why we started the company in the first place a little over six years ago. We started the company because we believe that deep learning, remember at the time, LLMs were not a thing, but deep learning was mostly applied to vision.
2:15We believe that modern day AI requires deep learning. It's a necessary component, but not sufficient. We believe that certain aspects of intelligence, this thing we often call reasoning, will not emerge purely from the statistics. And it's the sort of thing AI did back in the 80s. And we believe that we left money on the table and it's time to bring the two together. That's why we started the company. Now, fast forward today, what does the landscape look like and where do we fit in? So although I said that large language models, so very quickly we fell into LLMs. We were the heaviest users of GPT-3 when it came out.
2:56We decided to roll our own. And really, language is where the action is, because we often say that machine vision is a lens into the human eye, but language is a lens into the human mind, because there's no thought as intricate and nuanced as you want that can't in some way be expressed in language. Vision is a quote-unquote easy problem. Of course, it's not easy, but there's something to understand that this is a phone. I don't really care what the pixel is way on the side here. Always exactly true, but it's really primarily true. That's not true with language. Language connections matter terribly.
3:35You change a word here, the whole meaning of the sentence changes. In general, you can't escape semantics when you deal with language. And so it's harder, but if you crack it, that's gold. If you look at the enterprise, from the beginning, we were focused on the enterprise. 80 % of the data in the enterprise is text, mostly either not used or way underused. And there's a really good opportunity there. And that's kind of been our focus. So, of course, we're not the only people with large language models. We are one of the handful of companies that do really large, very capable language models. Our first model was called Jurassic One.
4:18It was going back a few years. It was not a most innovative model, but it was a good workhorse. It was a GPT-like, autoregressive left-to-right model. And at the time, it was slightly bigger, slightly better than GPT-3. Of course, both those models are by now eclipsed. We very recently released our most recent model called Jamba, which is very interesting in a number of ways. and we can dig deeper, but maybe at 30 ,000 feet, architecturally, it's different. It's not pure transformer model. It really is mostly based on structured space state model, SSSM as they're called. And we can speak about the advantages and disadvantages of those, but basically we took that architecture and added elements of transformers, the attention layer, to get the both of both worlds and you get performance that is as good as any model of its size, better than most of its kind of size group and extremely efficient.
5:23We have a context length that's larger than any other model of the size. We released, the version we released has a 250K context window length, although we trained it up to a million and yet it all fits onto a single 80 gigabyte GPU. And so your show is titled Practical AI. This starts to make it practical. That's great. And speaking of practicalities, you mentioned the focus on enterprise from the beginning. You also mentioned that a lot of data in the enterprise is kind of locked up in this unstructured text. I remember when I first got into data science, the focus is, oh, we're going to do big data and all of this cool analytics stuff with data warehouses.
6:09And I think that's sort of waned a little bit. I'm wondering if you could talk to that point, like why are enterprises, what types of value can they get out of this sort of text that's sitting around? And because I think maybe a lot of listeners, maybe they've tried these chat interfaces, whether it be ChatGPT or Gemini or whatever, but maybe they're less exposed to the workloads that enterprises are doing with LLMs. So could you give us a picture of how enterprises are unlocking value with that kind of 80 % of text data, maybe just by way of example or at a high level? Sure. And really the use cases are quite broad.
6:53The industries are very broad, whether it's finance or healthcare, education, or you name it. And the use cases are varied. But to pick some concrete ones, Let's say you have manuals. There are companies with thousands of manuals. And whether it's the end user wanting to, I recently did, I had a new sort of oven-microwave combination. And for the life of me, I couldn't find the relevant information in the manual. So I searched online and so on. It'd be really convenient to go and ask a question, get just the right answer. But even if it's not the end user, it could be the tech support person who themselves want to get quick answers.
7:35So that's an example. We call this contextual answers. Another would be summarization rather than response to a specific query. You have this 10K report that came out and you want a pithy summarization of it, maybe a summarization geared towards an aspect you care about. So that would be another use case. These are both ways of consuming data. Of course, Gen.ai is a terrible name, but we won't find that battle. You're stuck with it. Well, you know, you'll get me started. I'll start complaining about Gen.ai, about AGI, and so on. But certainly some use cases call for producing information, not only consuming information.
8:16So, for example, one of our use cases, very successful, our product descriptions. You have companies, retailers and e-commerce companies who have thousands of products that come online constantly. And writing a product description is labor-intensive, error-prone, expensive, time-consuming. And we're able to compress all of that dramatically. So these are some use cases. I'm kind of curious also, as you're looking at these opportunities in the enterprise and addressing these various use cases, as a company who is creating models and putting them out there for enterprises to use, for people who are not in the industry itself, how do you as a co-founder and CEO see your company as like, how do you say, let's go do this?
9:07We see the value in this compared to others that are making models. In other words, if you say, I'm going to make a model, what is it about that motivation which makes you think you'll make a difference in that enterprise market? And you're kind of representing all companies that do so, just to shed some insight on how a founder thinks in the space. I wouldn't purport to represent the entire industry, so I'll speak for ourselves. Fair enough. Overshot on my asking. No worries. But maybe somebody's comment to others. So first of all, the baseline is a general purpose, very capable model. There's a need for that.
9:44Now, there are companies who provide services using other people's models, and that's totally legit. If you actually own the model, you can do things that you wouldn't be able to do otherwise. And our emphasis, in addition to the general capability of the model, is in order to make it practical, there are two things, especially in the enterprise. So if you're using a chatbot to write a homework assignment, the stakes are low. A mistake doesn't carry a big penalty and probably nobody would read it anyway. But if you're writing a memo to your boss or to your prized client, and if you're brilliant 95 % of the time and garbage 5 % of the time, you're dead in the water.
10:33And so reliability is key. And as we know, large language models are these amazing, creative, knowledgeable system, but probabilistic. And so you will get, I don't like, here's another term I don't like, hallucination, but you'll get stuff that either isn't grounded in fact, doesn't make logical sense, and so on. And so you can't do that. You need to get high reliability. That's number one. I'll tell you a moment how we do that. But the other thing, it needs to be efficient. You know, for every customer query, you're going to pay$10 to answer it. It'll take you 20 seconds to answer it. That's not no good either.
11:10And so you need to address that also. So we have several things we're doing in this regard. The first is what we call task-specific models. In addition to our general purpose model like Jamba that came out, we provide language models that are tailored to specific use cases. You can think about it as a matrix. You have industries and you have use cases. And it turns out that while initially some, you know, you might think that, oh, I'm going to do a healthcare LLM or a finance, that's a little bit boiling the ocean. You want to be more specific. And one way to be specific is to think about what I'm going to use it for.
11:47These are the columns. So, for example, take summarization. That's a specific task. And now you can optimize your system. And I am deliberately saying system and not language models. I'll tell you in a moment why, but you can optimize that for that use case. So all companies now are experimenting with multiple solutions, as they should. And in this particular use case, a very large finance institution took several of their financial documents, several hundred, and tested various solutions. Our task-specific model and summarization and some of the general purpose models of other companies. and ours were just hands down better in terms of the quality of the answers they got.
12:31There was no hallucination, if you pardon the expression, very on point, very grounded and so on because it optimized for the task. But by the way, if the system is a fraction of the size of the general purpose model, so you get the answers immediately and the cost of serving is low. And this enables use cases that this latency and unit economics enable use cases that would just be unrealistic otherwise. So our task-specific models are one approach, and maybe I won't overload my answer with saying why it's not only models, but we'll get to AI systems. The other is, and it's related, having models are highly efficient.
13:16That goes to Jamba as an example of a model that's very capable, but not big. If I were to jump ahead and let's think about 2024, what are we going to see in this space? Among other things, you'll see focus on total cost of ownership of the reality of serving these models. You're going to see a focus on reliability. And you're also going to see focus on, not the term I hate, agents, but AI systems that are more elaborate than this transactional interaction with a large amount of tokens in. you know a few seconds token back thank you on to the next one more elaborate so this is I think what's going to happen technologically in the industry you're also going to see correlated with that the industry move from what today is mass experimentation to actual deployments we're seeing signs of it now and I think 24 you'll see this sort of phase shift there also
14:31This is a Changelog Newsbreak. On April 18th, Meta released the latest version of their open-ish large language model with state-of-the-art performance. The Verge rounds it up like this. Quote, Meta claims both sizes of Llama 3 beat similarly sized models like Google's Gemma and Gemini, Mistral 7b, and Anthropics Claude 3 in certain benchmarking tests. In the MMLU benchmark, which typically measures general knowledge, Llama 3 8b performed significantly better than both Gemma 7b and Mistral 7b, while Llama 3 70b slightly edged, Gemini Pro 1.5. End quote. What followed was your typical X-Bros posting N mind-blowing demos of what Llama 3 can accomplish, where N equals the number that a rival X-Bro just posted, plus one.
15:26Not very interesting, but two things that did stand out as interesting to me about this announcement. First, they didn't compare Llama 3 to GPT-4 at all, so we can only assume it still comes up short when compared to OpenAI's best. Second, they continue to call Llama open source, even though the license retains the commercial requirement of your business not being too big, which is 700 million monthly active users. So I guess Llama 3 is open for businesses of all sizes, depending on how you define all and sizes. You just heard one of our five top stories from Monday's Changelog News. Subscribe to the podcast to get all of the week's top stories and pop your email address in at changelog.com slash news to also receive our free companion email with even more developer news worth your attention.
16:16Once again, that's changelog.com slash news.
16:24so you have i love that you bring in this element of thinking about ai systems not just large language models or the model maybe that ties a little bit into what you were just talking about about more complicated workloads or automations that are likely coming as part of the solutions that people are building but i'm wondering if you could comment on that like where where does systematic thinking and the thinking about architecting AI systems fit within what you're seeing people do now and what you think needs to happen for them to get value out of these models? So the part of the answer that I'm comfortable speaking about has to do with what is out there already.
17:08And the others I'll speculate maybe at a little more higher level. So even if you look at task-specific model, they're really not models. They're little systems. So when you say you want to do summarization and you say I care about these elements, there's a little data processing and reasoning goes on before you call the language model. So you feed it. You don't just stick it into the context. You actually do some reasoning so you can steer the model in the right direction. And then when you get something back, you don't just spit it out. You don't sort of sample temperature zero and give the top answer.
17:42You get answers and you evaluate them with validators. And only when you're confident that the answer is legit, you return it to the user. And it may sound very expensive, but actually the operation of an LLM totally dominates in terms of the compute resources and time, these other elements. And that's an example of a system around the language model. But that's a baby step. what you're going to see is, and you're already seeing it now, but right now it's people touching parts of the elephant and doing it in a very ad hoc-y way, but you're going to see people stitching together multiple calls to a language model because a task may require multiple things.
18:26And it's not just chaining. It can be more complicated scripts that you're running, but you can't just do it. It's not like writing a scripting a script, a familiar scripting language and running it because the computing elements here are different. They're expensive and they're error prone. And if you just, for example, Cascade calls the language model, number one, it can be very expensive. And second, these errors compounds and you get at the end much more noise than signal. And so you need to worry about that. You need to execute differently. And so that's an example of what you'll see. And there are other aspects of these AI systems that you'll see come into play.
19:11The term orchestration is often used here. It means different things to different people, but very much you have these elements that are running either sequentially or in parallel. Somehow you need to execute this execution, kind of like an operating system, but an operating system with AI elements. And so we and other people use the term AIOS. Again, an overloaded term doesn't mean anything precise, but that's the spirit of things. I kind of want to get maybe to the roles that are interacting with this AI OS, because I think one of the things people are struggling with is how do I put the right talent in place to build these?
19:52Because you're talking about like programmatic, operational, systematic thinking, which is kind of like there's an element of engineering there, but it's not people that are necessarily building their own models. They're architecting these solutions and putting the right checks, the right validations in place. They're creating more than chains, these workflows. And there's some engineers coming to the table there, but there's also domain experts who maybe are able to speak into some of how the models are prompted. So do you have any kind of observations from your experience with how people are putting together teams to architect these solutions and these systems like you've just described?
20:38Is it, from your perspective, still going to be a heavy kind of engineering dominated type of process going forward? Or are you seeing a mix? What's your observation there? So my answer won't be based on an observation because the systems don't exist yet. They're baby solutions right now, but I don't think they represent what we'll see going forward. But in answer to your question, it very much will be a mix. There will be companies such as ours that will put in the foundational infrastructure to run these complicated flows. These will have to be extensible systems, and they'll be extensible in a variety of ways.
21:22Some of them, absolutely, you'll be able to have programmers write actual code and insert the code there, but there absolutely will be a role for low-code or even no-code specification of the flow you want on top of this framework. There will be a data scientist that will write validations of various kinds and data pipelines for sure. And so I think everybody from the developer to the data scientist to the business user who's somewhat savvy to the end user who just wants a system that works, everybody will have a role and interaction. And we haven't mentioned DevOps yet. DevOps here is going to be very important also.
22:08As we've kind of talked around the ecosystem a little bit and what, you know, about systems themselves, can we turn a little bit and could you tell us a little bit about as we're leading toward into Jamba, but I'd like to know a little bit about kind of where the company has been and some of the models that you have put out there leading into this one and kind of the heritage of how you've developed that. We'd really be interested in kind of how you've pursued that since you started the company. I can divide it into three periods in our long history of six years. That's an eon in AI these days.
22:44I had a different color hair when we started. As I said, we started by building Jurassic One. We just felt like we absolutely had to build it. And we innovated there, but in a minor way. We had a vocabulary that was five times the size of what was common at the time. It was rather than 50 ,000 tokens, we had 250 ,000. It was slightly larger than GPT-3, not to make a point just because it worked out that way, 178 billion parameters, a dense model. And that served us well. But the next phase in our sort of – we did many things. We had our own application called WordTune that had done very well, a reading and writing assistant using our technology.
23:29But on the models themselves, the next thing we put out are task-specific models, which basically is not really distillation, and it's not just fine-tuning. Like I said, it's putting a system around it, but at the end of the day, you get something compact for certain use cases, and that set is growing. That was our second phase. And the third phase was really seeking a way to make these models fundamentally more scalable, more efficient to serve, especially in this era of, you know, RAG kind of solutions. So you have stuff that you want to kind of bring in at inference time to influence the output of the system.
24:12And at some point, the system chokes. You know, we had a context window of 4K, then 8K, then 16K. Now, although some bigger numbers are thrown out, but most models choke at 32K, maybe 64K. That's not enough if you want to put it. So we wanted something that, now, if you were to run it on, you know, 64 H100s, you can do a lot of things. But that's not realistic. So the question was how to get something that's efficient, that can run effectively on a small footprint. And that's how we got to Jamba. With Jamba, you mentioned taking some things from kind of the Mamba architecture, the sort of SSM and adding in some transformer based things for those that aren't familiar with the kind of background with those types of models, maybe the kind of non transformer models that people were exploring.
25:07Could you give a little bit of context to that and why it was important for, I mean, you've already mentioned efficiency and other things, but why you felt it was kind of important in this generation of model to pull the trigger in a slightly different architectural direction? Sure. And for this, maybe we can double click a little bit about how these systems are architected. So at some point, the dominant architecture where the RNN, and then LSTMs, as you go left to right, the system doesn't remember the distant past. What it does, it carries with it the state that somehow encapsulates everything that it's seen so far.
25:48That's quite powerful. But as this path gets long, it gets harder and harder to encode and access that information that's been encoded. and it worked fine for vision because this in vision object recognition is something very local it's iconic iconic in the sense that what you see is what you get right like i said the phone you know this is a phone i don't care what's here so i go along i hit the phone so i don't need to remember but in language different and in fact if you looked at the benchmarks by the way another pet peeve of mine benchmarks are can be very misleading but that aside if you looked at the National Language Benchmark, they kind of puttered along with not much progress until Transformers came in.
26:32And Transformers, again, coincidentally, what is it, about six years now, they changed the architecture and they had the attention mechanism that says, no, I mean, as I'm going along, I can relate disparate pieces of information. And that allowed you to do things you couldn't do otherwise. And that's great. So the quality answer is shut up. You pay a price because the complexity is quadratic now in the context length. And that kills you, which wasn't the case with RNNs or LSTMs. There it's linear. I mean, you just, and so the question is, how can you have your cake and eat it too? Enjoy the benefits of being interrelated disparate kind of pieces of information and yet have something that's, if not linear, close to linear.
27:19And so Mamba, so first to say Mamba is a straight kind of left to right what's called SSM model and the structure safe space but its innovation was it was a version that allows you to actually parallelize the training and much more efficient but it still suffered from the lower quality of answers and so what our guys did was say okay we'll take this as a basic building block and Mamba is all of what four months old now. It's just Kadema recently. Yeah. But I said, that seemed like a really good idea, but let's now take elements of the transformer architecture and put it in. So every few, in our case, it was every eight or 16, depending on which version, layers, you put an attention mechanism.
28:08So you take a little performance hit, but not nearly as much as if you had transformers all the way. So that's kind of how it led to this particular architecture. Well, Yoav, you did mention that Mamba is only a recently released architecture and published architecture, but you've been able to move quite quickly. And I want to talk a little bit about Jamba and the release and all of that. But prior to that, it might be interesting for listeners. Most of our listeners aren't sitting in a company that is trying to be a foundation model builder, building these kind of more general purpose models.
28:44I'm wondering if you could give a picture a little bit behind the scenes, whatever you think would be interesting on what does it actually take to go from, hey, this idea we want to mix, kind of get the best of both worlds with Mamba and Transformers all the way to, hey, here's our blog post releasing a model. What were some of the challenges in that kind of middle zone? And what is that process like to determine, you know, from data set to exact architecture and sort of final training runs? So first I'll say that I don't think that everybody needs to be building foundation models. But as I said to somebody, if somebody's organizations are technical and wants to remain relevant, even if they're not building foundation models, they should understand how they're built.
Read the full transcript
29:39And if they really put their mind to it and their resources, they could build one because it really gives you a visceral deep sense of what's going on. Now, regarding the Jamba, we actually tried to be very transparent. You know, people, so this is our first open source model. And the reason we did it was that it is very novel. And there's lots of more experimentation to be done here. Optimization, serving the, you know, training these models can't be done on every type of infrastructure. Serving them similarly. and where you do serve them right now, we've had several years to optimize the serving of transformers.
30:21We want to enable the community to innovate here. And so we were quite explicit in our white paper, perhaps unusually so relative to the industry. So to the listeners who want to kind of get the nitty gritty, I really encourage them to look at the technical white paper. But I can tell you there's been a ton of experimentation of ablations that our guys did, trading off lots of, people use the term hyperparameters. It hides a lot of things that are very different from one another. But how many layers do you want? And, you know, how many Mamba layers, how many attention layers, batch sizes, all kinds of stuff.
31:04And what really makes a difference? It's hard to sometimes understand what makes a difference. And, again, we try to share the, for example, Mamba, I said that Mamba's performance doesn't compete with the performance of comparably sized transformer models. But that's at the, when you look at the details, it's actually quite competitive on many of the benchmarks. But then there are a few that it's really bad at. And that gives you a clue of why that's the case. It can latch on to surface formulations and syntax that the transformers managed to just abstract away from. And so we describe how, you know, you make this observation, you correct for it.
31:44There's lots of details that go into making these decisions. And then there's also pragmatic decisions. For example, we wanted a model that will fit on a single 80 gigabyte GPU. That was a design decision. And from that emanated a few things that, you know, we did put a bigger model and, you You know, certain contact windows will fit there, others won't. It's still, you know, 256K is humongous compared to the alternative. But we can also do a million and larger, but not on a single GPU. And so those are some of the design decisions and the rationale. Honestly, it is a process, although condensed, a process that involved, you know, hundreds of decisions that led to what we put out.
32:34That was a really great explanation. I appreciate that. As you were going through it, and I was thinking about the applicability for Jamba in the enterprise and kind of bringing the innovation, I'm curious is why, I know you had kind of alluded to the fact that Jamba early in the explanation was kind of the first open source model. And so I was wondering, as you're trying to enable enterprise innovation, what was the change in your thought process that made you decide to go open source with Jamba versus the earlier models? What was the thinking around that? I was curious as you said it and wanted to wait till we got to the end.
33:11Yeah, it really was very simple. We felt like if we were the only ones augmenting and pushing on this model, it wouldn't advance as fast as it could. and we saw that within days of our putting it out there, there was, I think today, I haven't tracked it, but when I looked about a week ago, the 30 ,000 downloads and I forget how many forks, but a large number of forks, some fine. So by the way, very important to say, what we put out is a base model, not a fine-tuned model. And we're very clear about it and we caution people for using it for production purposes or for user-facing application. And of course, we'll be coming out with our, In fact, we've announced that it's available for preview, our aligned model.
34:00But we felt like it was really important for the community to add value to this architecture. And that's why we did it. For those that are listening a little bit later on the podcast, so it looks like Jamba, at the time we're recording this, was released, at least on Hugging Face. Well, it was updated 15 days ago. And I see the blog post at the end of March, I believe. But now on Hugging Face, there's sort of 38 models I see with Jamba in the name. That's sort of not including those maybe that forked and just created their own special name. So already you're seeing this kind of explosion of a model family, I guess, which is quite interesting.
34:45I'm wondering over time as a company, you mentioned kind of not being the only ones working on the model family and wanting to see it become more. Is that observation kind of based on what you've seen in other model families, whether it be Llama 2 or Mistral and others? And there's sort of because when I look at a model like that, that's released, I almost immediately and I know people you mentioned DevOps, people have automated pipelines in place to create the quantized version of this or fine tune it for that on their data set. We had the noose research. We had a discussion about noose research and what they're doing in some of this areas as well.
35:26So what is the sort of innovation that you're hoping for with the kind of Jamba model family? Is it, you mentioned fine tunes or, you know, you're releasing the base model, there could be fine tunes, but I think also there could be much more than that. So what are you kind of hoping to see as people get hands-on with the model and try to explore various elements of how to use it? Yeah, fine-tuning is happening, will happen. Like I said, we have our own fine-tuned or aligned model, but that's not the reason we put it out there. The reason we put it out there is that people can contribute to the very model so others can benefit from it.
36:09And I think there's at least two areas where a lot of value can be brought. One is serving efficiency. For example, when you consume it on Hugging Face, it's less efficient than we consume it on our platform because we have optimized the serving and we'll continue to optimize. But a lot of smart people out there and we'd love for them to optimize it further and everybody will benefit, including us. That's one thing. The other thing is that I think we would really value it if this kind of model were able to be trained on multiple types of infrastructure, which currently isn't the case. And so I think by putting it out there, people now, they can look at the white paper, they can look at the model, and they can now enable further the training of such models, which will benefit everybody, including us.
37:03So as we start to wind up here, fascinating discussion. Thank you very much for taking us through all the insight. I like to wind up asking kind of where you think things are going and if you could address it potentially at two levels, both kind of where your own organization expects to go, what kind of thinking you have over whatever horizon is on your mind, but also give us insight into how you think the industry as a whole is progressing and how you expect that kind of servicing the enterprise need to evolve and you know with the strategies that are out there we'd love to understand how you're seeing the world in that way i think the key notion is reliability trust and reliability you need to have the same kind of trust in the system to be able to predict what they'll do, be able to understand what they did, as you do with other pieces of software.
38:01You know, we always have errors, you know. Even the Pentium had a bug. But that's an exception, whereas currently it's the rule for language models. So that can't be in the enterprise. And everything that I think about what's going to happen in the enterprise orients around that. I think you'll see special purpose models like our task specific models. I think you'll see AI systems increasingly sophisticated and robust. Right now they're not robust, they're experimental, but you'll see more AI system. And I think this may sound philosophical, so bear with me. But there's a question within the AI community, do these language models actually understand what they're talking about?
38:50They spit out this incredibly convincing stuff, very smart, something on point. And how can they not understand? And it's sometimes they're totally stupid. And everybody, we all have favorite examples. And I think we need to get to the point where we believe that the system actually understand what they're talking about. And what understanding is, is, again, it sounds philosophical. And there's a philosophical aspect to it for sure. but it has very practical ramifications. And so when I think about the future, all these pragmatic things, task-specific models, AI systems, but in the background, this notion of understanding.
39:32These systems need to really understand. That's what I'm looking at. Yeah, that's great. Well, I think as a part of the development towards that, certainly open models and innovation around these model families, like we talked about, I hope, is a key piece of that. And from a member of the community, I just want to express my thanks to AI21 for being a leader, both in terms of the thinking and infrastructure and innovation in this area, but also a leader in terms of putting things out there for the community to work on as a community. So thank you for what you've done with Jamba and really excited to follow AI21 and where you're headed next.
40:14So thank you so much for joining us, Yov. It's been a pleasure. Thanks very much for having me.
40:47slash community. Thanks again to our partners at fly.io, to our Beat Freaking Residents, Breakmaster Cylinder, and to you for listening. We appreciate you spending time with us. That's all for now. We'll talk to you again next time.
41:11Game on!
From the publisher
First there was Mamba… now there is Jamba from AI21. This is a model that combines the best non-transformer goodness of Mamba with good ‘ol attention layers. This results in a highly performant and efficient model that AI21 has open sourced! We hear all about it (along with a variety of other LLM things) from AI21’s co-founder Yoav.
Changelog++ members save 3 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
- Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs.
- Changelog News – A podcast+newsletter combo that’s brief, entertaining & always on-point. Subscribe today.
Featuring:
- Yoav Shoham – LinkedIn, X
- Chris Benson – Website, GitHub, LinkedIn, X
- Daniel Whitenack – Website, GitHub, X
Show Notes:
Something missing or broken? PRs welcome!




