#244 Yoav Shoham on Jamba Models, Maestro and The Future of Enterprise AI

27 Mar 2025 · 52 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Episode #244 Summary

Episode Overview

  • Title: Yoav Shoham on Jamba Models, Maestro and The Future of Enterprise AI
  • Host: Craig S. Smith
  • Guest: Yoav Shoham, Co-founder of AI21 Labs
  • Release Date: [Date not provided]
  • Sponsor: DFINITY Foundation - Aims to shift cloud computing into a decentralized state via the Internet Computer (ICP).

Key Themes and Discussions

Introduction to Yoav Shoham

  • Yoav shares his journey from academia (Yale and Stanford) to entrepreneurship, founding AI21 Labs.
  • His academic background includes game theory and logic, which has influenced his work in AI.

AI21 Labs and its Mission

  • Founded AI21 Labs to combine traditional symbolic AI with modern deep learning techniques.
  • Developed Jurassic and Jamba, focusing on efficient AI model architectures.

Jamba and its Innovations

  • Jamba: A hybrid AI model architecture that increases efficiency while competing with traditional transformer models.
  • Combines state-based modeling with attention mechanisms.
  • Achieves linear complexity, making it more efficient in terms of latency and throughput.
  • Successfully scales, allowing smaller models to fit on a single GPU.

Maestro

The Orchestrator

  • Maestro: AI21’s orchestrator designed to work with multiple LLMs and AI tools.
  • Aims to provide enterprises with reliable, predictable, and efficient AI systems.
  • Targets real-world challenges of enterprise AI beyond mere demonstrations.

Limitations of Current AI Models

  • Discussion on the limitations of LLMs and their reasoning capabilities.
  • Yoav emphasizes the importance of moving towards AI systems that integrate various components for comprehensive solutions.

Applications of Jamba

  • Real-world applications include banking, retail (product descriptions), and financial document processing.
  • Emphasizes the ongoing transition from experimentation to deployment in enterprise AI.

Philosophical Aspects of AI

  • Yoav discusses the philosophical question of whether AI truly "understands" and the implications for future AI development.
  • Engages in the debate about the nature of understanding and consciousness in AI systems.

Future of AI Systems

  • AI is moving towards systems that go beyond just being language models.
  • The industry is shifting towards reliance on AI systems that offer better control, reliability, and predictability, particularly for enterprise applications.

Key Takeaways

  • Jamba represents a significant shift in AI model architecture, providing efficiencies not seen in traditional models.
  • Maestro is positioned as an essential tool for enterprises to orchestrate AI applications effectively.
  • The conversation touches on the philosophical and ethical dimensions of AI, urging a thoughtful approach to understanding AI's capabilities and limitations.
  • The future of AI lies in creating integrated systems rather than standalone models, focusing on real-world applications and reliable performance.

Where to Access Jamba and Maestro

  • Jamba 1.6 is available with open weights, while Maestro is currently in preview. Interested parties can find more at [AI21](https://ai21.com).

Additional Resources

  • For more insights on decentralized technologies, explore the DFINITY Foundation and the Internet Computer.
  • Follow Craig Smith and Eye on A.I. for updates on future episodes and discussions in AI technology.

Contact

  • Craig Smith on Twitter: [@craigss](https://twitter.com/craigss)
  • Eye on A.I. on Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)

---

This summary captures the essence of the podcast episode while providing clarity on key concepts and discussions presented by Yoav Shoham regarding the advancement of AI technology, its applications, and philosophical considerations.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I think language models, even the latest so-called reasoning models, they have their limits of what they can do. and really the world is going towards AI systems now, away from just pure language models, no matter what the specific architecture is. So language models aren't going away, but they really are becoming part of a bigger picture. This episode of Eye on AI is sponsored by the DFINITY Foundation. DFINITY Foundation is a Swiss not-for-profit that is home to some of the world's leading cryptographers, computer scientists, and experts in distributed computing. Their mission is to shift cloud computing toward a fully decentralized state by supporting the Internet Computer, also known as ICP.

0:49If you don't understand anything about the Internet Computer, you can see my episode with Dominic Williams. ICP's vision is that most of the world's software will be replaced by network resident software that's an evolution of traditional smart contracts. To achieve this vision, ICP is designed to make smart contracts as powerful as traditional software, all while remaining tamper proof, unstoppable, transparent, and verifiable. Since DFINITY's launch in 2016, they stand as Switzerland's most extensive blockchain research and development initiative and have been awarded more than 500 research grants worldwide.

1:39DFINITY remains steadfast in its mission to drive the advancement of the decentralized internet. Network resident software can now be used to run AI models and RAG infrastructure, preventing them from becoming quote-unquote hot wallets from where data can be stolen and increasing resilience. Furthermore, the technology has been designed to allow AI models to spin up and modify running web applications and internet services solo by addressing several key challenges. If you're interested in reading more about the internet computer, visit internetcomputer.org. I also encourage you to listen to my episode with Dominic Williams in which he explains the internet computer.

2:33I got interested in AI when I was in college and went to study in the States and went to Yale. And then I was offered the job at Stanford as an assistant professor there. And we're talking back a long time ago, 1987. I'm an old guy, but I honestly didn't know that that's what I wanted to do, but I figured Stanford was a good place and I'll have a good time until I figure out what I really want to do. And 28 years later, I retired as a professor. And along the way, I ran the AI lab there for a while. And my work there was largely theoretical, not-for-profit research. I worked on logic and philosophy and game theory as it pertains to computer science and AI.

3:30but in fact I have an online game theory course together with some colleagues that's more than a million people view it already this is for a not particularly useful topic but pretty but not useful but I do have an applied site so I started several companies AI21 being the last one that's me briefly and how did you start AI21 and then we can talk about what you guys are doing there. AI21 is seven years old now. And the idea germinated. I'd sold my last company to Google and was thinking that AI started to put all the eggs in one basket which was deep learning. LLMs weren't a thing, and large language models weren't a thing, but deep learning was, and machine learning in general.

4:40And it sort of worked. I mean, there's a reason why people gravitated to it. But my feeling was that, and this is statistics, basically, statistics will never give you the robust reasoning of the kind that one needs and that you saw in AI early on. And we should kind of put together a good old-fashioned AI with modern-day deep learning-based AI. So that's kind of an odd reason to start a company, but truthfully, that's the reason we started the company. There's me and my co-founder, Ori Goshen, who's half my age, twice my brains, and another guy who's not that younger than me, but somewhat younger, but certainly very smart.

5:26Amrun Shahshua was well known for having started and running Mobileye, the companies, but he's also quite an accomplished computer science professor, and I knew him from the academic side of things. That's why we started the company. And you were combining going back to symbolic AI and expert systems and infusing them with deep learning. Is that what you were doing in the beginning? So we did a lot of experimentation. And once, in fact, I remember early on, the first 20 employees, I'd recorded lectures on language representation, and I forced the poor guys to learn about logic and about temporal reasoning and about frame systems and so on, because they knew modern AI well, but they didn't know that.

6:24But it wasn't so much to take experts and glue machine learning onto them, but to think more deeply, what's missing from modern AI and how best to imbue it. And for about three years, we did nothing but technology building. And coincidentally, we started the company right around the LLM revolution, because that's the year that the Transformer paper came out. And it was clear to us that the work was in language, not in vision. Vision is, quote unquote, an easy problem. So the object recognition, it's obviously not easy. And, but to know that this is a mug, I don't really care what this pixel over on the side is, but there's nothing local in language.

7:13Take a sentence and change a word over here. The whole, you can't get away from semantics. And so it's clear that that's where the work was. And then coincidentally, Transformers came out. And so we felt we had to be very good at language models. And we, if I can say so myself, became very good at that. When GPT-3 came out, we were, I think, the heaviest, most demanding users in terms of the use cases and how we banked on it. And at some point we decided we need to build our own. And we built our first model called Jurassic One. It's not a most innovative model. It's a standard left-to-right autoregressive model, GPT-like.

8:02It was slightly bigger, had five kinds of vocabulary. People who tested us found us a little better than GP3. But it was really, at the time, a novel thing to do, maybe even a dairy thing to do. And a small company doing and putting out this. and at the time we so we were playing around with technology and building this language model but at that point this is already three years into our existence or there about two and a half years maybe two years uh we didn't want to just be a research lab and we had and still have really an amazing collection of talent honestly i've had the luxury of working with very smart people but such a big kind of dense density of brain i'm always just the dumb person in the room um but we didn't want to just be uh and i i i i always mention uh deep mind as an example of a group i really admire but early on they really didn't have a there's not a big business in solving atari games so the question of what business should we be in and it was clear to us We wanted to be in sort of the enterprise kind of B2B world.

9:14But there was no business there then. And so we had built our own business. It was an application, which today is a very busy space of writing assistants. But at the time, it was quite novel. There was nothing like it. We had an LLM-powered writing assistant called WordTune. and it really was amazing. It still is. It crossed, it very quickly crossed 10 million users. It was a freemium model generating kind of a lot of revenues. But it was never meant to be our core business and it's still there, but our core business became the enterprise and it was somewhat audacious. You know, there we are, we're three years old, we have this application called WordTune, We have our own model.

10:10So we're, let's say, competing with Grammarly on writing assistants. We're competing with OpenAI on LLMs. And we're like 30 people. And then we raised some money, but less than OpenAI. So it was a little audacious, but that's what we did. And in time, we focused more and more on the enterprise. That's kind of roughly our arc. Yeah. When you say focus on the enterprise, I mean, there are a couple of things I'm interested in. I was doing some reading about Jurassic 2, actually. And I read that you guys used Amazon SageMaker to build that. Is that right? Our Jurassic models were trained. We actually trained on both AWS and GCP.

11:03we tended to use the bare metal not use much of the infrastructure on top of it we were quite portable our models starting with about a year ago were trained only on Amazon because they're broke new grounds there no there were no longer a pure transformer architecture the thing was transformers i mean why did transformers take is because up until then there were rnns and you know things that uh look at the input and didn't have a long context to look back into and like i said you know vision is kind of local that work but language isn't so the attention making transformer allowed you to relate different district parts of the input The thing is you pay a high price for that.

11:56It's suddenly becoming quadratic complexity. And if your input is like a thousand, a thousand squared is fine. But now we're pushing a million, a million squared is not fine.

12:12And so our guys created this new architecture, a model called Jamba. It's the Jamba family. In fact, we've just now released Jamba 1.6. It's amazing. And the way the model kind of is architected, it's mostly a so-called state-based model, specifically based on the Mamba model that came out of academia. We were the first one to really just scale it to an art model. So the advantage of state-based models is that they're more efficient. You're back to sort of linear-ish complexity. but you're not and you're not as bad because you try to remember the path and the state you carry along that's why they call state based models but that's not as effective at the attention mechanism of transformers what our guys did every so often every eight layers in that case put a attention layer and there were a lot of ablations about exactly how to do it but the end of it was a model whose performance in terms of the quality of the answers was competitive with the best model of similar size.

13:19And in terms of latency and throughput and memory requirements, there's no comparison, way, way more efficient. This is a mixture of expert models. So you have a total number of parameters and a smaller number active at any moment. So our Java small was 52 billion parameters, the total of which, oh my God, 12 billion, I think, were active. And the large model was 398 billion, of which 94 billion were active. So the small model would fit on a single GPU, which is mind-boggling. And the large model would fit a single 8GPU pod. And the throughput and latency are just no confidence. They're almost linear.

14:08and not quite linear because you have a little attention, but almost linear. So those are the models that we're putting out now. And like I said, we just released Jumbo 1.6, which is our latest. Can you talk a little bit about Mamba? I haven't done an episode on that, and a lot of people, a lot of the listeners have heard about it, but don't really understand about its development and how you integrated it? So first of all, I really encourage you to speak with the creators of Mamba. They're nice, smart guys, academics. I'm sure they'll be able to go into as much detail as you want. But like I said, it's a state-based model.

14:59So it's order regressive goes left to right and looking at the input token by token and predicting the next token and doing the stochastic gradient descent to update the weight. The difference is as it goes along, it carries state and updates the state. and now the state can't doesn't capture the you know entire history it can't but it captures enough of it so you have this great it's like a memory aid that the model has so that's how it operates and like i said it has certain advantages of speed and an advantage over you know more vanilla left to right kind of RNN type mechanism, but not quite competitive with explicitly every time looking at the entire history with the attention hits.

16:02I had Sepp Hochreiter, I can never pronounce his name the way he did, the creator of the lstm on fairly recently and he's pursuing that and of is building a company around it around sort of an improved lstm is mamba nothing like lstm or is it It's a completely new architecture. Completely different, it's hard to say. And I think what Mamba did, it opened the door for people to experiment, to be brave enough to either experiment with totally new kinds of models or tweaks. Mamba itself had had several versions. We ourselves didn't take Mamba as it is, but we did a lot of ablations. people have looked at now taking a second look at bi-directional models not these left to right auto-reversive one but you know with masking so i think i think something very healthy that people aren't intimidated by successive transformers and it'll be interesting to see what comes up i have to say i think language models even the latest so-called reasoning models you know, the O1s and R1s and so on, they have their limits of what they can do.

17:39And really, the world is going towards AI systems now, away from just pure language model, no matter what their specific architecture is. So language models aren't going away, but they really are becoming part of a bigger picture. So in this hybrid model, you have you were saying you have a mamba layer and then several layers above that of a transformer exactly you stack them with mamba transformer mamba transformer is that right yes except it's not mamba transformer it's there's a jamba block and by the way i really encourage people to go and read because we actually gave quite explicit description of jamba the architecture and how it was built we we actually open source the open weighted the model we felt it's particularly novel for for the community to help kind of improve things but in the jamba block there were eight layers one of which is transformer the rest are mobile layers and they're also a mixture of expert layers there yeah that's essentially what it looks like and so So how is it performing on all the various benchmarks that everyone is watching these days?

19:01So when we first make a general kind of pontificate generally on this topic of benchmarking, you know, we are somewhat cynical about the benchmarks. they're well-intentioned and they give you some signal about the quality of the model but it's a very weak signal for two reasons one which is sort of an unavoidable one is the correlation between the tasks that are encoded in the benchmark and what you see in the real world is not strong. And so you may perform great on GSM 8K or on MMLU or score high on, you know, LMSIS or Arena Hard or something. And that doesn't necessarily mean that when you go and actually try to work on a real problem in the real world, the model will be better than the model that scored less.

20:04That's one reason we take it with a grain of salt. The other is, unfortunately, these are really easy to game. and now we've been so by the way to answer your question we score well on these despite my cynicism there are a measure there are an aid in optimizing your training we never cheat but we take it with a big grain of salt and I can tell you that our customers also take it with a big grain of salt there's you know especially the sophisticated sort of people they'll always make a face when you discuss these benchmarks but I'm saying that we're okay yeah I mean you can train to those benchmarks can't you first of all yes and I do know about some others where people explicitly gave instructions, you know, we need to win this benchmark.

21:16And you could do it, even without cheating. I mean, the most blatant thing you could do is take the test data and put in the training data. But if you don't do that, you just spend more cycles in that area, mathematics, logic, you know, whatever it is, whether it's natural data or produce a lot of synthetic data in that area, you'll do better there, which doesn't necessarily mean you'll be better in real applications. So people do that. So if you throw enough flops at it and you throw enough data at it, you'll get good performance of the benchmarks. Back to DeepMah and Zalpha Go. oh, people have been using reinforcement learning directly to train the models, not to refine or fine-tune the models.

22:10What do you think about that? Did you use that approach with Jamba? By the way, this year's Turing Award winners are the two pioneers in AI of reinforcement learning, Andy Bartow and Rich Sons. Certainly RL is having its moment now. for a variety of reasons. RLHF is a misnomer. It's really not reinforced with learning. It's sort of reward modeling, in this case, by humans. But RL is playing a role in the jumbo models that's out there. We didn't use it. We didn't find the need there. Recently, we have. but when you're referring to R1 or O1 those quote unquote chain of thought models the RL plays a particularly important role there as it happens when you look closely it's not enough but that's really the primary way in which a training time not at inference time, but at training time, you try to guide the model towards these chains of inference, chains of tokens, the correct way to say it, that are more likely to be productive than others.

23:35And that's the role of RL in those models. But you guys, how did you train Jamba? We did, you know, I'll call it standard pre-training with a lot of quality data. was natural and synthetic. We then did a lot of alignment tuning of various kinds.

23:58And using really, I mean, SFT supervised fine tuning there. And that's primarily how we trade Jamba. So from your point of view, Jamba is this new hybrid architecture or hybrid model. Are you going to continue building larger models? What's the goal? You said at the outset that your interest was in working with enterprise. Right. So I'm not in a position to sort of announce what we'll do in the future. But I will say that we believe language models are and will continue to be important. It's very important for a company like ours to be very good at them. So I don't think that's going to go away.

24:55But increasingly, the emphasis is on AI systems. and in particular we've just announced maestro which is a planning based orchestrator that works with multiple llms as well as other tools and code and what have you and orchestrated to work in a way that really will serve the enterprise to give the reliability and predictability and increasingly we think that would be the emphasis in the industry and certainly for us. So are you going to continue coming out with larger and larger Jambas or is size of both compute and training not your goal here? Our goal is to produce the technology that serves the enterprise reliably.

26:01That's our goal.

26:05Language models are part of it, and the quality of the model is a function of multiple things. Size, unlike a sandwich, in language models, bigger is not always better. and it's not only the amount of training data. A lot of it is the quality of the training data and the ability to specialize the model to quickly, efficiently, reliably in certain domains and like I said, the efficiency of it, the way to serve it efficiently. That'll continue to be our focus but like I said, I can't speak specifically to what kind of models we'll be training. Yeah. And you were saying with Jamba, it is open weight and it has a particularly large context window.

27:07Is that right? And very high throughput, token throughput, compared to transformers. you were saying you can fit a model on a single GPU. So what are some of the applications that you imagine with this model? Well, I don't need to imagine, but we have customers using it for a variety of, you know, the types of applications aren't that mysterious. You see it across the board. We are working with a bank in a call center setting to do question answering and troubleshooting. We're working with a big retailer to do product descriptions. These online retailers have millions of products, and writing product description is time-consuming, costly, error-prone.

28:09So we're helping with that. we are working with financial institution to query financial documents. And so the applications really vary. And what we're finding is that enterprises realize there's a big gap between a fashion demo and something that actually works. And what we see in the industry is that we're in what we like to think of as the second phase of the modern AI revolution. So up until maybe two and a half years ago, so you had sporadic experimentation. The CIOs or the CEOs had just gone, finished their migration to cloud and they didn't have energy and go ahead and play AI if you want.

29:03And then suddenly somebody turned on the switch, and today there's no CEO in the world that doesn't say I'm an AI-first company or I want to be one. And so we're in the era of mass experimentation. We have kind of customers, partners with hundreds of use cases. but the drop off between an experiment, a PLC to a prototype to actual deployment is precipitous. There's one study by AWS that shows the drop of 94%, 6 % of the project's CO2 production. And we think we're on the cusp of the transition to the third phase of mass deployment. But for that to happen, you need to deal with these two issues of have the model be efficient, because the economics of LLM are not the economics of traditional software.

30:09It can be extremely expensive, certainly with these quote-unquote reasoning models. and the second is accuracy because these models are probabilistic machines and sometimes they give you wonderful, creative answers and then they give you total garbage and so if you're a high school student writing an essay, it's okay if you're occasionally wrong, maybe you have time to correct and if you don't, that's okay worst case you'll get a poor grade but if you're writing a summary of a patient checkup and you make a mistake, or if you're writing a memo to your client or your boss and you make a mistake, you can't be covered for that.

30:58And tackling that is required to get to the phase of mass deployment. And this is where AI systems will come in because no language model, no matter how good, how many parameters, how much data, how many quadrails you put there. They will never, we call it prompt and pray. So long as you're in the prompt and pray world, enterprise won't adopt it at scale. So most of your deployments on-prem where people have their own GPUs, or is this primarily being accessed through the cloud? And then what do you build around the Jamba model to protect against hallucinations? So first of all, our models can be accessed really anywhere.

31:55You certainly can get on the cloud. We're on all the hyperscalers. We have our own SaaS. but we also in fact were viewed as the best on-prem solution and we have it installed in some really significant companies so it's all of the above in terms of how do you get to high accuracy

32:25Again, it's a fool's errand to try to force the LLM to be deterministic. When you call the LLM twice in a row, you'll get different answers. This is the way it works. And so you need to build a system around it. Also, an LLM is not that right. You need to call tools. If I'm going to access my proprietary data, you need a rag-like system to access that. that. Certain things, God did not put neural nets on Earth to do arithmetic. HP gave us a calculator in 1970. We don't need to reinvent that wheel and get a crooked one at that. So you have these pieces of code that do stuff reliably. You want to use those.

33:11And so what you want is a system, not an AI, an LLN, an AI system that knows to work with these aliens, by the way, not just Jamba. So Maestro, our new sort of orchestrator, will work with Jamba, will work with others as well, and is just smart about when to use what, which for what, how to coordinate them, how to parallelize things, how to plan ahead and give some performance guarantees about how much will it cost them, And right now, if you go to the so-called reasoning models, they'll go and, you know, think, generate these sequences of so-called thinking tokens and come back after half an hour with an answer, maybe a bad hour, an answer, a good one, but you've just spent a dollar.

34:08You can't have that. Not in the enterprise, not at scale. And so that's what the world we're looking at now of these. orchestration systems that can work and in a production environment and really juggle everything in a way that transparency is used to give the user control when they need it and so that's the world we're and this is what uh i think certainly where maestro is but i think where the world is going in general so what i'm going to say is what you see now happening the enterprise before such an orchestration tool became available is that people do manual coding of static chains. So they would go and write a script, write a program.

35:02We'll call the LLM here, check its output, call another LLM or the LLM again, call some custom code. Somebody went and wrote that code. And that's, again, that's manual. It's a one-off for one use case. You have a different use case, a different workflow. You need to start it from scratch. And so that was the way, until now, people controlled for the unruly behavior of LLMs. But now we have AI to help control this through a very deliberate planning process. Yeah. What do you think of this GPT 4.5 that you let it run long enough and it comes up with an answer that's more accurate, less likely to be a hallucination?

35:56Is that a system, like as you're talking about, or is that a model? First of all, I have to say, to caveat it, they haven't really shared the data behind any of their models. But it clearly is a model. But it's a model of those chain of thought variety, where they've trained the model to predict, not just the very next token, but a long sequence of tokens that hopefully represent some useful thinking pattern or reasoning behavior culminating the right answer. And sometimes it works. The thing is, it's still a prompt and pray regime. It can improve the answers. It's costly, like I said. It'll cost you much more than a single call, even a single call to an LLM, which itself can be expensive, and you get no guarantees.

Read the full transcript

37:03And it's something of a misnomer to call these reasoning models because the thing that you see along the way sometimes corresponds to what you might think of logical or coherent reasoning, and sometimes it's just stuff. and so I call them large using models one of my colleagues Rao one of our professors also one of our consultants calls them chains of thoughtlessness and it's a little unkind because they do service a useful purpose but they will not give you the reliability and cost control and controllability in general that the enterprise needs. Maestro is out. Jamba is out. What's next in you personally, in your work?

38:09Me personally, it's complicated. I try to not interfere with smart people who are doing useful work.

38:19I will tell you, I have a pet project, which I do in my spare time on the weekends, together with my colleague Kevin Layton-Brown, who's a professor from Canada, brilliant. And also, he happened to be working after the company, but this is a totally academic work on, we call it understanding, understanding. What does it mean to understand something? And do LL really understand? It gets to interesting philosophical questions, but in my spare time, that's what I do. But in terms of stuff that actually matters to the world and the company, we think this world of AI systems that are based on explicit planning is, in its very early stages, there's a ton of work to do there.

39:14What we've just released is a limited version of our system. It allows you, here you are, you could call an LLM, Jamba, but you could call Opus, you can call Grok, you could call GPT, 4.0 even, or instead, you could call Maestro, who would use those respective systems, but amplify the performance dramatically. On average, it amplified the accuracy by 50%. But that's the early functionality that we're making available. There's much more work to do in an AI system on how to plan things ahead, both at training time and inference time. And so there's a ton of work to be done there. I wanted to ask, you mentioned understanding, and as you said, this is kind of a bottomless philosophical question, partly because no one has defined understanding in any kind of a scientific way.

40:27But Jeff Hinton certainly believes that LLMs understand and think. What's your view? And where does understanding... There's a gradient or a spectrum of understanding that ends in awareness. Where do you stand and where do you think the current LLMs are on that spectrum, or not LLMs, but AI models? And where do you think, do you think that it'll continue to climb that ladder or we've created something that has something that we call understanding, and that's as far as it's going to go. A lot of stuff packed into that. Let me first say that I have the utmost regard for Jeff. And for another guy, Andrew Ng, my colleague at Stanford, and both brilliant people, but I mention Andrew because Andrew and Jeff had an online conversation about whether LLNs understand.

41:52And agreed between them, they understand something to some extent. I think that is an ungrounded conversation because they haven't defined understanding. And what Kevin and I did to start out by saying, what does it mean to understand and how would you evaluate it? And then we can meaningfully speak about. And it's actually a technical paper with lots of kind of equations, but I'll give you the gist of it. In order to demonstrate understanding, first of all, you don't just understand. You understand something, a domain. You understand arithmetic. You understand human biology. It's something you understand, don't you?

42:45You understand motorcycle maintenance. Okay. Now, what does it mean to understand within that scope of that domain? So there are really two or three things that you have to kind of keep in mind. One is you've got to be competent. In other words, if I ask you, does the system or the person understand arithmetic? You're going to ask the question and get answers. If, by and large, the answers are wrong, I don't care what, that person or that system doesn't understand arithmetic. That's kind of one baseline requirement of general confidence. Let's call it the passing grade. You don't need to be 100 to be qualified in understanding arithmetic, but you can't get 15.

43:38That's one thing. The other thing is you can't be ridiculously wrong. And what's ridiculous is a matter of context here. But let's say I want to know if the system understands arithmetic. Or multiplication. and I'll give it some multiplication problems and it'll occasionally fail or maybe long ones but then I ask it to multiply 2 by 2 and it'll say 5 and I'll say are you sure? and I'll say oh yes and I'll give me an explanation of why 2 plus 2 equals 5 and maybe as some recent chatbot do try to gaslight me into believing that 2 plus 2 equals 5 That system does not understand mathematics. You've got to have general competence.

44:33You can't be ridiculous. And the third element that's really important, which is explanations. You've got to give good explanations. And that's interesting. It took us a while to understand why explanations are relevant here. And you see in high school, in school in general, in exams, a teacher will say, explain your answer. and why do they do that? You ask the student a small number of questions three or five or ten the set of questions you could have asked it's huge, not infinite but it's many millions of problems in arithmetic you could have given them so what you would like to know is that the answer they gave represented if you had answered others you want to have confidence you would have gotten good answers there also and one way is to ask enough questions that the statistics are such that it's unlikely that they would have lucked into a correct answer but the way to amplify that is to actually give an explanation but if I tell you I show me one and solve this one multiplication problem and tell me how you solve it and and I give you my multiplication procedure, and you believe that I didn't just memorize it, I actually applied it.

46:07That's a good question. Why should you believe it? But let's say you believe it. You don't need to ask me any more questions. Now you could have asked me any others, and I would have applied the same procedure, or you believe I would have. So those are some of the elements in Ghost you want me to really understand something. And when you look at language model today and take a domain that's not trivial, then you can show that they really don't understand the domain. What would it take to understand them? That's a longer discussion, but yeah. Now you asked another question and folded it, you said, will that give us, if we did understand, will that get us to awareness?

46:54And now I'll really take you really far afield, if I may. Stop me if you... When I was a professor at Stanford, I had a freshman seminar that I called Camp Confuses Think, Can They Feel? And this is the kids who haven't been corrupted yet, the smart kids. And I would start with six questions. Can a computer, let's see, I can reconstruct, can a computer think? Can they understand? Can they be creative? Can they feel? Can they have free will? And can they be conscious? There, I got to sixth.

47:43And El Mishik, you have to vote. No, it depends. and no hedging, yes or no. And at the end of the course, after we spoke about AI and machine learning and everything, they'd vote again and the answers would be different. But I think the most interesting thing that came out of it is that people realized that it's not so much that they don't know what the machines could in principle say, be aware, like you asked. If they start to question, what does it mean for people to be aware? And that's the thing that's fascinating about AI it forces us to think about ourselves, not just about machines. And I think that may be the most fun thing about AI.

48:28That's fascinating. I should have you for another episode just to talk about that. So I have more questions. But if somebody wants to try Jamba or Maestro, you say it's available everywhere. But the most direct way, I imagine, is to go to AI21. So what's the URL? So if you go to our landing page, AI21.com, all the information will be there. I just wanted to maybe clarify something. Jumba 1.6 is widely available. It's open weights. Unless you're a big company with more revenues and some threshold, then you should pay us because we paid a lot of money to develop it. So we need to recoup that money somehow.

49:31Maestro is in close preview right now. And so we're working with a small number of companies, and there's a wait list, and we encourage people to sign up. We'll gradually bring more people in, and at some point it will be an open preview, and we'll announce that. Okay. And it's AI21.com? Is that the URL? That's the landing page, and from there you'll see exactly where to go. This episode of Eye on AI is sponsored by the DFINITY Foundation. DFINITY Foundation is a Swiss not-for-profit that is home to some of the world's leading cryptographers, computer scientists, and experts in distributed computing.

50:19Their mission is to shift cloud computing toward a fully decentralized state by supporting the Internet Computer, also known as ICP. If you don't understand anything about the Internet Computer, you can see my episode with Dominic Williams. ICP's vision is that most of the world's software will be replaced by network resident software. that's an evolution of traditional smart contracts. To achieve this vision, ICP is designed to make smart contracts as powerful as traditional software, all while remaining tamper-proof, unstoppable, transparent, and verifiable. Since DFINITY's launch in 2016, they stand as Switzerland's most extensive blockchain research and development initiative and have been awarded more than 500 research grants worldwide.

51:21DFINITY remains steadfast in its mission to drive the advancement of the decentralized internet. Network resident software can now be used to run AI models and RAG infrastructure, preventing them from becoming quote-unquote hot wallets from where data can be stolen and increasing resilience. Furthermore, the technology has been designed to allow AI models to spin up and modify running web applications and internet services solo by addressing several key challenges. If you're interested in reading more about the internet computer, visit internetcomputer.org. I also encourage you to listen to my episode with Dominic Williams, in which he explains the internet computer.

From the publisher

This episode is sponsored by the DFINITY Foundation. 

DFINITY Foundation's mission is to develop and contribute technology that enables the Internet Computer (ICP) blockchain and its ecosystem, aiming to shift cloud computing into a fully decentralized state.

 

Find out more at https://internetcomputer.org/



In this episode of Eye on AI, Yoav Shoham, co-founder of AI21 Labs, shares his insights on the evolution of AI, touching on key advancements such as Jamba and Maestro. From the early days of his career to the latest developments in AI systems, Yoav offers a comprehensive look into the future of artificial intelligence.

 

Yoav opens up about his journey in AI, beginning with his academic roots in game theory and logic, followed by his entrepreneurial ventures that led to the creation of AI21 Labs. He explains the founding of AI21 Labs and the company's mission to combine traditional AI approaches with modern deep learning methods, leading to innovations like Jamba—a highly efficient hybrid AI model that’s disrupting the traditional transformer architecture.

 

He also introduces Maestro, AI21’s orchestrator that works with multiple large language models (LLMs) and AI tools to create more reliable, predictable, and efficient systems for enterprises. Yoav discusses how Maestro is tackling real-world challenges in enterprise AI, moving beyond flashy demos to practical, scalable solutions.

 

Throughout the conversation, Yoav emphasizes the limitations of current large language models (LLMs), even those with reasoning capabilities, and explains how AI systems, rather than just pure language models, are becoming the future of AI. He also delves into the philosophical side of AI, discussing whether models truly "understand" and what that means for the future of artificial intelligence.

 

Whether you’re deeply invested in AI research or curious about its applications in business, this episode is filled with valuable insights into the current and future landscape of artificial intelligence.

 

 

Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI

 

 

(00:00) Introduction: The Future of AI Systems

(02:33) Yoav’s Journey: From Academia to AI21 Labs

(05:57) The Evolution of AI: Symbolic AI and Deep Learning

(07:38) Jurassic One: AI21 Labs’ First Language Model

(10:39) Jamba: Revolutionizing AI Model Architecture

(16:11) Benchmarking AI Models: Challenges and Criticisms

(22:18) Reinforcement Learning in AI Models

(24:33) The Future of AI: Is Jamba the End of Larger Models?

(27:31) Applications of Jamba: Real-World Use Cases in Enterprise

(29:56) The Transition to Mass AI Deployment in Enterprises

(33:47) Maestro: The Orchestrator of AI Tools and Language Models

(36:03) GPT-4.5 and Reasoning Models: Are They the Future of AI?

(38:09) Yoav’s Pet Project: The Philosophical Side of AI Understanding

(41:27) The Philosophy of AI Understanding

(45:32) Explanations and Competence in AI

(48:59) Where to Access Jamba and Maestro

 

More from Eye On A.I.

All 266 episodes
#244 Yoav Shoham on Jamba Models, Maestro and The Future of Enterprise AIEye On A.I. · 52 min
Listen in VO