Orchestrating Smarter AI Systems with AI21 Labs’ Yoav Shoham | AI Basics with Google Cloud

10 Jul 2025 · 22 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: This Week in Startups - Episode on AI Orchestration

Episode Information

  • Title: Orchestrating Smarter AI Systems with AI21 Labs’ Yoav Shoham | AI Basics with Google Cloud
  • Host: Jason Calacanis
  • Guest: Yoav Shoham, Stanford Professor Emeritus and Co-founder of AI21 Labs
  • Date: [Insert Date]

Episode Overview In this episode of *AI Basics*, Jason Calacanis interviews Yoav Shoham, discussing AI21 Labs' contributions to the field of artificial intelligence, particularly their work on large language models (LLMs) such as Jurassic-2 and orchestration systems like Maestro. The conversation covers the challenges of enterprise AI, the nuances of AI terminology, and the potential for AI agents in automating tasks.

Key Topics Discussed

  1. The State of AI and AI21 Labs
  2. AI21 Labs focuses on building LLMs and tools for collaborative reasoning.
  3. Shoham discusses the current landscape of AI deployment in enterprises and the significant gap between experimentation and actual deployment due to reliability concerns.
  1. Reliability in Enterprise AI
  2. Challenges:
  3. High ratio of experiments to deployments (10:1 or 20:1).
  4. Concerns over AI reliability, particularly in critical applications (e.g., accounting, customer support).
  5. Hallucinations: Described as the tendency of AI to provide erroneous outputs, leading to distrust in enterprise settings.
  1. Orchestration and Maestro
  2. Maestro: A system designed to orchestrate multiple LLMs to improve reliability and output quality by having them validate and check each other's work.
  3. The orchestration process involves routing between LLMs, executing code, and verifying outputs systematically.
  1. The Concept of Agents and Agent Washing
  2. Agents are perceived as a solution for automating repetitive tasks but are often misrepresented in the tech landscape, leading to "agent washing."
  3. Shoham defines agents as systems capable of proactive and complex task execution, beyond simple transactional interactions.
  1. Small vs. Large Language Models (SMLs vs. LLMs)
  2. Small Models: Beneficial for specific, narrow tasks (e.g., legal or accounting) due to lower cost and latency.
  3. Large Models: Necessary for general-purpose tasks where varied consumer inputs are expected.
  4. AI21's Jamba model: A hybrid approach aimed at balancing quality with efficiency by reducing the typical costs associated with LLMs.
  1. Verticalization of Knowledge
  2. Discussion on the trend of developing domain-specific models that focus on particular fields (e.g., law, finance) to enhance accuracy and performance.
  1. Agent-to-Agent Protocols (A2A)
  2. A new development in interoperability between agents, aiming to facilitate collaboration between different AI systems.
  3. Challenges include defining shared semantics and incentives for cooperation among agents.
  1. Overhyped vs. Underhyped Areas in AI
  2. Overhyped: Some aspects of agent technology may be premature for practical application.
  3. Underhyped: The need for reliable workflows in enterprises and innovative approaches to education using AI.

Key Takeaways

  • There is a strong enthusiasm for AI in enterprises, but significant barriers remain, particularly in terms of reliability and trust.
  • Effective orchestration of multiple AI systems can potentially improve output reliability.
  • The distinction between different types of AI models and their applications is critical for success in various industries.
  • Education and workflow reliability represent promising areas for innovation in the AI space.

Conclusion The episode emphasizes the importance of understanding the nuances of AI technologies, especially for founders and innovators looking to implement AI solutions in their businesses. The conversation sheds light on the balance between excitement for AI capabilities and the practical challenges faced in deploying these technologies effectively.

Further Reading and Resources

  • [Future of AI: Perspectives for Startups Report](https://goo.gle/futureofai)
  • [AI21 Labs Website](https://www.ai21.com/)
  • [Maestro Product Overview](https://www.ai21.com/maestro/)
  • Links to relevant tools and technologies discussed in the episode.

---

This detailed summary captures the essence of the podcast episode, presenting the key insights and discussions in a structured format for easy reference and understanding.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:04All right, everybody. Welcome back to This Week in Startups. It's time again for our AI basics series. What is this? Man, founders ask us all the time the same questions over and over again. So we have done basics for legal. We've done it for marketing and growth tactics, obviously accounting. And here we are in the age of AI. People need to understand the best practices. And hey, it's moving pretty quickly. If you want to get caught up on your AI basics, go ahead and download this fantastic report by our partner, Google Cloud. It's called The Future of AI perspectives for startups featuring insights from 23 top AI experts.

0:43And today, one of them is with us, Yoav Shoham, is here. He's a Stanford professor, emeritus, as you know, and he is the co-founder of AI21 Labs, the team behind Jurassic 2 and Word Tune, building large language models and tools for collaborative reasoning. Welcome to the show, Yoav. Fun to be here. Thanks for having me. Let's talk a little bit about what we mean by reasoning, you know, here in 2025. Are these machines actually reasoning? And how do you get the best out of them? Right now, so many experiments going on in corporate America. So many people are testing and starting to deploy AI technology from large language models or with the basis of large language models.

1:29But there is some concern about the reasoning. Are these making the best decisions, hallucinations, etc.? So what is the state today? And maybe tell us a little bit about your company, AI21 Labs. Maybe I'll start with the latter. AI21 Labs, it's about seven years old. We're definitely one of the main LLM builders, although we took a slightly different tack than most people recently with our Jamba family, which is not a pure transformer architecture for efficiency reasons, and we can speak about that. most of our effort now is around orchestration and planning all these complex AI systems in particular a product we call Mindstrow which we released but if that's directly relevant to your question Jason about so you know we in AI are guilty of using terms that are so ill defined that they come back to bite us and we can draw a long list from AGI to agents to reasoning but if I step back from the actual terms, the issue, as you pointed out in the enterprise, is that there's a ton of experimentation.

2:38Like two years ago, nobody paid attention. Maybe three years ago, you couldn't get a CEO or a chief innovation officer to pay attention. Now, everybody's on top of that. But for all the hundreds of use cases in a given company that you see, the number of deployments very small and the main reason is the issue of well there are many reasons issues of compliance and safety and you know use cases new technology it's all good but the main reason is reliability and again the term hallucination may be not the best terms but these are probabilistic machines and they're for sometimes often they'll give you brilliant output but if you're brilliant 95 % of the time, not just wrong, but total garbage 5 % of the time, that may be okay in consumer land when not in the enterprise.

3:28And so there are many studies that show that the ratio of experiments in the enterprise deployments is like 10 to 1, 20 to 1. And so that is super relevant. What it means is there's an enthusiasm for the technology. But when it comes to mission-critical applications, if you're doing something in accounting, if you're doing customer support. Obviously, if you're a developer and you're pushing code to a server, it needs to be a lot more reliable. And that's why humans are in the loop. But with Maestro, I think the concept, and you'll correct me here if I'm wrong, is we have many different LLMs and having them, as crazy as it sounds, work in concert with each other.

4:11Interesting, concert Maestro. And having them check each other's work, maybe trying to get them to justify their answer, as it were, can result in better output and more reliable systems, correct? Largely, yes. Let me nuance this. Please. So someone would like to take the LLM, or as a new version of them sometimes called LRM, large reasoning models, which is really a misnomer, but system like you know 01, 03, R1 and try to get them to behave through guardrails and alignment efforts and everything and all that's good to do and we do that but that will never iron out the variance in the models and so what you really need to do is to put logic on the outside to orchestrate and it's not just a matter of routing between you know this LLM or that LLM Sometimes you run code.

5:14Sometimes you use a tool. You know, you'll access a database. You'll call, you know, a weather API, what have you. And something needs to orchestrate all that. And you're absolutely right that you want to do not only system testing, but unit testing. So every step of the way, you have an explicit plan. And every step of the way, you want to, as well as best you can, validate how well you're doing. And sometimes you'll do it with the language models, what's called a judge language model. And often you'll just, for example, you want an output to be 600 to 800 words long. Just do the damn counting.

5:49Let's talk about agents specifically. This promise has captured people's imagination because agents feel like a way for humans to stop doing chores. So much of what we do when we go to work, when we try to build a business, is the actual product and the service we're providing to people. but we all have to do our chores and cleaning stuff up, normalizing data, just work that humans find monotonous, most humans, and they don't like to do, digging dishes. The agents feel like the proper solution for those. How close are we to having agents at scale doing these repetitive tasks? How often are people actually building an agent that is in 2025 when we're recording this, getting the chore done reliably enough that humans can forget about that chore?

6:47Your mileage varies, depends on the store and depend on the quote-unquote agent. The problem is that people have been using the term agent now. It's so seductive for anything that smacks of any kind of automation. And it'll come back to bite us. I call this agent washing. And so you're absolutely right that the biggest bang for the buck is when you're trying to get the technology to take care of fairly simple, fairly mundane stuff, kind of like robotic process automation on steroids. So more and more stuff can be automated. Is that an agent or is this simply a program that you wrote that maybe is an LLM.

7:26We don't need to get anal about the definition, but typically, when we speak about agents, what do we have in mind? We have in mind a system that's not a transactional call to an LLM. There's something that's more ongoing. There's the AI system, the agent can be proactive, not just respond to our prompt. It takes, it executes complicated flows, not just like a one-step thing. It uses multiple tools. And as you do this, it gets dives here. Because if a single call to an LLN carries some uncertainty, when you start to compose them, at some point you get more noise and signal. And that's where I think maybe some people are getting a little ahead of themselves.

8:16So simple, rote, repetitive stuff, yes. More complicated stuff, we have work to do. It seems like there's been an investment in smaller, more narrow models and some debate about that. Some people believe the large models will figure it all out eventually. Other people think, hey, why not make a smaller model faster, cheaper, better, and more constrained? Maybe you could talk about the state of SMLs. I think the short answer is in consumer land, if you're looking for a very general purpose chat, let's call it a spade a spade, a chat GPT-like experience, there's probably no replacement for a very large language model that is being aligned, tuned to cover a huge variety of cases because you can't anticipate the variety of input you'll get from consumers.

9:07As you go to the enterprise and your needs are much more narrower, there's several reasons to go narrow. First of all, is just cost. It's cost not only in terms of dollars, but also latency. And so, you know, and as you know, we came up with this new family of models called Jamba. This is a hybrid state-based model and transformer that in terms of quality of the answer, you get it's competitive with the most sized model. In terms of latency and memory footprint, there's just no comparison, especially as the input, as so-called the context length increases, that kills you. The transformer architecture, this is what, in 2017, the famous paper from Google, thank you, Google, really moved the needle.

9:55Suddenly, stuff happened in language that hadn't happened before, which did happen in vision. And the reason is that the attention mechanism in Transformer allows you to, as the term suggests, tend to a very disparate part of the input. In vision, it doesn't so much matter. To know that this here is a phone doesn't really matter what the pixel way over to the side is. But in language, there's nothing local. So that made a difference. The problem is that it's expensive. It's a quadratic complexity in the input or the contact length, as we call it. Now, when we had input of a thousand, thousand squared is fine, but we're now pushing a million.

10:32A million squared is not fine. And so you need to deal with that. Part of it is smaller language models that doesn't quite deal with the contact length side of things, but then rethinking the architecture. So the state-based model, which is inherently linear and not quadratic, and mixing it with just a little bit transformer, it gives you the best of both worlds. What about verticalization of knowledge? Is there a movement to say, hey, this small language model is going to focus really on accounting to just put it in business categories? This one is just really amazing at legal concepts. And you know that when you throw this legal brief into it or on the accounting side, you throw this RFP into it.

11:18These are huge contacts windows dumping cases of case law and trying to process them. Is it just about the speed and the cost, or is it also about the accuracy? Because the model has been constrained to not have to worry about, oh, all the movies and songs ever written in the world and every blog post about those songs and music in the world. Is it actually going to result in a higher fidelity of content? Short answer, yes. I'm sorry, I'm always nuanced. But the longer answer is that there's some general common sense that the baseline model has learned that you want to retain, even if it's just mastering correct English language, grammar.

12:03And so as you attain more domain-specific language, it's okay to forget certain things, but not others. So the art here is to just remember the right stuff. Let's talk a little bit about agent-to-agent protocols. For people who don't know, A2A is a protocol for interoperability between agents. You know, technologists are always looking at what's around the corner. If you do get your agents working well, let's say the accounting department agent doing purchase orders and paying bills, eventually you might want that agent to interact between companies, maybe to put an RFP out to get five companies to bid for, I don't know, the new shed you're building or the new software that you want written.

12:51And agent-to-agent protocol is going to solve for that. But this is, we're talking within the last 60, 90 days, this stuff is all starting to be publicly released by Google, other players, and we're starting to see some consensus from different technology companies and data sources that this is the next big thing. Has anybody actually got this in deployment now, in your experience? What are the early results like? What are people thinking this will do? No, I think it's too early for anything to have been in production, even in the ideal scenario. So this is not a knock on A2A, just too early. So first of all, I think kudos to Google for shepherding this and a good start, but it's just a start and we have to realize the limitations.

13:41So the vision that there's several things that excite people when they hear about multiple agents coordinating. Part of it is the something from nothing. Oh, I don't need to think hard about the problem. I'll just build a bunch myself. We'll build a bunch of agents and then magic will happen when they come together. That historically has led to disappointment. Often the magic is and the glue is good to compartmentalize and and factor things out that's always a good but often the magic is how you put together things and what the algorithm is but i think as you said the promise is that it's not only my agent speaking to my agents it's my agent speaking to other agents in my company but that i didn't build but also outside my company and here i think not so fast.

14:35There are two fundamental problems. One is, if you look at the protocol, there's a part of what the agent communicates in JSON, which is its capabilities. And other stuff where the contract it does with other agents. The problem is it specifies the syntax, but not the semantic, not the meaning. and that historically has been the pitfall of distributed object systems that objects advertise their capabilities but there's no reason for me to for my agent to understand what you meant when you put in language i know how to find uh you know good flights well what is good flights is it mean efficient time you know so when you share semantics it can be done but it's a big undertaking it's a community kind of activity that's number one number two is shared incentives if you go outside the boundaries of even my own unit and company because we don't always share the same incentives even referring you know in the same company i should let alone so for example if i'm looking i have an agent that's trying to put together an itinerary for me and book a flight and it's speaking with your agent and you're maybe you know uh you know an agent from one of the alien companies we do not have the same incentive and so you need to put in some control for that and actually so i i spent in my you know wearing my academic hat i spent a good fraction of my academic work on the area of multi-agent systems in fact we have a standard textbook in the area and a lot of it has to do with crafting protocols for multiple agents to get them to play nice together, even though left to their own devices, they wouldn't.

16:23A lot of game theory and stuff like that. I mean, if you think about what we went through when we tried to have a semantic web exist, oh, hey, you're a chef and you have a really silly example, but you're putting your recipes online. We'd like you to make your recipes semantic. These are ingredients, these are steps, you know, and here's what the output looks like. And here's the origin, and here's the temperature. We want all this stuff to be semantic. It was like, okay, yeah, I'll do all that for you. And then, you know, a bunch of people just scrape all your recipes and you get less traffic to your website.

16:58And the promise was, oh, I would get more traffic. People would search for these three ingredients and, you know, my recipe might come up. So the devil is in the details there. And, you know, thinking it through is critical. Let's end on this. What's hyped? What's overhyped? You got a lot of founders listening here on this week in startups. They're building stuff. what's something they should be doing now that's obvious and going to pay dividends for their startup what are things that hey maybe agent to agent falls into this category you could become aware of there's an opportunity here but it might be a bit too early to get significant gains from it where should they be focusing their time and effort in your mind well present just for me to give a definitive answer because there's so many kind of degrees of freedom here but But I'd say that if you look for the maybe under-hyped opportunities, maybe I'll mention two.

17:47One is the boring stuff. You know, getting workflows to be reliable is, and again, I'm thinking enterprise. This is kind of the lens I put through. That is the biggest blocker in the enterprise right now, getting the workflows to be reliable and customizing them per deployment. that's hard work but that's where i think the real pain is the thing is it's not sexy if you give a demo that did something amazing once or maybe many times that's sexy but it's not sexy to show that things don't fail but that's where the real value is i think there's there's that's where i you know one area i would focus in yeah the area where i don't know that it's, you know, underhyped, but I think it's underserved is education.

18:38I was there in the early days of online courses, you know, Coursera and Udacity starting my corridor at Stanford. And, you know, I have an online course on game theory that's been seen by over a million people, but it doesn't begin to scratch the surface of what the real opportunity is in proactive per student teaching. The issue is not how to get chat GPT out of the classroom so people don't feed. The issue is how to get the technology in the classroom and rethink what education is really about, how we do it right with technology. So that's an area that I would really like to see kind of blossom.

19:25It is super interesting. The first step was getting all those courses online. and somebody who went to Fordham University, didn't quite hit the IVs. I was always jealous, like, what's going on, you know, at MIT, at Stanford? Like, how different is it than my experience? And I was feeling particularly under-resourced in macroeconomics, right? And I was like, I really want to understand this. And I just went to YouTube, and I found courses at Stanford, MIT, and I watched the courses, and I was like, wow, this is like alchemy. they're trying to figure out how macroeconomics works but the fact that it was available for free on youtube the same course that people were paying and had to qualify to be in the point one percent of people on the planet to get there was available for free and now you imagine if it was adaptive learning and you could answer some questions up front and then i said yeah you know you should really start with this third video or actually you're not ready for the first two videos, you should do this pre-calculus and maybe the statistics course first before you get in there, learn some basics about statistics so you can actually understand the material better.

20:38And that would be so amazing to have it be adaptive and not leave any, because as a professor, you wind up leaving some students behind. And at what fork in the road did they get disengaged is always the question, yeah? Absolutely. You know, it's typical. When a new topology comes on, And you try to use it the way you use the old technology. So television initially was televised radio. And then over time, you understood what the media was really good for. I think the time is ripe now with the flat world. Everybody has access to computers and networking to get AI to rethink education. Yeah, absolutely.

21:18And you think about the role of professors, creating courses, creating quizzes even. in. You can go into any LLM today, ask it to take Great Expectations, Charles Dickens, and give you a series of Q &A and be your coach. And you just give it that tiny prompt. And you could sit there with your phone and have a personalized tutor that you would have had to spend, you know, whatever days or weeks to find them and pay hundreds of dollars. And it's just available to everybody for free today. What an amazing discussion. Thanks so much, Yoav, for joining us. Thank you. It's really, really fun. Thanks for having me.

21:52You can learn more in Google Cloud's report, The Future of AI Perspectives for Startups. Go to goo.gle slash futureofai. That's goo.gle slash futureofai. Go check out Gemini. Man, I love that deep research. And everybody, if you want to get more of our Startup Basic series from legal to accounting to marketing and now AI, go to thisweekinstartups.com slash basics. Thanks again for listening. We'll see you next time. You

From the publisher

In this episode of AI Basics, Jason sits down with Yoav Shoham — Stanford professor emeritus and co-founder of AI21 Labs, creators of Jurassic-2, Wordtune, and the new orchestration system Maestro.

They unpack:

  • Why enterprise AI struggles with reliability
  • What orchestration really means (and why LLMs alone aren't enough)
  • The pitfalls of “agent-washing”
  • Small vs large models, agent-to-agent protocols, and where real opportunities lie

This one is for founders building with AI — if you're navigating hallucinations, chasing automation, or exploring multi-agent workflows, this episode is a must.

*

Timestamps:

(0:00) Yoav Shoham joins Jason to discuss AI Basics.

(1:05) What AI21 Labs is building — from Jurassic-2 to Maestro

(5:49) The overuse and confusion of the seductive term “agent”

(8:22) Small models vs large models: what's best for enterprise?

(10:53) The verticalization of AI: legal, accounting, and beyond

(12:17) The challenge of agent-to-agent communication and shared semantics

(17:08) What’s overhyped vs underhyped in AI — Yoav’s advice for founders

*

Uncover more valuable insights from AI leaders in Google Cloud's 'Future of AI: Perspectives for Startups' report.

https://goo.gle/futureofai

*

Explore Further

Google Cloud’s Report: The Future of AI

Get insights from 23 leading experts on how startups can leverage AI for real business impact.

👉 Read the full report

More AI Basics Episodes

Founders, operators, and builders — catch up on our full Startup Basics series (Legal, Finance, Growth & AI):

👉 thisweekinstartups.com/basics

Links from episde:

  • Google Cloud — Build smarter with AI tools built for startups
  • AI21 Labs — Makers of Wordtune, Jurassic-2, and Maestro
  • Maestro — Orchestrate AI agents and tools across workflows
  • Jamba — AI21’s hybrid language model with breakthrough efficiency


Learn About Google’s A2A (Agent-to-Agent) Protocol

The next evolution in agent interoperability and coordination:

  • *
  • Follow Yoav:

    X: https://x.com/yshoham

    LinkedIn: https://www.linkedin.com/in/yoavshoham/

    *

    Follow Jason:

    X: https://twitter.com/Jason

    LinkedIn: https://www.linkedin.com/in/jasoncalacanis

    *

    Follow TWiST:

    Twitter: https://twitter.com/TWiStartups

    YouTube: https://www.youtube.com/thisweekin

    Instagram: https://www.instagram.com/thisweekinstartups

    TikTok: https://www.tiktok.com/@thisweekinstartups

    Substack: https://twistartups.substack.com

    More from This Week in Startups

    All 653 episodes
    Orchestrating Smarter AI Systems with AI21 Labs’ Yoav ShohamThis Week in Startups · 22 min
    Listen in VO