#216 Nenshad Bardoliwalla: Inside Vertex AI - The Ultimate AI Toolkit by Google Cloud

30 Oct 2024 · 52 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Episode Notes

Episode Title

#216 Nenshad Bardoliwalla: Inside Vertex AI - The Ultimate AI Toolkit by Google Cloud

Host

  • Craig S. Smith - Longtime New York Times correspondent.

Guest

  • Nenshad Bardoliwalla - Director of Product Management for Vertex AI at Google Cloud.

---

Episode Overview In this episode, Craig Smith speaks with Nenshad Bardoliwalla about Google Cloud's Vertex AI, an advanced AI platform. They discuss the components of Vertex AI, particularly focusing on its three core layers: Model Garden, Model Builder, and Agent Builder, as well as the implications for enterprises using AI.

Episode Segments

  1. Introduction to Nenshad Bardoliwalla & Vertex AI (00:00)
  2. Overview of Vertex AI's Three Core Layers (01:52)
  3. Nenshad's Journey to Google Cloud (05:35)
  4. Choosing the Right AI Model (06:36)
  5. Google’s AI Infrastructure & Tensor Processing Units (TPUs) (08:00)
  6. Model Builder: Fine-Tuning & Prompt Optimization (10:15)
  7. Agent Builder: Building AI Agents with Tools & Planning (12:11)
  8. Model Evaluation & Prompt Management (17:57)
  9. Generative AI for Business Analysts (21:23)
  10. AI Model Modality & Use Case Selection (23:24)
  11. Popularity Distribution of AI Models (25:23)
  12. Prompt Optimization Tools (28:18)
  13. Building AI Agents: Real-World Use Cases & Ethical Safeguards (34:20)
  14. The Capabilities & Limitations of AI Agents (40:13)
  15. TPU vs. GPU (45:48)
  16. Future of AI at Google Cloud (50:33)

---

Key Concepts

Vertex AI Components

  • Model Garden:
  • A curated platform for accessing and evaluating various AI models, including Google's first-party models, open-source models, and third-party models.
  • Supports models like Gemini, Imagine, Claude, and others.
  • Allows users to compare models based on performance, cost, and regulatory compliance.
  • Model Builder:
  • Enables users to fine-tune models, manage prompts, and ensure that models align with brand values.
  • Includes features for monitoring model usage and safety.
  • Agent Builder:
  • Facilitates the development of AI agents capable of performing complex tasks.
  • Integrates tools like Google Search for real-time data retrieval.

Model Evaluation and Selection

  • Enterprises should evaluate models based on:
  • Use case requirements (e.g., summarization vs. content generation).
  • Performance metrics (accuracy, response time).
  • Cost considerations.
  • Nenshad emphasizes the importance of using evaluation datasets to compare models realistically.

Prompt Optimization

  • A new technology that optimizes prompts for different models to ensure consistent outputs and facilitate easier transitions between models.

AI Agents

  • Agents are described as goal-driven systems that utilize various tools to accomplish tasks.
  • There is a need for human oversight and checkpoints to avoid fully autonomous actions that could lead to ethical issues.

Ethical Considerations

  • The necessity for human oversight in AI design to prevent misuse and ensure safety.
  • The importance of setting clear boundaries and guardrails for the capabilities of AI agents.

Hardware Considerations

  • Tensor Processing Units (TPUs) are highlighted as superior hardware for AI tasks, providing significant performance advantages over traditional GPUs.
  • Google Cloud's infrastructure is tailored specifically for AI workloads, making it a preferred option for many model providers.

---

Key Takeaways

  • Vertex AI offers a comprehensive toolkit for AI development and deployment, catering to various user needs from model selection to real-time data integration.
  • The landscape of AI model usage heavily favors a small number of popular models, while many exist on the fringes with niche applications.
  • As AI technology evolves, it's critical to maintain a balance between automation capabilities and human involvement to ensure ethical and safe deployments.
  • Continuous advancements in hardware, particularly with TPUs, are pivotal in scaling AI solutions efficiently.

---

Conclusion This episode of Eye on A.I. provides valuable insights into the workings of Google Cloud's Vertex AI and the broader implications of AI technologies in various sectors. With the rapid evolution of AI, understanding these tools and ethical considerations is essential for practitioners and businesses alike.

---

Additional Resources

  • JLL AI Solutions: Learn more about AI in real estate at [jll.com/AI](https://www.us.jll.com/en/solutions/ai/?utm_source=eye-on-ai&utm_medium=pdcst&utm_campaign=am-us-bnd-cp-ai-102).
  • Follow Craig S. Smith on Twitter: [@craigss](https://twitter.com/craigss)
  • Follow Eye on A.I. on Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)

For more episodes, don't forget to like, subscribe, and turn on notifications!

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00One of the remarkable changes that has happened in the industry when it comes to generative AI is that software developers and even business people can interact with these models. so easily. They don't necessarily have to train anything to start getting results very quickly. And so a lot of the experimentation to see whether a model is good for my use case can actually be done by software engineers as well as business analysts and the like, because all they're doing is typing their prompt into a box and seeing, well, does this kind of look like what I was hoping for. Imagine a world where our spaces work smarter, not harder.

0:44Where technology and humans collaborate to create something extraordinary. At JLL, they're not just imagining this future, they're building it. JLL's AI solutions are transforming the real estate landscape, accelerating growth, streamlining operations, and unlocking hidden value in properties and portfolios. From predictive analytics to intelligent automation, they're creating smarter buildings, more efficient workplaces, and sustainable cities. JLL, shaping the future of real estate with AI. Welcome to a brighter way. To learn more about JLL and AI, visit jll.com slash AI. Again, to learn more about how JLL is shaping the future of real estate with AI, visit jll.com slash AI.

1:51So Vertex is a brand that was launched in 2021. And it represents Google Cloud's AI platform. So, you know, a combination of core machine learning services like training and inference and notebooks along with pre-built models including large large foundation models as well as more recently capabilities like enterprise search so the vertex portfolio has definitely grown over the last few years okay uh and so you'll talk about what Vertex is and what it does. And does it support different models beyond Google models? Absolutely. So we have a property at Vertex that we call our model garden. And the entire notion behind the model garden was that we would support not only Google's leading first-party models, so Gemini, Imagine, and the like.

3:04But we would also have a very strong representation for open-source models and, most interestingly, third-party models. So we are the only platform that provides that combination of our own models, open-source, and third-party models, and we allow all of them to compete freely on the platform. Yeah, so Claude and GPT-401, they're all there. So the anthropic models from Claude are there. Mistral's models are there. Meta's llama models are on the model garden. Today, we do not have a partnership with OpenAI. So those models you have to get from somewhere else. Yeah. Okay. So we'll start in a minute.

4:01Just one thing that I'm really interested in, just in the last month or so, and I've interviewed Srebris a few times from when they first came out and then with their second wafer engine. And then I ran across Andrew Feldman at a conference, and I interviewed him again. And then at that same conference, I interviewed a guy, Rodrigo Leung from SambaNova. and I'm really curious about beyond the models that you host, is Google using these other inference chips? Because it sounds to me, I mean, probably you're using Crock, but it sounds to me like there's a lot happening on the inference side. But I'll leave that to this.

5:08We've actually invented our own chips in that area. Okay. I'm happy to talk about that. Okay. So why don't we start, Nanshah, by having you introduce yourself. Tell us how you got to Google Cloud, what your responsibilities are. You can talk about what the Vertex platform is, and then I'll jump in with questions. Great. Well, thank you very much, Craig. Hi, I'm Nenshad Bardoliwala, Director of Product Management for Vertex AI. I've been at Google for almost two years, and I am responsible for the foundational platform components of Vertex, which include being able to train and tune models, to be able to experiment with models and notebooks to evaluate those models, as well as being able to do inference with those models, monitor those models, and steer them through their lifecycle in production.

6:16So that entire gamut of capabilities, it's probably the largest part of Vertex AI that's carved out as a separate area is my responsibility. I came to Google. Oh, please go ahead. No, no, no. Go ahead. I'll cut that out. No worries. I came to Google after a 20 plus year in enterprise software and data specifically. I have only always had three intellectual interests in my life Computers I was a hacker since I was a very young kid I'm talking Commodore 64 era So I'm dating myself The human mind And I'm also a very avid and passionate musician As you might be able to tell over my shoulder There are a number of noise making devices For my guitars So my entire career has actually been focused on data analytics and AI.

7:2710 years at large companies like Siebel and SAP. And then 12 years in the startup world, which, you know, a number of that time was spent at one of my companies called Paxata, which pioneered a self-service data preparation platform which we then sold to DataRobot which is a unicorn AI company and that really set me up for the privilege of being able to be part of the Vertex team here at Google. Yeah, so tell us about Vertex. What exactly is Vertex? So Vertex is a lot of things. So let's break it down into a few different layers, right? The first layer of Vertex is what we call our model garden.

8:18So we have a vision that customers should be able to choose any model that they think will be helpful for them to solve their use case, that meets their regulatory requirements, that meets their latency and performance requirements, and so on. And as you know very well, Craig, there is a massive number of models that are available in the market now. And so the Model Garden is our curated place in Vertex AI where customers can get access to state-of-the-art Google models. So, for example, Gemini 1.5002, which just came out last week, is available inside the Model Garden. Our Imagine 3 model, which is our text to image model, is available in the Model Garden.

9:10And a number of other Google-specific models are available there. But we also believed from the outset that choice is extremely important to our customers. And we show that openness by offering in the Model Garden a very rich collection of open source models. So you can find favorites like Stable Diffusion or Meta's Llama 3.2, which just also came out last week. And, you know, over 100 other open source models, including Google open source models like Gemma 2, which has done very well in the market. And then we combine that with our third party model support. And so we offer models from Anthropic, like the Claude family.

10:01We offer models from Mistral, like Mistral Large and their CodeStral model. We offer AI21's Jamba models, and there are more coming seemingly each week. So the base layer of Vertex's AI platform is the model garden and the ability for customers to choose, and this is what's really interesting, Craig, competing models. So if you look on the leaderboards of people who do comparisons of models, you will find Gemini and Llama and Anthropic all competing on the same leaderboard. And Vertex AI platform, our mentality is that customers should choose the best model for them, whether it's a Google model or not.

10:47So foundation layer one is the model garden. Layer two is what we call our model builder layer. So this is for being able to take a model and being able to fine tune it. So adding your own data to make the model respond in a way that makes sense for you. This is the layer that allows you to do a prompt management, prompt versioning, prompt optimization, prompt revisioning, and so forth. This is the layer that allows you to deploy those models with the fine-tuning capability, and it includes capabilities to monitor the recitation or the usage of specific code or word fragments from third-party sources to monitor the safety so that we ensure people are using the models in a way that aligns with the models.

11:50with their brand values and a host of other capabilities in that regard. So the model builder works hand in hand with the model garden to allow customers to interact with the models, tune them for their needs, and then ultimately make them available for their usage. Then the third layer is the last part of Vertex that we'll talk about today is what we call the agent builder. So you've probably heard, Craig, you know, agents are kind of the rage right now. Everybody seems to be talking about them. And at Google, we have some very unique assets that make agent building very powerful. we do use some open source technology in fact lang chain which is an incredibly popular library for doing agent building is offered as a managed vertex service but we also combine that with a really rich set of componentry from our enterprise search because if you think of retrieval augmented generation or rag the retrieval part is search and so we offer very powerful search capability We also offer the large language models that allow for the generation of the results.

13:10And uniquely to Google, we even offer the ability to ground the results of the agents using Google Web Search. So you can actually get very accurate up to the minute results from Google Search as part of your agent, which of course we are uniquely positioned to add. so three layers of what we call vertex ai today the model garden the model builder and the agent builder can you talk about uh yeah uh on the a couple of things so there you say over a hundred models and uh i was looking at the hugging face leaderboard the other day um how how did these models differentiate themselves how do you decide when you're in the model garden which model you should use is it like shopping in the grocery store and you know the brand that does the most advertising that's most familiar to consumers gets most the traffic or is it is there true differentiation among models there is so first first a note about something you said that's a good prop and reminder to me is that actually model garden is also directly integrated with Hugging Face.

14:46So in fact, even though Model Garden offers, you know, 150, 160 curated models, we also have a one-click deploy, literally one click from Hugging Face directly into Vertex Model Garden. So you can get access to thousands of other models that way. But let me answer your question about, you know, which model should I use? Because it's one that I get multiple times a day you have to understand a few few uh you have to ask ask yourself a few different questions about about what you're trying to do number one what is the use case what am i actually trying to do am i trying to summarize content am i trying to generate new types of content for example you know marketing a marketing brochure um what uh what regions and what regionalization do i do i need do Do I need to generate this content in Hindi or just in English, for example?

15:42I need to understand how quickly I need the response. Some models take much longer. I mean, there are physics involved here. The larger the model, the slower it's going to be to respond. So you want to be able to choose a model that is the right size for the latency and accuracy that you're looking for in the use case. What is the model good at? Some models are much better at doing reasoning. Other models are much better at being able to invoke different tools in agentic-type workflows. Other models are really strong at certain text classification-type tasks. So you have to know what you're trying to do when you're evaluating the model.

16:31And then there is, of course, cost, right? different models cost different amounts. We at Google have been particularly aggressive. Our new 1.5 models are priced at half the cost of some of the competing models in this space. And that obviously makes it very attractive for customers. So you have probably about 10 dimensions that you need to look at when you're evaluating what is the right model for me. And the goal of the model garden is to surface a lot of that context to make it easier for you to winnow down, like, what are the three models I want to try for this use case? Because there usually will be three.

17:17And then the next thing you have to do, if I may, Craig, no, no problem. The next thing you have to do is actually evaluate those models, right? So you have to evaluate the models, not just with the leaderboard information that you get from public sites like LMSIS, which are very useful, but generic benchmarks. But that doesn't tell me whether for my enterprise use case in my organization with my tuned data, whether the model that I'm considering will actually do a good job. So you have to build your evaluation data sets and use tools. For example, Vertex has a generative AI evaluation tool that allows you to actually compare the responses of multiple different models to see how close they are to what you really want to get.

18:10And what's really cool is that you can actually, you can have human rated information as part of that analysis. You can even have models serve, I mean, you can have AI serve as the evaluator and actually have that tell you how well it thinks the system is doing for response. So there are systems as part of Vertex and, frankly, other vendors' platforms that allow customers to take all of these factors into account, run their evaluations, and then ultimately say, okay, the most accurate model was Model A, but it also costs a little bit more. is that additional cost worth the accuracy? The customer can then decide.

18:57And if they say yes, they can go with model A and deploy that in production, or they can choose model B, which costs a little bit less, but maybe is not as accurate, but that might be totally okay. So this is the logic that we help customers with and the tools that we provide customers with in the process. Yeah. Who is doing the choosing and all of this evaluation within an enterprise? Is it just whatever software engineer is working on whatever project, or is this becoming a specialized job of understanding, keeping up with what models are doing and which ones do what? So there are a couple of interesting trends to reflect on towards this question.

19:56One is that in previous eras of machine learning with predictive ML, right, building regression models classification models and the like um that was very much a data scientist type task and model evaluation is very well understood in the data science world for for predictive predictive models so you do find a lot of data scientists who have the skill set of doing model evaluation, now performing that function for organizations for generative AI models too. That being said, one of the remarkable changes that has happened in the industry when it comes to generative AI is that software developers and even business people are able to interact with these models so easily.

20:53They don't necessarily have to train anything to start getting results very quickly. And so a lot of the experimentation to see whether a model is good for my use case can actually be done by software engineers as well as business analysts and the like, because all they're doing is typing their prompt into a box and seeing, well, does this kind of look like what I was hoping for? I can tweak that prompt, et cetera. So typically we find that people do a lot of experimentation who are not the data scientists. They kind of winnow it down to a couple of models. And then if they're happy with the results, they can just move forward.

21:36Otherwise, they can turn to their data scientists and say, I think these three models are kind of the right ones. Could you please run a more formal systematic evaluation? And we encourage customers to do that because literally every day, Craig, a new model comes out. It changes all the leaderboards overnight. And so you need a systematic, programmatic methodology for doing this in a repeatable fashion. Yeah. And you were talking about AI tools to help identify models for particular use cases. When someone goes on the platform, are they clicking on drop-downs that give options? And, you know, like Airbnb, I need two bedrooms.

22:30I need it in this price range. i needed or is uh does the do you state your use case uh to a conversational agent and it does it comes up and says these are the three best models you should try them out so um i think we're definitely headed towards that agentic world of helping customers to choose models. Today, what customers do is they get a series of choices that can make like what modality are you trying to interact with, right? That's another thing that you have to look at for models, right? So some models only handle text in and text out. Some models only handle code in and code out, like programming code.

23:18Some models only handle text in and images out, right? So you want to allow people to choose what modality they're going to be working with and also what task that you want the model to do for example text classification is very common as a use case and some models are much better at it than others so we try to make it very easy for customers using a series of filters and user experience affordances for them to be able to narrow down the set of models that meet their criteria. And then for the individual models, we provide what we call model cards. And those model cards give them really detailed information about like, where was this model trained?

24:08What kind of data was it trained on? What use cases is it really good for? Here are some examples of prompts that work really well to show the power of these models. And so I think of it as a sort of progressive choose-your-own-journey in that you start with the whole field of every model that's possibly available, but then very quickly you start narrowing down the decision space to be able to determine which set of models based on your criteria is actually going to be helpful for your use case. Yeah. And as you said, new models appear every day. Is the market big enough? Is the demand large enough that all of these models are getting used?

24:57Or is it sort of like you've got these beautiful rose bushes in the middle of the garden and then these straggling sunflowers off to the side that no one's really paying attention to? So our observation is that it seems to follow the market distribution seems to follow the 80-20 rule. 80 % of the inference that customers are doing are on a, let's call them less than half a dozen, you know, or half a dozen or so models, less than a dozen or, you know, between half a dozen and a dozen models that are very popular, that get tried very quickly, that have lots of tooling built around them, etc. etc.

25:52And those tend to be the dominant models. And whereas for many of the other models, they're just much more in many cases, they're much more niche. For example, they're only designed or trained on certain types of medical information, or they're only, you know, trained on certain financial information, still extremely valuable, but they're not broad, you know, broad purpose, right? They're only meant for a specific industry, and in some cases, a specific task so you will definitely find that the majority of the action in the models that we observe is you know again between half a dozen and and a dozen models that are very very regularly tried and used and put in production by customers and the rest are are just less so and when you say less so is it uh i'm just trying to get a sense of scale because it's it's the entire world that's uh that's using these models is it um you know there will be a model with like two or three people using it or enterprises using it is it does it get to that point or do most models have of at least a couple of hundred users?

27:16I mean, I have no idea what the scale is. There's definitely a long tail. We absolutely, you know, in our model garden today, you will see, you know, models that are used by millions of people daily. And you will see models that are used once a day, right? You will see models that are used by every customer that I could possibly name in Google Cloud and other models that because they're much more specific, they're only used by a handful of entities. So there is definitely a broad distribution between massive usage and then I'd say a very long tail, you know, that the curve kind of goes down like this and, you know, asymptotically starts to approach zero.

28:05once you get out of those top 6 to 12 models. Yeah. You mentioned a prompt framework. What was that? And can you talk about that? Sure. So our customers have told us that they struggle from a couple of pretty hard challenges. One is that they don't see consistency when they try the same prompts between models from one provider to another. For example, I may be a customer using something besides Gemini today. And then I see that Gemini is now 50 % the cost of the model I'm using now, which is significant savings. And so I would love it if I could literally just pull the plug from model one and then just plug into model two.

29:03But it doesn't quite work like that because the way every model responds to its individual prompts is different. So we have just introduced some new technology, our prompt optimizer, that can actually take your prompts and the outputs from model A and then feed it into model B into the optimizer and get very similar results to what you were getting with model A. So use case number one is this really supports migration between different models so that customers can have that choice and flexibility. But that also holds true for the same models in the same family. For example, when we come out with Gemini 1.5001, and then now we just came out last week with Gemini 1.5002, there will be differences in the way the models respond to prompts.

30:06And you can use the prompt optimizer to help in that use case, too, with the comparison being between 1.5001 and 1.5002. But you can see that this programmatic way to sculpt prompts and the system instructions to help me get the results I'm looking for is very valuable for customers in a world where the models and the cost structures are constantly changing. Yeah. In that prompt optimization, if you're moving from one model to another, so you take the prompt that's been working in one model, you go through the optimization into another model. Do you then work with the prompt that has been output by the optimization engine and build on that?

30:59Or do you keep working through the optimization engine because you're not really familiar with how prompting works in the new model? Yeah, so the way it works is actually really fascinating. What actually changes is not the prompt itself, but the system instructions that are fed to the second model. So what we have found, and this is one of the beauties and privileges of working at Google, is we get to work with the DeepMind research team arm in arm on a daily basis. And our cloud AI team also has a very strong research board. So as they find interesting customer problems and develop technology, we can really rapidly bring it out to market.

31:49And so in this case, prompt optimization was built by our cloud AI research team, although we have many other examples, for example, the Gemini models themselves that come from our collaboration with DeepMind. And this technology allows us to generate a system prompt that continues to learn based on the input prompts that we're giving it and the outputs that we're looking for. And so it kind of iterates the system prompt until it gets a set of instructions that allows Model B to emulate the behavior of Model A. it's pretty fascinating but the prompts themselves stay the same yeah uh and and do you also have like prompt libraries if someone's trying to do something fairly complicated that someone has already figured out how to do where you can does that exist it does so uh remember i told you there are three layers of vertex ai platform the model garden the model builder and the agent builder in the model builder layer we have a service called vertex ai studio and in vertex ai studio we have a number of facilities for allowing customers to save their prompts version their prompts so that they can you know reuse them and share them with other people and we also provide a number, probably 50 or 60 examples out of the box, depending on what type of use case, what type of modality, to show you different ways to get value quickly out of the models.

33:36So between the prebuilt library, which is static, and then within your organization, not static, but it evolves over time as we add more examples. And then your organization, where people are constantly adding new prompts and new examples, we have this really rich set of data that we can use to provide people with examples for how to prompt in their organization.

34:03Now, I'm going to skip to the agentic layer. Sure. How does that work? and what's the extent of activities that you can build an agent for at this point? Great question. So I'll start by saying that agents are still a nascent technology, although we do have customers who are actually deploying agents today very successfully for the use cases that they've designed it for, right? As an example, ADT, the home alarm company, you probably know them, you may even have an ADT device in your house. They've used Vertex AI to build a customer agent that allows the customers, millions of their customers, to select, order, and set up their home security.

35:01So, pretty sophisticated flow, if you ask me. what agents allow you to do is set a goal right the agent has a goal or a purpose one example i love to use is help me plan my trip to southern california right i need i need a travel a travel agent agent which i know is silly but it helps me to remember it but an agent to help me plan my travel so i need i need a goal right which is help me plan this trip i need a set of tools that the agent can call to help me in that journey. So who would I want to call if I were going to Southern California? Who are the places that I'd probably call? I have, you know, reasonably young kids, so I'm probably going to call Disney.

35:51I'd love to have a tool or extension to their reservation system. Probably want to call SeaWorld. I probably want a tool that works with Google Maps so I can chart out the distances and look at traffic information and the like. I have a tool that, you know, potentially could even allow me to enter information about, you know, my credit card or other ways that I want to pay. So the agent has to have access to multiple different tools in order to do its work, right? And then the third part of the agent is it has to actually be able to plan and reason about how it is going to go solve this problem, right?

36:33And when you think about it, if you say, I want to book a trip, which is sort of the highest level node, okay, where are you going? How many days? Monday, Tuesday, Wednesday, Thursday, Friday. Okay, so now I have the next layer of the dates that I'm going to go. Well, what do I want to do on each one of those dates? And each one of those dates, I'm going to invoke different tools. I'm going to have different constraints, like what's my budget? How far can I travel on a daily basis? And so forth. And so when you combine those three things, a goal-driven system that has a number of tools available to it and can plan and hierarchically decompose the problem to get you to that goal, you have an extremely powerful set of technologies that I think we're just scratching the surface of today.

37:22Yeah. You know, it's funny. I saw Yuval Hariri on Bill Maher's show the other day. and he told a story which i'd heard a different variation of uh of uh sort of trying to emphasize how dangerous or powerful these models are and he was set the The AI, he said, was set a task, and it ran into a roadblock because it needed to solve a CAPTCHA, you know. and the story goes i've got to track down this story that the ai went on upwork or one of these freelance platforms and hired somebody to solve the captcha for it and uh and the excuse the ai gave the the person on the other end is that well i'm uh i'm blind so i can't see the captcha i mean the story is to me is is uh this i've been hearing that story for a couple of years now before there were agents and i i just i just think it's bs and i wish i had been on bill marr's show because i would have called hariri on it but but when when an eight what are the modalities that an agent can use at this stage on your platform to interact with tools can it make phone calls using a robotic voice can it Can it solve a CAPTCHA or can it, you know, the world is a pretty complicated place and the digital world is even more complicated.

Read the full transcript

39:49Yeah. So what's the extent of the capabilities of these agents that are being solved? i understand agents like you know you say to alexa you know sent my thermometer at 68 degrees yeah to me that's very straightforward i understand how that works but anyway what's what's your answer so i think i think there's a maybe a broader point that needs to be addressed which is um and And then I'm happy to answer your direct question. The first thing I would say is, we have to be very smart about how we design these systems. And we have to put in not only guardrails, right, for what we want our agents to do and not do.

40:38But we also need to put touch points with humans in place. so that we make sure that things don't get automated beyond our level of comfort. So going back to that travel example, if an agent could do all of this for me and come back to me with an itinerary, that's awesome. But I absolutely want to approve that itinerary. I absolutely want to say, actually, I don't want to go through Santa Monica on Wednesday. I've already seen Santa Monica, but I didn't tell the agent that. that's why i came up with the plan like the you want to set up very clear touch points where the agent can interact with a human and depending on what you're trying to automate uh you you will increase that or or dial it down right there are more innocuous tasks right like alexa alexa and your uh thermometer right very simple very deterministic please do this it goes does this Once the agent starts to be able to plan and be able to invoke tools, like go generate these images for me or go call this location.

41:53And we have examples of these technologies at Google, which are pretty incredible. And the voices don't sound robotic. They sound like other people.

42:07So my general point, I think, is that I think technologically, we have the ability to do many, many things today with agents, with speech, with voice, with image, with video generation. You could have an agent that could actually create a video because the agent generates a prompt, it submits. So the possibilities are limitless. But I think the question that we always have to ask ourselves is, are humans involved enough in the process at the right points for what this agent that I'm currently designing is trying to do? So I can't attest to the veracity of the story that you shared with me. I have no idea whether it's true or not.

42:58But I would say that for people who are in this space actively designing these systems, we would have put many safeguards. We would add a combination of not only the language model being able to plan, but also rules in a deterministic way, like thou shalt not do these things. Yeah, no, and I understand that. My question is, today, could an agent do that? Could it run into a problem where it needs to solve a CAPTCHA? Could it, on its own, without a human saying, oh, it needs Upwork, so here I'll log into Upwork for it. Just on its own, could it go through all that steps, send the messages, hire an Upwork?

43:51I just don't believe agents are that capable yet. I would say that that strikes me as beyond what I think I have seen agentic technology do. Because the process you just described of hiring somebody, you know, identifying what the problem is, then, first of all, knowing that there is a service like, was it Upwork? Do you have a picture of a nerd? Knowing that that service exists, knowing that the purpose of that service is to get access to human beings, being able to convince somebody by the model lying that it is a human being who is visually impaired, this seems pretty fanciful to me. Yeah, yeah.

44:43That sort of thing really bothers me. I can understand why. ev and then everyone's like oh my god um okay so back to uh vertex uh we were going to talk about the uh the hardware that you have in your data that people have access to because i've been talking to you know cerebris and sambanova and and these guys about these new chip architectures that are built for inference are 10 times or more faster than traditional gpus obviously open ai and azure is married to to gpus but increasingly it seems there should be a market for this uh faster inference uh there most certainly is in fact uh one of the i think crowning achievements of google's infrastructure, which was born out of a very real practical need.

46:02If you look at our consumer properties, if you look at YouTube, if you look at ads, if you look at search, if you look at Gmail, if you look at Google Maps, we use AI extensively. We have for more than a decade across all those products. And we determine that the cost of inference, if we were to rely on the available technologies at the time we would not be able to scale offering ai to the world with the existing technologies that existed and so a pioneering group of individuals at google more than a decade ago created the tensor processing unit or tpu and tpus are uh if i might be so bold the precursor of many of the innovations that you are seeing come to the market right now.

46:56And quite a few of those companies are started by people who worked on TPUs previously. That's right. That's right. So we pioneered machine learning specific computation down to the silicon over a decade ago. We are in our sixth generation of that technology that we've announced publicly. It powers that those TPUs power Gemini. They power all the consumer sites that I mentioned before. They power Imagine. And I will also tell you that some of the world's largest model provider companies insist on coming to Google Cloud because that is the only place to get TPUs. And when you have to do training at the scale of building your own large language model or other foundation model, the economic benefits of a hyper-optimized infrastructure, which not only include the TPUs themselves, but the way they're networked, the way they're laid out in the data center, etc.

48:06We have a significant head start on the rest of the market because we recognized this opportunity a long time ago. And it's great to see new innovation in the space, by the way. There are some really exciting companies out there. But TPUs have been available in Google Cloud for quite some time. And they are an exceedingly popular offering because people who have used them know how absolutely powerful they are. Yeah. And I understand the lock that NVIDIA has had because of CUDA and, you know, in building applications for a certain chipset. But inference, it's irrelevant. You're just hitting a model and getting the answer back.

49:03So on inference, how much faster are TPUs than GPUs? Do you have a metric? Yeah, there isn't a public metric that I could share with you. And I also want to be very clear that NVIDIA is one of our largest and most strategic partners. We have made the decision, just go back to what I told you about Model Garden. We are willing for innovation to flourish in our company, even if there's competition. And so we view it and our customers view it as a great thing that we offer the absolute state of the art NVIDIA GPUs in our platform. We partner very closely with NVIDIA. They have specific requirements around data center, networking configuration, et cetera.

50:03And we work with them to do that. So it's not really a head to head comparison that we do. It's much more about fitting the use case to when the customer thinks that TPU will do a better job for what they're trying to do than a GPU. But the beauty is they have choice. Right. And you do offer TPUs in the cloud. Yes. Absolutely. Imagine a world where our spaces work smarter, not harder. Where technology and humans collaborate to create something extraordinary. At JLL, they're not just imagining this future, they're building it. JLL's AI solutions are transforming the real estate landscape, accelerating growth, streamlining operations, and unlocking hidden value in properties and portfolios.

51:06From predictive analytics to intelligent automation, they're creating smarter buildings, more efficient workplaces, and sustainable cities. JLL, shaping the future of real estate with AI. Welcome to a brighter way. To learn more about JLL and AI, visit jll.com slash AI. Again, to learn more about how JLL is shaping the future of real estate with AI, visit jll.com slash AI.

From the publisher

This episode of the Eye on AI podcast is sponsored by JLL.

JLL's AI solutions are transforming the real estate landscape, accelerating growth, streamlining operations and unlocking hidden value in properties and portfolios. From predictive analytics to intelligent automation, JLL is creating smarter buildings, more efficient workplaces and sustainable cities.

To learn more about JLL and AI, visit: jll.com/AI

 

 

In this episode of the *Eye on AI* podcast, we explore the world of AI at Google Cloud with Nenshad Bardoliwalla, Director of Product Management for Vertex AI.

 

Nenshad unpacks the three core layers of Vertex AI: the Model Garden, where users can access and evaluate a diverse range of models; the Model Builder, which supports model fine-tuning and prompt optimization; and the Agent Builder, designed to develop AI agents that can perform complex, goal-oriented tasks.

 

He shares insights into model evaluation strategies, the role of Google’s Tensor Processing Units (TPUs) in scaling AI infrastructure, and how enterprises can choose the right models based on performance, cost, and regulatory requirements.

 

Nenshad also delves into the challenges and opportunities of AI prompt optimization, highlighting Google’s approach to ensuring consistent outputs across different models. He discusses the ethical considerations in AI design, emphasizing the need for human oversight and clear guardrails to maintain safety.

 

Whether you’re in AI, tech, or curious about AI's potential impact, this episode is packed with insights on next-gen AI deployment.

 

Don’t forget to like, subscribe, and turn on notifications for more episodes!

 

 

Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI



(00:00) Introduction to Nenshad Bardoliwalla & Vertex AI

(01:52) Overview of Vertex AI's Three Core Layers

(05:35) Nenshad's Journey to Google Cloud

(06:36) Choosing the Right AI Model

(08:00) Google’s AI Infrastructure & Tensor Processing Units (TPUs)

(10:15) Model Builder: Fine-Tuning & Prompt Optimization

(12:11) Agent Builder: Building AI Agents with Tools & Planning

(17:57) Model Evaluation & Prompt Management

(21:23) Generative AI for Business Analysts

(23:24) AI Model Modality & Use Case Selection

(25:23) Popularity Distribution of AI Models

(28:18) Prompt Optimization Tools

(34:20) Building AI Agents: Real-World Use Cases & Ethical Safeguards

(40:13) The Capabilities & Limitations of AI Agents

(45:48) TPU vs. GPU

(50:33) Future of AI at Google Cloud

More from Eye On A.I.

All 266 episodes
#216 Nenshad Bardoliwalla: Inside Vertex AI - The Ultimate AI Toolkit by Google CloudEye On A.I. · 52 min
Listen in VO