AI Developer Tools at Google with Paige Bailey

9 Jan 2025 · 37 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: AI Developer Tools at Google with Paige Bailey

Podcast Information

  • Podcast Title: Software Engineering Daily
  • Episode Title: AI Developer Tools at Google with Paige Bailey
  • Description: Discussion on Google's ML, AI, and data science developer tools and platforms including Colab, Kaggle, AI Studio, and the Gemini API.
  • Host: Jordy Mon Companies
  • Guest: Paige Bailey, Uber Technical Lead of the Developer Relations team at Google ML Developer Tools.

---

Key Concepts and Discussions

  1. Introduction to Google’s ML and AI Tools
  2. Overview of Tools:
  3. Colab: A popular Jupyter notebook environment that runs in the cloud.
  4. Kaggle: A platform for data science competitions and collaboration.
  5. AI Studio: A dedicated environment for experimenting with AI models.
  6. Gemini API: A new API designed for various machine learning applications.
  1. The Evolution of AI and Machine Learning
  2. Historical Context:
  3. The conversation reflects on the evolution of AI, particularly the shift from models focused purely on text to multimodal capabilities (text, video, audio, etc.).
  4. The advent of transformer models (2017) has led to significant improvements in model versatility and performance.
  • Current Trends:
  • Multimodal AI models that can handle multiple data types (images, audio, video, code).
  • Tools now designed for easier integration and experimentation for developers.
  1. The Role of Different User Personas
  2. User Cohorts:
  3. JAX Users: Machine learning framework for building models, primarily for those with advanced programming skills.
  4. Gemma Users: Those fine-tuning or deploying pre-trained models, less focused on extensive programming knowledge.
  5. General Developers: Using APIs for various applications without needing deep knowledge of machine learning processes.
  1. Kaggle Generative AI Intensive Course
  2. Course Insights:
  3. The course attracted around 150,000 students, focusing on practical applications of AI tools.
  4. Content covered included prompting models, retrieval, embeddings, fine-tuning, and MLOps best practices.
  1. Gemini API Overview
  2. Capabilities:
  3. Supports various modalities (video, audio, text, code).
  4. Gemini 1.5 Pro model offers a 2 million token context window for processing large amounts of information.
  5. Gemini 1.5 Flash provides a cost-effective alternative for simpler tasks.
  1. Open Source vs Proprietary Models
  2. Gemma Overview:
  3. An open-source family of models that can be fine-tuned and deployed on local machines.
  4. Benefits include cost-effectiveness and customization for specific use cases.
  • Gemini’s Proprietary Approach:
  • While Gemini APIs provide high-performance models, they are proprietary, focusing on powerful server-side inference.
  1. Advanced Features in AI Studio
  2. Key Features:
  3. Code Execution: Ability to write and execute code dynamically to solve problems.
  4. Function Calling: The model can access external tools or databases to enhance responses.
  • Target Audience:
  • Designed for both experienced developers and newcomers, allowing a wide range of use cases from simple model queries to complex applications.
  1. Future of AI Tools
  2. Agentic Properties:
  3. Discussion on models being able to act autonomously with user permission and how they may enhance productivity.
  • Multimodal Learning:
  • Excitement about future advancements in AI models, particularly those that can generate and manipulate both text and multimedia content.

---

Conclusions and Recommendations

  • Encouragement to explore AI Studio for hands-on experience with AI tools.
  • Emphasis on the importance of learning and experimentation in the rapidly evolving field of AI and machine learning.
  • Follow Google Developers Twitter for updates on upcoming features and tools.

---

Additional Resources

  • AI Studio: [ai.google.com/studio](https://ai.google.com/studio)
  • Kaggle Course Details: Link to the course will be available in the show notes.
  • Twitter Handle for Updates: Follow Google Developers @GoogleDev for the latest news.

---

Final Notes This episode provides insight into how Google is shaping the future of AI development tools, equipping developers of all levels with the resources to engage with cutting-edge technologies.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Over the years, Google has released a variety of ML, data science, and AI developer tools and platforms. Prominent examples include CoLab, Kaggle, AI Studio, and the Gemini API. Paige Bailey is the uber-technical lead of the developer relations team at Google ML Developer Tools, working on Gemini APIs, Gemma, AI Studio, Kaggle, CoLab, and Jax. She joins the podcast to talk about the specialized task of creating developer tools for ML and AI. This episode of Software Engineering Daily is hosted by Jordy Mon Companies. Check the show notes for more information on Jordy's work and where to find him.

0:51Paige, welcome to Software Engineering Daily. Excellent. I am so excited to be here and really excited to have the opportunity to talk to you and also loved the questions that you are asking before we hit record. I think this is going to be a fun conversation. I do have a point to make at the beginning, because you're one of the owners of one of the funnest social media handles. You are dynamic web page, but I do have a question about it. Apart from being fun, have you ever done any dynamic web page design, web page loading that credits you with the honor of being the owner of such a handle? So I am not gifted in the web design space or the web app creation space for that.

1:31I look to all of my dear friends who are working on things like Next.js and all of the JavaScript and TypeScript libraries. I will say that I did have the pleasure and the honor really of working with the VS Code team for quite some time when I was at Microsoft. And that's not really web design, but it is very much kind of like the JavaScript TypeScript contingent. And I love and adore creating VS Code extensions just because they're super easy to create if folks haven't experimented with them previously. And they're also very, very useful in the sense that you can have VS Code extensions do a broad spectrum and variety of things.

2:06We were chatting about the fact that I had been following you for years now and that you, in my vision of the industry, you've been always in this AI space. Probably we would have called it ML or any other terms in the past. But I was thinking about my own career and I've always been in developer tools and DevOps platforms, stuff like that but i did have a short stint way back when like in 2013 14 if i'm wrong and what i at the time would call and probably it's still called lang tech industry so companies and products that i participated in the development products are machine translation but also in sentiment analysis and so forth and i'll bring this up not only to point out that i'm certainly not the expert in this field that's why you were here but also because it feels from my experience in that and following this field a bit from afar, but now quite close, that it all stems, this AI revolution that LLMs have put out there, it all feels like it stems from language, from written language, and spoken language, but written language, right?

3:07What are your thoughts on that statement? I will say I am glad to meet another kind of machine learning veteran. You know, I started building models, I think, around 2009, 2010. So it's been a wild and crazy ride since then. I will say that kind of the transformer models and things like GPT-2 and GPT-3, they originally started focused just on text and code, kind of the written word. But now we're getting into this really brave new world of multimodal models. So not just to this underpinning language backbone, but also really interesting capabilities in terms of video understanding, audio understanding and transcription, image understanding, kind of coupled with text and code as well.

3:49You can get a lot more out of the models, even apart and aside from text and from code understanding, which is very exciting. I'm sure you remember, like back in the day, even to just get a model to be adept at doing a single task, it took months of getting the right data in order and trying to experiment with different model types and trying to do hyperparameter tuning and then even just to get like the smallest percentages of improvements and now like all of these models can do relatively well for all of the tasks that we had been using single task models for out of the box so are we experiencing a step function evolution of word to vec sort of like that technologies that powered nlp that will be, I guess, the way in which I would classify that previous stage of text-based AI, ML.

4:39Are we experiencing just sort of like a natural evolution or does the underpinnings of what's happening with video, with multimodal models have a different nature? Yeah, it's a great question. I think, you know, everything kind of started with the Transformer paper around 2017. And then, you know, we're building a whole bunch of models building on the concepts expressed in those papers. But one of the coolest things I think now is that people are building these AI systems that kind of couple together different model types. Like as an example, if you're using Gemini, you're using kind of a mixture of experts model that is really, really good at multimodal use cases.

5:18But if you want things like audio as output or video as output or images as output, the model is not yet capable of that. it can generate text, but it is not going to be giving you kind of the image or video outputs that you would get from something like Imagine or Vio. So when we start kind of seeing these really novel new approaches, like I'm sure you've experimented with Notebook LM. Yeah. Yep. Where you can, for folks who might not be familiar, I encourage you to go try it out. But notebooklm.google.com, you can input a PDF or kind of a GitHub repo or anything, and it suddenly generates a podcast recording of two people discussing in great detail.

5:59You should point out at this point that this is not a notebook LM conversation. This is a real... That would actually be very hilarious to give the kind of transcript and to see how well the notebook LM folks. I was actually fiddling with some ideas to see if we could do something like that. Maybe in the next iteration of this conversation, this interview, but this is a real one. It's happening today, the 19th of November. Yeah, I love that these kind of AI systems that are increasingly multimodal systems give you the ability to create not only text and code as output, but also images and video.

6:34And it really kind of resonates with, you know, I'm thinking in particular of my cousins. They love watching videos. It's a stretch to ask them to, you know, read something a little bit more long form. So I think to really be able to engage with audiences and to help people learn and to understand and to really hit every single learning style, we're going to need to experiment with different modalities of outputs, not just inputs. Yeah, correct. So you work at Google. Google, in a very Google fashion, has joined, if anything, I mean, you've mentioned the papers that revolutionize this field. These mostly, if not all, came from Google.

7:11but has sort of like joined the release of models in an abrupt way in the sense that it's put out so many things out there. So give us a sense of what is Google doing with AI, specifically with this new generation of AI, and what kind of products and models do you focus on? Yeah, so my particular role is I'm the Uber TL for our ML developer tools, which is a new org that was created at Google just a few months ago. The products on this team are the Gemini APIs, AI Studio, Kaggle, Colab, Jax and the open source stack for Jax, and also Gemma, our open source model family. So basically everything that you can imagine from like a 3P facing ML developer perspective kind of lives in this ML developer org.

7:56The things that are top of mind for me for these tools are really kind of growing the number of students, researchers, and also early stage startups that are incorporating AI into their products. I think for our enterprise customers, there are a whole bunch of other, you know, like great tools that exist within the Google Cloud, like Vertex AI product offerings. But to really be able to move quickly and to experiment with the latest models, Gemini APIs, AI Studio are the place where you should go to try that out. and really the only place where you can get access to the latest Gemini models. Before we dive into the products, what is the typical persona that you're engaging with?

8:33Because I find fascinating the fact that we're talking about ML developers. Like, are there real people, and not one, five, six, but dozens, potentially hundreds and thousands of people that are able to not only train models like the ones that you mentioned in the Gemma family, but others, and able to also deploy them in a fashion that software engineers and developers without any first name or surname, are they able to do CICD with those things? Like, is such a figure, such a persona exists? It's a really different way of building software, I would say. And the personas for each of the tools would be slightly different.

9:14So as an example for JAX, JAX, for folks who might not be familiar, it's a machine learning framework that Google uses to build all of our models. And it took off like gangbusters. I think all of the papers being produced by DeepMind over the last few years are using JAX. It feels very similar to another kind of numerical library that you might be familiar with called NumPy. But it gives you the ability to build models or to build kind of physical systems and to dynamically scale them in a very straightforward way. So you can build a JAX model and with zero code changes, it can run on CPUs, GPUs, TPUs, and any arbitrary like hardware backend as a result of it using this thing called XLA, which is a machine learning compiler that was originally created for Google to be able to interact or to deploy models very efficiently on TPUs.

10:05So JAX, like when I think about the canonical JAX user, my brain is just like, oh my God, people who are building large language models or multimodal models, or who are doing like highly complex dynamic physical modeling, like that is the group. And that cohort is quite small. Like the number of people who are building models from scratch with JAX is quite small. Then when I think of the Gemma audience, the Gemma audience is slightly different, right? Like Gemma is a model that's already been created. You can either fine tune it or you could do continued pre-training on it, but you're probably using a high-level Python API to do that.

10:39You could also just take the model checkpoints and deploy them on multiple devices or deploy them in browsers. And the user groups for both of those aspects are a little bit different, right? Like the people who might be fine-tuning Gemma. Perhaps they're, you know, wanting to create evals or perhaps they're wanting to do some sort of research on it. Perhaps they want to use it as part of their product, but that's different cohorts than maybe from the building models from scratch, Jack's humans. But the beautiful thing about the Gemini APIs is that if you can make a REST API call, then you can call the Gemini model.

11:11And it's the same with OpenAI with Anthropic. We just recently released OpenAI library compatibility. So if people have already been preferring the OpenAI models. It's just a three-line code change to get the Gemini models being used instead. But the ML Ops process, in all honesty, feels a lot simpler than it did when you were having to worry about data versioning, model versioning, etc. If you're just making a REST API call, you do have to worry about which model you're calling. You have to worry about the format for your prompts. But there is a whole bunch other machine learning maintenance work that's just taken out of the equation.

11:50So it actually simplifies the DevOps process in a number of ways, as opposed to building your own models from scratch, deploying them and maintaining them. So I presume those three personas, even the first cohort that you mentioned, they're probably very acquainted with low level programming, despite that sort of like target architecture agnosticity of JAX that you mentioned, they must have this small cohort must have that knowledge. But have all of these three cohorts been present in the recent Kaggle workshop that you actually come from finalizing right now. Give us a sense of what's happened there and how people can know about future upcoming, if there's going to be another edition of that.

12:26Yeah, thank you for the question. We recently did a five-day generative AI intensive course on Kaggle, which is a platform at Google that originally was for competitions, but is now more of like a model hosting, data set hosting, learning platform. I think we weren't expecting so many folks to be interested in learning about the programming, but we ended up having, I think, around 150 ,000 students register. And everybody was kind of forking the notebooks, running them, asking great questions on Discord. The content was really around prompting models, retrieval, embeddings, fine-tuning models, and then also implementing evals and sort of these MLOps behaviors.

13:08And we had students that were really the spectrum from just getting started with the Gemini APIs to just getting started with Gemma. So lots of variation in terms of skill sets and backgrounds. But I think everybody, from what I can see, really enjoyed it. And I especially loved that the curriculum we designed was focused not just on the model calls, but also on all of the additional features that you need to have around the models in order to make these systems production ready, like setting up retrieval or prompt management or really designing strong evals, all of these things are very important to get the right outcomes from the models.

13:48Let's actually focus on that. Let's double click on that. So this is Gemma exclusively related, right? No, Gemma and Gemini, but the course was predominantly focused on the Gemini APIs. Okay. So what about those? Can you give us a broad overview of what the APIs are capable of? Yeah. So the Gemini APIs are kind of the recommended way to interface with our Gemini models. They support video, audio, text, code, etc. So all of those modalities that I was just describing as inputs. The Gemini 1.5 Pro model has on the order of a 2 million token context window, which means that you can send to the model a whole bunch of information right at imprints time.

14:29That means that you can analyze full videos, multiple code bases simultaneously, simultaneously, all of the above all at once, and be able to get sense out of it without having to go through the process of standing up a vector database or fine tuning. Gemini 1.5 Flash is our smaller version of Gemini. It has a 1 million token context window, which is still a lot, but it's also much, much faster and much, much cheaper than most other models out on the market. I think it's 7.5 cents per million tokens. And we also have a Gemini 1.5 Flash 8B version, which is around two-ish cents per million tokens, which means that you can record, as an example, everything that you're doing on your laptop screen, 365 days a year, you know, 24 hours a day, and it would still cost less than like a cup of fancy coffee to analyze all of the videos and to be able to make sense of all of the things that you're doing.

15:25So Google is really invested a lot in making sure that our models are performant, efficient, but still very capable. And also not really breaking price points for anyone. If you look at artificialanalysis.ai, the Gemini 1.5 Flash and 1.5 Pro models are always kind of the most cost effective frontier models on the board. Indeed, yeah, very affordable. Where does Data Gemma fall into this picture that you're describing? Yep. So our Gemini APIs, they're all proprietary models, which means that we haven't released the source code or the data used to train or the checkpoints or anything of that nature.

16:03They're just available via these REST APIs. Gemma is a family of open source models that we've, you know, kind of released all the things for. So you can look at the code on GitHub, you can kind of download them from Hugging Face, you can experiment with them, you can fine tune them. Our latest version of Gemma is Gemma 2, which comes in a variety of sizes. So 2 billion parameters, 9 billion parameters, and I believe 27 billion parameters. The smaller models are small enough that you can embed them within a browser. So embed them within Chrome or embed them on a mobile device like a Pixel. And they give you the ability to do a lot of interesting kind of text-only large language model work.

16:45So you can generate code, you can generate text, And then you can also fine-tune these Gemma models to do a broad spectrum of things. So like data, as an example, you mentioned data Gemma. There's also poly Gemma, which helps with multimodal understanding. So you can understand images. There's shield Gemma for security use cases. And I think the last time I looked, there were tens of thousands of Gemma fine-tuned variants on Hugging Face. So lots and lots of people kind of stretching them, fine-tuning them, kind of making them great for specific use cases. So what is the rationale behind releasing Gem as open source, Gemini as closed source?

17:22What is Google's stance on this rationale? Well, I obviously can't speak for Google, but from my perspective, I think it's really nice to have both options, to be able to call to a performant kind of proprietary model, send your data to a server. And then for other use cases, you might have different constraints, like you might be under different cost constraints. One of the nice things about open source models is that if you're running them locally, that's kind of free. You're just using your onboard compute. you might want to customize in ways that you would not be able to with a proprietary model, or you might be operating in an area where perhaps you don't have Wi-Fi connectivity, in which case having an open source model that's on board for your mobile device or for your laptop is kind of mission critical.

18:09You can't be sending your data elsewhere. There are some companies that also have data privacy constraints, and so they don't want to be sending their data off-site, which means that Rust APIs are kind of out of the question. And so having a version of Gemini that's not a mixture of experts approach, but is a much lighter weight, kind of very efficient model that's also open source so people can kind of tweak it, customize it to their delight is really powerful. Of the techniques that are more popular these days, like RAG, can you explain the differences between them? Like RAG, RIG, I believe, or RIG, I'm not sure how you pronounce that.

18:47Which ones are the most popular and what are the use cases why people use them for? So I think for folks who might not have experience with these different approaches towards retrieval, just think of them as kind of ways that you can get better performance out of your model's outputs and then also ground your model's outputs in data sources, which helps mitigate hallucinations and helps with accuracy of the model outputs as well. So as an example for retrieval, you might want to, one example that I hear quite often from customers is I would really like to ground the model's outputs based on my own company's internal data.

19:29So if somebody asks a question about, you know, HR benefits, or they ask a question about a specific club that is just internal to the company, I want to be able to source the outputs to not just use information that it might have learned from the internet somewhere, but to have it extract insights from my company's data sources and use those to guide the outputs. This is nothing really new. I think internal corporate search is something that everybody has been interested in for quite some time. But the retrieval phase is really kind of doing this kind of extraction from sources that might be relevant, giving that to the model, and then having the model summarize those insights as outputs.

20:13If you haven't experimented with, there are a couple of approaches for this that have been kind of baked in wholesale for the Gemini APIs out of the box. One is grounding with Google search. So you can turn on grounding with Google search. And if you ask the model a question, it will first kind of use the top 10 or however many results from Google and use those to kind of summarize and ground its answers, which gives you a higher confidence in the accuracy of the outputs. And then there's another feature that's only available through Vertex where you can say, I want the model's responses to be grounded in the data that I have located in this particular GCS bucket.

20:56And so you could say, hey, here's a pointer to all of my company's data. Hey, model, if you're going to be giving outputs, use these data sources to help with your summarization and then have pointers back to those sources. So just think of retrieval as a way to figure out what information to stick in the context window to help the model with more accurate summarization and its outputs. Would this technique work for the following use case? I'm a CTO, I'm a senior developer, hiring junior developers, and I want them to be constrained by, influenced by, and hopefully learn the company guidelines.

21:38So can I feed those assets into this retrieval technique and therefore allow for any junior developer to be able to be provided with answers that are fine-tuned to, again, the coding style of the company, the policies that need to be followed, etc.? Would it work in the same way? I think you could attempt it with retrieval, but dependent on how many guidelines you have at your company and also your stylistic guidelines for code bases, it might be worthwhile to first experiment with just putting that information into the context window. As an example, with Gemini, I had mentioned before that you can have 1 million tokens, 2 million tokens just kind of given to the model.

22:21If you do that with a repo and you say like, hey, here's my company's code base. Now please generate outputs aligned with the conventions in this code base, as well as any style guide or any kind of like guidelines that you might have. Gemini should be able to do that out of the box. And then oftentimes if you have stylistic constraints, if you do just kind of add that as a preamble in your prompt, you know, like if you're giving me code recommendations or if you're doing completions in this way, make sure to follow these stylistic conventions. Usually the model pays pretty close attention without even needing to set up something like retrieval or fine tuning.

23:01What about AI Studio? I haven't used it, but what I get from the name is that is this a playground where I can use all of this? Yeah, absolutely. So Aistudio.google.com is, and every time I mention it, I feel like I need to open up like a browser tab and start showing things. But it's Aistudio.google.com is a place where you can go, you can kind of experiment with the different Gemini models. So the Gemini 1.5 Pro family that I had mentioned before, as well as Flash and some of our newer model versions. You can also experiment with image generation within AI Studio. You can turn on features like function calling.

23:39If you want to do tools use, you can turn on code execution, search grounding. You can compare models against each other. You can also fine-tune models. You can generate API keys and kind of track usage over time, and all without having to kind of wrangle with the Google Cloud Console, which I think can sometimes be quite overwhelming for junior developers. So then I presume AI Studio is open to both all the cohorts that we've mentioned before, right? So those that have extreme expertise already in fine-tuning, actually developing models themselves maybe, to those new people that are just starting, right?

24:18Have you seen the most junior people start getting acquainted with AI Studio? What is the main use case they go about resolving for? Well, it's pretty much everything, right? Given that there are connectors to drive, that you can upload files, that you can kind of record yourself speaking or videos, you can basically just use it for any kind of model question that you might have. Like one example that I always like to show is like upload a video and then ask for, you know, extracting out all the logos along with the timestamps where the logos are occurring, transcribing all of the audio from the video, identifying all of the different speakers in the video, describing or summarizing the events from the video with timestamps, dividing it into chapters, identifying any like electronic equipment.

25:05And like all of these things are just things that you can ask in natural language within the context of AI Studio. All of these, it kind of makes me laugh because I'm sure you remember in the before times, there were all of these dedicated single task models that were sometimes available as things like cognitive services or like other specialized video intelligence APIs. Now, pretty much all of those you can just use Gemini for. And it's just a prompt as opposed to trying to figure out which API you should be calling and doing API key management for all of them. Of the latest features of the APIs, which ones are your favorite and why?

25:43I really, really love code execution and function calling. Just because, so code execution for folks who might not be familiar, it gives you the ability to say, Gemini, I'm going to ask you a question or I'm going to ask you to help me with a task. and you have the ability to write and execute arbitrary Python code in order to solve it. So it's setting up a sandboxed environment with the Python standard library, as well as a few other additional libraries, and then giving the model the ability to write and execute code for you. And if it gets it wrong the first time, it will just keep going and going until it gets the correct answer.

26:20And this is just available out of the box. It's a one-liner change. All you have to do is say like tools equals code execution and to turn it on. and you're off to the races. Function calling, likewise, very cool. It gives you the ability to identify tools that the model can call. So it might be like, hey, Gemini, you have access to this database, so you can write SQL code against the database. You have access to this weather API. You have access to this model that can do satellite image segmentation. So you could use that as a tool. And then you can ask highly complex questions and get Gemini to select which tools it needs to use in order to answer the question, as well as to execute any arbitrary code for those tools.

27:07So it's giving the model a lot of flexibility. Otherwise, it would not have. So this feels that it's going into the fascinating field of a sort of like agentic properties. But before we dive into it, in the code execution example that you just gave, how would the model know that it's achieved the right answer? Should the test be provided in the prompt? You don't have to provide the tests. I think for most code execution, it's just looking for a specific output, going to break the rules for podcast folks. So I will describe what I'm showing on my screen. I've just pulled up AI Studio and I'm selecting our smallest Flash models.

Read the full transcript

27:51So Gemini Flash 8B. I'm turning on code execution, which is just a little toggle button that you can share. And first I'm going to show what it looks like without code execution turned on. And then I'll show what it looks like with code execution. But you can ask questions like, please give me the dates of every single Monday in the year 2026. So if I hit run, the model will give me kind of a really troubling response that'll say, unfortunately, I can't provide a complete list of every Monday because this would require a calendar program or something similar. If I turn on code execution and then rerun that same prompt, the model recognizes that it needs to write Python code.

28:38And then it runs it for me until it gets the correct response. So you can see here that the first iteration of Python code, it ran, it didn't get the correct response. And then it just kept going. It said, I saw there was a bug in the previous code, and then it was able to get the correct response. Fascinating. Just for the record, we do have a YouTube channel. So this might be actually uploaded there. So for those of you intrigued about the interface of AI Studio, you'll find it there. Otherwise, in the URL that Paige mentioned, and it's quite intuitive. Excellent. It's obvious. Amazing. And there's also like one other thing that I adore about AI Studio is that after you do all of these really interesting explorations in the UI, if you hit this get code button, it gives you the exact code that you would need in order to rerun the experiment that you just did.

29:28And for tools use or for code execution, it's a one liner that just says tools equals code execution to be able to give Gemini the ability to write and to debug code over and over again. What supported? That's always my next question. So Go, Kotlin. Yeah. Didn't you announce something about Android Studio very recently? Yes. So Gemini models have been baked into Android Studio as well for code completion as well as code generation. So if you want to be able to use AI assistance within Android Studio or Colab or some of our other coding IDEs, that already exists and is powered by Gemini. But yeah, we saw on screen a minute ago, but for those that are not watching, so Curl, Python, plenty of languages.

30:14There's more, obviously. There's a myriad of them in the world, but a wide range of supported languages at the minute. Yeah. Following on the field, on the questions of agenticness, agency rather, how do you feel? What's your personal view? I know I'm asking now about the future. I don't want to get you to talk about the roadmap or stuff that is not shareable. But where do you see these things going, like models being able to act by themselves? This is a very broad way of describing it. But yes, how do you see that? Well, I think we're already getting into this world where, you know, the most interesting use cases for models, at least from my perspective, are these kinds of write, run, execute code, do it in a while loop until it works sort of scenarios.

31:00Though I think that in order to help people have confidence on these use cases, there has to be transparency over every action that the model is taking, as well as kind of overseer, like, yes, go ahead stage for folks, if there are any changes to the system that are going to be made. I will say one of the, you know, I had mentioned before that we're baking models into the Chrome browser. So if you want to try out Gemini Nano within the Chrome Canary release, that's available for you to test today. The Gemini Nano is also embedded within Pixel devices. But what that means is that... Within the Pixel device itself and the hardware, or rather in Chrome running in a Pixel device, or both?

31:42It's not embedded within the hardware, but the model is baked into the operating system. So it's running on device. But if you do have these models that are running on board, then suddenly you can start imagining really interesting step-by-step behaviors that the models might be able to make on your behalf. Like as an example, I would love to be able to say, hey, Gemini, please, you know, look on my calendar and find the next best time for me to go and like do yoga or something and have it be able to both look at the calendar for my yoga class, look at my work calendar and then try to figure it out for me and schedule time.

32:22Those are all things that could be done today in theory. It just takes someone kind of setting up those step-by-step calls to do with the model. And I wonder from a compliance perspective, if the model eventually will, after performing the tasks and hopefully correctly, will be able to deliver a sort of like a chain of thought proof of what the process has been. so that, again, someone verifying not only the tests eventually, but also that the process has been logical or compliant, right? That would be probably something that an enterprise user would be thinking of. Or before the model takes action, it could say like, hey, Paige, it looks like you have some free time around next Friday at 12.

33:06Do you want me to go ahead and book it? And in which case I could say yes or no. What else has you really excited about what's coming up? Also very excited about these multimodal paradigms. You know, Notebook LM was really enchanting for a number of reasons, but I think partially because, you know, not everybody is a text learner. I love to read. I adore it. I probably read too much. But many people prefer video content or they enjoy listening to books as opposed to reading them. So giving folks the flexibility of being able to learn in the way that is most effective for them, I think is really exciting.

33:44And then also just from the perspective of, you know, I could write books all day or like tell stories to my nieces and nephews, but I was never able to, I'm not gifted in drawing. So having the ability, it's a very challenging skill to learn. And so being able to generate videos or images is also pretty magical. We have a video model that is getting released through API as well called Veo, which gives you the ability to both describe a video and have it have it displayed in six second chunks or to seed the first frame of the video with just like a static image. And that's been really cool to see.

34:25where else can everyone find you where else can actually people know about the releases of that pertain Gemini the APIs Google AI Studio and all the things that we've been talking about yeah so I strongly recommend as you know I mentioned I'm dynamic web page pretty much everywhere but I strongly recommend following our our Google devs Twitter handle that should get you insight into the all of the latest new features that are coming for Google developers And then I would also recommend following some of the other folks on the team, Logan Kilpatrick, Chris Perry, who's the PM for Colab, and Matt Veloso, who's kind of the lead from Microsoft for the full team.

35:08So Google Developers, Twitter handle. My assumption is that everybody is going to have similar properties around Threads or Blue Sky or LinkedIn. in, but I will send the Google devs Twitter link via chat right now. So you'll have handy access to it. Is there any point in doing the Kaggle course that we've talked about and that I'll include a link in the show notes? Yeah, I definitely think so. So the Kaggle course, it's a five-day generative AI intensive course around one hour a day is the expected time for the coursework and then a follow-up hour to listen to the live stream. But each day includes kind of collab notebooks so you can walk through code examples.

35:48It includes podcasts summarizing a whole bunch of white papers or just the white papers themselves if you would prefer to read. It includes a live stream and then there's also a lot of great discord discussion about the course itself. So if you're interested in generative AI, both function calling, building agents, like prompting models, doing retrieval, interacting with embeddings, writing evals, like this should be a really great crash course for you to try. Yeah, the curriculum looks fantastic. And I really look forward to actually glancing over it and being able to understand at least, I'd be happy with 30 % of it.

36:22Awesome. Because I should point out that I'm not a developer. I think everybody can be a developer these days with these generative AI tools, to be honest. That's true. That's very true. And actually, these tools are felt we understand code bases, that C++ code bases that are way beyond my understanding. And I'm really happy that it's opened the gates of my understanding to, in this particular case, arcane and low-level programming languages and codebases. Anything else that we didn't touch upon that you would like to mention before we conclude? No, just the takeaway for everybody should be, if you haven't tried out AISudio.Google.com, go explore it, test it out on your own data.

37:03And we have a very generous free tier. So I strongly, strongly encourage you to take advantage of it. Well, thanks so much, Paige, and take care. Have a splendid rest of the week. Excellent. You too.

From the publisher

Over the years, Google has released a variety of ML, data science, and AI developer tools and platforms. Prominent examples include Colab, Kaggle, AI Studio, and the Gemini API. Paige Bailey is the Uber Technical Lead of the Developer Relations team at Google ML Developer Tools, working on Gemini APIs, Gemma, AI Studio, Kaggle, Colab

The post AI Developer Tools at Google with Paige Bailey appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
AI Developer Tools at Google with Paige BaileySoftware Engineering Daily · 37 min
Listen in VO