Building real-time voice applications with Live API

6 Aug 2025 · 40 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Google AI: Release Notes - Episode: Building Real-Time Voice Applications with Live API

Overview

In this episode of *Google AI

Release Notes*, host Logan Kilpatrick talks with Shrestha Basu Mallick, product lead for the Gemini API. They dive into the features and capabilities of the Gemini Live API, Google’s real-time multimodal interface for developers. The discussion explores how audio can serve as a unique and powerful interface, the latest advancements in the API, and the various applications developers are building with it.

Episode Details

  • Release Date: Not specified
  • Duration: Approximately 38 minutes
  • Main Topics:
  • Overview of the Live API
  • Unique aspects of audio as an interface modality
  • Developer use cases and feedback
  • Roadmap for future developments

Key Discussion Points

  1. Live API Overview
  2. Launch History: The Live API was launched in December 2022, gaining positive feedback for its real-time, multimodal interactions.
  3. Features: Includes capabilities for audio input/output, screen sharing, and webcam interaction.
  4. Recent Updates: Enhancements made to support additional languages, new voices, and session management.
  1. The Importance of Audio
  2. Natural Modality: Audio is a natural way for humans to communicate, often faster than typing, making it a dense source of information.
  3. Use Cases: The API supports various applications such as:
  4. Software co-pilots for complex applications like Photoshop
  5. Real-time interactions in vehicles (e.g., autonomous cars)
  6. Learning tutors, especially in language acquisition
  7. AI assistants for interviews and user research.
  1. Developer Use Cases and Innovations
  2. Software Co-Pilots: Real-time assistance for developers (e.g., coding agents).
  3. Educational Tools: Language learning applications leveraging audio and translation capabilities.
  4. Emerging Applications: Including interactive experiences in cars and potential use in customer support scenarios.
  1. Enhancements and Features
  2. Native Audio Models: Introduction of proactive audio and tone-aware responses to improve interaction quality.
  3. URL Context Tool: A new tool enabling the API to extract and analyze content from specified URLs to enhance the depth of responses.
  4. Async Function Calling: Allows background tasks while maintaining a conversation with the API.
  1. Developer Feedback and Roadmap
  2. Feedback Mechanisms: Gathering insights from developers to improve multilingual performance, session length, and turn detection.
  3. Planned Developments: Future features include proactive video responses and enhanced contextual understanding, aiming to create a seamless user experience.
  1. Market Outlook and Advice for Developers
  2. Voice AI Market Growth: Rapid expansion in the voice AI market, with numerous applications across various sectors.
  3. Getting Started: Recommendations for new developers include experimenting with the Google AI Studio and accessing available code samples and documentation.

Key Takeaways

  • Exciting Opportunities: The Live API offers unprecedented opportunities for developers to create innovative audio-based applications.
  • Importance of User Feedback: Continuous improvement driven by real-world developer experiences enhances API functionality and user satisfaction.
  • Call to Action for Startups: Encouragement for startups to explore and build applications using the Live API to fully harness its potential.

Demo Highlights The episode concludes with a live demonstration showcasing the Live API's audio capabilities. The interaction varied from different character personas (such as an over-enthusiastic blobfish and a snarky hedgehog) to provide a tangible experience of the API's responsiveness and adaptability.

Closing Logan Kilpatrick thanks Shrestha Basu Mallick for the insightful discussion, emphasizing the ongoing developments within the Live API and the exciting future of audio applications.

---

For more insights, you can watch the episode on [YouTube](https://www.youtube.com/watch?v=4xlwlU6h-wM).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We're talking about all things LiveAPI, which is our multi-modal interface for developers to build with. Audio is perhaps the most natural of the interface modalities. This is truly a revolution. This market has exploded. It's unlocked a lot of interesting use cases like software, co-pilot, learning tutors, real-time interactions in cars. We can have a coding agent that talks with you. The pair programming use case is the one that I continue to be super excited about. We've made builders out of so many people. Yeah, that's a cool thing. You really can build new things that you couldn't have built before.

0:33This is my call for startups to build stuff. We try to add more configurability and performance to the API. You can really prompt it to speak with a specific tone and style. Can you talk to me like an over-enthusiastic blobfish? A snarky hedgehog? An embarrassed panda? Did you know that? Okay.

1:00Hey folks, welcome back to Release Notes. My name is Logan Kilpatrick. I'm on the Google DeepMind team. Today, we're joined by Shrestha Abbasu-Malek, who's the lead for the Gemini API. We're talking about all things Live API, which is our sort of multimodal interface for developers to build with. Thanks for being here, Shrestha. Thanks for having me. I'm excited. Do you want to give the sort of high-level overview of, we've been working on the Live API now. I think the initial launch was December of last year. Yeah. Shipped a whole lot of stuff. Can you give, like, for folks who haven't been following closely, like, just the state of the world of where we're at with the LiveAPI?

1:34Yeah, happy to. So as you said, we released it in December and it had a really phenomenal response because you could now have real-time, bi-directional, multimodal, as you said, interactions with Gemini. And the thing that people liked, and we'll be talking about audio input and output a lot in this conversation, but people also loved that you could screen share and do webcam. So the screen sharing part is what was going viral online in December, which was crazy. It really was. And it's unlocked a lot of interesting use cases, such as like software co-pilots. So that was December. And then over the last few months, we've been steadily making updates to the live API.

2:19So around April for Google Next, we made a lot of updates to what we call our half cascade architecture. So this is native audio input, but the audio output happens through a text-to-speech model, same text-to-speech model as was being used by Notebook Ellum at the time. And that was the architecture we released within December. And, you know, around next, and we'll get into it, we made a lot of updates to, like, how many languages were supported, new voices, session management updates, and turn detection updates. And so we tried to add more configurability and performance to the API. And then now what's exciting is as part of I.O., we actually released native audio output.

3:07So now we have an audio to audio architecture. And so now you have, you know, more natural sounding voices and all the benefits of native audio, along with a couple of controls like proactive audio, which is, you know, the model being a little proactive. and how it chooses to respond, effective dialogue where the model picks up on the user's tone and sentiment, and you can turn on thinking with these native audio models as well. Before we dive into all of the details of sort of what people are building with the Live API, all the sub-bets new, I think audio is this special modality, and I'm curious to double-click on that and talk more about why it's so interesting.

3:53Yeah, you know, over the last few months as I've worked with the live API, I've also been learning a lot about audio, things I should have known. And, you know, somewhere audio is perhaps the most natural of the interface modalities. When you think about it, humans actually learn to talk before they learn to read. and even now most of us and this is important for the ui of the future most of us talk faster than we type right um and what that means is it it audio can be a really high like information dense modality uh the other thing you know which is kind of relevant in the context of now the native audio having thinking in it you know there's so many use cases like let's say you built a to-do app based on the live API, where as humans, we tend to think out aloud.

4:49That's one of the first things we do when we're problem solving, right? And so I think there's something very special about audio. And as we look into a world, which is post the chat experiences that generative AI started with, I think audio and talking computers will become really important. Yeah. How do you think about the sort of trade-off? And I think the thread of the model you know you can speak quicker than you can type I think there's like a there's like a precision trade-off which I think is really interesting um and I feel like that makes this um it makes the use case also very interesting in in how you like you yeah I'm curious to get your reaction to that but like how you think about that no I think that's true and you know that's that's one of the things where I often get asked the question okay so then should we go all in on native audio outward verses.

5:41And there is still things, as you said, related to this precision trade-off, related to even latency, where I think native audio needs to cook a little bit more to be like production grade. But I also think that the models, as they become more powerful, will also learn to work around this precision trade-off that you mentioned, because they'll learn to understand the nuances of human speech, the aberrations better. Yeah. There was many I.O. launch audio threads. There was VO3 getting native audio capabilities, which people have been super excited about. There was native audio in the Gemini models.

6:28There was also a bunch of text-to-speech stuff that we launched. That's right. Which isn't part of the Live API, different use case. But it's also through this like main thread of like we're going deep on audio stuff and making a huge investment in that space. Any I know you also worked on the text to speech stuff. Anything, any high level stuff that was exciting about that that you want to talk about? Yeah, I mean, now you have Gemini itself doing text to speech. Right. And I think that that's a really that's also native audio in a way. It's just not available through the live API. It's available through the chat experience.

7:04I think there's a few things that are cool about it. The first thing is, these are now controllable and promptable text to speech. So you can literally say, you know, talk like so and so, you know, and we'll do a little bit, tease a little bit on our demo later, perhaps if you have time. but like you can really prompt it to speak with a specific tone and style much more than you used to and then now we also have multi-speaker text-to-speech this one's really cool because this is how you can do kind of like uh maybe not the full notebook podcast experience but it is like the same it's the same models or similar now it is yeah now it is notebook lm's now on these models as well and yeah and you know like we haven't told the story a lot but it is it's exciting.

7:51I don't know how many people have like actually tried to use that product experience yet. Yeah. You know, I have a friend who's, who's actually taking these like sommelier courses, like he's, he's a techie, but he's learning how to sort of understand wine and identify wine and all of that. And the other day he comes to me and he's like, I'm using these like multi-speaker TTS to like, just like create podcasts out of, and they have to study a lot for their wine exams. And so he's using it to like convert it into these podcasts and having the speakers talk to him while just so that he's practicing for his exam.

8:29I love that. I feel like we should we should definitely talk more about TTS stuff. Coming back to the live API. We've we've obviously made a lot of progress. We've seen a bunch of people like actually start to adopt building with the live API in production. What are some of the things, like, I'm always thinking about, like, what are the use cases that developers should be, can be building with? I feel like some of this is, like, you have to show people the art of the possible because, like, these are, it is, like, truly, like, a new way to build AI product experiences. So what are some of the things that you've seen, like, use case-wise most successful from a customer perspective?

9:04Obviously, we see a lot of co-pilot or voice assistant type of use cases across a breadth of, you know, specific areas. So one of the interesting use cases that have started to emerge is, let's say you have a complex software like Photoshop, or we had a very successful demo at Next with Cloudflare. And this whole screen sharing allows the user to show the screen to the live API and get real-time feedback on how to navigate that software. But you've seen all kinds of other co-pilots and assistants like AI companion use cases, industrial equipment debugging use cases. So that's one category of use cases.

9:48Then the other category of use cases is learning tutors. And, you know, they could be specific to learning. There's also a specific sub-use case there around language learning, which leverages the excellent translation capabilities. then you have gaming assistants that are being built and then some of the more fun use cases that I've heard about uh is um um audio output and like real-time interactions in cars um I you know I was in a Waymo the other day I just just was in a Waymo today and I had the question that what happens if I leave my keys in the Waymo yeah you know if that happens in an Uber you know somebody sends me a message.

10:33But I had the question around, what do I do? And I was like, wouldn't it be great if I could just ask the car this question and the car told me? The car's calling you and being like, hey, I'm outside. I've got your keys. Yeah. Or don't leave your keys, right? They do give you that disclaimer. I came in and they were like, don't forget your keys. Yeah. But see, that's like, they're just like, that's just a PSA. Let's say I'm actually leaving my keys and the car is like, wait, hold on. Don't get out. You have your keys in the car. I think the other one that we've seen is actually conducting interviews.

11:06So these could be like just recruiting interviews, but I've also seen user research type of interviews being done using the live API. Yeah, I feel like that is a super interesting use. I feel like the cool thing is there's so many different use cases. The one that I personally want, I think this is the hard product experience to build, which is I want like full co-presence. I want it built into my operating system. I want to click a button. I want it looking at all my stuff. I want it like queuing up actions on my behalf when I'm like, oh, I missed this thing. Like actually let me, you know, fix that email or something like there's just so much cool stuff.

11:42Um, that's possible with the live API. I think that one of the mechanisms that's enabling that is like tools in the live API. I I think we shipped a bunch of new tools in the LiveAPI. It just like also generally has access to all of our tools. So is there, I'm curious about what that story looks like. Yeah. So, you know, one of the things that we did when we launched in December, which was also pretty well received, was that we shipped with tools. So at the time, the LiveAPI supported two first party tools. So Google Search as a tool. as well as code execution, which is an execution runtime as a tool.

12:23And we also, of course, supported third-party function calling. And what was special about the Live API at the time, and since then we brought it into our chat experience, is that you could chain these tools together. So you could technically get data from the internet and have the code execution runtime, do an analysis, and display that on the screen or talk about the results of the analysis. So that's what we've had. Now, recently, we shipped two new updates to tools. So first is this tool called URL Context. Which I love. URL Context, Chef's Kiss tool. It's great. We should have it. Everyone should be using it.

13:04It's awesome. Basically, what URL Context does is it allows you to get more in-depth content from a set of URLs. And why this is important is you can imagine use cases such as, let's say you search on a specific topic. and then you use this tool to get more in-depth content. You chain search with URL context and then going back to the data analysis use case that we talked about, now you have more in-depth data, more in-depth analysis and you can do sort of more things with it, right? And now what you have is your own version of a talking research agent that you can ask to do a deep dive on anything.

13:46Yeah. Yeah. I think people often like get the, they, people like have URLs as this like thing. And like, you know, I know, I know generally what's at this URL. I know there's some information there. And historically the models haven't been able to like really take action based on that. And I think it really is super powerful. It's like just such a common, I have URLs. I want the models to be able to know what's at the end of these URLs, bring that into the context. So I'm glad we launched that. That's also available outside of the live API. You can just use that in general, right? You can use it on the chat experience.

14:18Then I want to talk about one other tool. And this one is actually, like I said, the live API is two experiences right now, right? And we should talk about when should users use which one. But like we have the half cascade experience and the audio to audio experience. And so what we did on the half cascade experience is we introduced async function calling. and this is basically now it allows you to start a task in the background while you continue talking to the live API and then the live API notifies you when it's done using that task so it's it's early days for that tool and we're thinking about how we bring it to native audio but we wanted to put out a version of that tool and see what people do with it yeah I feel like that's a like in you know if you're to put this in production that's exactly what you need because the latency part around audio i think is like such a weird thing if you if the model is not actually doing things on your behalf um while it's responding and stuff like that this like takes me to um a couple of the other things that you had mentioned before just as far as like what's new that we've shipped proactiveness i think is a really interesting one like i think one of my gripes with the live api historically had been like anytime i said anything it would just it's like trying to respond it's like very eager to answer any sound or little noise that I've heard.

15:36And I think proactive audio helps with this a little bit. So you want to give sort of the rundown of what it does? Yeah. So proactive audio right now is basically the model chooses not to respond. So let's say you and I are talking to a real-time live API, and I do a sidebar with you. I'm like, Logan, can you dim the lights or something? And the model will know not to respond to that. So that's proactive audio. We also released effective dialogue. Again, early days for this feature. So we would love user feedback. But this is, you know, responding in a more tone-aware, sentiment-aware way to you.

16:21And the use cases could be more of, you know, the AI companion, the AI helper kind of use cases. This is like if you show up and you're like, you sound really sad, the model's not going to be like, oh, hey, have a great day. Like all this. It's like matching your level of like excitement almost kind of. Yeah. And it might even on occasion comment on it and say like, how can I cheer you up, Logan? Yeah. So, yeah. Interesting. That is kind of cool. Yeah. We do have some customers asking us for, you know, the specific use cases they're building with it. Let's talk a little bit about some of the some of the feedback that we've been getting from developers.

16:58you and I are in all the threads on Twitter. We see it all the time. Lots of great feedback. Also lots of great feedback from the customers that we're working with. What's some of the stuff that's top of mind from that perspective? Yeah. Some of my best user research comes from lurking on your Twitter thread. So thank you, Logan, for everything you do for us. It's a good exercise. So I'm happy it's helpful. Yeah. So one of the earliest sources of feedback that we used to get would be better multilingual performance, you know, especially languages like, say, German or Brazilian Portuguese. And I think that with native audio output, actually, we've improved on multilinguality.

17:39Again, with all things, these are super early days, we need to keep pushing on this. But the native audio, that's one of the benefits of native audio, seamless multilinguality. And actually, really quick, that's also what's powering the notebook LN experience, right? Like at the same time that we landed multimodality in the API, they also announced support for notebook LN to do all the different languages, right? It's the same model setup. It is similar, but you know, they use the text to speech, right? Because like there's the, but it is the same set of languages that are supported and native, as you say, native audio output.

18:14Yeah. The other one was session length, right? Because when we launched, you could do 15 to 20 minutes of audio and like video was very small, like less than five minutes. Again, work in progress. But what we've done is give developers a few controls, right? So now with the live API, developers can set up sliding window. So at least, you know, like you're not cut off, you know, your context window keeps sliding. We've also given developers the control to adjust their image resolution when they're screen sharing, because adjusting that, having a lower image resolution for certain use cases can give you a longer session length.

18:56Developers also have a control of should video be streamed only when audio is being spoken or should video be streamed independently? Again, has implications for session length. and we've given some configurability around how you do session resumption and if certain special turns should be kept in the session all the time, right? So all of these, again, we could be doing much more and all our work on context windows and all, but these already help to alleviate that problem a little bit. The third one, turn detection. This is a huge deal, right? Right. Because you want the model to not keep interrupting you all the time, as you said.

19:42This is the I think this is the hardest part. I think the times that I'm most frustrated is like the model will like weirdly do like a little half blip. It's like it's funny because in human conversation, it happens so much. Like you say some like affirming phrase in the middle of some when someone's speaking and like it's not weird. But then when the model does it, too, it is kind of weird in some way. It definitely is. I think that, yeah, and there's a lot of work we need to keep doing there. But a few things we've done is, you know, we launched with a server-side turn detection model. But now we've given developers configurability.

20:18And there's four controls along the lines of, like, how much sensitivity do you have to sound? How long do you wait after the human has stopped speaking to declare end of turn, for example? So there's configurability that developers can tune or they can just turn off our turn detection and bring their own. This is something that the community as a whole is investing in a lot. Like it is one of the problems in audio output right now that people are working on and trying to solve. So a couple of the other things that we keep getting, tool calls, improved function calling. That is our battle in progress, and we intend to get there someday.

21:05I think we're making good progress. I feel like with each of these, at least the last three, on the 2.5 Pro side, I feel like there's definitely been incremental progress in that direction, which is positive. I think we just need to make it happen across all the Gemini stuff. I agree. I agree. Yeah, no, we've definitely made a ton of progress. Function calling are also first-party tools. and we'll keep improving. And then, of course, you keep getting feedback around just the, you know, the trinity, performance, cost, and latency. Keep, reduce hallucinations. How do you, who do I get more sessions for, at an affordable price for developers who are building at scale that we want to encourage?

21:49And then latency, of course. Time to first token is so important in these applications. So we're obviously, we're getting tons of feedback. we're doing lots of stuff based on the feedback. How are you thinking about the things that folks are interested in or excited about, but like we're not going to build on the API side or that are like, you know, the model will eventually do that. So we're not going to build that scaffolding layer now because, you know, we think it'll end up in the main model. Yeah, I mean, there's a couple of things. It's not that we never say no, but, you know, we have to hear that feedback at scale and, you know, figure out like the cost of building that in.

22:24But for example, people ask the ability to bring their own fine-tuned models to the live API. We understand the use case, but that's just not something we're investing in the ability to do it right now. And then overall, I do want to make the point that, you know, some of these capabilities are moving down the stack. So we're at a point, and you know, the voice AI market exploded in H2 of 2024. So there's a lot of people building with this. And what often happens is in their individual applications, in order to solve a problem, they'll do a one-off implementation. But then if enough people have that problem, the orchestration library and framework providers build the solution into the frameworks.

23:09And then we'll hear about it and, you know, build it into the API, like, you know, the turn detection configurability. And then eventually, you know, going back to your point around, say, humans knowing when to interject with the filler phrase, whereas the model doing it in a way that's kind of disruptive for you. we would eventually like the model itself to be able to do semantic turn detection, right? And so then the model will be able to take in all of the information, not just what the human is saying, but how the human is saying it, how the human normally talks to decide whether it's a turn that the model should take.

23:46Yeah. One of the threads that's always so interesting in all these conversations that reminds, it sounds like a broken record, but I'm reminded of it every time. is just like how much the other parts of the Gemini ecosystem from a model perspective end up mattering for things like audio as an example of this. Like long context is like the main limiter. Like, you know, if the models had a hundred million token context, it'd be very different to interact with the model from an audio perspective and a video perspective, because you can just like keep everything in context all the time, bar latency and some things like that.

24:19But it is interesting that it's like you, yeah, it's interesting how connected it always yeah and you know that's also another area that i think we as a community need to improve on like model performance as the turns keep getting longer you know um another thing i've noticed is if i hem and haw a lot like let's say i'm like thinking out aloud and asking the model to do something for me but there's a lot of you know i'm just like a little random in my sentence construction the model kind of gets confused as opposed to if i just tell it in a punchy way like Here's what I need you to do. So, yeah, lots of things.

24:57Yeah, yeah. I'm super curious what's coming next. Quite a few things, but the things that we are immediately thinking about, speaking of thinking, one of the things that we should touch on is, you know, we did release the ability to think with the native audio model. And one of the things that we want to focus on next is how do we make, how do we get the right user experience for when these models think? Like, I imagine no one wants the model to be thinking out aloud. It can be sometimes painful to listen to humans as they think out aloud. So, like, what's the right UI? We also are thinking about this feature, thinking.

25:39We're also working on planning to launch this feature called Proactive Video. and there the idea is the model proactively is able to identify and respond to specific inputs in the video stream so a classic use case is let's say I put my key down somewhere in this room and I don't know where it is and let's say I just stream this room to the model and the model says oh at timestamp xyz I saw this the key and the evolution of that might be the model someday says, oh, I saw the key on that green sofa somewhere. I think this was part of the Astra glasses demo originally. They left something and came back to it, and they're like, where did I leave that thing?

26:23Exactly. People loved that. Yeah, it was a great demo. Yeah, I mean, I think Astra is a really cool product and a really powerful product. But our hope with this real-time live API is that people can build their own versions of Astra for their use cases, their specific needs. I feel like that's the interesting thing that I'm not sure that people appreciate, which is like the Astra story is, or the LiveAPI story is the story that you could go and build something like. Like Astra is meant to be the sort of research frontier of a lot of these capabilities, but most of that stuff ends up making it into the LiveAPI in some capacity, which is really cool.

26:59Yeah, and also going beyond Astra, right? We talked about your own talking research agent. Yeah. We can have coding agent that talks with you and pair programs with you. The paired programming use case is the one that I continue to be super excited about. I think it'd just be super interesting. This is my call for startups to build stuff, which is go in, you know, go and find a way to click a button inside of an IDE and have a paired programmer on demand to go and help you do stuff. I think that would just be such a magical experience. And through this lens of like, we talk a lot about internally, like the next 100 million developers sort of onboarding into this ecosystem.

Read the full transcript

27:38And I think it's going to, people are going to have to build tools like that in order to bring people in. And it's cool that it'll hopefully be live API powered to make that happen. I would hope so. And it's how incredible that you say the number 100 million, right? There was a time even a few years ago when we would say, oh, the total number of developers in the world is under 30 million. Yeah. And what's incredible about the times that we live in is we've made builders out of so many people. That's a cool thing. Let's actually talk really quickly about the audio market. I think there's like voice AI, voice audio market.

28:17There's a ton of stuff happening. I know you're super close to a lot of the things and you're at a lot of the stuff that's happening. Like for folks who aren't super connected, like any sort of high levels of, I feel like it, and the reason I think this context matters is I do think this is like what the future is going to look like. So it is interesting to sort of unpack where we are today. Yeah, I think, I mean, we've touched on a lot of this in the course of the questions, but I think I want to emphasize that it's, this is truly a revolution. This market has exploded in H2 2024, driven by a lot of capabilities emerging across the stack and a lot of builders being excited to build about it build um you know we've already talked about use cases another use case that i read in some report somewhere which really spoke to me is um being able to support our customer support agents by making after our calls right like a lot of times my problem is you know if i have to do like errand calls i don't get into to it in the middle of my busy workday?

29:17And wouldn't it be great if the humans could be supported through, you know, the AI being able to take at least the initial versions of the after-hour calls? So that use case really spoke to me. So yeah, people are exploring across the stack, you know, the model layer, the API layer, the orchestration frameworks. There's both horizontal and verticalized solutions, you know, for certain sectors like healthcare, finance, B2B industrial stuff, media. So it's really like, it feels like almost across every vertical, people are exploring how can voice help to, you know. Build stuff. Build new things.

30:06I feel like that's the cool thing, which I feel like is why there's been such an explosion is like you really can build new things that you couldn't have built before um which is always like the most exciting thread for me about ai i have one more question which is just around like um like advice for people who like they've done nothing in voice ai so far they haven't used the live api where to get started suggestions for like how to build anything like that for for folks who are like new to this yeah i mean i would say you know first get a feel for this on Google AI Studio, right? Play with the voice, both on the chat experience and on the live API.

30:46See all the ways, the natural sounding voices, the languages in which it gets things right, and all the places where, you know, it maybe doesn't do what you want it to do and let us know. And then, you know, you can, we have cookbooks, we have code samples, you can use that to get started are we have an excellent dev rel team that pushes out like you know little code samples of how you build various kinds of applications so yeah we've talked a bunch about the live api so far could we actually see a demo and hear some of the audio stuff live i think we'll focus on the hearing part of it yeah yeah let's do it so you get to hear a demo that you've never heard before i'm excited i love new demos

31:33Hi. Hello. What's your name? My name is Gemini. What does Gemini mean? Gemini is a constellation in the Zodiac and also the name of a NASA program that focused on space exploration. It also represents the twins in mythology. That's right. Okay, Gemini, Can you talk to me like an over-enthusiastic blobfish? Oh my goodness. I'd absolutely love to talk to you like an over-enthusiastic blobfish. Get ready for some really super... Okay, that's great. I would like you to keep your responses to two to three sentences unless I tell you otherwise. Is that okay? You got it. I'll keep my responses nice and concise.

32:24Okay, sounds great. Just two or three sentences. That sounds great. What's my name? You know, I don't have your name right now, but I'd love to know it. Actually, I'll make you guess it. So here's a clue. Oh boy, this is so... All right, all right. Don't get so excited. Here's a clue. Okay. This is a bit... Okay. My anticipation is through the roof. Let's see if I can guess your name. Okay, this is a bit embarrassing, but my name actually means the best in Sanskrit. It starts with an S and my gender is female. What's my name? Mmm, starts with an S, means the best in Sanskrit, and you're female?

33:19Could it be Shruti? That's a beautiful name. But that doesn't mean the best. My name means the best and it starts with an S. Oh, I apologize for the mistake. Let's try again. A name that means the best in Sanskrit and starts with an S. Could it be Shrest or Shrestha? I'm so excited to get this right this time. That's right. You got it right this time, but you should be embarrassed. You did not get it right the first time. So can you now sound like an embarrassed panda? Oh dear, I'm really embarrassed. Okay, you have a chance to make up. What is the meaning of Logan's name? Oh dear, I'm really embarrassed that I didn't get the previous name wrong.

34:09Please forgive me. Now, regarding Logan, it's a Scottish name derived from the Gaelic word meaning Little Hollow. I hope I've made up for my past mistake with this accurate answer. You kind of have. Robin, did you know that? Okay. So now, you know, this Logan guy, he like really likes shipping. Like he incessantly ships. And when we are talking about shipping, we're not talking about shipping packages. We're talking about shipping products like software, AI products. But like he ships all the time. Can you sound like a snarky hedgehog? And because Logan doesn't speak Hindi, can you make up a snarky poem about Logan's habit of shipping in Hindi?

34:59Because he won't know what we're talking about. Fine. I'll try to sound like a prickly hedgehog with a bit of an attitude. Honestly, all this talk about shipping is making my spines twitch. And you want me to make up a poem in Hindi for Logan who doesn't understand it? Fine. Whatever. Here's a little something. Logan ek aisa banda hai, shipping ka diwana hai. Product bas nikalata rehta hai. Aaram se bäjna usne na seekha hai. There. Happy now? Quite happy. Now, I actually want Logan to hear this. So will you translate this into English for Logan? Ugh, fine. Here's the English version of that poem about Logan's shipping obsession.

35:53Logan is this guy. A shipping fanatic is he. Always churning out products. Never sitting still, you see. There. Are you satisfied now? It's honestly the things I have to do. Well, yeah, I'm kind of satisfied. Could you make it a little snarkier, maybe? Seriously? More snark? Fine. I'll dial up the prickliness even further. Here's a slightly more barbed version for Logan. Logan, the guy who can't stop shipping. It's like he's missing a pause button in his internal programming. That's kind of true. Seriously. No, keep going. No, keep going. Please read out the full poem. Okay, okay, fine. Here's the complete snarky poem in English for Logan, with extra prickles.

36:47Logan, the guy who can't stop shipping. It's like he's missing a pause button in his internal programming. Seriously, does the man even know the meaning of rest? His shipping spree is beyond the pale. Honestly, a never-ending tale. One might think he'd learn to take a break, but alas, his shipping habit won't forsake. Happy now? Can we move on from this ridiculous hedgehog persona now? We should move on from this ridiculous hedgehog persona. Let's move on. Can you end us with a happy cow persona? Just like close this demo out. Moving right along then. As a happy cow, I'd like to say that this demo has been utterly delightful.

37:34It was utterly fantastic chatting with you. Now, if there's anything else I can help you with, please don't hesitate to ask. Have a wonderfully pleasant day. Did you say utterly fantastic? Yes, I did say utterly fantastic. It's like utterly, but with a halfy cow twist. Just a fun way to say it was absolutely wonderful. Thank you, Gemini. You're very welcome. Have a wonderful day ahead. Muchís thanks for the conversation. I love that. Thatís awesome. I feel like this is such an interesting example of just like the range of the conversation. It wasnít too snarky, was it? It wasnít too snarky. Whatís the bit onÖ Iím curious how itís coming up with personifying these animals.

38:27Like theÖ I guess the hedgehog. I think it just says like, oh, I'll be prickly in my responses and all of that. And so I just wanted to see if, or like at the end it says, Moo Chas Thanks. So like Moo, M-O-O Chas Thanks. How do you feel about Gemini's take on your shipping habits? I think it's a little bit overblown. I did just take a break. So I'm well charged and rested. So yeah, a bit of an extreme take, but I appreciate the sentiment. um this was a super cool demo i think a good example of like the range of everything coming together um also just like the coherence through the like it was a like through a long thread and the model able to like shift gears many many times like i think it'd be hard for um it reminds me of those like videos that people do where they like speed run through uh like not impersonating different voices, but whatever that, like, they do, like, act as multiple people very quickly, like one sentence.

39:29I feel like the model was doing that, which was really interesting. Yeah, yeah, exactly. And you know what was interesting is I didn't plan it like that, but like when it was reading out the poem about you in English for the first time, and I had a sidebar to you and it stuck, like it was able to go back and read that entire poem again. Yeah, that was awesome. Thank you for making it. I appreciate all the work that you and the team are doing in the live API. I think there's lots more cool stuff coming. We've done lots of cool stuff. Thanks, everyone, for tuning in and listening to this conversation.

39:59This is another episode of Release Notes, and we'll see you in the next episode. Thank you, Logan. This was really fun.

From the publisher

Shrestha Basu Mallick, one of the product leads for the Gemini API, joins host Logan Kilpatrick for a deep dive of Gemini Live API, Google’s real-time, multimodal interface for developers. Learn about how native audio alongside new capabilities like proactive audio and async function calling unlocks the unique power of audio as an interface.

Watch on YouTube: https://www.youtube.com/watch?v=4xlwlU6h-wM

0:00 - Intro
1:18 - Live API Overview
3:36 - Why audio is a special modality
5:07 - Speed vs. precision in audio
6:17 - Controllable and promptable TTS
8:31 - What developers are building with the Live API
11:14 - URL context and async calling features
15:02 - Proactive audio and affective dialog
16:55 - Addressing developer feedback
21:54 - Live API roadmap
23:49 - The role of long context
24:57 - What’s next for the Live API
26:41 - State of the AI audio market
30:10 - Advice for developers getting started with the Live API
31:16 - Live API demo
38:10 - Demo wrap up and closing

 

More from Google AI: Release Notes

All 30 episodes
Building real-time voice applications with Live APIGoogle AI: Release Notes · 40 min
Listen in VO