Inside Lyria 3, Google's music generation model

18 Feb 2026 · 37 min · 24 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Inside Lyria 3, Google's Music Generation Model

Podcast Overview Title: Google AI: Release Notes Host: Logan Kilpatrick Description: A behind-the-scenes look at Google AI, featuring discussions, interviews, and insights from AI pioneers and industry leaders.

---

Episode Details Episode Title: Inside Lyria 3, Google's music generation model Time Stamps and Key Topics:

  • 1:00 - Defining music generation models
  • 1:40 - Lyria as a new instrument
  • 3:05 - Connecting language and creative intent
  • 5:08 - Guest backgrounds and musical journeys
  • 7:57 - Demo: Instrumental funk jam
  • 8:29 - Bridging the gap for non-musicians
  • 12:03 - Demo: Exploring lyrics and vocals
  • 15:07 - The magic of iterative co-creation
  • 15:40 - Meeting users across the expertise spectrum
  • 17:01 - Empowering new musical expressions
  • 18:29 - Emotional and communal impact of music
  • 19:51 - Opportunities for developers and community
  • 21:09 - Real-time vs. song generation models
  • 23:23 - Creating experimental sonic landscapes
  • 25:08 - Demo: Capturing unexpectedness and energy
  • 28:33 - Evaluating music through taste and expertise
  • 31:30 - The diligence of music evaluation
  • 31:52 - The future of Lyria and AI-first workflows
  • 35:07 - Articulating creative vision through language

---

Key Concepts and Discussions

What is Lyria?

  • Definition: Lyria is Google's music generation model that creates original music based on user inputs (text and soon images).
  • Functionality: It allows users to provide artistic direction, resulting in high-quality sound output that can be further crafted.

Lyria as an Instrument

  • Lyria is likened to a new musical instrument. Users can shape sounds and express their feelings and intent through the model, even without musical terminology.
  • The model supports a practice approach, where users improve their skills over time, similar to traditional instruments.

Connecting Language and Music

  • The model emphasizes the connection between language and creative intent, allowing users to describe the music they want in nuanced ways.
  • Lyria can interpret various prompts to generate music that meets user expectations.

User Experience and Accessibility

  • Lyria aims to bridge the gap for non-musicians, empowering them to create music without needing deep expertise.
  • The importance of emotional and communal music experiences is highlighted, showing that music can foster connections and express feelings.

Demos and Real-World Applications

  • Various demos, including instrumental jams and lyrics exploration, showcase Lyria's capabilities.
  • The discussion touches on the potential impact of Lyria on music education and therapeutic uses.

Future Directions and User Control

  • The future of Lyria involves enhancing model control and granularity, allowing for layered compositions and temporal adjustments.
  • Encouragement for users to experiment and create innovative music, fostering a collaborative artistic environment.

Evaluating Music Quality

  • The evaluation of music generated by Lyria involves subjective taste and expertise. The creators gather feedback from diverse listeners to improve quality and adherence to user prompts.

Community and Developer Opportunities

  • The episode concludes with an invitation for developers and creators to explore the possibilities of Lyria, emphasizing the model's versatility in creating unique music.

---

Key Takeaways

  • Innovative Tool: Lyria is positioned as a powerful tool for both musicians and non-musicians, democratizing music creation.
  • Emotional Connection: The emotional impact of music is highlighted, with the potential for AI to enhance personal and communal experiences.
  • Continuous Improvement: As Lyria evolves, user feedback and creative experimentation are crucial for its development.

---

Conclusion The episode provides an in-depth exploration of Google's Lyria 3, emphasizing its innovative approach to music generation and its potential to empower users across varying levels of musical expertise. The discussions reveal the future directions for AI in music and the importance of community engagement in shaping the product's evolution.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to Lyria 3

0:00 to 0:31

Learn about Lyria 3's unique music generation capabilities.

“And with the tool you can start like building the sounds that you need for that to express that.”

Understanding Lyria's Inputs and Outputs

0:54 to 2:09

Explore how Lyria uses various inputs to generate music.

“Now, what are the inputs into that model?”

Using Lyria as a Musical Instrument

2:09 to 4:04

Discover how Lyria can be used like an instrument for creative expression.

“So even if you don't have the musical terms, you have the vision, you have the intent, you have the vibe, and you can use that.”

Personal Experiences with Music

4:04 to 6:15

Hear the guests share their personal music experiences and preferences.

“Like, whatever, like, literally, like, top 50 pop is the music that I listen to.”

Demonstrating Music Generation

6:15 to 7:58

Listen as the guests demonstrate how Lyria generates music from prompts.

“But then I got really into electronic music, you know, going to a lot of shows and festivals.”

Exploring Prompt Flexibility in Lyria

7:58 to 9:21

Understand the flexibility of prompts in generating music with Lyria.

“and there is a lot of detail here if I expand on everything that we introduced in terms of the instrumentation and everything.”

Advancements and Limitations

9:21 to 11:28

Discuss the current advancements and future possibilities of Lyria 3.

“really bring out the richness and detail.”

Analyzing a Specific Music Example

11:28 to 14:00

Delve into the specifics of a generated music example with the guests.

“We're only just barely scratching the surface of what's possible here.”

Iterating on Musical Prompts

14:00 to 15:00

Learn how careful prompting can guide AI-generated music creation.

“And I iterated until I got a result that I liked.”

Controlling AI Music Generation

15:00 to 16:00

Explore the levels of control available in Lyria 3 and how they enhance creativity.

“It was only until I started listening and iterating on.”
Show all 24 chapters

Facilitating Musical Expression

16:00 to 17:00

Discover how AI tools can unlock creativity in users who may not identify as musicians.

“And the great thing is we can meet the users where they're at and then help them grow or help them learn, right?”

The Emotional Power of Music

17:00 to 18:00

Understand the unique emotional connection people have with music and its significance.

“But this may spark a musical journey or a different entry point where music is part of that.”

Community Engagement and Creativity

18:00 to 20:00

Encouragement for the community to use AI for educational and therapeutic purposes.

“that like you actually don't feel in a lot of other things, or at least I don't feel.”

Integrating Real-Time and Non-Real-Time Models

20:00 to 22:40

Learn about the benefits of combining real-time music generation with traditional methods.

“I would love to invite the community to think about what educational use cases you can build with this technology.”

Exploring Unique Sound Combinations

22:40 to 25:00

Investigate how mixing different sound types can lead to innovative music experiences.

“know, you can not actually get music, but like you could kind of coerce the model into doing some of those odd things.”

The Role of AI in Musical Performances

25:00 to 28:00

Delve into how AI-generated music can become a part of live performances and collaborations.

“interesting music that you made with yeah yeah yeah here's one that i made recently i won't try to preface it by describing it but yeah let's just play it it's one of my favorites so far just because i like it Thank you.”

The Role of AI in Music Creation

28:00 to 28:35

Explore how AI tools might transition into physical instruments for artists.

“Because like it just sounds fun to jam, you know.”

Evaluating Musical Quality

28:35 to 29:27

Learn about the evaluation processes for assessing AI-generated music.

Understanding Music Prompts

29:27 to 30:13

Discover how prompt sets influence the music generated by AI models.

“And as the modeling team, you just know these prompts.”

Feedback Loops and Partnerships

30:13 to 31:26

Discuss the importance of human feedback and internal partnerships in development.

“And then that happened on one of the really hard prompts that I had.”

Future of Lyria and Music Editing

31:26 to 32:46

Anticipate future capabilities for music editing and generation with Lyria.

“I was gonna say there is something very beautiful about the fact that for music you can't shortcut it you can't rush it you gotta actually listen to it you can't just like speed it up right?”

Empowering Creators with AI

32:46 to 34:04

Examine how AI can help users become more expressive in their music creation.

“that, I can, you know, jam on the guitar a bit, and then I can bring that as a context and I say, okay, I want a backing drum beat, but I need these lyrics to hit in at this point.”

Articulating Musical Desires

34:04 to 35:20

Learn about the significance of articulating specific musical needs in AI models.

“And it's like, oh, we're all artists, right?”

Encouraging Musical Exploration

35:20 to 36:07

Understand the importance of experimenting and pushing the boundaries of music creation.

“being able to do that feels very interesting for this model.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I can take anything in the universe and convert it into something that this model can understand to generate a unique piece of music at that moment in time that is literally unique and has never existed before in the universe. That's awesome. And with the tool you can start like building the sounds that you need for that to express that. So even if you don't have the musical terms, you have the vision, you have the intent, you have the vibe, and you can use that. These models help you express that in musical terms. I feel like this hopefully will be sort of like the nano banana moment for people around music.

0:40Hey everyone, welcome back to Release Notes. My name is Logan Kilpatrick. I'm on the Google DeepMind team. Today we're talking about Lyria 3. I'm joined by Miriam, Jeff, and Jason, and we're talking about the new model available, all the product services that it's coming out to. But actually, what is Lyria for folks who don't know what the model is? Lyria is our name for our music model. So it generates music. Now, what are the inputs into that model? we have text we have image we'll soon have other ways that you can configure it but you're providing it with inputs to give it some artistic and creative direction and it outputs high quality sound audio music that you can then craft how you want that's effectively what it is i like to think of it as an instrument i love that how should folks think about like um if you've never used a music model before which i assume most people like haven't interacted with a music model like what's the mental model or the framing or maybe they should just go try it and experience it for themselves.

1:35Do you have like a quick description or explanation of what people would expect from one of these models? I'm a musician at heart and so I generally just think of this as a new type of instrument. You know there's a lot of technical you know details here and nuances and complexities but it is ultimately something that you're trying to use to make some sounds and express something in the way that you want and there's many different ways to shape and try to configure it and do all these things and you actually do get better at it over time much like practicing an instrument you also learn the nuances and you know how to kind of steer it in ways that are kind of unique right and so and you ultimately want to make it you know make things that are in your style right and and and your you know your likes that represent what you want to express and so in that way it's really like an instrument that you can play with but also practice right and perform with yeah I love that one of the things that we have learned a lot is that there are many people that have a something to say or like something that they want to express and sometimes it's just starting with that in mind so what is it that I'm trying to do or or how am I feeling or I want to do something for my family like a video like expressing how much fun we have over Christmas um and with the tool you can start like building the sounds that you need for that to express that.

2:55So even if you don't have the musical terms, you have the vision, you have the intent, you have the vibe, and you can use that. These models help you express that in musical terms. The kind of double down on the idea of this as an instrument, I worked on image generation for years. And one of the things as a person who had worked on language, like as a natural language processing researcher, is I really wanted to connect language to vision and have there be a tight connection between, you said you want this, this, and that. with these properties in your image and then you see it come out in the image.

3:25That's a really powerful thing and I think we really nailed that with Imagine 3 and then later on with Nano Banana. And so with Learia 3, a lot of what our effort was was let's really work on getting really rich captions that would go into the, to build the connection between like the way you're describing music and the actual music that comes out. And then what you get is this really nice generalization where you can start to really morph the language in order to morph the model's response to that. And then you get these really fine-grained kinds of details that you can express to the model and the model will kind of follow these contours And so we give a very long context of prompting ability to this model So I'm really excited about like seeing what people come up with with a model That's going to really like be able to be like clay that you can mold across the whole song I like to kind of mess with the model by putting in like strange punctuation or emojis You know I develop these and get it to react or respond in different ways sometimes unexpectedly sometimes awesomely Yeah, it's kind of fun.

4:17What kind of music do you listen to? Pop music, honestly, probably. Like, whatever, like, literally, like, top 50 pop is the music that I listen to. So, yeah, nothing good. Or good stuff, but, yeah. What's something recently that you listened to that kind of caught your ear? That's a good question. Honestly, I don't, like, I have a playlist from five years ago, and I just re-listened to the same old songs, but I honestly don't listen to music that much, which is why this is interesting to me, because, thankfully, I don't commute right now, So I feel like that's been, that would historically been the thing.

4:48And I don't, and I listen to podcasts when I work out, which had been the other thing where I historically listen to music. So, um, yeah, this will be a fun, it'll be fun to spend more time with Larry. Maybe this will broaden your horizons. Brace me back. Yeah, I felt this way. I've actually discovered some really cool music and some really good producers, um, from listening to podcast intro sequences. Interesting. Yeah. Um, Jason, what about you? What do you listen to? I listen to almost everything. Like I listen obsessively. My dad was a rock and roll pianist in Michigan in the 1970s and 80s.

5:18So I grew up with music around. My son plays classical guitar. And so and I've just basically like to make this clear, like my younger brother, Justin and I, every year we select our hundred favorite tracks that we listen to that year. So that usually is me paring it down from 800 songs from the thousands I listen to every year. So for me, like working on this music model is like tremendously fun because like I love so many different genres and I was like yes like let's go here let's go there and like let's see if the model can do this um so it's been really exciting for me in that regard as a music lover that's awesome Jeff we also I don't know you've got to plug your uh your own music and yes how do how do people find your music I didn't even know that you produce music and yeah yeah yeah this is C++ which is obviously a computer science reference as well like I'm on YouTube you know and my musical journey and and you all know this this is a dream job of mine right but I grew up playing piano, clarinet, and bass guitar.

6:13You know, when I was in my teens, I listened to a lot of alternative rock into my 20s as well. But then I got really into electronic music, you know, going to a lot of shows and festivals. And so in my 30s, I, you know, dove even deeper into it and started learning how to produce. And so I learned Ableton, and I make house music, and I still try to find time every single morning to just sit down in the DAW and just start crafting sounds and styles that I like, right? And that's ultimately what brings me joy. Well done. Miriam, what about you? I went to music school when I was little, studied piano and classical music for years, but I was more of a performer, so I did ballet and dancing and singing, a lot of singing.

6:58So for me, this project has been bringing back amazing memories and making me re-engage with my musical side that I had sort of, like, left behind. So it's very exciting. I listen more to like EDM, like chill EDM. It's my commute music. So it keeps me going. I've started to produce a little bit more these days. Again, like reconnecting with the musical side, but I, yeah, I love everything. And I love dancing. So anything that you throw at me, I will probably dance it. Marya, maybe we go and look at some examples of Lyria 3 and we can listen to some music and just hang out. I pre-generated a few examples and we can add more.

7:42But basically on the Gemini app, now we're enabling Lyria 3 and you can provide a prompt as big as you want, but also as detailed as you want with any sort of structure, what you're expecting to change, and it will do the generation for you. So in this example, we are generating an instrumental funk jam and there is a lot of detail here if I expand on everything that we introduced in terms of the instrumentation and everything. and let's just play it. Miriam, really quick in this example, like obviously you, like for somebody who doesn't make music, like I don't know if I'd have the terms to describe in this prompt, like maybe we can see some other examples after this, but like how flexible is the model if you're not like able to really articulate, if I can't prompt, if I don't have the prompt engineering skills yet for Lyria 3?

8:29Absolutely. I mean, there are many different ways to do it. I would use Gemini also to help me create a prompt. And so sometimes it's just that I'm advising the intent, what I'm trying to do, or I just say a genera and I iterate with Gemini until I get a prompt that I really like and I will try that. So that's another way to do it. And it helps you build knowledge in terms of like musical terms. So it's actually really, really helpful. Nice. There's musical terms, but there's also just artistic or aesthetic or creative terms, right? You can describe moods, you can describe emotions. Yeah, yeah.

9:01Using words, obviously. I feel like that's what I can do. Yeah, you can do that. and as eloquent or non-eloquent as a way as you want and of course you know you can continue playing with that. I love it. But we definitely give help just as we do with Imagine and Veo and all these things like you just say like I want a cat, I want a video of a guy jumping into a lake, and I want a funky jam. Like we have things that underlying will talk to the model and really bring out the richness and detail. Very cool. All right let's listen to the first song.

9:54Like when you cool studio headphones. Yes, exactly. In terms of really getting the stereo with headphones really helpful. Yeah, yeah.

10:05And so Mirren, does this do multi-turn editing as well? If I wanted to, say, take this song and riff on it? With limitations, you can do some edits and continue the conversation. And it may suggest different directions as well, so sometimes it's better that you prompt from scratch with the new prompt that it's giving you. Nice. Well, that was cool. Yeah, one thing to point out, as I was listening, I really appreciate how it really got some of the instrumentation around the different piano, this and that. But it's funny because you can always get better, right? Because I know for some of our examples, we've played it for the team.

10:42One of our teammates is a professional drummer and he'll be like, Oh, that snare sounded a little off. Or like, there's too many types of snare drums in that click. It's not a real drum set that only has one or two or whatever. But at the same time, maybe you do want to create something where like, a theoretical alien with eight arms is playing a bunch of snares. And so there's interesting things you guys think about there. But yeah, there's always things we can do better. You mentioned multi-turn editing. We're absolutely going to invest in being able to do that repeatedly and at a super granular level.

11:11So you can even focus on specific sonic elements. Of course, you want to be able to separate out different instruments. But ideally, in the long term, you want to be able to replace a certain instrument with a certain other one. And you want to condition it so that what that instrument does is literally accompanying what another instrument does and like figuring that out in the time dimension as well. So there's still so much work. We're only just barely scratching the surface of what's possible here. And we're really looking forward to making this even more advanced for everyone. I love that.

11:39And this Lyria 3 does support lyrics in songs as well. Cool. Do we have any cool, interesting demos of lyrics? I'm curious how natural it feels. I can generate some, I can bring one of my previously saved ones. Yeah, yeah. I'll read a lesson. Yep.

12:03I like this. It's one of my favorites.

12:21Pedal to floor but the heart's moving slow Nowhere to be and nowhere to go Just the harm of the tires on the wet pavement Finding the peace in a temporary arrangement City is breathing, I'm holding my breath Dancing with silence, avoiding the rest Neon reflections on the dash Life moves on in a purple flash We're just two heartbeats in the night Wrapped in the glow of the sea So vocals and singing are obviously critical to music It's a very, you know, critical element there And, you know, one thing we've obviously worked on is trying to like avoid the uncanny valley and like make sure that it sounds expressive and natural yeah and you know there's always more work to be done on that front but i'm really happy with the progress so far but then kind of harkening back to the previous conversation about the alien playing drums or whatever sometimes you do want you ultimately want maximum control right and so maybe you do want something that sounds very auto-tuned maybe you do want something that sounds not auto-tuned at all and is quite pitchy maybe you do want something with tons of reverb and tons of kind of breath noises and maybe you just want something that's super clean you know and and and so ultimately we want to provide that flexibility for people to choose how to make their music right if you listen to some of the music that i make it's like intentionally very robotic sounding voices because that's just kind of the vibe i'm going for that's awesome well and in this song that we were just listening to what got us to that so i actually had a very simple prompt for this one um it was a pop, rap, fusion, melodic, hooks, atmospheric production, late night, urban vibes.

14:16And I iterated until I got a result that I liked. Because originally it gave me male vocals and I want a female vocal. And I wanted some harmonies. I wanted some like voices in the background. So I managed to iterate a little bit until I got a sample that I, it resonated with me. And it sparked the, this is the thing that I would love to like right now listen to while I'm working on emails at night. Yeah. Is this a good example of like the controllability that you see with Layer A3? Like, could you actually, like in this example, you're not giving like who you want to be singing, but like in your, you said like you sort of did it multi-turn.

14:50Could you have just said like, I want, you know, a woman to be singing this or like that type of voice? I could have provided that level of detail at the point that I started creating some of the songs. I didn't have that intent. It was only until I started listening and iterating on. So it gave me, I could work with Gemini and with the music until I got something that really resonated with what I was after. And that's the key thing about making music, right? You kind of figure out what you want as you go along, right? Like by making the music, you start to craft it, you iterate on it until you get something you're happy with.

15:23But that's part of the process. It's part of the magic. And then you have a co-creation process, which is far better than just one thing out and then you get a result and then you're done. And now you can actually iterate on the idea and kind of just be aspiring and playing. Yeah. How should folks think just at the macro level about the levels of controls? And I think actually this is maybe somewhat product dependent as well. We can start with like at the model level. And then is there any nuance at like across the product spectrum of those controls? We want the model to be as powerful as possible, right?

15:53And then across the products, you have a range of, you know, musical expertise and knowledge and a range of technical expertise and knowledge. And the great thing is we can meet the users where they're at and then help them grow or help them learn, right? And so when it comes to the most advanced users and the most technical users that want to configure everything and have super maximum control, you know, we want to give that ability. And those by definition won't be a mass market consumer product, right? But that's why we've been working with our partners and with the industry and with specific artists as we build these models because we want that feedback to get to that level of quality and control.

16:33And so, you know, ultimately you can think of it in two dimensions. One dimension is sound, right? The different instruments, the different frequencies, the different, you know, layers of the song or sample you're making and time, right? You want to be able to control when and where that happens and the flow of things, right? And so that's something we'll always be working on. Yeah. I'm very excited about this launch because I think we're going to learn a lot from people that haven't had this sort of musical journey. But this may spark a musical journey or a different entry point where music is part of that.

17:10And I think that's beautiful. Not a lot of people know that they have something to express until they have the tools and instruments to do it. And we're trying to facilitate that. So I'm very, very excited about the learnings and how we can do more of that. Yeah, I have played around with the previous Lyria models and other AI music tools that are out there. And I feel like this hopefully will be sort of like the nano banana moment for people around music. It feels like with how broadly accessible the model is going to be, hopefully more people will experience that. Yeah, even in front of me, leading up to today, so many people have been coming to us excited about it, trying it out for themselves more so than before.

17:49Yeah. Jeff, I was reflecting in real time as you asked me about the music I like, and I feel like something is, and I'm curious, you three probably felt this and thought about it more than I have, but there's this emotionalness that you, visceral emotionness that you feel from music that like you actually don't feel in a lot of other things, or at least I don't feel. and I've also as someone who can't who does who's not musically inclined enough to sort of express the details of that I'm curious how you all think about like these tools and these models well yeah like I don't know the the relationship between those things from from Lyria and like how people will be able to do that stuff not to get too philosophical or we'll be here but I've always felt like music is special I think psychologically or even physiologically if you look at how the brain works and everything back from the early prehistoric humans with bun flutes and banging drums and using their voices and just you know it's a it's a it's a very personal thing but it's also a very communal thing right some of the peak moments in my life have been centered around a particular song or a particular chord progression or a particular you know gathering you know um we speak with so many different tongues but like you go to a concert and like it doesn't matter where you're from like music brings people together and so whether you're using a bone flute or you're using Ableton Live or you're using, you know, a model or tweaking all of them in combination with each other, like being able to express something that any other person or animal honestly can hear and emotionally react to.

19:23That's like super duper special. I just think it's magical. And I use music to regulate my moods, to give me energy for different kinds of things. So when I was working in my dissertation, there's a set of like trip hop and drum and bass albums that I listen to. So when I go back to writing code, I put those on and I back in it. If I'm writing a paper, I listen to Miles Davis and other things like that, or I have Brian Eno. And these just allow me to slot right in to that mental and emotional state I need to do that kind of work. I love that. Obviously, like the models available for everyone. Any advice for people who are like going and experimenting with Lyria and like trying to build a product around it?

20:01Yeah, I've had people come to me like, oh, what if you make this thing that analyzes this image and takes this other thing and combines it with the volcanoes erupting over there and converts that into a you know a formula and then makes a song or makes a track or makes a soundtrack based on that and i'm like that's awesome you should go build it the sky's the limit like i'm really looking forward to seeing what people come up with is like i can take anything in the universe and convert it into something that this model can understand to generate a unique piece of music at that moment in time that is literally unique and has never existed before in the universe so again I'm getting so philosophical, but I think it's just super cool.

20:35That is cool. I would love to invite the community to think about what educational use cases you can build with this technology. How do you help children to understand music in different sort of dimensions and levels when you may not have the right tools at your disposal? Or maybe your school can fund instruments. Also in the health space, I'm pretty sure there are plenty of things that we could do with some of this technology to help even train like music therapists. that do an amazing job for people. So I would love to invite the community to just start. We make it available to you so we can build some amazing stuff to help others.

21:10I love that. We had Lyria real time in AI Studio and I feel like it's been my go-to just playing around with it. I love putting in the prompt and changing the bars and I forgot what the name of that example is that we have, but it's cool and I feel like there's lots of interesting experiences and I've actually found myself like trying to remix that into other into other like music actually became something that I had at my disposal as somebody who wanted to build something which historically it hadn't been before we did that and I feel like this is gonna take it even farther I know that there's a small but but hopefully growing and hardcore audience of Lyria real-time fans so for folks who have played around that experience is that something they should think about or like how to think about the difference between real-time and yeah song generation.

21:57We're absolutely continuing to work on that model as well. And so the real-time model is generating music on the fly. So think of it as a soundscape that you're sculpting as it's coming out. And the cool thing about that is these work together, right? Because as music is being generated, you can always go back after the fact and try to edit it. And, you know, so I actually think, you know, real-time, non-real-time, these all are part of the bigger picture of like making music, editing music, producing music, performing music. It's all part of that spectrum. If we were to boil down to like just a raw audio capability, like what actually would make Lyria different from, you know, we have like speech generation models as an example.

22:36I'm sure you and others have like, you can do some like weird things where like, you know, you can not actually get music, but like you could kind of coerce the model into doing some of those odd things. I'm curious actually, there's any like interesting edges or like interplay between some of the other models in the model families that we have. Yeah, I mean, you can, I don't know, this is exactly what you're asking for, but you can have speech happening in the song. Yeah. Right. And so you can prompt for speech and then the model will speak through these things. You can prompt for train sounds and things like that.

23:06And it doesn't always come across with the train sounds yet, working on that. Yeah. But it can sometimes lead to really interesting music. Like, you know, like one time I put a prompt in, it's like a man with a really dry voice is singing as though he's a cracked kiln longing for something, blah, blah, blah. And then you end up with this really interesting kind of unique sound. And so the fact that like we're mediating all this stuff via language does get these really weird, interesting mashups in the model itself. But, you know, maybe back to your point, as you start to merge in speech data with music data and diegetic sounds, like the kinds of sounds you just kind of hear around you and so on.

23:43you could maybe really create interesting new sonic landscapes that are the kind of experimental variety that, you know, and it really gets you back to what people like Ryan Eno and others were doing in the 70s, playing around with new synthesizers and like various like just ambient sounds and so on. I'm hoping we'll see some cool stuff. Yeah, sonic landscapes is a much more ambitious way of framing where my head was at, but I think it's actually a good example of like, you know, there's like really subtle background sort of, we'd consider like background music in a movie, which like, you know, maybe there's not actually someone singing and there's kind of, uh, you know, yeah, there's ambiance, but maybe there's like, I don't know, like more sound effects also in there.

24:25So I'm curious if that's something that is like sound effects as like a use case and something that works well. Yeah. I mean, on the whole, like, uh, across the teams, we're all thinking about these things. Right. And if you think about it, the difference between speech and rapping and singing or foley and sound effects or noise or sounds versus music. These are all just labels that we as a society or humans have put on things. But it's a spectrum. And many of these things, all of these things can be combined in different permutations to make art. right and so that's ultimately our goal to empower people to to make that art jeff you have some interesting music that you made with yeah yeah yeah here's one that i made recently i won't try to preface it by describing it but yeah let's just play it it's one of my favorites so far just because i like it Thank you.

26:05Nice.

26:32Yeah. Jeff, you gotta give context on this. And I feel like something we were talking off camera, something that would help me for these songs is now, like, actually visualizing, like, instrumentally what's happening. Because I'm having, like, this song is actually a good example where I'm, like, having a hard time wrapping my head around, like, what even would be happening. Totally. And for me, I hear the bass guitar. I'm like, oh, I wish I could play that well. And I hear the drums and the different things going on. But yeah, it'd be awesome. You could vibe code an app or something that basically shows the spatial arrangement of these different instruments, what they typically do look like.

27:00And the saxophone comes in over here. Yeah. What intrigued you? And also, what were you trying to capture? Yeah, I was trying to capture some notion of unexpectedness. There's definitely some improvisatory elements there. you know and I just wanted to make something that sounded like an increasing kind of crescendo of energy in a somewhat kind of like you know the excitement you get right before you hit the launch button right yeah and could you did you were you like letting the model drive the sort of unexpectedness or is that something where like you were sort of you know dictating what unexpectedness meant in that context I didn't dictate what unexpected meant I'd specified instruments I I did say a timestamp roughly around where I wanted that to crescendo or increase.

27:44But otherwise, a lot of it was just playing, you know, letting it play with itself. Yeah. And honestly, you know, I forgot to mention, but I'm also in like a little alt rock punk rock band with some friends in Brooklyn. And so like now I'm like, wait, I want to take that and like, you know, turn it into tabs or sheet music and then like play that live, you know, physically. Because like it just sounds fun to jam, you know. do you think artists or just everyday people will like have a you know this type of like ai tool as a physical instrument that they're like working with because i imagine like it might be weird if you're like with your band and then you like whip out a laptop or something like that like i don't know it just seems it seems on the guitar yeah that's the magical thing about computers man like just figure out some way to get the inputs and outputs out and you can come up with the form factors you want you can optimize for the latency or the sound or the whatever requirements that you want right cool let's let's talk really quickly about evals i feel like um and also if there's like anything interesting and like how this sort of story has evolved from lyria to lyria 2 to even the real time to now lyria 3 of like what what are we actually hill climbing on is it just like taste of you know you all are vibing and listening you're getting paid to decide what music sounds good and what music doesn't or what's the um yeah how are we some of the very initial evals certainly but no not at scale and evals evaluations yeah so i mean you pointed out you know music is about tastes right so what counts is good music or it's bad music but ultimately we want to respect the user's intent with what they are trying to create right um but also when it comes to having human beings listen and give feedback on what the model was creating we intentionally went and got a very broad range of people including people with a lot of musicality musical expertise right because you want to be able to catch the details on all those things and so yeah i mean it's probably pretty obvious to state but the more the better and the broader and more expert the better yeah yeah and i don't know how granular you want to get with this but We actually followed a lot of the same approach that we did with Imagine and Beo, where you end up with sort of what we have as our vibe prompt sets.

Read the full transcript

29:55And as the modeling team, you just know these prompts. You know what that image should look like. You know what that video should look like. Or you know what that song should sound like. Or you know what the past attempts were. And so you build a mental model for a small set of prompts. And then you know when you're hearing certain things happening. Like, oh, the harp actually came in on this track for the first time. And then that happened on one of the really hard prompts that I had. I was like, okay, cool. This is actually happening now. Or you're looking for singing virtuosity and various things like that.

30:25So there are some things where no matter what genres you like or don't like, you can still assess the musical quality as a team. But then once you've kind of got to a point where you think, well, all right, this checkpoint seems pretty good. Let's fire it off to our human abouts. And then we have general pools, we have music expert pools and so on. And we run a lot of those. We look across many dimensions from like the prompt adherence, is this following what I instructed it to do, but also musicality wise from the vocals, from the quality, fidelity. And as we mentioned, we look at like various pools, but basically we're just trying to make sure that we have a good understanding on what the gaps are so we can get better at those.

31:05And sometimes it's looking at what distribution do we have in one checkpoint versus others and we keep hearing. what is fantastic as well is the partnerships that we have internally across google to help us kind of like listen because you can look at an image and say i like it or i don't like it but when you look when you listen into music sample you need to listen to the whole thing and then be very diligent to like this i like this didn't quite fit and make sure that it has some consistency so there is a lot of effort a number of hours into listening to to the songs and the evals and it's a beautiful task just to see us on the desk like but also like with the rest of the pools um that we work with.

31:43I was gonna say there is something very beautiful about the fact that for music you can't shortcut it you can't rush it you gotta actually listen to it you can't just like speed it up right? Yeah. What what comes next for Lyria? I feel like obviously like obviously folks will experience the model hopefully many more people will sort of including myself uh figure out, you know, become more in tune with music. And I'm sure we'll get tons of feedback. And there's a million things that we already know we want to do that the models will sort of need to get better at and evolve into. But I'm curious for all three of you, actually, like what's top of mind going into the future as this model sees the light of day?

32:24I think just as we saw with Nano Banana, where image generation turns into image editing and modification and having many more ways of controlling the model and the kinds of outputs you're getting. I think that kind of editing and generation for music is where we're going. We really want to be able to, and then what we're talking about with the model as an instrument becomes even more accessible to people and musicians and so on, because now I can bring that, I can, you know, jam on the guitar a bit, and then I can bring that as a context and I say, okay, I want a backing drum beat, but I need these lyrics to hit in at this point.

33:01And then the model will weave that together. So that's definitely where we're heading. Yeah, ultimately it's about maximum control and granularity and layered composition, both sonically and sound and temporally in time, right? And just giving that maximum control so people can make what they want. And in terms of user experience, like having something that is as natural and intuitive as the conversation we're having right now. So it's really like an AI first way to work with these models. DAWs are amazing, the digital audio workstations, but they require a ton of expertise. And Jeff knows how to do it.

33:38I don't have that expertise. But once you kind of do that, you have that accessibility of people to be able to work using language, using audio references, using visual references and so on to shape the sounds that they're interested in hearing them. That's awesome. That was going to be my point that the empowerment that you get through any starting point, being an image or being your own lyrics, you're writing something and you just want to use that to start creating something. So, you know, when it comes to thinking about music and referencing music and coming up with the right words to describe the music you want to create or hear, I think it's good for people to like, sometimes we like draw this line between artists and us.

34:19And it's like, oh, we're all artists, right? So let's all work together and let's all create cool things together. And we don't have to rely on naming some concrete thing that's like in a box over here or whatever. Right. That's like we are developing the vocabulary to describe the music we want to make and listen to. Yeah, yeah. I feel like that would be very helpful from like a... Because, yeah, because I feel like I certainly don't have that vocabulary. And I feel like there's lots of people who don't have it. So I think it'll be interesting to see like how these product experiences can like actually help people sort of like work towards building the vocabulary to actually be able to describe what they want.

34:51And we had a conversation with the Genie 3 team. And, you know, a thread of that conversation was like Genie 3, prompt engineering is back. just because it feels like we're in that era of the model story. And maybe that's less true for Illyria 3, where you could just vaguely describe what you want, and the model will figure it all out. But it does feel like there's actually a real edge in being able to articulate exactly what you want and the way you want it, and specifically in the musical context, being able to do that feels very interesting for this model. I think there's always going to be a raising of the bar, right?

35:27And you can always get better at it. And there'll be people that are motivated to spend the time and build expertise to get better at it. And that's great. That's working as intended. I think hopefully more people will try it out and get their toes in. Hopefully some people will decide to go even further. Maybe some people will decide to be professional producers or artists. And then we go from there. And so I think it's always great to have something that you can get better at. As we're developing it, I often told the team I want the model to be weird-able. So it's actually able to be weird. and that's one of the fun things when you have that kind of control we've been talking about where you can press and prod and use weird turns and then you actually get different kinds of sounds to come out and that's I think just an amazing way to pro these kind of capabilities and I love it make music weird, I love it yeah, key music weird, that's a good one Miriam, Jeff, Jason, this was awesome I'm super excited for folks to get their hands on Lyria 3 I'm super excited to see all the cool and interesting music I'll start listening to your music Jeff, I'm excited.

36:24We'll get you one more follower on YouTube to tune in and listen. Thanks, everyone, for tuning into this episode, and we'll see you in the next episode.

From the publisher

1:00 - Defining music generation models
1:40 - Lyria as a new instrument
3:05 - Connecting language and creative intent
5:08 - Guest backgrounds and musical journeys
7:57 - Demo: Instrumental funk jam
8:29 - Bridging the gap for non-musicians
12:03 - Demo: Exploring lyrics and vocals
15:07 - The magic of iterative co-creation
15:40 - Meeting users across the expertise spectrum
17:01 - Empowering new musical expressions
18:29 - Emotional and communal impact of music
19:51 - Opportunities for developers and community
21:09 - Real-time vs. song generation models
23:23 - Creating experimental sonic landscapes
25:08 - Demo: Capturing unexpectedness and energy
28:33 - Evaluating music through taste and expertise
31:30 - The diligence of music evaluation
31:52 - The future of Lyria and AI-first workflows
35:07 - Articulating creative vision through languag

More from Google AI: Release Notes

All 30 episodes
Inside Lyria 3, Google's music generation modelGoogle AI: Release Notes · 37 min
Listen in VO