Mikey Shulman: Suno and the Sound of AI Music (Encore)

29 Aug 2024 · 48 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Generative Now | AI Builders on Creating the Future

Episode Title

Mikey Shulman: Suno and the Sound of AI Music (Encore)

Podcast Overview Generative Now is a weekly series by Lightspeed, focusing on the emerging stories, strategies, and insights from pioneering AI companies. The host, Michael Mignano, interviews key figures in the AI industry, including founders and engineers, discussing their impact on the future of work.

Episode Description In this episode, the focus is on Suno, an AI music platform co-founded by Mikey Shulman, which simplifies music creation through prompts, aiming to democratize music production. The episode revisits a previous conversation with Shulman about the technology and vision behind Suno.

---

Episode Chapters

  • (00:00) Introduction to Mikey Shulman and Suno
  • (04:03) Inspiration from Transcribing Earnings Calls
  • (08:41) Creating a Unique Audio Dataset
  • (12:37) Text-to-Speech Hacks for Music Generation
  • (16:25) Exploring Product-Market Fit in Generative Music
  • (21:15) The Future of Music Formats with AI
  • (28:44) Suno’s Achievements and Milestones
  • (31:39) User-Centered Design in Music Technology
  • (38:26) AI and the Limits of Creativity
  • (40:48) Navigating Music Rights and Regulations
  • (46:19) Talent Acquisition at Suno

---

Key Themes and Insights

  1. The Vision Behind Suno
  2. Suno aims to make music creation accessible to everyone, allowing users to generate original songs through simple prompts.
  3. Music has been undervalued compared to other forms of media, but AI can transform how it is created and consumed.
  1. Technology Development
  2. The development of Suno was inspired by previous work in audio AI at Kensho, particularly in transcribing earnings calls. This helped them recognize the potential of audio data.
  3. A unique dataset was necessary for training AI models, leading them to create their own audio datasets since existing resources were inadequate.
  1. Generative Music Opportunities
  2. The current music landscape is heavily dependent on streaming platforms which may limit creative formats. Suno aims to explore new possibilities in music creation and delivery.
  3. Generative music can foster a deeper connection between creators and listeners, moving beyond passive consumption.
  1. User-Centered Design
  2. Suno focuses on making music creation intuitive, avoiding complex workflows suited only for professional musicians.
  3. The design philosophy prioritizes understanding how novice users think about music and how to simplify the creative process.
  1. AI's Role in Music Evolution
  2. The discussion highlights the potential for AI to transcend traditional music creation limits, enabling new styles and formats that have not yet been explored.
  3. Shulman believes that AI will not eliminate artistic expression but will instead enhance it by allowing creators to reach new heights.
  1. Navigating Music Rights
  2. The episode discusses the ethical and legal implications of AI in music, emphasizing Suno's commitment to not infringe on existing artists’ rights.
  3. Future developments may include collaborations with artists for generative music while ensuring fair compensation.

---

Conclusion The episode with Mikey Shulman presents an optimistic view of AI in the music industry, suggesting that generative tools can enhance creativity and accessibility. As Suno continues to evolve, it aims to redefine how music is created, shared, and experienced.

---

Stay Connected

  • Website: [Lightspeed Ventures](http://www.lsvp.com/)
  • Twitter: [@lightspeedvp](https://twitter.com/lightspeedvp)
  • LinkedIn: [Lightspeed Venture Partners](https://www.linkedin.com/company/lightspeed-venture-partners/)
  • Instagram: [Lightspeed Venture Partners](https://www.instagram.com/lightspeedventurepartners/)
  • Podcast Link: [Generative Now](http://generativenow.co/)

Additional Notes

  • The conversation reinforces the point that the future of music, shaped by AI, is full of unexplored potential. Mikey Shulman and the team at Suno are focused on creating products that inspire joy and collaboration in music-making.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:04Hey everyone and welcome back to Generative now I am Michael McDonough I'm a partner Lightspeed And for this week's show, we're revisiting one of our favorite episodes with Mikey Shulman, the CEO and co-founder of Suno. Now, since this episode first aired, Suno has launched its first mobile app, making the generation of music possible from wherever you are. In this episode, I spoke with Mikey about the hack that was key to making music with Suno, why AI will change the format and structure of songs, and what it means to disrupt the AI-generated music space. So take a listen to my conversation with Mikey Schulman.

1:07interests and passions growing up, they're very aligned with what you're working on right now. Yeah, definitely. I think I've definitely been doing music for much longer than I've been doing AI. Played a lot of piano as a kid. Taught myself a few other instruments. Played bass in a lot of bands in high school and college in New York. Some very small clubs that maybe you've heard of. and really fell in love with it. I think, you know, I remember very distinctly the first show I played ever was with a four-piece band. And I remember at the end of the night getting a very small stack of bills from the promoter.

1:50And I remember just thinking like, wow, this is criminal. I know that I had more fun than the audience did tonight. It wasn't a very good show, but it's just so fun to make music with people. and I guess it comes full circle having done a lot of stuff in between then and now being able to let people make music is a lot of fun and they enjoy it and it definitely feels good that we make something that makes people smile. It sounds like you were entertaining people. You were getting paid for it. So what about did you and I guess your bandmates, did you ever spend any time in the studio? Did you try to record music or only ever really playing out live?

2:30We did. We recorded one EP. I'll say, like, I'm certainly no seasoned studio musician, so take this all with a grain of salt. But, like, a lot of takes, it's, like, laborious and, at least for me, in many ways, less fun than just playing for fun. I actually remember one track. It was actually, like, a really good take, And I actually kind of like slipped off my chair a little bit, like at the very, very end. And we ended up having to basically redo the whole thing. It was like really, really sad. Yeah. And just thinking like, wow, that would never have happened if we weren't recording. So it seems like you were very focused, or at least for part of your life, really focused on music.

3:14But you studied physics at Harvard, which is a very, very different thing than music. And then you end up at Kensho, right? Tell us a little bit, maybe how did that happen? And tell us a little bit about Kensho, maybe for people that don't know about it. Yeah, total accident, honestly. Really a nice right place, right time type of a thing for me. I was in the last year of grad school, and I was introduced to some people at Kensho, one of whom is one of my co-founders now, Martin. And we went to lunch because I was local. And at lunch, they said, when do you want to come interview? And I said, I don't know.

3:52I'm a grad student. I can do whatever I want. And they said, how about right now? And so we went upstairs and I interviewed and I did very poorly. But they decided to give me a chance. And yeah, like I said, right right place, right time. A lot of stuff kind of luckily went right at Kentro. So Kentro was a company that did machine learning for financial services. We did a lot of NLP. You can think there's a lot of documents that are really financially relevant and had you kind of in an automated way makes sense out of a lot of those. Acquired by S &P Global in 2018. In some sense this was amazing.

4:28We got acquired by a company with one of the biggest, if not the biggest, trove of financial documents that you might want to play with. So you're kind of like a kid in a candy store. Got lots of good data to train on. We had a lot of fun, learned a lot, met a lot of great people. and sort of was the impetus for starting Suno is we did one audio project, which was learning to transcribe earnings calls, so do speech-to-text on public company earnings calls. So I don't know if you've ever read one, Mike, in your seat, the transcript. I have. I have. You know, when I was at Spotify, I would try to listen to most of them, but whenever I couldn't listen, I would read the transcripts.

5:11Pretty dense. A lot of tense information. Yeah, not real page turners, but there's a good chance then that you read an S &P transcript. Oh, okay. And a lot of people resell the S &P ones, and they're perfect, those transcripts. And pre-2019, they really couldn't be done with any automation or machine learning. And so people would ask, like, why don't you just go to insert your favorite cloud provider here and get speech-to-text from them? The answer is that it will come back with so many errors that it will take longer to fix than to start from scratch. And, of course, if you have 20 years of perfectly transcribed training data sitting in the basement and you're a machine learning company like this, you know, your eyes kind of light up.

5:55And that was our first foray into audio AI. And, like, the rest is history. We fell in love with it. It is, like, it's so beautiful. It's so much more human. It's messy. there's so much interesting stuff in it and it is so far behind images and text and I think that was definitely true in 2019 when we started on that project and you know in some sense it's more true now if you just think about everything that you've seen happen in images and text and there's nothing fundamental that says audio has to be behind forever and that was I guess the kick in the pants that we needed to say okay let's go and and let's do this right we love we love audio.

6:38We love AI. Let's do this. Yeah, I mean, that's an observation we had at Spotify often. And Daniel Ek, the CEO, has talked about this in that audio is something that everyone consumes. We all listen to music. We all, you know, many of us listen to podcasts, audiobooks, the radio. I mean, it is massive, massive in terms of the amount of attention that gets allocated to consuming audio content. But for some reason, it hasn't necessarily been valued in the same way as, say, video. I think things are just beginning to change now. But, you know, even rewind just six months or something like that. And audio looks a lot like text did in 2019, which is to say there's a lot of very unipurpose models, like very task-specific models.

7:27And so, you know, we look at text a lot to kind of understand the landscape. And so imagine, I know that we did this at Kencho, things like named entity recognition. It's like, here's a piece of text. Let me pull out all of the people and places and companies in this piece of text. And you'll train a model to do only this one thing. And you will always be extremely data limited. You're basically always just limited by how much label data you can get your hands on. And I think audio in those two tasks looks a lot like that. It's like, how much label data can I get on speech to text? or how much labeled data can I get on text-to-speech?

8:02And these models end up small and somewhat brittle because of the lack of all that data. And now go back to text. No one would ever think about training a named entity recognizer model. You would just paste some text into GPT, and you would say, give me the named entities in this piece of text. And it's because there was some sort of paradigm shift there of I'm going to forget about a single-purpose model. I'm going to train a very large self-supervised model on as much unlabeled text as I can get my hands on. And okay, then let's just do the same thing for audio. Let's train a large self-supervised model on as much audio as we can get our hands on.

8:36And then we can figure out how to make it do the things that we need to do. And I think you're just beginning to see that happen in audio. So if I could, maybe my understanding of what you're saying, probably in a much more simple way, because I'm obviously not nearly as close to the research as you are. It sounds like lack of training data, lack of data has been a big limitation. Gathering up a bunch of text doesn't seem that hard. There's, you know, a treasure trove of it sitting on the Internet. Audio, maybe not so much. Is that kind of what you're saying? Big time, big time. So there's no common crawl or pile for audio.

9:16And even if you had it, which you don't, it's really hard to work with audio. You can't easily inspect your data. You can't easily search over your data. You know, in text, if you just have two data points and you want to know, are these two data points the same? Are these two sentences the same? That's really, really easy to do. In audio, that's really, really hard to do. And so you just have to be much more careful with how you arrange and organize your data in a way that you just don't have to in text. And so, yeah, it makes people not want to do it. Got it. So how much of the focus of Kensho was on this work that you were doing around audio and transcriptions of earnings call?

9:55Was that just sort of a sliver of the work? And where did that happen along sort of your journey at Kensho? It was a part, but not the majority at all. Like most of what we were doing was text. And it happened maybe started a year after the acquisition, you know, after we kind of learned where all the data lived and what all the interesting projects are. And I think, don't get me wrong, there are a lot of future interesting things to do in financial services with audio. But I think ultimately there's way more stuff with audio outside of financial services that you can do. And that's like really what drew us to want to go do that.

10:36And I think financial services is a, for good reason, a somewhat more risk averse domain. And so I think there's so many interesting things to do with text there that it will be hard to look away from those text based projects. Got it. So so you're you're you're running machine learning at Kensho. Then obviously go through the acquisition. You're inspired by this this challenge around audio. Did you and your co-founders actually, I believe all your co-founders were at Kensho, did you leave specifically to go work on this audio project that became Suno? Pretty much. We knew that as much as we loved our jobs, pursuing audio in a financial service company doesn't necessarily make sense.

11:20And so I think we didn't know exactly the form that our product would take, but we knew that there was big opportunities. Yeah. So tell us a little bit about that sort of founding journey. You know, I'd love to hear a little bit about the early days of Suno. You know, it's funny. We looked around for a while and we talked to a lot of people doing things in audio. And in some sense, one of the biggest learnings there was what I said before, is like people really don't like working with audio. And I think maybe that's what made us special. We really liked working with audio, but like most people don't like working with audio.

11:52and we knew that the quote-unquote right way to do things would be to try to train a big self-supervised model, a foundation model is the buzzword. And we always said we wanted to do this right so that things are kind of future-proof and it gives us capabilities very easily compared to training things in sort of the old style way. And honestly, we weren't sure the instant we left Kencho that we would do music as opposed to speech. You know, we knew speech very well. A lot of people told us like speech is this really, really big market. Don't focus on music. Like that would be a really, really bad idea.

12:30And I think a couple of things happened. One is, you know, by virtue of all four of us being musicians and audiophiles, like we kind of couldn't help ourselves in some sense of starting to do music. And then there was another really big realization, which is that we put out an open source project called Bark, which is a text-to-speech model and it got good uptake by the community got a lot of github stars and what we found is that the we would ask people like what they were interested in on the there was a little signup form on the readme from bark and the biggest thing people wanted was music and this was like a big aha moment for us it's like here's this text-to-speech model it got a lot of popularity people like it and the biggest thing that everybody here wants is not text-to-speech, it's music.

13:19And so if you kind of can't help yourself from trying to do music and then everybody in the community is telling you to go do music, that's like a pretty strong signal. That was a little over six months ago. So we're like just past six months from releasing our first music model. We've released a few more since then. And yeah, haven't really looked back. So you come out with Bark and you're saying developers in the open source community are saying, hey, we want to use this for music. We want to build music applications with this model. That's right. And in fact, we had a little Discord server, and we would see people trying to pull music out of Bark, albeit poorly.

13:58And this is like a real clue of like, here's somebody really abusing your model to try to do something it's not meant to do because they really want to do it. Yeah, and what was that like? Like, did it work? Was it actually spitting out music? Or what did that even sound like? I will decline to say what is and is not music. No, it spit out like little bits of music sometimes, not terribly reliably, obviously. You know, and sometimes it would sing and sometimes it would be background music and it was a total crapshoot. But this was a model meant to do something very, very different. Got it. And maybe talk a little bit about the model and how it worked.

14:36And, you know, give us sort of the simplified version of how this model works, maybe compared to for some of the other modalities we talked about, text and video. It's no secret. We're big fans of Transformers. This is probably our text backgrounds. Not all audio is done with Transformers. There's a lot of audio that's done with diffusion. And these two methods have pros and cons, depending on the modality also. But we are big fans of Transformers. Some of it was inspired by some academic work, but there really was very little out there in the open source around doing things in audio with Transformers.

15:11And so there was a lot of kind of basic research or technology that we had to build. We had the good fortune of by picking something like Transformers where there's a lot of work being done in text, there's a lot of stuff from the open source that we were able to borrow. So if you go and you look at the source code for BARC, you'll see big capital letters we we thank and attribute a lot of the code to, for example, Andre Karpathy's Nano GPT, which is like a really, really easy stripped down implementation of GPT. And it's just like having resources like that available to us really let us focus on the bits that we love to do, which is really understanding audio and really trying to model audio correctly.

15:51And so Bark was a series of a few transformer models that you need to turn text into ultimately by sounding speech or sometimes in some abusive use cases, nice sounding music. But yeah, it's crazy to think that that was a year ago. I think we've advanced a lot in the community and the open source community and the research community has also advanced a lot since then. So developers wanted music. You guys as founders were very passionate about music. Sounds like you were all musicians. But maybe beyond the passion for music, why music from a product direction? Like, what's the opportunity for generative music?

16:34I think there's a lot. And I think we only focus on some of it. I'll tell you the corner that we focus on, which I think is the most interesting and the most fun, is making music that doesn't exist. And so I think, you know, the way we think about this is let's just look at the music landscape right now. most people have a streaming service like Spotify or Apple Music and they kind of sit there passively listening to music much of the time maybe not even paying attention to it and there is so much more to experiencing music than just that and one very big thing is creating music and I guess I am I'm lucky to have experienced that joy creating music with people but there's usually a pretty big barrier to doing that, which is like getting pretty good at an instrument or getting pretty good at some complicated piece of software like Ableton or Logic or even GarageBand.

17:33And you can think generative AI is a tool, a means to an end to letting everybody kind of take the sounds that are in their head and make really nice music out of it. One thing that we often think about is you look at like the gaming industry is 30 to 50 times bigger than the music industry. And, you know, if you want it to be really reductive about why that is, it's because like you are a very active participant in gaming. And so, you know, one thing to think about is like, what are the other 49 50ths of music that we've just yet to build a compelling product experience around? And it's not only like I sit in front of my computer and I make music, and then I share it with someone, although that is really, really big, but there's lots of things from doing it with people, like you often play games with people, and that can be some of the most fun gaming experiences.

18:24You know, we talk a lot about having collaborative concerts and stuff like that, or little jam sessions. There are, I think, totally unexplored social dynamics around what happens when you let people make music for their friends and share it the same way you might do with an image. So there's just so much that is, I think, unexplored and we're just beginning to explore the creation bit now, but there's a lot more to come after that. Maybe to set the context for, I'm sure, a lot that we're going to be talking about with this, what is the product today? Just talk about what Suno does specifically today.

18:57So today we let you, in a couple of ways, make songs that don't exist right now.

19:09So there is a song that's in your head. There's something that's going on in your life right now. And you can turn that into a song. The same way that you might write in a journal or take a picture of something, you can turn it into a song. And music is a really human form of expression that I think we're all really hardwired to want to do. You can do something as simple as type in, make me a reggae song about podcasting, or you can go a little bit deeper and you can bring your own lyrics and you can really understand the different ways you can prompt these models to do things, whether it's with playing with line breaks or saying this is a verse and this is a chorus or describing the music with adjectives that most people would not think to use to describe the music and out will pop a song.

19:55and hopefully you enjoy it. And I think most importantly, hopefully you've enjoyed the process of arriving at the final product. Now at scale, like when people are doing this at scale, what do you think the role of generative music will be in relation to non-AI generated music? Fast forward five years from now, are my kids, instead of listening to Taylor Swift, are they gonna be listening to AI generated music? Is there a hard line? How will these things coexist? I'm not necessarily sure there's a hard line because I think AI is a set of tools and those tools will come to artists as well. And while we are not focused on that, other people are.

20:34And I'm sure there's going to be a lot of work where artists do things with AI. But I think the answer is definitely both. And you like listening to Taylor Swift because you have some connection with her and her music and what she's writing about. and you will also like listening to the song that your best friend made for you because you have a very different kind of connection with your best friend and um i think you know when i said uh we're trying to figure out what the other 49 50ths of music really is that is to say i don't think that first 50th is going away it's like really really important um so yeah i i think i think there is a pretty big continuum here that we're trying to explore.

21:19How much do you all as a team think about the format of music as it exists today? Right now, music is delivered obviously digitally via streaming service, Spotify. It's organized into these algorithmic playlists that could deliver to people. As a result, the actual creative format of the music is often optimized by the artists, by the producers to play into that. It's something like two and a half minutes to three minutes, gets right to the hook to catch you as quickly as possible. How much do you think about the format of music as it exists today? And how do you think the format will evolve in a world where much of it is AI generated?

21:56It's a really good question. We think about this a lot. And I think the short answer is we don't know, but we know it's going to change and we know that it is changing. And I think, you know, like you mentioned, songs are now, you know, two and a half to three minutes and maybe five years ago they used to be longer and 10 years ago they used to be longer still and they're trying to get to the hook faster you know for for various reasons having to do with how streams are paid out um you have things like tick tock now where you're taking 60 seconds of an existing song and you're doing something with it and i think ai will certainly accelerate some of the changes in format but we don't exactly know what they are and we are playing around with them a lot you know one thing that i think pretty firmly is you know right now we make songs that are, you know, one to two minutes.

22:39But a one-minute song is very different from a one-minute clip from an existing song. And how do you condense the journey of a whole song that used to be three minutes into one minute is something that I don't think it's for us necessarily to define. I think it's for us to make the tools to let people make the music and then that new art form will kind of emerge from people making stuff. You know, and then kind of as a corollary, when you have a lot of one minute songs, not one minute clips of songs, what does a playlist mean? Playlist maybe now takes like different forms. Like maybe actually I make an album and my album doesn't have 11 songs on it.

23:16Maybe my album had 25 songs on it because it's still that listening experience. And it is a new way to tell stories in kind of smaller, smaller chunks. You know, it's like a book with three page chapters. Yeah. It seems like the delivery mechanism of music has so much to do with the actual creative format of music, right? And so it strikes me that however we're all listening to music will probably inform what people end up wanting to do with Suno as a tool. And so do you imagine a world in which, you know, a Suno-generated song is delivered in the same way that, say, a Taylor Swift-generated song is, and we're sort of listening to them back to back or are they happening in like completely different contexts and experiences i think it is certainly possible that you're listening to those things back to back but there will certainly be new experiences from this new art form that we are trying to figure out what they are and i think it would be maybe a little bit reductive of us to think it has to be all one or all the other.

24:20You know, like one thing we say is we are almost indifferent as to what you do with the songs that you create. You own the songs that you create with Suno. If you want to put them on YouTube, that's good. We're not going to try to prevent you from doing that. And we're not going to try to take a cut if you make a lot of money. But we're basically just as happy if you shared that song with your three college roommates and you all got a good laugh out of it. And then And that's where the song lives forever and ever. So and again, I think there's a lot in between Spotify and that three person group text message.

24:59It's interesting to think about where this all could be going. I mean, maybe just taking like image models as a comparison. It feels like the creativity and almost the art form of making an image will naturally change. Like the experience of making that will change as the models get faster and faster. Right now I go into mid journey. I type in a prompt. I get four of them at low resolution. You know, I think that experience has then inspires the way that I create for that medium. And I imagine as it gets faster and faster, I probably start to create differently. Right. And maybe in a more dynamic way.

25:38now going back to Suno and applying it to Suno, I imagine as your models and your experience gets faster, it will also impact the creative experience, but it might impact the consumption experience as well, right? I mean, can we imagine a world where music is almost dynamically being created on the fly for our own personal tastes? A hundred percent. And I think there's, you know, Right now, music gets streamed to you on the fly, and there's a really interesting discovery problem of figuring out what music out there is going to be most in line with my weirdo taste in music. In other modalities, like basically user-generated content makes for increasingly smaller and smaller niches that really appeal to people.

26:24And ultimately what Generative AI will let you do is let you create that really, really micro niche that has only one person who really, really likes it. And you can't support a whole industry just for that one person. But Generative AI can make you stuff on the fly, tailored to your exact tastes, and can keep learning from the things that you tell it to make that music better and better. So I think that is yet another experience in music that doesn't exist yet today that we are excited about eventually being able to tackle. So we talked about how the delivery mechanism inspires the creative process, but it almost feels like, based on what you're saying, that the creative process could inspire the delivery mechanism in the future.

27:07If music is so fundamentally easy to create in this near future and dynamic and personalized, maybe we don't listen to songs the way we do now, right? Where we just listen to the three-minute or even the one-minute song and then skip to the next one, right? It could be something else completely, entirely. A hundred percent. I think one of the big unknowns here, I'm a hundred percent sure you're right, and I'm also a hundred percent sure I don't know the answer to exactly what it looks like. Yeah, me too. I think one of the biggest question marks for us is ultimately what do the sharing dynamics and the social dynamics look like here?

Read the full transcript

27:41And so just to be clear about that, you know, right now when we stream music, there's a bunch of artists who make music. The rest of us just passively listen to it. We stream it. It's like very unidirectional. And there is some connection between artists and fans, but it's somewhat cursory. And then there are, and I would say that kind of resembles, let's say, the high end of Instagram, where you have people spending six figures making posts and they have a lot of followers and they make their livelihoods like that. But there's lots of things that are missing in music that exist in other modalities, whether it's the social graph that exists on Facebook or whether it's like that tail end of Instagram, where it's like someone with a locked account with 17 followers posting pictures of their kids.

28:26and that doesn't exist in music yet today and that can exist in music. And the exact nature of how these things get shared and liked and remixed and used for additional creative inspiration, we don't know those details, but they're gonna matter a lot in how these things ultimately get distributed. Yeah, it's fascinating. So what's gone on so far since you shifted to music? There's Discord, there's a website. I believe you also had a really exciting partnership with Microsoft that you announced recently. and maybe talk us through some of the highlights over the past couple of months. Yeah. So we started on Discord, and it was an amazing way.

29:04You know, we had a Discord from our open source work, and it was an amazing way to get our stuff out to the community and see how people liked it. In November, we released our first non-Discord experience web app, and people liked that even more. And I think since then, yeah, it's been super exciting. So the partnership with Microsoft is great. They've integrated our kind of free tier into Copilot, which is a set of lots of different creative experiences. And just like you may want to make an image as an outlet or for any other creative reason, you may want to do the same with a song. How did that come together?

29:43I mean, that's huge. Yeah, it's funny. I think I've heard at least rumors that somebody had shown the very thin web app to Satya over Thanksgiving. and he played with it and he really loved it. So obviously we were pretty stoked to have the opportunity to get our songs into Microsoft Copilot. Eventually that is an amazing way to spread the word about how enjoyable making music really is. We've just announced the little Valentine's Day experience. I think music is so important when we think about, let's say, soundtracking our lives and these different experiences that we can bring to people that are tailored to, for example, Valentine's Day and how do we get people making songs for one another and spreading love and spreading joy.

30:34Really, really exciting. People seem to like it a lot. So, yeah, I don't know. It's been a really fun couple months. And we build a product that makes people smile. Speaking of products, you know, so much of your background has been in machine learning, engineering, research. now in this role leading Suno, I imagine you're spending so much of your time on product, designing product, thinking about the product experience. How do you approach building product at Suno to connect users with this model that you've developed? Yeah, I'll caveat this with, I don't think I'm terribly good at product, but there's certainly people here who are.

31:14I think there's a few important principles that we think about. I'd say the first is, one thing we say is like aesthetics matter, which is to say it's very easy with a machine learning background to just be super obsessed with whatever quantitative metrics you have about how good your models are. And ultimately all of these quantitative metrics are flawed and they're especially flawed in music just because they're not very mature. And ultimately like these are the best judges of what's going on and so you need to listen to the stuff that is coming out of your model. And that is like a big driving ethos, both for the product and the whole company.

31:53Another big thing that we think about is just like, what are the workflows that people want to do that are intuitive for people to make music? and it can be tough actually with a lot of people at the company who really love music to

32:12tend to make things that resemble the workflows that professionals like to use to make music. And you know one thing like we always tell ourselves like my mom does not want to make music the same way a professional producer wants to make music. And for example I think once you've once you've talked about something like I want to stem this out and I want to take that kick drum and I want to take out the low end. Like, I promise you that's not how my mom thinks about music. It's gone too far. Yeah. It's gone too far. And so thinking about it very much from the perspective of here's somebody with very specific tastes in music who has never produced a bit of music in their lives.

32:46How do they intuitively think about this process? And either fortunately or unfortunately, it means that we end up building things that are often not suited for professional musicians. And it's just kind of not the goal right now. The last thing I'll say is that it is also very easy with a lot of academics to think that you can just first principles reason your way toward what is the best experience here. And I think this is an empirical science where, let's say, because we have backgrounds in music, we actually can't intuitive the way novices want to use it. And so we need to run a lot of experiments and we need to ship a lot of features and we need to see how people like them.

33:33And then we can kind of iterate our way there. Speaking of product and the product experience, you talked earlier about how now it's, you know, you can you can bring your own lyrics. You can just describe it. But it sounds like it's mostly coming from, you know, tech, a text to song approach. as you think about the product's evolution, what are the other ways in which you might enable people to create music outside of just text? Yeah, it's a really good question. I'll say before I answer it, I'll just say like, I'll be unkind to machine learning people because I am a machine learning person. Like, you know, people will think like, oh, this is a text to music model.

34:08And I think that is an extremely non-user centric way of thinking about what we're doing. And if all you can do is think about A to B, text to music, or tapping on the table to music, or it's like, we will never build something people want to use. And so we really try to think hard about it from the other end and think about like, what inspires people to want to make music? And how can we get them to express that? And so yes, describing stuff is great, but not everybody is in the habit of writing lyrics. Maybe there are other ways of describing things, whether that's tapping your pencil on the table and recording that.

34:43Or maybe you sang into your microphone and for some inspiration, I think maybe you can come with your own samples of music and kind of describe why you like them and start that as a process, kind of like you might do mood boarding in Pinterest. There's lots of stuff there. I think one of the workflows that we are really into is what we call soundtracking your life. And it's like, what are all of the random things that are happening to you today and how are you going to kind of show those to a model as inspiration for the sounds that are in your head and maybe it's a really loud car horn you know like whatever it is so we try to keep a pretty open mind there we've got a lot of stuff coming that I don't want to that I that I don't want to tease out just yet but I think I think we are very cognizant of the fact that text to music is going to be limiting.

35:38It's really, really cool. And then on the output side, you know, you've clearly made some very intentional choices, obviously, in terms of the quality, which is phenomenal, but also on the limitations, right? I believe, you know, you make it very difficult, if not impossible for anyone to make anything that may be infringing of existing artists' rights. Maybe talk a little bit about that. We always said we were going to do audio the right way and technically that meant kind of the foundation model approach and then you know ethically and legally that means not trying to infringe on existing artists and trying to do this in a way that is artist friendly and this is not only because it's ethical and legal and moral it's because it's also what we think is the future of how people want to do music it's that other 49 50th that we were talking about before of people aren't going to lose their connections to artists but there's just a lot of other experiences that we can do and so we don't let you say make me a Taylor Swift song actually our models don't even know who Taylor Swift is but instead of actually making something up that's just not we try to gently nudge users into the behavior that we think is the long-lasting one, which is making the original music that's in your head.

36:59And the same for other people's existing lyrics. And one thing that is actually nice for us is there's a lot of these cover-type things where maybe you could have Taylor Swift doing Enter Sandman by Metallica. And I think those tend to go very viral, and I think that is really not the future of how people want to do music. And so we're very happy to kind of let those things go viral and then kind of evaporate somewhat quickly. And I think the analogy that I have with your platform, exactly, exactly. So you can't do that with our platform. You can do that with someone else's. It is almost certainly illegal.

37:38It is very viral. And I don't think that is the long lasting use case. I think the analogy that I have in my head for making covers for Taylor Swift covering Enter Sandman is the first time you played with GPT and you made a Shakespeare sonnet about drinking coffee. And then you made another one. And then the third time you were like, this is cute, but this is not what I want to do. And that doesn't mean that GPT isn't extremely cool and useful because it is. It's just not for writing parody Shakespeare sonnets. The comparison to GPT makes me think, you know, one of the things that people knock it about is that, you know, it sort of, when you try to get it to do things that are really creative, it sort of hits this kind of upper bounds.

38:18And maybe that's because it's, you know, it's a it's you know, it's it's training on the world's knowledge. So it's sort of it bumps up against the limits of that knowledge and that information. Thinking about music, though, music is obviously so creative. And if you think about like the greatest artists ever, so many of them became the greatest artists ever because they did something that no one had ever done before. Right. Right. And so do you imagine a world in which Suno and maybe just music models in general can do things that have never been done before? Or is it always going to kind of hit up against the upper limits of everything that came before?

38:58Yeah, I think we can make music that is kind of above the limits of what we've seen right now. And I think that for a few reasons. One is, you know, humans keep making music that is above the limits that humans continue to see. And so they are, you know, standing on the shoulders of giants. You get inspiration from everything, not just from other musicians, but you get inspiration from the car horn that honked outside of your window. And there's a lot of like new ways that humans can be inspired by these models as well. So am I using autotune when I record my album? Like that did something that was impossible for a human to do to sing that perfectly.

39:37And then the tool let me do it. or that I use an effects pedal on my guitar. That was impossible to pull out of my amp, but that thing did it. And so I think this is kind of a natural progression. And I don't know where the line of which constituted AI and which was not, but I think that's kind of always been a part of this. And I'll tell you one thing that I'm actually pretty excited about. If you think about the progression of music, more recently progress in music looks like like things that are sonically interesting so like interesting sounds but not necessarily more interesting chord changes and I think AI has the potential to kind of bring that back where you do more interesting stuff melodically and harmonically it also obviously sonically but I think it is a way to like it is a way for humans to express the sounds that are in their heads.

40:34And sometimes that there's a block that kind of artificially sets the bar for humans too low. You know, one of the things that's come up on this podcast a couple of times is this notion of training and training on the entirety of the internet. Definitely feels like it's a little bit of a Wild West right now. You know, the OpenAI New York Times case comes to mind and everything, you know, everything that's being discussed there. Like, how do you think this shakes out? Are there changes coming? And what will that mean for the future of these large foundational models? Yeah, I think there definitely are changes that are going to come.

41:12I think some will be across the modalities. Others may be modality specific. I think music is particularly interesting because there are just much more established ways of interacting with rights holders. And just for example, on the, let's take that New York Times OpenAI lawsuit where the New York Times is upset that you can get an entire New York Times article at a GPT with the right prompt. there are you know if I have a restaurant and I stream some song there are like very obvious ways that are established that I can compensate the owners of that song the master and the songwriter and that doesn't exist in text and so I think it's not going to be blanket solutions that apply across the whole industry I think some things will and some things won't the thing that's going to be really tricky here I think is actually geographical distance geographical differences between, let's say, different countries.

42:13And I'm honestly a little bit worried about how things will shake out. Is OpenAI going to have to have different models in the EU and in the US and stuff like that? And just that that may get very difficult and kind of prone to, I don't know, hacking and people are going to be using like VPNs to access different countries' models and stuff like that. I think it's really, really early. I think people think this is going to get sorted out in the next six months. And I don't even think the OpenAI lawsuit won't get sorted out in the next six months, but that won't even sort everything out for the whole industry.

42:50So for us as Suno, especially with music being somewhat farther behind images and text, we watch this stuff very closely. We try to always do the legal and moral and ethical thing, and these things aren't always the same. Something that's legal may still be not artist-friendly. The thing that we stress is it's like super early and I don't think there are any scenarios that are amazing for us or terrible for us that I can foresee in the next year. It seems like OpenAI is doing licensing deals with media publishers. And you have to imagine a world in which those media publishers obviously then benefit in some way.

43:31Maybe, you know, maybe it helps drive subscriptions or I don't know, maybe there's a new version of these models where, you know, you can use certain ones that output New York Times and certain ones that don't. Do you do you imagine similar types of deals happening in music? Can you can you imagine a world in which I know we keep mentioning Taylor Swift, the Taylor Swift's of the world are collaborating directly with music models to create, you know, generative Taylor Swift music? Yeah, 100 percent. So, you know, Google is starting to do this with their Lyria project and they have a few big artists on board there.

44:06But, you know, if we think about it, like, let's fast forward a few years and the licensing climate is a little less uncertain. And maybe we can let you prompt a model with a Taylor Swift song. And like, this was the big inspiration. And there's something that is akin to the way people pay out for sampling now, but it's obviously a little bit different. and then all you had to do, Mike, when you wanted to make a Taylor Swift inspired song was like, pay for that sample. And I don't know how much it will be. And I don't know how much will actually go to Taylor Swift, but like, we will be able to let you do that.

44:42And we can actually do that right now. We just can't let you do that because we don't know how to pay Taylor Swift for it. Right. Right. So it's technically possible. It's just not the infrastructure for that from a royalty perspective is not yet developed. Exactly. Exactly. What's next for Suno? Where, Where do you go from here? What are you looking at over the next couple of quarters? Yeah, a lot of exciting stuff. So we've got a new model that we're going to release soon. It is kind of better in all of the ways that we think about, so we're really, really excited about it. Just can go tactical for a second.

45:14When we think about models' qualities, we think about audio fidelity. Like, does it sound like it was crisply recorded with beautiful hardware? We think about song quality. Is it catchy? Does it make me feel? and we think about controllability. When I asked for something, does that something come out? And because music is still so early, we're still making gains across all of these, so really, really exciting. Lots of new ways that we want people to interact with stuff, so whether that is different ways of prompting the model, whether that's different interfaces into our product, we're really excited about all of those.

45:54yeah those are those are the big ones there's there's a couple of other secret things coming down the pike that that we're not ready to talk about just yet but I think you know the the future the future of music is really really big and I think you know people think a lot about how streaming almost killed the music industry and people think that AI will almost kill the music industry and it's I think it's really really the opposite I think we we are quite confident that there's another 49 50th of music that we among others are going to try to uncover. Always got to ask, is Suno hiring? We're always hiring.

46:28So we're in Cambridge, Massachusetts, kind of a fun place for building companies. It's something we've done here before. Always looking for the best talent across software and machine learning and product and design and music. If you really love building stuff and you really love music, this is probably a good home for you. Mikey, thank you so much. This has been so much fun. We really appreciate all the time. Awesome. Thanks so much. This was great. Happy to be here. Let's do it again sometime. See ya. Thanks, Mike. Thank you so much for listening to Generative Now. If you liked what you heard, please rate and review the episode.

47:07That really does help. And of course, subscribe to the podcast on platforms like Spotify, YouTube, and Apple podcasts. If you'd like to learn more, you can follow us at LightspeedVP on YouTube, Twitter, LinkedIn, Instagram, and Generative Now is produced by Lightspeed in partnership with Pod People. I am Michael Magnano, and we will be back next week with another awesome conversation. Thank you so much.

From the publisher

Most of the hype around AI has revolved around its text capabilities, and the powers of LLMs like ChatGPT. But Suno is focusing on the unsung hero of AI - music. Suno is building a future where anyone can make good music - all you need to do is type in a prompt, and out will come a song that’s never existed before. This week, we’re revisiting a conversation with Mikey Shulman, CEO and Co-Founder of Suno. He joined Lightspeed Partner and host Michael Mignano earlier this year to talk through the intricacies of programming for sound, and what this technology could mean for music.


Episode Chapters

(00:00) Introduction to Mikey Shulman and Suno

(04:03) How transcribing S&P earnings calls inspired Suno

(08:41) There’s no Common Crawl for audio - they had to make their own

(12:37) Hacking text-to-speech to make music

(16:25) What’s the product market fit for generative music? 

(21:15) How will AI change the format of music?

(28:44) Suno’s highlight reel so far

(31:39) Designing with the end user in mind

(38:26) Can AI transcend the creativity ceiling? 

(40:48) How does Mikey think regulation and music rights will shake out?

(46:19) Is Suno hiring?


Stay in touch:

The content here does not constitute tax, legal, business or investment advice or an offer to provide such advice, should not be construed as advocating the purchase or sale of any security or investment or a recommendation of any company, and is not an offer, or solicitation of an offer, for the purchase or sale of any security or investment product. For more details please see lsvp.com/legal.

More from Generative Now | AI Builders on Creating the Future

All 90 episodes
Mikey Shulman: Suno and the Sound of AI Music (Encore)Generative Now | AI Builders on Creating the Future · 48 min
Listen in VO