In short
Podcast Summary: The Future of Sound: Udio’s Vision for AI-Generated Music | E2016
Episode Details
- Podcast Title: This Week in Startups
- Episode Title: The Future of Sound: Udio’s Vision for AI-Generated Music
- Air Date: [Insert Date]
- Host: Alex Wilhelm
- Guest: David Ding, Co-founder and CEO of Udio
Overview In this episode of *This Week in Startups*, host Alex Wilhelm interviews David Ding, the co-founder and CEO of Udio, a startup focused on AI-generated music. They discuss Udio's inception, the advancements in AI music creation, and the evolution of the Udio AI model. The conversation also features a live demonstration of Udio's capabilities.
Key Topics Covered
- Background and Inception of Udio
- David Ding's Background:
- Former researcher at DeepMind.
- Combines passions for technology and music.
- Foundation of Udio:
- Founded November 2023 with a vision to help artists create music using AI.
- The idea emerged from the success of AI in generating art and text, prompting Ding to explore music.
- AI Music Generation
- Functionality of Udio:
- Udio's AI models learn music generation by analyzing vast amounts of musical data.
- Understanding music theory, structure, and recording techniques is key to the model’s training.
- User Control:
- Current limitations include control over time signatures and specific musical elements.
- Future releases aim to enhance user input options for more precise control over music generation.
- Development and User Experience
- Model Evolution:
- Early versions struggled with certain functionalities but improved significantly with each iteration.
- Udio version 1.5 introduced features like key control and global language support.
- User Engagement:
- Aimed at music lovers and creators, facilitating an easier music creation process.
- Encourages collaborative music experimentation among users via platforms like Discord.
- Financial Aspects and Business Model
- Funding:
- Udio secured $10 million in funding, highlighting the importance of financial discipline in operational growth.
- Cost Efficiency:
- Focus on reducing operational costs associated with AI music generation through optimizing compute power and using cost-effective cloud services.
- Market Positioning and Future Outlook
- Competitors:
- Udio positions itself favorably against competitors like Suno, focusing on high-quality tools for music creation.
- Concerns of Replacement:
- Ding emphasizes that Udio serves as an enhancement tool for musicians rather than a replacement, suggesting a parallel evolution of traditional and AI-generated music.
Live Demonstration Highlights
- The episode features a live demo of Udio:
- Demonstrates lyric generation and music creation based on user prompts.
- Showcases the generation of a jazzy neo-noir rap song about dinosaurs.
Conclusion The conversation underscores the potential of AI in revolutionizing music creation while addressing concerns surrounding its impact on traditional musicianship. David Ding expresses optimism about Udio's trajectory, and the episode concludes with a call for creativity and exploration in music using AI tools.
Key Takeaways
- Udio leverages AI to assist musicians in the creative process.
- User feedback and collaborative communities play a significant role in Udio's growth.
- Continuous improvement and feature updates are pivotal for maintaining user engagement and satisfaction.
Further Information For more details, visit [Udio's Official Website](https://www.udio.com) and check out their latest features and updates. Follow David Ding on [X](https://x.com/daviddingai) and LinkedIn for insights into the future of AI in music.
---
This summary encapsulates the main discussions, insights, and developments presented during the podcast episode, providing a comprehensive overview of Udio's mission and the evolving landscape of AI-generated music.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Hey everybody, welcome back to Twist. This is Alex and we have a special interview for you today. I am an enormous fan of music. You may not know it, but I grew up playing classical and jazz. trumpet throughout my youth, and music has remained an absolute huge passion of mine throughout, really, my entire life. So when the AI revolution of the last couple years came to the world of music, I was incredibly curious. Two companies have really caught our eye here on This Week in Startups, UDO and, of course, its competitor, Suno. Today, we have David Ding, the co-founder and CEO of UDO on the show, to tell us what it's for, who's paying for it, and where AI-based music creation is going.
0:39This Week in Startups is brought to you by .techdomains. Don't miss our Jam with JCal contest. To apply and get more details, go to jamwithjcal.tech, brought to you by.techdomains. LinkedIn ads. To redeem a$100 LinkedIn ad credit and launch your first campaign, go to linkedin.com slash thisweekinstartups. And brave. If you're building AI and search-based applications, train your models with the Brave Search API. Get started for free at brave.com slash jason. We're going to talk about AI and music. David, hi, how are you? And welcome to the show. Hi, hello. Really excited to be here. So I want to start with some background stuff because I know you were at DeepMind for a while and UDO is a relatively young company.
1:29I think it was founded in 2023. so just give us what was the moment in time in which you said i have to leave where i am and go found this company because why yeah uh sure so yeah as you said uh udo was founded last year november of 2023 and before that i was a researcher at deep mind and uh throughout my entire childhood i've always been interested in two things primarily so one is technology and the other is music so as a kid, I always wanted to build computers that can simulate the way a human brain works. Maybe you can wire the neurons together and then try to model the brain. And then it turns out that when I went to college, this thing was starting to pick up traction.
2:15My first year of college, I took a machine learning course so that I can participate in this field. And my other passion growing up was music so i played classical piano and i played it for at least like 10 years like growing up uh before going to college and uh i always thought that would be really really cool if uh like a computer technology could compose and make music and so then fast forward uh to when i was working at deep mind generation modeling really took off you see technology like ChatGPT or DALI or MidJourney like Emerge that really revolutionized the way that computers can make art.
2:59And so at that point in time, I was like, hmm, what happens if we apply the same technology that I've been learning how to build and apply it to music to have a machine that can help people create music, ideate and create songs? And so this is why we left DeepMind to create a company that produces a product to help artists and soundwriters turn their ideas into reality. So I want to go back to the point about LLMs to image generation to music generation, because I mean, my day job is writing. So that's kind of what I know the best. And so to me, the idea of a large language model taking in a lot of data, and then helping kind of do next word prediction, admittedly, it's more complicated than that.
3:44But I can really understand it. I kind of get how we can use LLMs to do image generation. but when we expand the the the work done to music to me i i feel like i'm missing a link in in how the technology actually functions so without spilling any secret sauce if you will how does you know the the ai models that i best understand end up creating tunes because i it just seems to be like a like a real stretch of what was possible but it clearly works yeah so similarly to how um these large models learn how to produce images and text all models they learn how to produce music by listening to lots of examples of music so you listen to music and it tries to synthesize the common elements across music so like a music theory elements of music theory like which chords follow which other chords or how rhythm interacts with the overall structure of the song as well as other elements what does it mean to be country music which is rock music or how does a guitar string vibrate and how does the sound of a piano echo and vibrate around the room and finally how does this all interact with the recording technology how do you turn this sound and turn it into stereo and so this model because it's trained on the final output music and learns how to do everything.
5:14So from like the very fundamental music theory level all the way to how the sound is recorded by the microphone. Okay, so prepping for our chat today, I was playing around with you, Dio. By the way, I am now your most recent paying customer. Shout out. And I decided to throw a curveball at your software. And I said, okay, look, I wanted to do a progressive metal song that sounds a little bit like Periphery, a band that I love. but I'm like, look, let's do it in 6-8 time. Now, you're a classical train pianist. I'm a classical train trumpet player. You and I know that when it comes to time signatures in the world of music, 6-8 is not very complicated, right?
5:52We're not doing like 11-8 or something crazy. You count in sixes and then fives. This is pretty simple. And it kind of did it, but not perfectly. And I know this technology is still improving, so I'm not trying to be negative. But is this going to a direction in which i could tell a service like yudio like i want to do a song first half in six eight second half in seven eight and i want to do a chord change from c to c major and then like how specific can we get and then is that underpinned based on a very granular understanding of how music is put together or does the software better understand like broader chunks of it versus like down to the individual note level yeah so this is definitely a direction that we do want to support giving um users and musicians more ways of controlling the model like time signature key tempo bpm or instrumentation uh or even like dynamic levels like uh like start quiet uh start swirling and then like um uh and then uh die down again so um this is something that we definitely want to support a time signature is something that we do not support at the moment because because in our music, when we're training our models, we did not teach it the concept of a time signature when we were annotating the data.
7:12However, a key signature, the key of the song is something that we do support. And this is something that we did not support when we launched the model, version one, back in April. But it's something that we added in July because we recognize that users want to be able to control the key. And so then we annotated our dataset set to contain um oh this is d major this is a c sharp minor so that um now when you go to edio and you specify a key a minor it will produce a song in that key well a minor is now the most famous key in i think all of music things do mr kendrick lamont if you don't get that reference congrats for being offline for the last three months of music history wow this jam with jcal contest has been a blast so far i've had the opportunity to meet with four great founders from companies like CorePod, Ulama, Uptrends AI, and the Roam app, all because they all use.tech domains.
8:05And we have room for one more. Do you want to come on the pod and tell me what you're building? Well, you only need two things to enter. You got to be a founder with under 2 million in funding, and you got to have one of those awesome.tech domains. So head to jamwithjacowl.tech and tell me what you're building. And if you win, I will invite you onto this week in startups, and you'll get to share your vision with me and the world. I'm working with.tech domains because killer startups use them. You know, 1x.tech, Rabbit.tech, so many others. And guess what? We use it too. That's right. .tech powers our Founder Friday program.
8:38So tell me about your awesome.tech domain and startup. Apply for the Jam with JCal contest today at jamwithjcal.tech. We're picking the final winner soon. Okay, so it sounds like what I did there was I asked UDO to do something that it doesn't do quite yet, which is probably why it got a little bit funky. But you said something interesting there, which is data annotation. And that, I think, is the thing that I was missing, because it sounds like you guys, the human label, and help it understand, like, this is a rock drumbeat in 4.4. So does that create, like, a flag that then the software or the model can kind of go back to and, like, point to and understand?
9:16Yes. So by annotating the data in a training data set, you teach the model how to associate certain descriptive words with musical elements. So then it sees 3-4, the time signature, and it hears a song that's in 3-4. And it knows, oh, 3-4, it means you have three beats. And then the first beat is emphasized. When a user then asks the model to create 3-4 music, it can then like take its understanding of 3-4 and apply it to the composition of the song it's kind of like when a human learns um if you never teach a human oh this is 3-4 you can't ask the human hey create me 3-4 music even if it can produce 3-4 music it just doesn't know what 3-4 the words actually mean so it sounds like the data annotation then provides almost like a a connective layer between music and the user's request and kind of helps natural language inputs translate to something the computer can understand as a command prompt, essentially.
10:20Yes, exactly. And we aim to improve our model by giving it more annotations to understand more elements of music so that the model can produce these elements upon command. Okay, so I want to go back in time, though, because I've been playing with UDO since, and this is a true story, one of my friends started sending us funny songs he made for us in the group chat. And they were whimsical things like, Alex doesn't want to go to work tomorrow and stuff like that. And I was like, okay, where is this coming from? And it was from you guys. And so I got a kind of early look at the software and I've made different songs and I've gotten to play with the new model some.
10:57But back in the beginning, David, when you were first getting the 0.1 version of this out, one how good or bad was it and and how how easy was it to get from like proof of concept if you will to something you were confident that people might want to actually use yeah so uh funny that you mentioned like the uh the first version of our model uh like the baby uh the very very baby version when we were still debugging our overall uh code base and training structure we spent a couple weeks trying to figure out why our model couldn't produce any lyrics. You provide lyrics and the model just refuses to sing the lyrics.
11:37And then we spent a while looking at the model, analyzing different loss curves. And then eventually we found the reason to be quite simple is that when we were feeding the data set to the model, there was some kind of bug that cause the lyrics to not appear. And so the model never saw the lyrics and so therefore it couldn't possibly know how to turn the lyrics into a song. So essentially it couldn't run the engine of lyrics because there were no words going in. Exactly, yeah. And so this really goes to show how the process is quite dependent. You had to pay attention to detail and it's all about the input data.
12:20And so we fixed the bug and then after we fixed the bug, the model actually just kind of took off. Every week, we saw improvements. The first week, it probably knows the broad genres like rock versus jazz. As model training progressed, it started learning more specific keywords like energetic, what's hard rock, what is smooth jazz. Yes. And also the sound quality improved, starting from something that sounds like very noisy to something that's much more refined and more like what you would get from a studio. Yeah, no, the actual fidelity is pretty good in my experience. And one thing as a fan of heavier music in general, there are certain heavy metal subgenres that depend a lot on orchestral additions that are mostly programmed.
13:10And so I'm familiar with like the current state of the art for studio music with added digital elements, if you will. And we're not that far off from this just with UDO's own creations offer. So that's very exciting. But it sounds like from the point of inception of the initial like it works to go into market to 1.5 released in July is pretty quick. And that chart's going up in here's a model quality and fidelity and so forth. Do you think that trajectory continues for a long time? Or were there early winnings, David, that let you improve faster than you might be able to now and in the future?
13:47So obviously, there is a point where you start from zero and you get something. And so that's the biggest delta. And as you observed, our audio quality is actually pretty good, although there are still areas for improvement, which we are working towards. but the big focus going forward is additional controls for users so giving people um like more ways of controlling the music like um maybe you want to provide um uh like a guide like you had this a melodic line already and you want the model to follow this melodic line and and add musical elements to it or maybe you have this like musical style but you don't really know how to describe it using words so how would you um take this musical style synthesize it and um feed it as an example for the model to follow and so we want to enable these additional controls because we recognize that music creation the creator wants to have a lot of control over the music because that it's their own creation right and so um uh so that's the area that we really want to focus on going forward okay i want to i want to do some demos in a little bit to show people what we're talking about because you and I have used this, of course, a lot and they might not have.
14:59But one thing that I was thinking about is who this is for, because I am an enormous music fan. So to me, music is part of my day from kind of when I get up to when I go to bed. I'm either listening to audio books or music, right? And so to me, it's very personal, very important. And I know music theory and I love it. And it's key to me. Not everyone's like that. And so people have different music tastes, different consumption habits. And so I don't know, is UDO aimed for folks that want to create stuff for their own consumption? Is it more of a rough draft machine for artists as they explore new ideas?
15:35Is it a way to generate Muzak for elevators? So I guess kind of like who do you think this is for now and in the future? So we think that UDO is for people who love music, people like yourself, and also like artists and some writers who obviously love music as well. We just want to create a tool to allow, to make music creation a lot easier than before. Kind of like other tools that have come, that came before in the past. I've done that for like DAWs, sampling, drum machines. These are all innovations that turned something that was a little bit harder before and with the aid of new technology, just making this creation process easier so that more people can participate in the creation process and that existing artists can leverage this to try out ideas at a faster pace and and come out with our like music that incorporates these elements in creative ways that maybe even the creators of the technology never had in mind i think like one good example for this is auto team like when auto team came out you know a lot of people are they had qualms about using it it's like oh it's like cheapening the experience using it that's a play way of saying it yes but sorry keep going yeah like people were like oh this thing is like cheapening the uh the experience like it allows like people who can't sing to to sing and that's a bad thing but then like you know like um what really happened was like you know it like um really transformed uh the industry like people were using it and then people found ways of using it very very creatively like you know like bumping it up beyond the like the spectrum and embracing the art-witting sound as a musical style, right?
17:16And so we think that with these technologies, it makes music creation easier, and people will find creative ways of using it. Okay, so it sounds like for someone like me, a big music fan, I could use it to create fun things for myself to listen to. If I was a musician, I could use it to expand ideas and give me new ideas. But this doesn't replace you know, I don't know, my spouse's spotify account at some point in time this is more like distinct acts of creation in the future versus passive consumption exactly so uh i mean uh you might you say that you play a trumpet right uh like your trumpet doesn't uh replace uh listening to like uh like great trumpet players of the past uh on spotify right because you enjoy uh listening to music that other people create but you also want the like the joy of creating music yourself yeah no i i think that's right and what i what i like about the idea behind taking modern ai techniques and applying them to music is it just allows a lot more people to do stuff uh you know david like five years ago people talked a lot about low code and no code and there was this big chat about the democratization of software development and that's kind of worked out but i love the idea of more power to more people and this to me seems to fit into that now on the critical side though some musicians are worried that they're going to be replaced whole cloth or diminished in some way.
18:42I want to run my theory past you, which is that I don't think that's going to happen because the musicians that I love to listen to have their own very specific, sometimes experimental style that probably couldn't be replicated by even very intelligent models. So to me, this exists, if you will, side by side with kind of how music is made today. I'm curious if that's your view as well yeah so we so that's definitely my view as well uh we we so i believe that people will continue making music the way uh they've always made music and this is simply another tool in a toolkit that they can choose to use or they don't have to use it but then it just um something additional right like uh just because like electric guitar got invented doesn't mean that the acoustic guitar got completely replaced right it just just means that like there's yet another instrument that you can add onto your band yeah actually i remember um in my high school jazz band we had a song i think it was an old buddy rich tune and um it had a little bit of guitar by itself and our guitar player played electric and one time he forgot to turn his guitar up so we got to that part of the song and he played and no sound came out and my our director was like well you know he was a trumpet player and he was making fun of the electric guitar for needing you know help essentially and i was like i don't know that seems a little bit old-fashioned but this probably fits somewhere in there um i do want to ask a quick question though about uh where udo will come up because you mentioned daws or digital audio workstations earlier um very much now a well-known kind of entity in the musical world does udo ever become part of one of those a plug-in a an api that i can call does it leave the website and end up somewhere else?
20:26Quite possibly. We think a lot of our power users, they use Edo to come up with ideas, and then they download the individual stems, which is a feature that we allow. So people can download stems and then load the stems up in their DAW for further post-processing. So essentially, they take the draft, bring the stems over, and then you can do... That's pretty cool. Is it hard to do individual stems? because that implies that the model is making a collection of tracks that are then mixed together. Is that how it's always been, or is that a new change to how the underlying model works? So the underlying model always produces the fully mixed track.
21:12But then recently, with Vision 1.5, we added the ability for users to download the individual stems, which are separated post hoc from the mixture. Oh, post hoc. Interesting. Yeah. Oh, okay. So you, so it creates something that's mixed and then you isolate. I would have thought this the other way around, but that's why we ask questions. Yeah. Okay. So, uh, before we talk about money, stems are the individual tracks inside of a song for, for example, bass or guitar or piano or whatever. Um, I just want to make sure that everyone listening understands stems, David, is that how you would define them as well?
21:48Yep. Okay, cool. So if you don't know what stems are, now you do. Okay, there are more than 50 ,000 venture-backed startups in the United States alone. This means marketing has to be perfectly targeted. You got a lot of competition out there, or you're just going to fade into the background and your money will go with it. All your ad spend will be for naught. You got to make sure you target the right prospects. So how are you going to do that, especially in a business-to-business context? Well, the answer is obviously LinkedIn ads where you can precisely reach the professionals who are likely to find your ad relevant.
22:21Just think about it. Wouldn't it be great to target your ads by the job title or the industry, the location that that company is, you know, and maybe even a very specific company, maybe you got a list of 20 lighthouse customers that you want to bear hug, that you want them to know about your product or services. LinkedIn ads is going to help you do that by building relationship and driving results, LinkedIn is the environment where people are receptive to business. They're not there for food or politics or entertainment or music. They're there to do business. A billion members, 130 million of them are decision makers, and 10 million of them are C-level executives.
23:00So start converting your B2B audience into high quality leads today. Get$100 from your boy Jay Cow, linkedin.com slash this week in startups to claim that credit. Again, LinkedIn.com slash this week in startups, no spaces, no dashes, terms and conditions. Why? Because they're giving you a hundy. Okay. So UDO raised$10 million. That was earlier this year. And Jason Horowitz was in there, a number of artists, including the producer, Tay Keith, and love to see venture capital funds. Glad you guys raised some money. But my thought is this, I currently pay you$10 a month to use something like 1 ,200 song creation credits.
23:41I look at that and I know how much AI costs to run. People talk a lot about that. I feel like I'm burning through your bank account. So is it as expensive as I imagine it is to run the model to create music? Because it sounds very compute intensive. Yeah, obviously it's a balance for us. we want to make sure the price is set at a point where we allow people who are curious about the technology to try it out while being able to like make this process sustainable so um so we chose a price in a way to basically allow for this like to be able to sustain usage while um while not being like very expensive and going down the road we definitely want to optimize our models make them more efficient so that they can run at a cheaper cost because we want to maintain this commitment to users but we also want to run a sustainable business yeah so we've seen this with just to pick one example out there open ai's gpt family of models um when 4.0 came out it was one cost and then it's come down i think and we've seen that pretty frequently does that mean that you guys are able to extract a lot of efficiency from the underlying model and that this should get much cheaper to run over time?
24:58Or is there less low-hanging fruit because it's doing music, which is just, to me, seems harder than doing text? We think that there is a lot of room for improvement for sure. I wouldn't comment on whether or not it's on the same scale as OpenAI. OpenAI obviously has entire teams of incredibly talented engineers working on this, and we are a much smaller company. But we do believe that there are similar levels of efficiency gains to be had. Okay, so essentially, yes, you know, Udeo is a smaller company. I think OpenAI has over 1 ,700 people now, but with work, a similar ish curve. Okay. That's actually a really good question for me to ask.
25:40How big is the company today? What's your current staff size? So we currently have about 17 people. So we've grown quite a bit. How many people did you have before? when we when we launched the company uh or when we launched um our model back in april we only had eight people eight oh god yeah um and just because this is a startup show let's do some basics remote hybrid or in office um mostly uh mostly in the office but some uh some working remote okay and um just thinking about staffing for the rest of the year are you going to keep hiring as aggressively as you have or will that slow down that you've more than doubled in size uh we'll probably stay a little bit more steady okay so i know you guys raised uh from andresen i mean adventure firm that everyone watching this show knows and you guys raised 10 milli is that enough money david because some of your competitors have raised more and we are in the era right now of companies that use ai raising lots of money let's say so i'm just kind of curious why why 10 million was the number and also you know how soon are you going to be back on the show telling me about your shiny new round yeah so uh when we started we raised 10 million because we want to be disciplined in how we um spend the money we believe that like um there's some amount of truth in the idea that scarcity produces innovation and uh and so uh or we try to be super efficient in a way that we use our capital to develop our models so tell me tell me about that because you know developing a model i mean people talk about how models are eventually going to cost like a billion dollars to put together but that's for a very general purpose model and so forth so for for you guys how do you ensure that your capital expenditures on model creation and improvements are cash efficient yeah so one thing that we do is to is try to secure a cheapest compute power that's available like the the chip that's cheapest in terms of floating point operations per second versus dollars and so we ended up like choosing google cloud's gpus which we identified as offering significant savings over other chips like NVIDIA GPUs.
28:05Google's startup cloud program is one of the occasional sponsors of this show, so I just want to point out that no one asked him to say that. That was off the cuff, but there you go, so we're not being biased. Just to put it into perspective for me, though, because I don't get to go to those negotiations, how much cheaper was GCP for UDO compared to competing providers? Was it a lot cheaper, or was it more of a marginal differential. Yeah, I'm not sure if I should comment on specific numbers, but it is... Oh, you should. David, you definitely should. You should drop all the numbers you can right now.
28:38Yeah, but it's definitely quite a bit cheaper. And so that's one factor. And the other factor is that we have quite a few really talented modeling research scientists among our co-founders. And because they have a lot of experience training these really big, adjunct models they know how to make maximal use of the available hardware how to create really efficient programs and how to like design architectures that can train efficiently so if you're doing that work though because that's that's nitty-gritty stuff if you're doing all that already why not buy your own h100s or equivalent and and just run your own mini data center it It seems to me like if compute is going to be such a core element of what makes the digital brain that you use, why not own the neurons themselves?
29:33I guess for us, as a startup, we didn't really want to deal with the logistics of running our own data center. And we thought it would be simply to go with a cloud option. Do you see the company in, let's say, five years from now, just looking down the road so far that I know we're making kind of almost like a joke here. But do you think that you'll still be on a major public cloud provider in five years? Or do you eventually off-ramp when you have more money and staff and so forth and do your own data crunching? Yeah, it's hard to say. The cost of computation on clouds has been going down. I think there's some kind of new law, like Kwan's Law or something, that supplanted Moore's Law about the cost of GPUs over the years.
30:19the cost of a cloud company could go down like very significantly. And it's very hard to predict five years down the line, whether or not it will be more economical to buy your own chips or to lease them from the cloud. Okay, so Huang's Law, by the way, this was, of course, a reference to Jensen from NVIDIA, I presume? Yeah. Yeah, okay. So if you know NVIDIA, you know the guy. Huang's Law is, and I'm Wikipedia-ing this live, so this is not very lettered of me, but it's a general idea that as Moore's Law predicted that the number of transistors would double about every two years. Huang's law is that GPUs will more than double their performance every two years.
30:57So it's essentially an acceleration or a faster version of Moore's law for GPUs. That speaks very well for you guys, because that means that your gross margins should improve over time, just naturally as chip companies make better chips. That's kind of cool. that's a tailwind for you as a ceo yeah it's definitely something that we're very excited about uh like cheaper compute making even more powerful technologies are possible all right are you building the next great ai product well if you're doing that you know how expensive all these apis can be for model training data obviously and training ai is very expensive that's a fact we all know that so you have to try brave's new search api yes i'm talking about brave the privacy browser that I use every day and on my mobile phone, Brave's browser has 65 million users.
31:51And that drives a lot of data into the Brave search engine, which is the only global scale independent search index outside of big tech. And that index is available to anyone with a Brave search API. So you're going to be able to use the Brave search API to power your chatbot or train models, inform answers to real-time queries, and serve images, web results, even rich tech snippets. The Brave Search API features an easy to use intuitive data structure. So you're going to be able to get things done quickly. And its data is populated by real human interaction, not web crawlers, that's critical.
32:26And it's all done at a fraction of the cost of the major players free for up to 2000 queries per month. So you can try it on, play with it really sort of brainstorming and then plan sort of as little as$3 CPM. So here you go. If you're building next gen ai apps or chat box you've got to try the brave search api get started today at brave.com slash jason on the public cloud front i'm going to not ask about an individual provider because i don't want you to get in trouble but i have heard that there is a capacity crunch out there that there's not enough total gpu based compute for people that want it has has you do any issues getting the amount of compute capacity that it needs at any point yeah i mean um it's always a balance where eventually we'll be able to get the compute, but definitely at times it just takes a while for different cloud providers to be able to find the chips that are available.
33:21Okay. Now, I want to go from there to a demo so everyone can see the product that we're talking about from a compute perspective. So David, we drew straws before and you're going to drive because you told me that you have some new stuff to show off. So let's pull up Udeo. If you're watching this on youtube you can see what we're doing live if you're watching listening to this on spotify or apple podcast we will narrate as best we can but we are going to do a little bit of testing around here to show off what we can pull off so david uh talk to me what are you showing me yeah so i'm showing you um the udl create page it's a dedicated um you can think of it like a creation studio where you have a list of your recent creations uh on the right hand side and then on the left hand side you have a place where you can specify the type of music you want to create as well as any lyrics that you want to you want to have so maybe we can start very simple let's start with just creating uh i don't like rock music so here we can type rock and then for simplicity's sake I'll just create a song about New York and so this is a feature that we launched just yesterday actually where you can write you can ask the language model to write lyrics for you before you submit the song and can even tell it to give it suggestions on what to do like for example let's say we want to make it a little bit shorter
34:58nice okay so you can essentially tell it to get more verbose or less verbose and do other things as well i don't know uh like make sure to mention new york does it keep the last prompt in mind when you give it another instruction so is it still thinking keep this short as you add the make sure to mention new york in the update box oh uh yes Yes, so we actually have a prompt history that shows all the prompts that have accumulated so far. So we aim for this to improve upon our previous lyrics writing experience by giving people the aid of AI to help them come up with ideas when they might have writer's block, they don't really know what to do, for example, like me at this current moment.
35:47And so now that I have the genre and lyrics, I can hit create. and that will uh queue up uh this creation yeah and well while that goes i just i've been thinking about this because whenever i sit down to use udo or a similar product i tend to think in not genre terms but in terms of um bands that i love and kind of how they approach the world and um i'm kind of curious when you when you're using udo do you tend to stick more towards like like rock or do you get a hyper specific like make me a rock song with a touch of i don't know tom petty or something like that because you can pull in different influences yeah so uh i would say that i usually uh just stick with a genre information uh but for users who um have a specific artist in mind uh we give we provide functionality for a user to type in the name of the artist and we look up the style for the artist so we don't actually put the name of the artist anywhere in the prompt because we don't want to create something that sounds exactly like that artist but then for example when you like type in like led zeppelin it will like replace led zeppelin with uh the list of genres like he is that they are like you know like are commonly uh associated with so like you know like uh like maybe hard rock uh maybe uh male vocalist 70s um just things like that okay so if i put in i mean this is again a niche genre but like periphery it's gonna think progressive metal guitar forward male vocals so it'll essentially atomize an artist's name and so essentially then artists become shorthands for genre and style that's correct and um and we try to make sure that like the generated outputs um are definitely influenced by the style of that artist but it's not that exact style because we that's one thing that we really do want to avoid and we are dancing around the lawsuit here and i'm trying to deliberately ask questions you can't answer but yeah thank you for answering that can we play this let's let's hear it uh yeah sure so this is uh the first uh example that came out
Read the full transcript
37:58concrete jungle rising Lost in a new place When people rush right past me Lost in a parade In my heart I feel like I've ever dreamed so made
38:19It's better than something that I could write, so shout out to that. And then, for example, you can then add additional descriptors to this. so maybe um not just raw maybe you want to make it uh a minor uh and then uh you hit create again um and then queues it up and so uh you i'm curious about this because there is a there's a little bit of time that goes through uh from when you click create to when udo gives you the song which by the way to me is is no big deal um but it does seem to be variable david so what makes it a longer or shorter calculation process on the ui side Yeah, so on the UI side, so we submit the request over to the server, and the server reads the prompt like rock A minor and tries to figure out what to do with it before it sends it to the model.
39:10So we have some processing that goes on. We also run checks for every song that gets submitted. We take the lyrics and we do a copyright check to make sure that the lyrics are not copyrighted lyrics. and this is something that's probably a little bit overly strict right now there are a lot of public domain songs or things that should not be that are not really copyrighted that gets flagged but we want to do it on the side of caution rather than not flagging something that's actually copyrighted okay now we have this new song same idea but now in A minor let's uh let's listen to the new song Chasing the Pulse let's see
40:10i'm not sure if you have perfect pitch or not but like uh uh it's a little bit hard for me to tell but it's definitely a minor key. David, I'm not going to lie. I do not have perfect pitch. Indeed, if you had ever heard me sing a bedtime song, you would think to yourself, that guy plays music? Because it doesn't sound like it. So I can't tell if it's A minor or not. It did sound minor. But what hit me though is when the, in the chorus or maybe it was the bridge, when the harmony voices came in, wasn't on the first note. They came in a little bit later, which feels very stylistic and therefore and i mean this in the best possible sense like like human it felt like something that that would be a music editorial decision that a human can make to say hey we're gonna have lead singer and then the harmony come in and delay i don't know i i'm i'm always a little bit torn with this between going this is the coolest thing i've ever seen and oh god are are humans gonna lose just because you know i've sat in the guts of a symphony as we took Beethoven's fifth out of the studs and rebuilt it.
41:15I don't want a future in which we lose that, but at the same time, I'm going to click these buttons a hundred times because it's a lot of fun. So maybe I'm part of my own problem, I suppose. Yeah, well, I think that people who play in bands, they don't necessarily, they won't stop just because there's this additional source of music. One of our co-founders actually is involved in a band and he regularly tours the uk to to perform with his band and so um i think he continues to enjoy this right because it's just fun for humans to create music and we just want to like give people uh more people the opportunity like he can create music with his band but previously he couldn't create music in his bedroom uh or like lying down uh like on his couch and then uh oh i want a song and like previously what would you do like you can't even do anything and so now uh this is possible yeah and just because i'm going to be an enormous brat because i can um court one of our fine producers here at twist has given me a prompt he would like us to try so david if you're up for it in zoom chat there is a prompt entitled a jazzy neo-noir offbeat rap song about dinosaurs which is evidence that Court is Gen X but we'll leave that aside for now but can we give that a try?
42:40Yeah, jazzy neo-noir offbeat rap song about dinosaurs let's see what we get the person who requested this song for everyone who's listening to this later on was in a punk band once so there you go this is what a punk fan is going to put into the the udo uh generation process all right uh well this is waiting david one question i had written down just because you know i love this sort of thing what's the craziest song that you guys have seen people come up with because everything that i've done thus far has been pretty standard but i'm curious has anyone like really blown your head off well one of the examples of a song that took off pretty unexpectedly is a song called B.B.R.
43:30Drizzy. It's a song that one of our users created. He himself is not a musician, but he is a comedian, actually. And so he wrote the funniest lyrics, and he used EO to turn this set of lyrics into a song. And it ended up being sampled from by Metro Boomin, who created a beat. Part of the entire Drake and Kendrick feud, and challenge people to create voices on top of this. And the funny thing is Drake himself actually rapped on top of it. It's pretty amazing watching it from the sidelines to see your tool be genuinely immersed in pop culture. We'll come back to that in a minute, but I want to play everyone this song.
44:17So here is the first sample clip of jazzy neo-noir offbeat rap song about dinosaurs. Hit it, David.
44:26City lights, geno feet at the floor. Brasaurus, bruise, T-Rex on a roll. Pterodactyl fly, rhythm digging in my soul. Bones of the past, we dance like we're aces. History breeds a jazzy night. Ballot stages, echoes in the alley. Shadows keep time. Jurassic jazz notes in the moons climb. Dinosaurs in the urban globe. Rhythms of time, let the ancients show. Underneath the starry flow. Where the city and wild collide. We go. I mean, I'm not going to lie, that's not bad. That's not bad. Brontosaurus groove, T-Rex on a roll, pterodactyl fly, rhythm digging in my soul. That is actually probably better than some stuff that I listen to on Spotify currently.
45:11Yeah, our lyrics are very peculiar. It's an artifact of a language model. No, no, no, no. I meant all that. That was not sarcasm. I never thought I'd see Brontosaurus, T-Rex, and pterodactyl all within one rhyming couplet essentially yeah no uh again i don't know if you answered this but is the model that writes the lyrics the same model as what does the music generation or are those two different models that then are brought together for this final product that we just heard we use different models so um the model that creates the music is a proprietary model that we trained because there's nothing like that uh elsewhere and and the model that writes the lyrics we just use uh gpt actually oh simple enough yeah well i mean it works pretty well okay i want to talk about 1.5 a little bit and then we'll wrap on virality so 1.5 came out back in july um that brought key control improved i think was global language and then also audio quality so how has been the reaction to 1.5 and then what's next from UDO in the feature context?
46:25Yeah, I think people were excited by the changes. For a better global languages, a lot of our Chinese-speaking users remarked how the model suddenly became a lot better at producing lyrics that have Chinese in them. Our key control is definitely a feature that people have wanted for a very long time and people love the ability of like uh specifying the key and then modulating within the song so you can you know like specify a key for the free section and when you extend the section you can then specify a different key and you can kind of like um you know uh specify your own harmonic progression throughout the course of the song how long until that's like super visual like i can imagine myself like um let's say a song is three minutes and i'd like a line and i'm like this chunk should be in a minor and be up tempo xyz and then i want six measures of this that like does this become a visual tool versus just something that i prompt at least in my experience with words yeah so this is something that um going forward uh we do want to make more visual over time so we recognize uh there are some deficiencies in our current interface that make it a little bit harder than necessary and so we want to do like um user research to figure out how best to uh craft the interface in a way that's intuitive for musicians okay so because musicians are already super familiar with editing software and so forth so that kind of interface would be a second nature to them now i want to talk about virality because if we go back to when you guys announced your fundraise i think bloomberg reported and i have it somewhere in my notes here that you were seeing something like 10 songs created every minute or something like that i forget the exact um pace but how has the company been doing in usage terms in the last couple of months and how much bigger is it compared to that april june time frame yeah so it was actually like 10 uh sounds every second uh not a minute but uh people are um people still like uh like super uh engaged with the entire process we have a like a very dedicated group of power users who go on Discord all the time and they share the songs that they have created.
48:39There's actually a bit of a collaborative flow as well where people work on lyrics and songs together. And you see this because in the final output, they will credit each other. They will say, oh, this song created with the help of this other user. And it's really fun to see people uh working on music uh in this way it's kind of like how people would jam together in the past right and well people still jam together but like now we have another way for people to jam together and i think this is also part of what music is about like bringing people together with a common passion i agree with that entirely i was just thinking that you know you're right not everyone now has to go to a jam room which means that they don't have to get hearing loss like I did growing up in my skull band, which rest in peace did not make it big and turn us all into multimillionaires.
49:29I'm sad to say now on the virality point, you mentioned a discord, you mentioned power users. I learned about you guys from a friend, but I'm kind of curious is, is the product here inherently viral? Because, you know, I was sent a song about me and my friends. I like to use it and I've been playing with it, showing it to people. And so I'm just kind of curious if that limits your sales and marketing costs because people are almost taking your product to their own networks ambiently yeah i mean most of our growth uh almost all of our growth is we are completely organic channels where people are just like sharing uh amazing outputs that they have and then people asking each other how do you do that and then like uh just spreading this way so how has been uh like registered user growth at the company is it still as quick as it was before in like percentage terms or gross number terms how should i think about growth at the company essentially yeah i mean obviously there was a very large initial spike when we launched but we do still see like steady uh steady growth every single month and uh we believe that like once we uh launch uh new versions of the model there will be renewed excitement and um as people find new ways of controlling the outputs and using the model for their own production needs.
50:46Okay. All right. Well, I mean, I'm going to be watching with, with very close eyes because I'm a user and now a customer, but one thing I've heard from VCs lately, I think Sarah Tavill from Benchmark wrote about this. And she said that a lot of the AI, the big AI model companies, the open AIs and so forth, um, are going to go kind of up stack in time. And so that startups, not yours, but some startups that do build products using well-known commercial models, for example, might eventually get supplanted by their model provider essentially going up stack and taking their lunch. Are you at all worried as a company at one of the larger model companies, Amistral and Anthropic and OpenAI, I'm going, hey, music is cool, we should do that too, and then kind of bulldozing into your market.
51:32Yeah, so we think that inevitably in the future there will be more companies who enter this music creation space We believe that music is sufficiently different from text, and there's a significant product element as well. You want to have the right interfaces for people to interact with these models. So like a chatbot, like ChatGPT, is probably not the right interface for people who want to create music. And so I think there's actually an open question on how to best produce this type of product. is something that we are working towards. It's a tight coupling between the model and the product and getting the right level of controls in the model so that you can expose them to the user in an intuitive way in the product.
52:22Okay, and then just to wrap things up here, David, I just want to ask you one more thing before I let you go, which is, I think about UDO and Suno is the other company that I think people best know in your space. And so I'm kind of curious, where do you see UDO today in comparison to Suno? And how many of their engineers are you currently trying to poach? So I think we want to position ourselves as allies to artists and soundwriters and producers. We want to focus on giving them the highest quality tools available. So instead of focusing on meme songs in particular, we want to focus on the really powerful creation tools to help creatives make music and make high quality music that they're proud of.
53:05and maybe eventually they want to incorporate in their other music workflows. That was a very, very deft non-answer, but let me take another run at this. Do you consider UDO's music model to be the best in the market today? I would say so, yes. It's the only model that produces stereo music at like 44 kilohertz sampling rate. It's a lot higher fidelity than any other music model that's out there. It has a better understanding of genres than like almost any other music model. Okay, I'll take that. I do want to have you back though in, I was going to say a year, but given that you launched the product in April and it already feels like we've gone through two generations, probably sooner than that, because I'm curious to see how fast things improve, how competition evolves.
53:58And if you guys do decide to go back out into the market, because I think that given your traction, early monetization, and so forth, you should be able to raise more. So it'll be very curious. But David, thank you so much for coming by Twist. I really appreciate the information and the notes. And thank you for making a new tool for me to play with because I absolutely love music. Yeah, thank you for hosting the podcast. It was really fun. All right, everybody. Twist is back. We do live news. If you're not with us on YouTube, see you there. We're also on every single podcast platform you can possibly find.
54:26And we are always trying to find the best and most interesting founders to explain the market as it is. This has been Udio, David Ding and Alex. Hey, we're out of here.
From the publisher
This Week in Startups is brought to you by…
.Tech Domains. Don’t miss our “Jam with JCal” contest! To apply and get more details go to https://Jamwithjcal.tech brought to you by .tech domains.
LinkedIn Ads. To redeem a $100 LinkedIn ad credit and launch your first campaign, go to https://www.linkedin.com/thisweekinstartups
Brave. If you’re building AI and search-based applications, train your models with the Brave Search API. Get started for free at https://brave.com/jason
*
Todays show:
Udio’s David Ding joins Alex to discuss the inception of Udio (1:32), advancements in AI music creation (8:47), and the evolution of Udio’s AI model (10:32). Plus, David demos Udio’s capabilities live (32:47)!
*
Timestamps:
(0:00) Udio’s David Ding joins Alex
(1:32) David's journey and the inception of Udio
(4:26) AI music generation and user control over music elements
(7:52) .Tech Domains - Apply for the Jam Session with JCal contest today at https://jamwithjcal.tech
(8:47) Advancements in AI music creation and data annotation
(10:32) Evolution of Udio's AI models and early versions
(14:59) Udio's target audience and the future of AI in music
(20:06) Udio's potential DAW integration and music production terms
(21:52) LinkedIn Ads - Get a $100 LinkedIn ad credit at https://www.linkedin.com/thisweekinstartups
(23:17) Udio's funding and business model
(27:11) Financial discipline and GPU cost efficiency at Udio
(31:28) Brave Search API - Get started for free at https://www.brave.com/jason
(32:47) GPU-based compute challenges and a live Udio demo
(48:03) Udio's user interface, engagement, and community insights
(49:57) Udio's growth, virality, and competitive stance
(53:27) Udio's model quality and expansion roadmap
*
Subscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.com
Check out the TWIST500: https://twist500.com
Subscribe to This Week in Startups on Apple: https://rb.gy/v19fcp
*
Check out Udio: https://www.udio.com
*
Follow David:
LinkedIn: https://www.linkedin.com/in/david-fengning-ding-053b1282
*
Follow Alex:
LinkedIn: https://www.linkedin.com/in/alexwilhelm
*
Thank you to our partners:
(7:52) .Tech Domains - Apply for the Jam Session with JCal contest today at https://jamwithjcal.tech
(21:52) LinkedIn Ads - Get a $100 LinkedIn ad credit at https://www.linkedin.com/thisweekinstartups
(31:28) Brave Search API - Get started for free at https://www.brave.com/jason
*
Great TWIST interviews: Will Guidara, Eoghan McCabe, Steve Huffman, Brian Chesky, Bob Moesta, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland
*
Check out Jason’s suite of newsletters: https://substack.com/@calacanis
*
Follow TWiST:
Twitter: https://twitter.com/TWiStartups
YouTube: https://www.youtube.com/thisweekin
Instagram: https://www.instagram.com/thisweekinstartups
TikTok: https://www.tiktok.com/@thisweekinstartups
Substack: https://twistartups.substack.com
*
Subscribe to the Founder University Podcast: https://www.youtube.com/@founderuniversity1916




