In short
AI Today Podcast Episode Summary
Episode Title
Inside the Launch of Generative Music Model
Episode Description This episode discusses the implications of a new generative music model launched by Stability AI, exploring how this technology may disrupt traditional music workflows and enable innovation for musicians and creators.
---
Key Topics Discussed
- Introduction to Stability AI
- Background: Stability AI is known for its role in AI advancements, particularly with the introduction of Stable Diffusion for image generation.
- Recent Challenges: The company faced financial difficulties but is now positioned for a turnaround with new innovations.
- Launch of the Audio Model
- Overview of the New Feature: Stability AI has launched a music generation model distinct from vocal models, focusing on creating instrumental music.
- Comparison with Competitors: Rivals like Suno and Udeo exist in the generative music space but have faced criticism over copyright issues.
- Copyright Considerations
- Data Sources: Stability AI trained its model exclusively on content for which it holds copyright, utilizing royalty-free audio libraries and archives, thus avoiding potential intellectual property disputes.
- Technical Details of the Model
- Model Specifications:
- Size: 341 million parameters.
- Optimized for ARM CPUs, allowing it to run on mobile devices.
- Capable of generating up to 11 seconds of audio.
- Limitations:
- Primarily generates short audio samples and sound effects (e.g., drums, riffs).
- Lacks capabilities for realistic vocals or high-quality full songs.
- Limited musical styles primarily based on Western music.
- User Accessibility and Licensing
- Usage Terms:
- Free for researchers and hobbyists or businesses earning less than $1 million annually.
- An enterprise license required for businesses exceeding this revenue threshold.
- Future Prospects for Stability AI
- Leadership Changes: A new CEO has been appointed to navigate the company's revival, and notable figures like James Cameron have joined the board.
- Strategic Direction: Moving towards integrating audio with video, potentially creating comprehensive solutions for multimedia content creators.
---
Key Takeaways
- Innovative Disruption: The generative music model represents a potential shift in how music is created and consumed, democratizing access to music generation.
- Copyright Compliance: Stability AI's commitment to copyright-free music generation sets a precedent in the industry, but with certain limitations in quality and diversity.
- Company Resilience: Despite past struggles, Stability AI's new directions may revitalize its presence in the AI landscape, particularly in multimedia content creation.
---
Conclusion The episode highlights the intersection of technology and creativity, showcasing how advancements in AI can reshape industries like music. Stable AI's new model opens up opportunities while addressing ethical concerns surrounding copyright, making it a significant development in the AI and music domains.
---
Additional Resources
- AI Box: [AI Box Website](https://aibox.ai/)
- AI Chat YouTube Channel: [Jaeden Schafer on YouTube](https://www.youtube.com/@JaedenSchafer)
- AI Hustle Community: [Join the AI Hustle Community](https://www.skool.com/aihustle/about)
Call to Action Listeners are encouraged to leave ratings and reviews, and to explore AI Box for comprehensive access to various AI models.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the podcast, we're going to be talking about Stability AI and a brand new feature that has just rolled out. and that is the ability for them to do audio. So this is a new update that they've rolled out recently. And Stability is kind of an interesting company. You'll probably remember it just for the fact that it was one of like the leaders in the AI revolution. They literally invented a stable diffusion and the way that we use AI to generate images. And yet they really got left behind as a company that's had a lot of financial issues. But I think that they're about to make a big turnaround.
0:33And so because of this, I don't think it's a company that you should count out just quite yet. The one thing I did want to mention before we get into this, if you haven't tried it already, my startup, AIbox.ai, is officially launched. And our first product is the AIbox Playground. We have a beta out right now that essentially allows you to access the top 20 AI models all on one platform. You can chat with them all in the same chat. We have audio, image, and text all in the same chat for$20 a month. So you don't have to have subscriptions to 20 different platforms. you pay one time for that and then you get access to all the different platforms so you can check it out the links in the description aibox.ai all right let's get into what's happening with stability ai so the new update they have the thing that's really interesting about it beyond the fact that you know they came up with kind of like an audio model and i should preface this by saying they have a big announcement about an audio model but this isn't like a vocal model this is a music model.
1:29So specifically, it does music. There's a bunch of different competitors. There's Suno and Udeo that are doing this. But most of these ones that are kind of doing this generated music, people criticize them for the copyright. So they're like, look, these guys, they grabbed all of this data from the internet, they grabbed everyone's music, they trained a model, and now it creates music. So people are upset about kind of the copyright in the data set for this. Stability tried to avoid this essentially. And they did a couple of cool things. Number one, it's a really lightweight small model that actually can run on your phone meaning like suno and udio have apps that can run on your phone but obviously that's going up to the server to the cloud and running off of um you know their own their own websites and servers and stuff you have to have access to the internet with this application you technically could just do everything on your phone your phone is powerful enough to run this model and it can generate you stuff now i will put a caveat on this by saying this is not as good as suno or udo um it's just that's just the nature of the beast so stability trained this only on uh content that they had copyright for which is fantastic right they don't want any sort of ip risk uh involved with this when they're releasing it so they said that it's entirely made out of royalty free audio libraries and free the free music archive um and free sounds those are kind of their sources and they're allowed to do this which is technically great, except that it's not as good.
2:55So that's, I think, the big thing. It is really small. It's 341 million parameters in size, and it was specifically optimized to run on ARM CPUs. So ARM makes chips. These are built on, you know, this model was essentially built so that it's able to run on an ARM CPU right on a phone. These ARM CPUs are often put into phones. so the thing that it's specifically made for doing though is for quick um kind of shorter audio samples and sound effects so you can do drums you can do instruments you can do riffs um and it can make up to 11 seconds of audio you can do it on a smartphone and it takes about eight seconds to do this so this is you know definitely faster than your average uh udo or suno ai um piece but and i'm not saying it's bad i actually think it's fairly decent for what it can do but like it doesn't do vocals and so if you're trying to make a fully fledged song or honestly uh a really great song like suno and yudio are going to do a much better job in my opinion of making music i've tried both of the um i've extensively tried suno and uh it does incredible work makes amazing music people criticize that it was trained off of the copyrighted data i'm not too concerned about that that's not really my problem um you know i'm sure people get mad at me or criticize me for that but That's just my opinion is just like, that's their copyright issue to deal with the model so much better as a user and a consumer and someone that would like to create things.
4:16I'm gonna use the best model. So that's kind of what I'm getting out of Suno or UDO. All right, I wanted to give you a sample though because I'm actually quite impressed by what they have been able to produce. Completely copyright free, there's no issues there. So they have a couple samples of what it's able to actually.
4:36So you can actually go online, check out soundcloud they got a bunch of different samples and all of their samples are like much shorter but they are you know showing you exactly what it's capable of doing they could do some drums some music they have a bunch of limitations in addition to the all ones i've mentioned already one it can only do english prompts written in english so if you speak another language you'd have to translate your prompts into english and google translate or something like that it can't generate realistic vocals or high quality songs it's kind of low quality and it doesn't do a lot of different musical styles it was really just built on a bunch of kind of western they call them western biased training data so these free music libraries are not very extensive it's just mostly kind of like western music so it also has a little bit of restrictive usage I mean it's not the end of the world.
5:30You got to make money somewhere. So it's free for researchers and hobbyists and businesses that make less than a million dollars annual revenue. But if you're making over a million dollars, you have to pay Stability's enterprise license. This isn't the end of the world. And I think this is a pretty standard licensing kind of deal. Although, yeah, it feels like they'd be making something open source. So I guess some people are upset about that. Now, Stable Diffusion is a company that has had a ton of issues in the past. They've raised some new money last year. A bunch of their investors, including Eric Schmidt from Google, the Napster founder, Sean Parker, famously, who invested in Meta, were really trying to turn the business around.
6:11So Emod Mostak was their co-founder and he was kind of the former CEO. He apparently really mismanaged all of their finances, almost completely destroyed the company. Tons of staff resigned. There was a partnership they had with Canva that fell through. investors were super concerned about this so in the last few months they actually got a new ceo and they appointed james cameron to their board of directors which is interesting because typically this has kind of been famous as a image company and with james cameron you can kind of imagine where they're going with this is going to become a video company all these ai generated images are perfectly poised to create ai generated videos and they've also released a bunch of new image generation models so it seems like stability is on track to do some cool things i think specifically if we're looking at video, doing these sound effects and kind of these like smaller music bits makes a lot of sense.
7:02They want this in the background of if, you know, they're making music tracks to be able to, or sorry, videos, it'd be really cool to have also AI generated music in the background. So this makes a lot of sense with kind of their strategic direction. I'll be super curious to see where they go. This is a very prolific company. It's raised a lot of money. It's done a lot of interesting things, but again, it has faced a lot of challenges. So I'll keep you up to date on everything happening with stability. Make sure to leave a rating and review wherever you listen to your podcast. And again, if you haven't tried AI box already, there's a link in the description.
7:31I would love to have you try it. You can dump a ton of your subscriptions for$20 a month. You get access to all the top AI models. You can compare results side by side of different models. You can chat with all of the models in the same chat. You don't have to switch or, you know, not have the ability to keep talking to different models. And it's a lot of fun. So check it out, AI box.ai, and I will catch you next time.
From the publisher
Hear how this technology might disrupt traditional music workflows. This podcast covers how AI tools are unlocking music generation for all. Learn how musicians and creators can leverage this new model for innovation.
Try AI Box: https://AIBox.ai/
AI Chat YouTube Channel: https://www.youtube.com/@JaedenSchafer
Join my AI Hustle Community: https://www.skool.com/aihustle/about

