In short
Eye On A.I. Podcast Notes: Episode #188 - Edward Balassanian: The Evolution of AI Music
Introduction
- Host: Craig S. Smith
- Guest: Edward Balassanian, CEO and Founder of Aimi
- Focus: The development of generative AI music and its implications for the music industry.
Episode Highlights Aimi's Unique Approach
- Aimi stands for AI Music Initiative, founded to deliver generative music as a service.
- The goal is to transform music from an artifact into a service, making it accessible for both amateurs and professionals.
- Aimi distinguishes itself from other generative music startups by focusing on a bottom-up approach in music creation.
Technical and Legal Challenges
- Technical Challenges:
- Music is dense and multidimensional, requiring models that maintain musical and mathematical cohesion.
- Aimi emphasizes training models with ethically sourced, high-quality data rather than finished music.
- Legal Challenges:
- The music industry faces unique legal hurdles, including copyright issues and the pay-per-listen business model.
- Many generative music companies struggle with licensing content for training AI models.
Aimi's Music Creation Process
- Aimi’s proprietary scripting language, AmyScript, generates music in real-time, categorized into audio artifacts.
- The process involves:
- Combining and arranging small audio components rather than replicating finished compositions.
- Allowing user input to adjust musical elements, catering to both novices and professionals.
Market Potential for AI-Generated Music
- The demand for music as a service is growing across industries (e.g., gaming, social media, content creation).
- Aimi's model allows clients to use generated music without paying royalties or sharing revenue, differentiating from traditional models.
Real-World Applications
- Aimi targets diverse markets, enabling companies to integrate generative music easily:
- Pro Version: For individual creators to interact and produce music.
- API Integration: For enterprises like TikTok to incorporate music directly into their platforms.
- Live Streams: Continuous, non-repetitive music streams for retail, hospitality, and fitness industries.
Future of AI Music
- Edward believes that AI music is already at a producer-grade quality level, and the focus should be on how AI can enhance human creativity.
- Concerns about music spam in streaming services as generative tools become more accessible.
Key Takeaways
- AI Music as a Tool for Creativity:
- AI should enhance rather than undermine human creativity.
- Market Evolution:
- The generative music market is emerging, with numerous startups, indicating a significant opportunity.
- Consumer Interaction:
- Aimi's focus is on making music creation user-friendly for all skill levels, fostering broader participation in music production.
Conclusion
- Edward Balassanian’s insights highlight the transformative potential of AI in the music industry, emphasizing quality, creativity, and accessibility.
- Listeners are encouraged to consider the implications of AI in various creative fields, particularly in music.
Additional Resources
- Website: [Eye on AI](https://eyeonai.com) for transcripts and more information.
- Follow on Twitter:
- [Craig Smith](https://twitter.com/craigss)
- [Eye on A.I.](https://twitter.com/EyeOn_AI)
Sponsor Messages
- Vanta: Security and compliance platform offering $1,000 off for listeners at [vanta.com/eyeonai](http://vanta.com/eyeonai).
- Oracle Cloud Infrastructure: For robust AI needs, visit [oracle.com/eyeonai](https://oracle.com/eyeonai) for a free test drive.
This episode underscores the significant advancements and ongoing challenges in the dynamic realm of AI music, inviting listeners to engage with emerging technologies that are reshaping creative industries.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00I think the music coming out of Amy is producer grade and it's generative music so I think we're already there. I think the more important question isn't, are we going to be able to distinguish AI music from human music? I think the more important question is, are we using AI as a tool to accelerate and amplify human creativity or not? And I think the answer is resoundingly yes, I think we are using it to amplify human creativity. AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power.
0:39So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs. OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle. So now you can train your AI models at twice the speed and less than half the cost of other clouds. If you want to do more and spend less, like Uber, 8x8, and Databricks Mosaic, take a free test drive of OCI at oracle.com slash I on AI.
1:32That's E-Y-E-O-N-A-I, all run together. Oracle.com slash I on AI. That's Oracle.com slash I on AI. Hi, my name is Craig Smith, and this is I on AI. In this episode, I speak with Edward Balasanian, founder and CEO of AIMI, A-I-M-I, a generative music company that aims to transform music into a service. There are a lot of these generative music startups on the market today, but AIMI is different. Edward shares his vision for enabling an ecosystem of applications through AIMI's platforms, which combines various AI models in a proprietary scripting language to create high-quality, precision music.
2:24We discuss the challenges of AI adoption in the music industry, Amy's approach to building their platform, and the future of AI-generated music. Edward also demonstrates Amy's music creation process and explains how the company plans to empower creators and enterprises alike. I hope you find the conversation as fascinating as I did. Why don't you introduce yourself and give your background, your education background, if it's relevant in your work history. Sure. And then we'll get into AI music. Terrific. Well, my name is Edward Balasanian. I'm the CEO and founder of AIME. AIME is a generative music company.
3:14The name actually stands for the AI Music Initiative. We founded about four years ago, and our goal then and now has been to provide generative music as a service. In other words, to transform music into something that can be thought of as a service rather than just an artifact. My background is engineering, computer science degree, worked at Microsoft for five years and have been starting companies since then. I've always been sort of a platform guy. I love building platforms that enable an ecosystem of applications to succeed and thrive. And that's exactly what we're trying to do at AIME, is to turn generative music into a service that can be incorporated into any application, product, or service.
4:01Yeah, that's fascinating. I've spoken to a few people on the podcast about music generation. uh it's the market for me because uh i tend to look at at uh basic research more than uh specific startups or applications it i see a lot of activity and maybe you can help disambiguate that a little bit i know there's uh like google magenta ophonic like AudioCraft by Meta, Stability AI has Stability Audio. You know, I think Soundful's doing something, Hook Sounds, AI Studio. What's going on in that sector? Is everyone doing kind of the same thing with slight variations because it's a new market and there will eventually be consolidation?
5:07Or is everyone doing something very different? Well, I think music is one of the last mediums to really be fully engulfed or consumed by AI. And the reason for that is multifold. First of all, music uniquely represents legal challenges that the other creative mediums like photo, text, and even video don't have. There's an international body of incumbents that vigorously protect the status quo that will vigorously enforce copyright. And they also vigorously defend the business model of pay-per-listen. So it's difficult to create solutions in companies that succeed in the tech space where you have to line your business model around pay-per-listen.
6:00But I think the other The reason you're seeing so much momentum now is as machine learning has evolved and matured, music is a very dense medium. Not only is it very dense compared to something like text, but it's multidimensional as well. You know, it's something where you've got layers of instruments and then you've got a temporal aspect to it as well. And both of those need to be coherent. You need stylistic cohesion. You need musical cohesion. And you also need the mathematical cohesion, too. It needs to be correct from a music theory perspective. So it's been late in terms of AI adoption in the music space.
6:47Having said that, what you are seeing now is a lot of training of monolithic models on large bodies of music. And some of those models are trained on licensed content where the content was licensed specifically for training the model, and a lot of them have not been. And I think you'll see a lot of training in secret and a hope that we can kind of get away with it, if you will, building in mechanisms so that you can't identify the source inputs to these models. that uh that's i think where we're at now where a lot of these models are essentially trained on finished music our perspective is very different on that we just don't think that that kind of model training is going to be effective for truly creative uh for for true creation and creative expression in the space of music because music is made from the bottom up it's not made from the top down.
7:51And the best music making is done from the bottom up. When producers make music, they essentially mix and arrange audio artifacts and they sequence these in time to tell a story. That's a big part of music making. Just like if you were a chef and you wanted to make a soup, you can't just go up and taste a bowl of soup and try to recreate the recipe. You need the recipe. You need to know what the ingredients are. So yeah, our approach is very different, but I think you're seeing a lot of sort of sameness around the training of these monolithic neural nets right now. And this is using transformer-based models?
8:34It seems like a natural for transformer-based models since they're predicting the next token music is note-based. Yeah, I mean, what you're seeing is essentially these large language models that were designed for text being applied to music by essentially turning music into images that then get turned into text. So you're able to essentially transform the audio, which is very dense, into a representation as text. And then you can use a large language model to essentially generate music that way. One of the challenges you run into with this is you lose a lot of the nuance in music. And you can really quickly hear AI-generated music because a lot of the nuance that goes into music making that you get from building it from the bottom up is lost when you try to replicate the finished signal by just listening to it.
9:34Yeah, that's interesting. Well, presumably you would generate individual tracks and then stock them, right? You wouldn't generate an orchestra. You would generate a violin and generate a cello and generate, you know. That's the way our system works. So we actually are combining audio artifacts, sonic artifacts that are the compositional elements that go into creating music. So we're mixing and arranging small audio artifacts that we are applying effects to, that we're sequencing in time. It's very different than listening to a finished signal and trying to replicate that sound, which is how the large language models are being trained.
10:27They're being trained on a song and then trying to replicate it. So they would listen to a symphony and try to replicate that sound. But as you just said, you're going to lose the nuance of the five violins versus the four cellos versus the piano versus the harp. That nuance is super important. And you lose it when you try to just listen and imitate the music that you're hearing. Yeah. I had the founder of Soundful on the podcast probably a year ago, and they had sampled individual notes on individual instruments. They built up this massive library of individual notes of individual instruments and trained a model on that.
11:18And then the model combines those individual notes and then the individual instruments into a composition. Are you guys doing something similar? Kind of. I mean, the goal is really important to understand here. Our goal is really to be a creator tool. And when I say a creative tool, we want professionals all the way to amateurs to be able to make music. And that means giving the professionals the kind of creative input they need into the composition of the music. One of the challenges you've got with large language models right now is if you prompt it, it's difficult to reason about the output based on the inputs that went into it.
12:05If you go to ChatGPT and you put in the same prompt three times, you're going to get three different answers. and it's difficult to have the kind of creative input into that that you want. And what we wanted to do is really provide a system that would give you that kind of creative input without watering, without making it too complicated for those who just want to say, I need a hip hop song that's two minutes and 35 seconds and I want it to have the last four seconds tail off in an outro. So we want to give you the kind of input that you need based on your skill level but not relegate you to watered down music just because you're a novice or a pro.
12:44Yeah. And are you in your training data? Is it individual notes or individual tracks? I mean, I haven't. I was reading recently about what is the unit of audio? And it's, I think it's called a sample. Is that right? Yeah. Yeah, we use a lot of different kinds of audio artifacts. Samples are some of them. And we actually believe that in a lot of cases, you want a sample to reflect human creation, like someone singing. We call those hero samples. You know, sounds that need and should be human because the human being created them with emotion and flaws, you know, imperfections. That's part of what makes music beautiful.
13:42So we do use a lot of audio samples. We also have generative instruments that can play along with other samples. And these are programmatic instruments. So there's a lot of different kinds of audio artifacts that we're using at runtime. And we have a programming language called AmyScript, which is essentially a programming language for generative music. That programming language is mixing, mastering, and arranging music in real time, like a producer would. And our models generate that script rather than kind of generate the music. They generate, say that again, they generate that. Yeah, so think of AmyScript as the way you describe a recipe, right?
14:24So instead of trying to train our model to make soup, our model creates recipes. And then we have a programming language that executes that recipe. So we don't have to train our models to create finished music. Instead, we train our models to generate these recipes and these recipes generate the music. It also means that we're able to execute and generate music orders of magnitude more efficiently in terms of CPU utilization and time. because we're executing highly optimized TypeScript code. That's the basis for AIM script. I see. And when you said that you convert sound into text and then the language model generates a string of text that then is decoded into sound, is that right?
15:21That's not what we do. That's kind of typically how large language models are being used to train monolithic neural nets. So you essentially, large language models are really good at text. I mean, that's what they're designed for, right? That's why they're called large language models. So you essentially, the exercise is to transform audio into text. And you can do that by turning it into these visual representations and then turning the visual representations into textual representations, and then training the model on the text. So what are you guys doing that's different from that? Well, we have seven, I think, seven different models, actually, that do all kinds of different pieces of the construction of music.
16:11We have models that understand audio artifacts, know what key they're in, know what instrument they're in, know whether it's a female or a male vocal, know what tempo it's at. We have other models that separate sounds so that we know that there's two instruments, what those two instruments are. So there's a lot of collaborative models that work together to provide the framework that our scripting language can then use to generate the music. We don't fundamentally believe that music AI is a one model problem. So we employ and deploy models as needed to support the infrastructure that we've created.
16:50Yeah. And on the samples, the voice samples, do you have people in a studio making those samples? Are you pulling the samples from catalogs? Or how do you get the voice samples? We've commissioned a lot of artists to create vocal samples for us. we've also purchased a lot of vocal samples and we also generatively create vocal samples so again this just goes back to us not trying to apply one model to the universe of music problems that we're trying to solve instead we're very opportunistic about the way we leverage technology including models for the system so you know for example using lyric generators and then using those lyrics to generate vocals works it doesn't sound as good as humans, but in some cases that's okay.
17:47So it's fine for us to have audio artifacts that include human-generated artifacts as well as machine-generated artifacts. Yeah. And when you said you're providing this as a service, what does that mean? That's a great question. So going back to something I said earlier, the music industry has really sort of relegated music to something that you consume and you pay for the consumption of it. So if you listen to a song, you pay to listen. If you play music in an auditorium, you have to pay for the number of potential ears that would be in that auditorium. So our business model is very different.
18:30We charge for music as a service, meaning just like if you use OpenAI, they're not trying to take a percentage of your book sales if you use Chai Tupiti to write a book. And in the same vein, we don't try to take a percentage of your audio sales or music sales if you use our service to generate music. We want you to be able to use the music however you want. If you want to score a video with it, if you want to upload your own sample and have Amy make a song for you around your sample that you put on Spotify, and you become the next biggest hit great we love that but you pay for the service the generation of the music not the artifact of the service yeah i see and uh okay what i mean i've been talking to another guy who's in stealth right now but he's building uh an ai uh music uh platform uh you You know, not to generate music, but to distribute music, which is an interesting idea.
19:38How big, I should say, who's your client base right now and how big do you think this market could be? Well, I actually just talked about this at South by Southwest. So, you know, we spent four years building this underlying platform that is the backbone of annual music services. And we built a consumer app last year that we released. We built it mostly to show off the power of the system. We're not in the direct-to-consumer business, but you can download the app and listen to free music on it. And the music's great. Now we're releasing Amy Music Services, which is the service that powers it, so anybody can incorporate generative music as a service.
20:22And one of the things I talked about was this market has been slow to form, and it's not because of a lack of opportunity. I mean, music cuts across so many different industries, gaming, social media, content publishing, video production. It's everywhere, automotive. The problem is that the business models have been stymied. I mean, you've been forced into this play-per-listen model, and it doesn't work for the TikToks or the Metas or the Canvas or the gaming companies of the world that need music integrated into their product. So our vision of this is that music as a service is a massive opportunity.
21:03And by not trying to compete with published content, we're not trying to replace Taylor Swift. Taylor Swift and Drake and The Weeknd are always going to have a place. People want that human connection. But sometimes you just want audio for your video that you're going to upload on TikTok. And you don't want it to be Taylor Swift. You want it to be something uniquely created for your video. And that's where I think generative music can really play a role and where music as a service really can fit in. Yeah. And as a service, you mean that creators can log. log, they subscribe to the service, and then they can generate X number of hours.
21:47I would imagine they're tiers. Is that what you mean by service, or is the service that... That's exactly what I mean. Okay. It's not that you plug in Amy's API, and you can... That too. We have both. Yeah, and on the API end, you can integrate different music that's on offer into your project. Or is the API just pulling the music creation interface into your system? Yeah, so the music services is offered in three different versions. One is the pro version, which is designed for creators. It's a website, log in, you have a conversation with Amy, and Amy will make you music. You can upload your own sample to it.
22:39It'll build music around your sample. And you can see the construction of the music. You can actually interact with that construction, or the composition, I should say, which allows you to have a lot of creative input. So this is designed for amateurs all the way to pros who want to make music and they don't want to spend hours or weeks inside of a digital audio workstation. Separately, the same service is available by API to enterprises. So the TikToks of the world can integrate this directly into their platform and seamlessly include generative music for their creators. So if you're uploading a video, you can score your video very precisely with music that's exactly the right length and that follows the arc that you want.
23:20And then last, we also have what we call live streams. And these are just internet audio streams that you can integrate directly into your product when you want zero configuration. You don't get the input into these that you would otherwise have with our API or our Pro Tool. But for a lot of our customers, which actually are our initial customers right now, these live streams give them a way to just integrate in continuous streams of music that's not repetitive. It's high quality and you can just leave it on. You can play it, you know, 24 hours, seven days a week and not get tired of it. I wanted to jump in and give a shout out to our sponsor this week.
24:00When it comes to ensuring your company has top-notch security practices, things can get complicated fast. Vanta automates compliance for SOC 2, ISO 27001, HIPAA, and more, saving you time and money. With Vanta, you can unify your security program management with a built-in risk register and reporting, and proactively manage security reviews with AI-powered security questionnaires. Over 7 ,000 global companies like Atlassian, FlowHealth, and Quora use Vanta to build trust and prove security in real time. Listeners get$1 ,000 off Vanta at vanta.com slash ionai. That's I-O-N-A-I, E-Y-E-O-N-A-I, all run together.
24:55And that's Vanta, V-A-N-T-A dot com slash ionai, E-Y-E-O-N-A-I. Give them a try and get$1 ,000 off Vanta. your security program management with built-in risk register and reporting. Yeah. What's a use case there? Who would be a typical customer for the live stream? So that is offered mostly right now where our customers are B2B2B. So we provide it to businesses that in turn provide it to retail, hospitality, fitness, restaurants. So for example, one of our customers has integrated these live streams into their server product which slides into a rack at your hotel and it distributes music throughout the hotel so it's transparent to the hotel that it's amy's music and for them it's very uh the the service that we offer to the company that is our customer is empowering to them because right now they have to buy catalogs pay a lot for them and And when they distribute music to European customers, the European customers have to actually pay additionally, based on the size of their venue, to the performing rights organizations.
26:16None of that exists with Amy. There's a one-time per month fee for the service. You can put it in a closet or you can put it in an auditorium. We don't care. The service is what you're paying for, not the consumption of the music. Yeah, that's fascinating. And my son's in the music business, and he's been talking to a startup that is buying catalogs from not major catalogs from small musicians. I'm not, I don't know the language, but I'm wondering, is AI music going to change that market because these smaller catalogs?
27:21Yeah, I don't know. How do you think it's going to change the market? Look, we've seen this with kind of every creative medium. Like when iPhones first came out, all of a sudden everyone was a photographer. and anybody could take a picture. And all of a sudden, the quantity of pictures was overwhelming. And that didn't change the fact that good pictures are good and bad pictures are bad. It's the same thing with music. We're building a tool that allows people to make high quality, high production and high precision music. By high precision, we mean it's what you want. It's not just something that you know, and that spit out for you.
28:00So we see... an acceleration of good music being created with tools like this. But there's also the fact that a lot of music spam is going to start coming out. I mean, you already saw it with one AI company that was spamming Spotify. And as a result, Spotify had to pull down their content because there was just so much of it. And the unfortunate thing is these tools are being used to create music that's below good. So it's fun to make the music, but no one's listening. so you end up with a lot of music spam and music is expensive like to store it to stream and it's expensive so that's going to be a challenge i think you're going to see dsbs really crack down on music like that but again we're we're not in that business we're in the business of creating high quality high precision music as a service would you be able to uh share on on the uh on the screen the platform and just walk us through a simple music creation okay so this is the pro version of amy music services and again this is designed for creators and i know the word creator is really overused now but i like applying it here because we don't want to alienate people who are not musicians who are not professional producers from being able to express themselves musically at the same time we don't want to alienate uh producers by giving them tools that are too simplistic for them to use so let's take a quick look here on the right what we call the plan this is essentially an articulation of the plan that amy has the recipe for the song and i can interact with it from here i can say let's let's pick a reggaeton song for example and it'll show me a plan.
29:45I can also ask it to do it on this side. This is more of a conversational interface on the left. So I'm going to actually use this here to say, let's add a build up between the intro and the verse. So we'll go ahead and add a build up in there. And I can interact with this. You can see here, I can see the whole plan. I can see all the different elements in each of the different sections. We're adding the ability to get into each of these and choose the instrument. you can also choose whether there's vocals or whether they're male or female what language they're in so we can give you a lot of introspection of it so you can see here there's a buildup so we have in the intro bass harmony rhythm i'm going to actually get rid of the rhythm in the intro and i'm going to get rid of the percussion in the buildup so we have kind of a natural progression from intro to verse and then you can see here uh the verse adds in melody harmony in counterpoint i'm going to get rid of the counterpoint here i'm going to add it in over here and again i don't have to do this if i don't know anything about music i don't have to muck with any of this stuff i can just let any kind of do its thing now i'm going to change all this to g minor so let's change the key from d minor to g minor in all sections So Amy will go through and now redo the key for each of these different sections.
31:13I can also adjust the number of seconds that this is. I can add new sections here manually. You can see here it switched to G minor. BPM match, there's some kind of interesting tensions between how long you want the song to be what the bpm is and how many beats per section you have and those all have to be juggled when you're saying you want the song to be a certain amount of time right now we had a photographer at south by southwest with us who needed music for a 24 second video clip that he had and he wanted the last four seconds the outro to be only four seconds so he just went in there and click four so that was four beats and the outro became four beats so you get that kind of precision over this so let's go ahead and hit produce here basically i did an indie pop song i did a i did a hip-hop song i did a same one twice i did a techno song as well and it's really easy to do this you can see here these are some of the genres that we support we're adding edm We're adding pop.
32:23Super simple for us to add new genres. We actually sit down with producers who are very steeped in the genre and we translate their expert techniques into expert algorithms. We then source audio artifacts that are consistent with that genre. And that's it. Within a few weeks, we've got a new genre that we can generate music in. And when we have these audio artifacts, we're actually able to expand the universe of sounds just by using our own generative AI to replicate those sounds. So we can go from a small number of audio artifacts to a large number very quickly and support a massive range of music for any one of these genres.
33:11Are serious sort of producers, pop music or hip hop producers using Amy to do background tracks or that sort of thing? well this just got released at south by southwest so we are in the process of getting this out into the hands of we've worked with close to 200 plus artists over the past couple years they are effectively our audit committee if you will and one of the things that was really important to us is not to release amy until the sound was good to great and they were the litmus test until they gave us the thumbs up we weren't ready to pronounce this as being ready for for consumption so So this is now in beta.
34:00We're providing this to our artists. As I said, they can also upload their own samples. And a lot of our artists are focused more on the creative elements in samples than they are in trying to make arrangements, which is very tedious and oftentimes very prescriptive. If you want to make a song for Spotify, there's a specific formula that you need to follow. And that's not the fun part. The fun part is coming up with the cool melody or the cool bass line and that amazing drum beat. Those are the fun parts, and you can add that audio to this and place it anywhere in here you want, and Amy will build a song around it for you.
34:42Are there – and this will – music generated by Amy will be uploaded to Spotify and the various music platforms. Is that right? Yeah, that's correct. So we're going to allow people to use this music for whatever they want. Now, there are a couple of constraints and restrictions. You cannot copyright the output of Amy. You also can't use the output of Amy to train in the machine learning models. So the output of this content is either for an individual to use in their own publishing, for example, on Spotify, or for incorporation into a video. If you're an enterprise, you can use this audio on your platform.
Read the full transcript
35:24You can't use it to create a new music platform. But you get complete freedom in terms of the consumption of the audio. We do not try to track royalties, and we don't try to track any kind of rev share. The audio is yours once it gets created. Yeah. And does Spotify have a section for AI-generated music, or is there a requirement to label music as AI-generated, or is it just in the general feed? That's a great question. I haven't looked recently. I know that they had some serious issues with the spam, the music spam that was inundating their platforms. I don't know the answer to that right now, but it's pretty easy to spot AI-generated music.
36:17It's just right now, the quality level is poor enough that it's easy enough to identify. I think the bigger challenge here isn't so much should people be making bad music or not. People should be able to make whatever they want, right? Like we're not here to try to dictate whether people should be able to create bad music. And I think tools that make it easy and fun to make music, whether it's good or bad, are interesting. It's just not the business that we're in. We're really in the business of high-quality music, and that's our goal. Yeah. Yeah, that's interesting about the spamming. When did that happen?
36:54I wasn't aware of that. Just this past year. So, 2023, there was a significant amount of audio being dumped onto Spotify and other DSPs, and they had to basically shut it down, pull them off. And by one platform or just? There were a couple in particular that I think got caught in the spotlight of that just because they were being used so prolifically. And the issue was that these were getting like one or two listens. Like nobody was listening to them because they don't elevate to the level of being listenable to. Even the creators aren't listening to them. And it just ended up being spam on Spotify servers.
37:34They're storing it. They're serving it up. It's getting in the way of their rev share model. If you have 100 songs that people listen to once, it's not a big deal. But if you have 100 million songs that people listen to 10 times, that's a billion listens that you've now got to spread the revenue for. And that starts to break the business model. Yeah, yeah. You were saying that it's easy to spot. I would guess, as with all generative modalities, as these models either get larger or are fine-tuned, it's going to become harder and harder to tell whether something's AI-generated. Do you think AI generated music will reach that level of authenticity?
38:29I think it already has. I mean, I think the music coming out of Amy is producer grade and it's generative music. So I think we're already there. I think the more important question isn't, are we going to really distinguish AI music from human music? I think the more important question is, are we using AI as a tool to accelerate and amplify human creativity or not? And I think the answer is resoundingly yes. I think we are using it to amplify human creativity. I don't think we're using it in a way that compromises or diminishes human creativity. You can actually see this with chat GPT. Like I can tell when someone's written a marketing piece using chat GPT.
39:11So that's not bar, right? Like everyone's going to be at that level. So you need to be that much better to distinguish yourself. So I think it's just going to elevate the expectations that people have. And that's exactly what's happening with music as well. Yeah. What were you doing before, Amy? I've had an incubator that's been building startups for the past 20 years. So I've always been in the startup space. I didn't have any connection between music and tech. in any of my businesses. So this was really the first time that I brought the two of them together. Music is just fascinating to me because it's part art and part math and part science.
39:51It's unique in the creative space in that respect. Images are not mathematical in the sense that you can see an image that looks weird or funky and it's okay, but with music, if it's off-key or off-beat, it's hard to listen to and the tolerance that people have for bad music is much lower than for bad video or bad photos. Maybe bad is the best word to use there. Yeah, yeah. How much AI-generated music is there in the public space at this point? It's still a tiny fraction. I think it's significant, yeah. Well, it depends on what you mean by AI-generated because producers have been using generative drums and generative instruments forever yeah not forever for a while and yeah it's really like the full composition hasn't really happened until very recently and you know voice synthesis you saw that with that weekend clip that went up people thought that was an ai generated song it wasn't it was just the vocal that was the rest of it was hand produced by a producer and that's why it sounded so good so i i think you're going to start saying more and more that I think music has kind of come into its own this year.
41:08Last year when we were talking to customers, everyone was very nervous and scared. This year we're getting an overwhelming amount of customer outreach. We don't even have a sales team because people are coming to us at this point. And that wasn't the case a year ago when music was still this forbidden, unknown medium that everyone was scared of. Yeah, yeah. Yeah. And the startup space, are there a lot of people jumping into this or is it still fairly? Yeah. No, there are. I mean, I stopped even paying attention because one thing I've learned as an entrepreneur is you can't just look at whatever your startup is doing.
41:49It's like trying to win a sprint looking over your shoulder. You can't run fast if you're looking over your shoulder. So honestly, I don't look anymore. And there's a new music AI company popping up every other day. It's just the nature. To me, that's exciting. That means the market's here. When you're the only person in the market, you're early. It's good when you see this kind of momentum around the space. It also means that customers are going to be much more receptive. Yeah. Are you guys sort of following the ladder of models as they become bigger and better? Oh, yeah, absolutely. We have some of the smartest AI people in the world working at Amy.
42:37And they've been working on generative music and generative AI for decades. So this isn't a new topic for them or for us. We started four years ago. We're always on top of the latest developments in the AI space. And we use models, like I said, opportunistically. And there's seven of them now. We'll have more of them as time goes on. We see the use of AI as part of a bigger picture. That's our story. And that's always going to be the case. Yeah. Are you guys building a foundation model or are you relying on third party? We haven't needed to build a foundation model. We, you know, there's enough open source models out there that we're able to leverage.
43:20We've developed some of our own models for specific things like MIDI transcription, instrument classification, things like that are unique to us. And we have best in class technology for that. But we're not going to build a large language model. We're not going to build a foundation model like that. Yeah, I mean, the reason I ask is I've been talking to people about code generation. And as wonderful as the big models are in generating code, they're still kind of wonky because they're not trained on a curated data set of clean code. They're trained on everything. Yeah, on everything. There's a lot of junk in there.
44:06That's the data set problem. Same thing with music. Yeah. But those models are already trained. You can fine tune them, but you can't untrain them. But if you build a model from scratch with a highly curated, very clean data set, it's just going to be that much better. All of our models are. All of our models are. We don't use any models that have been trained on any material because we don't know the data provenance. We don't know the legality or ethicality of the data that was used to train those models. All of our models are trained on data that we have absolutely clear records on the data provenance.
44:49Right. But if you're building off of public language models, those already have training sets that are opaque, right? Yeah, we don't. We don't. I see. I see. So you're using open source models that you can train from scratch? Correct. Yeah. Our models are all trained from scratch, precisely because of the data set problem. Like we need high quality data, we need legal data, and we need ethically sourced data. And all three of those matter. And as you know, you can't really reason about the output of a model and try to figure out what input led to that output. So you can't retroactively try to go and fix these problems with these models.
45:41So we've been very careful up front about not incorporating any data in training any of our models that we don't have absolute confidence in. Are there any well-known hit songs or, I mean, you mentioned The Weeknd out there that have used Amy that I could point to? No, not yet. I mean, that's, as I've mentioned, we just released this. So hopefully you'll start seeing some more quality content being uploaded by people. But the examples I'll send you, you'll see they're pretty impressive. The level of polish in the songs is remarkable. Yeah. And on the vocals, is it just harmony or can you have them generate, you mentioned lyrics, I mean, someone singing specific lyrics?
46:41Yeah, both. So this goes back to the comment I was making about being opportunistic. we have vocal synthesis models that can take lyrics and generate audio from those vocals from that we have real human recorded samples as well we have the ability to do generative instruments which can play along with other audio samples so it's a combination of what we call audio or sonic artifacts being mixed arranged produced and generated in real time to match the brief that the user has given us. Is there something that I haven't asked about that you think listeners should know? No, I thought this was great.
47:25I'm very appreciative of the time. As I did mention, we've launched any music services that's available in beta now if your listeners want to go to our website. They can learn more about it there. Our consumer app is also available for download from the app store and also from the Google Play store. Those are really fun apps that kind of show you the value of generative music and how you can turn it into an interactive medium. And then we'll have hopefully some announcements shortly about enterprise customers that are using our technology to provide music into their platforms. Okay. And the app is just a limited platform compared to the online platform and the desktop?
48:15Yeah, the app is really designed to be a fun way to explore music for super fans, especially with electronic music. So you can download the app, you can explore these different experiences. And when you hit play, Amy starts making music from scratch in a specific genre. And you can give Amy feedback in real time to help it understand what you like. So it'll start mixing and arranging and producing music that matches your tastes. Yeah. Well, I'll give it a try. That's it for this episode. I want to thank Edward for his time. If you want to read a transcript of today's conversation, you can find one on our website, eye on AI.
48:58That's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may not be But AI is changing our world, so pay attention. AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud. Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs.
49:46OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle. So now you can train your AI models at twice the speed and less than half the cost of other clouds. If you want to do more and spend less, like Uber, 8x8, and Databricks Mosaic, take a free test drive of OCI at oracle.com slash ionai. That's E-Y-E-O-N-A-I, all run together. Oracle.com slash ionai. That's oracle.com slash IonAI.
From the publisher
This episode is sponsored by Vanta, The security and compliance platform trusted by more than 7,000 customers.With Vanta, you can unify your security program management with a built-in risk register and reporting, and proactively manage security reviews with AI-powered security questionnaires.
Listeners get $1,000 off Vanta at vanta.com/eyeonai
In this episode of the Eye on AI podcast, join us for an insightful conversation with Edward Balassanian, CEO and founder of Aimi, a trailblazer in the realm of generative AI music.
Edward takes us through his fascinating journey from engineering and computer science to pioneering a platform that transforms music into a dynamic, interactive service. Aimi's innovative approach to AI-generated music is redefining the industry by enabling both amateurs and professionals to create high-quality, customized music.
Discover the technical and legal challenges unique to the AI music space, and how Aimi's bottom-up approach to music creation sets it apart from other AI music initiatives. Edward delves into the complexities of training models with ethically sourced, high-quality data, and explains the importance of maintaining human elements in generative music.
Explore the vast potential of music as a service and how Aimi is revolutionizing industries like gaming, social media, and content creation by integrating seamless, non-repetitive music into various applications. Edward also shares real-world examples and future plans for Aimi, emphasizing the role of AI in amplifying human creativity.
Tune in to understand the transformative impact of AI on music and how Aimi is leading the charge in this exciting frontier.
Don't forget to like, subscribe, and hit the notification bell for more insights into the cutting-edge technologies driving the AI revolution.
This episode is sponsored by Oracle. AI is revolutionizing industries, but needs power without breaking the bank. Enter Oracle Cloud Infrastructure (OCI): the one-stop platform for all your AI needs, with 4-8x the bandwidth of other clouds. Train AI models faster and at half the cost. Be ahead like Uber and Cohere.
If you want to do more and spend less like Uber, 8x8, and Databricks Mosaic - take a free test drive of OCI at https://oracle.com/eyeonai
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI




