The Frontier of AI-Generated Music Models

3 Aug 2023 · 18 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The AI Daily Brief - Episode Summary: The Frontier of AI-Generated Music Models

Episode Overview In this episode of *The AI Daily Brief*, host NLW dives into the recent advancements in AI-generated music, specifically comparing Meta's newly launched AudioCraft with Google's MusicLM. The episode also highlights significant trends and predictions in AI investment and discusses recent developments within the AI chip sector.

Key Highlights

  1. AI Investment Predictions
  2. Goldman Sachs Forecast:
  3. Predicts AI investment could reach 4% of U.S. GDP by 2025.
  4. Global AI investment anticipated to approach $200 billion by 2025.
  5. Investment in AI is projected to grow from $93.54 billion in 2021 to $110.19 billion in 2023, with the U.S. representing a major share.
  • Investment Breakdown:
  • 2023: U.S. - $56.83 billion, China - $24.74 billion.
  • 2024: U.S. - $68 billion, China - $30 billion.
  • Key Factors:
  • Businesses need significant upfront investments in infrastructure and human capital for transformative AI integration.
  • Historical context: Previous tech booms (electricity and personal computing) followed similar investment cycles before leading to productivity growth.
  • Future Expectations:
  • 20% of Fortune 500 CEOs believe AI will decrease labor needs in the short term, while over 70% expect reduced labor requirements in five years.
  1. AI Chip Market Developments
  2. TenStorrent:
  3. Raised $100 million to compete with NVIDIA in the AI chip space.
  4. Led by Jim Keller, a veteran in chip development.
  1. New AI Models from Alibaba
  2. Alibaba Cloud has released open-source models aimed at competing with Meta's Llama 2.
  3. The models have some restrictions, similar to those of Llama 2, regarding usage.
  1. Google’s Search Generative Experience
  2. Google introduced updates to improve its AI-powered search experience.
  3. New features include multimodal results with images and videos, aiming to enhance user interaction.
  4. Speed improvements and clarity on source publication dates were also implemented.

Deep Dive

AI-Generated Music Models

MusicLM vs. AudioCraft

  • Google's MusicLM:
  • A text-to-music generator allowing users to create music through specific prompts.
  • Example prompts include crafting melodies evoking seasonal feelings or styles from different eras.
  • Meta's AudioCraft:
  • Combines three models (MusicGen, AudioGen, Encodec) to generate high-quality audio and music from text.
  • Focuses on various audio forms, including environmental sounds.

Comparison Outcomes

  • Both tools produce high-quality music samples from prompts, though reactions to each may vary.
  • The crucial aspect of these advancements lies not in which model is superior, but in how quickly and effectively they are integrated into user-friendly applications for musicians.

Industry Implications

  • Musicians' Perspectives:
  • Mixed feelings exist regarding AI's impact on music creation, with concerns about skill commoditization juxtaposed against excitement for new creative possibilities.
  • The prevalence of AI tools may democratize music creation without necessarily diminishing the value of skilled musicianship.

Conclusion This episode of *The AI Daily Brief* underscores the rapid advancements in AI technology and its implications for music, investment, and productivity. As AI continues to evolve, the conversation around its integration and impact on creative industries remains vital.

For further insights and updates, listeners are encouraged to subscribe to the newsletter and join the community discussions.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today on the AI Breakdown, we're comparing Google's Music LM with Meta's just released AudioCraft. Before that on the brief, major fundraising in the AI chip space, and Goldman Sachs makes a big prediction when it comes to AI and GDP. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Go to breakdown.network for more information about our Discord, our YouTube, and our newsletter. Welcome back to the AI Breakdown Brief, all the AI headline news you need in around five minutes. If you'd like to follow along, go to the AIbreakdown.beehive.com. You can also find that link at breakdown.network.

0:37Today, we kick off with a story that is another report with some big prognostications. This one comes from investment bank Goldman Sachs, and they say that by 2025, AI investment could represent up to 4 % of U.S. GDP. This comes from the article, AI investment forecast to approach$200 billion globally by 2025, that was published a couple days ago this week. Now, even in a world of bombastic AI reports, this one is bombastic. The article begins, Innovations in electricity and personal computers unleashed investment booms of as much as 2 % of U.S. GDP as the technologies were adopted into the broader economy.

1:16Now, investment in artificial intelligence is ramping up quickly and could eventually have an even bigger impact on GDP, according to Goldman Sachs Economics Research. So first, let's look at what they're estimating those investment numbers in AI are in the U.S., in China, and across the world. In 2021, the world saw$93.54 billion of investment in AI. That actually came down slightly in 2022 to 91.9. But in 2023, Goldman estimates that the world global investment in AI will equal$110.19 billion. Of that, they estimate the U.S. will represent$56.83 billion, and China will represent$24.74 billion.

1:55By 2024, Goldman sees the world number going up to$132 billion, with the U.S. at around$68 billion, and China at around$30 billion, and the lines just go up from there. Part of what Goldman's argument rests upon is the fact that people are understanding, widely speaking, just how much labor productivity could be increased by generative AI across a huge number of different industries and sectors. However, as Goldman puts it, for large-scale transformation to happen, businesses will need to make significant upfront investment in physical, digital, and human capital to acquire and implement new technologies and reshape business processes.

2:29So effectively what they're saying is that there will be this period of transition where a huge amount of upfront money will have to be spent in order to retrofit the business world for the AI era. That means investing in actual infrastructure, which might mean things like compute infrastructure, but it also means re-skilling employees, which of course comes with a price tag as well. When you view this$200 billion number not as some generically inflating investment amount, but representative of a key transitional period, I think it starts to look a lot more realistic. To give evidence of that, they point to previous tech-driven productivity booms, which they say have been driven by large investment cycles.

3:07In the case of both electricity and personal computing, Goldman shows how increases in investment in the underlying infrastructure. In the case of personal computing, that means information processing equipment and software investment. And in the case of electricity, that means manufacturing equipment and plant investment led then with a lag of a few years or a half decade to productivity growth that, broadly speaking, mirrored that level of investment. Goldman writes, AI-related investment is climbing from a relatively low starting point and will likely take a few years to have a major impact on the economy.

3:36Over the longer term, AI-related investment could peak as high as 2.5 to 4 % of GDP in the US and 1.5 to 2.5 % of GDP of other major AI leaders. They also point out that there has been a seminal shift just recently. Goldman writes, even though it will take time for AI to boost productivity, market interest in AI has already increased rapidly, with more than 16 % of companies in the Russell 3000 mentioning the technology on earnings calls, up from less than just 1 % of those firms in 2016. Roughly half of that spike came after the release of ChatGPT in the fourth quarter of 2022. For those of you who are listening rather than watching, Goldman's chart of the percentage of companies that are mentioning AI in their earnings calls looks like the type of up-and-to-the-right graph that gets VCs excited to get up in the morning.

4:21Now, what about who benefits from this$200 billion of new spend? Goldman says that they expected to be concentrated in four key business segments. The first is companies that train and develop AI models. The second is those that supply the infrastructures, i.e. data centers. The third is companies that develop software to run AI-enabled applications. And the fourth is enterprise end users that pay for those software and cloud infrastructure services. Now, when it comes to CEO expectations around how AI will impact labor, over the next year, a little over 20 % of Fortune 500 CEOs surveyed said that they thought AI would decrease labor needs, while over 40 % said that they thought that labor needs would be unchanged.

4:58However, zooming five years out, over 70 % of Fortune 500 CEOs said that they anticipated that lower labor would be needed, while under 20 % said that labor would be unchanged. The TLDR on this report is just another example of a major financial institution that thinks we are just at the very beginning of a major transformative AI cycle. Speaking of AI-related investment, next up on today's AI Breakdown Brief, we are looking at yet another firm that is trying to, if not dethrone NVIDIA in the AI chip space, at least provide some competition. Reuters reports that AI chip firm TenStorrent has raised$100 million in fresh capital from investors including Hyundai and Samsung.

5:36Previous to this funding, TenStorrent had raised$234.5 billion and was already in the Unicorn Club, and that interest has been driven in part by the fact that it's led by chip industry veteran Jim Keller, who previously developed chips for companies including Tesla, Intel, and Apple. Over in the world of LLMs, Alibaba Cloud has released two open-source models to compete with Meta's Llama 2. Now, interestingly, a couple weeks ago, Alibaba also announced that it would be supporting Meta's Llama 2, and so it appears from this news that Alibaba is not putting all of its chips in any one basket, even its own.

6:08It is worth noting, however, that we kind of have to put open source in brackets, as it's similar to Llama 2's version of open source, which one might consider as mostly open source, however with a few different restrictions. In the same way that Meta set some limits around needing special approvals and permissions should monthly active users exceed a certain number, Alibaba's QEN 7B models come with a similar restriction. Finally today, if there was any doubt where Google Search is moving, the company has just released a number of updates for their experimental search generative experience. On Wednesday, August 2nd, they released a blog post called Three New Things You Can Do With Generative AI in Search.

6:45Now remember, the whole point of SGE is that in addition to Google's classic set of little blue links from all around the web, there is a top section that is generated by AI that brings together information that perhaps lives within those links, but is custom curated by Google's artificial intelligence. The biggest update this week is a move into the world of multimodality. As Google writes, sometimes it's more powerful to understand something by seeing it. So we recently brought images to even more AI-powered overviews. For example, when you search for something like tiniest birds of prey, you'll quickly be able to reference what the bird looks like and get relevant information from the web.

7:20And over the next week, you'll begin to see videos within some overviews where it's helpful to see something in motion, such as a demonstration of a yoga pose or how to get stains out of marble. Now, outside of that major substantive update, Google has also increased the speed with which results populate, saying that they've reduced the time it takes to generate AI overviews by half. And they're also adding small features like making sure that you understand when different links that are being recommended were published to help you decide which little pathways you want to follow down. There's an interesting bifurcation happening right now where the announcements from companies like OpenAI and Meta around generative AI tend to be big and technical and have implications for the way that the field develops.

7:58Google, on the other hand, is extremely focused, it seems, on the productizing of AI for regular people. In other words, videos showing up in the SGE experience may not be as big of an announcement as something like Code Interpreter, but it might be relevant for a whole lot more people. Anyways, guys, that is going to do it for today's AI Breakdown Brief. I appreciate you listening or watching, and I'll be back soon with the main AI breakdown.

8:46Supermanage AI magically distills your team's public Slack channels into a real-time brief on any employee, any time. Catch up on contributions, work in progress, challenges they're facing, sentiment, everything you need to show up ready for a truly meaningful conversation. And it's completely free. Visit supermanage.ai forward slash breakdown today to start making the most of your one-on-ones. And thanks again to Supermanage for sponsoring the AI Breakdown. Welcome back to the AI Breakdown. I think when we look back at the cultural battles that are fought around generative AI, one of the front lines is going to be in the realm of music.

9:23Music has always had a particular place when it comes to the disruption of new technologies. And in many ways, the introduction of Napster and the battle that ensued with music lawyers made that industry more prepared than many others for what would come over the next couple decades of technology. I don't think it's an accident that music is the industry where the incumbent power, i.e. the record labels, probably maintains more control relative to the startups than just about any other space. What's more, we've already seen AI start to cause that sort of immune response once again. Earlier this year, hard on my sleeve, which of course was the beat that you heard over the intro of this episode if you're listening as a podcast, really scared the hell out of music executives because of how good it was.

10:08It was an AI version of Drake and an AI version of The Weeknd, and the song was an absolute banger, going completely viral, getting millions and millions and millions of streams and downloads, before effectively the entire music industry freaked out and started throwing legal weight around like crazy, getting it pushed off of platforms. But of course, it's a digital artifact, and it's still all over YouTube and everywhere else. More playfully but no less significantly, we are constantly getting AI remixes like this one of Frank Sinatra singing Little John's Get Low that become sensations on TikTok or YouTube or wherever they're premiered, and serve if nothing else to remind us just how good this technology is.

10:45Earlier this year, Google released research on something that they called Music LM. In the same way that a mid-journey or a stable diffusion is a text-to-image generator, Music LM was meant to be a text-to-music generator. For example, here's audio generated from the prompt, the main soundtrack of an arcade game. It is fast-paced and upbeat, with a catchy electric guitar riff. The music is repetitive and easy to remember, but with unexpected sounds like cymbal crashes or drum rolls.

11:17Now part of what made MusicLM exciting is that Google actually made it available in its AI test kitchen. For those lucky enough to get access, yours truly included, you can play around with prompting Music LM and seeing how it does. Here's an example of the prompt, a simple classical melody evoking the feeling of fall, emphasis on a melodic top line played by flute.

12:05So

12:17now, of course, the example I gave has an inherent element of subjectivity, given that I asked it to evoke the feeling of fall, but both of these options that it came up with did a pretty good job of one being a simple classic melody and emphasizing a melodic top line that was played by the flute. Frankly, subjectively too, I think they did an okay job with this evoking the feeling of fall piece. The way that the test kitchen works is once you've got your two examples, you give the one that you think did a better job a trophy to help improve the model. In this case, I think the second did a slightly better job.

12:47Let's try another. A driving 1980 style electronic synthwave track in a minor key that sounds like it might be a dramatic interlude in a video game.

13:36Both of these frankly nail it. I guess I'll give the trophy to the first just because I kind of like it better. And I think this is exactly what gets me so excited about this type of technology is that music is so much about a feeling and a vibe that being able to describe not just an instrumentation and a key, but an emotional register, and start to see even these very nascent examples actually achieve that is pretty remarkable. Well, just this week, Google's Music LM got some competition, and that is Meta's Audiocraft. Like Music LM, Audiocraft is a way to generate audio and music from simple text prompts.

14:13Meta's announcement post writes, Imagine a professional musician being able to explore new compositions without having to play a single note on an instrument. Or a small business owner adding a soundtrack to their latest video ad on Instagram with ease. That's the promise of AudioCraft, our latest AI tool that generates high-quality, realistic audio and music from text. Now, interesting, AudioCraft isn't actually just one model. Instead, it's a combination of three models called MusicGen, AudioGen, and Encodec. MusicGen, as you might guess, is a music generation model. Meta writes, music tracks are more complex than environmental sounds, and generating coherent samples on the long-term structure is especially important when creating novel music pieces.

14:53Now, AudioGen is for the generation of audio that isn't necessarily music, but is perhaps environmental. Now, Encodec is a little bit different. They call it a state-of-the-art, real-time, high-fidelity audio codec-leveraging neural networks. Encodec is trained specifically to compress any kind of audio and reconstruct the original signal with high fidelity. And of course, when it comes to these tools, what we really want to know is how it actually sounds. You had a chance to hear the Music LM version, so let's try those same two prompts in the Hugging Space demo environment for Audiocraft. So here we go, a simple classical melody evoking the feeling of fall, emphasis on a melodic top line played by flute.

15:42Not bad. And now let's try a driving 1980s style electronic synthwave track in a minor key that sounds like it might be a dramatic interlude in a video game.

16:06again pretty good although maybe a little bit more major key than minor key there perhaps unsurprisingly i think when it comes to which of these tools is better that may in spite of the name of this episode not really be the most important question the more important question is likely to me to be how fast are people going to put this code into end user experiences that musicians can actually get their hands on. In my conversation with musician friends, there is a lot of mixed feelings about this stuff. On the one hand, it's scary. People are worried about their skills being commoditized and reduced to something that a computer does with the press of a button.

16:39At the same time, there's a lot of excitement and enthusiasm to get their hands on these models and actually start to try to use them as part of composition. My strong, strong feeling is that in the same way that there are tens of thousands of times the number of people who can play guitar or play piano or sing, as there are people who can take those skills and turn them into great songwriting or composition, that just making it easier for anyone to dabble with music and audio creation isn't going to undermine the fact that the best outputs are going to likely still come from the same people who are creating music right now.

17:13This is obviously an extremely nascent part of the AI space, and I can't wait to see how it develops. That's going to do it for today's AI Breakdown. Come join us on the AI Breakdown Discord, drop in your best creations with these tools, and until next time, peace.

17:39Thank you.

From the publisher

On today's episode NLW explores Meta's newly launched AudioCraft compared to Google MusicLM. Before that on the Brief, Goldman Sachs predicts AI investment will reach 4% of GDP by 2025; a 9-figure investment for an Nvidia competitor; new Alibaba AI models and more.
Today's Sponsor:
Supermanage - AI for 1-on-1's - https://supermanage.ai/breakdown
ABOUT THE AI BREAKDOWN
The AI Breakdown helps you understand the most important news and discussions in AI. 

Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe

Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown

Join the community: bit.ly/aibreakdown

Learn more: http://breakdown.network/

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
The Frontier of AI-Generated Music ModelsThe AI Daily Brief: Artificial Intelligence News and Analysis · 18 min
Listen in VO