Meta's Voice Revolution: Introducing Voicebox Clone for Multilingual Speech Synthesis

2 Mar 2024 · 12 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI Today Podcast Episode Notes

Episode Title

Meta's Voice Revolution: Introducing Voicebox Clone for Multilingual Speech Synthesis

Overview In this episode, the hosts discuss Meta's recent announcement of VoiceBox Clone, a pioneering technology that can generate voices in multiple languages using just two seconds of audio input. This advancement holds significant implications for multilingual speech synthesis and audio editing.

Key Topics Discussed

  • Concerns of AI Voice Misuse
  • Previous discussions about the potential for AI voice generators to produce fake content, leading to scams or impersonations.
  • Tools like Eleven Labs have emerged to detect AI-generated speech to combat these issues.
  • Introduction to VoiceBox
  • Described as a versatile AI for speech generation.
  • Capable of audio editing, sampling, and styling.
  • High-quality audio clip generation and automatic removal of background noise (e.g., car horns and barking dogs).

Breakthrough Features of VoiceBox

  • Minimal Input Requirement: Can clone a voice from just two seconds of audio.
  • Audio Editing Capabilities:
  • Similar to image editing in Photoshop, users can highlight segments of audio and command the AI to remove specific sounds.
  • Cross-Lingual Style Transfer:
  • Can replicate a speaker's voice in multiple languages (e.g., French, German, Spanish, Polish, Portuguese) using the same voice characteristics.

Implications and Use Cases

  • For Creators: Allows for easier creation and editing of audio tracks, which could revolutionize content production.
  • For Accessibility: Helps visually impaired individuals by reading written messages in their own voices.
  • In Gaming and the Metaverse: Provides natural-sounding voices for virtual assistants and non-player characters, enhancing the immersive experience.

Diverse Speech Sampling

  • VoiceBox has been trained on diverse datasets, enabling it to generate speech that reflects real-world conversational patterns rather than textbook language.
  • This feature helps overcome the limitations of traditional AI that often utilizes rigid structure in language.

Privacy and Ethical Considerations

  • Meta has chosen not to open source VoiceBox due to potential for misuse, particularly in voice cloning, which raises ethical concerns.
  • The episode stresses the importance of verifying identities in communications to combat potential voice cloning scams.

Performance Claims

  • Meta claims VoiceBox is 20 times faster than existing models and outperforms single-purpose models through in-context learning.

Conclusion The episode wraps up with the acknowledgment of VoiceBox's potential impact and encourages listeners to stay informed about advancements in AI technology while being mindful of the ethical implications that come with it.

Resources

  • Invest in AI Box: [AI Box Investment](https://republic.com/ai-box)
  • AI Box Waitlist: [Join Waitlist](https://AIBox.ai/)
  • AI Facebook Community: [Join Community](https://www.facebook.com/groups/739308654562189)
  • AI in Music: [Learn More](https://musicalai.pro/)
  • AI Models: [Discover AI Models](https://aimodelspro.com/)

For more information on privacy, refer to [Privacy Policy](https://art19.com/privacy) and [California Privacy Notice](https://art19.com/privacy#do-not-sell-my-info).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00A recent theme that we've talked about a lot on the podcast is the fact that there is a possibility people use AI voice generators to generate fake spam content, you know, they can trick your family members into thinking it's you calling and asking for money, and all sorts of other nefarious things. In response to this tools like 11 labs have released AI speech recognizers, you know, you can put the AI speech in and it can tell you if it was generated by an AI. And so there's kind of this whole industry and there's this whole, you know, there's a bunch of issues, there's a space. But today on the podcast, I want to talk about a really cool breakthrough that is happening by Meta in this space and how it kind of ties into what we've been seeing in the past in this field.

0:43So Meta just recently released VoiceBox. So essentially, they're calling it the most versatile AI for speech generation. And of course, that is a bold claim considering things like 11 Labs are pretty powerful tools and are able to, you know, train off of your voice and also do a number of other different voices doing text to speech. So what exactly is VoiceBox? When you're looking at kind of the general, essentially, it's a generative AI model that can help with audio editing, sampling and styling. And this type of tech, they say could be used in the future to help creators easily edit audio tracks, allow visually impaired people to hear written messages from friends in their voices, and enable people to speak foreign languages in their own voices.

1:24Essentially what this is, is just like Eleven Labs, it is a text to audio generator. But I think there are some really interesting breakthroughs that they've made with this that I haven't actually seen in other places. So one of the things is that VoiceBox essentially can produce really high quality audio clips and edit pre-recorded audio. And they can do a bunch of this editing automatically. So things like removing a car horn or a dog barking, all while still, you know, preserving the same context of the clip and style of the audio. So essentially what this does is similar to how like Photoshop, right, they have their new AI feature where you have an image on there, you select a part of the image, and then you text, you know, write in like, you know, add a sun in the background or add a bicycle on the road.

2:12And then Photoshop can add those, you're doing the same thing with audio now with Facebook's new model, where essentially you have your clip, you highlight something and you say, remove the dog sound or ins, you know, I mean, essentially, eventually, I would say you probably going to start doing things where you're saying like insert this sound or that sound, and you're just typing, you know, it's like text. So you don't actually have to go get the audio and it's able to generate it, which I think would be interesting. But at the moment, you are able to edit and remove things, you know, remove this car horn here, remove this dog barking there.

2:40And it's not, you know, like I've used a lot of audio editing, I obviously use a lot of audio editing software for this podcast here. And usually when something happens, you may have heard in the background, you know, a kid screaming or, you know, a car honking or sometimes my landscaper walking outside my window with a leaf blower. And in those cases, if it's super disruptive, you know, I'll pause the audio and wait till the wait till the sound is gone. But sometimes when it's just a single sound, I'll have to go and try to clip it out or try to lower the volume on that one split second. And there's a bunch of different things that, you know, need to be done.

3:13So this is going to be really interesting where you're not actually having to manipulate the audio in that way, you're simply able to just select it, tell it what to do, and it will do that for you. So I think that's really cool. And I have not seen that anywhere else. Talking about this new AI model meta said in the future, multi purpose generative AI models like voice box could give natural sounding voices to virtual assistants and non player characters in the metaverse. They could allow visually impaired people to hear written messages from friends and read read by AI and their voices give creators new tools to easily create and edit audio tracks for videos and much more.

3:45So I think they have a lot of really cool use cases for this beyond just text to audio. And I think that they're thinking about how they implement this into their overall tech stack, their overall infrastructure, their building. So when they're talking about the metaverse, yeah, so you know, there's a game and there's a non player character in there and NPC, and they could have that thing talking in any type of voice. Now, what I think is really interesting with this is essentially this thing can take in up like as little as two seconds of audio. You get two seconds of someone speaking, two seconds of audio of that person's voice, and VoiceBox essentially can match the audio style and they can use it for, and they can clone a voice off of it.

4:26So this is something that I think is rather unprecedented. I mean, when I'm looking at using something like 11 labs, you can use less, but typically they recommend you use five minutes of audio. They say you're not going to get anything much better with more than five minutes, but five minutes of audio, I'll usually throw in there. I just to be safe on 11 labs, these, this is saying two seconds. That's absolutely insane. Um, it really has to leverage AI heavily because with just two seconds, there's like, they have to, they have to guess, they have to do so much guesswork about what you sound like saying different things.

4:58And they just grab from those two seconds. So really, really impressive. It'll be interesting to see how that plays out. The other thing, like we, like I said, is they do speech editing and noise reduction. And so you can identify the segment and get it to crop it out. So sort of like an eraser for audio editing. The other thing that it does is, and this has a major use case, I think in the real world, which is cross lingual style transfer. So essentially, when you give it a sample of someone's speech or them talking in English, it can then go and use their same voice to speak French, German, Spanish, Polish, or Portuguese.

5:34those are the the current languages that they're um supporting i think it's kind of interesting that they're not like the most popular languages spoken on the planet um you know like polish is in there and stuff so that's kind of interesting but we'll see what other ones they roll out to in the future but i think overall this is really interesting and this is really uh i think this has got a really cool use case right like i go and you know want to go travel to poland for example and beyond just having to use uh you know google translate or something like that if you know you had something that could literally speak for you, for example, you could talk and they could hear your voice in that language in, but it's still your voice, right?

6:14It's not a translator. It's not someone else's voice. I think this is really interesting for all sorts of things, right? Like I've listened to international speeches from, you know, people in tech or politicians from other countries and, you know, trying to listen to what they're saying. And of course, they always have this like dubbed over voice of a person talking. And sometimes I feel like some of the some of the I don't know, I feel like something is lost when it's not the actual person's voice. I know it sounds funny. But like, if I'm listening to this big tech CEO unveiled this amazing thing, and the voice that's speaking just seems kind of like, I don't know, I'm sure people will be offended, like goofy, right?

6:52Like, I don't know, sometimes the voice just does not match the face. and I think that that actually is a detrimental. It probably takes something away from the presentation and so if I'm able to hear that person's voice speaking my native language, I think that's gonna be a massive benefit to a lot of people and I think it'll be really cool. So even when the speech samples and the texts are in different languages, the capability could really be used in the future to help people communicate in a really natural and authentic way even if they don't speak the same languages and I think that's awesome.

7:22One other thing that VoiceBox does is they have what they call diverse speech sampling. So essentially having learned from diverse data, VoiceBox can generate speech that is more representative of how people talk in the real world and in the six languages listed that we're just talking about above, right? German, French, Spanish, Polish, and Portuguese. And so essentially what they're saying here is a lot of times when you hear like a robot or an AI trained on how to speak, it's speaking I mean, almost like how you almost like if I was to go and study Spanish and like the words and the phrases and how it's said in the textbook is not necessarily indicative of how the people in that actually country in the actual country speak.

8:04You know, their slang or some of their phrases or words may not have gotten into that. And so unfortunately, when an AI model is trained off of some of that content, it doesn't always carry over. So this is cool because it is it's been trained on what they call diverse speech sampling, which essentially just means that these things talk how actual people talk in the real world in those languages using their own phrases. And it's not just direct translation like, you know, like, how are you doing? You know, in France, for example, a lot of people would say salut, which is just like, you know, salute, like what's up, you know?

8:38and so um it i think it's really cool that they have that in there um and this should be interesting for the translations as well um and i'd be curious to see at what level you know something like google translate does this in any case um one of the big things that is grabbing people's headlines right now um well so okay one other thing i'll say that's really cool you you throw something in there um you know like you you put a piece of text in and then it's like would you like voice a b c or and you can listen to like all the voices saying the same thing and it's kind of cool. But in any case, something that is driving or that is grabbing a lot of headlines is the fact that unlike Meta, who usually will open source a lot of these AI tools and approaches they're using, they're choosing not to make the model or the code for this public because they say that the potential for misuse is too dangerous.

9:27So kind of what we're talking about at the beginning of this podcast with, you know, people using these for spam or scam or all sorts of other things. I think this model, they're going to have to get a lot of things right, especially with the fact that you can take two seconds of audio and clone someone's voice. It's going to be game over. I think this is an important time to, you know, remind everyone you know that AI can clone voices and everything that you hear is not necessarily someone's actual voice. it's important to verify especially if there's any sort of phone calls asking you for money definitely verify the voice and a lot of people have talked about the fact that uh in today's day and age where there is now voice cloning coming up with code words to you know verify that you're the actual person talking um is important and so i think that is an actual thing that you'll probably want to do and this can become more prevalent as time progresses because whether facebook releases this or not inevitably um more of these tools will come come forward.

10:28And this is going to be something that already is done at a fundamental level, it's already possible, but it's only going to get more sophisticated in the future, which will have some really cool use cases like, you know, you can get that you could get something like this to do a whole voiceover for a movie, or maybe I'd use it for my podcast. Although I kind of like actually talking, I feel like there's some sort of value to a human being actually talking to you right now. But, you know, you can see all sorts of use cases for this. And so I think that it's important to remember that this is only going to get more complex.

10:57Now, the one other thing I will say that is really interesting with this entire model is that Meta claims this is 20 times faster than current models and that it outperforms single purpose models through in context learning by 20x. So I believe they're taking a shot at 11 labs, not 100 % sure, but they're claiming to be 20 times faster than a tool like 11 labs, which I think would be absolutely phenomenal you know if this thing was launched and released and of course I would be super interested in using it as I'm a big fan of 11 labs and everything they're doing over there so if this is any better I think this would be very interesting to look at so this is an area we'll definitely have to follow very closely into the future

From the publisher

In this episode, we explore Meta's groundbreaking announcement of Voicebox Clone, a cutting-edge technology capable of generating voices in any language with just two seconds of audio input, revolutionizing the field of multilingual speech synthesis.

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Today

All 897 episodes
Meta's Voice Revolution: Introducing Voicebox Clone for Multilingual Speech SynthesisAI Today · 12 min
Listen in VO