#281 What Is AI Audio Separation with AudioShake CEO Jessica Powell

31 Mar 2026 · 39 min · 16 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI audio separation (source separation) that splits mixed audio into components (dialogue, music, effects, multiple speakers) so users get more control for editing, transcription/captioning, localization, and real-time broadcast workflows.

Guest backgrounds

Jessica Powell is CEO of AudioShake. She previously worked at Google for most of her career (started in London, worked in Tokyo, then Bay Area), including running communications. She studied at the University of Fribourg in Switzerland (pedagogy/education track) and learned French while there.

Key claims

AudioShake’s North Star is quality (not just speed). Separation is fully automated by deep learning (no human-in-the-loop for the audio processing itself). Better separation improves downstream ASR/caption accuracy in noisy media.

Notable examples

“Doctor Who” localization—separating dialogue/M&E from a director’s cut so German dubbing could be added. Live sports/broadcast use—isolating dialogue from crowd noise down to ~11ms. Real-time music removal for copyright compliance and rights-based replacement.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding AudioShake's Mission

0:45 to 3:01

Jessica explains the purpose of AudioShake and audio separation technology.

“So before we start, for those of you and for those of our listeners who don't know what AudioShake does.”

Jessica's Background and Journey

3:01 to 4:58

Jessica shares her professional history and the origin of AudioShake.

“And then when I was in Tokyo, though, did and my co-founder was there as well, did a lot of karaoke.”

The Karaoke Inspiration

4:58 to 8:02

The conversation delves into the karaoke experience that inspired AudioShake.

“I feel like the good answer would be something like as a little girl, I always wanted to, I don't know, I read Heidi or I don't know, right?”

Broader Applications of Audio Separation

8:02 to 10:09

Discussion on how AudioShake's technology evolved to address various media needs.

“But we did start with music because we're both musicians.”

Challenges and Innovations in Audio Processing

10:09 to 13:11

Jessica discusses the technical challenges and innovations in audio separation.

“So technically for you, and I want to talk about kind of the streaming component here as well, but this obviously is not like mission critical that it's fast.”

Challenges of Audio Separation

14:01 to 15:10

Learn about the difficulties in isolating audio in noisy environments.

“From a sound separation perspective, it's a lot of noise, right?”

Localization Use Cases with AudioShake

15:14 to 17:41

Explore how AudioShake contributes to localization in media.

“So the Doctor Who you just mentioned before, but can you tell us a bit more like how, what do you do in localization?”

AI in Live Interpretation

17:42 to 20:56

Discuss the challenges and potential of AI in live interpreting.

“Have you ever come across like an AI interpreting interpreting use case, like live interpreting?”

AudioShake Products and Solutions

20:57 to 23:14

Get insights into the various platforms and solutions offered by AudioShake.

“And there's no human in the loop in the audio, like audio shake process.”

Benchmarks and Evaluation in Audio Separation

23:15 to 25:57

Understand the evaluation metrics used in audio separation technology.

“And so data diversity also matters when you're training models and source separation.”
Show all 16 chapters

Impact of Generative AI on Audio Industries

25:58 to 28:00

Learn how the rise of generative AI has influenced the audio industry.

“and do a whole bunch of different tasks.”

Understanding Audio AI's Role

28:00 to 29:25

Learn about the misconceptions and realities of AI in audio processing.

“working in audio is like, yeah, duh, this is like, we have ears, you know, the same way that you would say video, like vision is an important input to a system, right?”

Investor Conversations and Market Dynamics

29:25 to 30:55

Explore the changing narratives around SaaS and AI in investor discussions.

“So, was this kind of understanding the world, was this part of the kind of investor conversation?”

Building a Customer-Centric Business

30:55 to 32:28

Discover how AudioShake grew through customer love and product-led strategies.

“Um, I think a lot of the structures that are also created are entirely created to optimize for VCE.”

Ethical Business Practices in Tech

32:28 to 36:08

Understand the importance of ethical considerations in building technology solutions.

“But what's really interesting, I think, about working in media is that people don't generally go into media unless they love media, right?”

Looking Ahead: Innovations for 2026

36:08 to 38:20

Get insights into upcoming features and innovations at AudioShake.

“I mean, a VC who really wants to dig a little deeper than just a five minute call would understand that this is like, It gives you a certain authenticity also.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jessica Powell:Basically what we're doing is we're separating audio and giving it to you in its components so that then you have much more control and usability over that audio.

0:13Hey everyone and welcome to SlaterPod. Today on the podcast we welcome Jessica Powell. She's the CEO of language tech platform AudioShake. So for those of you who don't know AudioShake, they specialize in AI audio separation. And I had to do a bit of digging and reading up on audio separation before this podcast to get myself up to speed. Very, very interesting company story, use case. So yeah. Hi, Jessica. And thanks so much for joining. Hi.

0:38Jessica Powell:Nice to meet you. We always ask where in the world our guests are recording is from today. So where are you based, Jessica? Where are you recording this from? I'm in San Francisco. San Francisco, the heart of it all. All right. So before we start, for those of you and for those of our listeners who don't know what AudioShake does. Give us the kind of elevator pitch, the 30-second, one-minute elevator pitch to set the scene. Oh my gosh, being in San Francisco, you have to elevator pitch all the time. So how about I not give you the elevator pitch, but some other version of it. What we fundamentally, what we do is we help make audio more usable for both human and machine workflows.

1:15Jessica Powell:That may still sound a little abstract. My parents always say it is. But what it essentially is, is we use audio separation. It's a field of AI known as source separation to separate mixed audio into its component parts. So that could mean in a broadcast that you're separating out the crowd noise, the sound of the squeaks of the shoes on the basketball court and the commentators. And meanwhile, there's a Kendrick Lamar rapping in the background. We can separate all of those elements out. And that just gives you a lot more control over what you do with that audio. So everything from workflows where you're just needing to boost or lower certain sounds, or you need to clean audio before it goes through other pipelines, think of transcription and captioning, of course, or where you simply need editing control in superhuman DAW heavy editing workflows.

2:08Jessica Powell:For example, you're separating dialogue and M &E for a production because you want to take that program into other languages. So basically what we're doing is we're separating audio and giving it to you in its components so that then you have much more control and usability over that audio. Now, there's a pretty interesting origin story behind the product. So maybe tell us a bit more about that, how you came up with that idea originally, and also just a bit about your background before co-founding AudioShake. I think you were with Google for quite some time. So just tell us a bit more about the background and then the origin story there.

2:45Jessica Powell:Sure. Yeah, I was at Google for most of my career. Started at Google in London and then ended up in Tokyo and then finally out here in the Bay Area with Google. Did a whole bunch of different things there. But my final years there, I was running communications across the company. And then when I was in Tokyo, though, did and my co-founder was there as well, did a lot of karaoke. You do a lot of karaoke when you're in Japan for business, for social, whatever it might be. And it was very, very fun. I discovered I love karaoke. But we'd always wish that we could sing along to any song, not just the songs, the re-records that are in those karaoke books that they give you.

3:28Jessica Powell:and so one of a thousand silly ideas that we had probably drunk while doing karaoke was like oh wouldn't it be cool if you could rip the vocals off of any song so you could karaoke to any song I don't want to pretend that the like heavens opened up and then all of a sudden we were like let's start a company to create no it wasn't anything like that it was just this idea that we had and then we both ended up back in the Bay Area I was at Google uh Luke my co-founder was at Platt, which is a fintech company out here. And he was already working on the AI side, leading data science there. And one day he came to me, he said, remember that, that silly idea that we had in Tokyo about being able to separate songs.

4:08Jessica Powell:Like, I think deep learning has gotten to the point where we could do that. And so just for fun, we started experimenting with that problem. And the first few passes that we had at it were, it sounded pretty terrible, But we were still really excited because very quickly we thought, well, if you could do this for a song, say remove the vocals, why couldn't you separate all the instruments? If you could separate all the instruments, why couldn't you separate all kinds of sound? And we got really, really excited because sound is working in sound. You're dealing with a ton of noise, a ton of frequency overlap.

4:40Jessica Powell:And if you can figure out, for example, how to distinguish a piano from a guitar, that means that you have a pretty sophisticated system that can then be applied elsewhere to other kinds of sounds. I want to take a step back even further before we go into the product in more detail. Like, I looked up your LinkedIn profile and you studied at Uni Fribourg in Switzerland. So I'm like, how did that happen? Tell us more about that. I feel like the good answer would be something like as a little girl, I always wanted to, I don't know, I read Heidi or I don't know, right? Like it'd be some, but it had nothing to do that.

5:16Jessica Powell:It had to do with like a terrible relationship and a boy, which was that I was dating a guy in Spain and we broke up. And I think I had like a hundred dollars and this was right after college. And I used that hundred dollars to buy a train ticket to where my aunt and uncle live. And they're in kind of in the Clia region. and um and I sort of showed up on their doorstep which I'm sure they loved and it was just like this moody heartbroken person and traipsed around the pre-alps for like two weeks just you know like wallowing in my misery and then my aunt was kind of like um you're 21 and you need a job like what are you what are you doing what are you doing here and uh and I was like I don't know like I didn't know.

6:03Jessica Powell:And, um, and she, she worked in Fribourg and so she had gone by the university there and seen that there was still a program that was open and it was in pedagogy, which, um, I'm sure for the European audience, everyone knows what pedagogy is, but it's not a word. Well, it's English. It's not a word you use that much in like in English in your day to day, at least in America. So I didn't even know what pedagogy meant. I didn't know it was just like another way of saying like an education degree. And I also didn't really speak any French. Um, Which was also a problem for doing a degree in French.

6:34Jessica Powell:But I did speak Spanish and the test to get in was a written test. And so I relied very heavily on cognates and somehow squeezed into a program where I ended up doing a do, which I guess is kind of like a. Anyway, I did some sort of degree or certificate at the university. It was a great experience. And I did learn French as part of that humiliating journey, being in a program to teach French when you didn't know French. I mean, yeah, Google's got to be present here in Zurich. Obviously, it's probably one of the top employers. So not in Fribourg, but here, close, right? One hour by train, one and a half.

7:12All right, let's get back to AudioShake and the product. So, like you mentioned the original use case that you had in mind in the karaoke, but how, like at what point did you realize that this was like bigger when you started the company, bigger than maybe the music use case? And then, yeah, tell us a bit more about like the other areas where it found traction. Well, we started again, I think from the very start, sitting out here in the Valley and

7:42Jessica Powell:working so like every single day with code, right? I think our heads pretty immediately went to what could you do with sound if you could work with it in a programming environment, right? If you had essentially structured data that would allow you to deal with audio in a reliable way. That was always sort of front and center in our minds. But we did start with music because we're both musicians. We were really interested in music. And again, we weren't really trying to start a company. It was like a fun hobby kind of thing. um but when we started first working uh we and again actually maybe this is relevant when we started we weren't we weren't thinking we were starting a company we thought it was just a fun party trick we showed it to friends that were in the music industry and we had already thought about things like remixing and stuff like that but they were like yeah but have you heard about sync licensing or you know anything the move that's happening towards like 5.1 mixes and immersive mixing that stuff was all again we played music as hobbyists but that stuff was all kind of industry type stuff.

8:42Jessica Powell:It wasn't anything that we were actively thinking about. But that got us excited that there actually was a market and an interest in stems and like separated sound. And so we went and spoke to some sync companies or sync departments at like the majors and everyone was immediately interested. And that was our entry point was music. And then, you know, from the sync, it then spread into other departments like A &R that would be working on things like creative remixing. And while we were doing that, the film people found us and they said, well, wait, if you could separate film, like if you can separate music, couldn't you separate like M &E tracks?

9:20Jessica Powell:And we were like, well, yeah, it's not entirely the same, but there's a lot of overlap there in that kind of problem. And so we very naturally got pulled into like M &E separation as well. And the very first project we did was actually Doctor Who, the kind of iconic BBC show, they wanted to bring it into, somebody purchased the rights in Germany, wanted to bring it into German, and all they had was the director's cut. So director's commentary, the original dialogue, and all of that great sci-fi soundtrack, all on the same track. And so that's where we came in, is we isolated the different components so that they could then put a German, a human German dub on top of that.

10:05That's so interesting. Can you just dwell on that a little bit? So basically the challenge was they just had the master giant file and yeah, it was just, okay, it was very, very hard. So technically for you, and I want to talk about kind of the streaming component here as well, but this obviously is not like mission critical that it's fast. It just has to be super high quality first, right? And I heard you talk on another podcast initially was really that the focus was on super high quality and then some of the other use cases where the speed was relevant came later.

10:39Jessica Powell:Yeah. I mean, quality is still our North Star on everything. When you're talking about our artistic workflows in particular, well, artistic workflows or a workflow in which separation is some component of a larger pipeline. And if you don't separate properly, you've all these downstream effects. We can talk about transcription and ASR accuracy in a moment, but that would be a really good example of that. Quality is paramount, right? That's why people use AudioShake is that we're the best at what we do. And we have our own models. And we've spent years investing in the research to build really, really great separation.

11:19Jessica Powell:But the initial workflows we were in were all post-production workflows. So again, on the music side, the track was already out there. And now you're just trying to find a new way to license it and splitting that track into the instrumental and the vocals or the different instrument stems now opens up that track for sync licensing, which is not always an open opportunity for older music, for example, kind of pre-2015. And then similar parallels, of course, on like the film and TV side, right? Then as we got better and better at building these models, we did start to get a lot of incoming around faster workflows.

11:59Jessica Powell:So examples of that would be kind of on the extreme end, you would just think of anything that's in real time where you're taking a streamed file coming in and being able to separate that. So for example, now we work with broadcasters where we are isolating dialogue. And there we can go as low as like 11 milliseconds. So we can take a live stream, like a live feed, isolate the dialogue and allow you to then, again, control the levels. And if you're ever watching a sports match, for example, and you can't hear the ref's call because the crowd is so loud. That would be a good example of where you might use it.

12:32Jessica Powell:Or there's bleed from, again, the crowd into the dialogue track. You can now split all of that. We can also remove music in real time. So music that's caught in the background. Again, if we're thinking about sports and broadcast, all of those things we have now models, special models, that allow it to work with streamed audio. So there, the performance aspect or the speed aspect is really, really important. You have to do different kinds of things to these models to get them to be like sufficiently lightweight to work in those environments. But we now today work both in post-production and live.

13:11Jessica Powell:So but they're sort of different model architectures depending on the use case. Yes. So when we did the research for this podcast, we also saw that you have a partnership with AI Media, the Australian company. So what you just described, that would be, I think you're integrated into the Lexi voice workflow. So the whole broadcasting, would that be where part of it would go via like an AI media partnership? They're unveiling a bunch of different things at NAB coming up. So I think I should leave, let them speak to that. But what I can talk about more generally that we do with them, which is already public, is we, this actually goes back to the ASR point that I've mentioned earlier.

13:50Jessica Powell:One of the big challenges with captioning media is that media, again, we think of it as really rich and beautiful and creative sound. From a sound separation perspective, it's a lot of noise, right? It's not a clean feed the way you and me talking right now on a podcast that's going to be tracked out where I have my recording and you have your recording. We're both mic'd up. We both have headsets of some kind. that's sort of like an ideal recording environment right but so much that happens in the media world is not like that so you can think of on set sound or again think of a sports stadium there's so much uncontrollable sound and frequency overlap if you try and take that raw audio and feed it into a transcription and captioning system quite a bit you're going to have a lot of errors and accuracy that you wouldn't have in a clean environment like this podcast right now.

14:46Jessica Powell:Even now, if you and I started talking over each other, that can create problems. But if we, again, think about where AI media is doing a lot of work, it's in these really rich media environments. And so if you can isolate the components of audio before it goes through transcription and captioning, you're going to have a clean input that means then that your ASR system is operating at its very, very best level. And so that's where we contribute. Got it. And I want to explore the film side as well and maybe the localization use case. So the Doctor Who you just mentioned before, but can you tell us a bit more like how, what do you do in localization?

15:25How is AudioShake relevant to these localization general use cases beyond having a rich catalog that needs to be backed up or, yeah.

15:37Jessica Powell:I mean, a lot of your listeners are bigger experts in this than I am, but I'll just speak to it from the perspective of the workflows that we get involved in. I'd say the most common ones we see are certainly dialogue and M &E separation. Certainly, that's kind of obvious why you would need to do that for older content, right? And that it's all on one track, that Doctor Who is a good example. There's no way, the way that, as I understand it, localization used to happen, is that you would basically silence the audio. So the way you would handle the fact that you didn't have multi-tracks or stems is that you would just silence all the audio and then have Florian speaking Swiss German across the track.

16:20Jessica Powell:And in some countries, as I understand it, they'd even have Florian doing all the voices of the track. um now you now are able to separate the dialogue the music and the effects from there it kind of depends on your workflow um i think if we're looking at um entirely human or primarily human workflows which is still the majority of what we see through audio shake is um uh even if ai is being used there's human in the loop but in these anyway and in these workflows you would have, you have your dialogue separated, you have your M &E, you're going to retain the M &E and then you're going to take say the Swiss German and you're then going to like stitch together the Swiss German dub and the M &E together.

17:07Jessica Powell:If you were doing an AI workflow, you would take that dialogue isolation, you would run it through transcription and voice cloning and so forth. And then again, you're going to take that outputted, cloned Swiss-German, stitch it back together with the M &E. And then I know that there's some that are like 100 % AI doing that and everything, but I think the majority of the work that we see tends to be either still 100 % human or human, but AI-assisted at different points. Very interesting. Have you ever come across like an AI interpreting interpreting use case, like live interpreting? For example, we did a research project some time ago for a client and it was a little early, maybe like one and a half, two years ago.

17:58But the problem was that when you have like real life, like remote interpreting sessions, sometimes the AI was like trying to translate the dog barking in the background and it just kind of produced gibberish. So, have you ever come across anything in that area? Has anybody ever approached you for that? I haven't seen anything on live interpretation, but from the way you're describing it, that feels very solvable.

18:21Jessica Powell:You need to be able to tell the machine what to listen to and what to ignore and then give it a clean input. Can you walk us a bit more through the products when we did the research? You have a developer platform, you have AudioShake Live, you have an enterprise platform. Just try to untangle this a bit. How would somebody come into you with a problem? How would they interact with you? Would they be on a platform? Would you, is it kind of a, is it a SaaS? Would they log in and use it? Just give us a bit more insights there. Sure. We have an on-demand platform that is largely used by enterprises and companies working across the media and tech space.

19:01Jessica Powell:Tends to be used more on the, I guess, media side. It's like a drag and drop. You're pulling in your asset. You're telling us what you want to do with it. And then we're separating it for you. Again, we have models that can separate instruments, sure, but also can isolate dialogue, music and effects, can remove music from an environment, but keep all the natural sound. So again, that's like a kind of a, how do you take older catalog where the licensing has expired and then make it so that you don't have the copyrighted music in it and you can put in new freshly licensed music. And then we have models that can separate multiple speakers.

19:38Jessica Powell:We also have models that are designed to boost the intelligibility of a speaker's voice. And so that's particularly useful in really noisy recordings. Got one through last night. Normally, this is all happening. I don't see the files and everything. But we had a special outreach from a customer that does a lot of work in unscripted and documentary. and they had some very sensitive recordings tied to different kinds of audio that was recorded via an Alexa or a Siri type device and wanted to see, could we separate all the different voices so that you could extract some very sensitive kind of material from that.

20:24Jessica Powell:So that kind of thing can all just be uploaded to the site. We separate it. We don't retain that audio. We always can do that on the API. And we also have, again, these real-time streaming technologies that people can implement as well. When you say you separate it, like there is a human in the loop at that point still somewhere? Or like that, and you just delivered a, no, not at all. We don't have any humans working at AudioShake, including me. Sorry. No, no, I'm human. The employees are human. And there's no human in the loop in the audio, like audio shake process. That's entirely deep learning.

21:07Jessica Powell:All of our human resources are going to building those models. When I was referring to human in the loop earlier, I was just talking about a lot of our media partners. You know, I think there was the hope when we first started, there was a lot of talk in the industry and customers that we would have about being able to fully automate pipe, like say localization pipelines. And I think that some of the companies we worked with that were in that space started off with 100 % automation and found that you couldn't get the quality that you needed because there were just too many points of failure in orchestrating all of these different AI systems together.

21:48Jessica Powell:And a lot of them moved into kind of human in the loop type scenarios where you then have, you might have audio shake at the start where it's isolating it. Then it's going to go through transcription. Your transcription is going to be boosted. But ASR still, even in clean, clean environments like you and me talking right now, ASR can still fail. Transcription can still fail. So having a human check on that, having a human check before, say, like an AI voice gets involved and after the voice is involved. So I think today what we primarily see are workflows that are either 100 % human, meaning the separation is happening and they're keeping the M &E file and then they're having a human Swiss German dub on top of it, or the human in the loop AI workflows.

22:33Jessica Powell:But yeah, the audio shape component of it is an automated component. When you're doing the separation of the voices, does the language matter at all or it really doesn't matter at all? Like if it's like a bunch of English speakers and you're separating or there are a bunch of, I don't know, Quechua speakers, would that matter? It can matter. Yeah. Because so it really matters when you're doing voice cloning, right? Which we don't do. But if you're doing voice AI, yes, like you would want to be able to train on Quechua as well as like Quebecois as well as Sutsu Deutz, right? Like you'd want those accents and those inflections that might matter a little bit less when you're talking about frequency overlap.

23:12Jessica Powell:But it does still matter. There are still cultural norms in how we speak and frequency differences in how we speak. And so data diversity also matters when you're training models and source separation. So are there like any kind of technical metrics and like eval benchmarks that you're looking at, that you're tracking and like, you know, benchmarking your progress? Because I think you posted somewhere that you came top in an evaluation of like models that Meta ran. But like what type of eval benchmarks are those? Yeah. Well, so generally, Meta created its own benchmarks. As they should. Well, so that's kind of a separate issue.

23:59Jessica Powell:But in terms of benchmarks that are used widely in the field, yes, we're state of the art. So we benchmark in source separation. It's something called the SDR score. That's what you use. We're state of the art on SDR.

24:14Jessica Powell:It's always helpful to have benchmarks because then you have some common thing that everyone's using and there's some good to benchmarks because there's something that a benchmark is evaluating that probably is important to your system. At the same time, if you're talking about sound separation, and I would imagine the same would be if you were thinking about other kinds of audio components. If I think about, again, voice AI is not our field, but if I was thinking about voice AI, I imagine there are similar constraints. When you think about sort of separation, of course it matters how another machine is telling you, you know, how did you score on this task?

24:53Jessica Powell:The problem is, is that there's also perception and what a human perceives as high quality may be different from what a machine perceives as high quality, no matter how robust you try and make those benchmarks. So it's totally possible to achieve great SDR scores, but not have output that sounds as good. So we've had models that scored higher on SDR, but that we don't deploy because we found another, like a better sounding model that doesn't score as high SDR, but that we know in all of our listening tests has much better perceptual quality. So we use both of those. Like, I think it's just sort of kind of, I think, part of running a business.

Read the full transcript

25:36Jessica Powell:Like, you need to know what the field is doing. You want to use, like, the same kind of language as the rest of your field. But you also have to develop ultimately for your customer and not for what a bunch of machines are saying is the best. So, but yes. And then the meta thing was a separate thing where they were creating, they were using generative models to separate sound and segment sound. and do a whole bunch of different tasks. And because of the nature of the task, they had to create entirely new benchmarks because the SDR would not actually be a good benchmark when you are essentially creating, inventing, generating audio.

26:18Jessica Powell:So that's probably like a little too into the weeds. But yes, on that particular meta evaluation, we did very well. So you guys started in 2020, right? that's when you started the company? I think it was later than that. We were, yeah, who knows what LinkedIn or other things say. We were playing around with models then. I'm trying to set up because basically, like, okay, there's starting the AI company before kind of ChatGPT and one after, right? So there's life before and life after. So like if you started before and you had to kind of power through that, like how was that if you did? Like you're building something in the AI space and then boom, this kind of revolution happens.

27:02Yeah, is that something that kind of shaped the trajectory or you were like kind of at that time when it hit, it was still very early days?

27:10Jessica Powell:I think it was a net positive for us rather than negative, though it had its, or it has its drawbacks too. I think on the positive side, when we first started, you know, I remember we went to raise money for the company out here in the Valley and everyone was just like, we hate music. We hate sound. We never want to hear about any of this. And like, why would you do this weird thing of separating sound? And what was really helpful, I think, in the generative boom was that all of a sudden, and this wasn't right away. I would say this has been more in the past two years. The same investors that were, you know, didn't understand why anyone would care about audio all of a sudden were like, oh, actually, if you're building a large generative model and it needs to understand the world, turns out audio is a pretty important input, to which anyone working in audio is like, yeah, duh, this is like, we have ears, you know, the same way that you would say video, like vision is an important input to a system, right?

28:09Jessica Powell:But that wasn't the case at all, right? Like everyone thought we were just doing this very weird niche thing when we started. And so actually having the large models come along and start talking about multimodal and talk about helping machines understand the world was, I think, useful into just generally furthering the understanding of audio AI, like the component of that. Where it was less helpful is, I suppose, you know, a lot of times people assume what we do is generative. It's not. We don't add any information. We're working with existing audio. It's a totally different area of AI. or they'll assume that we're working with models provided by the big tech companies.

28:56Jessica Powell:No, actually, those are our customers. Because again, what we do is just really, really different. So it has its advantages and disadvantages. But generally, I don't know, it's a pretty exciting time. It's pretty exciting to see what you can do when you can orchestrate these models together. And so it's cool being out here in San Francisco because you see it all kind of come together. So, then the narrative when you spoke to investors, yeah, as you said, changed a lot because you raised a Series A in October 2025 if our research is right. So, was this kind of understanding the world, was this part of the kind of investor conversation?

29:32Because right now we have this like bifurcation, right? It's SaaS and then, or it's AI and everything else, right? And if you're in the SaaS bucket, no, no, no, wait, no, hang on. We don't want to do SaaS. We don't want to invest in this, right? It's all going to be disrupted anyway. And if the eye bucket, probably, you know, they're waiting in line to write the check.

29:51Jessica Powell:I don't know. You know, I know I keep hearing SaaS is dead. But at the same time, I keep hearing that AI companies' revenue is fake. Like, I just think there's all these different narratives. We have a large number of enterprise customers that have essentially subscriptions to AudioShake that I think you would call that a typical SaaS business. We also have lots of users who are AI native that are used to more like paying for processing and like very and very similar kinds of structures to how your listeners might consume any of the if they're using an LLM at all. I think a nice thing about coming of age now and not inheriting a business structure that may be dated or it may not be is it gives you the agility to adjust.

30:46And I think it also removes the baggage and wisdom sometimes, which is unfortunate, of having these preconceived notions of how you have to structure a business and how you have to run it.

30:58Jessica Powell:Um, I think a lot of the structures that are also created are entirely created to optimize for VCE. Like they're arbitrary capitalist like structures, which like, I'm not going on some rant about capitalism here. I'm just saying like, in what world is it normal that I, like there's this brand of Muesli I like, and they're trying to push me into a subscription. Like I don't ever need to have a Muesli subscription ever in my life. And I'll buy the muesli and I'm happy to support the muesli company and I'm, and I will be a good customer, but I don't need to receive it every Monday. Right. And like, it's, there's so many bonkers constructs that happen across business just because of investors.

31:40Jessica Powell:And I think it's good for these different models to be thrown on their heads just to show how stupid some of them are. So I kind of welcome it. Ask me again in a year when I'm like freaking out about it, but right now it's, I think it's fine. Great take. Yeah. It's very, I don't need a Muesli subscription either, even though it's a Swiss German word. It's one of the very few Swiss German words that made it into the English language, right? So, I want to close on, like you mentioned that you have a large number of enterprise customers. How did you get them? I mean, what's the go-to-market, what's the sales process?

32:12Like, it seems like from the outset, very product-led, like one leads to the other. It's an actual problem you're solving. People come to you. you probably don't have a large sales force that's going knocking on doors.

32:25Jessica Powell:Yeah. We have, we still, I would say have grown largely by word of mouth. What's cool about like, you know, what's really cool about working, I'd say, particularly on the media and entertainment side, a little bit less on tech, or it's like a different, the different kind of transaction, I guess, on the tech side. But what's really interesting, I think, about working in media is that people don't generally go into media unless they love media, right? Like you don't, like people who work in the music industry love music because there's so much terrible, like just historical infrastructure baggage of the music industry, how it works of like the lack of rights, transparency, all this kind of stuff.

33:06Jessica Powell:Like you really have to love music to work in music. And it means that when people in music find something that they'd love, they tell you, like they tell their friends about it because they're so excited about it. And we really found that when we started, it was really, really important to us. I think as creatives, like I said, like both Luke and I play music. We're both published authors. Like we, we deeply value creative work. And it was really important to us in building something with AI that we build alongside industry and that we not ever have people feel like they were having something imposed on them.

33:49Jessica Powell:And because I saw when I worked at Google, I worked on book search and I had a team at YouTube and you would just see over and over again how sure these really cool technology

34:04Jessica Powell:innovations, like they could be quite disrespectful, I think, of creatives and the creative process And I think there's a balance because I think if everyone always signed off on everything, in some ways, things would not happen. On the other hand, that's a pretty violent experience, I think, to have as an artist where people are taking your work and taking something that you poured a lot of emotional and physical labor into and then just saying, well, this is the way things are now, right? At least without engaging with you in the conversation. So it was really important to us when we were doing something like separating content, right?

34:43Jessica Powell:Separating, which is content is art, right? Separating art that we took that technology first to artists and rights owners and that let them play with the technology and see how it could help them with reimagining work, with new kinds of distribution and monetization and see that there was really a there there. That was incredibly important to us and remains incredibly important to us. And I think that that sounds, you were talking about VC earlier, I think that kind of language sounds very airy-fairy, like intangible, like, you know, it doesn't sound like a money thing and it sounds just sort of philosophical and probably people would think that that would be a really bad business strategy.

35:27Jessica Powell:And maybe it was, maybe it wasn't. But I think where it really was, where it's been very helpful is I think people very much see us as true partners and want to build with us and work with us and collaborate and understand that we're trying to solve real problems in their industry. and off of the back of that, you know, the partnership then deepens. And that's been incredibly helpful. Again, it started very much just from a genuine philosophical perspective. Like we knew there was a way to grow much faster and we took the alternative path, which was to build with the people who build the content, right?

36:03Jessica Powell:And that's remained true to Audioshake. Like that's been there from the start. I mean, a VC who really wants to dig a little deeper than just a five minute call would understand that this is like, It gives you a certain authenticity also. Not a certain, like it gives you authenticity. Yeah, yeah, it does now. But I think, yeah. Yeah, no, absolutely. That's what I'm saying though. It was because the question was around like, how do you grow? Like the hypothesis we had was that, but it also came from an ethical point of view, right? Which is like, what kind of company do you want to build? Because like the easy business answer is go build the thing that requires no cooperation and no conversation at all, right?

36:45Jessica Powell:So if you were only optimizing for economic upside, that's 99 % of the time what you should be doing. But if you feel like there is an ethical component to what you're doing combined with you feel like it's smart business to be a good partner to the people that you're working with, then I think that that's a more sustainable business model too. But like it, it, it came out of just, again, like this is the kind of work that we ourselves do in our private lives. Like it just made sense to us intuitively to do it that way. Final question. What's in store for 2026? Anything you can disclose, any, any things that are, uh, up for launch, any new features, partnerships that you can, you know, tell us here.

37:30Jessica Powell:Yeah, we're pretty excited. We're headed to NAB, um, really soon, uh, in Vegas. Um, not excited about the Vegas part, but excited about the, you know, showing off some of the things. We have a pretty fantastic copyright compliance system that a lot of media companies are using to be able to remove, to detect and remove music, to get the accurate rights information from it, to be able to then replace that so they can then push that content to much broader audiences. And we can do that again in real time or in post-production. And we're also going to be showing a lot of our real-time tech at NAB and showing some of the new improvements there.

38:13Jessica Powell:So yeah, more separations, faster and faster is probably a way to sum up 2026. All right. Well, that was super, super interesting. I learned a ton. Thank you so much for doing this, Jessica. Thank you.

From the publisher

Jessica Powell, CEO of AudioShake, joins SlatorPod to talk about how AI-powered audio separation is making audio more usable for both human and machine workflows, and enabling new use cases across localization, broadcasting, and media production. 

Jessica emphasizes that early traction came from the music industry, particularly in areas like sync licensing and remixing. However, the company’s expansion into film and television happened organically as new use cases emerged.

The CEO explains that AudioShake’s core technology uses source separation to break complex audio into individual components such as dialogue, music, and sound effects. She describes how this allows users to gain precise control over audio for tasks like editing, transcription, and multilingual dubbing.

In localization, Jessica highlights how separating dialogue from music-and-effects (M&E) tracks enables both traditional dubbing and AI-assisted workflows, particularly for legacy content where original stems are unavailable.

Beyond localization, Jessica underscores the importance of clean audio inputs for speech recognition systems. In noisy environments like sports broadcasts or unscripted content, separating dialogue before transcription significantly improves accuracy.

Jessica also reflects on the broader AI landscape, noting that the rise of generative AI has increased awareness of audio as a critical modality. However, she distinguishes AudioShake’s work as non-generative, focused on extracting structure rather than creating new content.

The CEO discusses the current funding environment in the Bay Area and how the investor narrative has evolved leading up to AudioShake’s late 2025 Series A.

Looking ahead, Jessica points to real-time processing and copyright-compliant audio editing as key areas of innovation, as the company continues to expand its role in media and AI ecosystems.

More from SlatorPod

All 39 episodes
#281 What Is AI Audio Separation with AudioShake CEO Jessica PowellSlatorPod · 39 min
Listen in VO