ElevenLabs' Mati Staniszewski: How Voice Becomes the Interface for Everything

8 May 2026 · 27 min · 12 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode topic: Mati Staniszewski (co-founder, ElevenLabs) explains how voice becomes an interface for dubbing, voice agents, and music; why ElevenLabs built audio “frontier” models without huge compute; and what’s next (emotional intelligence, agent-to-agent communication, and trust/authentication).

Guest background

Mati Staniszewski, Polish (suburbs of Warsaw), co-founded ElevenLabs in 2022 with childhood friend Piotr; met in high school, later built a remote London/Warsaw research team focused on audio.

Key claims

Audio models can be built with smaller compute than other modalities if transcription/annotation is solved; early monetization helped fund R&D; emotionality and interactive turn-taking are harder than basic streaming; trust and watermarking will be critical when agents act on behalf of humans.

Notable examples

replicated Mati’s accent voice; “AI that can laugh” (Hacker News); viral multilingual speech (Javier Malay) and examples like Narendra Modi, Zelensky, and Matthew McConaughey speaking Spanish/Portuguese; Ukraine government voice agent for frontline info; Deliveroo restaurant call automation; Deutsche Telekom inbound sales via voice; Masterclass interactive Gordon Ramsay and Chris Voss negotiation-by-phone.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Human Story Behind Eleven Labs

0:46 to 2:30

Exploration of the personal journey and motivations behind founding Eleven Labs.

“And part of what started 11 Labs is inspiration from where we are both from.”

Audio Experience in Poland

2:31 to 5:08

Discussion about the peculiar audio experiences in Poland and their impact.

“Is that even possible in 2026, et cetera?”

The Problems in Audio Technology

5:09 to 8:05

Identifying the challenges in the audio domain and the potential for improvement.

“I think a lot of customers see 11 Labs through their narrow needs, right?”

Building Eleven Labs: Strategy and Approach

8:06 to 10:49

Insights into Eleven Labs' unique approach to building audio models and company growth.

“So we started getting those out and that was the moment for us because we made it to the top of Hacker News with the first AI that can laugh model, which was a very proud moment for us.”

The Evolution of Eleven Labs' Products

10:50 to 12:39

How Eleven Labs developed its suite of audio models and their functionalities.

“It doesn't replace the entire experience but takes and amplifies part of that experience.”

Wow Moments in Audio Technology

12:40 to 13:21

Key moments that showcased the capabilities of Eleven Labs' technology.

“which every citizen can access and get information about what's happening.”

Voice Agents: Current Trends and Opportunities

13:22 to 14:01

Discussion on the state of voice agents and their potential in various industries.

“and you can learn physics with them on the headphones while you are teaching that subject or learning that subject.”

Building a Small, Agile Company

14:01 to 17:40

Learn how a flat organizational structure enhances agility in a startup.

“So while you were in the kitchen, he can shout at you effectively to get better.”

The Future of Voice in Negotiation

17:41 to 19:56

Explore how voice agents may change negotiation dynamics in the future.

“Are you seeing people deploy voice agents to actually negotiate on their behalf?”

Trust and Communication with Voice Agents

19:57 to 21:40

Understand the importance of trust in the interaction between humans and voice agents.

“in the future where agents do more and more of the work.”
Show all 12 chapters

Audio Models and Their Challenges

21:41 to 24:04

Discover the challenges in developing effective audio models and their applications.

“Andrei spoke earlier about jagged intelligence.”

Creating Value Through Audio Technology

24:05 to 26:46

Learn how to combine audio models with user-centered design for impactful solutions.

“I'm a big fan of yours and they love labs.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:02So I love line charts and bar graphs as much as the next guy, probably more. The story of 11 Labs is also interesting from a human perspective, but as you started a company with a childhood friend. So maybe take us back to 2022 or earlier and just tell the human side of the 11 Labs story to start.

0:22Mati Staniszewski:I have the most luck in the story of 11 Labs because, well, it started in 2022. It feels like it started 17 years ago when I met my co-founder, Piotr. All the names in Polish are complicated, luckily, for us. But we met in high school, became best friends, took all the same classes together. And then through the years, did everything together. So we traveled together, studied together, worked together. And time is on our side. We are still best friends. It's working out. And part of what started 11 Labs is inspiration from where we are both from. We are both from Poland. suburbs of Warsaw. And there's a very peculiar thing in Poland.

1:01Mati Staniszewski:If you watch any foreign movie in Polish language, all the voices, whether that's a male voice or a female voice, get narrated with one single character. So as you can imagine, pretty terrible experience. You have literally one voice narrating everything. It usually also, on purpose, is kept in monotone. So you are meant to interpret your own emotions for that content. And while we grew up with this, this is still happening today for majority of content. And that kind of opened our eyes into one of the clear things across the domain, across audio domain, across the future, will be this ability for everybody to speak any language with the same emotion, the same intonation.

1:41And we started diving deeper into that problem and realized the problem of audio exists in so many other domains too.

1:47Mati Staniszewski:whether that's narrating the content around us, whether that's the books not being available in audio form, whether that's the news articles that we could read, whether that's that language barrier, or in the future, as we heard in the previous conversations, the future where human noise, the robots are around us, the voice will be the primary interface to a lot of the technology and something we would love to fix and solve. Excellent. And Eleven Labs builds frontier models for audio. I think there's a paradigm now where to build a frontier model, you have to start with hundreds of billions or billions of dollars and then figure out the rest later.

2:23Eleven Labs did not take that path. May I talk a little bit about your approach towards building this company, why this hasn't been replicated? Is that even possible in 2026, et cetera?

2:36Mati Staniszewski:I think that continues that great lack in timing because we started in 2022. to for those of you working the domain at the time. That was a year of crypto and metaverse. Nobody was still working on the AI side. Even further, people were starting to work, of course, on the text models, on the visual models. But audio as a domain was still considered a big niche. There's so few researchers in the space working on that work. So for us, that was a good part of picking that domain where, A, we were excited about where that future is called. We felt that people around just didn't realize the value of that domain.

3:14Mati Staniszewski:But three, the requirements of what you needed to solve were very different. The audio models were smaller. So you don't need as much compute as you need for some of the other sister domains. The data needs are big. But while there's a lot of audio data, we knew that the thing to actually get that audio working, you will need to figure out how to transcribe a lot of that data and annotate all the data, which we knew we can do. And then ultimately, it all boiled down to architectural side of can we solve that part in a good way. And here, my co-founder is one of the smartest people I know and a great researcher and has been able to assemble some of the best people in audio to help us.

3:52Mati Staniszewski:And we took a slightly untraditional approach at the time. We started in London. We had a lot of people between London and Warsaw and started a company in a completely remote way. So we wanted to hire the best researchers wherever they were. where we were going through the classic GitHub scraping and trying to reach people based on their work instead of based on their presence. And based on that work, we would reach out to those people. We would always share our samples and try to get them to join the team. And that's how we assembled the first set of people who we think are some of the best researchers in that audio domain.

4:27Mati Staniszewski:And through the years, they still help us crank a lot of those models into production. Then we launched the product. I think the slightly different approach we took was monetizing very quickly. So trying to get some of the revenue stream back so we can fund a lot of the work in the models. We tried to stay healthy on the margins so we can continue investing with the assumption that it's better for us to figure out that stream and be able to be independent in that development. But then as the ambitions grew, we knew that we needed to train models. So we, of course, brought a lot of money externally as well.

5:04Mati Staniszewski:And I think like projecting to today, one thing that's clear for us is there's still so many of those niches that people don't tackle that you can start with and then step by step start opening them up. I think a lot of customers see 11 Labs through their narrow needs, right? Maybe take a zoomed out view. Like what is the suite of models that 11 Labs works on? How do you prioritize them? How do you organize R &D, et cetera? So we started with the first text-to-speech model, so the model that could finally understand the context of what's being written. And based on that context understanding, get the right emotion, the right intonation from text.

5:45Mati Staniszewski:So if it was a happy sentence, you get that happiness out. If it's a dialogue, it can pronounce the dialogue out. And then continuously started adding that. So it started with the problem of breaking down language barriers. is the things you need to solve dubbing is transcription, so understanding, then the translation, and then text to speech. So we first solve text to speech. Then we knew we needed to add the other component, which is speech to text and being able to transcribe content in a great way. Then how we combine those models together. So that's kind of what the first three models in the first couple of years.

6:20Mati Staniszewski:And then, of course, the other thing started happening across the space, which is that a lot of the reasoning models started becoming quick enough and smart enough at the same time where you could imagine those interactive experiences being possible. And that's where we started launching more of the real time streaming models across audio, and then combining those into conversational experiences. So we added effectively all the stack, all the turn taking and orchestration to create a voice engine for a voice agent. And then on the other side, as we realized that the emotionality is something we can solve, we added some of the hardest modality in audio, which is music and being able to produce music.

6:58Mati Staniszewski:So today we span entirety of the research of audio, whether it's text-to-speech, speech-to-text, combining those models together in both localization with dubbing, with orchestration, with voice engine, and then being able to do that across music as well. And all those things and all that interesting development work, Was there any oh wow moment in terms of what these products are capable of that you can remember? You know, there's so many and it's kind of the bar changes for all of us. The first moment for us was, well, first moment for us, they always use my voice as a testing voice because it has this weird accent.

7:36Mati Staniszewski:And the first time was like when we could replicate my voice based on a good sample. That was like a first wow moment to myself. And you always go for this moment like this is not how my voice sounds like. and then you listen to yourself side by side and it's like definitely how it sounds like. Unfortunately. Then the second moment was where we first got it to laugh and people were like, okay, this is actually the thing that makes the whole experience more human. The laughter, the pauses, the ums, the ums, the imperfections. So we started getting those out and that was the moment for us because we made it to the top of Hacker News with the first AI that can laugh model, which was a very proud moment for us.

8:18Mati Staniszewski:And then, of course, through the years, kind of that extended where you might remember in 2023, 2024, there was a Javier Malay speech that went viral where he could speak other languages. It was translated into English. And it was the first time where he could still hear his voice out there. So that's the kind of continuous wow moment that was something that's completely impossible. And then we saw that happen time and time again with Narendra Modi, with President Zelensky, all the way to recently one of the, I feel like, pinnacles of the voice performance, Matthew McConaughey giving his newsletter and his iconic lines in Spanish and Portuguese, where for the first time his family who speaks that language could hear him speak those languages too.

9:06Mati Staniszewski:But for more recent pieces, the two ones that we are excited about bringing to production, I think the first one is finally figuring out the emotional intelligence in that interactive experience. So in the voice agent experience where it doesn't only get the right intonation emotion, but can understand the other side. So if somebody is stressed, it gets and delivers that soothing, reassuring emotion. If someone is excited, maybe it matches that. If someone speaks slowly, it makes sure to slow down. And that emotional intelligence is something that we are finally seeing internally a path to solving, which will be just a continuous step change to what's possible.

9:48Mati Staniszewski:And then the second one, which will apply there but also apply into general audio spaces, audio general intelligence where you can combine audio models together in one stream. So you could theoretically have a model that narrates, then pauses, and let's say starts singing with that same continuous voice. and that's something that's extremely hard to combine today and something that would be possible, I think, very soon. In voice, you mentioned voice agents, and it seems like everybody is, at least on the customer side, everyone's buying a voice agent. And I think intuitively you think customer support, the old phone tree replacement.

10:26What's actually going on in the world of voice agents and what do you think are the most interesting, overlooked opportunities, spots where startup founders should focus?

10:35Mati Staniszewski:Yeah, of course, the customer support is probably the one that everybody heard and knows about very well. I think the second thing and the second thread we are seeing is increasing shift to revenue-generating opportunities where voice agents can act in sales, whether it's inbound or outbound of sales. It doesn't replace the entire experience but takes and amplifies part of that experience. Maybe a good example is Deliveroo, where Deliveroo will have voice agents that contact the restaurants to capture their opening times. And based on their opening times, they can update the riders and drivers, and of course the people ordering on when to get to that work, all the way through to the inbound sales, where increasingly people, that's a good example of Deutsche Telekom, will be contacting to inquire about the service, inquire to buy a product.

11:24Mati Staniszewski:And instead of going through the dropdown, instead of going through the form, you can speak with the voice agent to leave that information. We do it ourselves, too, so we have a good metrics of an understanding of what's happening there. One, of course, so much simpler and quicker to go through instead of going through that form. But the second thing that started happening in that inbound sales flow is we had a lot more information that people started leaving because they would speak about the use case they're coming with, but then where it's not working, where it's working, some of the other use cases that they are evaluating, which we can combine and then just deliver such a much better experience afterwards.

11:58Mati Staniszewski:On the overlooked side, I think my favorite example is the citizen support education and healthcare will completely change. on the citizen support, like all of us would benefit from just generally better government access, whether that's understanding how to fill in the taxes that I think many of you went through earlier this month, all the way through to just learning what is the policy for travel abroad and how that might affect the space. We recently seen that work deployed in government of Ukraine, who we think is one of the most advanced governments on that front. We traveled to Ukraine working with their team.

12:39Mati Staniszewski:And what they are trying to solve is they have a government app, which every citizen can access and get information about what's happening. But given the war, given the front line and lack of that access, they wanted to figure out a new channel for people to be able to call in and get that information. So they created a voice agent effectively, where you can call in and get the information about what's happening on the front line. And you can get education help and some of the lectures delivered to your kids all the way through to proactive engagement about staying safe and staying out there. And maybe last example on education front, and that's probably my favorite one as I think about that changing.

13:19Mati Staniszewski:It's just how incredible would it be to have someone that is an incredible teacher available 24-7 where you can ask him questions, whether it's Karpathy, all the way through to Richard Feynman. and you can learn physics with them on the headphones while you are teaching that subject or learning that subject. And that's something that we are seeing pockets of. Like a great example is Masterclass, where Masterclass, of course, collaborates with incredible teachers to deliver static lectures. But recently, they launched an interactive version of that. So I don't know if that will be a good reference for this audience, but we recently worked with them on bringing Gordon Ramsay that can teach you cooking.

14:01Mati Staniszewski:So while you were in the kitchen, he can shout at you effectively to get better. Or maybe a better one, there's a Chris Voss where you can, of course, learn negotiation, but you can learn by negotiating with Chris live on the phone to get better, which I thought was a phenomenal subject. Having negotiated against Madi a number of times around financing rounds, I understand now. I think it helps you to say this, but I think the opposite is true. I have some more questions. I want to save time for the audience as well. Maybe one, as Constantine mentioned, more than$100 million of net new ARR in Q1.

14:36Obviously, the business is going very well. And you're sort of pioneering the startup founder, building a foundation model, applications. Any counterintuitive lessons about building a company in this era that for the founders in the audience, they might want to take home with them?

14:52Mati Staniszewski:So we are, just for reference, we are just over 400 people, over 400 million in revenue, but still keep the teams extremely small. So it's like rough, arbitrary, a little bit. It's less than 10 people. It's for each of the research product. Even the go-to-market ops talent teams are all smaller than that size. Most of people will have 10 direct reports or so. So it keeps it relatively flat and allows us to move a little bit quicker. One thing that we've done, which is in this model, and very surprisingly, this is a very similar model that we've seen actually with the government of Ukraine. Each of the teams, even the teams that aren't technical teams, will have engineers within them.

15:33Mati Staniszewski:So our people team, our go-to-market team, our legal team will have an engineer in that team that helps to build, of course, automation, upscale, uplevel the rest of the people. And recently, that really helped, because as I'm sure many of you are going through, So everybody will be vibe coding and coding a lot of the help, even if they are not technical. So now that kind of shifted the responsibility, but shifted the requirement of how good the review needs to be for a lot of that work. There's security, infrastructure implications. You will want to make sure that the output is right. And I think on the engineering side, you can put that expectation.

16:09Mati Staniszewski:On the non-engineering side, the ability to do that is relatively hard. So that technical resource in those teams helped us a lot to figure this out. And in general, there's just so many incredible work you can do by having that, whether that's the scraping on the hiring and recruiting front and analyzing what worked in the past to improve in the future, whether that's upskilling the legal team on how to use those tools and then figuring out ways of, we recently introduced this scoring system. For those on the go-to-market on sales side, you frequently will end up in this negotiation with your sales team of, can I give indemnity provisions?

16:46Mati Staniszewski:What's the liability cap? Can I give the set of clauses? And then you kind of need to draw the line of how many things you give. And I ended up being in so many of those conversations that we gave already a lot or we didn't. So now we introduced the scoring system that you can give per size of the customer. You can just give a few of those points out and in. We just made it so much easier. And of course, that's fully automated now with how we work across that team. So that was one of the unintuitive of small teams, bringing technical talent in the non-technical teams, keeping relatively flat. We also have no titles, which allows us to bring people and really optimize for impact that they are having.

17:26Mati Staniszewski:And then you can grow as quickly as you want. The tenure will not define this. And many more. So we'll see. It's a four-year-old company, so we'll see if that helps. JOHN MUELLER - Any questions? Oh, no. OK, Sonia. Are you seeing people deploy voice agents to actually negotiate on their behalf? And then when you, are you starting to see agents actually negotiate with agents? Sorry, I do three-part questions. When that world happens, do you think the agents are actually talking to each other the way that humans talk to communicate and negotiate? Or do you think it's beep, boop, beep, boop? Do you think it's, you know, it's all done instantaneously?

18:08Like, how's that world going to look like?

18:10Mati Staniszewski:So one, early inklings of that, we haven't seen any truly successful on the negotiation front. It was like more, you know, kind of order taking, what's the price, can we capture that, and then kind of goes back to the team, so not real negotiation. But there's a few startups that we see, especially on any, like, organizational shifts of, can I organize this event, calling, calling out of places, getting the price, and then calling again with, like, our budget. So that is happening, and I think this will shift. I think emotional intelligence, this is the big part that will start being important in a lot of that work, where it's not only the content that matters, but how you deliver, when you pause, that work.

18:47Mati Staniszewski:And then maybe the extreme version of that, which agents are not, like most of the people wouldn't do it, and they are not good at that. Today you will see a lot of interruptibility built in, where a human can interrupt the agent, but with negotiation you also want the opposite, where an agent will interrupt the human. It's kind of the extreme version of that. On the second part, on the agent-to-agent part, some of you might have seen this. We did a hackathon over a year and a half ago, and that was exactly the case, where agent was speaking with another agent. They detected that they are both agents, and they swapped over to a different language, and also get more efficient transmitter of information than just the classic spoken word.

19:32Mati Staniszewski:And I think this will happen 100%. Like the big question, will it be really voice? Will it be other transmission of information? And depends truly on what the infrastructure is built for. And I think this will define that experience. Yes.

19:52Mati Staniszewski:It's you at the catch box. Hey. Curious how you're thinking about the need for voice in the future where agents do more and more of the work. So basically, what are the kind of use cases maybe we're human conversation. I think it's more of a follow up to the last question. Like first, you all of us will have so many different devices around us. And step from that, you will have robots around us. So of course, voice will be such an important interface to instruct and be able to interact with those devices. In many ways, I feel like we see a lot of developments of intelligence. But then the real bottleneck of the future will be how we communicate with that intelligence.

20:33Mati Staniszewski:And I have voice and visual part will be a big unlock to be able to actually get the most of that intelligence value in those settings, which isn't yet possible. But on the flip side, the value of the human-to-human interaction will only increase. So whether that's the events like this one, whether that's events with your favorite artist, will increase in value with that ability of having voice all around you. But the trust will be such a big part and something we optimize for like in between the agent and human of you know in the future where all of you will all of us will have a voice agent for example to call and book a restaurant or give information to a health care appointment.

21:18Mati Staniszewski:and all of that will require such a high degree of trust that this is you and authenticated you. So there'll be like a level of encoding and decoding for real, then encoding, decoding for watermarked opted in human. And then by default, everything else will be fake, which is kind of the opposite of how it is today. You'll detect for AI, but it will detect for real authenticated AI in the future and assume it's fake.

21:50Andrei spoke earlier about jagged intelligence. Do you see similar odd places in audio where models are good and are bad that you might not expect? And what are they?

22:03Mati Staniszewski:There's still so much on the bad side. I think that we spoke a little bit about where we see the voice agents working. So this combination of the models together and support settings works really well, works reliably. early sales starts working, but the moment you start swapping to a true emotional interaction, not yet working. It doesn't get the emotion that well. It's slightly too slow. So that is still, I think, a big step change that should work. Same will apply in a very different domain on the music side. I think in the music side, you can get good production music. You cannot get top charts music, even with artist input.

Read the full transcript

22:47Mati Staniszewski:I think this will change over the next year or two. Yeah, of course. Andre's take was that the reason for that was that the labs were basically training for the stuff that had economic value. Where you're training your models, is that true of you? Are you basically training for the things that make the most money? Or is it that there are some challenges that are genuinely harder than others? We try to train the models, build the product and the ecosystem that will derive, of course, is the biggest impact for all our customers, all users, which should correlate, of course, with the revenue in the long term.

23:20Mati Staniszewski:So that long term perspective, it's going to be minimal in the next few years, so not next year. So frequently, we will train the models that might not provide that value in the short term, or even step before, we'll spend so much time labeling the data, not only the what of audio, but also how of audio. Like, what emotions did I use? What is my voice described as? What is this music described as? So we assembled a team of now 1 ,000 plus people that have been voice coaches, musicians, artists before that can help us annotate that behind the scenes. And that will not provide value in the next 6 to 12 months, but we think it will in the next 12 to 24.

23:57Mati Staniszewski:And then you often need to collect that data, which frequently just isn't that accessible as well. OK, last one, and then we'll go to lunch.

24:10Hey, you hear me? Thanks. I'm a big fan of yours and they love labs. Thank you. What do you think, from the model air perspective, what do you think are the modes here with audio models? The labs are going there, not going there. What are the kind of, in this sausage making of making a real good frontier audio model? What are the main defensible parts there?

24:37Mati Staniszewski:So of course, we do a variety of models. and recently had the pleasure of meeting Jensen, and he was commending on a few of those models, and he said that our speech-to-text models are technology, and text-to-speech is artistry, and we are all artists. So he gained a client for life. But, of course, we do believe there's a little bit of that to really fix text-to-speech and fix that emotionality. You will need to be really focused on that space. You really need to get in front of users, collect the data, collect the preferences, use that to fine-tune the models, And then there is a domain specificity in how you actually bring those models to production.

25:14Mati Staniszewski:And health care, very different than in financial services, very different than in education or experiences. So that's on the model layer. I think there will be continuous advantage that if you actually care about the quality, actually spending the time on the model work will help you keep that advantage. That's to your point. The models and a lot of use cases will use a model as just a small part of their stack. And that's where we spend a lot of time beyond the research on the product side of how you understand a user's problem, the workflow that they need. And voice agents is combining the audio models with knowledge and bringing that inside of the system, how you bring it outside with telephony systems so you can interact across channels, how you evaluate, test, and monitor.

25:57Mati Staniszewski:And then as you create, whether that's in the agent space, whether that's in the creative space, that same understanding, you build the ecosystem. And that's what we hope to build across 11 Labs, a place where, whether it's distribution and brand, that people can trust, the platform where you have pre-existing set of work that you can start off, whether it's a template for creating an agent, template for creating a workflow in a creative space, or whether that's a voice. And we had the pleasure now of having over 20 ,000 voices that people created, contributed, that you can use across language, styles, and voices.

26:29Mati Staniszewski:And I think that will be an increasingly important layer of how you are able to cater to that diversity and make it easy for people to start and really understand that work well. All right. I'm going to hand it back to Konstantin. Madi, thank you. Andrew, thanks for being a partner. Amazing. Thank you, guys.

From the publisher

Mati Staniszewski, co-founder and CEO of ElevenLabs, joins Sequoia partner Andrew Reed at AI Ascent 2026 to talk about how a four-year-old company built a frontier audio AI business with just over 400 people and over $400M in revenue. He explains why audio was overlooked in 2022 when the rest of AI was chasing text and images, why ElevenLabs chose to monetize from day one rather than raise indefinitely, and why he believes voice will be the primary interface for agents, robots, and the next generation of computing. Also: why emotional intelligence is the next frontier in voice, and what happens when one voice agent realizes it's talking to another.

More from Training Data

All 110 episodes
ElevenLabs' Mati Staniszewski: How Voice Becomes the Interface for EverythingTraining Data · 27 min
Listen in VO