ElevenLabs’ Mati Staniszewski: Why Voice Will Be the Fundamental Interface for Tech

1 Jul 2025 · 1 h

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: ElevenLabs’ Mati Staniszewski - Why Voice Will Be the Fundamental Interface for Tech

Podcast Overview

  • Title: Training Data
  • Description: A podcast hosted by Sonya Huang, Pat Grady, and other Sequoia Capital partners discussing AI technologies and their implications for technology, business, and society.
  • Episode Title: ElevenLabs’ Mati Staniszewski: Why Voice Will Be the Fundamental Interface for Tech
  • Host: Pat Grady
  • Guest: Mati Staniszewski, co-founder and CEO of ElevenLabs
  • Episode Description: Discussion on how ElevenLabs focuses on audio innovation and reshapes text-to-speech technology, contextual understanding, and emotional delivery.

Key Themes and Discussions

Introduction to ElevenLabs

  • Founding Vision:
  • Started from a personal story involving language barriers while watching movies in Poland.
  • Aim to revolutionize audio experiences and make content accessible in original voices, eliminating poor narration experiences.

Audio Innovation Focus

  • Technical Focus:
  • Voice AI is different from text AI in data and model architecture.
  • Emphasis on building unique audio models that understand context and emotion.
  • Differentiation in the Market:
  • Staying focused on audio amidst the rise of multimodal foundation models.
  • Investing in building superior research models to compete with larger labs.

Transformative Products and Applications

  • Viral Moments:
  • Notable products include:
  • Harry Potter by Balenciaga (2023).
  • Darth Vader's voice in Fortnite using original clips from James Earl Jones.
  • Real-time translation features showcased in interviews (e.g., Lex Fridman with Indian Prime Minister Modi).
  • Product Innovations:
  • Development of tools for creating audio books, dubbing films into other languages, and building conversational agents.
  • Highlights on ElevenLabs’ capabilities to perform laughter and emotion through voice synthesis.

Industry and Societal Impact

  • Voice as a Primary Interaction Mode:
  • Vision of voice becoming the main interface for technology (akin to the historical importance of spoken language).
  • Applications in healthcare, customer support, and education, enhancing human interaction and automating mundane tasks.
  • Future Predictions:
  • Potential for immersive learning experiences through voice agents.
  • Advances in breaking language barriers with the concept of a universal translator.

Challenges and Considerations

  • Engineering Challenges:
  • Differences between text and audio models, particularly in understanding emotional context and producing human-like delivery.
  • Latency and reliability crucial for conversational agents.
  • Ethical Implications:
  • Concerns regarding voice impersonation and the development of moderation mechanisms to prevent misuse.
  • European Perspective:
  • Advantages of access to passionate talent and a multicultural environment.
  • Challenges include navigating regulatory landscapes and learning from successful tech ecosystems predominantly found in the US.

Final Thoughts and Future Vision

  • Mati posits that the future will involve voice technologies seamlessly integrating into daily life, allowing for more intuitive and engaging human-computer interactions.
  • Ambitious Goals:
  • Aiming for human-like voice interactions and exploring the potential of neural interfaces for real-time translation and communication enhancement.

Important Mentions in the Episode

  • Research Papers:
  • "Attention Is All You Need" - foundational paper in transformer models.
  • Technological References:
  • Tortoise-tts - an open-source text-to-speech model.
  • Cultural References:
  • The concept of the Babel Fish from "Hitchhiker’s Guide to the Galaxy" as an analogy for universal translation.

Conclusion The episode encapsulates the innovative spirit of ElevenLabs under Mati Staniszewski's leadership and the potential of voice technology to transform how we interact with machines and each other. The focus on audio innovations positions ElevenLabs uniquely within the AI landscape, offering powerful solutions for breaking language barriers and enhancing user experiences across various sectors.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Late 2021, the inspiration came from Peter was about to watch a movie with his girlfriend. She didn't speak English, so they turned it up in Polish. And that kind of brought us back to something we grew up with, where every movie you watch in Polish, every foreign movie watch in Polish, has all the voices. So whether it's male voice or female voice, still narrated with one single character, but not on a narration. It's a horrible experience and it still happens today and it was like wow We think this will change this will change we think that technology and what will happen with some of the innovations Will allow us to enjoy that content in in the original delivery in the original incredible Voice and and and let's let's let's make make it happen and change it

1:04Greetings. Today we're talking with Maddie Stanishevsky from 11 Labs about how they've carved out a defensible position in AI audio, even as the big foundation model labs expand into voice as part of their push to multi -modality. We dig into the technical differences between building voice AI versus text. It turns out they're surprisingly different in terms of the data and the architectures. Maudi walks us through how 11 labs has stayed competitive by focusing narrowly on audio, including some of the specific engineering hurdles they've had to overcome, and what enterprise customers actually care about beyond the benchmarks.

1:46We also explore the future of voice as an interface, the challenges of building AI agents that can handle real conversations, and AI's potential to break language barriers. Maddie shares his thoughts on building a company in Europe, and why he thinks we might hit human -level voice interactions sooner than expected. We hope you enjoy this show. Maddie, welcome to the show. Thank you for having me. I have first question. There's a school of thought a few years ago when 11 labs really started ripping that you guys are going to be a roadkill for the foundation models. And yet here you are still doing pretty well.

2:26What happened? Like how were you able to stave off the multi -modality, you know, big foundation model labs and kind of carve out this really interesting position for yourselves? It's exciting last few years and it's definitely true. We still need to keep on our toes to be able to keep winning the fight of foundation models. But I think the usual and definitely true advice is staying focused and staying focused in our case on audio both as a company, of course, the research and the product, but we ultimately stayed focused on audio, which really helped. But the, you know, probably the biggest question on the old question is, is through the years we've been able to build some of the best research models and and out compete the big labs.

3:14And here, you know, credit to, to, to, to, to micro funder, who I think is a genius Piotr, who has been able to both do some of the first innovations in this space and then assemble a rock star team that we have today at the company that is continually pushing what's possible with Nodio. And as I, you know, when we started, there was very little research done in all of you. Most people focused on LLAMS. Some of you focused on image, you know, a lot more easy to see the results, frequently more exciting for people doing research to work in those fields. So there's a lot less focus put on to audio and the set of innovations that happened in the years prior, the diffusion model, the transform and models, were really applied to that domain in an efficient way.

3:57And we've been able to bring that in in those first years where for the first time, the text to speech models were able to understand the context of the text and deliver that audio experience in just such a better tonality and emotion. So that was the starting point that really differentiated our work to other works, which was the true research innovation. But then fast following after that first piece was building all the product around it to be able to actually use that research. It's, you know, we've seen so many times, it's like, it's not only the model that matters, it also matters how you deliver that experience to the user.

4:34Yeah. And in our case, what is narrating and creating audio books, what is voice -alvers, what is turning movies to other languages, what is adding text to speech and in the agents or building the entire conversational experience, that layer keeps helping us to win across the foundational models and hyperscalers. Okay, there's a lot here and we're gonna come back and dig in on a bunch of aspects of that, but you mentioned your co -founder Peter I Believe you guys met in high school in Poland. Is that right? Can you kind of tell us the origin story of how you two got to know each other and then maybe the origin story of how this business came together?

5:11The I'm in the I'm probably in the lockiest position ever. We met with 15 years ago in high school and the the we We started an IB class in and Poland and Warsaw and took all the same classes. So, so kind of everything and we headed off pretty quickly on some of the mathematics classes. We both love mathematics so we started about sitting together, spending a lot of time together and now kind of morphed from outside the school and time together as well. And then over the years we we kind of did it all from from living together, starting together, working together, traveling together and now 15 years and we are still best friends you know with the time is on our side.

5:56I was building a company together strengthen the relationship or there are ups and downs for sure but I think it did I think it did I think it's it's a battle tested it definitely battle tested it I know it's like you know when when the company started taking off we it's hard to know how long the horizon of this intense work will will happen. Initially it was like, OK, this is the next four weeks, we just need to push trust each other that will do well on different aspects and just continue pushing. And then there's another four weeks and another four weeks. And then we realized actually this going to be in the next 10 years.

6:36And there was just no real time for anything else. We would just do 11 laps and nothing else. And then over time, and I think this happened organically, but looking back at it's definitely helped. We now try to still stay in close touch in what's happening in our personal lives where we are in the world. And spend some time together still speaking about work, but outside of the work context. And I think this was very healthy for now. I know here for so long. And I kind of seen him evolve personally for those years. but I can still stay in close touch too. It's important to make sure that your co -founder and your executives and your team are able to bring their best self to work and not just completely ignoring everything that's happened on the personal front.

7:26Exactly. And then to your second question, part of the inspiration for 11 Lops came. So maybe the longer story. So there are two parts. First, for the years, when he was at Google, I was a parent here. We would do a Huck Weekend projects together. OK. So like trying to explore new technology for fun, and that was everything from building recommendation algorithm. So we tried to build this model where you would be presented with a few different things. And if you select one of those, the next set of things you're presented with gets closer and optimizes closer to your previous selection. Deployed it a lot of fun.

8:02Then we did the same with crypto. We tried to understand the risk in crypto and build like a risk analyzer for crypto. very hard. They're fully work, but it was good attempt in the first, one of the first crypto heights to try to provide like the analytics around it. And then we created a project in audio. So we created a project which on the last how we speak and gave you tips on how to win was this early 2021. Okay, early 2021. That was kind of the first opening. This is what's possible across audio space. This is the state of the art. These are the models that do with the erasation, understanding of speech, this is where the speech generation looks like.

8:42And then late 2021, the inspiration came from, and like the more of the aham moment, from Poland, from where you're from, where in this case Peter was about to watch a movie with his girlfriend, she didn't speak English, so they turned it up in Polish. And that kind of brought us back to something we grew up with, where every movie you watch in Polish, every foreign movie, or it's in Polish, has all the voices. So whether it's male voice or female voice, still narrated with one single character, like none of that is narration. It's like a horrible experience, and it still happens today. And it was like, wow, we think this will change.

9:21This will change. We think that the technology and what will happen with some of the innovations will allow us to enjoy that content in the original delivery, in the original, incredible voice. and let's make it happen and change it. Of course, they're an expanded since then. It's not only having realized the same problem exists across most content not being accessible in audio, just in English. How the dynamic interactions will evolve. And of course, how the audio will transmit the language barrier to. Was there any particular paper or capability that you saw that made you think, okay, now is the time for this to change?

10:02Well, attention is all you need is definitely one which which which which you know was so so crisp and and clear in terms of what's possible. Yeah, but maybe to to give like a Different angle to the to the answer I think the interesting piece was was less than the paper. There was this incredible open source repose That was like slightly later and as we started discovering like is it even possible and And there was a Tortoist TTS effectively, which is a model, an open source model that was kind of created at a time. It provided incredible results of replicating a voice and generating speech. It wasn't very stable, but it kind of have some glips into like, wow, this is incredible.

10:52And that was already as we were deeper into the company. So like maybe a first year in So in 2022 and but that was that was like another element of like Okay, this is this is this is possible some great ideas there and that of course we've spent most of our time like what other Things we can innovate through start from scratch bring the transforming diffusion into the audio space and And that that kind of heal that's just another level of of like human quality where you could actually feel like it's a human human voice. Let's yeah, let's talk a bit about how you've actually built what you've built as far as the product goes.

11:34What aspects of what works in text port directly over to audio and what's completely different, different skill set, different techniques. I'm curious how similar the two are and where some of the real differences are. The first thing is, you know, that there's kind of those three components that come come into the model that there's the compute, there's the data, there's the model architecture. And the model architectures is high -sem ideas, but it's very different. But then the data is also quite different in both in terms of what's accessible and and how you need that data to be able to train the models.

12:13And then compute the models are smaller, so you don't need as much compute, which allows us to, given a lot of innovations, need to happen and on the model side or the data side, you can still out compete foundation on models rather than just the use of data. They compute disadvantage. Exactly. But the data was, I think, the first piece, which is different, where in text, you can reliably take the text that exists and it will work. In audio, the data, first of all, there's much less of the high quality audio that actually would get you the result you need. And then the second, it frequently doesn't come with transcription or with a high accurate text of what was spoken.

12:59And that was the kind of lacking in the space where you need to spend all of time. And then there's a third component. Something that will be coming across in the current generation of models, which is not only what was said, so the transcript of Dario, but also how it was it said. What emotions did you use, who said it, what are some of the nonverbal elements that were said? That kind of almost doesn't exist, especially at the high quality. And that's where you need to spend a lot of time, that's where we spent a lot of time in the early days too, of being able to create effectively more of a speech -to -text model, and like a pipeline with additional set of manual laborers to do that work.

13:41And that's very different from text where you just need to spend a lot more cycles. And then the model level, you effectively, you know, you have this step of in the first generation that takes the speech model, you have understanding the context and bringing that to emotion. But of course, you need to kind of predict the next sound rather than the predict the next text talking. And that both depends on the on the on the on the prior but can also depend on the on the what happens after. Like an easy example is like, you know, what a what a wonderful day. And let's say it's a passage of a book.

14:22Then you kind of think, okay, this is positive emotion. I should read it in a positive way. But if you have a what a wonderful day. I said sarcastically, then suddenly it changed the entire meaning and you kind of need to adjust that in the audio delivery as well. Put a punchline and different in the different spot. So that was definitely different where that contextual understanding was a tricky thing. And then the other model thing that's the very different you have, the text to speech element, but then you have also the voice element. So the kind of the other innovation that we spend a lot of time working on is how can you create and represent voices to a higher, accurate way of what was in the original.

15:01And we found like this decoding and coding way, which was slightly different to the space. We weren't hard coding or predicting any specific features. So we weren't like trying to optimize is the voice male or is the voice female or was the age of the voice. Instead we effectively let the model decide what the characteristics should be. And then I found a way to bring that into the speech. So now of course when you have the text to speech model, it will. will take the context of the text as one input, and the second will take the voice as a second input. And based on the voice delivery, if it's more calm or dynamic, both of those will merge together and then give the end output, which was, of course, very different type of work than the text models.

15:50Amazing. What sort of people have you needed to hire to be able to build this? I imagine it's a different skill set than most AI companies. There's a... and it kind of changed over time, but I think the first difference, and this is probably less skill set difference, but more approach difference. We've started fully remote. We wanted to hire the best researchers wherever they are. We knew where they are. There's probably like 50 to 100 great people in audio, based at least on the open source work or the papers that they release or the companies that they worked in. That's a bit of admire. So you, it's at the top of the funnel, it's pretty limited because so much fewer people worked on the research.

16:36So we decided let's attract them and get them into the company wherever they are and that kind of really helped. The second thing was you know given, given we want to make it exciting for a lot of people to work, but also we think this is the best way to to run a lot of the research. We try to make the the researchers extremely close to to deployment to actually seeing the results of their work. So the cycle from being able to research something, to bringing it in front of all the people is super short. And you get that immediate feedback of how is it working. And then we have a kind of separate and research.

17:14We have research engineers that focus less on like the innovation of the entire kind of new architecture of the models, but taking existing models, improving them, changing them, deploying them at scale. And here frequently you've seen other companies called our research engineers researchers, given that the work would be, would be, ask, ask complex in those, in those companies, but, but that kind of really helped us to create a new innovation, bring that innovation extended and deploy it. And then the layer around the research that we've created is probably very different where we effectively have now a group of voice coaches, data labelers that are trained by voice coaches to understand the audio data, how to label that, how to label them motions and then they get re reviewed by the voice coaches, whether it's good or bad.

18:03because most of the traditional companies didn't re -support audio labeling in that same way. But I think the big difference is you need to be excited about some part of the audio work to really be able to create and dedicate yourself to the level we want. And we're a special at the time, small company, you would be willing to embrace that independence, that high ownership that it takes, that you are effectively, you know, working on a specific research theme yourself and, and of course there's some some interactions, some guidance from others, but a lot of the heavy lifting is individual and creating that work, which takes a different mindset.

18:51And I think we've been able to now we have like a team of 15 research and research engineers, almost, and they are incredible. What have some of the major kind of step function changes in the quality of the product or the applicability of the product been over the last few years? I remember kind of early, I think it was early 2023 -ish when you guys started to explode or maybe late 2023, I forget. And it seemed like some of it was on the heels of the Harry Potter, Bollinciaga video that went viral where as an 11 labs voice that was doing it, it seems like you've had these moments in the consumer world where something goes viral and it traces back to you, but beyond that, from a product standpoint, what have been kind of the major inflection points that have opened up new markets or spurred more developer enthusiasm?

19:39You know, it's there, what you've mentioned, it's probably one of the key things we are trying to do, and continuously even now, we see this is like one of the key things to really get the adoption out there, which is have the prosumer deployment and actually bringing it to everyone out there when you create a new technology, showing to the world that it's possible, and then kind of supplementing that of the top down, bringing it to the specific companies we work with. And the reason for this is kind of twofold. One is these groups of people are just so much more eager and quick to adopt and create our technology.

20:14And the second one frequently when you, when you, when we create a lot of the, the, the, both the product and the research work, the set of use cases that might be created. We have of course some predictions, but there's just so many more that we have what wouldn't expect like the example that we didn't have come to our mind that this is something that people might be, be creating and trying to do. And that was definitely something where we continuously even now when we create new models, we try to bring it to the entirety of the user base, learn from them and increase that. And it kind of goes in those waves where we have a new model release, we bring it broad, then kind of the prosumar adoption is there, and then the enterprise adoption follows with additional product, additional reliability that needs to happen.

21:05And then once again, we have a new step release in a new function and kind of the cycle repeats. So we tried to really embrace it. And through the history, the first one, the very first one was when we, when we had our beta model, so that you, you're right, just like when we released it publicly early 2023, okay, late 2022, we're iterating in the, in the beta with subset of users. And we had a lot of book authors in that subset. And we had this like literally a small text box in our product where you could input the text and get this speech out It was like a tweet length effectively and we had one of those book authors copy paste his entire book inside this box download it Then at the time you wouldn't you it was most of the platforms bandai content, okay, he managed to upload it They thought it's human He started getting great reviews on that platform and then came back to us with set of his friends and other book authors saying like, hey, we really need it.

22:10This is incredible and that kind of triggered his first like many virality mind with book authors. Very, very keen. Then we had another similar moment around the same period where we that there was one of the first models that could that could laugh. Okay, we released this blog post that the first AI that can love. And people picked it up like wow this is incredible this is really working we got a lot of the early users. Then of course the theme that you mentioned which was a lot of the creators and I think there's like a completely new trend that started around this time where it shifted into no um no face channels effectively you don't have the the creator in the in the frame and then you have narration all that creator across something that's happening.

22:52And that started going like like, wildfire and then the first six months of the work where of course we were providing the narration and the speech and the voices for a lot of those, a lot of those use cases and that was great to see. Then late 2023, early 2024, we released our work and moved other languages. That's one of the first moments where you could, you could really create the narration across other most famous European languages as an our dubbing product. So that was kind of back to the original vision. We created finally a way for you to have the audio and bring it to another language while still sounding the same.

23:33And that kind of triggered this other small virality moment of people creating the videos. And there was this, you know, the expected ones, which is just the traditional content, but also unexpected ones where we had someone trying to dub singing videos, which the model we didn't know will war on. And it kind of didn't work, but it gave you like a drunken singing result. So that is what Vyutai is viral too for that result, which was just fun to see. And then in 2025, at the early time now, and we are seeing I kind of, it's recurrently now, everybody is creating an agent. We started adding the voice to out of those agents.

24:17and it became both very easy to do for a lot of people to like have their entire orchestration, speech to text, their alarm responses, text to speech, to make it seamless. And we had now a few use cases which started getting a lot of traction, a lot of adoption. or most recently we worked with Epic Games to recreate the voice of Darth Vader. Yes, well done. Which players, there's just so many people using and trying to get the conversation of Darth Vader and Fortnite, which is just immense scale. And of course, most of the users are trying to have a great conversation, and use him as a companion in the game.

25:03Some people are trying to like, stretch whether he will say something that he shouldn't be able to say. So you see all those attempts as well. But luckily the product is holding up and it's actually keeping it relatively both. Performative and safe to actually keep him on the rails. And thinking about some of the dabbing, in the case of the one of the viral ones, was when we worked with Lex Friedman and he interviewed prime minister Narendra Modi. and we turned the conversation which happened between English, Prolex and Narinamodi spoke Hindi, and we turned the conversation into English, so it could actually listen to both of them speaking together.

25:40And then similarly, we turned both of them to Hindi, so we had a lot of legs speaking Hindi. And that went also extremely viral in India, where people were watching both of those versions, and then you asked people were watching the English version. So that was like a nice way of tying it back to the beginning. But I think they, especially as you think about the future, the agents and like just seeing them pop up in new ways is going to be self -requient. Like both early developers building, building everything from stripe integration and being able to process refunds through to the companion use cases all the way through to the true enterprise is kind of having probably a few viral moments ahead.

26:22Yeah, same more about what you're seeing in voice agents right now. It seems like that's quickly become a pretty popular interaction pattern. What's working? What's not working? You know, where are your customers really having success? Where are some of your customers kind of getting stuck? And then before I answer, maybe a question back to you. Do you see, do you see a lot more companies building, building agents across the companies that are coming through, Sikra? We, yeah, we absolutely do. And I think most people have this long -term vision that it's sort of a Agents style avatar powered by an 11 labs voice where it's this human -like agent that you're interacting with and I think most people Start with simpler modalities and kind of work their way up So we see a lot of text -based agents sort of proliferating throughout the enterprise stack and I imagine there are lots of consumer applications for that as well But we tend to see a lot of the enterprise stuff It's it's similar definitely what we are seeing both on the new startups being created where it's like everybody is building an agent And then enterprise like too, it's like, it's so helpful for the process internally.

27:24And like taking a step back, what we think and believe from kind of the start is, voice will fundamentally be the interface for interacting with technology. It would be one of the most, you know, it's probably the modality we've known from when the human, when genre was born, was the kind of first way the human's interacted. And it carries just so much more than text does. Like, it carries the emotions, the intonation, the imperfections, we can understand each other, we can, we can, based on, on the emotional cues responding in a very different ways. So that kind of, that's where our, our start happened where, where we think the voice will be that, that interface and build not just the text to speech element, but seeing our clients try to use the text to speech and do the whole conversation application can we provide them a solution that helps them abstract this away.

28:19And we've seen from the traditional domains, and to speak for a few, it's like in healthcare space. We've seen people try to automate some of the work they can do with nurses, as an example, a company that's a bit of a critic will automate the calls that nurses need to take to the patients to remind them of taking medicine, ask how they are feeling, capture that information back, So then the doctors can actually process that in a much more efficient way and And voice became critical where a lot of those people cannot be reached otherwise in the voice calls just the easiest thing to do then Very traditional probably that the the quickest moving one is customer support so many companies Both from the call center and the traditional customer support Trying to build the voice internally in the companies companies, whether it's companies like Deutsche Telekom, all the way through to the new companies, everybody is trying to find a way to deliver better experience and now voice is possible.

29:24And then what is probably one of the most exciting for me is education where they could you be learning through having that voice delivery in a new way. I'm a, I used to at least be a chess player or like a amateur chess player and and we work with chess .com where you can, I don't know if you are a user of chess .com. I am, but I'm a very bad chess player. Okay. Okay. So maybe, so that's a great cue. One of the things is we are trying to build effectively a narration which guides you through the game so you can learn how to play better. And there's a version of that where hopefully you will be able to work with some of the iconic chess players where you can like have the delivery from Magnus Carlson or Garrikas Parvarki, Karna Kamura to guide you through the game and get even better while you play it Which would be phenomenal and I think this will be one of the common things we'll see where like everybody will have their personal tutor for the subject that they want with voice that they relate to and they and they can get closer And that's on the enterprise side.

Read the full transcript

30:29But then on the consumer side too, we've seen kind of completely new ways of augmenting the way you can deliver the content. And like the work of the Time magazine where you can read the article, you can listen to the article, but you can also speak to the article. So it worked effectively during the person of the year release where you could ask the questions about how they became person of the year, tell me more about other people of the year, and kind of dive into that a little bit here. And then we ask the company every self and are trying to build an agent that people can interact and see the art of possible.

31:03Most recently we've created an agent for my favorite physicist or one of the two. I am with working with his family, Richard Feynman, where you can actually... It's my favorite too. Okay, great, great. He's amazing. He is such an amazing way to like, both delivered in knowledge in educational, like simple way and humoristic way. And just like the way he speaks is also amazing and the way he writes is amazing. So that was, that was amazing. And I think this will like alter where maybe in the future you will have like, you know, his, his cultic lectures or one of his book where you can, you can listen to it and his voice and then dive into like some of his background and understand that better.

31:50like the truly your jocke miss or Fey Men and like dive into this. I would love to I would love to hear a reading of that book in his voice. Yeah, that'd be amazing. Yeah, 100 % for some of the enterprise applications or maybe the consumer applications as well. It seems like there are a lot of situations where the interface is not the interface might be the enabler, but it's not the bottleneck. The bottleneck is sort of the underlying business logic or the underlying context that's required to actually have the right sort of sort of conversation with your customer or whoever the user is. How often do you run into that?

32:27What's your sense for where those bottlenecks are getting removed and where they might still be a little bit sticky at the moment? The benefit of us working so closely with a lot of companies where we, why would we bring our engineers to work directly with them frequently results and us kind of diving into seeing some of the common bottlenecks. And when we've started, you know, that's your thing about a conversational AI stack, You have the kind of the speech to text element of understanding what you say. You have the LM piece of generating the response and then text to speech to narrate it back.

32:57And then you have the entire tail taking model to deliver that experience in a good way. But really that's just the enabler. But then like you said, to be able to deliver the right response, you need both the knowledge base, the business base or the business information about how you want to actually generate our response and what's relevant in a specific context, and then you need the functions and integrations to trigger the right set of actions. And in our case, we've kind of, we've built that stack around the product. So companies we work with can bring that knowledge base relatively easily, have access to Rugg, if they want to enable this, are able to that on the fly, if they need to, and then of course build that the functions around it.

33:41And the sort of very common, common themes is definitely coming across where the deeper enterprise you go, the more integrations will start becoming more important, whether it's simple things like Twilio or Sip Tronking to make the phone call, or whether it's connecting to the CRM system of choice that they have or working with the past providers or the current providers where those companies are deployed like Genesis. That's definitely a common theme where that's probably taking the most time of like how do you have the entire suite of integrations that works reliably and the business can easily connect to their logic.

34:23In our case, of course, this is increasing in every next company we work with already benefits from a lot of the integrations that were built. So that's probably the most frequent one, the integrations itself. Knowledge -based isn't as big of an issue, but that depends on the company. Like if we work with a company that you know it's it's it's it's um we've seen kind of it all from like how Well organized and knowledge is inside of the company if it's a company that has been spending all of effort On on digitizing already and creating like some version of source of truth where that information lies and how I realized relatively easy to onboard them and then as we go to More complex one and I don't know if I can mention anyone But it can get pretty gnarly.

35:08And then we work with them and like, okay, that's what we need to do as the first step. Some of the protocols that are being developed to like standardize that, like MCP, is definitely helpful, something that we also bring into default. You don't want to spend a time on all the integrations if the services can provide that as an easy step away. Well, and you mentioned in Therapeutic, One of the things that you plug into is the foundation models themselves. And I imagine there's a bit of a co -op petition dynamic where sometimes you're competing with their voice functionality. Sometimes you're you're working with them to provide a solution for a customer.

35:48How do you manage that? Like how does that? I imagine there are a bunch of founders listening who are in similar positions where they work with foundation models, but they kind of compete with foundation models. I'm just scared. How do you manage that? I think the main thing that we realize is most of them are complementary to like work like conversational AI. Yeah. We're trying to stay agnostic from using one provider. But I think the main thing is true and happen over the especially last year. No, no, no, no, no, don't think about it. Is that we are not trying to rely only on one. We are trying to have many of them together in default and that kind of goes to both.

36:23Like one. What if they develop into being a closer competition? where maybe they won't be able to provide the service to us or their service becomes to blurry or we, you know, we of course are not using any of the data back to them, but could that be a concern in the future? So kind of that piece. But also the second piece is the, when you develop like a product, like conversational AI, which allows you to deploy your voice AI agent, all our customers will have a different preference for using the LLM. but frequently or even more frequently, you want this cascading mechanism that what if one LM isn't working at a given time, go through and have the second layer of support or third layer to perform pretty well.

37:10We've seen this work extremely successfully. So to large extent, treat the mass partners, hobby to be partners with many of them, and hopefully that continues. If we are competing, they'll be a good competition to. Let me ask you on the product, what do your customers care the most about? One sort of meme over the last year or so has been people who keep tabbing benchmarks are kind of missing the point. You know, there are a lot of things beyond the benchmarks that customers really care about. What is your customers really care about? The, in a very true on the benchmark side, especially in audio, but if our customers care about free things, quality, both like how expressive it is and both English and other languages.

37:50And that's probably the top one. Like if you don't have quality, everything else doesn't, doesn't, doesn't matter. Of course, the thresholds of quality will depend on the use case. It's a different threshold for narration, for delivery and in a, in a gender space and dubbing. Second one is latency. You won't be able to deliver a conversational agent if the latency is in good enough, but that's where the interesting combination will happen between the quality versus latency benchmark that, um, that you, that you have. And then the third one, which is especially useful on that scale's reliability.

38:25Like, can I deploy at scale like that? I put games example where millions of players are interacting with and the system holds up. It's still performative, still works extremely well. And time and time again, we've seen that they're being able to scale and reliably deliver that infrastructure is critical. Can I ask you how far do you think we are from highly or fully reliable, human or superhuman quality, effectively zero latency voice interaction. And maybe the related question is, how does the nature of the engineering challenges you face change as we get closer and inevitably surpass that sort of threshold?

39:11The ideal, like we would love to prove that it's possible this year, this year, and like, cross the touring test of speaking with an agent, and you just would say, like, this is like speaking another human. I think it's a very ambitious goal, but I think it's possible. Yeah. I think it's possible. If not this year, then hopefully early in 2020, six, but I think we can do it. I think we can do it. You know, you probably have different groups of users too, where some people will kind of be very attuned and you will be much harder to posituring tests for them. But for majority of people, I hope we are able to get to that level this year.

39:55I think the biggest question, and that's kind of where the timeline is a little bit more dependent, is, will it be the model that we have today, which is a cascading model where you have the speech to text, a lamb text to speech. So I can free separate pieces that can be performative. Or they have like the only model where you train them together truly duplex style where that delivery is much better. And that's effectively what we are trying to assess. We are doing both. We are now then the one in production is the Cascading model.

40:34Soon, And I think the main thing that you will see is the kind of the reliability versus the expressivity trade -off. I think latency we can get pretty good on both sides, but similarly there might be some trade -off of latency where the true duplex model will always be quicker. It would be a little bit more expressive but less reliable. And the Cascade models, definitely more reliable, can be extremely expressive, but maybe not as as as as contextually responsive and then latency will be a little bit harder. So that would be a huge engineering challenge. And I think no company has been able to do it well, like fuse the the modality of a lens with audio well.

41:15So I hope it will be the first one, which is the internal internal big goal. But we you know we've seen you've seen probably the opening eye work, the method work that are doubling in there. I don't think it passed the Turing test yet. So hopefully, hopefully it will be the first. Awesome. And then you mentioned earlier that you think of, and you have thought of voice as sort of a new default interaction mode for a lot of technology. Can you paint that picture a little bit more? Let's say we're five or 10 years down the road. How do you imagine just the way people live with technology, the way people interact with technology changes as a result of your model getting so good?

41:56I think the first, like, there would be this beautiful part where kind of technology will go into the background so you can really focus on learning, on human interaction, and then you will have it like accessible through voice versus through the screen. I think the first piece will be the education, I think they will be like an entire change where all of us will have the kind of the guiding voice whether we are learning mathematics and are going through the notes or whether we are trying to learn a new language and interact with a native speaker to guide you through how to pronounce things. And I think this will be the first theme where in the next five to the years, it will be the default that you will have the agents, voice agents, to help you through that learning.

42:43Second thing, which will be interesting how this affects the whole cultural exchange around the world. I think you will be able to go to another country an interactive another person while still carrying your own voice, your own emotion, intonation, and the person can understand you. There will be an interesting question how the technology is delivered. Is it the headphone? Is it neural ink? Is it another technology? But my do will happen. And I think we hopefully can make it happen. If you read Hitchhiker's Guide to Galaxy, there's this concept of bubble fish. I think bubble fish will be there, and the technology will make it possible.

43:23So that'll be a second, a huge, huge theme. And I think generally, like, you know, we've spoke about this personal tutor example, but I think there'll be other set of assistants and agents that all of us have that just can be sent to the performed tasks in our behalf and to perform a lot of those tasks you will need voice, whether it's, you know, booking a, booking a restaurant around or whether it's jumping into a specific meeting to take notes and summarize that in the style that you need. You want to be able to perform the action or whether it's calling a customer support and the customer support agent responding.

44:01So that would be an interesting theme of like agent to agent interaction and how it's authenticated, how do you know it's real or not. But of course, voice will play a big role in all free, like the education. I think, and generally how we learn things will be so dependent on that. The universal translator piece will have voice at the forefront and then the general services around the life will be so crucially voice driven. Very cool. And you mentioned authentication. I was going to ask you about that. So one of the fears that always comes up is impersonation. Can you talk about how you've handled that to date and maybe how it's evolved to date and where you see it headed from here?

44:47Yeah, the way we've started and it does like a big piece for us from the start is for all the content and generative 11 loves. You can trace it back to the specific account that generated it. And so I have a pretty robust mechanism of tying the audio output to the account and it can take action. So that provenance is extremely important. And I think we'll be increasingly important in the future where you want to be able to understand what's the AI content or not AI content, or maybe it will shift even step deeper where you will run an authenticating AI. You'll also authenticate humans. So you'll have one device authentication that, OK, this is Matty calling another person.

45:30The second thing is the wider set of the moderation of, is it a call trying to do fraud and scam, or is this a voice that might be not authenticated, which we do as a company, and that kind of evolved over time to what extent we do it and how we do it. So, moderating on a voice on the text level. And then the third thing, kind of stretching that what we've started ourselves on, like the provenance component, is that how can we train models and work with other companies to not only train it for 11 labs, but also open source technology, which is of course prevalent in that space, other commercial models.

46:07And it's possible, of course, as open source develops, it always will be a cat and mouse game whether you can actually catch it. But we worked a lot with other companies or academia like University of Berkeley to actually deliver those models and be able to detect it. And that kind of the guiding, especially now that the more we take the kind of the leading position and deploying your technology like the conversational AI, soon a new model we try to spend even more time on trying to understand what are the safety mechanism that we can bring in to make it as useful for good actors and minimize the bad actors.

46:47So that's the usual trade off there. Can we talk about Europe for a minute? Let's do it. Okay, so your remote company, but you're based in London, what have been the advantages of being based in Europe? What have been some of the disadvantages of being based in Europe? It's a great question. I think the advantage for us was the talent. Being able to attract some of the best talent. And frequently people say that there's lack of drive in the people in Europe. We haven't felt that at all. We feel like these people are so passionate. I think such an incredible team. We try to run it with small teams.

47:25But everybody is just pushing all the time. I'm so excited about what we can do and some of the most hardworking people I had a pleasure to work with and such a high caliber of people too. So talent was a extremely positive surprise for us of like how the team kind of got constructed and especially now as we continue hiring people whether it's people across broader Europe, central Eastern Europe, like just the caliber is super high. The second thing, which I think is true, where there's this wider feeling where Europe is behind. And likely, in many ways, it's true. Like AI innovation is being led in the US.

48:14And kind of it's in Asia, our closely following, Europe is behind. But the energy for the people is to really change that. And I think it shifted from over last years where it was like a little bit more cautious over when we started the company. Now, like, we feel the keenness and we want to be at the forefront of that. And I think getting that energy from people and that drive was a lot easier. So that's probably an advantage where we can just move quicker. The companies are actually keen to adopt increasingly, which is helping as a company in Europe. really as a global company, but with a lot of people in Europe, it helps us deploy with those companies too.

48:55And maybe there's another flavor and last flavor of that, which is Europe's specific, but also global specific. So when we started company, we didn't really think about any specific region, like we are, you know, a Polish company or British company or US company. But one thing was true where we wanted to be a global solution. Yeah, and not only from deployment perspective, but also from the core of of what we are trying to achieve where it's like how do you bring audio And I make it accessible in all those different languages So it kind of was that through the spine of the company from the star from the core of the company and And that definitely helped us where now when we have a lot of people in all of different regions They speak the language they can work with the client and that I think I likely helped that we were at Europe at a time because we were able to to bring up people and optimize for that local experience.

49:46On the other side, what was definitely harder is in the US there's this incredible community of you have people with the drive but you also have the people that have been through this journey a few times. And you can learn from those people so much easier. And there's just so many of people that created companies, access companies, led a function at the different scale than most of the companies in Europe. So it's kind of almost granted that you can learn from those people just by being around them and being able to ask the questions. That's much, was much harder, I think, especially in the early days to just be able to ask those, like not even ask the questions, but know what questions to ask.

50:25Yeah. And of course we've been lucky to partner with incredible investors for the years to help us through those questions. But that was harder, I think, in Europe. And then the second is probably the flip side of, you know, while I'm positive there is the enthusiasm now in Europe. I think it was lacking over the last years. I think the US was like, excitingly taking the approach of leading over, especially over over last year and creating the ecosystem to like let it flourish. I think Europe is still figuring it out and that's, you know, whether it's the regulatory things that like you AI act, that I think will not contribute to us accelerating which people are trying to figure out.

51:17There's the enthusiasm, but I think it's slowing it out. Yeah, by the first one is definitely the bigger. Yeah, should you do a quick fire around? Let's do it. Okay. What is your favorite AI application that you personally use? And it can't be 11 labs or 11 reader.

51:37It really changes over time. But a perplexity is it was I think an is one of my one of my favorites. Really? What is what is and for you what is proplexity give you that chat GPT or Google doesn't give you? Yeah, the chat GPT is also amazing Chagy PT is also amazing. I think for a long time it was Being able to go deeper and understand the sources That was a hesitated a little bit of the was is where where I think Chagy PT now has a lot more of that of that of that component So I tend to use both in many of those cases this, you know, for a long time, and non AI application, but I think they are trying to build AI application like my favorite app would be Google Maps.

52:21I like, I think it's incredible. I'm, it's such a powerful application. Let me put my screen. What other applications do I have? Well, sorry. What are you doing that? I will, I will go to Google Maps and just browse. But yeah, I'll just go to Google Maps and explore some location that I've never been to before. It's 100%. I mean, it's great as a search function of the area too. It's great. As a niche application, I like FYI, this is a will I am startup. Oh, okay. Or just like a combination of well, it started as a communication app, but now it's more of a radio app. Okay. Like when a curiosity is there, cloud is great too.

53:07I use cloud for a very different things then then then to GPT like any deeper coding elements prototyping I was used Cloud and I love it. Actually, no, I do have a real like more real recent answer which is loveable. Yeah, yeah, I loveable was the use it at all for 11 labs. You just use it personally to Ah, that's true. I think more like you know, my life is 11 labs like it's so the trophies all those applications. Yes, I it's like all of these I use partly for big time for 11 laps too. But yeah, I love a ball I use for for 11 laps. They're like exploring new things to you every so often. I will use I was lovable which ultimately is tied to 3 11 laps but that it's great for prototyping and like pulling out the quick demo for a client.

54:00It's great. Very cool. So, I know, not related, I guess. All right. Who wrote us your favorite one? My favorite one. You know, it's funny. So, yesterday we had a team meeting and everybody checked with ChatGPT to see how many queries they submitted in the last 30 days. And I'd done like 300 in the last 30 days. I was like, oh, yeah, that's pretty good. I'm like, pretty good user. And Andrew, similarly, had done about 300 in the last 30 days. Some of the younger folks on our team was a thousand plus. And so not not only I'm a big DAU of chat you be seeing I thought I was a power user but apparently not compared to what some other people are doing.

54:36I know it's a very generic answer but it's unbelievable how much you can do in one app at this point. Do you just thought as well? I use cloud a little bit but not nearly as much. The other app that I use every single day which I'm very contrarian on is quip which is which is Brett Taylor's company. I know from years ago that got sold the Salesforce and I'm pretty sure that I'm the only DAU at this point But I'm just open Salesforce doesn't shut it down because my whole life is in quit. I use that pond here I like who was good. It's really good. Yeah. No, they nailed the basics. Yeah, like they nailed the basics Didn't get bogged down in bells and whistles to build the basics great experience All right, who is uh who in the world of AI do you admire most?

55:18These are like hard crap another I'm not a rapid fire questions, but I think I really like Demis Kassabis. Tell me more. It's, you know, both his, the, I think he is always straight to the point. He can speak very deeply about the research, but he also has created for the years so many incredible work himself. And I was, of course, leading a lot of the research work. But I kind of liked that combination that he has been doing the research and now leading it. And whether this was the A &R Alpha fold, which I think is truly a new... It's like everybody, I think, agrees here by a truth frontier for the world.

56:05And taking what, while most people focus on part of the AI work, he is trying to bring it to biology. I mean, Daria Amadeus, of course, trying to do that too. So it's going to be incredible, like what this evolves to. But then that he was creating games in the early days. It was an incredible chess player, has been trying to find a way for AI to win across all those games. It's like the versatility of how he, he both can lead the deployment of research can, is probably one of the best researchers himself. Stays extremely humble. And just like honest, intellectually honest, I feel like, you know, you were speaking with with them as he or Sir Dammis you would you would get an honest answer and Yeah, I said see see he's amazing very cool.

56:54All right last one Hot take on the future of AI some belief that you feel Medium too strongly about that you feel as under hyped or maybe contrarian I I feel like it's an answer that you would expect maybe to some extent. But I do think the whole cross -link you all aspect is still totally underhyped. If you will be able to go to any place and speak that language and people can truly speak with yourself, and whether this will be initially the delivery of content and then future delivery of communication, I think this will change the world of how we see it. I think one of the biggest barriers is in those conversations that you cannot really understand the other person.

57:37Of course, it has a textual component to it, like be able to translate it well. But then also the voice delivery. And I feel like this is completely underhyped. It's like, no, nobody is. Do you think the device that enables that exist yet? No, I don't think so. There won't be the phone, there won't be glasses, might be some other form factor. I think it will have many forms. I think people will have, you know, glasses. I think headphones will be one of the first, which will be the easiest. I'm glass of for sure will be there too, but I don't think everybody will wear the glasses. And then you know like, is there some version of a non -invasive neural link that people can have while they travel?

58:19There'll be an interesting like attachment to the body that actually works. Do you think it's underhyped or do you think it's hyped enough? Let's use case. I would probably bundle that into the overall idea sort of ambient computing where you are able to focus on human beings, technology fades into the background. It's passively absorbing what's happening around you, using that context to help make you smarter, help you do things, help translate whatever the case might be. Yeah, I think that absolutely fits into my mental model of where the world is headed. But I do wonder what will the form factor be that enables that?

58:57I think it's pretty what are the enabling technologies that allow for the business logic and that sort of thing to work starting to come into focus What's the form factor is still to be determined? But I absolutely agree with that. Yeah, maybe that's maybe that's the reason it's not hyped enough that you don't Yeah, people can't picture it. Yeah, yeah Awesome, Mati. Thanks so much. Pat, thank you so much for having me. That was great conversation. It's been a pleasure

From the publisher

Mati Staniszewski, co-founder and CEO of ElevenLabs, explains how staying laser-focused on audio innovation has allowed his company to thrive despite the push into multimodality from foundation models. From a high school friendship in Poland to building one of the fastest-growing AI companies, Mati shares how ElevenLabs transformed text-to-speech with contextual understanding and emotional delivery. He discusses the company's viral moments (from Harry Potter by Balenciaga to powering Darth Vader in Fortnite), and explains how ElevenLabs is creating the infrastructure for voice agents and real-time translation that could eliminate language barriers worldwide.

Hosted by: Pat Grady, Sequoia Capital

Mentioned in this episode:

Attention Is All You Need: The original Transformers paper

Tortoise-tts: Open source text to speech model that was a starting point for ElevenLabs (which now maintains a v2)

Harry Potter by Balenciaga: ElevenLabs’ first big viral moment from 2023

The first AI that can laugh: 2022 blog post backing up ElevenLab’s claim of laughter (it got better in v3)

Darth Vader's voice in Fortnite: ElevenLabs used actual voice clips provided by James Earl Jones before he died

Lex Fridman interviews Prime Minister Modi: ElevenLabs enabled Fridman to speak in Hindi and Modi to speak in English.

Time Person of the Year 2024: ElevenLabs-powered experiment with “conversational journalism”

Iconic Voices: Richard Feynman, Deepak Chopra, Maya Angelou and more available in ElevenLabs reader app

SIP trunking: a method of delivering voice, video, and other unified communications over the internet using the Session Initiation Protocol (SIP)

Genesys: Leading enterprise CX platform for agentic AI

Hitchhiker’s Guide to the Galaxy: Comedy/science-fiction series by Douglas Adams that contains the concept of the Babel Fish instantaneous translator, cited by Mati

FYI: communication and productivity app for creatives that Mati uses, founded by will.i.am

Lovable: prototyping app that Mati loves

More from Training Data

All 110 episodes
ElevenLabs’ Mati Staniszewski: Why Voice Will Be the Fundamental Interface for TechTraining Data · 1 h
Listen in VO