ElevenLabs’ Vision for Voice Interfaces | CEO Mati Staniszewski

27 Oct 2025 · 1 h 5 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: ElevenLabs’ Vision for Voice Interfaces | CEO Mati Staniszewski

Episode Overview

  • Host: Joubin Mirzadegan, Go-to-Market Operating Partner at Kleiner Perkins
  • Guest: Mati Staniszewski, Co-founder and CEO of ElevenLabs
  • Air Date: [Insert Air Date]
  • Episode Description: Discussion on how ElevenLabs is innovating in voice AI, focusing on emotional representation in synthetic voices, applications in various fields including healthcare and creative industries.

Key Themes and Discussions

Introduction to ElevenLabs

  • Background: ElevenLabs was founded in early 2022, aiming to transform how we use voice in technology.
  • Initial Inspiration: The need for diverse and emotionally resonant voices in Polish dubbing, where one voice narrates all characters.

Vision for Voice AI

  • Emotional Representation: ElevenLabs focuses on creating voices that can express emotions, enhancing user interaction with technology.
  • Applications:
  • Healthcare: Personalized healthcare communication tailored for different age groups.
  • Creative Industries: Voiceovers, dubbing, and narration in filmmaking and media.

Company Structure and Early Growth

  • Team Composition: Early hires included a mix of researchers, engineers, and marketing staff.
  • Remote Hiring Strategy: Aimed to recruit top talent from around the world, leading to a diverse and skilled team.

Technology and Innovation

  • Research and Development Focus: Combining deep technical knowledge with practical applications in voice AI.
  • Voice AI Capabilities:
  • High-quality synthetic speech that represents human emotion.
  • Ability to recreate a variety of human voices, accents, and dialects.

Challenges in Voice AI

  • Uncanny Valley: Discussion around the challenge of making AI-generated voices indistinguishable from human voices.
  • Current State of AI Voices: Despite advancements, there are still instances where AI voices are easily recognizable as non-human.

Market Dynamics and Competition

  • Position in the Market: ElevenLabs is positioned amidst significant competition from large players like OpenAI and Google.
  • Response to Market Changes: ElevenLabs aims to innovate continuously and maintain a competitive edge in foundational research and application development.

Future Outlook

  • Long-term Vision: The goal is to create a generational company that outlives its founders.
  • Anticipated Growth: Expected to double its workforce to support expanding projects and market needs.
  • Raising Capital: Recent funding rounds focused on secondary offerings to ensure long-term alignment among team members.

Personal Insights and Reflections

  • Mati’s Philosophy: The importance of creating impactful work that contributes positively to society and human experience.
  • Work-Life Balance: Despite the high demands of leading a tech startup, Mati expresses an optimistic view about the future of voice technology.

Key Takeaways

  • Innovation in Voice Technology: ElevenLabs is at the forefront of voice AI, merging emotional depth with technological capability.
  • Diverse Applications: The potential applications of voice AI are vast, ranging from healthcare to entertainment.
  • Competitive Landscape: The voice AI sector is evolving rapidly, with growing competition and the need for continuous innovation.
  • Long-term Focus: Building a company that offers lasting societal benefits and transforms the way people interact with technology.

Conclusion Mati Staniszewski’s vision for ElevenLabs reflects a profound understanding of the potential of voice AI to shape various industries. As the company navigates the complexities of rapid technological advancements and market competition, its commitment to innovation and emotional representation in voice technology stands out as a key differentiator.

---

Connect with the Guests

  • Mati Staniszewski: [X](https://x.com/matistanis?lang=en) | [LinkedIn](https://www.linkedin.com/in/matiii/?originalSubdomain=uk)
  • Joubin Mirzadegan: [X](https://x.com/Joubinmir) | [LinkedIn](https://www.linkedin.com/in/joubin-mirzadegan-66186854/) | Email: [grit@kleinerperkins.com](mailto:grit@kleinerperkins.com)

Additional Resources

  • [Kleiner Perkins](https://www.kleinerperkins.com/)
  • [ElevenLabs](https://www.elevenlabs.io/)

Feel free to reach out to the respective social media handles for more insights or to connect with the leaders in voice AI technology.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If you watch a movie in Polish language, all the characters get narrated with one single voice. That was the initial trigger point of like, hey, we know this will change. Knowing what's possible across audio space, decided to start 11 labs. And of course, the vision quickly expanded from there where it's not only that that will change. It will be the entire interaction of technology around us that will change. AI is not going to replace the entirety of the process on creating a movie, but it will replace parts of it where you will have started with an idea and then it will help you draft the first set of scripts.

0:28And then you'll refine that and then you will repeat this process a few times. And then you have an end result, probably a combination of human and AI, where this is indistinguishable. You are going through a set of iterations on selecting the voices, redoing the lines, sound effects. Maybe you created the voices that you care about and you want to bring those voices inside of the production flow. You need a pretty solid platform to be able to do that. And we kind of combine all of that in one place for you.

1:02Welcome to Grit. I'm Jubin, partner at Kleiner Perkins, a show where we go beyond the highlight reel and explore the personal and professional challenges of building history-making companies. Today on the show, we have Mati, co-founder and CEO of Eleven Labs, the audio company building infrastructure for how we speak and listen in the world of AI. They are in the eye of the hurricane of all AI hurricanes right now, and it's an absolutely incredible story of starting in Europe and building what looks to be one of the preeminent businesses in LLMs today. Enjoy the episode. So your early employees, how much of it was like, I don't know, researchers slash scientists versus these roles start to get very blurry and fuzzy.

1:41But researchers in our mind, these are like people that can create new work, new paper, get it out to the world. And this is like a true innovation that it could make any paper at any of the conferences and would be one of the best paper. So that's the researcher. Then you're right, there's standard full stack engineers. But then we also have like a category of research engineers. So people that can take an existing paper and improve it or do some of the modifications on the inference side to make the model serving a more optimal and quicker or better. So kind of in those three categories, initially we would optimize for researchers and research engineers.

2:19So that's why we started very much remote at the very beginning where we wanted to hire anyone who was the best in the world, wherever they were. And we knew that the only way for us to get to the best talent would be to hire the researchers from, we had a person from Korea, from Europe. And that was like kind of the big part of like, let's hire them wherever they are. Let's assemble the best researchers in the world in one place. And then, of course, full stack engineering came kind of in parallel at the same time to start building the product. Okay. And so your first, let's say your first 10.

2:53Yeah. How many were each of those responsibilities? So my co-founder, the smartest person I know, also an incredible researcher. So that's what he used to do at Google and previously at university. So he was one of the researchers. But on top of Piotr, we hired two people in the first 10, one, two people, then additional one research engineer. But the research engineer, like in the early days, you kind of do everything, whether it's data, whether it's improving the model. So it kind of infuses. But I would say from the first 10, including me and my co-founder, four of the people would be doing research or research engineering.

3:31Two of three people would do engineering and three people would go to market growth. So pretty quickly, we started looking into how we can get the technology out there to both the self-serve distribution and then sales distribution through working with some of the companies. Yeah. And before we get too far ahead, maybe for those that don't know, what do you all do? So we are 11labs. We do a combination of research and product work, product deployment work focused on AI voice space. We are on a mission to be the voice of technology, and we do it for two QAs. One is an agents platform offering. We help people build conversational agents across customer support for personal AI, across media entertainment to make interactions interactive.

4:15And then second is our creative platform offering where we help people create narration, voiceover, dubbing, add music to their work in one holistic suite. So really stretching the use cases where voice can empower and now even more on the omnichannel experience and agent and omnimodality experience on the creative side. What's crazy is the company started pre-GPG3, right? And you were, if my math is correct, like 25 years old at that point, living in Europe and recruiting research researchers. like uh that doesn't i'm trying to like that doesn't really add up to me like how does that how does that like what are they even what were they working on like what research were they even doing yeah how did that how did that work no it is you know it's like in a way it was um also the best in many ways it was the best we started a company in 2022 it was it was kind of still the year where crypto metaverse were kind of on their right final final downhill moment and so AI like nobody was working on AI which gave us an advantage of two things one of us we we could we had a good amount of time to start work before the GPT moment where like a true viral moment where that gave us like kind of advantage of working before before others and to the people we were able to hire in those days were true missionaries and excited people about the space you mean you're not competing with zock back then yeah you are not competing with other big players although happy to compete with them we you know we we we i think we have a great spot for a lot of the researchers to do the best work of their lives in uh in in audio and beyond um but the the main thing is that all the people that join you at the time are like a true missionaries they really believe in a mission, really believe in what the voice can do for the world.

6:12That was still like, currently, it became normalized that voice is going to be the future interface. It's an important part. Back then, nobody really cared about AI. Nobody really cared about voice. And me and my co-founder, so we started a company. We knew each other for 15 years. So we are best friends from high school, took all the same classes. We both loved mathematics and everything around it. And then for the years, did everything together. Worked together, traveled together, live together and are still best friends. So the time is on our side. And the idea came from Poland. We are both from Poland and there's a very peculiar thing.

6:47If you watch a movie in Polish language, like a foreign movie in Polish language, all the characters get narrated with one single voice. So whether it's a male speaker or a female speaker in the original movie, in Polish, you get just one voice narrating. Is that unique to Poland? It's unique to the post-communism country. There's like five countries in the world that do it this way. And Poland is like on one of the few, which is, it's like almost hard to imagine, but it's a terrible experience. And it still happens for most of the work today, most of the content today. So that was the initial trigger point of like, hey, we know this will change.

7:23We know that with the invention of Transformer a few years back, with the fusion models, audio will be able to get created in that way that's actually heavily emotional and has a good intonation. You can recreate any voice you want. And with knowing what's possible across audio space, decided to start 11 Labs. And of course, the vision quickly expanded from there where it's not only that that will change. It will be the entire interaction with technology around us that will change. If you can create synthetic speech that's high quality, the voice will be such a powerful medium for interaction.

7:59And started 11 Labs in 2022. too. We also, through the years prior, so here at Google, I was at Palantir, we would meet and do Hack Weekend projects together across a variety of different domains. I'm curious how you stumbled on your idea as well. But in our case, we were exploring new technologies and seeing what they can do and how we can implement some of the mostly ideas around them, less so commercially focused ideas and one of those was in audio. But that gave us a good kind of glimpse into how we can work together. We're so excited about going full-time and dropping our relevant companies and starting LevenLabs.

8:43Crazy. What an insane story. And then when, I guess, what was it, GPT-3 came out? Was a year later? So the ChudGPT came December 2022, January 2023. Not long into the journey. Not long into the journey. It's like... Did you basically scrap everything? No, it was... So in a way, a lot of our ideas were based on the foundational stack anyways. So they were like, just validated how relevant they are. So like, it was almost a parallel invention path, just a different technology. We of course have a lot of the transformative fusion ideas are coming into our work, into our work too. So when we started in 2022, we would use some of the most relevant work recreating what should be done in text-to-speech.

9:37And then when ChatGPT moment came and a lot of the technology around us, this was A, validating that we are doing the right thing because the ideas were giving good results. And two, it kind of gave you this amazing overlap of what's possible, where you could have chat GPT now and you could have voice on top of it. Suddenly, you can do content in such a much easier way where 11 labs has even more advantage. I guess, you know, like at the time, so in 2022, most of the synthetic speech just was so poor. It was still very robotic. It's still, I don't want to name some of the big other companies, but you can imagine the hardware devices around.

10:22It's still not good. And back in 2020, it was even worse. So it was truly robotic speech. And nobody could figure out how to make the speech sound human, actually represent the emotions. And that's where the 2022 for us was all focused. How can we get the text, understand the context in a much better way, understand the emotions around the text? Is it a happy text? Is it a dialogue? And bring that in the same way that the voice actor would or human would. When you read, you understand that context and then you bring those emotions across. So that was the big innovation we brought in 2022 where text-to-speech sounded human.

11:00And the second big was being able to recreate voices. So how can you effectively encode a voice and the characteristics of the voice? And then when you produce synthetic speech, it still carries the style of the voice, the accent, the dialect, whether it's narrative voice or conversational voice. How do you bring that all across? And we figured this one out. But a lot of those ideas were based on similar ideas that LLMs would carry as well. And then, of course, a lot of the nuance that were required for the audio space. One of the things that I've been thinking about that I wanted your opinion on, and maybe this is just like a very stupid thought, but like we have this concept here in Silicon Valley about like crossing what's called the uncanny valley.

11:45And the uncanny valley is this concept where basically machines become indistinguishable from humans. You cannot tell if it's an LLM talking to you or writing to you versus a human. And so far, certainly in text, I can, within 10 seconds, tell if an LLM has written it or not. It's very, very obvious. Based on the M dash or? Based on the M dash, based on just the way that it writes. It's not even just the M dash, based on. No, I agree. Like not even random, like forget about the bolding and the M dashes, right? Because nobody actually writes like that. Just the way that it actually writes does not sound human to me.

12:22and it's very obvious the minute that somebody sends me something immediately i'm like you didn't write it just like something inside my brain is like that's not a person that wrote this okay and so you know maybe audio if it's coming off of text is a similar paradigm today and everybody continues to say oh no no it's happening like it's gonna happen for sure and one of the questions that i've been asking myself like you know these models have gotten a lot better and they're still getting better. And everyone's like, well, as long as it keeps getting better, it's definitely going to not sound like it's going to sound indistinguishable from us.

12:58And, you know, I've listened to the like podcast LM from Google. Like I've listened to all of this stuff. I've done the use 11 labs to clone my voice. And I'm not even saying it's not a knock. Like I'm not even saying it should or shouldn't sound like me. I like it. Maybe it shouldn't. Like maybe if you're from India, your dialect should be Indian. Like maybe if you're from Iran, your dialect should be something else. Maybe if you're from machines, your dialect should be machine. You know, like, I don't, I don't know. Like, and, and, and on the, on the lamb quick back question back, is it across all use cases you can tell, or you can like, you know, like if it writes a copy for your email or writes a copy for your post, or is it what, what, across all settings or just specific settings?

13:41I think unless they, somebody has to really, I think the only way that I generally can't tell is if somebody writes down what they first thought, then puts it into the LLM to tweak what they thought and then goes back and then like unbundles the things that sound like an LLM. Right. But I think that's across every use case. I have seen, I have seen that. Yeah. Do you agree, by the way? No, I don't think. I think, yes, in many settings, it's obvious when somebody uses an LLM and some of the things you mentioned make it obvious. But of course, there's something about a style. If somebody goes the lazy effort of, hey, write this for me, immediate tell of it's easy.

14:25But I see more common the last thing of what you mentioned. So somebody will write a thought, get the LLM, rewrite the thought, so they are more human. and I don't know if we are classifying that in the same category, but in those cases, I think it's very not, it can be very good. I agree with you. I wouldn't classify it in that same category. I'm just saying if you tell an LLM to write something for you, just like if you tell an LLM to speak something for you, is it always going to sound like an LLM? Yeah. So I think if you just use like a default, likely yes. I do think now you can kind of like pre-prompted where it becomes a little bit less of a tell.

15:03And in some settings, it's as good or better than the human would. And let's take a customer support setting where we work a lot. If you're working in customer support, the baseline you're comparing it to is usually a IVR flow or like a decision tree of if this, then that. Or a potentially human that has maybe not had as much of the training and will try to help you, but might not be able to help you. In that setting, I find our labs can provide as good, if not bad, just on the tech level, if better than human experience where you get the support and you're comparing it to something that's not there.

15:47I totally agree with you. If you're using Sierra and you're using 11 labs on the backend and Sierra is partnered with United and you're comparing the experience that United would give you pre and post, it's night and day. Yeah, but I'm thinking even on the human side, if you are comparing, you know, just, so even let's ignore voice and I have my take on voice as well, where just on the text level, if you're comparing customer support done by AI agents versus humans, I think many people wouldn't be able to tell and it's probably a better result altogether. If we are thinking about like more of a creative case and like a narration for a movie your conversational dialogue.

16:29I think there, yes, it's a little bit harder to get on the high-end part to get something comparable. On the voice side, on the Uncanny Valley piece, I think in the narration space, I think you can get the results from voices that people won't be able to tell are not narrated by human. And we've seen this across so many places where an audiobook created is just so good that that people are both immersed. And I don't think if you ask them, point back, is it better or worse? They would be able to tell. In the conversational setting, I don't think we passed uncanny value for all types of use cases.

17:06I think we passed it in call center application. I think we passed it in some of the customer support experience use cases. But let's say in immersive media our gaming or personal AI type spaces, I don't think it passed that kind of vary or like a Turing test in all settings. Like we are speaking now, we can interrupt each other at any point, we can move pretty quickly, pause, we can, there's like a reasoning and voice that's combined to get us one. That moment has not yet, I think, happened for conversational setup. But you're convinced it will. I think it will. Yeah, I think it will. Hopefully we make it happen.

17:44So yes, I'm biased as well, I guess. Yeah, you're biased. That's fine though. But even with voice, so you think in that conversational setting, you would be able to tell today, I presume, if you are messaging with an AI straight away that this AI or like after a few exchanges. Yeah, so I - And you don't think it will happen? So I have two points that I'm making. One is certainly on, in most cases, I can tell. Okay. Second, I don't even think it's a bad thing that I can tell. I guess what I'm really questioning is why do we want them to sound and feel like a human. Meaning if I am chatting with United, okay?

18:21Or if I'm on a phone call with United, I don't actually care if it's a human or an LLM on the other side. In fact, I would rather know. And what I care about is how quickly can you get to resolution? That's all I care about. And in fact, it would maybe give me a little bit of peace of mind if I know what's on the other side of the person that I'm talking to. Oh, no, I 100 % agree on that part. Like ideally you do know it's preempted. And it's a good experience. Right. So I think like, yeah, I would decouple those two points too. It's like you want the resolution as quickly as possible. And then two, does being human voice or human experience help you to get to that resolution?

18:59And in some cases it does, in some cases it doesn't. It's like in customer support, no. You want the content, you want the actions, you want the intelligence to be well-connected. You want to know the information about you. You want to know you're speaking of AI agent to help you resolve that case. And then let's say you are in a gaming context. you are in an immersive game in Fortnite. That's a kind of example of a game we worked with, with Epic Games to bring Darth Vader experience to the game. You want to make it feel as emotional as possible, as real as possible. Just like Darth Vader. Just like Darth Vader, exactly.

19:31And, or like, you know, by emotional here, you want a pretty raspy, a pretty serious tone of the experience. And customer support, you don't. But then there is a very different experience depending on who you are. So like, you know, let's say you or I want you want to ask quick to the resolution, quick responses across the conversation. We recently worked with a customer support company in healthcare in Japan. And they had this case where a lot of their population calling in is older. And then they have a small proportion of people that are younger. It was effectively a widget on top of the website.

20:09and they had the information about who is reaching out for help. And both the voice part, but also just the style of how the information gets delivered was drastically different between those groups and led to much different results between those groups. So if you had a younger person reaching out for help, you wanted it relatively quick, you wanted it relatively short in style, straight to the information. If you had an older person, you wanted to make it a slightly more calmer, less emotions, a lot longer responses along slower responses. So in this case, it really mattered depending on the context of who was calling in.

20:50And in the same way as you think about a human, if you have a human responding on customer support in that same client, what you really wanted to do is to be able to alter depending on who you are speaking with in what type of experience you deliver and then what leads to the resolution and that's where um that's where i think it matters that you know you make agents versatile to help across all those different use cases i don't think we should force them to be emotional or human in any of those cases it's more of like how does it be able to be it's able to adapt to the to the use case at hand and kind of bring it to the resolution in a way that's understandable.

21:26Yeah, it makes total sense. Maybe the other point that I'm making on that, because I think we're in agreement on the, should it sound like, should I know who I'm talking to on the other side? But I also think that there is a thing happening now where everybody in my world just assumes that, including you, that we are definitely going to pass the uncanny valley. It is going to happen. There will be a world where we cannot, if we choose not to, it will be indistinguishable whether we are talking to a human or a machine, both in text and audio format, and potentially even in video format. If you're going through YouTube shorts or whatever, and you can't tell the difference if this is a real or not.

22:07And all I'm saying is that hasn't happened yet. Like the YouTube short videos have infiltrated when I scroll through, I can always tell. The audio, I can pretty much always tell. The text, I can almost always tell. For purely LLM created and generated, right? The models are getting way better, right? They've already gotten way better. We haven't crossed it yet. Everybody is telling me we're obviously going to do it, assuming that scaling laws continue to persist. And all I'm asking is like, is there a possible world where that's not true? That's like kind of my only like, is that like, is that like, I feel almost dumb asking the question because it's like, it's like sacrilegious around here to question the scaling laws, but I'm just curious.

22:52Yeah, so two answers there. I think one is, it's like, depending on what you are solving for, answering this question is important or not. In a similar way, as we think about delivering the best experience, in many cases, whether you pass on Cannibale and all those - Doesn't matter. Doesn't matter. And you should go out there. Agreed, because the bar is talking to some United agent that's already terrible. Exactly. You wait in queue for 10 minutes. So I'm not disputing the value of it. But so I guess on the kind of like, as I think about a lot of Silicon Valley companies, I don't think the goal in itself should matter.

23:29And I think there's a second piece to it, which I think is important, which is in the same vein, I think it's like a lot of obsession about the agent or the AI, kind of the buzzwords kind of proliferated the mission of the company itself in many of those places, where at the end of the day, it doesn't really matter across all of that. It really matters about are you solving the problem for the customer in a quick and good way. So that's kind of the part one where I think I agree with the hidden general trope that many seems to be focusing on solving without understanding what you're solving for.

24:03But to the second thing, I think it's going to happen. I do think that uncanny value is in many cases already crossed. I think you might be also more attuned to it given you are living and briefing that every day than most of the general population. I would say it's almost more like, how do we make sure the whole world is prepared where you aren't able to distinguish and you want to get information? I think this will be super important that you know that there is an AI agent on the other side, that content might be created with AI. You want to know that information or at least know that this is generated by human and kind of assume that everything else is AI.

24:38So I'm like 99.9 % that in most of the cases, the uncanny value will be crossed. And we need to kind of work on preparing on how you make sure the society functions as a whole. And by the way, I believe you. I mean, I do think you're talking in your own book, but I also believe you. And that's why I'm asking you. You've been engrossed in this. You would know better than anybody. I honestly think even if it doesn't cross for kind of mission of our company, we still will be - I completely agree. That's why I think it's like not us kind of tied in that side. But so you think there's more than 50 % chance we'll not cross on Cani Valley.

25:15So I do want to make a point about 11 labs, because I do believe that you do not need to cross the uncanny valley in order to be a huge company. No, but do you think that uncanny valley will not be crossed in a lamp setting and voice setting? It's more like, OK, if it is crossed, then every movie moving forward will always be generated by an LLM. Why in the world would you create a movie that it's just so much more work to do all the things if you can just tell it what you want it to do? But it's like, you know, like I don't think it will happen the way that, and so maybe we are like answering a different thing.

25:51I don't think it will be end to end in any short perspective of time. I think it's going to be more like middle to middle where AI is not going to replace the entirety of the process in creating a movie, but it will replace parts of it. That's why the middle to middle where you will have started with an idea and then it will help you draft the first set of scripts. and it will help you vocalize the first set of scripts, find the voices that are right for that script. And then you'll refine that and then you will repeat this process a few times and then you have the end result. That's in the end, probably a combination of human and AI.

26:26And I think this will end in a result in the category two of what you started with, where you kind of have AI or like human AI, say human, AI, human, where this is indistinguishable. I think in that case, this will happen. I think this is where the 99.9%, I think it's in that category. Will you be able to generate a movie from scratch at level of quality and creativity, ingenuity as the human would? In many cases, I don't think you will, but there will be definitely use cases where, yes, that still will be possible. yeah yeah i think it's going to be like a you know like a um like same same like in the current i think let's take movie as an example like even currently you you have rare this kind of sparks of incredible creativity where you're watching something that's completely new or different um and i think this will still remain like you will need a lot of idea of process generation from the human side but then you will be able to generate a longer tail of content that's okay it's still enjoyable to watch, but it might not have required that kind of entirely new idea.

Read the full transcript

27:39The interesting question that I keep circling back to is like, if you've crossed the Uncanny Valley for a three minute clip, why can't you cross it for two hours? And if you do cross it for two hours where I cannot tell, then why wouldn't a director just give their entire script that they wrote to an LLM. And then it can be like, put it in the style of Goodfellas with the punchiness of whatever, right? And it just does it. And then maybe there's some tuning or whatever. I'm just saying like, and if that's true, if that is possible, that would be insane. No, but, and I agree. And I think this will be possible, but coming up with Goodfellas or writing the script for Goodfellas or like that book itself, that's going to be extremely hard.

28:31Agreed. That's going to be extremely hard. And then there's another piece, which is like, it will be easier to create software, to create content, to create movies. So now you're kind of even more attached to the storyline, to the history of something than ever before. And that history will frequently go back to, has it happened in real life? Has it happened? Is it the story I used to follow when I was a kid? Is it a universe of characters I know and love and follow and other people do as well so we can kind of speak with it? So I think it's going to be possible to do it, but it'll be harder to create the attachment to the human across all of those use cases.

29:12Totally. What percentage of the revenue-ish is B2C for you versus B2B enterprise? It's closely approaching 50-50. Oh, interesting. Yeah. So it's a lot on the B2C as the creative use cases, narration, voiceovers, dubs, music. And then on the B2B, it's huge conversational AI push with our agents platform. So everybody's building for conversational use cases, making in voice, but increasingly also something not only voice, where like an omni-channel experience where you do voice and text interactions. So that's kind of skyrocketing across. Yeah, it makes sense. I mean, on the B2B side, like one of the great use cases that we have seen is support.

29:59Yeah. It's great because it's generally a cost center. And it's great because LLMs are very good at taking that unstructured nature and putting it into a structured format that it can automate away. And in so many cases, like it makes total sense to me that you're powering that workflow. Yeah, and 100%, it's kind of the place where we see just such a quick adoption. But an interesting thing that started happening is this shift from this being a customer support place where you reach out as a user when you have an issue into increasingly being more of a kind of a front-facing experience, customer experience across the entirety of your product or user journey.

30:44So from, you know, like in a traditional way, if you've got an issue, you would reach out through an email. Now we see companies working on bringing that as part of their website experience itself. So recently we worked with a company in Italy, the biggest real estate company, where they effectively placed an agent that as a buyer and seller of a company, you can reach out to to get information about that real estate. And they had like a historically a huge supply and demand mismatch. So the seller of a property wasn't available to pick up the phone to give the information about the property. And on the other side, you had a lot of the buyers that were just not the right buyers for the property.

31:27So now you have an AI agent that you can call in and reach out and get information about the property. And then that information gets aggregated, given to the seller of the property, and then they can reach out to the right ones. helped improve the qualification rates from like 20 to 60 % of people that are progressing to the next step. And before, you would only reach out if you had a problem going through that experience. So it kind of is interesting where increasingly it's like kind of the part of the experience itself. Insane. So insane. Even in our own work, we love trying to do it. Like how can we elevate our experience?

32:01We of course had, because of the B2C and B2B angle, we had so many people reaching out for help, credits how to use the product so we shifted it where now you have our agents both on the website itself so you can speak with an agent about 11 labs the moment you open the website of 11 labs then when you log into the product you can you have an agent that can help you navigate through the product experience will open items for you speak with you about how it works and and then of course on top of that you can reach out to support which is also powered by our agents So now you have like a full fleet of from understanding to using to reaching out for help.

32:39And to even if you reach out for like sales part, we work with our own agents to help with that process. And it's been amazing. It's been both amazing for the experience itself elevating and of course solving the issue. But second, and you kind of spoke about this briefly earlier, we noticed a behavior change where if you know you can speak with an agent, your behavior changes. And let's say you're talking about your use case, you are much more keen to tell you everything about the use case because you know you are not judged. You can be a little bit more expansive. You have any time in the world to tell you about the use case.

33:21And this has, in turn, meant that we can provide you better experience. even in the classic kind of BDR, SDR categories where you can reach out to 11 labs, all usual flow, you leave your details, you can request a use case. But then if you want, you can accelerate and give more information by talking to AI BDR and then we loop you into the right person a little bit quicker. And that has seen a great adoption as well. How big is the company now? So we have 330 people now. All everywhere? everywhere we have the biggest basis so we we have um still run it remote so everybody can can work remote but we started investing into hubs where you can work in person and our biggest ones are in london and new york san francisco warsaw and now we are starting some in tokyo bengaluru in sao paulo in brazil as well and you just raised it what was it like six six six 6.6 billion or something?

34:236.6, yeah. And you let people take secondaries. Yes. Actually, there was only a tender. Yeah, only secondary. Only secondary. Only secondary. 100 million? 100 million. 100 million of only secondaries. Yeah. Why? Was it a hard decision? Not hard decision. The goal is simple. We are here for the long term. We want to build something that will stand the test of time, but also will require the next five, 10 years of us continuously pushing. And we want people to be aligned on that long term. And by making it true that A, equity is valuable. Two, you can have early liquidity. So you take some risk.

35:05You get some of the reward from that risk, but you are incentivized to work with us on an even bigger outcome. And I think it really helps. And I'm sure many companies are saying that too, But I do think we have so many extremely passionate missionary-like people that are in here for that long term. And for them, the change in the capital will not change their motivations in terms of work, which is, I think, the most common piece that people mention. And we haven't. We've seen people are even more aligned. How do we get us now as a proxy of the impact we want to create? How do we get 11 laps to 11 squared or$121 billion company?

35:51Where I think it's possible. And I think we can. And now people are here for the long term. Where do you draw the line? I think it's tricky, right? I'm sure at some point it will be tricky. Yeah, the current one, in the same way like in sports, if you hire the best high-performance athletes, for them, if they are doing well, they want to be part of the winning team. That's kind of one. That's what we need to do. But two, the money equation doesn't change. You still want to perform until of your game for the years to come. And I think the same applies to if you set the culture in the right way, and then you have the people that are performing next to you in a great way, then it shouldn't change that motivation either.

36:44I think the harder one, as you think about that, it's harder if it's not going well. And of course, for a short period of time, if you've said it well, you can withstand that. But if it's a longer period of time, then I think it will get challenging because then a combination of you taking, although maybe not because you're still having, you know, a big incentive to make it great over long term. Yeah. Well, and you're not taking up rounds. You're not doing up rounds at that point if it's not going well. Exactly. So you kind of are. Yeah. So even in that case, you shouldn't. I'm sure the motivation will be a little bit harder, but I really believe in our team.

37:23I think they will be able to go through the hard and, you know, the hard times are never as, so the bad times are probably never as bad as they are or as they seem in the moment and the good times aren't as good as they seem. So you kind of are a bit closer to the median or the average of the times. And I think it's true here as well. So I, so on the bad times are never as bad as they seem. I agree. However, I know that I've internalized that. I deeply believe that everything that I've ever been anxious about basically in my entire life either hasn't come true or was not a big deal. And knowing everything that I know, still, I run into, let's just say with the company that we're working on now, I run into something and I'm like, oh my God, like I do not, like, it just feels terrible.

38:10Like I can't in that moment face a bad time and be like, eh, this isn't going to feel bad. Like I know. But I think it's good. I think it's good. It's like you, you, you care. That's why I think it's like the same way for me too. It's like, and for many people around, it's like, you still care. So every single kind of direction takes additional emotional impact. I think, you know, you can still stay level-headed and make the decisions. I think that's the important, important part, but, but it seems, you know, that, yeah. Yeah. My, my, my observation is that there are so many decisions to be made, right?

38:48So many decisions to be made, that like, each one of them on their own, you don't even notice that you're making these decisions. But over time, they just kind of like, it's just like, like, it's like a quarter percentage of energy on every single decision, you know, and then over time, you just like sum all of those up. And you're like draining or depleting some reservoir of decision making energy that you have in order to like, kind of like, when the one comes that's like oh you know you're like well i just made like 50 decisions yesterday you know it's like yeah it's it's we i've went for a few of these conversations recently with a few of people that were i sees for a long time and now are leading parts of the work and they just feel that from being able to what felt like meaningful work they're replying to constant things or tiny things across.

39:44And so first of all, for sure, I see like, you know, the kind of the more you get, the more you will get asked for decisions across. But then I think the two things are important to help out. One is you have, of course, the different segments of the company that applies in different ways. I think in our case, we have an amazing set of leads across that will effectively take a lot of those micro decisions and kind of trust them completely. They will do extremely well. That in turn means that the bad things get filtered out always back to you. And then the second thing is that I think there's this kind of balance of like which decisions you should really take versus the individual should take.

40:32and it was like plus minus 10%, they should just run with it. In our case, what we are trying to like instill for everybody that we are small teams, this will only work if people take ownership and try to make a lot of those decisions themselves. We are assuming and we are trying to put them in a position where we fully trust them and then they can flourish. They won't be punished by doing that and running ahead, but they feel empowered to like run towards the fire, take decisions, fix it, not try to ask. they're like you know fuel these up on like is it right and kind of go through that iteration cycle relatively quickly very of course now as a company we run the company on micro teams so we have 330 people on the product side it's like 20 teams of 5 to 10 people working on different segments of the product work and they are executing and then you kind of need to live with the part that you need to take more of decisions.

41:37But then in turn, it works well. Are you holding your breath every big model release? Like what I find fascinating about what you're doing right now is that your growth is fueled by the fact that you are in the eye of the hurricane, right? Like you are in a middle of a market that really matters. And it turns out when you're in the middle of a market that really matters to people and consumers, especially consumers, that also really matters to OpenAI and Anthropic, right? And Google. And, you know, like we saw this in our world, like, for example, like coding really mattered to Anthropic and OpenAI.

42:25Now, you know, Cursor has so far made it out the other end of that kind of, and maybe you're the cursor of this world and maybe cursor will continue to be ahead of them but these are not the uh incumbents of old right right like these are like it's just these are big giants right that are very talented that if you're in one of the key areas it is it just i'm not even saying oh they're gonna do it or not. It's just got to be a crazy time and somewhat nerve wracking and like knowing that you're right in the sweet spot of something that they really care about. Because consumers really care about it.

43:08You probably get the annoying VC question of like, why won't OpenAI do it? That's not really my question. It's more like when you're in this eye of the hurricane, it's insane. Look, we will fight and I think we can win. It is eye of a hurricane. And so far, we've been able to win time and time again. But I think we are in a unique position where we do both. We do foundational research model and the application around it. So in turn, this means two things. One, to your kind of main point, on the research side, we need to stay competitive. So that are better than OpenAI, Gemini models. And I think we can.

43:51We did it on text-to-speech. Then we repeated this on speech-to-text. Then we did it on music, orchestration, bring that all together. And we will continue battling them out and hopefully winning. It's a tricky one. But that's one. So definitely checking the mark and holding the breath when things happen. But at the end of the day, models only one piece of it. The second piece is as important, which is that application side how do you create the best platform for creating an agent beyond having the best voice or the best lm you need how does it integrate with your crm system with sales for service now goal drive how do you build functions in it how do you decide when to refund or take an appointment scheduling how do you then deploy that in production how do you test evaluate version control and then kind of we spoke about this is the platform functionality but what most customers really care about is the outcomes themselves.

44:47How this will solve, how will that impact my CSAT score and customer support if I do deploy that agent, help me understand that it was the cost benefit ratio of doing that work. And none of those other companies usually do that on the agent side. And that's where we focus and spend a lot of the energy. So now we are in this interesting spot where benefits of a lot of the foundational work ultimately end up benefiting the agents. And we hope to do both. We hope the product is going to be best and the incoming audio foundational layer that we do will be best too. Same applies very much on the creative side.

45:23If you are creating a voiceover or a movie, you are going through a set of iterations on selecting the voices, redoing the lines, adding combination of voice, sound effects, musing the track, maybe reaching to some of the assets you created from the past Maybe you created the voices that you care about and you want to bring those voices inside of the production flow. And then you want your other teammates to check and check whether that's good. And suddenly that process became a pretty complex flow to go through. And you need a pretty solid platform to be able to do that. And we kind of combine all of that in one place for you.

45:58Yeah, but now you're fighting all these flanks. We are. And that's good. You have all these flanks. No, it's great. I just, I can't, do you sleep well at night? These are a lot of flanks. You know, it's one way of looking at a lot of different battles you need to do, but it's also there's fewer points of failure. Because if you have an amazing research, amazing product experience across view, then even if your research isn't amazing and product is, you can still benefit from other research in this space. And the opposite is also true. So, you know, it's the way we, like, as I think about the flanks and small teams and winning, it's like, if we win one of those, we are an amazing company.

46:41If we win many of those, we can create a generational company. And the teams feel that too. Like, if we all do great work, then, and I think there's no other way. You cannot create something generational if you do, like, you know, one or two problem solves. You need to solve a double-digit amount of different problems to really create something special. And I think we are headed towards that. Why do you care about doing something generational? Like, why does it actually matter? I don't think, yeah, it's a good question. I don't think in itself, you know, like in a way, the work we create has like a positive impact on the world.

47:18You want to extend that impact as much as possible. kind of your creation should outlive you in many ways. And if it can compound and give the value to the ones around it, like what better way of impact of humanity where you kind of extend the technological advancement for everyone out there. So like in my mind, the generation, I'll just shorthand for that, where all of us can create something that will kind of go beyond what we can do ourselves in the future, which is just an exciting proposition of leaving something for the world. Yeah, fair enough.

47:55How, like on the research side, is it as we all read and see, like, is it as intense? Like, pretty soon we're going to need like an agency to manage, like a CAA to manage the like contracts of these researchers. Is one of the flanks that you're protecting Zuck not getting into your research, your lab? I'm joking, but I'm also being serious. You know, in a way, I'm very happy that for the world this happened where the IQ and the research talent is valued in the way they are. It's like in similar, like for me, it's the most amazing athletes you can imagine. And yes, they should be valued in a high way.

48:51There is a tricky piece and a similar in sports where one is more on the micro level, but the next years are the most critical for that world of research and how that innovation will be adopted. So the value is definitely there for the next years. The interesting thing will be, to what extent will the models themselves be able to produce that over time? I think very hard. You still have the set of innovation breakthroughs you will create. But no, I am one of big flanks. I think beyond the flank of compensation, you need to also create a flank within that of like how do you set them to be able to create the best work of their eyes and then release that work and see that work used.

49:43And I think we've been able to do that very well where a lot of the research work that's done at 11labs gets you to the users almost immediately. It's like such a short iteration cycle between the real world application and what you are creating. And then it's a super small, very high density team where there's like, everyone is amazing. Can I explore the like timeframe point on the researcher thing? Because I think it's super interesting. Is what you're saying like, hey, right now for the next several years, whatever that timeframe is, they are crucial to advance the state of the art on frontier models.

50:23But at some point, the models will get so good that their value will relatively diminish because the models will start to train themselves. I'm not putting words in your mouth, I'm just trying to clarify. Yeah, it's a variation of our previous conversation. So less so on the comp part, but most of the researchers that are in the team and most of the researchers even outside of the team when we speak with them, many will believe that AGI will happen in the years to come. And in that world, the whole compensation question is an interesting one. I don't know how this will work in many of those aspects in that reasoning thought.

51:00but that's where a lot of the kind of the thinking process goes through, where the current work and towards AGI will just change the dramatic outcome. Do you believe that? I'm less... Are you allowed to say that you don't believe that, even if you don't believe that? From like the self-serving perspective or like AGI being against me, if I say that? From not AGI being against you, from the... Yeah, I guess like self-serving, but also like... I don't know, I just haven't really heard anybody that's in the eye of the AI hurricane not be maximally bullish. I'm maximally bullish there will be a lot of breakthroughs, but I'm more on the longer-term spectrum.

51:43I don't think it will happen in the next two or three years. So that's where I would disagree as I think about the conversations with a lot of the research folks, but they know more than I do, so I should trust them. But I do think it's going to... A lot of the innovations on the foundational stack will happen and then bringing that to the real world will take a much much longer time frame just like the um the the business process of bringing that into an organization people adapting to how to work with this will just take such a long time so you know it's it's it's it's like closer five to ten years in that dramatic shift um uh and then it's it's like you know Now, what's the distribution of compensation then look like?

52:26It's an interesting one too. Is the knowledge work that's replaced or is it the creative work that's adjusted or is it something completely else? I'm here is more in a category where I think there'll be more of the middle to middle where the super smart people or super creative people will use those tools to their advantage. So that gives you that superpower. so I think you still will be in that category but you will want to you will need to want to do that work too have you had any crazy insane recruiting stories like have you had to do things that you were like I cannot believe I just had to do this to win or save this candidate the many but I don't know which ones I can mention no names but the AGI piece of like the world will change is yes That has happened a number of times.

53:20So we're just, which is, so the previous part was like, we had people, it's like, hey, how, one we didn't hire, but it was, hey, I want to maximize what I can get in the next three years, because I think in three years, the world will not lose. So it's a part of their negotiation strategy. Because they also view their own work, along with everybody else's, as diminishing. Yeah, no, exactly. If we achieve some artificial general intelligence. Exactly. That's insane. That is, that was interesting. That's really dystopian. Dystopian and very like, so we didn't end up proceeding there, but it was impossible to convince the other person that it, you know, like how it has impact for everybody else in the company, how that kind of sets up a bad precedent, et cetera.

54:07But it did, it was like, it's very hard to convince someone that AGI will not happen if they are spending, you know, most of their work and life in that space. But that's kind of the question that I have. Like, it is very hard to convince the smartest people who are the most entrenched in this technology that AGI will not happen in the next few years. Whatever AGI is, we can figure out that definition. But like something that does not exist today. But like, I guess like, why don't you or I believe them? that like that's that's the question that i keep coming back no i i part of it's like there's some percentage that i do it's it's you know it's uh it's a big question where i think there's probably much more knowledgeable people than i that can take it but um no i think there's like a good degree of percentage percentage points that that this might happen and they are fully right it's insane it's insane what do you take what's your take i'm probably one of the uh least disqualified people to talk about this relative to the folks that that you work with every day but my point is like man i'm still trying to get an llm to not sound like an llm like i'm still in the world where i'm getting sent text that sounds shitty like an llm and everybody else is talking about this world where compensation won't matter in a couple of years and i'm like a lot will have to go right for us to get from this point to that point, you know?

55:43And so I'm like, all right, like there's a few things, at least in the enterprise that have worked really well. Like coding has worked really well. Support has worked really well. Like there's not that many things that have actually worked really well yet. And then on the consumer side, like I've seen some behavior change, certainly my own, where I have a sounding board basically at all times. But I'm not, it has not, it's amazing. It is profound. It makes me more productive. But I don't feel the way that I think about my compensation or the value of my real estate or, you know, like cancer getting solved.

56:34Those have not, I haven't changed my milestones in my head. And maybe I'm like some laggard, but like I live in Silicon Valley and like at KP. So if you asked a lot of other people that are not here, I don't know, maybe they would reflect more of that sentiment of like, I don't know. I don't know, that's my take. I'm more in the longer timeframe because I think it will like so long to adopt, Even if you have like a, you know, like a knowledge, knowledge, like the smartest human in your pocket available that you can, you can speak with at any point and the kind of thousands of them. It's still like before you can realize all the outcomes for the world will take a longer period of time.

57:13Like even in company building, everybody's probably telling you like, oh, you can, you don't need 330 people. You can do this with a hundred people. and then what ends up happening like reality hits you across the face opens open eyes like hiring diagram what is it it's it's like scaling as quickly as it ever ever did ever did right yeah and by the way uh even in that case they're like you know 50 or more of all of software is probably being written by ai within that company and they're hiring a ton of engineers no 100 i it's so so that's yeah case in point these are like i mean i just i don't like if everybody that's smarter than me is telling me this is definitely inevitability then why isn't the behavior of the places like why is no no i that's i think we agree there i think it's going to take much much longer time before it's adopted and across all products across all dimensions and in uh your 330 you're just crazy to think to me as well we we were just um a year ago we So it's like a lot of people joined in the last 12 months.

58:17But the deployment of the technology, like what is go-to-market, what is engineering. So on the go-to-market side, working internationally across all the companies and being able to speak through the process, be a partner into how to use that technology is a lengthy and time-consuming process. You really need people there. And on the product side too, which goes back to that point of like, you can have the best foundational model, but to derive value from any of those models, you need a pretty solid product layer on top. So I'm with you on the longer timeframe. Fascinating. Like in your neck of the woods, where do you live?

58:53So I'm living between London and New York at the moment. Okay, London and New York. Like, do these conversations even exist? Like - They do, but less, you know, it's an interesting one. For NSF, they are definitely more frequent, yeah. I guess that's like, you know, the density of founders and companies and investors. It's just much higher. So I'll use that as my litmus test of whether the AI ecosystem is growing in those two. That's right. That's right. Or Waymo comes over there. Or Waymo comes over there. I heard Waymo's coming to New York or is in New York. That would be a good test for Waymo.

59:32That's about as good as it gets. That's about as good as it gets. There's WAIF in London, which is great, which is self-driving software for cars. and it works really well. Yeah, yeah. Anyway, and maybe on the headcount piece. So like, as you forecast your headcount for next year, are you gonna double? So two answers there. One is we mostly think about headcount as like a corollary to the output we want to achieve. So if we know there are specific projects we want to run or scale, then we'll increase the headcount. They are not like a goal in itself. um but as we look about kind of where we want to be in the next 12 months yes i think we will double double the head count i think i think somewhere between somewhere over 500 people in the next 12 months is is is likely as we think about scaling especially now we work very internationally so we have a lot of our clients here in us but then also across south america asia europe uh so that's going really well to lots of flanks that we need to win across and and those products especially in enterprise setting um are starting to get more complex and we want to be able to deliver to any complexity so that will require more people um and then research too like we see increasing move from from from from the models we create into how do you combine them with other modalities so we can derive that value.

1:01:01So that will also evolve a little bit. Do you sleep well? That is like a genuine question. I love it. Yeah, I think so. Like you're enjoying this. It's not like overwhelming you. So I sleep well. Caveat, the travel makes it a little bit harder because I travel a lot. And it's like I do a quarter without travel, almost a quarter with a lot of travel. And that's tricky because then you swap times on so much that even though I'm trying to like measure and get asleep in the right cadence. I don't know whether it's good for me, but just from emotional perspective, look, it's a once in a generation shift that's happening.

1:01:43AI is changing the world around us better. Regardless of the timing of a lot of those things, it's changing so much in the process work that we do and getting us smarter, quicker across the entirety of our lifetime. And we can be the voice of that change. We can be the voice of technology, which is a rare opportunity and chance that both I and I think the entire team have luck to be working on. For sure, man. You do quarterly travel. I do. I try to be in a one continent in a one quarter roughly and then maximize my travel during that quarter. And then another quarter focused on the deeper work.

1:02:23That makes sense. I appreciate you doing this. Thank you. Congrats on the startup. I'm excited for you too. Thanks, man. I appreciate it. I can't wait to see. I cannot wait to see how this unfolds. And I cannot wait to see your role in it. I cannot wait. I think regardless of how that unfolds, we need to do a test of checking whether you can actually tell the uncanny value for some of the voices. I think the quality is pretty good. Well, the amount of freaking podcasts that I've done, you have plenty of training data. There's plenty of things that you can use. You can work together. There's plenty of things that you can use.

1:02:59Are you hiring? What are you hiring for? Where are you hiring? We are hiring. We are hiring constantly. The current top roles would be helping us on a lot of the product work. So if you are excited about building on the frontier of agents, frontier of creative work, we need people to bring it to the enterprise, work deeply with the customer and figuring out how to deploy that across some of the Fortune 500s. So whether A, you are just in a great full stack or a great deployment engineer, forward deployed engineer, that's great role. Two, on the go-to-market side, similarly, I do think the role of how close you are to GPT 3.5 can be extended to other parts as well.

1:03:45But if you're excited about the technological shift and doing zero to one, then on GoToMarketSide, we love technical, highly analytical people to join the team as well. Yeah, this would be top two. I love my Cs. When you hear the word grit, what do you think of? When the hard times are there, are you still fighting and going after it and try to do the best you can to go through it? Thank you, man. Thank you. That's it for now. If you liked the episode, please leave us a review or go back into the archives where we've done more than 200 episodes with some fantastic folks. This podcast is a Kleiner Perkins production, and I'm Juven.

1:04:28Thanks for listening.

From the publisher

Before AI became a buzzword, a few true believers were already building.

Since early 2022, Mati Staniszewski and his team at ElevenLabs have been among them, working to create voices that “actually represent emotions.”

He shares with Joubin Mirzadegan how voice AI is transforming diverse fields, from delivering personalized healthcare for different age groups to amplifying creativity in filmmaking.


Guest: Mati Staniszewski, co-founder and CEO of ElevenLabs


Connect with Mati Staniszewski

​

Connect with Joubin

​

​Learn more about Kleiner Perkins

More from Grit

All 63 episodes
ElevenLabs’ Vision for Voice InterfacesGrit · 1 h 5 min
Listen in VO