The world of voice AI, with Mati Staniszewski of ElevenLabs

14 Apr 2026 · 1 h · 29 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode explains how voice AI works and how ElevenLabs builds and deploys it, from audio model representations (phonemes, spectrograms, waveform) to text-to-speech, speech-to-text, and voice agents. It also covers why voice assistants lag behind text LLMs in real-world products, and what’s needed for “voice Turing test” conversational behavior.

Guest

Mati Staniszewski, co-founder of ElevenLabs (founded 2022), scaled it to an AI-audio leader valued at $11B; credited with innovations around realistic emotional inflection and controllable, context-aware voice generation.

Key claims

ElevenLabs’ “two big innovations” are (1) better context-driven prediction of speech units (phoneme/token-level) and (2) open-ended, non-hardcoded control of voice traits (accent/style/emotion) so “Britishness” can emerge rather than be selected from fixed parameters. They invest heavily in custom labeled audio data and use diarization + keyword detection for transcription.

Notable examples

“11 Reader” for reading uploaded PDFs/text with celebrity-like voices (e.g., Sir Michael Caine, Sir Richard Feynman); healthcare use cases like restoring a patient’s voice (ALS/throat cancer) and an operating-room scenario; a Guinness pub-price voice agent in Ireland; and enterprise adoption with Deutsche Telekom/T-Mobile/Revolut/Klarna/Meta/IBM.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Audio Models

0:45 to 3:00

Exploration of how audio models are developed and function.

“Then that progressed into trying to create effectively like a digital signals for speech.”

Advancements in Speech Technology

3:00 to 6:00

Discussion on modern speech generation techniques and innovations.

“But kind of what happens before and after comes into the equation, and you need to bring that across.”

Data Annotation and Model Training

6:00 to 9:30

Insights into the importance of data annotation for training voice models.

“So it's similar to how you would operate on the token level on the text side.”

The Business of Eleven Labs

9:30 to 12:45

Overview of Eleven Labs' products and services in voice technology.

“from customer support, sales, hiring training, all the way through to marketing and storytelling for our creative tools.”

Challenges in Voice Technology Adoption

12:45 to 14:01

Discussion on the gaps and barriers in deploying voice technology.

“And one thing I note is that LLMs are amazing.”

Voice AI Readiness and Deployment Gaps

14:01 to 15:00

Discussing the current state and challenges of voice AI technology deployment.

“In the lived experience of people day-to-day, like they're using series transcription, which has gotten better.”

Advancements in Voice Models and Contextual Interactions

15:01 to 18:16

Exploring recent advancements in voice models and their application in different contexts.

“I think that's only recently became possible and where we've seen the big adoption across the enterprises leading on the technical side.”

The Challenge of AI Audiobooks and Accessibility

18:17 to 21:20

Discussing the effort to create a platform for AI-generated audiobooks due to accessibility issues.

“but I don't know about you, it just doesn't work.”

Personalized Voice Transcription and Recognition

21:21 to 24:26

Examining the future of personalized transcription technology and its potential.

“Customers save their details once and then they can check out in seconds across more than a million businesses with safe credentials.”

Innovations in Speech Generation and Emotional Control

24:27 to 27:33

Highlighting breakthroughs in speech generation technology and emotional responsiveness.

“We can already diarize speakers extremely well.”
Show all 29 chapters

Cascaded vs. Direct Speech Models

27:34 to 28:00

Delving into the differences between cascaded and direct speech processing models.

“The edge cases of how you want to describe it is pretty large.”

Speech-to-Speech vs. Cascaded Models

28:00 to 29:20

Explore the differences between speech-to-speech and cascaded approaches in voice AI.

“When I say speech-to-speech, is that the idea that it doesn't go through text as an encoding in the intermediate set?”

User Interaction with Voice Technology

29:20 to 31:10

Understand how users engage differently when using voice AI compared to traditional forms.

“And maybe in the future, just to finish that part, you will have some version of combination of the models.”

Second Order Effects of Voice AI

31:10 to 33:50

Discuss the broader implications of advances in voice AI technology on communication and accessibility.

“Because it's like an open-ended adventure game.”

Business Aspects of Voice Models

33:50 to 35:50

Learn about the economic factors and operational costs involved in developing voice AI models.

“And for the first time, she could replicate the marriage ceremony and speak the vows together, which was like such a heartfelt moment.”

Voice Model Development and Market Dynamics

35:50 to 38:10

Delve into the challenges and strategies for developing and scaling voice AI technologies.

“And then there's this kind of ever larger CapEx going into, I mean, a lot of it is inference these days, but also training.”

Conversational Agents and Their Applications

38:10 to 40:50

Explore the various use cases and future directions for conversational agents in businesses.

“It's still usually like not as reliable.”

Integrating Voice and Text Interactions

40:50 to 42:01

Examine how voice AI can complement text-based solutions in business interactions.

“But you are more of a partner in their AI transformation part.”

Voice Interaction and Customer Support

42:01 to 43:36

Learn about optimizing voice and text interactions in customer support.

“And so how does this break down between, you know, we had Dev Trainer from Intercom on here and they have Finn, their agent, and it's a thing in the website that you can go talk to.”

Revenue Growth and Business Expansion

43:37 to 45:30

Discover how ElevenLabs achieved rapid revenue growth and what drives their enterprise success.

“Voice is usually a big part of those interactions.”

Self-Serve Strategy in Tech

45:31 to 47:26

Understand the advantages of offering self-serve technology in a competitive landscape.

“And I think largely that technology that powers a lot of their agentic interactions just became reliable at the same time as high quality over the last year, year and a half.”

Feedback Loops and Product Accessibility

47:27 to 49:22

Explore the importance of feedback loops and making tech accessible for developers and SMBs.

“And on the completely other spectrum, we have the high-touch forward deployed engineering working side by side with the customers to customize the entirety of their work together.”

Organizational Changes with AI

49:23 to 51:42

Learn how AI advancements impact organizational structure and team dynamics.

“in the foot by not, like did you guys self-serve on Stripe or did you?”

Using Software Effectively in AI-Native Companies

51:43 to 56:05

Discover how ElevenLabs uses software to improve processes and enhance team efficiency.

“And so that could be about what the scaling factor is of like the number of people you need to do the work.”

Enhancing Business Efficiency with Voice AI

56:05 to 56:55

Learn how voice AI can assist in creating customized solutions for business meetings.

“So you have a pre-populated deck with the right numbers that is customized to that customer, which you want still the person to go through and develop, but ultimately is in there.”

Supporting Citizens in Ukraine through Technology

56:56 to 58:00

Discover how AI-driven solutions are providing essential services to citizens in war-torn Ukraine.

“So, of course, in Ukraine, with ongoing work, they need to rethink a lot of how their development, their systems, their support works for the citizens across the country.”

Building a Culture of Agency in AI Development

58:01 to 58:59

Explore the importance of agency and culture in scaling AI teams and fostering innovation.

“They actually have the same in every of the ministries.”

The Impact of High Agency on AI Success

59:00 to 59:20

High agency individuals are more likely to thrive in the evolving landscape of AI.

“that I think works well for the AI world is agency.”

The Joy of Innovation in AI Workspaces

59:21 to 1:00:09

Learn how a combination of agency and enjoyment fosters a thriving workplace culture at Eleven Labs.

“That was probably the biggest validation and happiness.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:02Mati Staniszewski co-founded Eleven Labs in 2022 and has since scaled it to the$11 billion leader in AI audio. He's credited with capturing the humanness of speech through realistic emotional inflection, and they're now expanding into everything from agentic workflows to music.

0:17Mati Staniszewski:Thanks for having me.

0:21Let me go place to start is, describe to me how, like, I know how an LLM works at a high level. describe to me how an audio model works. Like if we were Carpathia style looking to build a toy one from scratch, how does it work?

0:36Mati Staniszewski:In the early days, you try to replicate it exactly like you would replicate it with the human body. So you would try to completely try to reproduce a machine, analog machine, that will create a vocal tract effectively. Then that progressed into trying to create effectively like a digital signals for speech. Bell Labs was one of the first to try to like create a structured set of signals that would represent the speech. And that is a first precursor to what we would do today. Then you would try to stitch in phonemes, effectively different sounds of how we would speak humans, and then try to concatenate them together.

1:10Mati Staniszewski:It's another important part in that equation where you would, based on the most probabilistic approach of the next word, you would effectively try to bring the phonemes from your library of phonemes and bring them together. And then down to the modern history, where now we effectively do similar neural nets in other domains. So you predict the next sound based on, of course, the context of the previous sounds. If it's a streaming speech, if it's, let's say, a context of audio, you will use a combination of predicting of the phonemes, but you also use the contextual text element of that work. And here, credit to my co-founder, Piotr, who effectively came with that new idea of how you can now create voice models which are both reliable, high quality, quick, where you would bring a lot of the ideas from transformer models, from diffusion models, into the speech space.

2:01Mati Staniszewski:So that prediction of the next token on the phoneme space wasn't something that was possible. You might be, you always talk briefly about this, of like how you kind of operate on the text, on the waveform space, there's also mal-spectrogram space. So like usually you do text, mal-spectrogram waveform. So what's a spectrogram space? It's like a visual representation of how the speech sounds, across pitch, across energy, and then you transform that into a waveform. Got it. So like when WaveNet came along and Tachotron models, they would effectively use text to male spectrograms, so that visual representation, and then how you decode and encode that into the waveform to bring that across.

2:36Mati Staniszewski:And Piotr figured out how to like abstract some of those steps and decode and encode them a lot, a lot better. So that predicting of the next phoneme was one of the big piece. And second big piece was how do you bring that context into the equation? So what I mean by context is if a voice actor was reading a textual copy, you would know that, okay, this is a dialect sequence. I need to produce a dialect. If it's a happy sentence, I might need to pronounce it as a happy sentence. But kind of what happens before and after comes into the equation, and you need to bring that across. And then there's a last big piece.

3:09Mati Staniszewski:So voice model has the sound of how you intonate the given fragment. But the second big part is the voice itself of the characteristics of accents, of style, of prosody across that voice. So when you actually try to vocalize something, when you create that voice model, you turn text into audio, you need the text. You also need the voice reference of how you want it to be spoken. So here is kind of the second big innovation. So apart from context, it's how you decode and encode those features. So when Bell Labs came with their initial representation of speech, The big piece there was you would have effectively hard-coded parameters for that speech.

3:48Mati Staniszewski:With 11-lapse models... Hard-coded parameters for enthusiastic speaker, British accents. Exactly, exactly. That kind of stuff. Like the set of pitch elements that you can select, set of energy spectrograms you can select from. And in our approach, effectively, you would give the model open-ended ability to select what those parameters should be. So it's not going to be British, Polish, Spanish, English speaker. but the model will deduce them themselves. The same for other set of parameters that are not hard-coded, whether it's the enthusiasm, whether it's the sadness, etc. You're saying kind of Britishness is an emergent property in your voice models.

4:26Mati Staniszewski:Exactly. Yeah, and those kind of those two big parts, so it's encoding and decoding of how you create the voice. Super hard problem before and figured out too. How you then construct that in the sense is how you get the context across so you can predict the next phonemes. So how you bring them together in a reliable and stable way while doing it quick. And these were kind of the two first big innovations in the voice models that continue to today. But OK, so if LLM's reason about text and words and parts tokens as the way they think about the world, what is the equivalent of a token in the voice model?

5:00You mentioned phonemes a bunch. Like, what is that representation?

5:04Mati Staniszewski:So we do, we store the voice embedding effectively for the speaker. so you need that reference when you produce and create the speech. Of course, in the input to the voice model, you still get the text and you bring the speaker and coding. And then when you produce speech, you do operate on the waveform or effectively on the phoneme level of that speech. And then when we kind of go the opposite, so of course... Sorry, what is a phoneme? Fill in my understanding. It's like a syllable deconstructed even to smaller elements. and these are effectively the human sounds you can produce. Got it. So these would be the most close to that representation.

5:44Mati Staniszewski:But of course, in our models, now it's going to be a combination of not only operating on phoneme level, you also operate on the text level. You operate in both in sync because when you are predicting the context, you need to understand how that sentence will get constructed and especially if it's more of a streaming real-time use case in a voice agent setting. You need both parts to work across. So it's similar to how you would operate on the token level on the text side. We operate on the token level on the audio side. It feels like a big part of the magic of 11 was your voices were much more human sounding.

6:19How did you accomplish that?

6:21Mati Staniszewski:So I'll kind of give you a quick synopsis of how we think about the models on the text-to-speech side today. In any model, you need architecture, you need compute, you need data. So architecture innovations were one thing. The data part was the second big thing. With audio, you will have a lot of audio data available, but frequently you will not have it annotated in the right way. You won't have which speaker is speaking when. Some of the what is annotated, but the how isn't. So like as we are speaking now, what's the emotions that we use? What are the actions that we use? So we would invest a lot internally on effectively creating our own data labelers, our own team, to be able to create those data sets that will be better.

7:05Mati Staniszewski:And that was a combination of, of course, semi-automatic techniques and then manual techniques. And actually, a lot of the models that we did afterwards actually spun out from a lot of that research, too. So speech-to-text model initially was a model we did for ourselves because the models on the market just weren't good to annotate that data. And then another brilliant researcher on our team was kind of being able to construct it so we could span it out as a model that we brought to the customers. So you've just been doing useful stuff in voice, and that has emerged with a whole bunch of products that you might not have expected because you find you're building useful stuff.

7:39Mati Staniszewski:Exactly. Exactly. And that kind of combination of data of being able to do it automatically, create a team that's coached on voice on how to describe it. Yes. Because most of the labelers out there just aren't as well versed on understanding the audio and voice. Helped us a lot to bring that back. And then, of course, deploying those models in production, seeing how customers interact with them, having them annotate all of the data helped us refine those models over time. A very interesting thing on the side. So we spoke about kind of the speech representation. The first guy who created the speech representation is a guy called Kempelen, von Kempelen.

8:13So he created this analog machine that would represent effectively a human vocal tract and try to produce that sound.

8:21Mati Staniszewski:He had spent decades on that and kind of started producing vowels. But that's the same person that created a chess machine, the first viral, let's say, chess machine, that would kind of simulate playing chess. Is this a mechanical Turk? It was called Turk. Yeah, yeah. Exactly. But the kind of crazy thing behind it was operated by a human. Yeah, yeah, yeah. And it was all a fluke. And that's where the mechanical Turk formed, which actually we use in that kind of data labeling production to make that work there. Yeah, yeah. And sorry, we kind of jumped right in. But if you describe the Eleven business today, people think of you as the speech company.

8:59How should they actually think of your business to the extent you can describe the big areas? Text-to-speech, speech-to-text, voice agents, just like break down the business for us.

9:08Mati Staniszewski:Cool. So in like the nutshell, I'll describe Eleven Labs as a research and product deployment company. we build foundational audio and voice models and then build a platform for businesses to transform how they communicate with their customers, with their employees. And that will apply through AI agents from customer support, sales, hiring training, all the way through to marketing and storytelling for our creative tools. And in that set, we've created all types of foundational audio models. So text-to-speech models for producing speech, speech-to-text models that work over 100 languages and happily beat others on benchmarks, all the way through to conversational models of how you loop them together, to music, to other domains of audio.

9:57Mati Staniszewski:And then, of course, beyond the models, when you actually bring them to production, that's where the second level of the platform comes in, where that meets the businesses on the specific use case. So on the agent specific example, it would be how you now connect those models to the knowledge base, to telephony, to the integrations that you need to perform the actions, how you evaluate and monitor the agent that behaves in the right way, how you build the right safeguards. On the creative side, on the marketing side, it's how do you create a good ad so you can create a good video voiceover for one of the campaigns?

10:32Mati Staniszewski:How you create an article that's narrated with a specific voice that represents the brand in a good way? So that's where we combine the models and understanding of the customers we work with into one policy platform. Every platform company has this question about how far they go into applications. So how do you think about where you go horizontal and power the whole ecosystem versus where you develop applications? Because you can imagine there being a whole ecosystem of closed captioning tools that grow up that, again, are built on the 11 Labs tech. It's not necessarily a space that you would have to go after yourself.

11:06Mati Staniszewski:I think the big difference between, in your kind of question, today we see ourselves as a platform where if you're building a horizontal use case in your business, a great place to come. If you have a lot of domain specificity, that's where I see a lot of kind of application companies forming over time, where that's specifically not the spaces we will go into. And I think it also is interesting when the tech is moving as quickly as it is here. It's one thing, you know, with SaaS where you get these like vertical specific providers. but I would imagine one of the biggest risks for you guys in being intermediated is if there is, like in this example, a closed captioning service that is on a two versions, old version of 11 labs and hasn't upgraded.

11:52That's a problem because you want people to be using the latest and greatest model that you've developed and you'll be kind of deploying new capabilities every week. And I presume that's part of your thinking is that just when it's moving that quickly, you need to go direct in a lot of cases.

12:04Mati Staniszewski:That's right. In the closed captioning, here already now we know that our services is going to be able to tackle 99.9 % of the cases that customers have. And then there's added benefit of we work with healthcare customers where we will create custom models for those customers where we'll get that transcription perfectly. The context is a tricky thing in closed captions where we talk a lot about a lot of technical stuff on this. Yeah. For sure. And that's where you need effectively like a dictionary of words that you detect beforehand, which as we work with the businesses, we know we need to embed in that creation process.

12:42We're talking a bit about kind of products here. And one thing I note is that LLMs are amazing. And, you know, you have the usage stats of CheckGPT and Gemini and all the popular LLMs where they're working and people use them a ton. It feels like there's a big product overhang when it comes to voice, where the leading edge voice models are incredibly capable. And yet I was driving home the other day and I needed to read a PDF while I was driving. And so I said, okay, I'll just have my phone read the PDF to me. And you can kind of try and hack it with like iOS screen reader, but it doesn't really work with the scrolling.

13:17And then in theory, you can upload a Gemini, but you're trying to get it to not summarize it. And it actually just hung when I tried to press the, like, read this to me button. And so there was no way I could get my phone to read me something, which seemed like a fairly basic feature. And all cars advertise voice control, and yet it sucks separately. If you want to input something to the navigation, just no car has a good version of that yet, maybe Tesla does, and, and, and. And so why does it seem like with LLMs and cloud code and everything, we are using all the capabilities of the intelligence, whereas with voice, we're like living 10 years ago somehow?

13:57Mati Staniszewski:Well, I'm thinking what I agree with, the premise that we are 10 years behind. In the lived experience of people day-to-day, like they're using series transcription, which has gotten better. And it's still way behind the leading edge. Yeah, like there is definitely a piece of like, I think the technology in many of those cases is ready. There is a deployment gap to what you are saying. It's like an automotive or some of the big companies are not adopting that quickly enough or bringing that into the production. but like plenty of different problems that you need to fix along the way. I mean, the quality of voice models for them to actually sound good, like this is only like last three years thing.

14:34Mati Staniszewski:Yeah, those three years. It's a three years thing. Cars have over-the-air software updates now. So that's three years for the first voice model that can narrate text async. Two years ago, you can start seeing the real-time version of that and not really like it's, I think the real break was like a year ago where you can start seeing that in production. And then I think over 2025, the big piece that hasn't been possible is how you connect now the real-time voice interaction with something which I think you are referring to. It has context of what you want to do, what is the material that you want to read, how does it connect to set up your preferences from the past, and gets that across.

15:09Mati Staniszewski:I think that's only recently became possible and where we've seen the big adoption across the enterprises leading on the technical side. I think this year it should be in the automotive site too, or some of the applications. Okay, so you think we'll start seeing kind of great voice models in cars this year? This year for the on-cloud use cases, like on-car, in-car, so without connectivity, not yet. There's deployment, of course, gap of how you bring that into the gaps. But I think the next two years, three years. How about the PDF reading use case? That should work, yeah. Well, how should I have done it?

15:48Mati Staniszewski:So back in the day, I'll preempt this with a story to cue 11 reader, but we had this problem. We have so many audiobook authors come into 11 Labs. So 2023 released first software. We had a lot of creators and then a lot of audiobook authors or book authors that couldn't afford professional narration and wanted to create an audiobook. However, none of the companies accepted AI audiobooks. And you can't sell an AI audiobook on Audible or something. Exactly. So Audible would block AI content. So we had no choice. Like we need to create an avenue for them to bring them. Because there was no distribution for AI audiobooks.

16:24Mati Staniszewski:Exactly. So we created an 11 Reader. And that kind of came with functionality where you can upload your PDF, you can upload your text and have it read out loud with a number of incredible voices. So whether it's Sir Michael Caine all the way through to a state and working together with Sir Richard Feynman. So this is you're working with the Sir Michael Caines of the world? Exactly. And then you can actually read it out loud. And that kind of works extremely well. So that works. Now, how can you do it? Actually, I do want everything read to me by Michael Caine. It's a great voice. Yeah. Shouldn't you guys have a consumer app where I can just do the common voice things?

16:58Like, I want to be able to have an 11 app on my phone, and then if I upload a PDF to it, it can do the common things that I would like, such as have it read it to me. Yeah.

17:07Mati Staniszewski:That's exactly your own reader. So that works. The phone makers allow third-party keyboards. Do they allow third-party transcription engines? Will they, do you think? The phone makers, you said, right? Like Apple and Google. Yeah, they... So the OS makers. Yeah, not all of them. Android, with Android, you can work through it. There's like, you know, variations of that. Nothing, the tag, and others. But again, I feel like if you had a popular 11 app that allowed for transcription, people would use it a bunch, and maybe eventually Apple would say, oh, we should allow third-party transcription engines if that's what people want.

17:39Mati Staniszewski:I mean, it seems like they might be going in that direction, right? And I recently announced that we'll open up the LLM ecosystem. Hopefully they will do the same with voice ecosystem, which is kind of similar. But again, I think it's rational to do when it's moving so quickly. Yeah. The voice assistant paradigm is one of the oldest paradigm, you know, UI paradigms in computing. Like the open the pod bay door as hell from 1969. Yeah. I will claim it's not working yet. So Siri doesn't have the intelligence. And then on Gemini and ChatGPT and those apps, I mean, I want to use the voice mode, but I don't know about you, it just doesn't work.

18:19And so like sometimes I'll be using my phone and I'll use the iOS keyboard transcription to type in the field and then like say a bunch of stuff and then send it off. But this suggests to me that consumers really want voice mode that works. and yet it's just not working yet for the major LLM apps or for anyone. Why isn't it working yet?

18:41Mati Staniszewski:It is pretty hard to do because you want two things. You want to be able to say things that you want, but you want sometimes for it to execute it, sometimes to wait for you to finish and add something in a sentence. Sometimes you want it to be interactive so it asks you questions back to clarify and get some of the additional detail. And all of that is actually pretty hard. That's where the magical, like ideal version of a voice agent for us comes through where you need the speech-to-text element, you need the transcription side, unique, you need then kind of the turn-taking mechanism. So like when do you finish sentence?

19:14Mati Staniszewski:When is it likely based on silence, based on the context? And then sometimes you want it to speak back and clarify or at least give you the text back to clarify and then maybe execute a set of instructions. So that problem is still very hard research. So I agree with the claim that like this orchestration side has not passed a true conversational agent Turing test where it behaves as you would expect from another person. That's the simpler way of saying what I'm saying is that we have passed the Turing test with text LLMs a long time ago, and we're actually nowhere near that on voice LLMs. It's going to be interesting how that's a final frontier.

19:47Mati Staniszewski:Yeah, I feel like it's going to work in specific domains. In customer support, call passes the voice Turing test. It works well. Let's take another spectrum of that. an interactive gaming experience, like a truly interactive as you would have with another human in that game, it's like, you know, so hard and further out there, we haven't passed it yet there. Yes, yes. Yeah, but I think that's a combination of like, you know, like even like a simpler version of within that, like sometimes you might give a response immediately back. Sometimes you need a tool call to get additional information from the database.

Read the full transcript

20:22Mati Staniszewski:So how you orchestrate that. So like that's probably the most common thing we see as we work with some of the companies out there is you want those systems to orchestrate extremely well, where if it's a conversational use case, pretty simple. You can root the agent to speak with. But if you need to authenticate, if you need to pull additional information from the database, what do you do? How do you handle that graciously? That's where you're going. And yeah, and to that extent, I would agree, that's just getting, getting, getting there. And we'll hopefully see that our goal is to pass the voice Turing test in all those cases, or the Turing test for all conversational agents outside of voice too.

20:56Mati Staniszewski:and I hope we'll all be there in the next year or so.

21:01For subscription businesses, a lot of revenue is lost in that last few seconds before the checkout. Someone has to get up and find their wallet or they mistype their card number or they hit an error. They just give up and you lose the sale. For a company like 11 Labs, adding hundreds of thousands of subscribers, even a tiny bit of friction like that, it would really add up. But that's why 11 Labs uses Link from Stripe. Customers save their details once and then they can check out in seconds across more than a million businesses with safe credentials. So if you want a faster checkout for your customers, you should turn on Link from Stripe.

21:37Are you guys working on personalized voice transcription, where it feels like part of the way we're making it hard for ourselves is when I speak to Siri, I have a bit of an accent, and so it sometimes has a hard time understanding me, but my accent doesn't change. And so it could just get good at listening to John. But my understanding is it's not. It's just like running the global voice recognition model. And I'm guessing it's the same for 11 Labs where you're running the global voice recognition model. But again, you have an accent. And so if someone's understanding, like if you walked up to someone in a coffee shop and said two words, they might have a hard time understanding it because they're not putting it through their matty Polish accent filter.

22:17And so where's this going with like actually interpreting the person that you know to exist on the other side?

22:24Mati Staniszewski:Yeah, I have a very tricky one to detect. So my voice is frequently used in the test. Ah, you're a part of the test suite. For text-to-speech, for speech-to-text, for like everything. Yeah, yeah, yeah. I'd be like, yeah, it's pretty, pretty. But again, trying to parse your voice in a global model is just making life hard. It's like have a Matty-specific model. Yeah, so on the speech-to-text by transcription, exactly. Like the big part now that we are bringing in is you have two parts. One, effectively like a person or a voice-specific detection. which is true for the accent side, but it's also true for a crowded room.

22:57So that's where we have an incredible research team

23:00Mati Staniszewski:that's able to continually do both the accuracy high, but also add things like speaker detection, of course, noise reduction. But then the second part is also keyword detection. So there are specific words that you would want to say in those settings that you want to effectively monitor for. So we spoke about, let's say I'm going to the coffee shop and order things. The set of actions the coffee shop would expect me to do is pretty limited. There's information theory. It's like they can just listen out for the coffee words. Exactly, and then try to match it to the closest proximity. So both things will help.

23:35Mati Staniszewski:In a setup where you have my voice, perfect. You can decode it and code it on that. If you don't have my voice, or even if you want to double amplify it, we already support effectively a keyword detection, which is useful for real-time setting and async setting. So back to the cheeky pine transcription, you're going to effectively pre-generate that from the previous podcast and look for a set of words that you would use traditionally in that. And so how hard, okay, so you did the keyword detection already, but how hard are the, I want to get superhuman transcription performance by feeding it an hour of Matty audio before it listens to Matty, and then it should be able to do a much better job transcribing.

24:16Is that just a really hard research problem? No, solvable.

24:20Mati Staniszewski:We think we can roll it out in one of the next versions, which is hopefully in the next months. Oh, so you think this year you're doing person-specific transcription? Person-specific transcription. We can already diarize speakers extremely well. So if we are speaking, we can, of course, dissimilate who is speaking when, which is like in transcription side, apart from accuracy, diarization is one of the harder problems. And we do that extremely well. And now it's going to be effectively what you were saying, fine-tuning based on the speaker that I want to listen to, which we know will be important.

24:53Mati Staniszewski:I mean, in healthcare setup, such an important part. You're in an operating room, you're a doctor, you want to say a command, then you want to really be able to listen to that one person-specific piece. You have a hardware device at home, let's say it's a pilot, that helps you control the TV. Here too, you will want that to listen to you versus, let's say, the family roaming around. Or maybe you want it to everyone. So you could decide that, but in many cases, you want to be able to specify that. Okay, that's really exciting. It's great because there's still so many unsolved research problems.

25:27There's just breakthrough after breakthrough coming in the domain of voice models. How about on the flip side when it comes to speech generation, the Zoom touch up my appearance feature? I've always thought about that in the context of voice, where should you offer a de-accenting filter for voices? or even this one podcast that I like to listen to, but the voice is a little mumbly, and I always thought they should put it through a de-mumbling filter just to make the enunciation a little better. But all these things, again, like Photoshopping an image, there's no reason that the... Have you thought about voice-to-voice, basically, rather than voice-to-text or text-to-voice?

26:07Mati Staniszewski:Yeah, so there are kind of two big parts. One, on the speed generation side, similar. So many ovations still there. there's like a wider piece and that's like the we released the v3 model that i kind of we're solving that for the first time is like can you control speech so you can have the text to speech you generate something that sounds emotionally great previously until until end of last year effectively you would rely on model to decide what's the best performance you could regenerate it but that ultimately model decides the best performance so that's where the controllability came in where we can finally give it cues of say it in a slower way or change how you deliver the dramatic pause or kind of any cues that you give and to be able to do that you need the architectural changes and the data that we kind of created over time where you annotated what was said and how it was said so you can actually train the model to do that so today finally you can have both speed generation or entire voice agent experience um with with what we call expressive mode where the agent knows the emotions on the other side so if the person is stressed, it can react and be reassuring.

27:11Mati Staniszewski:And that's generating a limb response on the reassuring side and response in that set of emotions too. And that breakthrough was super hard to do. And that, of course, stretches to a lot of what you said. It could be some version of speech enhancement, either real time or in a post setup to change how that's delivered. And that's relatively recent innovation. And we know it can still be so much better. The edge cases of how you want to describe it is pretty large. So that's one. And then the second part of the question, which is a huge question, that's speech-to-speech models. So as you said, our approach, as you think about voice agent, conversational agents, is effectively a cascaded approach.

27:52Mati Staniszewski:You use transcriptional speech-to-text, LLAM, text-to-speech, and orchestrates all of that together. And then you have a speech-to-speech, which kind of goes directly from speech, and there's a speech response on the other side. When I say speech-to-speech, is that the idea that it doesn't go through text as an encoding in the intermediate set? Oh, interesting. For performance reasons, for accuracy reasons? You usually do it for latency. Okay. For latency. It is faster to run a model that does not have to transcribe and then generate. Exactly. It's quicker, but on the flip side, you lose reliability.

28:23Mati Staniszewski:Yes. You lose all visibility into the parts of the pipeline. In emotionality, we think you can deliver both on both sides extremely well, and maybe you can make it more controllable too. So today we are optimizing heavily on a cascaded approach. I'm sorry, a cascaded approach is? Is the speech-to-text, going through the text layer. Oh, okay. Going through the text layer. And as we work with all of the businesses and enterprises, they will need that visibility into what happens. They will want to execute certain tasks on top of that. They want a good visibility into each of the steps and great accuracy of all the models.

28:55Mati Staniszewski:But beyond that, they can abstract away what's the LM layer, what's the intelligence layer. The integrations are easier in that system. So that's where we are betting a lot of the research work of how you can make that great. And we think we can make that great. And speech-to-speech, as you think about maybe more of a companion version of the applications, that's where that will flourish because maybe the hallucinations aren't as important, but the latency is a little bit more and maybe hallucinations are even a feature. And maybe in the future, just to finish that part, you will have some version of combination of the models.

29:25Mati Staniszewski:That for low-complexity, easy models, you will have speech-to-speech. And for higher complexity, you will have the cascaded. Okay, so I was going to ask about this. You know the way there is research on how the invention of writing changed humans' brains and just changed the neural pathways in ways beyond the actual written language. Do you observe that speech-to-speech models think differently than cascaded models? It sounds like they're dumber. They are definitely dumber. You need smaller model. You cannot. But that's interesting, right? That like forcing models to reason about text. I mean, I know they just have much more in there as well, but they're smarter.

30:10Mati Staniszewski:Yeah, but it's like, you know, like if you are going speech to speech, usually you will use smaller models so it's still quick. So you can like say that. I see. So it's also just a model size thing. Okay, but are there interesting differences beyond like correlates like size? What I can say is like slightly different to your question. The people interacting through voice and the performance we see for how they interact with the business changes just by nature of interacting with voice. A good example, you can contact 11labs and register for your interest. You go through the form. And at the end of that, we supplemented that.

30:45Mati Staniszewski:Instead of going through the form process, you can speak with our agent and leave more details. And what happened there are two things. One, people were actually much more keen to leave the forms through speaking with the agent. So we would go through the form a lot easier. But second, they would be a lot more open-ended in terms of what the use case are. So they would start giving us information about the wider set of use cases, the complexity of the use case. So the writing out was tedious and tricky. Because it's like an open-ended adventure game. Open-ended. You could ask follow-up questions.

31:15Mati Staniszewski:You could clarify. But people were just more at ease and could trust the system while doing that, that it's working. And that kind of helped us a lot. and then free, which maybe is more of a technological barrier, it also works across all languages. So now we have leads from all parts of the world coming in and leaving their details. So we did that use case, and now we have a few different companies building their HDR versions of that too to help them capture the leads coming in from banks all the way to actually one of the automotive companies that leaves that where people are just more keen to speak through voice.

31:50I want to ask about this kind of a second order effect. You have, you know, you've talked in the past about how growing up in Poland, I guess the dubbing of TV shows, they were cheap and so they would only have one voice actor for a TV show. So no matter all the parts, male and female, they're like, I love you, I love you too. You know, there's like one voice actor doing all of them. And now, you know, thanks to better voice models, you'll be able to just have like really good voices, AI generated for all the dubbing. Because again, it's not like it's taking jobs from great dubbing that was happening previously.

32:21It was like awful dubbing happening in Poland previously. So that's like one example of the second order effects. What are the other second order effects you're seeing of ubiquitous good text to speech, speech to text? It seems like across a broad array of languages, because whatever about in English, just this didn't exist in Polish or Irish or, you know, pick your language.

32:44Mati Staniszewski:One, breaking down the language barrier, you know, kind of the inspiration came from the movie side. But it also applies in any communication setup. Like, could in the future, could I travel to another country and speak Polish or speak English? And that language isn't being understood in the local native language. Like from Hitchhiker's Guide to Galaxy, this version of the Babelfare. Exactly. That you can actually understand the world. And voice, of course, will be an interaction layer. But similarly, all of us will have our own kind of extension and voice agents that can help on our behalf. And there is like very clear and great examples of that of people that lost their voice and can get it for the first time back.

33:23Mati Staniszewski:We see that everywhere, whether that's people that lost it due to ALS or throat cancer that can get it back. Just recently, there was an example of a patient that had Neuralink. We worked with them to bring the voice that that person could speak with their own voice back with the family around. We worked with the lady that lost her voice before she got married. And then finally technology became possible. We were able to recreate that voice. And for the first time, she could replicate the marriage ceremony and speak the vows together, which was like such a heartfelt moment. Probably like the most important from all the work that we do.

34:03When you guys talk about voice agents, is a voice agent just the idea that you have some long-running or persistent agent that is going out and interacting with the world through voice? And so customer service being one example of it. In the other direction, your claw going and making you a restaurant reservation and actually calling up the restaurant. Is that kind of how I should think about voice agents?

34:30Mati Staniszewski:That's right. It's exactly, whether it's like the reactive side of being able to interact with the customer or the proactive to call it back. We recently had a very interesting one, Topical, because it was a Guinness-related one, where there was a developer developing a Gindex, effectively. Oh, I saw that. They were calling all the pubs in Ireland to check on the price of a pint. You could ask that or report information. The Gindex is built with 11 Labs technology. It was built with 11 Labs, too. So people could actually do both sides, could proactively reach out, reactively reach out, all was captured through voice.

35:05Mati Staniszewski:And then kind of 3 ,000 different entities could report their prices and get that across. Have you, by the way, hooked up your OpenClaw to 11 Labs? Is the OpenClaw-11 Labs combo something that a lot of people at 11 are doing? So, as you know, the OpenClaw will kind of look for the most popular tools frequently where it tries to hook up. So 11 Labs is one of the recommended ones. It's the top option for voice. Can you tell me a bit about the business of voice models? I think people have an intuition around big LLMs, where there are these very expensive training runs. And yes, they kind of appreciate it quickly, but there's so much usage that all of the models trained to date have paid off their training runs and then some.

35:51And then there's this kind of ever larger CapEx going into, I mean, a lot of it is inference these days, but also training. And so people have some intuitions from the LLM world. I'm curious just how I should think about voice. For one, how expensive is training the voice models? Is the expense in the researchers? Is the expense in the training runs? And I mean, the economics is presumably kind of simple, where it's just per usage. But yeah, just talk us through the business.

36:21Mati Staniszewski:Yeah, definitely cheaper than the LLM and image video models. It's a really smaller models. Okay, so the models are smaller. Smaller. What's a parameter count for a leading-edge voice model? A few billion to low-tenths of billion parameter models. And for context, I think the, I mean, kind of like, you know, CPUs moved away eventually from gigahertz as, like, the metric as they moved to more cores. I think we've mostly moved away from just raw parameter count, but I think the leading-edge LMs are in the hundreds of billions of parameters. I think the leading ones, yes, but, of course, you know, you have the variations that you will use at lower scale.

36:58Mati Staniszewski:So CapEx is still pretty high. We've, of course, raised recently a half a billion at 11 billion valuation. Makes sense. Makes sense. To continue being able to build the best models in the world. Researchers, you know, of course, you want the best people in the world. I think we have those people working in audio and my co-founder who is leading that work. So that's definitely a big piece of like, not financially, but even like how you keep the ambitious deployment. So you kind of continue building leading models, helps you attract more talent and building that. And then on the how we serve it.

37:37Mati Staniszewski:So of course, inference is correlated with how the models are used. And for us, like we've seen incredible, incredible growth across the work. Mostly this is charged per, if it's input text or text to speech, it's usually per text token. If it's voice agent or transcription, then it's per minute. And we see that kind of being the bigger part. But usually, like broadly, it's per token basis. And of course, as we work with businesses, it's like an annual agreement. The bigger the spend, the bigger the comments, the bigger the discount to get that gross. The way we usually do is like when we have a new model, we try to give it at cost to a lot of the customers so they can experience the best.

38:17Mati Staniszewski:It's still usually like not as reliable. And the newest thing is often the most expensive, whereas you make the newest thing the most economically attractive one? We try to make it attractive so the customers are, like, you know, like, it's more expensive for us than any previous generation. We don't, like, the quality is higher, so we try to keep the prices still competitive to that. I see. You subsidize it, but it's inherently more expensive. Exactly. It's a bigger model. Exactly. Exactly. And over time, we might do some tricks to optimize it, but, like, we want the customers to, like, experience.

38:44Mati Staniszewski:Because of research, the big thing that we've seen is the reliability of the model in the early days might not be there. And then two, people don't even know what's possible with that model. So you kind of want the widest set of distribution so people can show the world what's possible. So you can have it, of course, as a distribution mechanism, learn yourself what to improve, what to change, and then get it out there. Are the voice models just getting bigger and bigger? Like, will we have voice models in the hundreds of billions of parameters? Or have we found, like, it seems like for certain types of model architecture, there's like an upper limit on like the natural size.

39:24Have we found that upper limit for voice models?

39:26Mati Staniszewski:It feels like for specific use cases, like say audiobook narration, you probably found that size. You probably don't need to stretch it too much bigger to make the quality as much higher. But for certain use cases, that will probably grow. The thing that's, you know, like I hesitated on the question is, in a cascaded approach, you probably will not see like dramatic size changes. You inherently want the models to be quick and reliable. You want to orchestrate them in a smart way. In a fused approach, probably that will get into like tens, hundreds, billion-parameter models because you kind of combine, of course, the LM side and the voice side.

40:03Mati Staniszewski:So that will get bigger. But on the just voice, I think it will keep being small. Okay. But there are certain domains where we'll see bigger models. That's so interesting. It is amazing how, it does seem fun from a research point of view, how there are still these various unsolved aspects and how you guys are just making technical breakthroughs and then releasing them down the product pipeline. That's like a really fun stage of a company's lifecycle. For sure. It's fun because it feels like we can do innovations on both sides. There's so much on research side, so much on product side. And then ultimately the biggest path is how we deploy it to the customers.

40:41Mati Staniszewski:Where SMB will have a very different dynamic than the enterprise. It's not vendor-sass relationship where you just give the product out there for the biggest companies out there. But you are more of a partner in their AI transformation part. So you want the resources to work alongside them to work on the frequently very new use cases that were impossible to help create and bring those voice agents to production. So that's like a big, big, big shift. But the biggest focus is how we bring the conversational agents out there to the businesses around the world. So when you say bring conversational agents is the biggest priority, is this for customer service type use cases?

41:21What are the most popular use cases for conversational agents?

41:25Mati Staniszewski:Yeah, we want to be a partner for full interactions between businesses and their customers or their audience. I'm saying their audience because that will apply in support. Support is the easiest one because that's where it's most ready. And that's maybe the big difference to how we see ourselves to some of the other companies in the space is this can also apply to sales. You can have the proactive side of reaching back. You can have AISDR versions of that. And then you can have all the way to the marketing use cases where we are your partner for working on even outside of like the conversational agent space of how you create a great marketing campaign.

42:01And so how does this break down between, you know, we had Dev Trainer from Intercom on here and they have Finn, their agent, and it's a thing in the website that you can go talk to. And he described a very similar phenomenon that you described, which is you start maybe thinking, oh, this will help me answer customer support queries. But it becomes like a generic UI for the website, where it's a box you can type in to go do things and understand things. And so why wouldn't you read the docs and design your integration that way? You know, whatever. And so will I have like one for text and then one for voice?

42:38Will you guys do text too? Will just, how does that? Because it seems like this is also succeeding at the text level with Finn and Sierra and all these things.

42:48Mati Staniszewski:The places where we know we will be able to provide the biggest value is like where ultimately today you will have either a big portion or most of the interactions coming through voice. So if that kind of intersection is there, that's where we can provide higher value. And of course, like if you need a text chatbot there, that's like if you fix the voice agent, you'll have fixed text piece or like inherently as well. But the place where we do optimize today is going to be like, how do you select the right voice for the right customer interaction? How you pull that in the pretty complex case of what you mentioned earlier of like how you orchestrate that to pause or look for something deeper into the docs, how it can be extension of entirety of the business.

43:31Mati Staniszewski:So not only in support, but across entire of the user journey. But the bottom line is like, We want to be able to provide you across entirety of the interactions. Voice is usually a big part of those interactions. And yes, we need to solve the integrations. We need to solve the knowledge. We need to solve text as part of that. But we wouldn't, for example, go into what I think will happen in a lot of those cases, very deeply into a reasoning version of those use cases where you maybe need to multitouch. Yeah, yeah, yeah. And a lot of complex actions. A lot of financial analysis. of like, that would be not something we optimize for.

44:08Can we talk about your revenue ramp, where you're just one of the fastest growing startups period of the past few years. What's your most recently announced revenue figure?

44:16Mati Staniszewski:Most recently announced was end of 2025. Whatever number you want to give us. So most recently announced was 350 at the end of 2025. Yeah. But the best proof of the technology working, so recently we announced our work with Deutsche Telekom and T-Mobile, with Revolut, with Klarna, with Meta, with IBM, a wide set of use cases. And this quarter was kind of one of the best for enterprise growth, where we had the first quarter hit 100 million in an additional ARR growth, which is crazy. In net new ARR. In net new ARR. Okay, so if you're saying this quarter was 100 million in net new ARR and 350 million at the end of the year.

44:52I'm no mathematician, but it's up in the 450 million range. And that's versus this time last year. That's a several fold increase. Just what's working? Like from the outside, I would assume that there is really strong cohort growth within accounts. And then you seem to have self-serve and enterprise businesses that both contribute a lot. I don't know how big self-serve is, but as a user, I like to be able to fiddle with 11 labs and not have to go talk to sales. But maybe you can just talk about what works to reach 450 million plus of ARR so quickly.

45:28Mati Staniszewski:Yeah, so exactly. So we are over 50 % is now sales led on enterprise. And I think largely that technology that powers a lot of their agentic interactions just became reliable at the same time as high quality over the last year, year and a half. So frequently you know this extremely well. You will start the account and then of course it continues expanding. And we see there's definitely land and expand motion across 11 apps. What does that expand look like? Is it like new departments? Is it just the usage starts taking off when a customer expands? Both. But usually the first part, too, it's like we try to make it very easy for our customers.

46:11Mati Staniszewski:Maybe that kind of against ourselves where we give the technology a pretty attractive economics because we so much believe in the technology providing value. So you can actually try it and test it. And then within that one apartment. And you think you'll make it up in usage, basically. Exactly. That usage, the kind of content continues increasing because you know it's providing value. And then it's so much easier to make that a choice. Yes. And then, of course, cross-department pollination is there too. And it's like, you know, our work with digital comes sort of marketing side. So we did Magenta work and Pogta's generation.

46:44Mati Staniszewski:Then it kind of expanded to customer support. And then it expanded to us working in agent across the entirety of the network so people can call in and have the agent. So you could see those step changes across. But we are now 470 people as a company. So we keep on growing. But some of the things that stay consistent is small teams. So we have less than 10 people teams for each of the product or research initiatives. Or even as you think about sharding some of our go-to-market strategy, those will be smaller teams understanding the industry in depth, understanding the market in depth, and going independently and going quickly.

47:23Mati Staniszewski:So that definitely contributed largely to that. Two, especially on the biggest enterprises, what we found works is, and it's like we have the full spectrum, self-serve, PLG motion that helps drive distribution, drive kind of awareness of 11 labs. And on the completely other spectrum, we have the high-touch forward deployed engineering working side by side with the customers to customize the entirety of their work together. Why did you guys do self-serve? Because I presume you have a lot of competitors where they have tech and it's behind a contact sales form and you have to go talk to an SDR and then talk to an AE, blah, blah, blah, blah.

48:01And you guys just offer the tech available on the site. And I'm a huge believer in this. I mean, a huge part of Stripe's growth has been driven by the fact that we just made Stripe available to anyone and built a lot of product around that adoption pattern. But so many companies seem to skip it.

48:17Mati Staniszewski:So I'm curious how you guys came from. So many reasons. So many reasons. I think the quick ones that come to mind is feedback loop. You can have immediate understanding of how good your technology is. Two, which is an extension of that. We stand behind our tech. We believe it is the best in the world for models, for voice agents, for deployment. So we want people to experience that. And I think you do that the same in Stripe, where the best version of the technology is available to everyone, which is so attractive to actually try it out. We always try to make everything we build for the highest end use cases, bring it back to the ecosystem free.

48:53Mati Staniszewski:Frequently, the newest of the use cases, for enterprise, you will need reliability, you need compliance, you need the scale which we deliver. So frequently, as you develop new technology, it might not be ready for a lot of those parameters, but it's definitely ready for developers and SMBs. and we love what they are doing because they are showing us the future and effectively helping us find a trajectory of where 11 apps should go. I'm totally convinced, I'm just always amazed that more companies don't pursue it or it feels like they're really shooting themselves in the foot by not, like did you guys self-serve on Stripe or did you?

49:27We self-serve on Stripe. Yeah, for example, you know, 11 is a huge company and yet you started on Stripe on a self-serve business.

49:33Mati Staniszewski:You kind of like initially, and it's like, you know, we were two of us at the beginning. You try to see what's working in the industry, but you try to think from first principles. So you want to try it out. You want to understand how it works. So the more friction elements before you're trying it out, the less you trust whether it's available, whether it will be additional payment that's hidden behind some of those steps. So you don't want to go through them. So it's so much. Speaking of Stripe, do you have any Stripe feedback for us? Anything you want us to fix? My most common feedback until recently is like, why don't you give us pay-as-you-go user-based billing type version?

50:07Mati Staniszewski:but one of our finance lead, Maciek, I know I was speaking with your team and that was the day before he was great, he was like thinking about it for a long time, he's great He said like, you guys should buy Metronome You should buy Metronome, and then the next day Metronome acquisition was announced so now you have it, so that was my most common feedback, and we'll be launching that's a good announcement for this for this podcast, we'll be launching user-based billing to everyone. I'm shocked you Oh, as in previously you had it on an enterprise basis, but everything on the self-serve basis was like plans?

50:42So we had the subscriptions, yeah.

50:44Mati Staniszewski:Subscription plans, you can go over them, but now we are launching a full pay-as-you-go experience. So you can just try out voice engine, which is effectively this whole orchestration loop all the way through to any of the models directly. Going back to self-serve, I think a new thing in AI is that all self-serve products should have pay-as-you-go as an option. Maybe you want to have like a subscription with some unlimited tiers, but I don't know if you had the experience of like you're using Cloud and like you're typing away your queries and eventually you hit some rate limit and it's like, sorry, you've hit your usage limit and you want to be able to do the thing that you can do a Cloud code, which is just pay per API.

51:20It's like, I'll pay for it. And it's kind of very funny as a consumer to not have the option to pay more, to use the product more. And so, yeah, I think every AI product will need, you know, they probably want to have some all you can need, most of what you can need is subscription with limits. and then the ability to pay for overages. So it sounds like that's what you're doing.

51:37Mati Staniszewski:Yeah, exactly. That's what we're doing. The only thing I wanted to ask you about is, I feel like all CEOs of larger companies today are trying to figure out how do all these AI advancements change the nature of the organization, and how do you redesign your organization a bit around all this new intelligence? And so that could be about what the scaling factor is of like the number of people you need to do the work. But it also should be like, do you need more senior people because they're better able to direct the AIs and the AIs are maybe you can do the work of, well, previously would have been junior people.

52:15Do you need more junior people because they're going to be more AI native in how they work? Do you want smaller teams? Do you want bigger teams? How do you actually go do the process engineering of your finance team should be using Cloud extensively? But like finance teams do not historically, you know, have a lot of home-built software. And so there's all these questions that are floating around. And you have very rapidly built a much more AI-native company. And so I'm curious what lessons we should all be learning from Eleven Labs as a large business recently built. And so without the baggage of decades of how we've always done it.

52:52Mati Staniszewski:Yeah. Yeah, we started in 2022, which is a year when the two topics of the day were crypto and metaverse. So just before. And then, of course, AI flow started. Exactly, exactly. But we could like had the privilege of like kind of scaling through the world when it was all happening. For us, Woodworks, and we really believe in that being the big part of the future. The first is small teams, like keeping the teams small and super flat. So like, can you have, both me and my co-founder will have over 15 direct reports each that we'll work with. And most of those people will have that same scale of direct reports.

53:28Okay, so your span of control is way larger than the traditional company. Normal would be eight. Exactly. You have double that. And obviously that's an exponential.

53:34Mati Staniszewski:Exactly. And of course, there are some teams which in the short term might not do that. But ultimately, that's where we think it's going to be headed. It's like roughly 10 team size within each of those work items. And startups. Pretty close. No offense, but startups often have pretty wacko management ideas. There was a funny tweet, Laura Grant made the confidence of an early stage startup founder blogging about their management theories. But you think this is not a startup effect? this is an AI effect where basically no it's definitely a little bit of startup effect I think it's like it's hindsight hindsight benefit I'm cancelling our stripe changes yeah no no it's like I need to pre-end that I'm not kind of you know it's the hindsight of this may be working we'll see in the next five to ten years in their places much flatter org much flatter org so it works for us it might not work for all the companies and there are some parts where like go to market we still are trying to figure out what's the best way but smaller teams flatter org and i think they're two paradigms but like generally people being more technical or if not technical even in non-technical teams having a technical resource so you know we will have a person in ops or in talent that will we have effectively a tech lead for that team yes that helps them automate a lot of that that that work and helps up level the rest of the team too yes so there are kind of two parts that are helping okay so talking through this in talent or something like that.

54:55Is it that you are building your own software where other companies might have bought software like a Workday or a Greenhouse or something? Is it that they are using the existing software you have better? Is the process that would be spreadsheets in a traditional company are built with software? How do you kind of use the software in these sorts of organizations?

55:16Mati Staniszewski:Yeah, like sometimes, but we still use a lot of like the traditional vendors. Like one pattern is of course, is LM-ifying everything, like making the data explorable for you to be able to interact with it, of like who's in the pipeline, what worked, who does the best references, like all of that works, so you can double down on that. But two, it's frequently things that you manually do that a lot of the current, like there is a gap between where the agents are today versus what you could do if you have the technical skill set. And a good example is like, how do you scrape all the right profiles to be able to reach out to the right candidates?

55:52Mati Staniszewski:so you like analyze whether it's you know how much i should want to say but but uh the like try to detect specific things that we know worked so you'll bring that across to the to the to the to the people on go to market side like there's just so many things you can do with with with additional amplifiers you know it goes from understanding what case studies are relevant and creating a good pre-read for you before you go to the meeting, through creating the AISDR experience that we spoke about, to creating an entire deck experience. So you have a pre-populated deck with the right numbers that is customized to that customer, which you want still the person to go through and develop, but ultimately is in there.

56:35Mati Staniszewski:So there's plenty of those additional things that you know will amplify the work of the people around, potentially replace some of those easier tasks that are done. And then there's like, you know, we wanted for people to explore the culture at 11labs. So we created a voice agent that people can speak with and see what's the culture, but also get prepped for the interviews. I think across many of those teams, like additional benefit of what they can do. Interesting piece. So, of course, in Ukraine, with ongoing work, they need to rethink a lot of how their development, their systems, their support works for the citizens across the country.

57:11Mati Staniszewski:And people are in the war zone. They don't have the same access to the information. They cannot rely on the same phone lines. They cannot rely on the same physical services around the country. So they've developed effectively a central... So your employees in the Ukraine? We had a few, but they reached out because they were developing their central map called Dia. They developed it over the years, but now with war, they were doubling down on how this can be a way of supporting the citizens. And of course, there's an easy part of how you create a first-agenda government where you have help with the benefits and what's happening on the front line or education.

57:47Mati Staniszewski:So that's delivered to everyone. Or healthcare, so you can book your checkup or appointment. So how you create all of that. And of course, we traveled to Kiev. We worked with them on bringing that and making that available for voice so everybody can access it. But the thing we've learned while being there was that model of what we speak about where you have technical resources in each of the teams. They actually have the same in every of the ministries. So every ministry had technical resources working on creating that agentic version of their work. And then it was like a central digital transformation team that would like assemble this all together to deliver that for the central citizen support, which I thought was brilliant.

58:25Mati Staniszewski:That's very tacked forward by Ukraine. So take forward, like the most advanced set of work we've seen. So we got a little bit validated, like, okay, maybe technical resources in each of the teams is a good idea. And that works heavily for us. And, you know, you mentioned some of the other parts, like do you hire the senior or younger? Like main thing we try to filter for, of course, the culture piece is so important. You can scale people, but scaling culture is much harder. So, like, you want to optimize for that being right. In our case, it's first principles, taking ownership, striving for excellence, but staying humble.

58:58Mati Staniszewski:And the main thing that's kind of in that ownership part that I think works well for the AI world is agency. Like if you have that agency to explore, regardless of where you are in the experience cycle, it's going to be a tremendous amplifier to your work. My biggest takeaway from all this has been that around agency, where I feel like high agency people are the winners of the advances in AI and within organizations, low agency people will lose out. Yeah, completely agree. Probably the most proud thing that Piotr and I are is as we scale the 11 labs, the people that are at 11 labs is being like just the culture and seeing the expansion of the culture, where culture builds the company now rather than any single person or any single product builds the company.

59:45Mati Staniszewski:That was probably the biggest validation and happiness. and there is kind of the other angle of that where I think people are like striving to be incredible in their craft and their work but at the same time have fun in a lot of their work and that kind of combination of agency and just enjoying what you do is probably the best thing we've been able to do today at 11 Labs. Well, it sounds like a really fun stage like we were saying. Interesting research breakthroughs, really fast growing business so I'm sure you're enjoying it. Andy, thank you. John, thank you so much. A red light car with a Bikram Jake In the riddle Ken дорage11 Canondin 11 If they aren't 28 Canondin subtitles

1:00:38subtitles Thank you.

From the publisher

Description

Mati Staniszewski is the co-founder of ElevenLabs, the research company making audio accessible across languages and voices. He sits down with John to discuss the "voice Turing Test" and why AI has conquered text but still struggles with conversational speech. They discuss the future of human-computer interaction, including why we still can't get our phones to read a PDF properly and the massive potential for voice agents in everything from farming to healthcare. Mati also opens up about ElevenLabs’ rapid ascent to an $11 billion valuation and gives a behind-the-scenes look at how Ukraine is using their tech for digital government services.


Timestamps

(00:00:27) How audio models work

(00:08:52) ElevenLabs business model

(00:17:50) The conversational Turing Test

(00:21:01) Link by Stripe

(00:26:02) Cascaded vs speech-to-speech

(00:31:53) Universal translation

(00:51:41) Designing an AI-native org

More from Cheeky Pint

All 35 episodes
The world of voice AI, with Mati Staniszewski of ElevenLabsCheeky Pint · 1 h
Listen in VO