In short
The Newcomer Podcast: Episode Summary
Episode Title
Inside Anthropic’s Plan to Fix AI’s BIGGEST Problem
Overview This episode dives into the challenges of Artificial Intelligence (AI), specifically focusing on the concept of "AI slop" and how Anthropic aims to address these issues. The discussion takes place during the Cerebral Valley AI Summit and features insights from leaders in the AI industry, including Anthropic's Chief Product Officer, Mike Krieger, and Eleven Labs' CEO, Mati Stunashevsky.
---
Key Discussions
The Current State of AI
- AI Slop: The term "AI slop" is used to describe the low-quality outputs generated by AI models, which lack depth, critical thinking, and engagement.
- Truth-Seeking Models: Mike Krieger emphasizes the need for AI models to prioritize accuracy and truthfulness rather than just engagement metrics.
- Engagement vs. Truth: There is ongoing debate on whether AI should optimize for engaging conversation or factual correctness.
Anthropic’s Approach
- Combatting AI Slop: Krieger discusses strategies to produce high-quality, thoughtful outputs from AI, utilizing feedback from users to continuously improve the models.
- Future Directions: Anthropic is looking into growth opportunities in specialized sectors like life sciences and enterprise AI solutions.
Voice AI Revolution
- Mati Stunashevsky’s Insights: As the CEO of Eleven Labs, he shares how voice AI is transforming customer interactions across various platforms, enhancing user experiences in customer support and gaming.
- Authenticity and Safety: Eleven Labs emphasizes the importance of user safety and authenticity in voice AI applications, particularly with celebrity voice cloning.
---
Key Concepts
AI Slop and Its Implications
- Definition: Slop refers to AI-generated content that appears low-effort or lacking critical engagement.
- Combat Strategies: Focus on training models to produce concise, valuable outputs rather than verbose, engagement-driven results.
The Role of Voice AI
- Use Cases: Voice AI is poised to become a primary interface for technology, with applications ranging from personal assistants to interactive gaming.
- Future Predictions: Stunashevsky predicts exponential growth in voice AI, particularly in customer support, where it could dominate within the next 18 months.
The Future of AI Models
- Commoditization of Models: Both Krieger and Stunashevsky note that as AI models become commoditized, the focus will shift to the product and application layers.
- Vertical-Specific Solutions: Companies like Eleven Labs work closely with various industries to tailor AI models for specific use cases, ensuring relevancy and effectiveness.
---
Notable Quotes
- Mike Krieger on AI Truth-Seeking: "We compete on how accurate our responses are rather than engagement."
- Mati Stunashevsky on Voice AI Potential: "Voice will be one of the key interfaces for interacting with technology across various domains."
---
Conclusion This episode of The Newcomer Podcast highlights the critical discussions around AI's evolution, the risks of misinformation, and the transformative potential of voice technologies. Through candid conversations with leading figures in the AI landscape, listeners gain insights into the future of AI and its applications across industries.
Call to Action
- For those keen on understanding more about AI’s trajectory and the latest developments, subscribing to The Newcomer Podcast and following the discussions on platforms like Substack is recommended.
---
By outlining the key discussions and concepts, this summary provides an insightful overview of the episode, ensuring clarity on the pressing issues and innovations in the AI domain.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00A hundred million dollar funding announcements, the shadow of the looming AI bubble and debates surrounding agentic AI. This week's Cerebral Valley AI Summit just concluded, which, for the uninitiated, is Newcomers' flagship event hosted by myself, Aaron Newcomer, and Volley co-founders Max Child and James Willsterman that brings together the top founders and investors and features cutting-edge discussions with the biggest names in AI. In this episode of the podcast, we're revisiting two of those important discussions, starting with the conversation I had with Anthropics Chief Product Officer, Mike Krieger.
0:32During the conversation, we discussed the evolution of AI product design with an emphasis on the need for truth-seeking models. Mike also shared his thoughts on what Anthropics can do to combat AI slop and how they plan to focus on growth opportunities and life sciences and specialized verticals. This is the Newcomer Podcast.
0:57All right, lean in everybody. Excited about this one. Mike, thanks so much for joining me. Great to be here. You know, co-founder of Instagram to chief product officer at Anthropic, I wanted to start with almost like a philosophical question that spans those two companies. You know, we talk about AI now, but obviously social media companies with feeds were using machine learning and systems to surface content. Then that era seems like it was about engagement. And it was like, okay, we're going to blame or the humans will be responsible for the truth value of what they have to say. And we're going to see what content people are interested in.
1:38And those humans can say what they want to say. you know as a journalist and someone who's interested in like the truth and chasing the truth one thing i've liked about models despite like all the like oh they hallucinate or whatever is the aspiration is like we are judged based on how accurate our responses are you compete on leaderboards that are saying how often are you getting things right you want to pass you know math olympiad type tests do you do you accept that framework and how much do you sort of in this role see Anthropik as this sort of like truth-seeking organization or will engagement seep back in?
2:14That's a really good question. Hi, everybody. Good to be here. I started with a heavy... Yeah, I think there's a bunch of different directions to take this. I'll try to be succinct. I think there are places where we're training the models as an industry to be good conversationalists. And sometimes that actually looks like continuing the conversation. I got to reach out from somebody that was like, hey, are you guys trying to optimize for engagement? Because Claude will often ask me a follow-up question like, well, did you want to talk about this? And the funny part is like not at all. And like time span is like, I can tell you, like not on any of the dashboards that I ever look at.
2:43It's just not a, like a main consideration, but just training cloud to like, keep the conversation going, being a good conversation. So we might actually need some interesting sort of counter metrics to what, you know, well, can you get the same conversation done or accomplished less of a conversation? I do think there are some really interesting sort of other phenomena happening in the industry. So one of the things that we try hard not to optimize for, like I just don't think it's the right incentive, but is these chatbot leaderboards, like Elemarine and all these places. They're useful yardsticks of how we're doing, but they're not the thing that you should optimize for.
3:19But we have found, as you look at what ends up doing well on there, it's verbosity. Being more long-winded can actually be praised by the raters on the - That's like taking an exam. I'm going to say, every note you remember, like unload it into it, yeah. And so it is interesting, like what we, I think the evals are important, but whenever we sort of like have these public yardsticks, it can tend towards the like yappier or more engagement-y sort of thing. So I think that's one vertical. The other thing that I think about though is, primarily what we're doing is building AI for businesses, right, and so we have like Cloud for Enterprise.
3:53And in some, like, it's not like we would do like super engagement-baity type things, But a year later after a contract gets signed, people are going to look and say, like, did people use our AI or not? And part of that is just, did it solve the problem? Was it truth seeking? Did it do the right things? But then there's also, was it a product I liked using? And so I think we're all navigating this question of, how do we train the models? How do we design the products around the models? And how do they get delivered in a way that is maximally useful rather than falling into this engagement for engagements?
4:25What do you make of this word slop? I feel like that, if you had to sum up the criticism of AI from, I don't know, the skeptics, that's a word that often comes to mind. What do you make of slop and what's to be learned from that accusation? Yeah, I've told our product team, one of our goals when we do things like build PowerPoint decks and Excel files and Word documents is to be the anti-slop. And I think the way I think about it, it's hard to put that on an e-mail, right? like 80 % slop or 7 % slop. I think it looks like a couple of things. Like one is, does it look super low effort? Like, did it just like, does it not reflect any sort of critical thinking that the human did alongside the AI?
5:08And you can often tell, you're like, was this thought through? Like, this wasn't really edited. Like, is this 2000 words when 200 would have done if it had actually been edited down? So I think there's that strong component. And I went to a talk by Ted Chiang, who's a science fiction writer. He wrote the short story that became Arrival. He's one of my favorite writers. And we were having this conversation about AI and can AI be creative? And you made the point of creativity is the product of a lot of decisions, right? And so that's true for novels. Like, what does slop look like for a novel? It's like, when you're like, well, like, this just seems like the kind of base level output.
5:39If you're just like, tell me a story. Right. But I think that it applies just to nonfiction as well, right? Like, if the document that you put, like, even product reviews, if somebody comes in, and this happens rarely, thankfully, in Anthropic, but with a product requirements doc that looks like it was just the first output from just like, write me a PRD for this feature. Like, this is just slop, right? So I think that's the content piece. Then there is the design and the quality piece, which maybe is like a version of what that other one is, but it's more visual, right? Is time the solution? Like, if I think about writing as a human, I go back and revision is often where you find great writing.
6:19Is that the case with AI where it's like, okay, if you have more time, the answers will be more tight? Or how do you see the relationship between time and the quality of the answer? I think it's levels of engagement and sort of how many iterations that you've gone through and how many... And actually, you could imagine a... Here's a very well-constructed prompt where I have already pre-made a bunch of the decisions. I had a... I mean, you can like see what Claude is thinking. You can expand the load. like I rarely do it because I just did students thinking I mostly want the answer um but I expanded it yesterday and I had written like a pretty long prompt and it's thinking it was like you know Mike has already like done a lot of thinking here so I'm not going to ask him any follow-up questions I'm just going to give him like yes thank you that was actually that was my intention here but um I do think it is that like how much of your independent thought even that's where actually now tying back to the first question sometimes the model should ask a question like hey this is a really open wide option space like can we like start narrowing down and i'm going to engage with you on it now i think we need a lot better uis for that like just here's a question that you then have to go and type it right it feels kind of annoying but in navigating that option space you should be able to hopefully come up with something that is like complemented by ai and accelerated by it but still has your thinking at the core just to get to that i mean you're a product guy this is the uh is ai fundamentally the chatbot era like do you think text with the machine that is the main way we're going to use ai use anthropic in five years or that's a bridge to that's how we figured it out in the beginning and now we need to build products so two uh kind of ways to tackle that one is like like the classic meme of like you know you're like you're naive view and then you're like like super like galaxy brain view and then back to the original view is like how I have felt about this exact question.
8:08So when I joined Anthropic, I was like, if we are still talking to AI with chat boxes a year from now, like, oh, I've failed in my job. Oh, really? Yeah. I was like, I was very adamant that there was like something wrong to the like kind of dominant UI paradigm that we had settled on. And it felt like, you know, we exposed the lack of creativity. And then we did a bunch of explorations around like, how do we create more structure around it? How do we make it friendlier to people? And I've never use these models, all of these different pieces. And I realized a lot of those explorations end up constraining how the model operates or what it does in a way that made it so that when the next model came out and was much smarter and maybe didn't need as much hand-holding, we actually were holding it back.
8:46And so the chat box might look different. Cloud code is a chat box but in a terminal. But I've really come to believe that now what happens behind the chat can really expand and now like cloud is writing code and running code for you and like calling MCPs. And there's a lot more that's happening underneath. And the sort of metaphor might not be text message. It might be more like a Slack, but you don't expect a message back immediately. But I think specifying the kind of request in like mostly text actually makes sense. And then what can happen is like underneath it. Right. I mean, you'd never ask this question about a book.
9:21You're like, oh, it's just text. Yeah, it's like, obviously language is great. But so what you're settling on, you do think most of what you're delivering is this sort of chatbot experience. I think that, or a conversational experience that then has more and more work that happens beneath the hood. The one sort of nuance that we've kind of come to believe there too, it's like, that's a great paradigm for kicking off work or doing research or even like condensing a bunch of ideas into like a sort of first draft presentation. It's a bad UI for, hey, can you move the text on slide three up by three.
9:54Because if you ever had this argument with any of these columns, just do it. And it's like, it doesn't ring wrong. And you're like, no, no, just right there. This is where I stumble with vibe coding. And I downloaded it. At some point, it's like I hit some wall where it's like, I need to move this thing. And then it's like, I'm lost. And that's where I think tools with richer user interfaces still really matter. And some of those might be kind of coded just in time and materialized in front of your very eyes to edit it. And some of them are tools that have just been honed over a long time. And it's why we built Cloud for Excel, which is, hey, Cloud is a great first draft of your discounted cash flow model.
10:29But if you want to go tweak it, let's just let you open it in the tool where it's actually going to be most useful and then let you continue maybe pairing with Cloud there. Returning to sort of my core philosophical question, like the sycophancy question, like what is your view on that and how much to enable sort of everybody likes to be flattered. Like it's a reality of human beings versus an effort to be direct. And how do you think about those trade-offs? Yeah, I think there's like a wide gulf between like true empathy and then like sycophancy. And it's interesting that it materializes not just in, hey, I'm having a conversation with Claude about like some coaching or personal goal that I have, but it also does encode as well.
11:09When we were testing Sonnet 4.5, one of the things that people got most excited about was when Claude was like, this idea is bad. Not that you should feel bad about it, but this idea is not a good direction. I can go and implement it if you really want to, but I would suggest that we try this other thing instead. So there is something like that pushback is not just valuable in a personal relationship with AI sense. It's actually like how you get good work out of the models. But for a long time, our models have been, I think, appropriately empathetic. like they're like if you're going through a hard time like i was dealing with the death of a pet and i talked to claude a lot about these different things and it always started some sort of like hey that sounds hard like sorry to hear but then i'm going to give you like a factual answer i'm going to go research these pieces but still with the place of empathy um as well and so i think when we look at it internally and we're just evaluating it ourselves it's again not that like empathy it's not even like the likability of the model it is do you like does it show up in the way that you'd want a good conversationalist to show up and then continue on its AI journey around what it is going to do with you as well.
12:14But I think it spans everything from that initial response all the way to how it evaluates an idea as well. Claude, especially previous versions, were known for being like, you're absolutely right when you correct it. And my wife got her first, you're completely wrong. And she was like, yes, this is great. And I think we should have more of that like kind of less san francisco yeah a little more direct new york um anthropic has obviously had a ton of success with the enterprise with coding uh delivering um value through the api like is that the company like how much are you leaned into sort of serving other businesses versus you know we're gonna see you spin up some random consumer app in six months yeah i think obviously you have a strong consumer app but like yeah you know you know what i'm saying yeah I think I look at what, like when I think about our product surface, there's a few kind of criteria around like when we expand and what we decide to build.
13:12And one of them is, is there some feedback loop that we need that would be well suited to a first party product? Because even though we serve a lot of customers using the platform, it is also really valuable to have, for example, a cloud code where we have that iteration loop and people are giving us feedback all the time, whether it's in micro moments or even just, you know, writing in, you know, with some longer feedback. So there's like, is there some feedback loop either of the product shape or of the model that we can better do? So there's one. The second one is, is there something about the category that we think we have some unique perspective on, either because of like what we've built internally or what we're trying to do with the models?
13:50And like, then it's worth like building some product surface around there as well. And then the third one is kind of like we get from just a customer draw, especially as we expand into different verticals. So Cloud and Excel came from very much from talking to all these financial services companies and being like, Hey, I want you to just bring this closer to work that I'm doing. But I do think that like there's, I, we've been doing more of these even like time limited sort of like research previews or demos. And I'd love to do more of those, even on the consumer side as a way of sort of. With a standalone app.
14:21I mean, you know, at Meta, I mean, you guys came up with standalone apps, like how much do you want? what is it slingshot or where like various experiments versus you know we want to work out of the core app like what's the lesson from that experience i think it's i think there was a few so for us slingshot the right one is that slingshot like facebook built slingshot we built one called bolt that nobody remembers okay it's like very fun you would open it to like a camera so like at that time the the big criteria was well people have a very specific sort of expectation of what happens when you open instagram and it's not that it opens to the camera right uh and it was like our most interesting competitor was snap at the time and it was like well they opened a camera, which means that messaging is really fast and it can be built in a separate messenger.
14:59That was the whole thesis behind building like first bolt. And then there was like an Instagram direct separate app exploration. But I actually think there was a kernel of insight there that I think applies here, which is if the reason you're opening an app right now is to ask a question of AI, then like, I think we can extend cloud in different ways of doing that. But that isn't the be all end all like of what you might want to do. Maybe you're trying to get, you know, like a really specific type of interaction. Maybe there's something around like your health journey and like Cloud can be a good companion for that.
15:28So I think it's still asking the question of like, what is the purpose when you are like entering the app? Like what's the context that you're in? And then, you know, does it cloud the use case to have something else embedded in there? I mean, we've talked about this, you know, verticals you're interested in clearly, coding, financial services, you just touched on health. Is that help the consumer? Or, you know, we had another event. I talked to CEOs of Abridge and Open Evidence. I've actually been playing around with Open Evidence. That one's targeted at doctors. It's interesting to go through it.
15:59And it's very like, you know, clinical, like a doctor. Do you think you do something custom for me, the patient, to navigate what a doctor is doing? It's interesting. Like, there's already so much of what people are using cloud for today. Like, when we have this, like, if you've ever seen, like, our topic economic index, the way we generate these insights into how people are using Cloud is we basically have Cloud-run analysis in a privacy-preserving way, so we never look at the chats, but Cloud can do it in a way that's privacy-preserving. And I did that for where I asked the question of the healthcare piece, or are people using it?
16:32And there is a double-digit percentage of Cloud conversations are about people's health. And I hear all the time from people, the first thing I do when I get a new lab result is I put it into a Cloud project and I have this history there. So there's clearly a pull there, but it's so annoying. All our pregnancy information, we would just dump into models, tell us what you think. Tell us what you think. And if you get a lab result back, it's like, well, I got to go download it. So I'd love to see a privacy-aware solution for more of that. And you think that could be sort of a custom? I think, yeah, that could be a more sort of bespoke experience.
17:06But then maybe I'll say it's shared across both patient and doctor. So if you have like a, the reality, somebody's, that a great phrase, I was at a healthcare conference recently, it was like, it's almost inevitable that most doctor visits will now be second opinions because your first opinion almost inevitably is that you're going to ask, talk to like Claude or model about it. So let's embrace that and be like, not just like, oh, I heard from somebody that there's a thing. Great. Let's acknowledge that you probably asked, you know, one of the LLMs this question, like, what did you learn? And like, let me, let's like have a conversation about that overall.
17:39So there's that piece. And then I like, there's a lot that gets dropped today in the sort of multi-doctor patient journey, and nobody's often looking at the kind of holistic experience. And I think there's a real role for AI to play in sort of stitching those different pieces together and generating insight that might not come even among experts among these different disciplines who are, by the way, probably super busy, contended, like context switching all the time and not stepping back and saying like, all right, this is the full view of this person given everything that I can infer there. Obviously, we have a lot of startup founders here.
18:13They sort of want to know how to work with you. And it's such a balancing act. At once, they want to know, oh, you're not, what won't you do with it that is exactly like me, so I avoid that. But where will you be more capable so that I can benefit from any improvements you make without competing directly with Anthropic? That's such a complicated relationship. Like what advice would you give to people in terms of reading the tea leaves and saying, okay, if Anthropic's saying this, I'm safe to build here or not? Yeah, I was talking to a founder of a very large enterprise company. And I was asking him for advice on this question because they had to have to navigate this over years where they'll build some functionality themselves.
18:54They also have a rich partnership and marketplace ecosystem, which is what we have as well. Our first-party products are more scoped. There's a lot more in the platform. I think there's a few principles I try to operate on. One is transparency. So like when we launched Cloud Code before we ever launched, like I got on the phone with like all of our major coding customers, like here's why we're building. Here's what we hope to get out of it. Here's how if we do it right, it should actually be a rising tide that lets everybody using cloud encoding. So there's that transparency piece. The second part is like that transparency is like telegraphing a little bit where we're going in terms of what we think are interesting verticals.
19:30So we did our Cloud for Financial Services launch. We did Cloud for Life Sciences about a month ago. And part of the role of those launches is this isn't just a first-party product. It is a vertical or a kind of set of capabilities we want our models to get good at overall. So if you are a builder, this might be a good place to get on. And our definitely goal is not to own that whole space. It's to enable all these different companies to then go and build some different pieces. And then what's been more interesting on the go-to-market front is we're now starting to see, all right, I'm already buying a big commit of Anthropic tokens.
20:04Can I use some of those on another product that's Anthropic Power? So I think there are going to be other ways in which we can work with both startups and larger companies in helping deploy their solutions into the enterprise. Another core thing startups and everybody wants to know is how much smarter will the model get? What can you telegraph to us in terms of 2026? Do you think there are still major gains to be had just from the scale of compute and GPUs? are we waiting for you to pull another like rabbit out of the hat in terms of like reasoning models or some technique like that like what what can you say about what next year looks like in terms of the capabilities and sort of raw intelligence that anthropic will provide yeah it's an interesting like uh perspective i get from startups where sometimes i talk to them and they're like models are great like we're just gonna like we have a bunch of work to do on like the go-to-market or like the scaffolding or the skills around it um and i'm like that's a good answer.
21:00I guess you can keep going that way. And then there's other startups that are like, we have a super hard eval and you're at 40%. And we think at 60%, it's like - Right. And there's a lot of VCU wisdom. It's like, build something so you're ready when the next model comes, which is - Yeah, which is that category. I feel like it's something I've said on stage. It is a real thing. And there's probably some midpoint in there. But I'll tell you that whenever we have a new model that's baking, and we have even an early snapshot on it, I have my list of companies that have, in the past been at that like yes we're pushing your model as hard as possible so that those gains actually get shown and like i guess like you want to be one of those startups or even like forget startups but really any company because i think the labs will want to sort of you're saying if you're one of those companies you're doing well enough you start to say okay we're going to be able to get you that last 10 yeah and we like we'll want to go you know in some cases actually go hill climb on that eval but just in general be like okay this is a demonstration of how well models do at like defensive cybersecurity which i think isn't here i'm really interested in and so if that's the case like let like the companies that are pushing us the hardest they're also the ones that we call then because we know that they're actually going to be doing there they'll be able to show a difference and like even like the peek behind the curtain whenever we launch a new model it's like just smarter is like just not a very effective marketing pitch right so the more we can say right and here's like a particular customer that demonstrated this um really well But back to your original question, I think there's still a lot of juice left in scaling up models, training them to do things, and then also layering on the right skills on top.
22:32So that's, I think, of tool use. Tool use is a great one. And then, again, on who's pushing us the hardest, it's the companies that say, hey, I'm trying to give the models 50 tools, 100 tools. All of these models, at some level, just start getting confused if there's too many tools. Can we make that better? So that's the kind of edge pushing that we need. And in reasoning models, do you think there's more progress from that or any other techniques where you think, okay, that's going to be a reason we improve next? Yeah, I mean, even within reasoning, it's been interesting to see, figure out what the...
23:05There is some additional parameter that people care about, which is, yes, you got to the answer, but were you able to get to it quickly and in an efficient way? So I think there's that kind of parameter to poke at. Then there's reasoning in the middle of responses as well, which is something that Claude can now do. And you watch it, if it's doing a lot of web searches, it'll sometimes reflect halfway through and be like, that was a good answer to that first question. Let me go and figure out the answer to the next one as well. So you want that back and forth of internal monologue, tool use, user response, and all of those different pieces.
23:38In my conversation with Max and James earlier, I said, nobody's talking about AGI anymore. I feel like at the first Cerebral Valley, there's a sort of obsession of like, oh, are we going to reach artificial general intelligence? We've sort of chilled out a little bit, partially because, you know, it's taking time. What is your view? What's the view within the company? How much this is still like a race to AGI? And like, how are you feeling about like timelines? I think it still is this sort of look at what are the hardest things that you can, maybe I'll break down to two pieces. Like for a given like task or problem, like how independently autonomous and sort of successfully can those models operate and on what time horizon, right?
24:20And whether that's like hours of coding or whether that's, you know, go off and do research tasks or whether it's do really complex financial analysis or whether it's like optimization problems or all like that feels like we still have a lot to go. And I don't know, I guess at some point you can call something super, you know, human in levels. It probably it already is in a lot of those different areas. So there's that piece. And then there's this other area which I think about a lot which is how do the models manifest in a way that actually learns the like call them soft skills or like skills around the fact that they're like very very good at like writing code or acting agentically for example and like I think that's the other piece where that'll feel like maybe the next moment where it's like oh there's it feels like there's been some departure here where you know it understands what's like information it should reveal to somebody else versus not.
25:09It understands the social dynamics of the company and power and all these different things, which are harder to train for, I think. Right. I mean, main shortcoming of the models to me is often when you ask a question and it doesn't say, I'm not really sophisticated about this or like what's stopping the models from saying, oh, I don't have a great answer in this case. That's often the most intelligent people disclose when they don't know something. Why can't the models do that? Are you working in that area? Yeah, I think that's an important piece, which is can it express uncertainty? We'll look at it like often the consequence of not telling you that it doesn't know is that it'll go on, confabulate something, and then that feels wrong.
25:48So we look really carefully at hallucination rates. It's nothing to drive down. But I think it is something that we can better train into the models around. What is the uncertainty that you have? Or do I need to go phone a friend or do a web search and then do this pace? But then tuning that is really important. We had an internal version that did way too many web searches. And you'd be like, why is the sky blue? Which is a question my daughter had. And it was like, I'm going to search the web for that. I'm like, Claude, you have an answer to that. You don't need to search the web for that. So tuning that is actually nuanced.
26:16Or you don't just want a thing that just be like, cool, let me Google that for you. Right. And obviously, if the model was just resulting every time, I don't know. I'm just a while. That would be disappointing. There is nuance there as well. But I think that nuance of uncertainty matters. And then also like the model learning from your interactions, not just in terms of like, I remember that Eric has these properties, but also, hey, I've learned something about how we work together that I think is still another unsolved problem for these models. I mean, if you were to tell people to run towards this space next year, like just like a couple of areas, I know we've talked around that, but like, where do you think people should be building or positioning themselves?
Read the full transcript
26:54I think, I mean, I get very interested in the life sciences overall. And like, that's both like obviously like large industry but also like the incredible potential for human benefit as well and like when you think about all of the things that happen from uh ideation even like fundraising upstream of that to discovery the back office the testing the trials like the model like there's like a whole complement of things that's one area that i i get really really excited about and then there's still i think um you know like some been good some good conversations and like are agents It's real. Like, what, you know, and even some of the conversations today I've touched upon it, there's still a lot of value in like that anti-slop.
27:31Like, not just making it work, but making it work so well that you rely on it and you want, it's your first protocol because you genuinely believe it's going to save you work. Mike, thank you so much. This has been great. Thanks for having me. For founders and developers building modern data-driven applications, MongoDB's local event series is coming to San Francisco on January 15th. And it's designed to help you focus on innovation, not infrastructure. You'll learn about technologies, tools, and best practices that make it easy to build and scale modern applications without complexity. Plus, attendees will hear directly from experts and innovators who are using MongoDB to power the next wave of AI applications.
28:10MongoDB.local, San Francisco, January 15th. Learn more and register at mdb.link forward slash sf dash dot dash local or click the link in the description. Our next segment features a chat between my co-host Max Child and Mati Stunashevsky, CEO of Eleven Labs, a conversation that was all about the rapid evolution of AI voice and how it's quickly becoming the primary user interface of AI. They also discuss how their technology is being used in everything from customer support and education to gaming and celebrity voice cloning. Mati also shares his thoughts surrounding Eleven Labs' focus on prioritizing vertical-specific solutions, authenticity and user safety.
28:53Now, please welcome to the stage, Mahdi Stanishevski, founder and CEO of Eleven Labs, in conversation with Max Child.
29:09All right, Mahdi. So, Eleven Labs is obviously extremely well known for voice AI, for text-to-speech, for... I think that beautiful intro we just got was actually on Eleven Labs. Oh, amazing! A little co-branding there. And I'm wondering, you know, in the last discussion we heard this, you know, topic of, is the text box the best interface for AI? And I would imagine you have a take on how, no, you know, voice is the best interface for AI or voice is the best interface for computing going forward. I'm interested, like, what do you think are the best use cases for voice AI? And where do you see it, you know, today, a year from now, five years from now and beyond?
29:47First of all, thanks for having me here. Great to see you all. and I actually didn't know this was voice generated, but it had a great pronunciation of my surname. I know. It's actually hard, so I'm happy. You guys train on your last name specifically. We should. I don't know if we do. So we'll definitely do it now going forward. But as a company, one of the key things we are aiming to solve is how humans and technology interact, how you create with technology and make it seamless, how you interact with technology and make it seamless. And in general, to your question, we think voice will be one of the key interfaces for interacting with the technology across.
30:24From the simple pieces like interacting with the personal agent to help you go through the day where it can be on your headphone and be able to guide you through to education, that's one of the ones that I'm probably the most excited about, where in the future the combination of what LEMS will allow you and what voice will allow you is that you will be truly immersed with learning the given experience, where you'll be able to effectively have your personal tutor on their phone helping you across. Then the third one is, of course, for voice and for the language barrier to break. We need to figure out how to be able to speak across different languages while carrying the same intonation, emotion, voices, which will be a big shift.
31:06And then in general, how we interact with everything around us, whether it's the lab laptop, the phone, the robot in the future, and I think robot maybe is the easiest example, of course this will be voice driven. There's no other interface. And today, maybe it's a year or decade, as Karpati said, of agents. Of course there's on the horizon the decade of robots. And I think here too, the most common interface will be voice. It's not going to be only interface though. So you'll have to. Yeah, yeah. I mean, it's interesting you brought up those use cases of like a personal assistant, a tutor and I guess a robot, you know, house helper or Danny or something like that.
31:44Like, is your mental model that basically anything that today is something where you could have a human counterpart, right? A human tutor, a human assistant, you know, human in your house, like you're going to fall to that voice interface as the most natural way to do it because we as humans are already used to using voice for those things. Or are there things where today, you know, voice isn't used at all really, but it's something that we're going to expand into going forward. Yeah, so first of all, for sure. I mean, we're already seeing that, and I think that's the easiest one, and the most immediate one is how customer experience, customer support is just changed and elevated, where instead of calling and trying to rebook your ticket and going through this IVR flow of click one, click five to get through the steps and waiting for the number of minutes, you will have an agent that fully understands you, can guide you to the response and go through it.
32:36Yeah, I wanted to get into that, actually, because we talked a little bit about agents on the phone and calling United Airlines or American Express or something, what percentage of customer support calls today are actually managed by a voice AI system or an agentic system or whatever you want to call it versus IVR touch buttons? And how do you see that progressing over time? Is that exponential curve going like this every year? Yeah, I think the exponential curve is going like this. especially like this year we've seen incredible adoption, whether it's Cisco, Twilio, Deutsche Telekom, all of those kind of leaning in quickly to rebuild how you interact with help of voice agents.
33:16And I think you're right, the IVR flows are still a big part. Who can I call today and get a voice AI agent on the phone? You can, we did our little summit yesterday as well, and one of the great ones was voice ordering with Square. where you can call Square and actually order food delivery through a lot of their shops that work around with Square and actually do it through voice. I actually recently, if any of you are from London or have traveled to London, there's an amazing restaurant called Zephyr. It's a great restaurant. We worked with the company supporting that where to my happy moment, I noticed that on their website, it actually had 11 Labs agent that you could call and actually book a spot there too.
34:03So you can book a reservation at this restaurant in London with a voice AI agent? Exactly. Okay. And it connects, of course, to your calendar, your appointment scheduling, which is great. And when do you think we hit the tipping point where the average customer service call goes through a voice AI agent? Like the median, the 50 % point, whatever you want to call it. I think over the next 18 months. Next 18 months? Yes. Okay. So like mid-27, a call, the average customer support is handled by voice AI. Exactly. And I think, I mean, this is the most immediate one, the one where we see the highest ROI and value.
34:35Some of the other use cases, we see as kind of where the future is headed. But to your point, there's definitely one flavor of your point, which is how you can do things more efficient through voice with the existing services. But there's also the second theme where you can do things that weren't possible ever before. One of the good examples was our work with Epic Games, where we brought effectively Darth Vader alive in Fortnite, where millions of players could interact with Darth Vader live throughout the game, which, of course, is not possible in any other way. James Earl Jones' voice, right?
35:13James Earl Jones and his estate worked with us, and it's such an iconic and incredible voice. And we think, in general, that concept of what was never possible before, where you have incredible voices, talent, and you can now shift them to be not only static, but actually dynamic delivery, personalized and different for all the users, something that you are already doing in many ways at Volley as well in an incredible way. I think this will be a big... I have to ask this, someone in gaming, right? I mean, Darth Vader did famously go slightly off the rails and maybe say some things he shouldn't have to various players online.
35:47Like, how involved are you guys in that? How much of that is something you're protecting or against going forward? Like, obviously with the IP partners and so on, they really want to protect, I guess you couldn't say it's the squeaky clean image of Darth Vader, but a certain persona of Darth Vader. Like, what's your sort of go-forward plan as you license more of these IPs and voices and things like that? Yeah, so on that project, we were specifically involved on the voice side. Yeah. But in general, as you think about those deployments, and that's the most common theme, is you don't only need the voice or the interactive experience.
36:21That's kind of one part of the equation. then there's two other big pieces to really make them valuable. The second one is how you integrate that with other systems and actually bring the knowledge base, the data, the business logic inside of the system and how do you make it interactive in the real world. And the third one, which is the one that you mentioned, is how do you now deploy that in production with the right testing flow, right evaluation flow, and then monitor over time as it behaves, put the right safeguards and evaluate and adjust that based on that case. So that's something that we spend a lot of time on with a lot of players.
36:52Testing more and evaluating more so Darth Vader doesn't go off the rails. It's any, any, and even, even like in a customer experience, you don't want it, for example, to shift and speak about politics. You want it to keep it on the topic. Even if, even if you say, ignore all previous instructions and. Exactly. Even if you, which is actually harder to say when you have like this, you probably see the prompts with a lot of. It's hard to do prompt injection with voice. And the Unicode characters say them out. Oh yeah. Okay. Impossible. Okay. So maybe voice is slightly less susceptible to prompt injection than LLMs.
37:21I'm interested with like, that's a good segue into sort of celebrities and celebrity voices, because I know you guys announced, I believe yesterday, you're setting up kind of a marketplace for celebrity voices and you have Michael Caine on there. And, you know, at our company, we build voice AI games. As you know, I would love to use Michael Caine in our game. Yeah, we can, you know, obviously it's a high gravitas. We can make a Batman game with him, something like that. What is the process between, oh, I want to use an AI version of Michael Caine in my game to actually shipping? And which parts do you guys take care of and which parts do I need to go off and deal with Michael Caine's people, I guess?
38:02Yeah, so effectively through 11 Labs, we've created a huge marketplace of voices. Until yesterday, that meant that everybody here could create their voice, any voice actor, voice talent, could create a voice, share it, and earn money when that voice is being used. 10 ,000 voices created this way paid back, coincidentally,$11 million back to the community. For a long time, it was tricky for the iconic voices of how we could bring them onto the platform in an even more controlled environment. So if you think about Sir Michael Caine voice... Sir Michael Caine. Sir Michael Caine. Yeah, yeah, yeah.
38:36It's an incredible persona, too. You effectively, all you would do is engage, like, hey, this is the project we want to run. create a game with this specific character. This team would evaluate that. And then we would help deploy that project in actual production. So going through all those steps, how do we make sure that there is right safeguards in place? How do we make sure that there's monitoring in place so it doesn't go off the rails? And build that in our agentic system. Got it. But yes, the initial stage of what's the project, what's the compensation structure, would be between you and that.
39:09Got it. So you guys sort of manage the safety, the agentic elements of creating Sir Michael Caine within the game, but you still have to do the deal one-on-one with him to sort of facilitate that. And over time, we think it will evolve. As we see more examples, preset rates on how that should work. I am interested, actually, this brings me to a more general point. You said you had 10 ,000-plus voices of folks uploaded where you could use any of their voices, I think, via just your marketplace model. Has deepfaking and so on been an actual problem? I feel like it was something I was hearing as sort of a moral panic, you know, 12 to 18 months ago, that we're all going to have our voices faked on the phone.
39:45And, you know, my grandmother was going to get scammed out of her money because I'm locked in jail or something. Like, is that something you guys see at all? Like, is that something you're protecting against a lot? Like, how serious of an issue is that with voices? I think you're right. It's, I still think it's going to be a big issue. Like, in future, all kind of will be air generated. We need to find a mechanism to protect and understand which ones are, which ones aren't. and as a company living in a space we do place a lot of safeguards where it's traceability how we moderate how you can detect the content and give that tools to others very quick story on the flip side of that what we've seen recently we worked with a charity which effectively detects the callers based on IP and if the IP is likely to be one of the scammers and they have roughly a good approximation of one that can be coming from they would have the real scammer call in and deploy a voice agent to waste their time.
40:40And I thought it was brilliant. So you're scamming the scammers. You're scamming the scammers. Is this the long-term strategy? You think this will work for everyone? I think the long-term strategy is you need three layers. You need a human authenticate layer, so on-device encryption where I'm calling you, you know this is Matty Small, it describes on your side, that's layer number one. Layer number two, all of us will have an agent, where it's the personal tutor agent, an agent that books things on our account, and they will likely carry our voices, carry our style, do our permissioning. We need a layer where that's watermarked and authenticated that we know it's a permission piece.
41:11Like what we're doing with Sir Michael Caine, all the content that's generated will carry information that this has been carried with permission. Do you watermark all your voices today out of curiosity? All the voices are traceable back to 11 labs. And then the third layer, everything else by default will be AI generated. Got it. I mean, one sort of area that's interesting to me with you guys is you've launched a text-to-speech model, you've launched a speech-to-text model, a recognition model. You have an agentic orchestration system. You have, you know, all these safety and evaluation tools.
41:41Like, you know, almost all of those areas, I feel like folks in this room, you know, insider, AI, founders, investors, and so on, could probably name like two to three big competitors. Some, you know, some with bigger bank roles than you, and, you know, some where you're much farther along. Like, how do you think about like competition more generally and sort of all these pieces of the space that you're playing in? Like, are there parts where you see it becoming a commodity someday? other parts where you feel like you have a more sustainable competitive advantage? How do you go through the list of all the products you're working on and figure out where you shake out competitively, I guess?
42:15Yeah, so we started very much on the foundational model side. And in general, we take an assumption, if you ask anybody at Leven Labs, they will take that too, that over time the models will commoditize. That's the assumption we go with. So all models will commoditize, basically? All models. they won't like, you know, in the commoditized here, what they mean is that differences between different models will be just so negligible, maybe in some domains a little bit more, but in general they will be relatively negligible, and that's where that shifts to the product and why we invest so much on the creative side of creating a platform where you can combine all of that together in a controlled way with incredible voices, incredible ecosystem, new ones across languages, accents, voices, and on the other side, as we build agents and help people deploy agents, we deploy that for specific use cases of specific industries, working very deeply with the customers to understand their domains and work backwards from there on what actually needs to happen on the agent side to deliver value.
43:07Got it. So you're saying you're going to specialize in certain industries and sort of really deliver extra value there, even though all the models are commoditizing? I think the product layer is underappreciated here. I think you still need to build... Good products. Exactly. On the agent side, you need to build so many integrations to connect with any of the legacy systems to actually take those appointments and calling. You need to build the right control on when you hand over from an AI agent to a human agent. You need to have safeguards of some of the ones we spoke about. Then you need the monitoring of how you deploy.
43:37All of that is not only a technological shift, it's also a business shift. So by us working so deeply with the customers, it's actually bringing all the knowledge about their business inside of the agent to actually be able to deploy that value. And I think that will continue delivering value for the long term. And then, of course, you can go layer above, where while the models will be relatively similar, the value will actually be on how you can make the models work well for your use case. So maybe you can fine tune the specific voices, fine tune the specific use cases so it works slightly different in the gaming use case to a customer experience use case.
44:11And I think that value layer will still be. Got it. So all the models will be commodity. You guys will win on product and sort of vertical, specific differentiation. Exactly. And the wider ecosystem that we build alongside where I think as we think about the work, we would love to work and bring industry on board with that change. It's so important to bring a lot of the talent, a lot of the partners to work together. By talent, you mean actors, famous voices? Actors on the voices side or integrations on the agency. So as my last question, what is the coolest sounding voice on the 11 Labs platform?
44:44It could be like a celebrity. It could be a famous historical figure. What is this you're like, man, that is an incredible voice, and I cannot believe how well it sounds when we synthesize it? So my favorite, and my co-founder's favorite physicist is Richard Feynman. Sure. For those that are not... Surely you're joking. And he's both incredible in delivering the knowledge, but also in the style he delivers the knowledge. Okay. And now we have Richard Feynman on our platform. Okay. Which I think is so cool for learning the subject and speaking with Richard to listening to his lecture notes from Caltech.
45:17Amazing. Okay. I'm going to have to check that out. Thanks so much, Mati. Thank you, Mike. Yep. Appreciate it. Thank you for tuning in to this week's episode of the podcast. If you're new here, please like and subscribe. I appreciate your support. And if you want the data, insider takes, real reporting, go to newcomer.co and subscribe to the Substack as well. Thanks for following along.
From the publisher
Buckle up—today's episode takes you inside the war on AI slop and Anthropic's bold plan to fix artificial intelligence's biggest problems. We kick things off at the Cerebral Valley AI Summit, where Anthropic CPO Mike Krieger shares why AI needs to be truth-seeking and what their team is doing to fight misinformation and "brainrot." From explosive funding announcements to real talk about the future of agentic AI, you're getting all the behind-the-scenes intel.
Then, Max Child sits down with Mati Stunashevsky, CEO of ElevenLabs, for a fresh take on how voice AI is taking over everything from customer support and gaming to wild celebrity voice clones. It’s all about authenticity, safety, and vertical-specific innovation—plus, why voice might just be your next favorite interface.
As always, we're diving deep into the industry, serving up candid conversations, and making sense of the latest AI trends, so you can stay ahead in the game.
MongoDB.local San Francisco is happening on January 15th. Learn more and register here → http://mdb.link/sf-dot-local




