In short
AssemblyAI’s voice AI infrastructure and why voice is scaling fast, including new context-aware voice models, enterprise adoption, and how voice will expand beyond phones into consumer hardware and robotics.
Guest
Dylan Fox, CEO of AssemblyAI. Background: long-time builder in developer tools; taught himself programming in college; early interest in voice interfaces (Echo) and NLP/ML; AssemblyAI focuses on voice models and scaling infrastructure.
Key claims
Assembly handles 120M+ voice conversations/week (2M+ hours), up 800% over 3 years; ~1M developers and ~100M API calls/day. New voice models can use environment context (e.g., “McDonald’s order,” ignore background kids). TAM is expanding 100x because coding agents let “anyone” build with APIs. Assembly differentiates by 100% focus on voice infrastructure (models, inference scaling, orchestration, developer experience), not applications.
Notable examples
Granola; Tollands healthcare AI companions; Athelis/Commure AI scribes; police body-cam transcription; drive-through ordering; humanoid robots struggling to disambiguate multiple speakers.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOExplosive Growth of AssemblyAI
0:00 to 1:06
Learn about the rapid growth of voice conversations handled by AssemblyAI.
“The amount of weekly conversations that Assembly handles through our APIs every week is up over 800 % over the last three years.”
Building with Voice Infrastructure
1:20 to 2:08
Discussion on the importance of voice technology and its applications.
“specifically like use the transcripts that I capture from recordings into different kinds of formats.”
The Voice API Revolution
2:08 to 3:16
Explore why voice applications are surging in demand and how AssemblyAI handles it.
“So the demand for voice applications and the applications that companies are building on our infrastructure is over the last two years in particular really started to explode.”
Dylan's Journey and Vision
3:16 to 6:11
Understand Dylan's perspective on the future of voice AI and its growth potential.
“And so we're seeing that through our platform because we're the infrastructure really under all of it.”
Macro Trends Driving Voice Tech
6:11 to 10:01
Discover the three key trends that are driving the growth of voice technology.
“that are accelerating that further beyond just the core technology getting better.”
AssemblyAI's Unique Differentiation
10:01 to 11:28
Learn how AssemblyAI differentiates itself within the voice AI landscape.
“and across just like so many different use cases.”
The Secret to Scalability
11:28 to 13:50
Explore the engineering efforts behind AssemblyAI's scalable infrastructure.
“we have a ton of software at the orchestration layer that companies can leverage.”
AI-Native Company Structure
14:49 to 15:50
Discover how AssemblyAI leverages AI to optimize internal processes.
“developers or on the outward facing side in marketing, that kind of thing.”
AI-Native Company Structure
15:55 to 16:28
Discover how AssemblyAI leverages AI to optimize internal processes.
“Hey, go create a new version of the deck.”
The Future of Voice AI
16:29 to 18:23
Explore the evolving landscape of voice AI and its potential applications.
“and we try to be intentional about that.”
Show all 17 chapters
Voice as a Data Capture Modality
18:24 to 22:46
Understand the significance of voice as a reliable form of data capture.
“I think you're going to have, so here's like what I genuinely see, like voice is a, I would think of the word like modality or interface, but I don't even think those capture it.”
Challenges in Humanoid Robotics
22:47 to 25:15
Learn about the challenges humanoid robots face in voice recognition.
“And they, like, love to mess with the Matic robot.”
Multilingual Voice AI Models
25:16 to 28:33
Delve into the complexities of developing multilingual voice AI models.
“And it's less about, it's less about like the science of it.”
The Future of Voice AI
29:28 to 39:58
Exploration of AI and voice technology, including challenges in natural language understanding.
“So you just plug in some USB stick and put it on your hand and then it just reads your mind.”
Innovations in Voice Models
39:59 to 42:00
Discussion on the latest advancements in voice models and their applications.
“And I think we're going to want that to change because I don't know about you, but I want to know, is this a human or an agent?”
Advancements in Voice Models
42:00 to 42:47
Learn about the latest developments in voice modeling technology and its potential.
“they think like, oh, yeah, it should like match like a human.”
The Impact of Relationships on Performance
42:47 to 45:00
Discover how personal relationships influence entrepreneurial success.
“Like if you had models at McDonald's, they don't know they're at McDonald's?”
Transcript
Automatic transcript. May contain errors.0:00The amount of weekly conversations that Assembly handles through our APIs every week is up over 800 % over the last three years. On a peak week, there'll be something like over 120 million voice conversations going through our platform, over 2 million hours of voice. It's a little over 4x the amount of daily volume that's going to YouTube. So they're getting more accurate, they're getting faster, more controllable, more capabilities. The newest voice models that we just launched, you can give them context about the environment they're operating in. If you drive through a McDonald's and you order with your voice, that voice AI system has no clue that it's taking a McDonald's order.
0:33But with our models, you can give them context on, hey, you're taking a McDonald's order. You want to focus on the person that's ordering. Ignore the kids that are screaming in the background. And so we're seeing this huge inflection now of enterprises, small businesses that can build with API infrastructure now. Our TAM has just increased by 100x because we're not just selling to engineering teams within product companies. It's now like anyone.
1:06Dylan Fox:Dylan Fox, welcome to Sorcery. Yeah, thanks for having me here. So we're going to have a very fun conversation on voice. We were just gabbing off camera about how much I love voice and how I've used it to start this podcast and do all these interviews. specifically like use the transcripts that I capture from recordings into different kinds of formats. I started it for investing and so I would take that I would repackage it I would make like open source expert calls and open source memos and all this kind of stuff just from capturing that data. Now you have built the infrastructure for how that even works so I want to get into that I want to get into everything from the business model how you built out the products You've been doing this for a very long time and there's a lot of new entrants to the category.
1:51Dylan Fox:So it'd be great to learn what is BS and what's not, as well as like areas that you're most excited about. But before we start, I guess it'd be great to kind of capture and illustrate just how big Assembly AI has gotten, how much data you train on. Yeah. So the demand for voice applications and the applications that companies are building on our infrastructure is over the last two years in particular really started to explode. One stat I was just looking at before I came over here was like the amount of weekly conversations that Assembly handles through our APIs every week is up over 800 percent over the last three years.
2:32Oh, my God. And so now, you know, on a given week, on a peak week, there'll be something like over 120 million conversations, voice conversations going through our platform, over 2 million hours of voice, which as of December of this past year was 4x or a little over 4x the amount of daily volume that's going to YouTube.
2:55Dylan Fox:Which when I found that was like, wait, I need to double check this because, you know, that can't be right. But yeah, the volume is huge. There's almost 100 million API calls a day coming against our API, about a million developers, a little over a million developers on the platform now. And 40 % of those developers signed up to the API last year. And so while we started this a long time ago, it's really over the last two, three years where voice and it's accelerating and inflecting voice is becoming just a core part of software and increasingly hardware too. And so we're seeing that through our platform because we're the infrastructure really under all of it.
3:34Dylan Fox:So if you look back over time, you guys went through YC. Yep. Did you think that this would be the reason why voice would take off and people talking to their phones and talking to their computers and that kind of thing? Like, what did you think? Yeah, I mean, I did, which is why I've spent so much time on this because we were actually the very first AI batch in YC when we went through YC. And so it was Daniel Gross, who, if you know Daniel now at Meta, he started the AI batch at YC back when we went through in 2017. And it was me and like five other companies. And we got, you know,$100 ,000 in GPU credits.
4:13Dylan Fox:That was like our perk, which now seems like cute, right? Right. It's like, wow, you know, that's nothing. But back then it was like, wow, you know, $100 ,000 of like NVIDIA K80 usage. Like, this is amazing. But the, you know, the AI ecosystem back then was just like in its infancy. Like I was going to the very first like TensorFlow meetups, just to put it into perspective. Like it was, no one was really using AI in production yet. But for me personally, I had gotten an Amazon Echo and was like really into voice interfaces and talking to this hardware that worked well. Because my experience with voice prior had been, everything was terrible.
4:51And so then I went looking to find like APIs that I could just in my spare time back then like build with and I couldn't find anything. And all the technology kind of sucked at the time. But I felt like, okay, over the next 10 years, this is going to get so much better. And when it does, if it's easily accessible to developers like a Twilio or a Stripe, you're going to have so many creative people build really cool things. And so now, you know, that's what gets us excited when we see, you know, our customers like Granola, right? Granola building these like amazing products that are helping companies and professionals like run their businesses.
5:29You know, we have customers like Tollands that are building, we were talking about it before, like AI companions that hundreds of thousands of people are talking to for like hours every day in the healthcare space. Athelis, Commure, it's a big health tech company building AI scribes for doctors and And it's really cool to see what people are building. So I think that to go back to your question, like I always viewed this journey as similar to self-driving cars where it wasn't a question of like, oh, is the product market fit gonna come? It was like, no, the technology just sucks. And as it gets better and better, you're gonna continuously cross these thresholds where you open up more tranches of the market.
6:07And that's still happening. And there's new macro trends that are happening that are accelerating that further beyond just the core technology getting better. I can talk about those, but that was very much the thesis behind the company. For me, it was like, this technology is going to get better. It's going to enable so many new applications and voice to be something you can really build with and a core modality in our lives. And it just felt like something really cool to work on.
6:39Dylan Fox:So what are those core macro trends? So there's really three things that we see that are causing, you know, everything, all our metrics like inflect. The first, which is what we control, is voice models are getting just much better. So they're getting more accurate. They're getting faster, more controllable, more capabilities. You know, the newest voice models that we just launched, for example, you can give them context about the environment they're operating in. And to give you an example of that is like, you know, today, if you drive through a McDonald's and you order with your voice and there's a voice AI system there, that voice AI system has no clue that it's taking a McDonald's order, right?
7:18It's like, it could be, it has no context. It's like fresh. But with our models, you can give them context on, hey, you're taking a McDonald's order. You want to focus on the person that's ordering. Ignore the kids that are screaming in the background. And those types of capabilities make them much more accurate. And then they're also faster and they're lower cost, so they can be more widely deployed. So you have these voice models that are just getting better across these key dimensions. That's like what we control. But the other two macro trends, the one is number two in this list of three is like you have this ecosystem of AI infrastructure getting built out.
7:55So you have these reasoning models, you have vector databases, you have other AI models across other modalities. And so now when you're a company wanting to build something like drive-through voice ordering or AI note-taking or healthcare, you don't just have the voice models. You have the rest of the infrastructure you need to go build the thing you're trying to build. And that has been accelerating over the last couple of years. And then the third, which is new, is coding agents. So we're now seeing companies sign up, spending a lot of money on API that like not even two years ago, like 18 months ago, like I would never think of them as our target customers because we're API platform.
8:42So you would think like, oh, you're selling to engineering teams, product teams. But now everyone's an engineer because they have coding agents. And so we saw this small business lawn care chain sign up. And we're like, what do they do with the API? And a lot of these small businesses are automating parts of their back office using Lovable or Cloud Code or Cursor or whatever, Replit. And those coding agents are building our infrastructure. And so we're seeing this huge inflection now of like you have enterprises, you have small businesses that can build with API infrastructure now. So I think about this as like our TAM has just increased by 100x because we're not just selling to engineering teams within product companies.
9:31It's now like anyone. anyone. And we're seeing some global enterprises now come in and build their own software for a bunch of different use cases. And they're doing that on top of our infrastructure. And when I go talk to them, like, why are they doing this? It's like, well, they have two people working on this with a bunch of coding agents. And so those three things, it's the voice models are getting better, like the broader ecosystem of AI technology. And then the third is anyone can build this stuff now because of coding agents are really causing voice to be way more widely deployed and across just like so many different use cases.
10:11Dylan Fox:So for people that are new to this space, it would be helpful to explain your differentiation between the likes of like 11 labs. Obviously, Sierra is in a different bucket too. But like, how do you fit in within that world? Yeah, so we are 100 % focused on voice AI infrastructure. So we don't do anything at the application layer. Where we focus is the models. So we create models that are amazing, voice models are amazing for all those use cases I spoke about. So healthcare, we have medical-focused models, drive-through voice ordering, contact center, and note-taking. We then build out the inference around those models.
10:54So to handle 4x the amount of YouTube volume on a day, 120 million conversations a week, and growing over 100 % year over year, there's a lot of infrastructure we have to build out to make sure everything scales, is available, is super fast, is low cost for customers. So we build out a ton of infrastructure on inference around our models. And then we also do a lot around the orchestration layer. So if you want to build a voice agent, if you want to understand speakers, if you want to translate data, we have a ton of software at the orchestration layer that companies can leverage. And then above that, we have this amazing developer experience, agentic coding experience, so agents can easily build with our stuff.
11:42And when we say infrastructure, I think a lot of times people think like, oh, you're just creating the model weights, right? Like you're just creating like models. That's a part of it. But those other parts are equally important to us, the inference, the orchestration, the developer and agent experience. And so we're 100 % focused on that. And so and across like all the different use cases that we focus on and making sure everything's like tailored for those. And so the reason most AI note takers are using Assembly as the voice infrastructure is because our models are the most scalable. If you have tons of people, millions of users that you're doing live note taking for, we have the most scalable models in the market that are super low cost, high accuracy, very fast, can burst to huge amounts of traffic.
12:32You don't have to stand up servers as a customer. So all of that is like the area that we're focused on. And we're enabling companies to build on top of that.
12:41Dylan Fox:How did you get it to become so efficient? You know, years of engineering work. People like Ben at our company. We have an amazing team of engineers, of researchers that just operate so closely to customers that they really understand like how this stuff is being deployed. I think that's the biggest difference between us and a lab at a bigger company. So we have a lot of researchers, research engineers that will come to assembly from a bigger company. And they're so far removed from the customer. So they just create their model, they benchmark it, and then it's like, okay, I'm done. But you have to know, all right, who are the customers?
13:21How are they using this? What do they care about? What errors break their application and what errors don't matter? And then how do you optimize both your models and your infrastructure for that? So we've just spent years and years building out the infrastructure to make this stuff work. So infrastructure that's like cross-region and cross-cloud and stuff that just will really scale. But it's probably like half our work is spent just making our infrastructure more and more scalable.
13:50Dylan Fox:This episode is brought to you by Brex, my favorite. You become what you spend on. And I refuse to spend my time on work that shouldn't exist. Expense reports, receipt chasing, and manual closes. The companies building what's next from Vercel, OpenAI, Anthropic, Granola, and Deepgram all made the same call. They all run on Brex. Brex is the intelligent finance platform that combines cards, expenses, and banking into a single stack with agentic finance built in. AI agents that handle expenses automatically, enforce policy before spend happens, and close your books in minutes. That's why Sorcery runs on Brex, so I can spend time on building and not busy work.
14:35Dylan Fox:It's time to get Brex AF. Learn more at brex.com slash sorcery. That's B-R-E-X dot com slash S-O-U-R-C-E-R-Y. Bye. I feel like between like, well, the SaaS apocalypse is one thing, but even leading up to this point with AI helping companies restructure, whether it's on the kind of internal side with developers or on the outward facing side in marketing, that kind of thing. I'd be really interested to hear how you re-architectured the company internally to do that. Yeah, it's a great question. I mean, we're, we're AI native company. Like everything we do is, is AI native. And so what I mean by that is like, I have even me personally, like I have my own agent that like I built, right.
15:28I call it like my Dylan claw on my computer and it has my meeting notes and it can make slides for me and it can do it, do everything. And so I just onboarded someone new to the company and I was going through my slide deck and it's like, Oh, I was talking through a part of the deck that was like a bit clunky. And then because my agent has access to the slides and can make new slides and was, you know, has the notes from that meeting powered by assembly. I was like, Hey, go create a new version of the deck. That is a bit easier to walk through and like, look at the transcript from that most recent onboarding to make some changes.
16:04And I want to like merge these two decks together. And then like 20 minutes later, it's done and it's perfect. And I'll have to edit it. We have an agent that has access to all the company information. So our metrics and the whole company uses it. It can submit PRs in GitHub across our products. And so I think like, you know, we've always stayed very small. You know, we're only, I think we'll be 80 people at the company pretty soon. So we have huge scale for the size of company that we're at. and we try to be intentional about that. But, you know, I think for us, like, we've just really leaned in to try to, you know, use AI, like, across everything.
16:43Like, you know, the most recent example is we had Cloud rebuild our whole website off of Webflow and just deploy it on Vercel. And now anyone, you know, people in Slack can just make changes to the website. So you, like, get off a customer call, you're like, oh, this part needs updating. And it's like, boom, you can just change it. So it's things like that that have, I think, helped us leverage AI across the company.
17:05Dylan Fox:Sovereign AI has been a huge topic. And you mentioned some of your customers that are in categories that are a bit more critical and sensitive for data. But we are seeing this across the board because you can get copied by the labs. Yeah. So you can deploy our platform on-premise. You can run it all self-hosted, and then it's totally private. And we do see an uptick in that for sure. I think that it's this mix like in voice is so early where you know we we the models and the capabilities are getting better so quickly that it's still hard to like to branch out you know what I mean by that like like we've seen some customers maybe six months ago or a year ago they'll take an open source model and fine-tune it or something and then you know they're now like okay this is like shit this is like really behind and it's like nightmare to maintain and I don't want to work on this.
18:01And so, you know, they'll come to assembly. I think voice is much earlier where we haven't seen as much of that as you're, you know, you're probably talking about in like, yeah, if you're a bank and everything's going into cloud or something, it's like, you know, wait, all my proprietary kind of knowledge is going out. But we are working on ways for companies to create like their own custom models, deploy them on our inference or run them, you know, on their own cloud but it's it's earlier there for voice for sure i was thinking about it because when we were
18:27Dylan Fox:at the raise summit this was a hot take that max cook who's a sector head at code 2 said he thinks the keyboard and mouse is over we've been through so many different form factors of computers yeah human like machine interface that's like no we gotta close the books on the computer keyboard and mouse it's funny when you go into offices and everyone has like the gooseneck microphones and they're like whispering over. I think you're going to have, so here's like what I genuinely see, like voice is a, I would think of the word like modality or interface, but I don't even think those capture it. Like voice is a reliable form of, let's just call it data capture now.
19:11Like for the first time probably ever in the past year, it's like a reliable form of data capture. And so what does that mean? That means like you can actually talk to your computer now. You can actually have it listen in. You have your computer listen in on your meetings and take notes for you. Doctors can use it during your patient encounter to update the charts. We have a startup. It's called Ciro AI. I don't know if you've heard of them. They're building on the platform. They're creating an AI for field service tech. So if you're a plumber, HVAC technician, They have an app you can run and it will listen in on your service visit and then give you feedback after for like sales coaching.
19:49And it's helping those field tech salespeople be so much more successful because like they're actually getting feedback and they never were before. So they're earning like 20 % more take home pay or something. You know, we're seeing it with like toys and with hardware where you can like talk to toys now, which is cool for kids. You know, me as a parent, seeing that stuff is really cool. So voice is this reliable form of data capture. And I think when people talk about voice AI, a lot of times they're talking about like AI voices, right? And that's a part of it. That's a big part of it. Like you have voices that sound great and they're amazing and they're so much better.
20:25But the other part is like, can the computer or the machine understand you? And that voice capture part, is that reliable? Is that solid? Does that work at scale? But that is the case now. and so I actually think you're going to see you know like touchscreens didn't replace keyboards it's like you have both you have the you know on a coffee machine like there's the rotary dial like and there's the button the touchscreen menu and I think voice is going to be another dimension that's probably the way the way I think about it voice is going to be a new dimension and but it's going to be an additional dimension and so for me like I like to type but I also like to talk to my computer I like to do both things and I think for you know where I get excited is like consumer electronics that you're going to be able to talk to.
21:11So like your TV, like, it's still crazy that like most TVs, you can't just be like, Hey, put on like open up Apple TV and put on, you know, this, this episode of whatever.
21:20Dylan Fox:They are always listening. I don't know if you saw the succession. No, I didn't see that. No, always listening. Yeah. But like, it's crazy that doesn't work yet. And so I don't think that means that computers are just going to to be like a screen and you just talk to it because actually we'd get tired of talking but but I think it's going to be a dimension that's added to everything and so that over the next couple years will happen and I think that's really cool because I think that will allow for us to be freed of like our addiction to screens right because your input to most computers now is like the this like touchscreen it's like smaller and smaller touchscreen that like just captures you.
22:04But with voice, that enables more passive hardware and more passive computing experiences where it's like you can talk to things and you can be walking around and you can be free from being like a prisoner to the device. So that's what I think is really cool. And that's for sure coming. We have some like big consumer electronics companies like actively engaged with us, like implementing our technology right now for things like that, which is like really cool.
22:33Dylan Fox:What about humanoid robotics? Those two. You know, I think that, you know, every time I think about, like, the Neo or the figure, I think about, like, my, I have a four-year-old and a three-year-old. And I think about, like, one of those in the house with them. Because we have a Matic robot. And they, like, love to mess with the Matic robot. They broke one. No. It was, like, because they always go run over it and, like, turn it on. Because they realize, like, you can press the button and turn it on. and then they like sit on it and they mess with it and try to ride it around. And yeah, I don't know what the weight is of a Neo or something.
23:07But yeah, I think you'll have humanoid robots. Obviously, the way you're going to want to interact with a humanoid robot is by talking to it. But this is where what we do is so important is because I'll tell you one of the main problems the humanoid robots face today, because a lot of them are using our APIs. If you have three people standing next to the robot, it doesn't know who to listen to. and it has a hard time disambiguating who's saying what. And so it all just gets kind of jumbled and merged. This is a big problem.
23:36Dylan Fox:Because it's using listening data, not video data? Yeah, exactly. And even if it's video, like if you're not facing it, maybe you can't see, right? So like it does have the cameras, but you really need to be able to, like for you, like for me, like if I close my eyes, I can still tell who's saying what, right? But for voice agents, even if they're over the phone, this is a big problem with them today. you know i i last week or last month like called a voice agent and i called a restaurant it was a voice agent that answered right so it's like ai that i'm talking to and i know it's an ai even though it tries to trick me because like i build this stuff but my son's in the back and he's talking and it's like tripping up the voice agent because it doesn't know that that's a background speaker it just hears speech and it's like i need to respond to this and those problems are still not solved yet these voice models and that's going back to like the sovereign ai stuff i think right now we're still in the phase of like, I need this stuff to work.
24:29And I just need to like, cause there's, there's such an opportunity if you can be the first to get it to work. And we're still in that, like, get this stuff to work phase. Probably that will be the next, you know, 18 months. And, uh, and yeah, I think that like, yeah, across humanoid robots, consumer electronics, like you're going to see voice as this dimension that you just like expect. And what the example I think about is, you know, when you see pictures of like kids, kids trying to swipe on TVs or something because they just expect everything is a touch screen. I think it's going to be similar where in five years, kids are just going to talk at things expecting them to be able to understand them.
Read the full transcript
25:10And that's the shift with computing in general that we're going to see.
25:14Dylan Fox:What has it been like to expand the models into translating different languages? That's a hard, it's a hard problem. And it's less about, it's less about like the science of it. And it's actually more about like the culture of it, right? So like our head of research is from, is Japanese from Japan. And there's like a lot of, you know, there's, there's a lot of like opinions about like how, you know, like how you like think about it as like policy alignment, policy alignment for how you want it to show up for a native speaker. My wife is Danish. And so I always have her test our Danish models. And I'm like, well, is it good?
25:52You know, like, try it. And she's like, yeah, but like this, you wouldn't really do. And it's like these details. But they're important details. And so I think it's less about the science. And it's more about, can you get those details right? And more and more, you want to have a single model that can do all these different languages. And so our latest model, for example, it's a single model. It can do 20 different languages. You don't have to tell it. It just will figure it out. It can switch across those languages in real time. But if you fix this one thing, then you might break this other thing.
26:28And so it's that equilibrium that you're trying to make in these models. And you have to have this kind of deep understanding of each language. And it's why, actually, if you look at local voice AI vendors in certain countries, they They actually are typically the best because it's, you know, they speak the language, they understand it and they can more quickly identify like which data is good and stuff to train on. So that's a big part of it is like that like policy alignment and like the language expertise per language that you really need.
27:01Dylan Fox:Would you acquire those companies? Like how would you get that? Yeah, I mean, a big a big part is like, you know, we we try to build a team of those experts. We have a lot of customers that help out and service those experts and give us that feedback. We partner really closely with customers. A lot of them are not working in sensitive applications, so they opt in to letting us use their data to improve models and their feedback, which is great. So it's a mix of all those things. But that's really where there's still a lot of opportunity to differentiate at these models. And I think for voice, a lot of people will think like, oh, voice, like what people have been telling me forever.
27:43It's like, oh, voices, you know, isn't that solved? Like, isn't that just like commoditized technology? But like, it's definitely not because there's so many gaps across languages, across different verticals and applications and domains that there's big rooms for improvement in still.
28:01Dylan Fox:If you're building what's next in AI, you need to know MongoDB. The database platform developers love and built for the agents you're running. MongoDB stores searches and reasons over your data in real time with vector search and embeddings from Voyage AI all in the same system. No separate pipelines, no stitching together 10 different tools. It's why 75 % of the Fortune 100 and leading AI native startups run on MongoDB. Build and scale from your first user to billions of vectors. Go to mongodb.com slash AI to learn more. That's mongodb.com slash AI to learn more. Bye. Assembly AI is a voice AI infrastructure layer millions of developers build on.
28:48Dylan Fox:They build the industry's best speech-to-text, voice agent, and speech understanding models that serve as critical infrastructure for companies like Granola, Haygen, Ashby, and ClickUp. Their speech-to-text models lead the industry in accuracy and quality, and their speech understanding models help you go beyond transcription by uncovering insights, identifying speakers, and highlighting key information from voice data. You can get started today at assemblyai.com slash sorcery and get$50 of free credits to start building voice AI products. That's assemblyai.com slash S-O-U-R-C-E-R-Y. I have a hard question.
29:28Yeah.
29:28Dylan Fox:how far away i'm gonna try to deliver this without smiling
29:37Dylan Fox:how far away are we from translating dogs and cats dogs and cats yeah i don't think voice will be the interface i think it's gonna be i think it's gonna be more like uh like mri scans or something like something sub vocal that's the thing i really think about too will we talk about computers, I think probably the ultimate interface is, it's just reading your mind, right? So you just plug in some USB stick and put it on your hand and then it just reads your mind. And then it's just like, it's unlimited bandwidth to the machine. So that's probably the solution for the dogs and cats. I don't know.
30:12We both said, we both watched Project Hail Mary and he did translate an alien yep so i know yeah i was telling you this before but you know as i was watching that part i was like this is way too easy for him to trade this model basically like he would need
30:30Dylan Fox:like a lot more data uh but yeah so where did he go wrong with that with uh with rocky with with rocky well he solved it i mean he yeah they were able to talk uh but yeah he did they they did turn Rocky into more of like as we were talking about like a dog that I think like a very intelligent space-faring entity yeah that was a weird realization because I was like thinking about it I was like they both left their planets I mean Rocky's supposed to be like 100 or 200 years old yeah right and like he's an adult and he has a lot of responsibilities but like I mean both of the characters were quite playful and I think that's why the movie was so great yeah but but it was definitely super weird.
31:15Dylan Fox:Yeah, yeah, yeah. Why was Rocky the dog? Yeah, they made him like Grace was like the parrot, but it's probably good for viewers. Now I'm like totally off tangent. Okay, so I guess like to get deeper into how you train based off of the hundreds of millions of hours of data that you have, how does that really work? Yeah. My experience has been, it's really probably like 75 % of it is like the data that you're training on. Like I would say for any AI model, it's like, there's always like, you know, like these, these like step functions and like algorithms and architectures and stuff, but the data is just so important.
32:02And so a huge, we spent a huge amount of time just on data, trying out different data mixtures, training different models. And we release model updates like every couple of weeks. And that's the big benefit that our customers have when they're building on our platform. It's literally like constantly getting better. It's not like every six months or every year there's an update. It's like there are constant improvements going out. We have a ton of different model versions. We have a ton of different APIs. We just launched like a different API today for different use cases for dictation and like push to talk type use cases.
32:35So we're constantly innovating. On the model training part, so much of it's about the data. Is the quality good? Is it aligned with what users want? And I'll just give you an example of that. like for for the speech to text task like some applications don't want to pick up the background speakers some do so like we have some customers are you know looking at police body cam footage for like legal discovery they're running that through our platform and they want to know like everything that was said by every single speaker like leave nothing out right but if you're mcdonald's taking a voice order you don't care about anything other than like the person in the front seat placing that order.
33:15If the kid in the back screen screens like an extra large frosty, it's like, you know, you want to be smart enough to ignore that. And so there's all these different expectations of these models. And that's where I think the alignment of data to the use cases is really the process knowledge that is required to build like really good technology. And it's why like if you go online, you look at some of these open source benchmarks, it's like, oh, wow, all these models are performing similarly, but it's very easy to optimize for a public open source benchmark. It's very hard to optimize across all these real world applications.
33:50And so we spend a ton of time on evals. We have a whole team internally that works on evals, constantly evaling our models across a million different metrics, a million different data sets, across all these languages. And so much of it's about the data.
34:04Dylan Fox:How is the team constructed? How many employees do you have? So we're, yeah, we're almost 80 people at the company. It's pretty much like almost entirely like product research and engineering and forward deployed engineers that are just like working really closely with customers. We didn't really go into your background too much. Sure. Yeah, yeah, yeah. I mean, it's interesting. I sometimes I'm like, wow, I've been doing this a long time. I like really where it goes back to is I taught myself how to program in college. And actually, like way before that, as a kid, my brother would build computers like in our basement and I would be on like IRC, if you know of IRC.
34:43Totally. Yeah.
34:44Dylan Fox:Yeah. No, these like I would be playing these like video games and in these like chat rooms with like random people and like building, you know, websites for these like teams we make. And so I was like always on computers and then kind of got into programming and taught myself how to code in college. And I loved the idea of being able to build something and these hard problems in computers. So that got me into natural language processing and machine learning. And that's where I really got interested in this natural language understanding voice area. And it just turned out to be this really cool emerging space where we get to work with all these cool innovative companies and build new technology.
35:27And so for me, that's the really cool thing. like building this. I think every company has a DNA and like a reason why they started or why it started. And for me, it was like, I love developer tools. Like I'm developer by background. Like that's kind of always what I was passionate about and giving like other people tools to build on and seeing like what they can do with it is so fun and cool. And every time we see, you know, some startup build something on the API and then they like raise the series A and stuff. It's like really cool to see. At some point I want to try to figure out like what's the cumulative total funds raised of like all the startups on assembly.
36:11Cause just to show like how much demand is going towards all this stuff. But yeah, my, my background's really about building things.
36:17Dylan Fox:So that's why your infrastructure layer versus application. Yeah, exactly. I think, you know, companies have their focuses and their, and their DNA. And for us, it's like infrastructure and we're really good at scaling stuff. We're really good at creating technology. And that's our DNA. And that's ultimately like what our customers are buying from us. It's like, hey, we're this amazing infrastructure platform. We're like the AWS for voice capabilities. That's the product vision that we have. I know some people can't think this far in advance because AI has made things, time window is very short.
36:54Dylan Fox:Yeah. But what are you most excited for in the next 12 months? Yeah. So 12 months is, it's actually, I feel like it's easier to predict like 24, 36 months than it is 12. I think over the next year, we'll see a lot of, a lot more consumer applications, hardware and software where voice is a core dimension so toys games um consumer electronics and i think that's going to be really cool because i think that's gonna you know like i one like free us from our phones to create this more futuristic environment where technology like blends in more you know because right now a lot of technology a lot of technology like it's kind of like when you see a fridge that has the custom panels you know in a kitchen it like blends in and you don't notice the fridge it's like where's the fridge versus like the fridge that's just like stainless steel and this like object there and I think that voice is a way to create like technology like that that just kind of exists because you don't need to be able to access it.
38:05You can talk to it and it can talk back to you. So I think consumer is going to be an area where there's a lot of innovation. We're, for example, working on on-device models. So models that can run on a phone or on a really low-powered piece of hardware, like a remote control for a TV or something. And that's going to expand a lot of these applications. So that's part. And then I would say there's something that we all in the voice industry have to reckon with, which is this, you know, there's this interesting difference where every text-based agent that you talk to, customer support in an app, everywhere, at least my experience has been, I've not encountered a single one that's like trying to convince me that it's a human.
38:57It's always like, hey, I'm, you know, the AISDR. I'm the AI support agent. I can help you with this. Voice is different where today when you're building a voice agent that you're talking to over the phone or something, for the most part, you're trying to trick the goal is to try to trick the human into believing that it's also a human. There's exceptions to that, like Tollands that we talk about, where it's very clearly a, you know, it's like an AI avatar. Like you, it's like a character that you're talking to. And I think that as an industry, that's something that you kind of have to figure out.
39:41I think right now you're trying to trick the user to believe you're a human because, you know, that's the way to mimic intelligence. Everyone's used to these like really shitty voice experiences that just like don't work. Yeah. So they bail right away, right? And we see this. Our customers, if they're building a voice agent and you disclose up front that you're an AI, people just hang up. Versus if you don't, people continue. And I do think that will change. And I think we're going to want that to change because I don't know about you, but I want to know, is this a human or an agent? I had this really weird call.
40:17Dylan Fox:Would you act differently? Yeah, I would. I had this really weird call, and this is now starting to trip me out, but we recently had to move some stuff. We were at a moving company, and I got a call, and it was like, I was like, after, I was like, that was a weird call. I was like, was that it? I think that was an AI, but it was like, it was like with an accent. The AI had an accent. It was like really trying to trick me that it was a human, and it was just, the result was like, this is weird. that I felt. And so I think that with voice agents in particular, there's something we still have to figure out about the UX for AI that you're talking to, to balance the, hey, this is intelligent, but I'm not trying to pretend to be a human and trick you experience.
41:14Because if you ever talk to a voice agent and then like two minutes in, you realize you're talking to AI, it's just like a weird experience. And I think that's something that we still have to figure out.
41:24Dylan Fox:I feel like at some point we should just like assume everything is AI though. Maybe. Imagine you call some like medical helpline and then you're like asking for help. When our kids were born, I used to call the like nurse line for the hospital to ask questions. And like, imagine now I'm doing that. And then like two minutes in, it's like, Like, you know, I realize I'm talking to AI and be like, wait, what is going on? Versus like, no, I want to call like and talk to like a human nurse. You know, I want to know, you know, I want the options. But I don't think we figured out how to balance that today.
41:58And I think today when the market and the industry talks about a good voice experience, they think like, oh, yeah, it should like match like a human. And I think you can have good pacing and a good conversation. It can be natural, like with a robot. That should be, there's examples in science fiction where that's been achieved. And I think we could figure out a way to do the same.
42:23Dylan Fox:So I know we covered a lot of topics. Is there anything that we didn't cover that you want to touch upon? I feel like we covered a lot of it. Yeah, if you, yeah, the new models that we're building that have the ability to take in context are like really cool. They're the first type of voice models that can do that. But no, it was great to cover the full range. It seems kind of obvious, though. Like if you had models at McDonald's, they don't know they're at McDonald's? Like that's crazy. Yeah, yeah. The voice models haven't been able to do that yet. And, you know, we shipped like really like the first one that can a couple of weeks ago.
43:02Dylan Fox:That's super cool. Yeah. Okay, so as we close out, there's one question I have to ask. because this is a Brex question, because they're all about performance. Spending is smarter, moving faster. I like to think that for personal performance, it's kind of who you surround yourself with. Some people say it's like the five closest people, like who's either mentored you, who's a close friend, who's been inspiring. Who are those people for you? Yeah, I really put it in like three buckets, like family, friends, and then it's going to sounds so corny, but like really are investors and, you know, on family, like my wife's an entrepreneur too.
43:41She's a founder. And that's been like a blessing because, you know, that, yeah, always able to like get advice from her and talk to her friends. You know, I've been able to meet over the last couple of years, like other founders that are in the same stage and phase. And that's been amazing. And I think having like finding peers that you can just like be super open with and transparent with is super helpful. But then in terms of investors, like we have an amazing group of investors that has really been along for the ride. You know, and I think about like Keith Block and Smith Point that invested in our company at our Series C.
44:25They're operators from Salesforce. They've started their own VC fund. And like they're just, you know, so I think like Steve from Excel, Steve Laughlin, Rebecca from Insight, like they're just always like pushing the company and pushing me and in in great ways. And they're they're amazing people. So that I'm not even trying to like be cheesy when I say, oh, I don't feel like that's a hot take.
44:50Dylan Fox:I feel like a lot of founders are like, you know, I don't need investors. Like I'm good. Yeah, they're great. And I mean, I think we've been really, really lucky to have a great group of people around the company. So yeah, I think friends, family, you know, investors have been super helpful for me along the journey. Amazing. It's a good place to end it. Well, Dylan, thank you so much. I appreciate you sharing everything on Assembly in the early days to even the craziness that's happening now. And I hope pretty soon you can start translating dogs. Yeah, we're working on it. April 1st, come check it out.
45:26People, yeah.
45:28Dylan Fox:Thank you. Yeah, cool. Thanks for having me on. Huge thank you to the entire RAISE team for an incredible event. And thank you to Brex, MongoDB, and Assembly AI for making this trip and series possible. If you enjoyed this conversation, you're going to love the rest of the RAISE series with Tony Kim from BlackRock, Scott Wu from Cognition, Andrew Feldman from Cerebris, Rodrigo Yang from Salmonova Michael Hurlston from Lumentum CJ Desai from MongoDB and many many more like our hot takes that we did at a secret location that you can find on X, YouTube and Instagram subscribe to Sorcery on YouTube for more conversations with the people shaping AI and join the free newsletter you can also do paid at sorcery.vc for weekly insights on AI, robotics enterprise software consumer, semiconductors, did I say AI?
46:20Dylan Fox:AI again, and everything that's coming next, like funding announcements and all big things in tech. Thank you. Bye.
From the publisher
Dylan Fox is the founder and CEO of AssemblyAI, the voice AI infrastructure platform behind AI notetakers, medical scribes, drive through ordering, contact centers and humanoid robots. AssemblyAI went through Y Combinator's first AI batch in 2017 and now serves over 1 million developers with roughly 100 million API calls a day.
We get into the volume numbers, why weekly conversations are up over 800% in three years, the three macro trends pulling voice into everything, why coding agents turned small businesses into API customers, how data mixture drives 75% of model quality, the humanoid robot speaker problem, on device models, and the disclosure question the entire voice agent industry is avoiding.
"It's very easy to optimize for a public open source benchmark. It's very hard to optimize across all these real world applications."
We cover › 120 million weekly conversations and 4X YouTube's daily voice volume › Coding agents expanding the TAM by 100X › Context aware voice models and the McDonald's ordering problem › Why humanoid robots can't disambiguate speakers › Whether voice agents should disclose they are AI
Dylan Fox: https://x.com/YouveGotFox
Molly O’Shea: https://x.com/MollySOShea
Sourcery: https://x.com/sourceryy
𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊
YouTube: https://youtu.be/v6z_eoUX5LE
𝐒𝐏𝐎𝐍𝐒𝐎𝐑𝐒
• Brex—The modern finance platform, combining the world’s smartest corporate card with integrated expense management, banking, bill pay, & travel. https://brex.com/sourcery
• MongoDB–Millions of developers and more than 65,200+ customers across industries, including ~75% of the Fortune 100, rely on MongoDB for their most important applications. With integrated capabilities for operational data, search, real-time analytics, & AI-powered data retrieval, MongoDB helps organizations everywhere move faster, innovate more efficiently, & simplify complex architectures. https://mongodb.com/ai
• AssemblyAI–Millions of developers use AssemblyAI to power their voice ai applications and features. One API gives you access to best-in-class speech-to-text, voice agent, and speech understanding models for both pre-recorded and real-time audio. Granola, ClickUp & HeyGen are scaling with AssemblyAI - get $50 of free credits today at http://AssemblyAI.com/sourcery




