Building AI Voice Agents with Scott Stephenson - #707

28 Oct 2024 · 1 h 2 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The TWIML AI Podcast Episode #707: Building AI Voice Agents with Scott Stephenson

Episode Overview In this episode, host Sam Charrington interviews Scott Stephenson, co-founder and CEO of Deepgram, discussing the advancements and challenges in building AI voice agents. The discussion focuses on the integration of perception, understanding, and interaction in voice AI, the role of multimodal LLMs, and practical applications of AI voice agents.

Key Concepts and Discussions

  1. Evolution of AI Voice Technology
  2. Historical Context:
  3. Scott shares his background in particle physics and the development of Deepgram over the last nine years.
  4. Advances in audio models have been significant, especially with the introduction of models like Whisper, which improved accessibility to audio processing.
  • Current State:
  • The audio AI space has seen increased specialization and the emergence of numerous models focusing on audio processing.
  • The conversation highlights the shift toward a more multimodal approach where models can handle different forms of input (text, audio, etc.) more effectively.
  1. AI Voice Agents: Perception, Understanding, Interaction
  2. Key Components:
  3. Perception Models: Focus on the ability to recognize and process audio input.
  4. Understanding Models: Analyze and make sense of the recognized input.
  5. Interaction Models: Facilitate communication back to users, including audio output and user engagement.
  6. Scott emphasizes that these models should be viewed as a unified system rather than isolated components.
  1. Model Development and Challenges
  2. Current Gaps:
  3. Although advancements have been made, gaps remain in achieving generalizable models that can handle diverse acoustic environments and real-time interactions.
  4. Scott notes that models have improved considerably in accuracy, but challenges such as low-quality audio and background noise persist.
  • Future Directions:
  • The discussion touches on the potential for models to learn and adapt in real-time, reducing reliance on supervised learning and enhancing the user experience over time.
  1. Practical Applications of AI Voice Agents
  2. Use Cases:
  3. Scott outlines potential applications in various sectors, including:
  4. Healthcare: Assisting with patient care and information dissemination.
  5. Food Ordering: Streamlining processes in quick-service restaurants.
  6. Call Centers: Automating routine inquiries and improving customer interactions.
  • Deepgram's Voice Agent Toolkit:
  • Scott discusses the newly released agent toolkit from Deepgram which aims to empower developers to create voice AI products efficiently.
  • The toolkit focuses on enabling developers to leverage Deepgram's infrastructure without needing extensive technical expertise, allowing for integration into existing systems.
  1. Future of Voice Interaction
  2. Adoption and Resistance:
  3. Scott addresses skepticism around voice technology, suggesting that quality improvements will lead to broader acceptance.
  4. He emphasizes that voice interfaces will not replace text input but will coexist to enhance productivity.
  • Long-term Vision:
  • Future voice agents will be increasingly sophisticated, capable of handling complex interactions and learning from continuous use.
  • Scott expresses optimism for the evolution of AI agents and their integration into everyday tasks, enhancing productivity across various domains.

Conclusion The episode concludes with Scott expressing excitement about the future of AI voice agents, their potential to improve human-computer interactions, and the ongoing evolution of the technology landscape.

Key Takeaways

  • Unified Approach: AI voice agents should integrate perception, understanding, and interaction models into one cohesive framework.
  • Real-time Learning: Future models will need to adapt and learn in real-time from user interactions.
  • Expanding Applications: Voice technology is set to grow in various sectors, providing solutions to everyday communication and information dissemination challenges.
  • Coexistence of Modalities: Voice, text, and other modalities will complement each other rather than completely replace existing methods of interaction.

For more insights, listeners can refer to the [complete show notes](https://twimlai.com/go/707).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We shouldn't be talking about speech attacks models anymore, LLM models or TTS models. We should be talking about perception models, understanding and interaction models. And maybe they're all one thing, too. It depends on the goal that you're trying to accomplish. If you're doing perception, if you have an understanding step and then you're interacting with the world somehow, you're changing a database entry, you're generating audio and speaking back to a human or whatever, this is an agent.

0:37All right, everyone. Welcome to another episode of the Twimble AI Podcast. I am your host, Sam Charrington. Today, I'm joined by Scott Stevenson. Scott is co-founder and CEO of DeepGram. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Scott, welcome back to the pod. It's great to be here. It's great to have you back. It has been quite a while since we spoke, about seven years, if you can believe it. A very long time in AI years. Yeah, absolutely. A lot has changed in the last seven years. For sure, for sure. Just to give folks a quick refresher on your background, if I remember correctly, you came out of physics?

1:22Yeah, that's right. I was a particle physicist. I built deep underground dark matter detectors. um so essentially there's particles whizzing through our galaxy and we built detectors to try to capture them um and and that's called uh the what you're looking for is dark matter and then yeah it started deepgram about nine years ago with a few other physicists and now we're about a hundred person company that's serving audio ai models to over 400 b2b customers that's amazing when I think about the audio space and when we spoke back then you know there were a handful of folks for building specialized audio models and at least the model side of the equation has changed pretty significantly I think whisper kind of offered in a new era at least from my perspective in terms of ease of accessibility of audio, not fully eradicating the complexity problem, like played a lot with trying to get diarized audio with Whisper, and there's a lot of components involved in that, and the results can be challenging to get right.

2:33But even more recently than that, it seems like there are new models every day, Quinn and others specializing in audio and even the off-the-shelf multimodal models like Gemini can do a pretty decent job of transcription and even diarization. I'd love to have you riff on how that has played out from your perspective and what it means for a company that specializes in audio AI. Yeah, a lot has changed, but there's also a lot more to go. And so that's something to talk more about. But it's definitely true that more models have been focusing on the audio area. There was a boom over the last few years for text.

3:23But once you start to get a handle on that, then other folks look for other modalities. So, for instance, images, video, and audio. And I like to think about it from a perspective of typing, tapping, and talking. And, you know, we click on things, we tap things, you know, or we type or we talk. And this is the way that we interface with machines now. And it only just in the last couple of years started to become unlocked that the that the accuracy was good enough. The speed was good enough. The cost was good enough. The accessibility was good enough. And just also all around the world, people, developers and product builders are paying their tuition on how to build a voice application.

4:05And so all of this is happening at once. and we're starting to see a lot of players come into this space now. But yeah, it's very welcome from my perspective. When we started the company, we thought, hey, typing and tapping is taken care of, but talking isn't yet. And it's a major modality for human interaction with machines. Let's get it going. The hockey stick will come soon enough. And now it's here. Yeah, for sure, for sure. with all of the activity around models, certainly models are only one part of the equation, but do you feel like, I'm going to say, is it solved? I know the answer is no.

4:50What are the gaps, do you think, from a pure model perspective to getting where we want to be? A, from your perspective, how do you think of where we need to be from an audio perspective? And then what are the outstanding gaps? Yeah, the goalposts always change with time as they should with AI. So, you know, seven years ago, the accuracy just for a standard speech to text, don't even think about real time. Don't think about other languages. Just think English and think podcast audio, you know, so like really high quality, clean audio. The accuracy was not very good, you know, seven years ago. And it took a lot of development in order to just get that to work well.

5:35But now you start to add in confounding factors like look at low-quality audio. So like phone call audio or phone call audio with a lot of background noise or that type of thing. Then if you go back seven years, the accuracy was about 50%. It was like every other word was incorrect. So it was not very usable. You were experimenting with those models and you had to specify your channel. And the results varied pretty dramatically based on your channel. And the implication being that there needed to be a match between what was trained into that model and what you were using on. And I always suspected that part of the problems that I would run into is the gap between what my recording environment was and what those models expected.

6:22Yep, absolutely true. And so you want the models to get more and more general over time. And I think of this a little bit like a university or going to high school and then going to university. You want to create models that are well-versed in the world and can tackle many different problems. But then I think another analogy there is important that once somebody graduates university, they go and specialize in their job. And this is like the next phase of AI from my perspective, where you have really good general models, but then you allow them to adapt to their use cases. And so this can be agents, it can be just specifically like speech to text models getting used to an acoustic environment, but doing it in a way where it learns by itself.

7:05So it used to be that you had to use pure supervised learning in order to make these models better. That's not necessarily true anymore. You can generate synthetic data, you can use self-supervised techniques. And just going back to what you were talking about before about the progression of models. There used to be Caldi, which was a speech-to-text framework that had a very pipelined approach. And then there was Wave2Vec came out maybe four or five years ago, something like that, from Facebook. And that changed the game a lot because they started to incorporate self-supervised learning. And some of the similar things that we see for LLMs now where you don't necessarily have to have perfect labels for everything.

7:50You can still learn a lot. And then Whisper came along as well, utilizing some of those techniques, but then also improving the underlying model architecture. But this train is going to keep on moving. And at DeepGram, we have a research team that pushes the frontier here as well. And if you look at the accuracy and scalability and cost to implement, you know, then there's like a difference between open source type models and then B2B specific models as well. And this is where that like adaptation comes in as well. So it's not that all of your audio needs to be adapted to. Some of it's going to be handled by the model really well.

8:31But some of it you might want the adaptation for in order for the model to get better. And then especially adaptation when you're starting to think about accomplishing tasks from like an agent perspective where you're trying to shape, you know, the outcome and take in data and learn from that and get better over time. And so when you talk about that kind of adaptation, are you talking about very domain specific applications? You know, for example, if you want to enable, you know, doctors to be able to speak their patient notes or something like that, then clearly they're going to be using a lot of specialized language.

9:12You know, does it need to be that specific for you to need to have that kind of adaptation? Or, you know, are there more quote unquote pedestrian or B2B kinds of use cases where you also want to go there? I would think of it more like a modality and maybe use like interacting with a human as an example. um so for instance if you talk with another human and they pronounce your name a certain way like my last name is stevenson but it's pronounced s-t-e-p-h-e-n-s-o-n a lot of times people say stefenson you know if i wanted to i could correct them and say well it's actually stevenson not stefenson right um that that that's just a little like edit you know but when you're talking about a uh like an a call center agent like an ai agent if it keeps mispronouncing your name over and over like you're going to get irritated by that.

10:04So these little things that seem really tiny, but actually what it is, it used to be an open loop before, and now you can have a closed loop feedback right in the conversation. And so pronunciation is just one simple example, but if you're trying to solve a problem with an AI agent, you can direct it and say like, hey, I don't care about those things that you're asking me about over here. I care about these other things that I'm trying to solve? Can we focus on that and go down that path? And the answer to these models or for this next generation of models is that they will be able to do that.

10:41And I think that's probably one of the biggest fundamental shifts. The models of the past 10 years basically have been trained mostly from a supervised learning perspective. And they're trained on the like month to two years like timeline or something so they might be two years out of date one year out of day you know that maybe maybe a quarter if you're lucky um but that's not going to be the case in the future it'll be real-time updating and learning and um i i so i don't think of models necessarily necessarily as there's like one model um or 10 models there will be billions of models and you'll, some of them will, you know, fall off and never be used again.

11:32And some of them will continue to get better and better over time. And there'll be more specific to whatever you're trying to accomplish. Some of what you're saying suggests or calls to mind, you know, fine tuning and, you know, Laura adapters, maybe these are these billions of models. But that is typically not a real time closed loop automatic process. Like, uh you know close that gap for me you know is you know how is fine tuning used in audio what do you see today and how do you expect that to evolve yeah i think of it a lot like uh memory systems on a computer so you have your and you know back in the day you used to have tape and you'd put things on tape then you have like rust spinning discs that you can store things on and And actually, this hierarchy still stands today.

12:26Tape is the cheapest to store a lot, but it's the slowest, basically. And then, you know, rust spinning disks, you know, are the next cheapest. But they're a lot faster than getting tape, you know, seeking data on your tape. But if you want to seek very quickly, you'll use an SSD. And if you want to seek really, really quickly, you'll use RAM. If you want to do it like insane speeds, then you'll use the cache on the computer. And it'll be L3. that's your largest and slowest cache that's on your CPU. And then L2 is a little smaller, but even faster. And then L1 is a little smaller and even faster.

13:00And that's kind of where it ends. But that's like a huge range of big and slow to super fast, but small. And so it's going to be a very similar process that happens with building these models. And so that's from a speed perspective, but also the capability of changing and then how much it costs in order to run them as well. And so, for instance, when you're in a conversation, you might, at least today, how a system might be built, you'll carry along the context of the conversation in the prompt that is fed to all the different models. That will not be how it's done in the future. Okay, but at least right now, you'll just have this ever-increasing prompt for that conversation that you're having.

13:42And if your conversation goes on two hours, your prompt, everything will start to fall off the end of it. And it'll start to lose context over time, basically, right? Because conversation just keeps getting longer. That might be five minutes, might be 10 minutes, might be 20 minutes, etc. But certainly an hour or two hours into it, it's falling off the pipe for most systems today. That's okay in a lot of ways, because if you want to solve a problem with a call center or something, it might take only five minutes and that's okay, right? But again, that makes it seem like the model learned because it's holding everything in context.

14:16but actually the model weights didn't get updated. So if you call back, it'll be just as dumb as it was when you first started your conversation before, right? So you need to take this in context, like folks like to talk about in context learning. Some of that sounds like an engineering problem. You could build a system that like refreshes your context or some of it at the beginning of your conversation. And like, in fact, I don't know about you, but I've been griping about this for years. Like why when I call, you know, the telco and I put in my phone number, you know, the next agent asked me for my phone number.

14:53Like that's not a model problem or a technology problem. That is a, you know, a product problem. Yes. And well, what it all boils down to in those types of situations, though, is a cost problem. You know, how much does it cost to get it to do that thing? And so the call centers have determined that it costs too much money to give this person that information and have them learn about it, look you up, etc. Instead, they'll just dump them blind into the conversation and hope for the best, basically. But with AI agents, you're not going to have to do that anymore because the price to do an incremental amount of work or analyze an incremental piece of data, it's not going to take five minutes to do it.

15:36It's not going to be really expensive. you'll be able to cram more context into it. But yeah, you're right that a lot of this is engineering. But accessing the different echelons is not engineering. So if you want to move away from just providing bigger and bigger context and instead having the model actually learn over time from lots of people using it, then you basically have kind of a federated learning problem. You have 10 ,000 conversations going on at once. everything stemmed from this one AI brain agent, you know, and it broke out into these 10 ,000 different conversations. How do you include what you learned from all of them back into the same model so that when you fire up a new, you know, your 10 ,001 conversation that includes all of the other information in it, or at least some of it doesn't have to be perfect.

16:23Right. But like over time, you just want it to get better. And this is really the promise of intelligence. You know, Truly, the intelligence revolution is that we don't have to think about machine learning systems as only working in this supervised learning kind of way, like offline trained, etc. That's going to be the intelligence revolution 0.5 or whatever. The next version will be it can update and learn throughout time. And then, yeah, so the way that that works is, you know, it's a frontier research stage now. But what you do is you bite off pieces, you know. So can you focus on some of the most important things and then be sure that that gets learned in the model?

17:09And forget about all the details for now, right? And then you say, can we expand that, you know? And so you test it over time and, you know, you start to explore your territory more and more. And this is what all of the foundational research teams like DeepGram, like OpenAI, like Anthropic, et cetera, this is what they do all day, every day. Are you seeing any indications or what indications are you seeing of that future, you know, starting to come to pass? Like, you know, there's definitely federated learning types of use cases that I've seen that are, I think, you know, much more limited than the world you're describing.

17:50There's like research into like continuous learning, but doesn't typically have like that federated aspect to it. Like, are you seeing any movement towards this vision that you've outlaid, laid out? yeah but it's kind of it's it's how much of it is like you know you can see it and like grasp at it yeah in ai terms it moves at a glacial pace but that still is like two years from now it'll definitely be here um so so there's like these so you know there's there are these fast moving things that are being iterated on with uh like architectures that people already know and and you know and like modalities that people already know and that happens at a really fast pace um But when you want to stretch and move into a new way of training the models that isn't, that the papers haven't been written about them yet, you know, there aren't examples out there in the world all over the place, then that moves at a relatively slower pace.

18:53And the reason for that is you might have 30 ideas on how it's going to work, and only one of them is actually going to work. And you need to go do all 30 experiments, though, and they need to be on a large enough scale to tell you if the model will actually work. So they take a long time to train, they take a lot of data labeling, etc., in order to figure it out. But once you get something, you're like, oh, this is going to work. So for instance, end-to-end deep learning for speech, that's one. Transformers for text, that's another. And then all the individual components as well, like residual networks and stuff, all these little things build up to providing the structure for how everything is going to be built later.

19:35And then once you have enough proof points and you have enough marketing around those proof points, then everybody jumps on them. Right. And so I expect it to be this fairly slow burn in AI terms, you know, over the next two years, basically. But then it'll be like, oh, it happened overnight. But really what was happening is behind the scenes, everybody was like figuring out, like beating down the uncertainties, basically. Yeah. So maybe going back a little bit to this idea of kind of the model, the core models evolving quickly, you mentioned that you do have an internal research team. That said, are you finding that you're investing as much or more or less into core models?

20:19or are you finding that core speech-to-text diarization? Is that commoditized and you're fine using something off the shelf even though it's your business? Yeah, it's the first of those two. So we're seeing that there are maybe... I'm trying to do the combinatorics here. There are several different ways that you can build an AI company. So you can be an AI company that has a consumer product that builds your own models and you do the infrastructure for it all as well. You can be an AI company that only produces models. You can be an AI company that only does infrastructure. You could be an AI company that only produces apps.

21:09And you could do any blend of those two. So you could, and DeepGram is in the category of, we are a model builder, foundational model builder, and we provide the infrastructure. And then there are other companies that do different things. But what we notice is that, especially in the voice world, it's very complicated. Just audio and video is very complicated in general. And if you only provide the infrastructure, you will notice so many holes all over the place. So for instance, I can give you an example, like in building voice AI agents now, how do you know when somebody's done saying what they want to say, and now it's your turn to talk?

21:56It's a very difficult problem, this turn-taking ability, right? But what signal are you going to get from just standard speech to text that tells you they're done talking? It's a loose signal, like a period, like you don't really know, Right. It's and and also how long are you going to wait, you know, until you say, OK, I'm pretty sure they're done talking, because one thing you could do is like look for silence. Right. You could say, well, they don't say anything for two seconds. It's probably my turn to talk. Right. That misses a lot of context that you might be able to take advantage of if you're actually have intelligence and part of that, you know, have a model there.

22:30Exactly. So if you're a foundational model builder, you can look at that problem when when people are trying to build voice AI agents and say, hang on a second. this really should live in a perception layer and that perception layer should have speech to text yes but it should also have diarization it should also figure out topics it should also figure out you know if it's if it's a real-time model it should figure out if it's somebody else's turn to speak you know um etc and um so just this this loop this this uh like at least for voice the the loop goes something like this there's a perception understanding and interaction and so the perception is everything i kind of listed there there's a whole bunch more too um but then the understanding would be like, okay, now that you just wrote down all the facts and told me the facts, like, what do I do with them now?

23:11Right. And so should I say something? Should I go do something like send an email? Should I, you know, whatever it is. Right. And then if you determine that you should say something, then, okay, here are the things that I want to say, but it's not just, here are the words. Here's how I want to say it. You know, do I want to say it in a somber tone slowly? You know, that type of thing. This is all context, right. That needs to be passed along. And so, like right now, if you wanted to train or sorry, build a voice AI agent system that truly just felt like a human, like you really couldn't tell the difference, you just couldn't with current architectures.

23:44You know, you have to build models or or build new capabilities into models and essentially blur the lines between the models over time. So if you're only an infrastructure provider, then you're always going to be behind, essentially. And if you're only a model provider, then you have to go convince all the infrastructure, you know, people to use your models. And they won't exactly know how to use your models because they're trying to serve a large audience. And so instead, you just, you know, this is what we choose to do. We blend it together. We build the models and we host the infrastructure as well.

24:15And so if we need to build our own infrastructure from scratch, like we have our own data centers. And if we need to do that in order to serve a certain type of model, then we will. But what we're looking at is the use case. So going back to that 20 ,000 seat call center, that's a serious amount of compute that is needed. And you can't just stand up the current models of today and hope for the best, basically. So we blend those together as a company. And I think it's a really strong way to build because you get the feedback from your customers, you know, and then you could put it straight into the model building.

24:55But, you know, there are downsides to it too. When you have a consumer app, like OpenAI and Anthropic do a great job with this. They have a consumer app that drives a lot of awareness for their company. And they also learn a lot from that as well, right? So, you know, you have to sort of choose your battles though. Yeah. One thing that called to mind was I mentioned, you know, out of personal experience, the complexity of diarization. I'm imagining that if you are building a model and train it to like predict turn taking, there is a loop that you can close with regard to diarization to like refine that or to refine both.

25:39Absolutely. And diarization, you know, as it is typically implemented separate from the speech to text is, you know, it's challenging. Yeah, weird. Yeah, right. Yeah. I'll tell you the reason that diarization isn't great in the world right now, and it just isn't. Any open source thing, even from DeepGram, you know, our diarization is world class, but it's not great. But I'll tell you why. People don't like to pay a lot of money for it. it's just true and and so what happens is folks are willing to pay a lot of money for high accuracy speech of text for low latency real-time speech of text for a high quality tts voice you know that type of thing and then the diarization use cases folks are a lot less to you know spend in that area or uh app a less you know they're less apt to spend in that area um and so what that that it gives a signal to all of the model builders and infrastructure builders like, Hey, it should work, but like, don't spend too much time on it because all these other things we care about more.

26:43But you know, call, is that because the, is that because the UI to correct it is relatively simple compared to like, you know, correcting word for word accuracy or just the use cases aren't there. Yeah. Just to use cases and, and the dollars behind them. And so you have to just look at like the low hanging fruit, basically. Um, and diarization is complicated, but I can tell you, like if our research team just said, we're going to solve diarization, like it wouldn't take too long to do it, but we're also working on all these other things, you know? So that's the, that's the real trouble is that, uh, everything just over time, things tend to get prioritized more, you know, than diarization up until now.

27:25Um, although, but what happens is over time, uh, the model companies and infrastructure companies figure out better and better tooling for themselves. So it's cheaper for them to produce models over time. And so, you know, as you get larger as a company to like DeepGram has grown immensely in the last few years, then you can focus on more things as well. One thing, though, that we've chosen to do is focus on adding TTS to our lineup, adding audio intelligence to our lineup, and then adding an entire voice AI agent to our lineup. But, you know, next on the list, you know, so far, because I can't, I can't, exactly, exactly.

28:04But I can't see an expansion beyond those because that sort of encapsulates all of, you know, the, the low hanging fruit for audio. So I want to dig deeper into the agent stuff. We've kind of, you know, skimmed, uh, across it a couple of times here, but before we do that, you mentioned the, uh, TTS and that reminded me of some of the work that i saw you do with grok um you know some months ago that uh you made a comment earlier about how like you know we're not at like this real-time um interactions where you know that kind of passed the audio touring test if that's a thing but like it it does seem like we're getting there with um you know you guys are approaching real time you Talk a little bit about your capability in that area and closing that loop.

29:01Yeah. So you need to get the perception, right? And you need to do it in a very short period of time. And so that typically takes like 250 milliseconds or less. That's like your time budget to get your perception right. Then you need to get your understanding and then audio generation right as well. and your time budget might be 250 milliseconds for all of that, or maybe, you know, up to 500. So like three quarters of a second for everything, essentially, right? That's still a little too long. But nevertheless, that'll still feel pretty human-like. And that's especially true if you have the types of end-of-thought or end-of-turn models that are starting to predict that this next...

Read the full transcript

29:47Like, it sounds like what they're doing is wrapping up, and it's probably going to be my turn soon, as long as these next few words keep, you know, coming in and at the spot that they that I think they will, then I can already predict that I should I can fire out and say, hey, start thinking about the next thing to say. And this is like what we do as humans, again, like using using humans as a model, gives you gives you a good idea of like, where you should go with building these underlying models as well. But so so yeah, you you have this like, maybe 500 to 750 millisecond total budget to respond.

30:18And I can safely say that at least now with reasonable spend, so not$100 an hour, not$10 an hour, but$1 to$5 an hour in that range, you can hit that mark. And so that's starting to get around like what a call center agent in the Philippines or India would make. It's about$2 to$5 an hour is typically what they would make. And so this is where it starts to get interesting. If the cost is similar, then you might start to use them to take some of the lower hanging fruit in the beginning of the call and then pass it along with context to others later. I can tell you, it's already solved in some ways.

31:01But again, that Turing test of am I 100 % fooled? There's no way I could tell that this wasn't a human or not. That's not true. You can definitely trip it up. You can try to get it to pronounce weird things to pronounce. You can try to take it off track, et cetera. And it's acceptable to some of those things now, but it'll be, you know, the next two years, it'll get very robust to that type of thing. And so the, you know, we've, we've talked about speech to text and text to speech. Is, is that the only way to do it? Like, it strikes me that, you know, text is a bit of a bottleneck there and for you know some of what we're trying to do like you know as the model quote-unquote understands what's being said it could generate without having to go through that you know at least what i'm thinking of as a kind of limited bandwidth channel of text like is that the right way to think about it or great way to think about it yeah yeah you're essentially flattening it down and like trimming off all these uh ends of a bunch of different distributions and just like, you know, forcing it down to this one text representation.

32:09You know, a good way to visualize that is just think of any sentence, a simple sentence or, you know, hello world or whatever it is. It can be said by young, old, you know, any type of background, dialect. It could be said far away from the microphone. It could be said close to a microphone. It could be said with all types of background noise, et cetera. But if you were to transcribe it, it would all be transcribed to the same thing, right? And so you can already see that there's like this massive underlying in audio, there's this massive underlying like expansion of possibilities to quote unquote mean the same thing.

32:48But when we listen to it as a human, we know that they don't mean the same thing. If you hear somebody in a car that's telling you something about the context, right? If you hear excitement in their voice, that's telling you something about it, et cetera, right? So you're throwing away all of this information when you cast it down into just text. So you're 100 % right. However, that is the best way to first build a system. And one of the reasons for that is humans need something to debug. They're kind of going through the system and saying, okay, I need to hook up a real-time, something that works in real-time in order to get it to do what I want.

33:25So at least historically, that has been speech-to-text up until now. I'm just going to take audio. I'm going to convert it into words. and then I'm going to hook it up to an LLM and then I'm going to hook it up to TTS. But once people get sort of bored and sick of that and they pay their tuition on it, they're going to start saying, hey, this thing is losing context. How do I get context across these boundaries, basically? And there are ways to do that through like audio intelligence models that can help you do that. But the way that I would look at it is we shouldn't be talking about speech-to-text models anymore, LLM models or TTS models.

34:00we should be talking about perception models, understanding and interaction models. And maybe they're all one thing too. It depends on the, it depends on the goal that you're trying to accomplish. So if you don't need it, so for instance, if, if you have really challenging audio, you might need to beef up your perception model to understand what is happening inside it. But if all it needs to do is make a determination about like this category or that court, a category, you know then you can have a really simple understanding model and then the tts portion of it could just be like a single voice it doesn't need to uh replicate all human voices ever you know it doesn't have to be a super complicated model that could that is capable of doing that it could just be flattened down into one voice and speak back to you in a normal way that you that you have chosen it doesn't have to generate all sorts of other background noises or anything like that so you can simplify your output stage you can simplify your understanding stage but then you need like a you know a a more beefed up um input stage basically there may be times where you want something that is totally different um like when the audio is less challenging um and you need it to think really hard you know and you want it to be to speak in any language you know do translation do all sorts of things then you'll beef up the last half of that and so this is one of the reasons that you if you just if you just think and say well there's going to be one model to rule them all.

35:18Well, maybe there's a way to do that, but it'll cost so much money relative to what it could cost. And this is why from like a B2B lens, it makes so much sense to think about models that cross the boundaries or think small, medium, large for certain parts of the stack, and then the ability to adapt each of them to your use case. And then you deploy it at scale. Because in the end of the day, like this is what all that I'm also always thinking about with our customers is they got to scale this up to 10 ,000 concurrent connections. And scaling up like a 100 billion parameter model that is sampling every 80 milliseconds in a conversation, it's going to cost too much for them to do that.

35:59So you find other ways in order to make that work. Although I think in the future, like meaning five years from now, 10 years from now, there will be adaptable, controllable, truly speech-to-speech models. And this is part of the research that were doing at DeepGram as well. But in the meantime, you want the controllability with the tools that you have on hand as well. And you need the ability to debug it as well. So I feel like over the next two to five years, you're going to see a massive deployment of what I would call Voice Agent 1.0, where it is mostly speech-to-text, LLMs, maybe RAG, or something like it, and then TTS.

36:39Because it's a very understandable system for many people to deploy. And then there will be a new revolution on top of that a few years afterward, where you have widespread, you know, voice agent 2.0. And that'll be a very contextual AI based modality. Yeah, yeah. It's an interesting way to think about the progression and that like so much of the challenge of deploying, you know, LLM's deep models now is like interpretability. like are we even ready to lose the interpretability and predictability of putting everything to text first um it's a good you know you know forcing function for being able to introspect the model and then um like i could totally see like how the modular approach is a great next step uh after that but like more integrated where you're not forcing everything through text and then you know you apply the this idea that like some of the greatest gains we get are when everything is trained end-to-end and so like what does that mean when you like end-to-end train these different modules um you know but that you know sounds like you're saying at least five to ten years away.

37:57Well, I think there will be specialized use cases where I'm thinking five to ten years away. Simplifications. Yeah, I'm thinking five to ten years away for everything working that way. But no, it could certainly be six months from now and specific use cases are trained that way, but they will not be controllable in the way that... I think it's tempting for folks to say, oh, that means I'll be able to use that in any other instance. And it's like, no, actually, they're going to be very trained to do specifically what they're going to do. So for instance, Google for a long time has valued voice as a channel.

38:34And a lot of folks don't necessarily think about it this way, but this is the reason that they created their speech team at Google was they're looking around saying, hey, how can we be attacked in the world? you know, how can we be attacked? And who's going to come take our ad dollars, basically, and how could they possibly do it? And they've zoomed out into that typing, tapping and talking way of thinking about the world and said, well, we've got we've got typing and tapping like pretty, pretty well down people come to us default through text. But the talking, actually, this is a really big modality, like people spend most a lot of their day talking.

39:16And so if there is an interactive voice way to interact with a search engine, then that's a really big problem for us if we're not the leaders. And so this is why like in 2008, they started investing in it and then, you know, pushed out along, et cetera. And actually it's, it's, it's funny when you look in the designs of these other consumer products, like, like, like opening eye chat, GPT, et cetera, you can tell they, you know, just squint a little bit and think about it. It's like, oh, they're doing a very similar thing. You know, they, they know that if somebody else releases a really good voice agent, then all of that traffic is going to be pulled over to them.

39:51So now they have to do it as a defensive mechanism. But one of the keys here is they're not building it for all of everybody else's use cases. They're building it for themselves. And then they're saying, hey, maybe we can sell this to other people. But then they do a common trick that the hyperscalers do, which is they do a lot of marketing to say this is the only way. um and i just again you don't have to take my word on it for like uh today's technology just look at you know 10 years ago technology this is exactly what uh google and microsoft and others were saying then um but they're not the ones who are winning the ai revolution now uh in voice or text or any of these other ones they were just trying to quench any of the startups from coming and attacking them and now they're getting attacked you know so um that that's just that's like what you're seeing now yeah yeah it's interesting you know i'll often read people um i don't know this is like maybe hacker news comments or reddit comments and you know people will be talking about like voice search or you know voice aided development or any number of things and we'll be like yeah i don't really want that like i can't stand talking to my machines whatever and i can't help but think that's just because it's really bad right now but like do you would you not want to i don't know i guess i should speak for myself like if i could you know have an ide where i can you know reliably say yeah where it says this change it to that like uh and then have the lm go off and do its thing and like oh i think you should maybe supplement at this like yeah make that change like i would love that and you know i'm not you know the best typist like there are faster typers and some typists and maybe that's part of it but like i i guess my point is that i think there's a lot of resistance to voice because it's not really that great yet.

41:59But it's good enough that compared to what we spoke seven years ago, you can see the promise. Yeah. And in a very short period of time, due to synthetic data, due to everything just stacking on top of each other and the ability to make these better, faster with time, it's not going to be a long time until some people are very happy about it. But it does make sense that there are amazing engineers that are extremely deeply contextualized and talented in the thing that they're doing. And if some other person tried to talk to an LLM to get it to do the same thing that that person is doing, that's 100 % true.

42:43It's not going to happen. 100 % true. It's not going to happen. Right now, at least. Maybe in five years or something. But right now, that's not what's going to happen. I think what you see most of the adoption now in people who are kind of in between, they might know how to code a little bit. They might know some HTML. They might know how to edit some CSS. They maybe have done a Python script data analysis thing five years ago or something like that. And now they're thinking, I've always wanted to know how to code and actually build things for real. But where do they start? You know, oh, well, go back to college for four years and do the hard road and all that.

43:18It's like, no, that is not how it's going to happen in the future. Now, yes, you can just type or talk to one of these assistance agents and just build it. And, you know, I have plenty of examples internally in DeepGram. A lot of our company is not technical, but a lot of them now use these tools in order to build things to make automation in the company so much better, etc. And, you know, we haven't we're a hundred person company, but we're like punching well above our weight, partially because of that, because of the automations that folks have done internally. and yes it's engineers and it's researchers doing that but it's also just you know like our like our ops people and they didn't know how to code before and now they do because they've learned through these mechanisms and I think that will help coding was just an example though of that like I think the broader point was that like you know just like coding like document editing via voice or like you know command line via voice or hey shoot off this email via voice like I guess I wanted to get your sense of, do you think like voice will always be like the redheaded stepchild of like human computer interaction or inefficient or will it become like the primary mechanism?

44:39I think it will become a favorite. Are there inherent limitations? You know, that kind of thing. There are inherent limitations, though. So like, for instance, if you get on the subway, you're not going to be, you know, checking your email through voice or doing base bank transactions through voice and, you know, that type of thing. Audio is limited in another way, which is it's basically single threaded, you know, so you need to be in a situation where that makes sense, like driving a car, like, you know, you're in a quiet room and you don't mind talking to yourself or talking to a computer or something like that.

45:10So I wouldn't look at it like it's going to take over all interactions ever, and this is all we're ever going to do. I don't think that's true. I think we're still going to be typing, tapping, and talking. But I think they will be pretty equally weighted. So there will be times where the bulk of what you produced was produced just by you talking to some digital system, whereas before it may have taken you 10 hours to produce that. It might take you an hour or two just by talking to the system, and this is in the next year or two. And so then people are going to start to really like that. And I can tell you, I already do it like in hiring for executive roles at DeepGram.

45:47I used to write down my notes afterward. And now I focus on the conversation that I'm having with the folks. And then right afterward, I schedule time to just talk, you know, and say, here's what I thought about the person. Here's what they said. And by the way, here are some of the nuances. And by the way, those nuances, I would be, you know, personally too lazy to write down, you know, before, because like, I'm not going to write down all these things, right? It's too crazy. But if you, but it takes almost no effort to speak it. So you can just say all these things and then allow the AI to make the notes and pass them along to others.

46:21And so it's, I see it cropping up all over the place. And maybe I don't know if I should sneak peek, say this or not, if our marketing team or product team will get upset about it. I don't know. But we recently acquired a company in the application space called Poised that helps with folks who do meetings and they would like to make their pacing better, say um and ah and all that type of thing less and just get better at speaking. And one of the core challenges to that team now is to add functionality a lot like what you're talking about into this application, because we see opportunities all over the place.

47:03And I can tell you what I'm talking about, what I'm speaking to and having it write my notes for me, it is that application. We already have it working internally inside DeepGram. And we've already seen the aha moments. And so I think it's coming. I think it's going to become a favorite, but it's not going to become the end all be all. Everything's going to have to work in concert, basically. Screens are still going to be a thing. You're still going to be typing and poking and that type of thing. Text is still going to be a thing. But voice was like, voice used to be the original thing. It was telegraphed for a while, but then it was voice for everything for a long time.

47:37But then it was too complicated to make it work with the machines. And so text and digital stuff came along. But voice is back. It's back again. All right. So agents, sounds like something that you're working on and are excited about. So before we go any further, I want to ask you to define that because you're at this nexus of call center agents, which has one implication, and AI agents, which has another technical implication. When you think about agents, are you referring to something on the one side or the other side? I imagine, at least over time, a blending of both. But how do you think about that?

48:20Yeah, how we define it internally is that full loop, the perception, understanding, and interaction. And so if you're doing perception, if you have an understanding step, and then you're interacting with the world somehow, you're changing a database entry, you're generating audio and speaking back to a human or whatever, this is an agent. Now, the different capabilities of the agents and their integrations and everything are all very different, but this is how we think about it. And so what are you seeing and kind of building to kind of thinking about that, you know, that loop being closed? Yeah.

48:55So there are going to be many agents built and there will be many voice agents. There will be many non-voice agents built. We are not trying to build all of them or make the platform to build all of them. The voice agent API that we just released, we're not intending for it to be the only voice AI agent or API in the world. Really, it's a demonstration of what's possible. And in many use cases, it will be enough to do the job. And we anticipate that and we already see the growth in that area. But what we really do is we build the tools for folks to, you know, for developers to build products, voice products and, you know, the cutting edge voice products.

49:36And so we're still hyper-focused on that, but the agents push us as well. When we're building our own agents, that pushes us. This is similar to the analogy that I was talking about before between building models and building infrastructure. When you're forced to build your own voice AI agent, then you're forced to reconsider what your models are doing too before your customers ever even ask you or it ever even bubbles up to you that this is a thing, it's a concern or a problem or an evolving market. it. So for us, we look at it that way. One of the best ways to think about it is DeepGram, our North Star of North Stars is to increase the productivity of the world.

50:15And also, we just think of ourselves as a learning company. And so the productivity of the world, I'm thinking about that from the economic sense. When you have revolutions happen, technological revolutions, The productivity of the world goes up. That gets spread over time to all people. And that is our goal to do that. But also, we are a learning company. We build models that learn. Our humans learn. We are educating our customers and helping our customers learn, et cetera. And so from this mindset, we're building the AI agent to learn ourselves to be a durable product over time as well. But it helps us make better fundamental components as well that fit into other systems that other people are building.

51:00Our voice AI agent is not going to fit for everything, but our perception layer probably will. Our understanding layer probably will. And our interaction layer probably will. It sounds like it is kind of a framework that is intended to, in a sense, demonstrate capability and inspire as opposed to necessarily be, you know, this is the tool that you need to use to build the thing you're trying to build. Ultimately, the perception, understanding, and interaction, those are the things that you're trying, those are your offerings. And this agent toolkit is a way to tie them all together so folks know how to use them to create these experiences.

51:43Yes, exactly. It does both. But I would say look at like ChatGPT as well. This is an application. And our voice AI agent is not an application. It's an API. But nevertheless, it serves two functions, right? It makes money. For an average everyday consumer, they can pay$10,$20 a month, et cetera. It does that, but it also educates them about what's possible. Okay. And so what are some use cases that either folks have built using it or that you envision folks building using it that illustrate its capability or what's possible with these components? Healthcare is going insane right now from an AI perspective.

52:21They were like, they, you know, they're so understaffed over so many areas they have there and they finally have figured out sort of their regulation side of it and like what they care about and how to make it happen. Veterinarian, you know, just the healthcare world in general is starting to build a lot around this just for just to aid in patient care, not even necessarily to deliver the care, right, but just to help educate around like, okay, can we set appointments for you to come in? Can we educate you about what you need when you come in? After the appointment, can we educate you about how, you know, we're just going to restate and educate you about like how you should be taking care of yourself, et cetera, that type of thing.

52:59And so that's an obvious, really large, fast-growing use case. Food ordering is another one that is taking off with quick serve restaurants, you know, fast food restaurants. And we have some announcements that will come out later this year about this. But there have been some announcements in the last couple of years saying it doesn't work. It works, okay? These systems work. And it's a really big cost to just take the orders at these restaurants. And so think billion dollars in revenue a year type scale market. It's not an absolutely huge market, but it is a big market, just that one thing. um and so we see but you know honestly we see uh uh vertical specific voice ai agents pop up all all over the place now if you just look at the last like yc batch or the last two yc batches there are so many different uh applications um that are coming up i i really look at every text box as at the very like if you just want to boil it down to its minimum you know just every text box should have some kind of voice capability along with it but that doesn't take into consideration into consideration all the things that you would do if you didn't have to go through a text box and so those are building up as well but yeah can you kind of go to the next level of detail in terms of the what this agent framework does you mentioned previously like some of the more technical side of things like rag and some other things like does it incorporate all of that or is it intended to be used with a line chain or something else that does the rag bits and you're doing the speech bits like.

54:42Yep. Yep. It is a full system. So you would, you would specify what you need from our system. So you put in the prompts, you put in, you know, a reference documentation, you know, that type of thing. And then you use our, our systems in order to do it. However, you can choose like which LLM you want to use and, and that type of thing. So, but, but then we are the one running the infrastructure in order to keep the latency low, serve at, serve at high availability, et cetera. So it's more an infrastructure offering than necessarily a development framework. Is that the idea? Yes. Yes, exactly. Both of those.

55:18Yeah. With, you know, constantly improving models being injected into it, you know, to improve like end of thought and, you know, that type of thing. Yeah. And is the development experience like, you know, Python developer, they're importing libraries and they produce some artifact that's deployed to your system? Yeah, exactly. So we provide SDKs and we provide, if you want to just start the WebSockets yourself as well, we provide the documentation for that. But we also provide SDKs in several languages. And soon we're partnering with one of the largest developer-focused communications platforms.

56:04I don't, I think we just did marketing with them as Twilio. But, but anyway, but, but I think we have some more announcements coming out, but partnering with them to allow you to get a, a phone number and all of it. So like, you know, it can be as easy or as complicated as you want, basically. When you say you provide the infrastructure, like I need to write an app, I'm calling your API using my app. And then I need to like create some kind of artifact, like a container, like. Does the SDK do that? Or am I using something else, deploying it to like a Vercel or like a Docker DigitalOcean or something?

56:43Or am I deploying that to you and you're running that part of it also? Yeah, you are injecting the DNA. You're injecting the prompt or relevant documents or that type of thing and some of the control structure. But nothing, at least for the agent, nothing is running on your infrastructure. I say that, but like the deep gram also runs on-prem. So if you'd like the whole thing to run on-prem, you could do that. But, but really that's not the design philosophy. The design philosophy is that all of the, like the parts, you know, Hey, how long do I wait until after somebody said this before I interject?

57:19And what if they interject again? And, you know, all that, you don't have to figure that stuff out. That's all in the, in the framework and it's hosted by us. And so the, really what you need is you just need a, you need audio coming in and audio coming out so it's just a fully duplex web socket is how you connect to the to the system and then everything else uh is is handled um by deep gram um of course to your specifications but yeah okay and so is it more of a like a no code low code kind of thing like i'm putting prompts into text boxes and not like i'm creating some python package or something yes i mean if you want it to be that way.

58:02But for instance, if you create an agent that answers the phone, but you know the number that's calling, you might inject the CRM information into that prompt and say, oh, I already know their name. I already know where they're calling from. I already know their order history, that type of thing. But if you want to do that, then that's typically like you're coding that. So when you initiate that WebSocket connection, you will include all of that that payload that has that information with it. And then the agent will go out and carry the task with that, carry out that task with that information in mind.

58:35And does that part look like a typical tool use, tool use kind of model, meaning you tell it like you have access to this, like customer lookup tool and then the agent will, if it thinks it needs that kind of call out to that tool or something. Yes. With, you know, since there's like infinite tools, you know, with limited capability there, right. You can call out to a few things right now, and then we expand that capability over time. Mostly just because you want to, it's actually not really because we can't support it. It's that you want to make sure that your call doesn't hang for like 10 minutes while you're waiting for a search lookup on some other system.

59:17Our system is fast and ready to go, but you don't want to wait on this other thing, etc. So anyway, we're building more integration and relationships and partnerships there in order to make that possible. Right now it's fairly limited. But so I would think, you know, like a really good use case is, hey, I have a store after hours and or I service a whole bunch of stores that after hours I would like when somebody calls our local number instead of it just going to voicemail or nothing happening. Instead, pick up, have a nice agent, answer the phone and give them information about the store and what the store does and what the hours are and all of that stuff, you know.

59:54And so for the 12 hours of the day or the 16 hours of the day that we're not there, you can still have somebody answer the phone basically. And then you expand from that type of thinking. So you think low hanging fruit and then expand from that and expand from that. I wouldn't go straight to your most complicated. I want a lawyer on the phone to answer. Just think smaller, think simpler, think low complexity. And the fewer integrations, the better, or at least the higher the probability of the success basically. awesome awesome very cool well Scott it has been wonderful catching up we need to make sure not to let it go another seven years but thanks so much for sharing a bit about what you've been up to yeah absolutely it was great to hang out absolutely

1:01:04Thank you.

From the publisher

Today, we're joined by Scott Stephenson, co-founder and CEO of Deepgram to discuss voice AI agents. We explore the importance of perception, understanding, and interaction and how these key components work together in building intelligent AI voice agents. We discuss the role of multimodal LLMs as well as speech-to-text and text-to-speech models in building AI voice agents, and dig into the benefits and limitations of text-based approaches to voice interactions. We dig into what’s required to deliver real-time voice interactions and the promise of closed-loop, continuously improving, federated learning agents. Finally, Scott shares practical applications of AI voice agents at Deepgram and provides an overview of their newly released agent toolkit.

The complete show notes for this episode can be found at https://twimlai.com/go/707.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Building AI Voice Agents with Scott Stephenson - #707The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 1 h 2 min
Listen in VO