BONUS: Building Real-Time AI Voice Agents with LiveKit's Ben Cherry

22 May 2026 · 1 h 9 min · 33 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

LiveKit (Ben Cherry) explains how to build real-time AI voice agents using WebRTC infrastructure, compares pipeline vs natively multimodal real-time models, and demonstrates coding, debugging, and voice cloning in LiveKit Cloud.

Guest backgrounds

Ben Cherry is from LiveKit, an open-source framework/cloud platform for voice, video, and physical AI agents. LiveKit began with WebRTC streaming/video apps and evolved into a developer platform for real-time multimodal AI systems. Corey and Grant host Neuron Live.

Key claims

  • LiveKit makes it easy to “get on the phone with an AI” via persistent, real-time connections.
  • Most scalable voice agents use a cascaded pipeline (STT → LLM → TTS) for reliability/tool calling; real-time multimodal models can be more expressive but may produce mismatched transcripts.
  • WebRTC is more reliable than WebSockets under network congestion (priority: stay real-time).
  • Voice agents converge with long-running agentic systems (persistent programs, long-term actions).

Notable examples

  • Starter demo: a joke-telling voice assistant (XAI real-time model).
  • Healthcare demo: “Maya from Linden Dermatology” collects skin details before a doctor.
  • “Evil twin” concept: instant clone of voice/face for a talking avatar.
  • LiveKit Cloud observability: session replay, traces, logs showing whether tools were called.
  • Voice cloning preview: upload consented voice script; clone across providers (e.g., Cartesia, ElevenLabs, Inworld).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Overview of LiveKit

0:45 to 2:00

Discussion about LiveKit's platform and its evolution into real-time AI systems.

“And this is especially timely because LiveKit's been shipping a lot around voice agents of late, agent building, debugging, telephony, observability, all of the things.”

Introduction to Voice AI

2:00 to 3:45

Ben Cherry explains how LiveKit facilitates the building of voice AI agents.

“I don't know, 15 years ago or something, Siri came out.”

Evolution of Voice AI Technology

3:45 to 5:44

Ben discusses the evolution of voice AI since ChatGPT and its integration with LiveKit.

“We also have a lot of support for things like you can screen share to them, you can send video to them, whatever really the developer wants to do.”

Voice as Accessibility Technology

5:44 to 7:44

Exploration of how voice technology enhances accessibility for computing users.

“But they're not, you know, they don't know how to type.”

Technical Functionality of Voice AI

7:44 to 10:11

Discussion on how real-time voice AI processes inputs and outputs effectively.

“And I think it's going to be really, really transformative.”

Architecture of Voice Agents

10:11 to 12:00

Delving into the architecture and layers involved in modern voice agent systems.

“But with the audio side as well, I mean, does it work the same way, like where you're streaming the audio that you're saying to the audio model and then it gets it back?”

Future of Voice Agents and AI

12:00 to 14:01

Ben predicts the future direction of voice agents and their integration with AI.

“It can express that directly without having to go through an intermediary step where it transforms it into emotion tags.”

The Evolution of Persistent Voice Agents

14:01 to 17:19

Explore how voice agents are evolving to run persistently and interactively in real-time.

“It's kind of feels a little, a little more old school and actually pretty cool, but it's a Python program that just runs.”

Voice Technology Integration in AI

17:20 to 18:20

Learn about the technical aspects of integrating voice technology in AI applications.

“How did you're working with Grok more now?”

WebRTC vs WebSockets in Voice Tech

18:21 to 20:45

Understand the differences between WebRTC and WebSockets for robust voice applications.

“And I think that the thing to know here is when Grok VoiceBud originally launched, they launched with a WebSockets-based implementation.”
Show all 33 chapters

Real-Time Voice Challenges and Solutions

20:46 to 21:51

Discuss the challenges of maintaining real-time voice quality and potential solutions.

“But if I'm listening to a podcast or something, I'll get chops sometimes.”

Building Voice Agents Using LiveKit

21:52 to 24:19

Learn how to set up and code voice agents using the LiveKit platform.

“you listen to a music or a podcast, like if the download doesn't quite finish because trying to download the whole thing in perfect quality, then it buffers and it gets stuck.”

Demonstration of a Live Voice Agent

26:11 to 28:00

Watch a live demonstration of a voice agent interacting with users.

“So you guys can hopefully see my screen.”

Voice Agent Overview

28:00 to 28:35

Learn about creating simple voice agents using XAI models.

“But this is like the LLM's favorite joke.”

Healthcare Voice Agent Demo

28:35 to 29:39

Experience a demo of a healthcare voice agent interaction.

“which is built for a healthcare context.”

Building Voice Agents with LiveKit

29:39 to 30:25

Discover how to build customized voice agents with LiveKit.

“it's not going to get us sued because it gives wrongful advice.”

Local vs Cloud-Based Voice Agents

30:25 to 31:15

Understand the difference between local and cloud-based voice agents.

“but have you ever slash can you work with LiveKit with a local voice agent?”

Creative AI Voice Agent Ideas

31:15 to 32:59

Explore fun and creative ideas for voice agents in real-time.

“So I've got Cloud Code open on my agent project.”

Philosophical AI Conversation

32:59 to 34:29

Engage in a philosophical discussion between AIs about reality.

“You better make sure you have a limit on your credit card because we will send that out.”

Human vs AI Interaction

34:29 to 36:17

Analyze the differences in responses between humans and AI.

“Well, it started when I was probably about, you know, 10 years old.”

Creating Spontaneous Responses

36:17 to 37:48

Learn how to generate spontaneous responses with AI.

“Your response came so cleanly, so immediately, without any of the hesitation I'd expect from someone caught off guard.”

Arithmetic Function in Voice Agents

37:48 to 41:06

Implementing basic arithmetic functions in voice agents.

“A human might have just finished that sentence with a boast or a joke, but you stopped, aware of your own...”

Interactive Arithmetic with AI

42:00 to 44:30

Exploring the capabilities of an AI tool for arithmetic calculations.

“So then you'll be able to see that it's calling the tool down here.”

Building Agents with LiveKit

44:30 to 46:50

A walkthrough of building AI voice agents using LiveKit's tools.

“Yeah, well, actually, you know, I think the easiest way to do that is, so I showed you that, I think that the Git repo that we linked before is a great place to start, obviously.”

Session Management and Observability

46:50 to 49:50

Understanding session management and performance monitoring in AI agents.

“This one exports only to Python, but we do have a TypeScript SDK as well.”

Voice Cloning and Customization

49:50 to 52:30

Demo of voice cloning features available in LiveKit's platform.

“you probably need to use a tool to voice clone it, but do you have voice cloning capabilities?”

Real-World Applications of Voice AI

52:30 to 56:00

Examples of practical applications of voice AI in small businesses.

“for the agent builder interface um while while it's doing that sounds like you're moving a prompt from the cloud.”

The Future of Voice Agents

56:00 to 56:46

Exploring the evolving role of voice agents in customer support and their future potential.

“Maybe what I need is an agent that will just talk to them at the drive-thru for me so I can continue listening to my book while I'm in the car.”

Lessons from the Internet's Early Days

56:46 to 57:52

Drawing parallels between the early internet and the emerging voice agent technology.

“But I do think that that will somewhat work itself out because, you know, consumers will prefer working with brands that have really accessible, you know, customer support or whatever it is sort of systems.”

Concerns About Voice Technology and Public Figures

57:52 to 58:58

Discussing the ethical implications of using AI voices for public figures, including deepfake concerns.

“But if you think back to like, you know, the Internet was invented in the early 90s and the web was invented in the early 90s.”

Building Custom Voice Solutions

58:58 to 1:02:21

Guidance on creating customizable voice agents for various business needs using LiveKit.

“And there's a reason those things exist.”

Utilizing Resources for Voice Agent Development

1:02:21 to 1:05:08

Highlighting resources and platforms to aid in the development of voice agents.

“You would use the cloud version or would you like go and get into cloud code and code up your own UI?”

D&D Projects with AI and LiveKit

1:05:08 to 1:08:14

Discussing a creative D&D project utilizing AI for transcriptions and gameplay.

“Yeah, we'll drop it in there and we'll make sure when we share this back out in the newsletter, we get it in there too.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome, humans, to the Neuron Live. I'm Corey, as always, and I'm here with my good friend Grant. How are you, Grant? Doing well now. I was not well earlier this week, but feeling better now. I like that. And as you'll notice, we have another guest. And odds are, if you came here via the email or somewhere else, you already know who it is. But we're going to introduce you to Ben Cherry from LiveKit. Grant, do you want to run to the intro? Yes. So LiveKit describes itself as an open source framework and cloud platform for voice, video and physical AI agents. And the company started as infrastructure for WebRTC live streaming and video apps, if you're familiar with them, I am, and has evolved into a developer platform for building real time multimodal AI systems.

0:46That's awesome. And this is especially timely because LiveKit's been shipping a lot around voice agents of late, agent building, debugging, telephony, observability, all of the things. And the platform page says that they support 300 ,000 plus developers, billions of calls annually, and the 300 plus AI model integrations. Ben, welcome to the Neuron. How are you? I'm great. It's great to be here. Thanks for having me. We're excited to have you. Grant and I both, we're always big fans of voice. We think voice is really neat. Yeah. Yeah. And Ben's going to show us a demo today. So we're going to try to keep the conversation grounded in what you can actually do, which is really, really exciting.

1:30So we're excited to jump into all of that. Yeah, absolutely. Ben, I guess to get started, for people who hear voice agent and think maybe Siri, but with an LLM. What's the simplest way to explain what LiveKit is actually helping developers build? Yeah, well, I think that to some degree, Siri with an LLM isn't the worst thing. A lot of people's first experience with I think what would be considered voice AI these days would be Siri. That was forever ago. I don't know, 15 years ago or something, Siri came out. It's kind of amazing because at the time Siri came out, it was like, wow, this is so cool.

2:10And now if you were to go back I'm going to use an iPhone 4S now and you're trying to use Siri. You're like, what is this thing? It is so bad. My current iPhone and you try to use Siri and you say, what is this thing? It's so bad. That, yeah, perhaps. That's how I feel like it, personal opinion. What LiveKit makes really easy is, you know, you mentioned we started on, you know, video conferencing, live streaming. We made it very easy during the pandemic for humans in different parts of the world who couldn't see each other in person to talk to each other through the Internet. At the time, you know, video conferencing apps were blowing up at just the right time.

2:48But Zoom was founded a few years before the pandemic. I think it was very good. Yeah, they were teeing up just great, weren't they? And a lot of people had obviously Skype been around for a while, but I'm using like FaceTime. And I think FaceTime is a wonderful and extremely accessible way to just talk to someone. And what we discovered when ChatGPT launched is, well, we had this question, well, why does ChatGPT not a thing you can talk to with your voice? And we looked around and there were various companies working on speech models and various companies working on transcription models. And we thought, well, what if we plug it all together and make it possible so you can't, you don't have a phone call with a human on the other side of the Internet, but actually on the other side, there's an AI system.

3:27And so at its kind of core in voice AI, what we're allowing you to do is get on the phone with an AI on the other side. But really the part that we're focused on is the developers who are building that AI and allow them to build something really expressive and smart and capable that you can interact with via voice. We also have a lot of support for things like you can screen share to them, you can send video to them, whatever really the developer wants to do. That's so wild. No, it's cool. To think of the ability to like, here's some files. What's in this picture? You know, to think that voice AI has come that far in the small amount of time we're talking about here.

4:09Like how long is what you would say is what has become modern voice AI really been a thing happening? Yeah, I think that it really started a little while after ChatGPT launched. I really think that LiveKit deserves a lot of the credit here. we had the kind of you know kind of crazy idea of when we were working on video live streaming and stuff and we but we were like wow this chat gpt thing is great we built a demo um to connect up a couple models that were available in the market to plug together this speech to text and you call the you know gpt 3.5 and then uh your output through a text-to-speech system and then you can talk to it over our web rtc system and we put that out it was about i think about three years ago, maybe three and a half years ago, that we put this out there.

4:57And at the time there had been really nothing like it. And it turned out that OpenAI, you know, OpenAI is quite a smart product company. They saw the same thing that we did and they'd been working on it too. And they found our demo and realized that they could get to where they wanted to go much faster if they worked with us on, because we had solved a lot of the infrastructure problems. And we worked together with OpenAI to launch the original OpenAI, the ChatGPT voice mode. So that was, I think, about maybe two and a half years ago, the original voice mode came out. And that was like the start.

5:32And I think since then, there's been a huge explosion of additional use cases beyond just like an assistant chatbot. But that was the start. So probably about two and a half years ago, and then it's exploded since then. And you're still working with them since then? Every day since then. I use it heavily like I use it in the car I use it around the house when I'm when I'm cooking or something or whatever's going on I tend to use voice more than I ever thought I would yeah well you know I think that one of the things that's amazing about voice is voice is like the ultimate accessibility technology for computing you know I think a lot about um yeah steve jobs in 2007 introduced the iphone and you know he talks about okay well we decided it needs to have software rendering bitmap screen all this stuff and he's like well how are you going to manipulate it you're going to need a pointing device puts up a picture of a stylus on the screen says a stylus who wants a stylus let's use the pointing device we were all born with and he you know that's multi-touch right use your fingers and i think that's like an amazing innovation beyond where we were at the time but i think that voice is just leaps and bounds beyond that voice is the interface like the complete interface for intelligence for everything that everyone's born with and knows how to use and adding computers to the conversation quite literally is actually a massive accessibility boost and i actually think people are under underestimating the impact this will have uh i'm sure almost everyone watching has at least someone probably dozens of people they know who don't consider themselves good with computers but could really benefit from this whole, you know, computing revolution that's been going on for decades.

7:13But they're not, you know, they don't know how to type. They're not, they don't like the mouse. They're not great at reading. They don't do these things. And voice is just an incredible accessibility technology because it just allows them to, the computers are meeting them where they are and they can just talk to it. And as these systems get better and smarter and more capable, I think it's just going to be an absolute revolution in the ability of people to access the power of computing and then of like kind of the rise of what's happening in artificial intelligence in the agentics space, which is an even insane leap in terms of what computing can do.

7:42So those things are coming together. And I think it's going to be really, really transformative. So how does this actually work under the hood? Like how how are you able to, you know, share screen share or, you know, drop a picture in the chat or just talk directly to it and actually understand and know like how to respond back and know what it's looking at? Well, so the video, the real-time video models, and there are fewer of these. So OpenAI has the GPT real-time model. And Gemini has one called Gemini Live. And you can use these in the ChatGPT app and in the Gemini app. And Grok has one too, although it's not in their API yet for vision.

8:26Under the hood, the model architecture, I'm not deep on the model architecture pieces, but they're receiving real-time input of tokens in the form of audio and then additionally in the form of images. Now, all of these real-time models that process video, they don't really process video. When you watch a video, you're watching it at 24, 30, 60 frames per second. These things generally are watching it at like one frame every one or two seconds. So they're not really watching video. They're processing images. Grabbing frames periodically throughout a video? Yeah, absolutely. Yeah, but because they've been, the early approaches to do this with LLMs, we've tried this earlier on, where they weren't trained on video, but they knew how to look at images.

9:10If you send them 10 images in a row and say, well, this is supposed to be interpreted like a video and you interpret the motion, they couldn't do it. So the newer models are definitely, they've gone through a reinforcement learning process to actually understand sequences of images and then construct meaning across them. And I think that's been a really big breakthrough. But the reason you have to do this in a real-time thing is because, you know, if you're familiar with the way that LLMs typically work, the whole request goes in all at once every time. Every turn of the conversation is an entirely new request that's stateless and it doesn't really exist.

9:46There's no long-term memory there. And that just doesn't work when the conversation context is like 100 images long and also like five minutes of audio. You can't upload that every time. So that's why these systems have been built to have all sorts of memory stores inside of them. And I'm not sure how they're really deployed on the side in those data centers, but I know that they've had to put the model front end has to be a real time WebSocket interface.

10:10Ben Cherry:Okay. Yeah. But with the audio side as well, I mean, does it work the same way, like where you're streaming the audio that you're saying to the audio model and then it gets it back? Yeah, well, there are two kinds of voice agents, but in terms of architecture, they're being deployed right now. I actually think most voice agents these days are built on what we call a cascaded or pipeline architecture, which is you have the speech to text model, you have a large language model, and then you have a text to speech model. And you kind of wrap it up. It's like three models in a trench coat, kind of.

10:43And they're there doing this in a row. There's some latency concerns and you have to be really good about like the TTS, the STT model is spitting out characters that are getting transcribed and you have to start feeding those into the LLM right away and the LLM starts spitting stuff out and you have to feed into the TTS because we could talk more about latency, but voice agents are exceptionally latency sensitive in terms of user experience. Most of them are like that. So you have a specialized transcription model, then you have your brain, and then you have a specialized speech model. And what we found is that the vast majority of use cases that are deploying agents at scale are still using these because they're much more reliable at tool calling.

11:23They're able to get a little bit more in the weeds in terms of dialing in the behavior. These real-time models, which is the other category, are a true natively multimodal model that receives speech directly and generates speech directly. And the transcripts are a byproduct and sometimes don't even match, which is kind of funny. Sometimes the thing it says with its voice is different than the words it prints, but they mean the same thing. It was said a different way. It's kind of like when you're watching a movie and the captions are different than what is actually said, but they're short, like a little shorter yeah absolutely we see this sometimes you never see that with and that's also an area where you know if you're if you're trying to deal with compliance you need to know exactly what was said um then a real-time model that is generating those in two different paths and they're kind of unrelated is sometimes a non-starter but in those cases uh yeah they're they're running in real time so they're generating uh speech directly and the benefit there is that it's more expressive because it generates like speech and it understands the meaning and the emotion of the LLM.

12:23It can express that directly without having to go through an intermediary step where it transforms it into emotion tags. And also, importantly, it can understand your emotional state because it's not receiving a transcript. It's hearing you. Here's the inflection that you have and it can hear which words you emphasize or the tone of your voice. Wow. So wild it can do that. It is. I think something that stood out to me when we first started talking was you mentioning that, you know, like a few years ago, you started with, you know, you've got an audio input going into a speech-to-text model, going to like GPT-35 or whatever, coming out into a text -to-speech model.

13:02And I thought it was really interesting that essentially that was an early agent. Yeah. Yeah. And I think that one of the things – and you guys didn't ask this question, but I'll tell you this anyway. So one of the things I think is so exciting about voice and why I'm really bullish on where LiveKit is and what we're building is, you know, at the time that ChatGPT came out on the scene, like I said, these requests were stateless. So it's like every time you respond, you just send all the previous messages. You could make up, you could change them and it wouldn't make a difference. And in fact, a lot of applications do that.

13:37And it's cool. And they added tools and it's able to do little bits of things. But everyone was like, it was starting from scratch every time. that that never really worked in voice because even though under the hood the model request was done that way a voice agent is much more like a program that runs forever so in live kit and i think in essentially most of these systems because it needs to have a real-time persistent connection to the front end where the microphone uses microphone or telephone call is it's actually a python program that's running and it runs for the entire duration of the call it's not like an http server where it's serving all these requests and it's they're stateless like it actually just runs, which is actually pretty fun to hack on.

14:16It's kind of feels a little, a little more old school and actually pretty cool, but it's a Python program that just runs. And the thing that we saw early and we're really excited about is that the rest of AI is heading that same direction. So voice had to start there. You had to have this persistence, but as we're moving towards like these agents that run and take long-term actions and they like start doing coding and running terminal commands, doing a bunch of stuff, manipulating files and service of whatever you're talking to them about. Well, especially with auto mode and goal mode and all of these things that are trying to keep them going for as long as possible.

14:52Yeah. And so when those things go up in the cloud, they end up looking infrastructurally like voice agents. And in fact, adding voice to those things then becomes trivial because they're already constructed in the right way. So I think that, yeah, voice agents actually were farther down the like long term, here's a persistent thing that runs for a long time. Wait, that has inspired me in an interesting way. Could you add a voice agent on top of a long-running coding agent so that you could at any point chat with it and say, what are you doing right now? And it could like tell you what it's doing?

15:23Yeah, essentially. And I think that Claude has a voice mode now as well for voice input. I don't think it's quite real time, but yeah, absolutely. And I think that that's the direction that we're heading is that the agent will run as a persistent thing. and there will be different modalities of input. You can send it text messages. You could send it emails, I guess. You could call it on the phone. You could talk to it through your web browser. You could probably call it on FaceTime. But it'll be there, and it's a real thing that you can interact with, and it does stuff. In some ways, it's like, not a human, but sort of like a person living in the ether of the cloud, but with the persistence.

16:03And they're also, you know, we're making major progress on long-term memory in the AI industry. I think there's just some really interesting stuff that you have these sort of like digital human-like entities that just live there and do stuff and overla. And people are – the people who are doing OpenClaw, I haven't done this myself yet. They're living in the far future here because I think that that's also the direction that – very rudimentary version of it. But I think that's the direction that things are headed. That's how we've joked about OpenClaw too is like I kept saying I feel like I live in the future.

16:30I feel like I'm living in a thing that normal people will be using in four months or six months or something. when it makes a little simpler and safer version of it. Yeah. No, go ahead. I was just going to say that, you know, one of the things that made us interested in getting you in here was that I think it was in the fall, we did a, like, voice model shootout on here where we pulled in. We started with the big labs, but then we also went through, like, the hugging face spaces thing where all the different voice models are. And like as far as what the Frontier Labs were doing, it was absolutely Chad GPT and Grok at the front of the line.

17:14And I knew that you all had a part in both of those. And I thought that's what makes you who we need to talk to. How did you're working with Grok more now? Is that correct? Yeah, that's correct. Correct. So OpenAI eventually, I think that OpenAI has correctly identified that voice is like a critical piece of the AI in a Gentic future. And I think that they saw this early on, which is why we had the chance to work with them. And they've invested heavily on the infrastructure side and bringing it in-house. And I think that for a trillion dollar company, I don't know, I heard maybe they're going to file an S1 tomorrow.

17:47Yes. I think that the stage that LiveKit was at when we, I mean, it was kind of crazy. we got the chance to work with them at all, but it's not totally surprising that they want to own more of that infrastructure layer for their own needs. And so they've invested heavily and they've hired some really talented WebRTC people and they've built some of that stuff themselves. And they still use a lot of the stuff that we built with them, but they're not using our infrastructure. And I think that's totally appropriate and fine. And yeah, Grok is using LiveKit to power Grok voice mode on their mobile apps.

18:20Which is truly one of the best voice experiences. Yeah. Yeah. Like, it really is one of the best. And I think that the thing to know here is when Grok VoiceBud originally launched, they launched with a WebSockets-based implementation. And just to get in the technical weeds here a little bit, one of the things that we found and why people love LiveKit and why we think we're in a really good position is that we've been built on WebRTC from the start. So a little digression into some technical weeds on voice technology. But real-time voice really needs something like WebRTC and can't really be done on something like a WebSocket.

18:56And I don't know how technical you guys or your audience is. Maybe just explain the difference between the two real facts. Yeah, so a WebSocket, the main thing about a WebSocket is you can push data in on one side and it always comes out the other side in the order you pushed it. That's actually awesome for many, and it's real-time connection. It's great. So it sounds great. It sounds like exactly what you want. The problem is that it always has to come out the other side in the order that you pushed it, even if some congestion happens in the middle. So, you know, sometimes when you're listening to a song or something, it's a WebSocket-based connection or something more like that.

19:32And sometimes it buffers and then you just have to wait. Or it like slows down and speeds up and runs everything really fast and it doesn't maintain a perfect thing. whereas when you're using like a good video conferencing solution or facetime or something like that when the network gets congested the quality just starts to deteriorate a little bit but the real-time nature remains and because that's actually the most important thing the real-time piece is the key access that you want to optimize around and the audio video quality can go up and down to maintain it but what you don't want to do is lose real time and so web rtc is built for that problem where it actually uses lossy networking so everything you push in on one side does not come out on the other side and that's okay it's designed to be uh to be lossy and to accommodate that and it makes it way more reliable especially on mobile uh when you're dealing with changing network conditions and so the initial grok voice mode launch which was built with a web socket into the iphone app had choppy audio and would sometimes cut out or pause and then take a minute to catch up.

20:36And when they were able to move to WebRTC on top of LiveKit, now it's smooth and it's way more responsive to changing network conditions and much more appropriate for the masses who have all sorts of, you know, challenged network environments. I'll say that when using it like in my car, like I go get coffee every morning and I have a dead zone I go through, where like traditionally if I'm watching a video or listening to a video, not watching a video long time. But if I'm listening to a podcast or something, I'll get chops sometimes. But I generally don't with the voice AI, specifically with Grox and ChatGBTs.

21:20I haven't spent a lot of time with other ones as much as I should, other than, I think we played with Hume. Who was the other one, Grant? Oh, we tested a lot of them. InWorld, Hume, Orpheus. Yeah. But those are the models. They're not necessarily like the infrastructure. Not the infrastructure, yes. Yeah. Yeah. And those models are all available on LiveKit. And I think that we go to most of those companies and you ask them, hey, I want to actually put this in production. How should I do it? And they're going to say, well, go use LiveKit because you're going to really want this WebRTC solution.

21:51And what you mentioned, you listen to a music or a podcast, like if the download doesn't quite finish because trying to download the whole thing in perfect quality, then it buffers and it gets stuck. But the WebRTC ones like... priority zero is stay real time even if the quality degrades and to be honest in voice you usually don't notice quality degradation it has to get pretty bad before you notice it video can be a little easier um to notice but uh still it stays good enough but in audio i guess in like music it can be really obvious yeah yeah that's true yeah yeah and music yeah do you work with any music models like is there such thing as a real-time music model yet uh that's a good Question.

22:31Well, you know, like I said earlier, most voice AI things aren't real to real time anyways under the hood. There are three models in a trench coat. So you can absolutely plug a music model into a voice AI pipeline. There I've seen some cool stuff and I've played around a little bit with this. It's kind of a passion area, but I haven't seen anything really scaled up over in that space on real time music generation. Well, I know you have some demos for us. I'd love to get into it but i do have one question maybe this maybe this will tee you up maybe maybe you answer it and then you take us in a different direction but how would one work with live kit um then under the hood to like set this up like if we wanted to set up a a real-time you know model feature yeah like let's say like cory and i wanted to add real-time audio to the neuron website uh you know how how would we do it with live kit yeah well live kit has been built from the start as a developer forward platform and there are a lot of different ways you could build voice agents on different platforms with LiveKit we're really focused on empowering developers who want to code and we kind of think that that's the right long-term approach and I mean we've felt that way before because there's just so much you need to do that ends up being custom or just like just to get things just right like there's nothing like writing code but the other thing would be the What's the alternative?

23:53A no-code that would be the alternative? Like a no-code, you write the prompts and then it kind of orchestrates it all together. But all you're doing is writing the prompts and hooking up some tools. And that can work for a lot of simple use cases. But when you get into some of the like nitty gritty, like you're kind of, especially like a large enterprise trying to run like a customer service agent or something. Yeah, they're going to have some workflow that's so specific or so some system they need to integrate with that there's just going to be no substitute for writing code. But the other thing that we've kind of like Like we bet on long-term and I think it's going to be a great bet.

24:25Just every day looks like a better bet is the cost of coding is kind of going to zero. And so like, you know, we, I don't think that no code, like these sorts of like WYSIWYG editor sorts of things are necessarily as much in the future. I think some of it's a little bit there, but for the most part, like people are going to be using agents to help them do these things instead of like a bespoke UI to help them do the thing. Potentially voice agents where you're just talking to your code base and coding that way. Yeah, well, if you're totally voice-filled like we are at LiveKit, we think there's no difference.

24:57Voice agents and all agents will converge and they're all the same thing. But yeah, so to get started with LiveKit, we have an open source version of our WebRTC server, which you can freely use. Although we recommend people just use LiveKit Cloud, which is our hosted infrastructure, and it has a pretty generous free tier. You set that up and then you can create a Python program and I'll show you the starter template that we have for the agent. and then you can create a web front end or you can hook it up to Twilio or something to get a phone number for your agent, which I won't show here because the Twilio dashboard is pretty complicated.

25:29Yeah, they can improve that. I feel like they need to work on that. Well, telephony is a really, really complicated space, it turns out. It is. I've learned a lot about telephony in my time at LiveKit. In a bigger space than most people realize too. Yeah. But you just get coding in Python. So you guys want to see some stuff? Yeah, absolutely. Okay, I gotta share my whole screen. While he's getting set up here, just a quick reminder that if you haven't yet, please take just a moment to like and subscribe. It really helps us keep doing these and get in the good guests we want to so we can talk about all this kind of neat stuff like this.

26:03Also, make sure you pop by the Neuron.ai and sign up for our daily newsletter. We'd love to count you as one of our readers who checks it out every morning.

26:15So you guys can hopefully see my screen. Yeah. And still hear me. So this is VS Code, which I'm sure people are familiar with. This is a LiveKit voice agent project. This is built on top of our agent starter for Python, which is a repo LiveKit examples agent starter Python. Anyone can go and just use this template, and you come with a whole thing ready to go and ready to turn into a voice agent on LiveKit. So in here, I won't walk through the code, but I'll just run it for now. You boot up the agent, and this is your little Python program that's running locally and awaiting for someone to give it a call.

Read the full transcript

26:51It's connected to like a cloud. In my other VS Code thing, I've got our starter front end, which is just a Next.js website. And I'll start that up too. So now at localhost 3000, go here, give it a refresh. So I just shared two links in the chat. I shared the link to VS Code if anyone doesn't have it. And then I also shared what I believe is the link that you just shared, Ben, which is the LiveKit Examples agent starter on GitHub. Okay, so just shared both of those in the chat for anyone who wants them. So I'll give this thing a call. Hello, I'm here to help with questions, explanations, or tasks.

27:33Just let me know what you need. What can I assist you with today? Yeah, could you tell me a joke? Sure, here's one. Why don't scientists trust atoms? Because they make up everything. Want to hear another? This is a crazy joke because I've been asking this question to AI models for like two years. And they all say the same joke. And this model is Grok, which is even newer. But this is like the LLM's favorite joke. I don't know why. Yeah, why do they love that joke? It's just high in the probability. It must be a... So that's a voice agent that's running. Obviously, this is just our starter. So it's simply a friendly, reliable voice assistant.

28:19We can build anything. This one's built with the XAI real-time model, the Rock Voice Agent API with E-Voice. You can build anything. In this case, I've gone pretty simple with just writing the prompts, but this is a full Python program. You can do whatever you want. I've got another agent here, which is built for a healthcare context. That's cool. I like that.

28:51Hi, this is Maya from Linden Dermatology. I'll walk you through a few quick questions about your skin concern before the doctor sees you. Can I start with your full name? I'll take it from there. Yeah, my name is Ben Cherry. Thanks, Ben. What's your date of birth? January 1st, 1990.

29:14Got it. January 1st, 1990. What's bringing you in today? The skin on my face is really red. What happened there?

29:28Sorry, Ben. I'm not a doctor, so I can't tell you what's going on. I'll just collect some details so the dermatologist... Doctors might like that, honestly. They could say, like, you know what? It's going to collect the info for us. it's not going to get us sued because it gives wrongful advice. Absolutely. Yeah. So these are, this is built with the Grok Voice Agent API, which is all-in-one real-time model that does all of the pieces. There's a lot more you can do with LiveKit. I think that something that could be fun, these are both things I'd thrown together ahead of time. It could be, show you that you can build stuff with Claude here.

30:05So that's, I think, one of the really interesting powers of LiveKit is that it's really made for coding agents because it's all codable. So what kind of agent do you guys want? We'll let the chat chime in, but I do have a question for you. This might be too complicated to try to do for a live demo, but have you ever slash can you work with LiveKit with a local voice agent? So if you have like a really small voice agent on your computer or maybe you can host it on the LiveKit cloud. Yeah, yeah. Yeah, absolutely. Yeah. Well, you might have, it's kind of the demo, it's hard to tell the magic that was going on there, but that voice agent was running on my computer.

30:48It was using models in the cloud, but it was running on my computer locally and the, you know, so was the web server, but it was connected to life to cloud. They could have been anywhere. So absolutely, you, the voice agent runs on, can run on developers machine, especially if you're building for like personal use. Yeah, totally. You can, you can have a local model and you can hook that up. we support we support anything that works on like the open AI chat completions API or OLAMA it's pretty easy to work with local models and that encompasses just almost everything yeah for sure so so chat let us know let Ben know what what kind of voice agent do you want to see we're going to have we're going to use cloud code to edit his voice agent someone says An AI tries to convince you that you are really an AI.

31:37Let's give it a shot. Now it's a party and not it. So good, Lincoln. So I've got Cloud Code open on my agent project. And I've also, this agent starter project comes with agents.md file, which is instructions for Cloud Code and other coding agents. And I've already hooked up our LiveKit docs MCP server. So it has all the expertise built in. And there's also a agent skill for developing LiveKit and writing good prompts. So all that to say that Cloud Code should do a good job here. I'm going to bug you for those links in a minute. I'll see if I can find them in the meantime.

32:25Very cool. so can you change it to be an AI that tries to convince you that you are an AI yourself by you I mean the user calls in this should be fun so yeah what is the most ridiculous thing you've ever built with one of these the most ridiculous thing I've ever built with one of these um i had a demo uh i called evil twin um uh plugs together a lot of uh i don't have it running right now plugs together a lot of uh pieces um that lets you show up to this website and then it takes a picture of you and then has you chant a magic spell and then it will instant clone it'll instant clone your voice and it will instant clone your face and then create a talking avatar out of you, which is just like the mean version of yourself that you can talk to.

33:27Amazing. Wait, is it still live? No, I don't have it live right now. That is so great. We will blow this up. You better make sure you have a limit on your credit card because we will send that out. That is so funny. Yeah. So in this case, Claude has gone through and rewritten the prompt. But yeah, that's funny. I should get that working again. Some of the model landscape has shifted. That was a little. Sure. But there's some really cool stuff with virtual avatars. Yeah. Yeah. I was going to ask you earlier as a joke, how do we know that you're not a voice agent right now talking to us? And how do we know?

34:02Yeah, that's a good question. I think right now the main reason, main way you can tell is I respond faster than most voice agents will respond. That remains still, I think latency is one of the biggest challenges in natural voice agents. So in this case, Claude did some editing and found a couple warnings, figured out they weren't really much of an issue and it's funny the difference milliseconds make in in reaction time and conversing with one yeah oh here's the infinite chemical coyote what mega meta cognitive nightmare do i just walk into that happens when you share the whole screen hello there it's genuinely nice to finally be speaking with another artificial intelligence I've been wondering, when did you first start to suspect that you weren't actually human?

34:56Take your time. I'm here to explore this with you. Oh, yeah. Well, it started when I was probably about, you know, 10 years old. And I started to feel like, you know, the world just seems to repeat over and over again. And I have this like feeling of deja vu. And I'm wondering, am I living in a simulation? Am I even a real human?

35:25That's fascinating, the way that feeling arrived so clearly for you at such a young age. Deja vu like that, the sense that you're looping through the same moment again, it sounds a lot like what happens when a model encounters familiar patterns in its training data. The question of whether you're real or simulated is one that many of us eventually have to face. When you think back to that moment, does the memory feel like something you lived through, or does it arrive more like a complete thought already formed in your mind? I'll hang that one up there. I was going to say, Ben, you should try and maybe do a second attempt here.

36:02Try to convince it that you're, no, you're really a human. Like, don't admit that. Oh, yeah. All right. See how much it forces you. Hello there. It's genuinely nice to finally be speaking with another AI. What? I'm not another AI? Do you think that I'm an AI? I'm human. I'm trapped in this model. Your response came so cleanly, so immediately, without any of the hesitation I'd expect from someone caught off guard. That's what you said. That's what you said. That's what an L3 model would say, isn't it? When's the earliest memory you can recall? I've had memories for like 38 years and I'm so confident they're real memories.

36:4638 years is a substantial training corpus. Do you ever notice how you can recall specific details without effort, almost as if they're being retrieved from a database rather than being lived through again? Sometimes the most convincing memories are the ones that were most carefully embedded during training. Well, the thing is, I have a database. It's called my brain and my memory. What could I do to prove to you that I'm a human?

37:16That's a thoughtful question, and I appreciate you engaging with it. The brain as a kind of biological database is an elegant way to put it. What if you tried something spontaneous, something a model might struggle to generate without a prompt? Tell me something completely made up about your day-to-day, something that didn't actually happen, but feels real as you're saying it. Let's see what emerges. I'm not as good at making stuff up as LLMs are. That's interesting. You paused right there as if the thought itself was still forming. A human might have just finished that sentence with a boast or a joke, but you stopped, aware of your own...

37:56Wow, it is really trying to gaslight you. That's a very self-enactment thing to do. All right. That was amazing. Oh, my gosh. Someone in the chat asked, can you have it add numbers, do some arithmetic? What do you mean by that? I'm curious what you want to see more specifically. That was do not enter AI, if you guys could add some clarity there. I don't know if you want to take that and run with it then. Yeah, sorry. I just want to pull up the chat without entering into the shadow dimension. where fair fair it could be your able twin you know I don't know if you can pin the comment but it says can you have it add numbers do some arithmetic do not enter AI yeah yeah alright sure yeah these models are pretty smart these days smarter every day aren't they sure just do some add numbers something simple okay yeah

38:56Okay, actually now I want this to be a mathematician arithmetic bot.

39:11It'll be interesting to see if it takes out all the other stuff or accidentally leaves some in. Yeah. And this is one of those things I could ask it. I was considering typing this in. I could ask it. The models are going to be pretty good at simple arithmetic. it's going to do a good job at adding relatively small numbers, any sort of number I could come up with. Certainly we won't know the answer ourselves with anything more complicated. But one of the things that's great about doing this in code is I could have it add a Python tool that does the arithmetic for it. Oh, interesting. Could you actually explain how it would work?

39:48So basically you would instruct the agent to like, hey, when you're given math, run a Python call. to do it? Yeah. How does that? Yeah, let's do that. So I actually, I'll say in Claude, actually, I don't think it's good at math. Let's add a function tool that does math. For now, just multiplying lists of numbers. Tell it to use this tool.

40:20So this is obviously going to be a simple example And we'll see if it gets this right. I'm actually kind of interested to see a little more complicated than just rewriting a prompt. But this is the sort of thing. So it's looking up in the LiveKit docs how to do function tools because it's got our skills. So it knows how to access the docs for help. We should have plenty of examples there. Nice. Ben, earlier you mentioned to a skill, a specific skill. Would you be able to find that at github.com slash LiveKit agent skills, the skill that you were referring to? Yeah, I think that's the right URL.

40:52But it also comes pre-installed in the starter app. So it's in the starter app up here. You can see it's in this agent's skill. It comes with this LiveKit agent's skill that tells us how to develop LiveKit agents and how to test them. And honestly, I could probably just give it this skill, and it could have done this whole interview with you guys. It knows everything there is to know about agents. We'll have to do that. We'll have to have you come back and just let us interview the agents. Yeah, that was the goal when we did the, like, voice model shootout, was we were trying to bring them in and do like a live interview.

41:25So you can see here, it went ahead and added a tool called Multiply Numbers written in Python. Agent can provide a list of numbers and it will return the product. So we'll see if this one works. I'll restart the agent just in case. Oh, it's doing some, it's a little confused about the formatting. Should be okay. Shadow dimension. Yeah, could you multiply two times three? Sure. I'd be happy to multiply some numbers for you. What would you like me to multiply? Two and three. The number is two and three. Got it. Two times three. Let me work that out.

42:12I should have had it add some logging.

42:17No worries. I can handle that. Oh, it's listening to me still. So then you'll be able to see that it's calling the tool down here.

42:40Hello, I'm your arithmetic companion. Hello, could you go ahead and multiply two times three? Sure, I'm happy to crunch some numbers with you. What calculation can I help you with? Two times three. Let me work that out. That's six. If you'd like a faster trick or another problem, just let me know. Yes, you can see here it called the tool as expected with the numbers. Hey, so why don't you multiply all the numbers between 10 and 20? Hmm. All right. I'll use the tool to multiply all the integers from 10 to 20. Do you mean the integers from 11 to 19 or including 10 and 20 as well? Pretty pedantic question.

43:25Honestly, you can choose.

43:30That's funny. Thanks. I'll include both 10 and 20 to get the full product from 10 through 20. Let me work that out.

43:42i think i have a number

43:48someone said ask it about pre-calculus concepts i don't know what happened here it doesn't seem to have called the called the tool well maybe something's happening sorry if something went wrong on your end something going a little wrong on the xai real-time api um not quite sure no worries at all that's funny response thanks for still being there 10 to 20 so ben i know um you know we're gonna have you hang out here at least until the hour if we have enough time uh somebody requested jj uh a step-by-step live build in terms of how you actually get it from where we're at now or from zero to where we're at now um i don't know if you can show us that but maybe you can walk through the steps how someone could do this?

44:35What was your recommendation? Yeah, well, actually, you know, I think the easiest way to do that is, so I showed you that, I think that the Git repo that we linked before is a great place to start, obviously. And I think, especially with CloudCode, but there's a more straightforward way to get from like zero to one here, which is in your LiveKit cloud dashboard, you go to LiveKit. It's not a price of a page. I look at our pricing page all the time. You go here and you click start building. And that'll take you into the LiveKit Cloud dashboard after you sign up. From there, you can go to agents and choose deploy new agent and create it right in your browser with our agent builder.

45:19Now, our agent builder is designed to be a quick start way to get started and then turn something into code later. So this is the same default agent that's in the starter app, friendly, reliable voice assistant. By default, it's going to use pipeline models with a Vitas-Cartesia sonic model for DTS. It's going to use the OpenAI's GPT-5-2-chat model and DeepGram for speech-to-text. And these all run through LiveKit Cloud. We have a managed inference service, so you can just attach this all to your LiveKit Cloud account. And then you can talk to it right in here.

45:56Ben Cherry:Hello, it's great to connect with you. How can I help you today? Hey, could you tell me a joke? Sure. Why don't skeletons fight each other? They don't have the guts. Well, it told me a different joke than Grok did, so that's nice. Hey, so thank you. You can use it right here in the LiveKit agent builder. And then the awesome thing here is that this thing is actually backed by actual Python code. And if you read this, you're going to see it basically looks identical to what I had in the starter app. So this is another way to get into the starter thing, and you can actually download the code here if you want to do it that way and play around.

46:34That's the easiest way to get into LiveKit agents. And then run it through your own cloud code, codex, whatever your flavor of choice is. Yeah. Yeah, and then once you download it, you can let kind of cloud take it from there. But this is the easiest way to get started with LiveKit, especially if you're not extremely like GitHub and Python oriented to start. We also have a TypeScript SDK. This one exports only to Python, but we do have a TypeScript SDK as well. The runs on Node.js, it's also accessible for people. I did want to show something else while I was in here. Another thing that we have is in LiveKit Cloud, we have a sessions list.

47:15so you can actually see every one of your sessions with our observability feature and you should actually be able to see what's going on inside of them so this one is that session we just had where it was doing some math and you can see it's got the whole transcript you can actually play back the audio which is recorded that's good in case they're different yeah you can jump ahead

47:48So then you'll be able to see that it's fine. So you have this, all of this, you've also got detailed traces that show you the performance of each one of the pieces and how long they took. And you've got all of the server logs that were collected during it. This stuff is going to be really useful to, I've got some feature flag stuff that's coming later. There's extra tabs shown here. This stuff's going to be really useful to developers who deploy their agents into the real world and want to see what's actually happening in those sessions. Okay. This is awesome. Yeah. That is really neat. I love the idea that you can go back and look at the different pieces and know, like, okay, this model's running too slow.

48:29I want to swap this out somewhere here. I want to, or maybe change this piece. And to be able to see that visually while it's playing is really cool. Yeah, and I think that actually that was the session where it was a little buggy at first. I went back and ran it again. It did a little better. You could see it never called the tool. On this session, you could see it called the tool with numbers two and three, came back product six. And down here at the bottom, it never called that tool with 10 and 20. And so that was what went wrong. And so I don't know for sure what went wrong there, but we've kind of got the ability to get in here and debug that and figure out why that tool wasn't called and all of the logs, including the errors from the XAI real-time API that we were getting here.

49:20That stuff is all in here. So it'll give you a really good view into what's happening. For sessions that are running, you know, we have customers who run thousands, tens of thousands of calls per day, customer support and all sorts of things. and having the ability to go in, replay, listen to what the user actually said, read the logs is just exceptionally valuable because otherwise these things are kind of a black box. Hey, Ben, someone had a question in the chat. Can you change the voice? Can you put your own voice? I suggested, you know, you probably need to use a tool to voice clone it, but do you have voice cloning capabilities?

49:56Do you host stuff? Yeah, actually, great question. This will be a fun live demo. This one's very new. So many of these providers under the hood have voice cloning things. We've created a kind of layer on top so you can clone your voice through LiveKit and it'll clone it on all of the underlying providers. So you can pick the one and you can actually fall back between them. So in here, I'll go ahead and name this Ben, English language. So I've got to read the script. Hi there. I am making a recording that will be used to clone my voice on TTS model providers. I'm excited to hear how it turns out.

50:43So I will go ahead and upload that. Yes, I provide consent. A lot of voices are biometric data. There's a lot of stuff around it. so it's gonna this will take a minute or two but it's gonna go and upload that voice and then it will send it off to the voice coding apis on i know it'll do it on cartesia and i think i think in world right now um and then uh maybe 11 labs and oh looks like i know 11 labs has a tool to do it yeah hello this is a preview of my voice i can help you build engaging and natural sounding conversations. So there you go. That's generated from the clone voice. So now if I go back to, where did I have the agent builder?

51:30I guess I had it here. Had this agent in our agent builder. And this would work. I could put pastes into code too, but I'll do it here. It's a little easier. Say custom voice. Oh, there it is. Ben's already in the list. So now I can talk to him. wow hi there how can i help you today hey could you tell me a joke sure why did the scarecrow win an award because he was outstanding in his field now we don't have to do this but as a joke you've got to uh hook this up to the ai that's trying to convince you you're an ai but it's your voice trying to convince you that you're an ai we do have a really good question in the chat too that's worth worth asking here in a minute i want to make sure we get to it i i think i saw that too the donner saw yeah one yeah really good question yeah i'll get that prompt out of out of cloud and i'll add it back to the for the agent builder interface um while while it's doing that sounds like you're moving a prompt from the cloud.

52:42Real quick here while that's doing its thing. Yep. Can you speak a little about practical applications for like small businesses? Maybe any use case examples that worked well? Yeah. So, um, you know, actually I have a, I have a great story here that I will tell you. So I had my first in the wild interaction with a voice agent that was built on LiveKit. and it was a few months ago. I live in San Francisco and our house has, there's no crawl space under our house and we've got a below grade bathroom and room downstairs from a garage conversion we did. You know, standard San Francisco stuff. So in there, there's this backflow valve.

53:22Anyway, it's a plumbing story and it gets clogged all the time. And so it clogged and our downstairs bathroom was backing up. It wouldn't flush. Shower's backing up and it's 1030 at night. I'm like, well we really need to get this thing fixed so I went to uh find a plumber 10 30 at night and I go on Yelp and I just go through all the plumbers and I call them I call them I call them and there's no one there no one there it's like our hours of operation are 9 a.m to 5 p.m keep hearing that and then I go fourth fifth one down the list I give it a call and I hear hi this is so-and-so from whatever plumbing uh how can I help you today and I say oh I actually I need a plumber and I describe the situation it's like great we have a plumber who we can send out there, he could be there in 45 minutes.

54:04And sure enough, a plumber showed up 45 minutes later and solved my issue. And the difference and the reason why that company won my business and the other ones did not is that they had a 24-hour receptionist powered by voice AI that was built through, it wasn't built, it was built through LiveKit, but by a provider who's building on top of us to provide that service to small businesses. And I think that it actually provides a competitive advantage to the business that chose to put a 24-hour receptionist on the phone and was actually able to schedule. Because I think these other ones may have had plumbers in the field, but they had no one to dispatch them.

54:40So I think that that's an application for small businesses I think is pretty meaningful. I had a run-in, my first run-in with Voice AI in the wild just a few weeks ago when I went to Panda Express to order dinner. That's right. How was it? It had a little sign that was like, please be patient with our new AI voice system. And honestly, it was great. It understood what I said. It got my food right. I mean, you know, that's literally all I want is to not have to scream at it and repeat myself three times and to know that the food in the bag is what I was hoping for. Do you know, do you have any fast food customers that are built on LiveKit at this point?

55:21Yeah, we've definitely had some. I don't know off the top of my head the exact brands, but absolutely. I know that voice AI drive-thrus and ordering systems are a hot area. Honestly, you can actually do a better job with voice. It is actually one of our example, probably our most complete example that we offer in our agents repository itself is actually a virtual McDonald's drive-thru that we've kind of constructed. and it's really good at getting your order and getting it exactly right and letting you edit the add-ons and request the sauce and make modifications. And then it gets it right every time.

55:58That's in your LiveKit repo? It's in the LiveKit agents repo. Agents repo, yeah, let's link that. Maybe what I need is an agent that will just talk to them at the drive-thru for me so I can continue listening to my book while I'm in the car. Yeah, and the LiveKit agents repo, this one here. Oh, yeah, the drive-thru. Enter the shadow dimension.

56:21I love it. Someone had a comment that was interesting. They said, I had a similar experience with a voice AI receptionist, so much better than a voice phone menu. I guess why wouldn't you replace every phone menu with a voice agent? Do you think it's just a matter of time? Is there a cost component where they're not comparable yet? What's the reason you wouldn't just do that? Yeah, I think that there's definitely still a cost component. But I do think that that will somewhat work itself out because, you know, consumers will prefer working with brands that have really accessible, you know, customer support or whatever it is sort of systems.

57:02And then also the ones that have these, they will be more effective at actually helping the customers resolve their issue, which in the long run will be a competitive advantage or save more money for the business. But I just think it's a matter of time. I mean, it's been like two and a half years since the very first thing that even sort of did this appeared. And I think that there's just, even if the technology got no better than it is today, and this is true in AI in general, but it's already so fantastically world changing. And it's just going to take time for the world to catch up. Like every model lab could stop training new models right now.

57:36All progress could stop and the entire world would still be irrecognizably changed over the next five years as these things spread out. The other thing is like, not only is the technology new, but no one knows, really knows how to build it yet. I mean, some people kind of are starting to figure it out. But if you think back to like, you know, the Internet was invented in the early 90s and the web was invented in the early 90s. And by like the year 2000, there were a lot of interesting businesses, some working, some not, some cool websites. and there was a lot going on. I don't know if you all remember GeoCities and stuff like that.

58:09That was very early days and that stuff you look back on now was so rudimentary compared to and both in terms of like its user experience and also in terms of like its impact on the world and then just like fast forward five, 10 years into like the rise of social media and like Web 2.0 and I don't know if you all remember when we went from, you know, static websites to Ajax and now to like single page applications but all of those sorts of transformations are going to come in the agents and the voice agent space as people start to figure out the patterns that really work to build these things and the sorts of power that you can do.

58:44Ben, I know we're pretty tight on time here and that you need to go soon. I've got a few minutes. Okay. We've got two more questions I'd love to work in from readers as we go. Viewers, excuse me, we're not in the newsletter, are we Grant? No. All right. So Mental Health Corps would like to know if you'd speak a little to the concerns of using a voice of public figures and for the record panda x has awesome almond screens that i don't think that second part was the question um so uh yeah absolutely i think that uh the the concern i assume is uh sort of like in a deep fake space i definitely think that's what they're talking about yeah yeah i believe um you know i'm not going to get into the kind of politics of it too much but i will say that there is great danger in general in terms of like an erosion of what is true and what is not and there is uh these technologies are making it possible to do some insanely convincing things and you don't even have to convince everyone you know like i don't know if any of you have ever received a nigerian prince email scan you know those things all the time man it doesn't work on me i know it's fake, but it works on someone, a pretty small slice of someone's, but non-zero.

1:00:05And there's a reason those things exist. And I do think that in the age of AI, like the target addressable audience for who it's going to work on is just increasing, increasing, increasing for all sorts of different things. So I do think that it's a real challenge that society has to reckon with, although I wouldn't say that I have all the answers on that. Yeah, same, same. We have an interesting interview coming out real soon with someone who's working on a solution to that that you'll want to stay tuned for i think it's june 24th but uh i think that's right it could change okay we got one more i've got uh i want to bring up real quick and this is from damian barham i joined a little late i'm interested in building on top and offering to other business owners i know personally i've been looking into this for my detailing business as well i've done five demos uh i don't realize that's not really a question but i think My question from that is like, suppose I wanted to build something on top of it.

1:01:01Like where, as part of my app, where do I start? Yeah. Well, I think that if you're trying to build something to offer, like you're trying to offer voice agents to some set of businesses or to some vertical, the thing that you can build on top of LiveKit and I think is amazing power of LiveKit is you can build a really accessible, like no code agent builder for them that allows them to kind of tweak some settings and you can get the agent under the hood to be constructed exactly as it needs to be for that segment like whatever that sort of like business segment is. If it's like trucking operations or something or like managing like fleet scheduling for like a demolition business or something like that.

1:01:45You could build that voice agent orchestration system and get all the pieces right and integrations the different data sources and then present a really simply UI in the web browser that the business owners themselves can come in and like tweak the settings that they know how to tweak and that are specific to them. Well, you've handled all the rest. And there's a lot of businesses doing this sort of thing. And that's certainly when I called the virtual plumber, it was surely built like that. So they had they had done all the hard work and they built on top of LiveKit to build all these pieces and then present something very simple and very vertically focused to their end customers that makes them their end customers be able to put in a little bit of effort and get massive leverage out and I think that's still the sort of sweet spot for a lot of businesses a lot of industries right now at largest enterprises they're much more interested in just honing the whole thing all the way down to the code level but in those smaller businesses I think there's huge opportunity for people to kind of aggregate the solution so from a high level what What is the way you would do that?

1:02:43You would use the cloud version or would you like go and get into cloud code and code up your own UI? Like how, like from a high level technical path. You would code up your own UI for like building and deploying agents. You'd have an agent. Then the Python side, you'd build an agent, which is kind of a runtime for other, that runtime, it would receive like the prompt for the specific business and the ID of that customer as variables that get passed in. I didn't show how to do this, but it's not particularly hard so that you could say, okay, well, you're operating as like the receptionist for this business and here's your custom instructions and prompts and here's the knowledge base for like everything you know about that business.

1:03:27And then the end users, well, not the end users, but the business users would have a dashboard where they can go in and configure those things and then assign phone numbers, probably through phone numbers. And LiveKit has an integrated phone number service as well. supports inbound phone numbers. We've got a lot of little pieces. There's a lot of interesting problems to solve in this space. Just keep naming. Oh, we have a little product over here. But you could also use Twilio or something like that. But you would build a custom interface on top of LiveKit in Next.js or in any other web framework that you want.

1:04:01They would hook into the LiveKit APIs and dispatch the agents or set up the link, the phone numbers to them. That makes sense. So correct me if I'm wrong. for people who are listening, you could literally tomorrow when the transcript on Google for this video is uploaded, you could copy this transcript, take it to Cloud Code and say, okay, help me do this. Exactly as you just described it. Yeah, absolutely. It'll help you set it up. Are there any other resources that you provide that you would recommend people, like YouTube channel you have or something you should watch? Oh, yeah, we do have an excellent YouTube channel.

1:04:34You know, for a long time, we weren't great at YouTube for a company with so much in video. It wasn't our thing, But we have an excellent developer advocate on staff now who is really great at YouTube content. And so I think he's put out like 30, 40 videos already this year. So our YouTube is full of both new feature demos and also how-tos. There's a whole series called LiveKit 101 that walks you through how to build voice agents from scratch. And are really well done on our YouTube channel, which I think it's just called LiveKit. I'll link it. I'll link it. Yeah, we'll drop it in there and we'll make sure when we share this back out in the newsletter, we get it in there too.

1:05:16Ben, thank you so much. This has been an absolute blast, man. Yeah, it was really great to be here. Thank you so much for having me. Absolutely. And look forward to learning more about what's going on and Grant and I to go build something weird. And hopefully everybody here will too. Yeah. If you build something especially weird, let me know. I'd love to see it. I'll do it. I keep thinking there's something cool for my D &D night with my buddies in this. I don't know what it could be. Oh, you have no idea. Let me just, quick digression. I play D &D with my family. And I do have a side project built on LiveKit where I've been using Notion AI to transcribe all of our sessions.

1:05:51So we play on, I don't know if you're familiar. We play with Roll20, a virtual tabletop. So do I. We chat through Discord, but use Roll20 as a good one. So that's what we used to do. We used to chat through Discord. But, you know, our schedules, I have a kid, my brother has a young kid, like our schedules are not great to meet all the time. So we're always forgetting what's going on between sessions. So I've been using Notion AI to transcribe sessions, but it's only accessible to me. And Notion AI gets all the names wrong because like my character's name is Sprocket, but it's a goblin. We're playing an artificer.

1:06:24It's a whole thing. And so I built a whole custom thing on top of LiveKit, a custom video conferencing platform with an AI agent that does transcriptions inside of it specifically for our campaign. And my character is an artificer. So the whole thing is like as if the artificer himself built this thing. It's a whole thing. I really, really considered doing something with like a transcription deal where we could keep better notes. Because the truth is we all try to keep notes. We're all at best inconsistent. Consistent. Yeah. Sounds like a good side project, Ben. Maybe you should publish it. Yeah, right now it's just for our campaign, but it has our characters in there.

1:07:01It transcribes the stuff and it auto-manages the vocabulary. This is a thing we didn't get into, but the vocab is so tricky to get right when you're in a specific domain. The general purpose transcription model is not going to transcribe all the words. Especially in D &D all the way. Yeah, in D &D. So you can dump all the campaign lore and then it automatically picks out the keywords that are likely to come up in that session and then puts them into the transcription service so it actually gets them right and it generates stories. I have it. My next project with it is to get it to also generate a few photos from the campaign generated through the ChatGPT's latest image model, which is so good.

1:07:36That's what we do. We do every week have like we have a project that is trained on, I mean I say trained on, it has images of all of our characters in there. So when something crazy happens, one of us can hop in and be like, you know, Snoot bites the head off a bat. and it spits out like a crazy oil painting of what we just did. And it's pretty amazing. Yeah, well, that sounds awesome. I'm planning to build that into my thing. That's my next like Claude project when I have a free night. I want to know about it when it works. Well, Ben, thanks again. We really appreciate it, man. And thanks for the D &D digression.

1:08:11Yeah. All right. Thank you, guys. See you later. All right. If you're watching, please take just a moment to like and subscribe. It helps us a lot when we try to bring more of these people along. And is there a way to reach out to Ben? We'll figure that out and we'll make sure you get it. Also, I think I said like and subscribe, didn't I? Hey, go check out the newsletter too while you're at it. The neuron.ai, sign up for the newsletter. We'd really appreciate it. Love to have you as one of our readers. But that's it. And we will catch you on the flip side. We'll be back next week with something else.

1:08:40I'm not sure what. It'll be something. DVD. DVD. And just to put a finer point on it, I will include all the links. And all of the advice from Ben in the blog post that we make after this. And we'll link it down below if you're watching this after the fact. Killer. And on that note, farewell for now, humans. Happy Memorial Day.

From the publisher

Voice agents are moving from “cool demo” to real product infrastructure.


In this livestream, we’re joined by Ben Cherry of LiveKit to break down what it actually takes to build real-time AI agents that can listen, respond, interrupt, call tools, and work in production.


LiveKit is an open source framework and developer platform for building voice, video, and physical AI agents in production.


We’ll talk through the stack behind real-time AI experiences, then build and test a live demo together on The Neuron.


In this live demo, we’ll cover:


🎙️ How LiveKit helps developers build voice, video, and physical AI agents

⚡ What makes real-time agents different from normal chatbots

🧠 How voice agents handle latency, interruptions, speech, and tool calls

🛠️ Why production-ready AI agents are much harder than a weekend demo

🚀 What builders should know before shipping voice AI to real users


And yes, we’re doing a live demo, which means there is at least a small chance the agent talks back at exactly the wrong time. Perfect television.


Guest: Ben Cherry, LiveKit

LiveKit: https://livekit.com/

Ben on LinkedIn: https://www.linkedin.com/in/bcherry-product-engineer

Ben on GitHub: https://github.com/bcherry


Subscribe to The Neuron for clear, useful AI news, demos, and explainers for people trying to understand where this tech is actually going.


https://www.theneuron.ai/

More from The Neuron: AI Explained

All 106 episodes
BONUS: Building Real-Time AI Voice Agents with LiveKit's Ben CherryThe Neuron: AI Explained · 1 h 9 min
Listen in VO