In short
Dev Interrupted Podcast Episode Summary
Episode Title
AI Gets Eyes and Ears | LiveKit’s Russ d’Sa
Episode Overview In this episode, hosts Andrew Zigler and Ben Lloyd Pearson engage in a deep conversation with Russ d’Sa, the Co-founder and CEO of LiveKit. The discussion centers around the evolution of AI technology as it begins to perceive the world through sight and sound, moving beyond traditional text-based communications. The implications for this shift are explored, particularly in high-stakes environments like emergency response.
---
Key Themes and Discussions
- Transition from Text to Multimodal AI
- Historical Context: AI has traditionally communicated through text and clicks, but the future is shifting toward real-time sensory input (voice and vision).
- LiveKit's Role: LiveKit provides the infrastructure for machines to interact using real-time voice and video, impacting applications from emergency response systems to everyday software like ChatGPT.
- The Rise of Slop Squatting
- Definition: Slop squatting is a form of vulnerability arising from AI-generated code that may hallucinate non-existent packages, leading to potential security risks.
- Cautionary Tale: As AI tools become more accessible to developers and non-developers alike, the risk of supply chain attacks increases, urging developers to be more vigilant about dependencies.
- The AI Wars Heat Up
- Google Chrome Divestment: A judiciary ruling requires Google to divest from Chrome, prompting interest from major tech companies to acquire it.
- Competition in AI: As AI technology evolves, competition among companies like OpenAI, Yahoo, and DuckDuckGo intensifies in their quest to adapt to the changing landscape.
- Vibe Coding and Agentic Coding Techniques
- New Approaches: The podcast touches on methods like "Chain of Vibes" to enhance coding practices, emphasizing the importance of human oversight in AI-driven environments.
- Coding as a Product Management Task: The shift requires developers to think like product managers, focusing on the broader impact of their code rather than just the technical details.
- High-Stakes Use Cases
- 911 Emergency Calls: LiveKit's technology has been pivotal in modernizing 911 services, allowing dispatchers to receive real-time video and audio feeds, significantly improving response times and saving lives.
- Security Implications: The conversation highlights the unique challenges of ensuring safety and reliability in real-time systems, especially when human lives are at stake.
---
Key Takeaways
- Importance of Vigilance: As AI technology becomes more integrated into software engineering, security measures must evolve to address new vulnerabilities like slop squatting.
- Emerging Paradigms: The shift toward multimodal interactions necessitates a fundamental change in how developers approach software design and user interactions.
- Real-Time Technology Challenges: With the advent of real-time, AI-driven applications, engineers must consider new failure modes and ensure systems are designed for reliability under critical conditions.
- Future of AI: The episode concludes with insights into the future trajectory of AI, emphasizing its potential to become more embedded in daily life as a co-worker rather than just a tool.
---
Additional Resources
- Explore AI Code Reviews: [An Engineering Leader's Survival Guide](https://linearb.io/blog/ai-code-review)
- AI Collaboration Style Survey: [Discover Your AI Collaboration Style](https://linearb.io/survey/pbl92vsc50i/bT79ARJ9)
Follow the Hosts and Guest
- Hosts:
- [Ben Lloyd Pearson](https://www.linkedin.com/in/benlloydpearson/)
- [Andrew Zigler](https://www.linkedin.com/in/andrewzigler/)
- Guest:
- [Russ d’Sa on Twitter](http://x.com/dsa)
- [LiveKit Website](http://livekit.io)
Support the Show
- Subscribe to the [Dev Interrupted Substack](https://devinterrupted.substack.com/)
- Leave a review on [Rate This Podcast](https://ratethispodcast.com/devinterrupted)
- Follow on [YouTube](https://www.youtube.com/c/DevInterrupted) and [Twitter](https://twitter.com/DevInterrupted)
Conclusion This episode offers a rich exploration of the evolving landscape of AI technology, articulating both the opportunities and challenges it presents for software engineering leaders and teams. As the industry adapts to new paradigms, ongoing discussions and education will be crucial for navigating this transformative period.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Welcome to Dev Interrupted. I'm your host Ben Lloyd Pearson. And I'm your host, Andrew Ziegler. This week has some pretty weird news. We're talking about a new problem called slop squatting, introduced by Code Generation, the AI turf wars, and buyers literally lining up to buy Google Chrome, and some new vibe coding techniques that we've been reading and trying out. What do you want to talk about first, Ben? Well, as much as I love talking about vibe coding, particularly because I just started doing it this week for the first time, this word slop squatting just really has me going. So let's talk about that one first.
0:42Okay, this one's really interesting. So slop squatting is a variation of a vulnerability, a security vulnerability during production called typo squatting. This happens when you think you're installing a package, maybe like a really popular or big package into your project. but maybe you make a typo or you forget that hyphen or you use an underscore instead of kebab case. Whatever happens, you know, you can end up downloading the wrong package and inside of that could contain malicious code from a third party. And slop squatting is the newest variant of that. It happens when you ask an LLM to generate code and it does so so diligently, but in doing so, it maybe hallucinates a package that before didn't exist.
1:25and it can do this using really common combinations of really popular packages during everyday development of common things that you might ask it to do and when these things slip into your project suddenly you're opening the door to a malicious user or you know even like a hacker to get into your application so ben you know what do you think of this kind of phenomenon evolving as people start exploring vibe coding like yourself i really think this is like a cautionary tale for anyone that's running a software team. You know, as more developers and frankly, non-developers to get access to AI tools for scaffolding and creating code, the surface area for attacks like supply chain attacks is exploding.
2:08Unlike traditional attacks like typosquatting, for example, this doesn't rely on human error. You're potentially doing everything right. It's just exploiting something that or an assumption that you can make about how AI operates? And it raises some pretty critical questions for this AI-driven agent-native future. If agents are writing our code, what guardrails are we implementing to stop them from doing things like importing malicious packages? I think what we're really seeing is there's still a lot of unknown unknowns in the security space when it comes to AI. We're learning a lot of new things about how it can be applied to take advantage of you.
2:48And, you know, it's you got to stay up to date on what's happening. And of course, I started vibe coding this week for the first time. Like I am now taking such a fine tooth comb to every single dependency that it tries to bring into my project, which is a good practice. You should always do that. But it also made me realize that even when you know that this is a challenge, it still can be kind of difficult to validate that these packages are legit. Like it's pretty easy to fake a package page on a repository. you know so just like until we have tools that make this a lot more consistent and automatable i think we do really need to be conscious of when we use ai and when it does things that bring in new dependencies into our projects you know we'll talk a little bit more about this later too but it really highlights the skill sets and the in the mindset shift that you have to adopt when working with code generation in this way because you're spending a lot of time now looking up the packages that are going into your project and really understanding the building blocks of it.
3:47And that's because, you know, maybe some of your time that would have been dedicated to coding is now freed up to do this better, higher level understanding of your application. So it all plays together and really highlights the importance of understanding the code you're shipping. Yeah. So I want to talk now about the AI wars, because this is really starting to seem to heat up. So what do we have for that this week? Oh, yes. So there's a whole rumble right now in the tech world because Google has to sell or divest from Chrome. Judges ruled that as part of a monopoly that Google can no longer keep Chrome as part of its portfolio.
4:23And this is causing a lot of large tech companies, people adjacent to search and AI to swarm the scene, literally somewhat like vultures, trying to immediately buy the Chrome browser. Now, why is this important? Obviously, we all know Google is a large company, very large reaching tech portfolio, powers a lot of the modern world that we live in. And when you talk about a fundamental tool like Google Chrome, which is almost ubiquitous now with accessing the Internet, there's a lot at stake. It's a large user base. And this is happening at a critical time when we're completely re-evaluating what it even means to go online and use an application or search for information.
5:04and the ways in which people are doing this are kind of flipping on its head. You have AI that's kind of coming in to search. You know, we've been covering this a lot on the pod. We've been talking about how there's these encroaching wars between going to ChatGPT to search a query and going to Google and getting your AI-generated responses. They're both playing in each other's ball yard. So all of that is to say is that you've got companies like Yahoo, Perplexity, DuckDuckGo, all lining up. OpenAI, of course, wanting to obtain some of this for themselves. So Ben, what do you think of this circumstance?
5:39Yeah, well, first of all, as someone who used to work for Yahoo, I love that they're trying to find a way to make sure they remain relevant in the AI era. And also, I just want to say, I just want to have it out there. I would also like to buy Chrome. I might end up being like TuckTuckGo and not actually being able to afford it. But that's besides the point. I would like to buy it if that's possible. You know, we'd have to talk with them and ask. Yeah. But personally, you know, I would love to see it hosted like under a foundation similar to Mozilla, but it's probably like an unreasonable thing to expect.
6:10And maybe they'll become like an independent company. But it really does not surprise me at all that we have all these vultures circling to scoop it up once Google is forced to sell it. I mean, I think everyone would like to own Chrome. So it's kind of like, yeah, you and everyone else. Thank you. I think, too, is that the key thing here is that the tool we're going to use to search or go online and go to websites, I think the tool we're going to be using in a few years from now probably hasn't even been invented yet. The Internet and the way we access it completely reinvents itself very frequently.
6:41The browser wars, like that used to be a thing. We weren't sure which was going to emerge ahead. You got like Mosaic. Then you had Netscape. Then you had Internet Explorer. Then you had Firefox. And you had Chrome. And then Chrome became huge. And, you know, what's next? I don't know. but I'm sure it'll be something different. Yeah, and I would definitely love to hear more about OpenAI's vision about an AI-first browser because that was something they mentioned as a part of this. So like, it's probably going to be a cool experience. I would really love to see that happen. But this wasn't the only thing that we saw in the AI wars, right?
7:12Wasn't that we have another story involving Cursor and Microsoft? Oh, yes, we do. So there's something that happened very recently on the VS Code extension for C and C++. This is an extension that powers a lot of the tooling within VS Code to work within that language as a very popular part of the ecosystem, right? But recently, Microsoft released Agent Mode for GitHub Copilot. And so this is GitHub Copilot starting to move into things like Cursors territory. Going back to our turf wars, what we just covered with search and with AI, you have this happening as well with code assistance. And so by blocking the VS Code extension from working in things like Cursor, which is entirely within their realm and their right, because this is an extension they have developed.
7:57And if they want to make it for their own applications, they can. They haven't stopped Cursor from using or adopting a different type of tool within the ecosystem because there's lots of open source alternatives. And they have since done so, right? So it really shows that Microsoft might try to close a door to make less of an opportunistic space for their competitor. But because they also control so much of the ecosystem, they get to make moves like this that are quite fascinating that other companies can't quite do. Yeah, I think the AI space is going to get really cutthroat. And we're just going to see this competition heat up, especially with how much money is on the line.
8:34Also, given that nobody has established a defensible moat in this space yet. So everyone is prone to being disrupted that operates here. and I doubt it's a coincidence that Microsoft times this blocking with or blocking the VS Code extensions with the release of their agentic mode for co-pilot and I think on one hand it's a lesson that all startups should learn like it's Cursor's fault for relying on a competitor's product to build upon I also understand why they made that decision it's a tiny company they don't have very many employees they bootstrap their way to to where they are so it's a natural that they're going to encounter growing pains like this.
9:10And it also makes sense to me why they would adopt open source as the solution, you know, because it's a lot lower barrier to entry. Yeah, completely. Yeah, so this week in VibeCoding, we have yet another story on VibeCoding. So what do we have, Andrew? Okay, so this is a really great one that I read. I also shared it on LinkedIn. This comes from Pete Hodgson, and it talks about a prompting method within an agentic coding tool like Cursor or GitHub Copilot agent mode called Chain of Vibes. And really the thesis here is that AI isn't ready for unsupervised coding. And Pete found that a workflow called Chain of Vibes lets him maximize how much he can lean on the AI for coding while still having a really good rigid process with a human in the loop.
9:55The approach consists of driving the AI through a series of very separate and very distinct but fully autonomous vibe coding changes that build upon each other. And between each of them, there's a thorough human review to make sure that things haven't gone off the rails. I myself have had a lot of luck with this kind of technique and definitely encourage anyone experimenting with agentic coding to check that out. But Ben, have you tried anything like that? Yeah, I know you've been using agentic coding for a few weeks now. It's been fascinating to me just seeing secondhand, like learning from you what you've been doing with it.
10:29And as I mentioned earlier, like I just got into like vibe coding or agentic coding this week with Cursor. And it's frankly been an eye opening experience. You know, I still haven't figured out really how to maximize the use of things like rules to facilitate better agentic coding. Like this is kind of what Pete is describing in this article, right? It's like having more discrete chunks of work that have rules that guide it and a human overseeing the process. but it really does like mirror a lot of the approach I've taken with AI, just even outside of coding. You know, my focus has always been on breaking down complex work into a series of smaller steps that can easily be handled by individually purpose-built GPTs.
11:14And the reason I do this is because it improves consistency, but it also makes it much easier for you to inject human judgment into the loop to make sure that the output from one GPT is good enough for the input of the next. And there's a couple of other tips that really stood out to me that are things that I've already adopted. Clear context frequently. Like personally, I treat almost every interaction with a GPT as ephemeral. Like the moment that I'm done with it, it is deleted from existence where I never think of it again. Or if I feel like it's going off the rails, like I just start over from scratch with a fresh GPT prompt.
11:50And then he also mentioned using the right AI for each task. Like when you start breaking down these complex chains into discrete tasks, you might end up finding out that one model works for a certain task much better than another model, but then for a different task, the reverse is true. So, you know, I really, like this approach really makes it easy to set yourself up for success, but also experiment with the tools that are out there. Yeah. And this article does a really good job of kind of setting the perspective correctly of what you should be thinking when you open a tool like cursor or when you go to amazon queue in your command line when you open an ide now that has these capabilities you know you're putting on more of a product manager hat than you are a developer hat instead of focusing on the exact code lines that you're going to write you're instead focused on what is the purpose of what i'm building what's the intended impact and what do i need in order to get there.
12:48And then your goal becomes really as that product manager to rally all of that information together and get everyone on board. And when I say everyone, I mean all of these ephemeral little chats that you're about to have, like what Ben just described with the IDE. And another guide that are two sets of guides actually that go really well in depth with this. They're from Joffrey Hunter. They're on his website. We've actually linked them before in the download. So we'll be sure to link back to those as well if you haven't checked those out. It breaks it down into a really repeatable method, similar to what Pete Hodgson is kind of touching on here.
13:24And you start really by understanding what it is that you're building and writing very clear documentation for yourself and for the agent to use. Then you spend some time writing the bounds and the rules and the constraints of what you're building. In cursor, this might even mean writing rules for the project you're doing. And then finally, you're going to iterate. across all of those ephemeral sessions. You're going to reference that well-written spec. You're going to point to those strong, well-defined rules. And you're going to iterate until you get the results you want, reviewing every step of the way to make sure that you don't, of course, get slop-squatted, like we mentioned earlier.
13:59So it really kind of changes your whole perspective about what you're doing when you go into the IDE. And it really makes me wonder, to our listeners, are you experimenting with agentic coding right now too? I think a lot of people are trying it out for the first time and we want to hear about your experiences. I'm personally posting on LinkedIn every week and learning with others in the open about how to use things like Cursor, Amazon Q, Windsurf, all these other tools. So come learn with us and share your experiences too. Yeah, everyone's learning right now. And this is a good time to remind our audience that we have a Substack newsletter where we share a lot of these stories as well as a LinkedIn community where we're sharing this as well.
14:39So if you want to be the first to find out about all these great guides and get these tips delivered straight to your inbox, make sure you're subscribed to our sub stack. It's the best way to get all this news. So we've got a bonus story, too, and I wanted to include this one because I like to walk and I use crosswalks frequently. What is the story this week, Andrew? Okay, so this is a really weird one. Recently in Seattle, some of the crosswalks appeared to have been hacked. And people at first were not quite sure how this happened. But when you would go to the crosswalks on some intersections and press the button, instead of it telling you to wait or to cross, instead you were greeted by a deep-faked voice of Jeff Bezos or Elon Musk or other tech billionaires.
15:24And they'd be talking to you about all sorts of stuff. And this is a neat assortment of weird technology all coming together to create what was ultimately a viral moment. Because when you boil it down, it's a very simple vulnerability, actually, that they took advantage of. It turns out that a crosswalk is a lot like a router. It comes with factory settings, including a factory login over a technology like Bluetooth. So it really was probably just as simple as somebody who knew the model of the crosswalk and had their Bluetooth open and was able to connect with a default password and username to it.
15:58And then once they're in, They could upload, you know, an MP3 of what it should play when the crosswalk button is pressed. But what makes this a very impressive and very interesting usage of technology is ultimately they used modern tools like AI to deepfake these voices. They used very modern tools and the things that we're all worried about in terms of like misinformation to rapidly create a high impact message that then hacked something more than the crosswalk. it hacked social media because it became a viral story suddenly you had people going out to these crosswalks and recording it and sharing it on x and on blue sky and all these places look what this crosswalk is saying and then the whole world was talking about it now we're talking about it and that's ultimately what this kind of hacking in the wild is all about is making a message known using the tools available so quite a fascinating one what do you think ben yeah so first of all i want to just say put this out there like safety first like if you're going to do something like this please don't screw up the crosswalk so that like a disabled person can't use it like it's already dangerous enough being a pedestrian in many american cities but you know this is a tale as old as time like somebody's setting up a piece of digital equipment and just not bothering to change the default password like come on that's it's like security 101 like nobody should ever be doing that in this day and age.
17:22But the AI brings a very new twist to this. And you and I have been talking about this concept of code as art. As code gets more accessible to more audiences through AI, I really get the sense that we're going to see more emergencies of people just using AI and using code in unique ways that we had never really considered before. Again, going back to our slop squatting example, there's lots of unknown unknowns out there. AI is going to keep opening up new attack vectors. Stay tuned to Dev Interrupted. We're going to keep you on top of all the stories as they develop. So, Andrew, who's our guest this week?
17:58Oh, yes. This week's guest is a really interesting one. We have Russ DeSah, the CEO and co-founder of LiveKit. And he's really pulling out the curtain on where the future of AI is heading. And I don't mean just in the chatbots that we interact with now or even the agents that people are building and everyone wants to talk about. but instead interacting with us on our terms, being able to see, hear, feel, and understand from the real world using multimodal inputs. And the impact of this technology is profound. Talk about saving lives even on 911 calls. So stick around. You don't want to miss this one.
18:34Are your code reviews slowing you down? With Linear B Automations, you can transform the way your teams review code. With automatic AI-powered PR descriptions and code reviews, your developers get instant feedback on every pull request. Combine that with smart AI orchestration, and you can cut the noise and boost your productivity. LinearB AI flags bugs, suggests improvements, and keeps your team focused on what really matters, building great software. Say goodbye to review fatigue and hello to faster, higher quality delivery. Head over to linearb.io to learn more about incorporating AI into your code reviews.
19:12Hey, everyone. Joining us today is Russ DeSau, the co-founder and CEO of LiveKit, where they're building the nervous system for multimodal technology that interacts with the world through voice, vision, and real-time understanding. And what they're building at LiveKit, it's not just about its applications in AI. It's about how all modern systems are starting to sense, interpret, and act in real time. And to give our listeners a sense of this technology's impact just right off the bat, you know, right now, LiveKit delivers voice to millions of ChatGPT users. It even saves lives on 911 emergency calls, and we're going to talk a bit about that.
19:51It's also used in live streaming and robotic systems across the world. And in today's chat with Russ, we're diving into the major shift that's happening in software development that's currently underway and learning how your team can embrace those multimodal opportunities in your work. So Russ, welcome to the show. Thank you so much, Andrew. It's lovely to be here and hi to all of the listeners as well. It's great to have you. And you know, Russ, you and I had such a great chat and preparation for this. You know, your take on technology, it was really refreshing. It stuck with me. I've been really excited to bring you onto the show, especially this idea that you put in my head about how we're outgrowing a traditional request response model.
20:33And I want to dive into that a bit in our conversation today. But first, I think to benefit our listeners, let's zoom out a little bit because we've been talking about the shift into a real-time multimodal world. There is this huge shift happening with technology and not something that I foresaw when we started LiveKit, but then something that kind of ended up happening. For me, that was what we started to do with OpenAI and building voice mode with them for ChatGPT. And it was at that moment that I kind of like looked up a little bit and I was like, well, where does all this go? Right? Like, this is a really cool feature.
21:08I can talk to an AI, you know, using my voice, but where is all of this headed? And I think the thing that I realized back then, this was like August, 2023 is that the internet or let's say the web really was not built for multimodal real-time audio video, right? When you type into the browser, HTTP colon slash slash, like what you're typing in is you're typing in a protocol there. And that protocol, HTTP stands for hyper text transfer protocol. What you're doing is you're taking text and you're transferring it from one computer to another computer. It's not hyper voice transfer protocol or hyper video or vision transfer protocol, it's hypertext.
21:59And so, you know, just that word is telling, like the way transferring high bandwidth data, like audio and video over a network, it fundamentally requires a different approach, a different paradigm than transferring text over the network. The way that we interact with computers today, for the most part, is we open a browser, not every single use case, right? Like we use Zoom sometimes and we use Discord and all of that stuff. So it's, we are definitely transferring audio and video for some of these applications that we use. But predominantly, what you're doing is you're in a web browser, you're navigating to a web page, you're clicking into a form, you're filling out some fields, you're clicking submit or other buttons on that page.
22:43And that kind of paradigm or interaction model is the stateless web application model, right? So you click a button, some text is transferred to a server somewhere, that server receives that text and what you want to do, it looks up who you are in a database, it has some side effects that are generated. So it runs through some business logic, say that you're trying to get a reservation at a restaurant, right? And it's looking up like, are there spaces available at that restaurant for your party size? And then if there are, then it's going to put a record in the database for that time slot for that restaurant saying that you have the reservation.
23:22And then it's going to send back some text that gets rendered on that website that you're on saying, you have booked your table, show up at this time. And so that's not really a latency-sensitive application. And most of the applications on the web today are not. The way that you're going to interact with computers in the future, right, at a high level, what we're trying to build, I guess, you know, in society now is AGI, right? If we're trying to build AGI, like what is AGI? In my opinion, humans are tool builders, right? We create these things like hammers and nails and screwdrivers and planks of wood and all this stuff so that we can construct things and solve problems.
24:04And what we're building now is we're kind of building the ultimate tool, which is a tool builder. We're building ourselves, right? The mirror in some ways. And it's a computer that can behave like a human being, talk like a human being, maybe in its ultimate form is indistinguishable from a human being. I'm not trying to do any scare-mongering or anything like that. It's a computer that very much behaves like and mimics a human being. And when computers were kind of these not as intelligent machines. We had to create something like a keyboard and a mouse and adapt our own behavior. I know you're probably a bit younger than me, but I'd take a keyboarding class in school, right, where I need to learn how to type on this layout of keys.
24:53We have to adapt our behavior and learn how to give the computer information so that we can use it as a tool to help us, right? The bicycle of the mind. Yes. To use Steve Jobs' kind of phrasing, now we're almost creating the mind itself. And if you are going to build a computer that is as smart as a human, you no longer have to adapt yourself. The computer actually will adapt itself to you. And the way that you give information and communicate with other human beings is the way that you're probably going to communicate and behave or give information to that computer that is very human-like. And so humans use eyes, ears, and mouths.
25:35The computer of the future is going to use cameras, microphones, and speakers. Those are the equivalent sensors to kind of the humans, eyes, ears, and mouths. If you're going to build a computer that takes in natural human input and output in the same way or a similar way, you have to almost change the way that you build applications for that computer, right? That run on that computer. It's not where I click a button and then I wait and then like a database kind of looks up information and some actions are taken and then a response gets generated some number of seconds later. This is more similar to how you and I are communicating with one another right now.
26:17Like you are constantly listening to me. I have your attention. You're figuring out whether I'm done speaking or not, or whether you should interrupt with the next question. You're keeping this rolling context growing in your mind of everything that I'm saying, and you're committing some of it to memory, and you're going to come back to maybe something that I said or focus in on it. It's this constant connection, and the data is streaming to you throughout this entire time that we're talking, and you're processing it in your mind and also at the same time listening to new information that's coming to you from my mouth and through my movement.
26:53And so if you're going to build applications for that future, they're going to look just very different. So when we're talking about shifting how people are looking at, are we going to be using technology in the future? It's really a fundamental shift from getting out of the box that we put ourselves in, of using text and converting everything in our world to text, to talk to a machine, which is the traditional model, and more about opening it up and bringing, elevating the model or rather the computer up to our level, allowing it to participate in our own senses, have that same kind of understanding of the world.
Read the full transcript
27:28And then it unlocks new use cases, more ability for that technology to have impact on our world. And everything that you're building so far, of course, there's the trend of what you're describing of ultimately a very intelligent machine is going to need very intelligent senses to interact with our world. But along the way, there are so many use cases as well that evolve and unfold once you bring sensory input into the conversation. And you're sitting right at the helm of that. You're seeing it being applied and used in a lot of different places. What are, do you think, the strongest or most impactful use cases that you like to tell people about that this technology is transforming today?
28:14So my favorite use case is actually the one that I think would tie in very closely or most closely with what you were just saying about how this kind of new paradigm or where you interact with the computer and what kind of abilities it is imbued with now that it has kind of human level senses of vision and hearing. I think it's self-evident in this 911 use case. So LiveKit today is used by about 25 % of 911. So 911 calls, right? There's a company out there called Prepared. And what they do is they deploy LiveKit server or media server in dispatch centers around the country. So about a quarter of them or maybe close to a third of them.
29:05And, you know, it was started by three folks that were college kids at Yale in a very different kind of generation from me. So when I had just graduated college, Bowling for Columbine had come out and it really kind of brought to light this tragedy that we have in this country around school shootings and safety at these institutions. And, you know, that was for me. And then like, I think maybe 15 years later, the founders of the, of this company, they were at Yale and they're going through school shooter training and there's a school shooting happening every, I think three times a week or something like that around the U S it's a really shocking, uh, kind of amount of this happening.
29:54And so they were going through this school shooting training. And one of the things they thought about was like, well, what if I could actually broadcast what's happening live? Like any student could take out their phone and start to broadcast and provide situational data to the police or the authorities that are coming to deal with and handle this issue and, you know, restore safety to the campus. and they were talking to one of the officers that were there on campus during the training about this thing that they were working on. And that officer said, well, you should go talk to the dispatch center in New Haven for 911 because they have been built on the telephone system from like the 70s or 80s.
30:37The technology hasn't changed. You still have to get on the phone and tell them where you are and explain the situation and they don't have any eyes. you know they don't get gps data streaming to them about location or anything like that and so they've been dreaming of having this richer situational data in the dispatch center right for these 911 calls and so they started they did a pilot in new haven and i think they did another one in nevada and within the first week someone had a heart attack and called 911 the dispatch agent sent a text to their phone with a mobile URL. They tapped on the user who was calling, tapped on that URL, and they were in a mobile browser streaming audio and video and GPS data to the dispatch agent.
31:25And they had someone hold their phone and the dispatch agent coached them on delivering CPR to this person who had had a heart attack and was unconscious. And they saved that person's life. And now every single week, a very similar story to this where someone has a heart attack and is revived through a video call with a dispatch agent through LiveKit, of course. So it's very rewarding and impactful for me. It saves a person's life every single week. And so, you know, it's but it's something that just was not possible before we had this ability to really like stream audio and video on demand in real time with low latency and kind of like teleport that dispatch agent to the scene of the emergency.
32:08So just an incredible use case and emblematic of this new world that we are entering into. What an amazing, impactful story, like an incredible first usage or rather like, it goes to show how quick the impact of that came. Like they put that in place and then it was immediately useful and immediately saving lives. So it really speaks to like this really big need in our world and a disconnect between how we live in it and how we try to use technology to influence it. So it's really kind of hinting at here how LifeKit is bridging things that before were not bridged. And with that, you get these huge gains, like in this case, saving someone's life or in other cases, responding to real-time scenarios.
32:53Being able to use this technology to drive that kind of good, I'm sure, has been really rewarding. The way that you describe it too, I love your passion for how people would use this technology and what it would mean. And I also like how you're talking about the evolution of, you know, we took our brains and we turned them into text in order to use computers. And now we need to bring the computers to us. We need to help them understand the world that we live in. And it reminds me so much of when mobile really came on the scene and about how everything tried to get shoved into mobile, right? And we tried to figure out how do we use everything on mobile.
33:27And now it's a growing pain. We turned things that did and didn't work into mobile experiences. We learned not to just make that standalone mobile website. And we all figured out responsive design. And that required so much effort. And now it's innate to our world. And our world is, in fact, almost built mobile first. So it really kind of hints at these possibilities of being in a multimodal first world, where by default, you're interacting with it that way. Do you see that in our future? I definitely do. I think that technology adoption kind of moves in these phases, right? Where you have these early adopters and then you kind of have this period of time where people are taking like the paradigm that's already at scale, I guess.
34:12And then there are these S-curves, right? And you're kind of taking the paradigm that's already at scale and you're adapting what's at scale to work in this new paradigm and adopting the technology in as almost like you know it's not meant to be like reductive or diminishing the value of doing this because we all kind of do it at first which is you know you kind of bolt it on right like are you you figure out a way to integrate that new technology into what exists already but then after that wave and the technology or the paradigm shift kind of gets popularized to a degree and you're firmly in like the adoption kind of neck of the S-curve or in that middle section, then you start to have products and companies get built, which are kind of native to that new paradigm, right?
35:04And so in mobile, it was like, it was Uber and it was Snap, it was Instagram. And it was, it was these companies that were kind of like mobile only, right? Or mobile native, thinking about what they were building as mobile being the only entry point into those services and how the interaction should work and how the product should work. And so I think we're definitely going to see that with AI as well. I think it's on an accelerating timeframe though, which is kind of fascinating. Like everything, I think all technology is an accelerating timeframe. So any point along the curve is also going to happen a lot faster as these things kind of move along.
35:45the progression that I see, I think that there's, there's kind of like three things. I think the first one is we're kind of in this like co-pilot moment right now. I don't think co-pilots are actually going to go away just to be clear. I don't think it's like the co-pilot is the bolt-on thing. And then like, it gets completely replaced by something else. I think that it starts off with this co-pilot use case where your co-pilot is your virtual voice assistant in the chat GPT world or it's cursor that's sitting in your code editor, or it's another experience where you have this assistant that is there to like help you with stuff.
36:22And I think at first it will help you by augmenting your abilities, like inserting itself in certain parts of your workflow and, uh, you know, generating content or information for you that you can leverage, right. Even just thinking as basic as going into notion and having it rewrite a paragraph that you wrote or something like that. That's like a co-pilot use case. And I think if you take that all the way to its end, I think what you end up with is something like Jarvis from Iron Man, where it's like, well, hey, like I'm envisioning something here instead of it being more engineering driven.
36:55Now it's more design driven where it's like, hey, I'm like the maestro and I'm telling you kind of what I want. And then you're kind of generating and doing a lot of that mundane work for me, putting something in front of me. And then I'm more of the editor, right? I think that that's kind of where the co-pilot use case or implementation is taken to its logical end. But beyond that, what ends up happening is you end up going from co-pilot to co-worker. That's more of like this agentic thing that's going on or that people talk about. That term is kind of, you know, we have a product called agents too.
37:27It's like the term doesn't mean anything anymore. But this agentic kind of workflow thing is you're instead giving like instructions or what you desire to this AI entity. And it's going and it's doing the whole thing itself and then coming back and you're having kind of a meeting with it. And you're like making sure that you're in sync or you're helping give it some more information if it needs to refine the work that it's done. And so I think that that's going to be like the next stage of all of this. And then I think the third stage is we now have a model or a set of models that are so smart, they can go and do work themselves and all of that stuff and execute on pretty complex tasks.
38:07well okay now I need to embody these things and put them inside a thing that's shaped like us and can move like us and then it's like it's literally like a co-worker or a friend or a companion that can go and and navigate the physical world and do things in the physical world that help me and you know increase my productivity or maybe just increase my happiness. Yeah I think it'd be a really interesting journey I agree with you about the natural end of co-pilot I think that even like the parallel that we just drew between this and like mobile, the same thing can be said about like chat assistance with AI and mobile.
38:45It's the same idea like, oh, it's our first taste of the technology. We're just trying to shove it into the quickest to delivery format that we can that we're used to. And right now that's a chat conversation. That's like the most intuitive way to extract things from it. But it will obviously evolve into other forms that are more specialized and that kind of have more intention behind them. And I want to take the conversation at this point and kind of shift it into how this kind of thinking is going to impact engineering leaders in our space. And about how they're going to have to really rethink how they build software, how they train their teams, and also how they are going to bring their software through its entire, you know, SELC.
39:26Like how does working with this type of technology impact it? And so I have some questions that are kind of top of mind for you. They're kind of things that might be burning questions in our audience mind as well. One of them is really like when you're working with a real-time technology, I think the stakes are a lot different. And failure, real-time failure looks a lot different than like a server going down somewhere on like a request response architecture. I'm wondering from your standpoint, what are some fundamental things about engineering that you and your organization have had to rethink when building something like LiveKit?
40:03I think that there's kind of two sets of challenges here, right? There are challenges and questions and approaches for engineers, developer teams, companies that are kind of building applications for this new world. And then there are the challenges that we've kind of had to solve as a company or as a product for our product. And these are kind of two distinct buckets in a way, just because the problems that we've had to solve for LiveKit around failure modes, technical challenges, scaling, all of that, reliability, those problems we had to solve independent of AI. just because what LiveKit provides is not too dissimilar from what we had to provide during the pandemic when the open source project started.
40:52It was that we were building infrastructure that made it really easy for a human to connect with another human anywhere in the world, right? And the way that human is connecting with that other human is using cameras and microphones from either of our computers and we're teleporting that data over some wires to the other person. That's what we're doing right now as well. And what we're doing is we're leveraging the same technology, but you're connecting those cameras and microphones with the machine. You can use the same technology because the machine now takes the inputs in the same way that you do, right?
41:27Looking at a camera that is capturing something that I'm seeing or a camera that's pointed at me. So either a camera pointed at me or a camera pointed out at the world. and so it's seeing what I'm seeing and it's hearing what I'm hearing, right? When I speak to you, you're like listening to me with your ears and the computer is also doing something very similar. So the same technology can be used for connecting humans to other humans as it can be for connecting humans to machines. And thus the problems to solve there around like making sure it's as low latency as possible, making sure you have servers closest to the edge for any user or AI model to connect to or agent to connect to, making sure that if a server goes down, you can transparently fail over to another machine so you can kind of horizontally scale these servers, and then making sure that you can load balance across these servers and they can scale up and handle millions and millions of concurrent connections.
42:22All of those problems that we had to solve for video conferencing and for live streaming before this whole AI thing happened, And there's almost 100 % overlap with kind of this proliferation of AI and voice and vision interfaces to an AI model. So the problems we've had to solve from that perspective have been the same. Now, there are new problems that we've had to solve that are aligned with kind of what I would say the challenges are for the application developer or the engineering teams out there that are using something like LiveKit to build an application. There, it's really about how do you make sure that you do this safely?
42:59Right. When I talk to another human being, you know, there's, there's a few things that let's say with an AI model, if you were to draw a parallel to a human being, like if I said something to you and you like nodded your head and then you just said nothing, or you shut off your camera and disconnected from this call or whatever, that would be a pretty weird experience. Right. So like reliability is definitely something that we have to care about. Make sure that, you know, you have that failover and and and there's this continuity between the conversation that I'm having with an AI in the same way that you want that continuity when you're in a Zoom call with someone.
43:31unexpected scenario. And so you got to make sure that those fundamentals work. And that's, you know, what we take on as our job. But then what we will help developers do, and we're going to be building stuff in this direction, but that developers also have the burden of solving as a challenge is how do you make sure that the model is like saying the right thing, right? Or how do you make sure it's not hallucinating? And let's just say like swearing or saying something inappropriate? Or if it's a model that is generated a video model, how do you make sure that it's not like generating inappropriate images or things like that?
44:07I think that what is tricky with real-time and with audio and video in particular is that when it's real-time, you do not have a lot of room for corrective measures, right? Because this information is being generated on the fly. The second problem versus text is that you can make assertions about text, like they're strings, right? Of characters. And you can, we've been writing assertions about strings of characters in code for a long time. But now you have to suddenly like have assertions around like waveforms or around images and how fast can you, can you make those assertions, right? You got to like take audio that is coming out of a model and you have to maybe like send it to another model that is looking at or somehow determining whether that there's appropriate, you know, it's inappropriate or not and flagging it.
45:00And then like, if it's inappropriate, like, how do you correct for that? Have you already delivered the audio bytes to the user's device? Like, do you kind of do like an ask for forgiveness thing or like an ask for permission? And, and so that's a challenge, And one of the ways that you can kind of solve that challenge is through like simulation and evaluation where you have like one model or a human even sometimes like go and run through like these testing calls or you're spinning up voice agents that are going and like having a call with another voice agent and making sure that it's doing the right thing.
45:36What you're ultimately trying to do here is build a statistical confidence that whatever you are going to ship to users is going to act and behave appropriately. That's fundamentally what you're trying to do. Between two humans, the way that you do that is through, I don't know, like reading their resumes or making sure you have common LinkedIn connections or whatever, right? Like social proof. There are all these mechanisms that humans use to kind of build trust, including just looking at someone, right? Like we make these judgments just based on like kind of how people appear. And then, and then there, then, oh my gosh, it like enters into this other question about like, how do you make sure that you're not like having some kind of inherent bias, right?
46:18But also like making a good judgment around safety based on the input that you're getting. It's, it's really tricky. It's like a world that we're still figuring out, but it's something that is a real thing to consider when you're shipping these systems. I really like how you explained that in terms of at the end there, especially where you have like a test case, right? Like we write assertions on strings, on math all the time. That's a solved problem in tech. And that's also like table stakes, right? Like you're making an application and you're shipping it. Like you better have some unit tests in there somewhere or someone's probably going to be unhappy and it's probably going to be you and your users.
46:55And so in the case of like using multimodal tech, you're kind of solving from zero on a lot of those test case scenarios and figuring out how would a human validate this kind of input from another human? You know, you and I are having conversation, like you said, like I'm nodding my head, you're nodding your head. We can see that from each other. But if I change my tone of voice or if I close my computer or if I turn off my camera, like it's going to fundamentally shift the conversation. Right. And so being able to react to those in real time to understand when that needs to change behavior of an application is really important.
47:32But really just kind of solving that totally unsolved problem of testing a voice conversation. I like the innovative idea of spinning up a voice agent. And it's almost like a test case. And it's like, here's your script. And that is like the test, the unit test or something. And then you're going to go through this flow. And then that allows you to scale it. And I think that's ultimately the big thing, right? As you build all of these solutions for keeping your technology secure, keeping it delivering, going from zero to one whenever you come across one of those new scenarios. Yeah, exactly. Very early days for it, right?
48:09Like we're still a little bit in the stone ages of all of this and trying to figure out how do we actually scale these systems in a way that feels safe? Because the other issue is if you, the cost of getting it wrong is very, very high. It's not as severe as say like a self-driving car killing someone. When that happens, you know, everyone's like, well, we can't use this anymore. Like even if a self-driving car is safer than a human driver, there's something weird about like when you give your agency up to this machine, if it gets it wrong, then all of a sudden we kind of have this expectation that machines are deterministic and that they get it right every time.
48:49But AI, the magical thing, the magical and scary thing about this new wave of AI is that this approach to it is probabilistic. We have a probabilistic computer, right? I'm not going to say it learns exactly the same way that we do or thinks exactly the same way that we do, but there are similarities. And ultimately, at a high level, we are probabilistic machines and so are these. And so we have to kind of adjust our expectations as well as a society for what the outputs might be and start to get a bit comfortable with some of the adversarial cases that may come up. But at the same time, devise ways that we can start to build more and more statistical confidence and kind of bring some determinism or at least statistical determinism to these systems.
49:42Right. We can't throw away those things that have worked for us before just because the tables changed a bit. We still have to care about being certain. We're working with technology that is uncertain. It's probabilistic, like you said. And with that opens up so many opportunities for it to do things that a deterministic machine cannot. but with it also so many pitfalls, so many first-time errors or issues that you have to encounter, so many learning issues. And when you're talking about technology that can interact with our world, you can talk about even like the stakes being as high as like someone's life and whether it's like someone's life is getting saved on a 911 call or someone's in a car that's driving itself.
50:20It opens up a lot of interesting like societal questions about agency, about who is responsible for the machines that make these decisions Is it the users or the people who build it? Is it somewhere in between? Because it can't just be the machine itself, right? And so these are a lot of questions that I think we're all still figuring out. It's really interesting to learn how those questions kind of are imprinted in the technology itself as it's getting built. Like you're testing the safety of a voice conversation. You're putting things right on the edge. That way you can have as low as latency as possible.
50:56You know, you're solving the problems in different ways. with them, there comes different problems. I think something too that really stands out to me is how it impacts high stakes environments. And like the 911 use case is incredible. I think those are huge frontiers. And just in general, this technology, I think this is just the start of it. I think we're talking about something that's still in really early stages. I'm excited to see where LiveKit goes, how this kind of technology takes off. And, you know, maybe we can talk again more in the future as that journey kind of continues, because I think it's going to be a really interesting one.
51:27And this was a total blast for me to have you on the pod, to kind of dig into your brain about how you're thinking about building technology, the different ways you and your team have to orient your mind around delivering this kind of software. Russ, where can people go to learn more about LiveKit and what you're building? You can check out a few different places. So the website, of course, is livekit.io. I think also important, LiveKit is an open source kind of project and company. We built our commercial offing around our entire open source stack. So github.com forward slash live kit is where there's all of the repos and the code and you can self-host it and build with that without necessarily giving us money.
52:11Though if you want to give us money, I won't complain. And of course, live kit is also on X. So x.com forward slash live kit. You can also DM me on Twitter. I'm at DSA, three letters on that keyboard. that you won't be using in a few years from right to left. But it was really a pleasure to chat about this stuff. And I'd love to do a check-in down the road and see where we've ended up and how things have progressed. Yeah, absolutely. You know, we'll definitely be staying in touch and we'll get those links added to our show notes as well so that our listeners can go check out the project, can learn a little bit more about its impact.
52:49Thanks for joining us on today's episode. Clearly, if you're hearing me say this, that you really liked it, you stuck around to the end. So thank you so much for making it this far. If you are listening, please be sure to go and like our podcast wherever you are listening to it, as well as read our Substack newsletter. It comes out every Tuesday. It's where we deliver our podcast. If you're only listening to the podcast, you're only getting about half the story. So be sure to check out that newsletter. Be sure to interact with us on LinkedIn. You know, Russ gave you a lot of great places where you can go find us.
53:17You can also find us on there. We'd love to continue the conversation and we'll see you next time. Thanks so much, Andrew. Bye.
From the publisher
We've spent decades teaching ourselves to communicate with computers via text and clicks. Now, computers are learning to perceive the world like us: through sight and sound. What happens when software needs to sense, interpret, and act in real-time using voice and vision?
This week, Andrew sits down with Russ d'Sa, Co-founder and CEO of LiveKit, whose technology acts as the crucial infrastructure enabling machines to interact using real-time voice and vision, impacting everything from ChatGPT to critical 911 responses.
Explore the transition from text-based protocols to rich, real-time data streams. Russ discusses LiveKit's role in this evolution, the profound implications of AI gaining sensory input, the trajectory from co-pilots to agents, and the unique hurdles engineers face when building for a world beyond simple text transfers.
Check out:
- AI Code Reviews: An Engineering Leader’s Survival Guide
- Survey: Discover Your AI Collaboration Style
Follow the hosts:
Follow today's guest(s):
- Website: livekit.io
- GitHub: github.com/livekit
- X: x.com/livekit
- Russ: x.com/dsa
Referenced in today's show:
- The Rise of Slopsquatting: How AI Hallucinations Are Fueling a New Class of Supply Chain Attacks
- Has the VSCode C/C++ Extension been blocked?
- OpenAI tells judge it would buy Chrome from Google
- Chain-of-Vibes | Pete Hodgson
- Seattle crosswalk signals with deepfake Bezos audio may have been hacked with just a cellphone
Support the show:
- Subscribe to our Substack
- Leave us a review
- Subscribe on YouTube
- Follow us on Twitter or LinkedIn
Offers:
