In short
Dev Interrupted Podcast Episode Notes
Episode Title
Your keyboard is the real bottleneck | Wispr’s Sahaj Garg
Podcast Description Dev Interrupted is a podcast that focuses on software engineering leadership, featuring discussions on strategies, struggles, and stories from high-performing software teams. The hosts—Andrew Zigler, Ben Lloyd Pearson, and Dan Lines—interview industry experts and cover relevant industry news.
Episode Description In this episode, Andrew talks to Sahaj Garg, co-founder and CTO of Wispr. They discuss the limitations of traditional voice dictation and how Wispr is revolutionizing the way developers communicate their intent through contextual models, minimizing the need for corrections. The episode covers the engineering challenges of turning raw thought processes into clear, actionable artifacts, the importance of sharing context in teams, and Sahaj's framework for experimenting with new tools in the fast-evolving tech landscape.
Key Themes and Concepts
- The Limitations of Traditional Voice Dictation
- Historical Context: Traditional voice dictation tools have been unreliable, leading to user distrust.
- User Experience: Past experiences with inaccurate transcription have caused many to abandon voice technology.
- Wispr's Approach
- Contextual Models: Wispr seeks to understand the user's intent by incorporating context, allowing for more accurate voice-to-text translations.
- Zero Edit Rate: The goal is to create outputs that require no corrections, enhancing user trust and efficiency.
- Engineering Challenges
- Speech Models: Most speech models are simplistic, merely converting audio to text without understanding context.
- Two-Part Solution: Wispr combines voice recognition with contextual understanding to improve clarity and usability.
- Communication and Workflow
- Shared Context: Effective communication within teams involves understanding shared context, which Wispr aims to facilitate.
- Asynchronous Communication: Voice technology can enhance async work by allowing developers to express incomplete thoughts more easily.
- Leadership and Experimentation
- Adapting to Change: Leaders must encourage their teams to experiment with new tools and share their findings to foster a culture of innovation.
- Continuous Reinvention: Companies should aim to reinvent their processes every three months to stay competitive and relevant.
Key Takeaways
- Voice Technology as an Amplifier: Voice tools can significantly improve communication efficiency, especially in remote or hybrid work environments.
- Focus on Intent and Clarity: Understanding user intent is crucial for developing effective voice interfaces.
- Encouragement of Experimentation: Leaders should promote a culture where team members are incentivized to try new tools and share their learnings.
- Simplification in Tool Selection: Leaders should opt for simple, effective tools rather than complex solutions that complicate workflows.
Conclusion The conversation emphasizes the need for evolving communication technologies in software development. By leveraging voice technology effectively, teams can streamline their workflows and enhance collaboration. Sahaj highlights the importance of continuous experimentation and being adaptable in the fast-paced tech environment, urging leaders to embrace these changes for greater productivity.
Follow the Show and Hosts
- [Dev Interrupted Substack](https://devinterrupted.substack.com/)
- [Dev Interrupted LinkedIn](https://www.linkedin.com/company/linearb/)
- [Dev Interrupted YouTube Channel](https://www.youtube.com/@DevInterrupted)
Follow Sahaj Garg
- [Wispr Flow](https://wisprflow.ai/)
- [Sahaj on LinkedIn](https://www.linkedin.com/in/sahajgarg/)
Additional Resources
- LinearB Offers:
- [Start Free Trial](https://linearb.io/start-free-trial?utm_source=podcast&utm_medium=referral&utm_campaign=devint-shownotes&utm_content=shownotes)
- [Book a Demo](https://linearb.io/book-a-demo?utm_source=podcast&utm_medium=referral&utm_campaign=devint-shownotes&utm_content=shownotes)
This episode provides valuable insights into transforming communication in tech through innovative tools, setting a course for leaders and developers to explore new methods in their workflows.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding the Shift to Voice
0:46 to 2:14
Discussion on the shift from traditional coding to voice technology and its implications.
“if the limiting factor in your organization isn't actually like a context window, but how quickly you can get the right context out of someone's head?”
Building Trust in Voice Technology
2:15 to 4:06
Exploration of rebuilding trust in voice technology and achieving zero edit rate.
“And it's actually one of the reasons why we almost never built the product that we did.”
Engineering Dynamics and User Experience
4:07 to 6:04
Insights on balancing model precision and user experience in product development.
“So I'm curious in how when you solve the problem of like rebuilding trust and doing it in this two part way and acknowledging that it's a nuanced engineering problem.”
The Role of Feedback Loops
6:05 to 8:01
Importance of feedback loops in product development and user experience enhancement.
“And the other half is the user experience, right?”
Adapting Tools for Different Roles
8:02 to 12:16
How different roles utilize voice tools for effective communication and project management.
“This is what we call like no brainer bets or scaling here internally.”
Communication Challenges in Remote Work
12:17 to 14:02
Addressing communication pitfalls when using voice tools in remote environments.
“But I think those two kind of give a sense of what it looks like.”
Caution Against Communication Anti-Patterns
14:02 to 14:40
Discuss the importance of clear communication when using tools.
“So that's the one thing to kind of caution against when you start using tools like this is, hey, what's the way in which it could impede really important live and productive conversation.”
The Nuances of Communication Tools
14:40 to 17:40
Examine the complexity of user needs and communication styles in tech.
“And you're trying to follow and trying to understand.”
Voice Technology and Contextual Understanding
17:40 to 21:05
Explore how voice technology can improve context understanding in communication.
“is not trying to do everything for everyone all at once.”
Empowering Leaders with Technology
21:05 to 23:20
Discuss the role of leaders in leveraging technology to enhance team communication.
“If I think about five years from now, what does computing look like, right?”
Show all 17 chapters
AI Adoption and Team Engagement
23:20 to 26:00
Strategies for leaders to engage their teams in adopting new technologies.
“So I'm curious, how would you equip a technology leader right now to get fired up and turn back to their engineering team and get them to start operating in these new ways?”
Evaluating Tools in a Rapidly Changing Landscape
26:00 to 28:03
Strategies for leaders to evaluate and adapt to new AI tools effectively.
“But now it's fast forward and it's 2026, right?”
Evaluating Development Tools for Simplicity
28:03 to 29:17
Discover the importance of simplicity in choosing development tools and workflows.
“Quantitatively, it's super hard to know whether like cloud code or a cursor works better.”
Customer Experience and AI Integration
29:17 to 31:29
Learn how to define and optimize customer experiences using AI systems.
“I think in the world of SaaS software and people just vibe coding a replacement to something instead of renewing a vendor, I'd much rather in that world be something like grep or git.”
Understanding Problems Before Solutions
31:29 to 33:36
Understand the necessity of clearly defining problems before seeking software solutions.
“Because a lot of times, historically, what software leaders have had to do is go out and shop and buy and get the closest to what they need and then conform it to what the company's goals are.”
Embracing Change and Reinvention
33:36 to 34:39
Explore the need for continuous reinvention in a rapidly changing industry.
“This has been such an insightful conversation, but I want to ask before we kind of wrap things up, to that leader who does want to operate that way, you know, any final advice that you would give them?”
Call to Action for Embracing New Tools
34:39 to 35:18
Encouragement to adopt voice-to-text technology for enhanced productivity.
“like, maybe some things that I think about is like, nobody knows the answers right now.”
Transcript
Automatic transcript. May contain errors.0:04Sahaj Garg:Today, we're exploring the way engineers are relearning how to build software from scratch with Sahaj Garg, co-founder and CTO of Whisper Flow. And Whisper is a company that's been on everyone's lips lately, and for good reason. They're at the center of a shift that many teams are feeling, moving from voice to text, from never worked to a possible workflow that people now use every day. all day. And this change in the bottleneck going from different parts of the process to actually being your keyboard has led many leaders and senior engineers to discover that they can express their context, taste, and intent faster, and help their teams move quicker with voice.
0:45Sahaj Garg:Because what if the limiting factor in your organization isn't actually like a context window, but how quickly you can get the right context out of someone's head? So today we're going to be talking about shared context as being a valuable currency, and all of the different changes that this introduces to how people can communicate with software. And we're going to get really practical about that as well. So Sahej, welcome to Dev Interrupted.
1:10Andrew:Thank you, Andrew. It's fantastic to be here. Really excited for this conversation.
1:14Sahaj Garg:Me too. And I want to start at the top by just mentioning a little bit about using Whisper. I've been using Whisper Flow very recently, and I've totally fallen in love with the type of software that it is. It's a true delight to use. And I can see why it's dismantling the way that engineers have traditionally approached working with code on their keyboard. I myself have definitely written like a whole novel at this point. And Whisper even tells me as much. And so I'm actually so blown away by how much I can trust this technology to express what I'm trying to say. And that's what I really want to explore today.
1:46Sahaj Garg:Because I know me and our listeners, too, we've all been burned in the past by like bad transcripts. or like, you know, that voice of text didn't really capture what I said, or that voice note had a crazy typo in it. And those little tiny burns, they add up over time and people walk away from the technology. But, you know, we're seeing a shift now where you're rebuilding trust in a technology that many had dismissed. So what has that been like for you at Whisper and how did you approach that challenge?
2:15Andrew:Yeah, it's a fantastic question. And it's actually one of the reasons why we almost never built the product that we did. Because in some ways, I actually thought using voice for communication, for typing, for interaction in a computer was like a fundamentally doomed thing after 20 years of being disappointed by every product there was in the market. And I think the thing that we learned as we built it, and especially the first version was like, oh my God, when this works, it's magical. And that is the thing that we have so consistently heard from people who use the product over time, even now, which is when it works, it's magical.
2:48Andrew:and all the work that we try to do here is to expand the settings and the context and the places in which you just get it right for the first try. Because like, as a user, what I want is a system that just gets me intuitively. Like, I don't have to explain myself a bunch of times. It should know, like, if I'm talking to Cloud Code and I'm talking about a.end file, what that actually means, right? Not show up in a completely nonsensical way. I'd say, like, the biggest framing for us of this problem, right, is we want to build you a voice interaction where you never have to go back and fix the mistake.
3:22Andrew:For us, we call this zero edit rate, right? Something where you have to fix no mistakes with what you're doing. And that means both getting everything you said right and figuring out what you actually meant to convey so that we can help you fix that up on your behalf in a way that sounds just like you and that you can actually use downstream.
3:42Sahaj Garg:So really, it's like a two-part equation because you have like the traditional layer of like, Like, oh, yes, we can turn this voice into the words it is. But then there's also this contextual layer of understanding what are you operating in? What are the words around you? And what have you said recently? What are the things that matter to you as the user? And combining those two is like the formula, I think, that whispers is getting right and is what's letting people work very quickly with it. So I'm curious in how when you solve the problem of like rebuilding trust and doing it in this two part way and acknowledging that it's a nuanced engineering problem.
4:18Sahaj Garg:And so you're approaching this with your teams. And, you know, we're now in a green field because we've acknowledged that there's a lot more nuance in how we solve this. So what assumptions there in that world do you throw away to ultimately arrive at an application or a system that's a delight that gets that zero edit rate? You know, I'm curious, like what becomes the secret levers?
4:39Andrew:Yeah, yeah, yeah, yeah. There's a bunch of different ones. I think the first one is this idea of like, don't treat speech models as dumb things. Like right now, speech models are mostly dumb. They just take an audio. They just produce text. And it's kind of like listening to like a three second voicemail from somebody who you don't know. It's like very hard to actually figure out what a person's saying. and so the really fundamental assumption to break down is like speech models actually need the kinds of things that people use to make sense of each other which is that context that memory that understanding of the other person which helps make sense of so much more and so there's this part of it and then there's like how do we actually build products like this and i think products in general right now is there's two types of things that we do there's one type of work where we do where we just really really care about precision and scale right so this is like build the world's best speech models and spend and pour all of our energy into making it better and better and better and this is like a monumentally difficult effort in terms of how we actually approach it and requires very sustained persistence it's not the kind of thing where you can like vibe code in a day and have the thing work and like it's going to solve all the problems it's very much you know Use the coding agents to build out all the infrastructure, to train these models, to create a self-reinforcing feedback loop to make it better and better.
6:03Andrew:That's one half of our work. And the other half is the user experience, right? How do we help people build this habit? How do we make it really good? And that's where we actually do a lot of experimentation. Build a version of a way you can handhold a user to learn how to do this. And ship that every two hours, right? Because with being able to go from we built a thing to we saw what happened, now we learned something to now I can express that idea of what I think we should do next into a coding agent, it completely changes the feedback loop of how we build a product like this. So much of that is experiment to learn, experiment as quickly as possible, and then get that to be something that billions of people use.
6:44Exactly.
6:46Sahaj Garg:And you're calling on something that's really smart that actually a recent guest, we had an engineering lead from Codex at OpenAI talk about by capturing the effects of research and using that to drive downstream engineering decisions and then feed that back into the research to create this really amazing feedback loop. And we recently talked with Jeffrey Huntley, who talked about with the Ralph loop, it's all about capturing the back pressure of working and finding the outputs that are most useful for that next input. So what you've described here is fascinating because it's like you identified and created that compression event.
7:19Sahaj Garg:You are doing the iterative work, that persistence, you said, which I love, on the model level, which is necessary. But then you're approaching the UX level with the level of experimentation and rapid iteration that's needed to really survive and be effective as a user tool. And those two things are a constant, probably, balance. And it's a new kind of engineering dynamic that I think a lot of product leaders are still kind of wrapping their head around. I think you've you painted a really clear picture of how those things work together.
7:51Andrew:The way I think about it is like, what's the hat to be wearing right now? Is this the kind of thing where we know a certain approach is going to work and we just have to hammer away at it to figure out how to make it work? This is what we call like no brainer bets or scaling here internally. Or is it an experiment? Because if you treat an experiment the same way that you treat the no brainer bets, like you're going to get nowhere. You're going to get nowhere at all. And so the thing that I always ask myself is like, hey, put on the right hat. What's the hat to be wearing right now? And what's the right way of working to accomplish that?
8:26Sahaj Garg:Absolutely. So I want to dive into a bit how this shifts the technology itself. I want to talk a little bit about Whisper as something that now has been created in this compression chamber capturing event that's taking all of the best of research, all of the best of UX experimentation and combining it into what we know as Whisperflow, that's transforming how people work. And if we take that one step further into the downstream effects for how teams are now communicating and sharing their context and engineering, both like in an IDE or with an agent, along with what you've built. And I want to frame that actually in a really fascinating way.
9:06Sahaj Garg:It ties back to an article I just recently read from Steve Yage about how the economics of the software era are changing. And the types of software that are useful and survive now in this new era are just fundamentally solving a problem of cognitive burden and couldn't possibly be replaced. And voices like that, like conveying your thoughts and being able to do so effectively and accurately, most people are not going to have the throughput or the energy to solve that problem. So there you discover your moat in this new AI era. So for folks that are now using these types of tools, like how do you think they would utilize your tool compared to someone who wouldn't to get further ahead and to share their context best?
9:54Sahaj Garg:Like when you model your users and your downstream engineers, for example, what does that look like for you? Yeah.
10:01Andrew:So I'll talk maybe high level about how things are changing and then for different types of people what this means. So in terms of how things are changing, right, the better and better AI gets, the more of the work it's doing. But the one thing it can't do is figure out like what's in my head, what's in your head and the unique things that we observe in the world. Right. And so that's the one core thing where we want to help amplify people's voice. We want to amplify the thing that they can get into communicating with an AI or into communicating with other people. The two types of places where people are going to be communicating right now.
10:33Andrew:there's a lot of different ways in which it amplifies different people so if you're a developer and you're building software the most important thing to actually get these tools to build a software that you want is to express with clarity what you actually want to achieve and to go back and forth on brainstorming with it if you don't do that then you're not able to actually build the right thing and build the right plan and actually go and execute on that plan and so fundamentally right that system is bottlenecked by like your ability as a person to express the things that are on your mind into the tool.
11:04Andrew:The other like really powerful thing about voice is even if you don't have your thought fully formed, you can still express it into the tool because it'll help you think through what you haven't fully understood yet, right? So that's an example for developers where it's like you vibe code the wrong thing, debugging it and fixing it takes weeks, right? But if you get it right on the first try because you gave it all the context upfront, it like saves you all of that downstream pain. The other example is kind of managers and leaders. Like a lot of what managers and leaders do is they have context across an entire company or an entire organization.
11:38Andrew:That's one of the unique vantage points they have. They know a lot of different dots in the team that they're leading. And a lot of their work is connecting those dots. And connecting those dots often means just getting the right information to the right person at the company at the right time. And so if you can do that faster, right? Because you can get a message and instinctively just say the reply instead of having a meeting, right? Avoiding the meeting when you don't need to and avoiding the pileup of like 100 messages in your Slack, it becomes an extremely strong amplifying force, right? Because then one leader can help unblock people so much faster and keep doing more.
12:15Andrew:Those are like two examples. There's plenty more for people who communicate with external clients or customers or legal industries and so on and so forth. But I think those two kind of give a sense of what it looks like.
12:28Sahaj Garg:totally and so for you know engineers doing for example async engineering work and where does voice meaningfully improve that uh like things like around reasoning and exploring your code and architecture and where does it possibly introduce ambiguity because there's obviously lag in communication between teams and people and kinds of async and remote coding environments i'm curious too i've seen the really cool posts of people like in their engineering offices with the gooseneck microphone you know everyone's whispering in the engineering room i love that like i want to be there and i want to be like it's so quiet too is what i hear so like i love that story but then there's also the engineers plenty of us who like me are you know work in our living room and so how does exploring and molding with voice in this way like what are things they should
13:14Andrew:keep in mind yeah versus like in person um i actually think they're the if you work remotely or in a hybrid capacity it's like the easiest to adopt tools like this because when you're not around other people, it's so easy to just speak to your computer all day. And it's really delightful too, right? And so there's like that aspect of doing it. And there's that aspect of using it to kind of catch up on async quicker when it comes to communication and so on. The one pitfall is I've seen times where people on our team are both using flow back and forth to each other, DMing each other back and forth in a Slack thread.
13:46Andrew:And it's like, I'm speaking to you, you're speaking to me, but like, we're both using a tool to turn it into text back and forth in for real time. And that's not good. Like at that point, you should hop on a call, right? And I think it's really easy, especially in hybrid and remote environments to accidentally fall in that trap sometimes. So that's the one thing to kind of caution against when you start using tools like this is, hey, what's the way in which it could impede really important live and productive conversation. But outside of that, it's like, I use the tool definitely the most, It's like between 10 p.m.
14:19Andrew:and midnight when I'm just able to work at home and really get in my full state. Yeah.
14:27Sahaj Garg:And I think it's funny, too, that you talk about these communication anti-patterns that kind of emerge. I think, I don't know about you, but I've definitely at some point in my life been at the receiving end of like a stream of consciousness voice note from somebody. And you're trying to follow and trying to understand. And now it's like we're in a world where that could be a Slack message that comes to you across from somebody. maybe even your boss and they're trying to get you to do something so it's more important than ever that like these tools don't act as obstacles they don't make work for the recipient you talk about a zero edit rate but also like a zero work rate for the recipient if they are on a receiving end of one of those to understand because like we talked a little bit about like the obviously talking to your computer total delight but being voice to text it can be used in a whole bunch of ways so you have to think about a really wide user span right as like a as a pretty big challenge especially as like a cco like how do you how do you you know bucket and understand that it goes beyond and just like you know when you go to onboarding flow and you're like oh i'm a coder i do engineering work like it there's actually more to that it's like it's like how do you tackle
15:32Andrew:that problem it's actually tremendously tremendously nuanced so there's two parts of it right so i think about it as okay right now we're a tool to amplify our communication eventually it'll be a tool, but also takes actions for you. But let's just talk about communication right now, right? The way we think about the product spec for our language models that are going from what you said to what you want to communicate is there's two tasks, right? One is to make it something that represents what you said in a way that's true to you and true to the communication you want to convey. And the second is to make sure it's intelligible as the person receiving it, right?
16:08Andrew:Because I don't know about you, but I've never spoken for two minutes straight and been able to produce a perfectly coherent email when I speak. But the recipient of this kind of a conversation, you'll understand me when I speak for two minutes, right? People listening here will understand a two-minute dialogue. And so there's definitely a way to satisfy both of these constraints. And where it gets even more kind of challenging is like the way that this should happen depends so much on who's communicating with who. Like when I text my co-founder, some of the employees on my team, my wife and my parents, it's all going to look pretty different.
16:46Andrew:And the shared understanding, the shared context, they're super, super, super different. And so what we really try to do is basically build ways to automatically infer and learn your intent. Understand what you're trying to do right now. Give you as a user control. right so if you wanted to be more like verbatim to what you said or if you want to be more interpreted on your behalf give you that control and then learn your preference kind of automatically over time right because the the way that i see it is like hey you fix a mistake that we produce we should learn that and not do it again right why should you have to tell us multiple times we should be able to learn all those things about how you might want it in different settings like different from how somebody else does.
17:34Andrew:And so maybe the biggest challenge from an actual engineering and tactical perspective is not trying to do everything for everyone all at once. Because if we do that, it's very hard to actually tackle all these problems. But really methodically working through it in a way that's both specific and then over time generalizable so that we can make it work actually for everyone. Right.
17:59Sahaj Garg:Right. And in thinking about the two, you know, and thinking about your wide breadth of users and the different contexts that they find themselves within, then even within that, there's a level of granularity of the type of handoff that it is because there's, you know, in cases where there's a human talking to their agent. And then, of course, like hearing something back from the agent and then working with it versus working with a co-worker. Right. But in the world we live in right now, all of that text swims together. And so it's a fascinating kind of problem to solve for. But I'm curious to know, too, like when you are able to pull the context out of people's head, what form do you think it should best live in?
18:39Sahaj Garg:Because obviously when we capture those thoughts and then they become markdown documents and these things accumulate, right? How do you, as somebody who's, oh, I'm capturing my thoughts and working with them, not create garbage or instead of creating useful artifacts?
18:54Andrew:um it's a great question and i think there's like garbage and two ways that you can create right one is the stream of consciousness that goes to somebody which is that's really dangerous as you said uh and then there's a stream of consciousness that you output into your notes app and i think everybody everywhere has never had a great experience with finding a way to like dump all of the crap that's going on up here into something that uh is useful over time right and so I think one of the unique things about voice is the thing it's best for is frictionless capture of the jumble of ideas in your brain.
19:35Andrew:It's so much better than anything else for that kind of problem. And then the question for us will become, hey, how do we help you organize that information for yourself? How do we help you proactively make sense of that kind of information for yourself? And those are two problems that we're very actively exploring right now and prototyping different solutions because, you know, this idea of being able to offload my thoughts into a second brain is something that people have, I'd say, tried quite a bit and never really found a way to make it stick. And like voice is that interface to which we can do that.
20:15Andrew:And yet today, like our best tools kind of just tack on some AI after the fact to like try and help you maybe do something which is not really what people ultimately need or want to solve that problem.
20:29Sahaj Garg:Which is why going back to perhaps what you alluded to, the idea that you solve this problem now for voice and in understanding that context and then the, you know, the idea of what someone's invoking, you can then go one step further and solve the action problem too. And because you understand the context in which they're operating in, in a much more intimate way, because of that's the, that's the benefit of voice. So is that where you see this type of technology evolving, where, you know, I could today, of course, talk to any sort of type of tool to operate things for me. But do you think that maybe even that collapses one level further?
21:07Andrew:I think it does, right? If I think about five years from now, what does computing look like, right? I'm probably just going to be either expressing my intention to my computer in response to a decision that it asked me to make or like spontaneously because there's something that I want to put into it. And like the things that I want to happen and will actually happen, right? Like that's the magic that technology is supposed to promise us. And unfortunately we're now in a like world where we're like twiddling away tiny thumbs on a tiny screen, far, far from that part of that promise, right? But that's what I expect it to be.
Read the full transcript
21:39Andrew:And I think the biggest things that are missing right now on that path are a thing that actually gets what your intent was and what you were asking on the first try. Because like, if you have to go and fix that up all the time, is not going to be a tool that you ever trust as a primary interaction, right? And the other thing is like the right user experience and workflows around this. Like lots of people have built voice interfaces that drive actions. And a lot of people don't use them besides for setting timers. And it's not due to a lack of trying to use it for more. And I think it's because we haven't come up with the right interaction patterns and interaction paradigms with the right quality of underlying technology for anybody to be able to trust it and do stuff with it.
22:27Andrew:Like, I'm never going to remember the 200 different commands I could execute with my voice. Like, that's the reason why we have UIs. And so, you know, a lot of the work that we're doing, there's both just improving dictation, making it better and better and better, but also thinking about, hey, what's the right user experience for me to express my intent into my computer and for it to just get what I want and help me do it.
22:54Sahaj Garg:Exactly. And I love how you brought us here to kind of like this lack of, you know, we currently don't have these workflows, these realities, these supporting structures that help teams operate in this way and in the way that they need to. And that's ultimately a burden that falls on, you know, the leader. So I really want to talk about the leader's role in all of this and how they can take things like getting unblocked by communication is just one small step, but then also understanding the compounding factors of technology like voice-to-text that allow them to amplify the work that they can do and how it is their responsibility, right, as a technology leader within their own company to create these pathways, these highways between their teams and how they work for this context to not pile up and for it to be useful and for people to feel supportive with amplifying their work in this way.
23:45Sahaj Garg:So I'm curious, how would you equip a technology leader right now to get fired up and turn back to their engineering team and get them to start operating in these new ways? What are some of the first things that you would tell them?
23:57Andrew:Yeah, I think about this is like getting AI enabled across the board, like not just with Whisper, right, but with lots of different tools that are part of the stack. I think the beauty right now is it comes down as simply to people just trying and using it. it's so easy to talk about it and to theorize about it. And it's like kind of useless to do those things. So it's like, well, these ideas are promises, right? And what actually matters is what happens when people deliver on that promise and the degree to which they do, right? And so if I were kind of an end leader, and I am within Whisper, like, the biggest thing I'd be doing is making sure that not only I am trying all of the new tools, but that everybody on our teams are proactively trying all those new tools and just sharing what they learn on a daily basis.
24:44Andrew:Because if it's just you as the leader who's pushing the charge, it's actually going to move way slower than if everybody on the team is pushing the charge. And you can create a system for actually amplifying the people who are really, really keen to figure out these new ways of working. that's the that's the biggest thing that i've noticed as a way to kind of scale ai adoption because like if it comes top down not much is gonna happen right but if it comes like from everybody being like oh my god like i feel lighter work feels easier i can have more impact this is fun i get to do all the things that i just wanted to do but like felt like i was stuck because i just couldn't get all my thoughts out or couldn't turn my creative idea into code fast enough like that's the kind of thing that really unlocks it um and then you know your job as a leader is to basically synthesize right synthesize all those different ways of working develop some yourself right based on the kind of work that you are doing and like spread that knowledge and teach others we're very early in this transformation anybody who's doing that's going to have a huge it should impact.
25:58Sahaj Garg:I love the picture you painted. And I think it aligns with what a lot of experts like yourself in the last year have come on the show and kind of painted about how that experimentation should look, what you should measure, obviously elevating your champions and making everybody drive the effort. But now it's fast forward and it's 2026, right? And a lot of teams have been doing that now for a year. And you've accumulated maybe a huge buffet of AI tools and everyone kind of has picking their own poison and everyone has their favorite things. And so now as a leader, how do you look for the high signal tools, the ones that are most useful?
26:32Sahaj Garg:And like, what are the maybe kind of even metrics that you would advise a leader in the situation to use the narrow down the effective ones?
26:40Andrew:Yeah, it's a great question. I think actually the funny thing about this is experiments that were run three months ago, in some ways, probably should just be completely discarded today because the conditions of all of this has changed and so like the best way to do a thing today might just look nothing like what it was like three months ago and there are just a lot of intermediate stepping stones along the way which are necessary right but um i think the first first and kind of funny thing is like don't attach yourself to anything that was a specific thing that you learned at some point right if it's like oh okay great like the compaction window once you cross a certain amount it's gonna be bad so make sure like everything that you do is about avoiding that that's important today that'll be important for a month maybe two and then like nobody will care like there's no way that that's going to be the problem that persists for like six months in AI tooling.
27:44Andrew:And so if something feels like a blip along the way or a way of doing a workaround to a thing that's going to improve, like my general perspective is look for the simplest possible solution that solves the problem at hand. The simplest, most elegant possible solution that solves a problem at hand. Quantitatively, it's super hard to know whether like cloud code or a cursor works better. But like, at least for me, one of them is simpler for one task, which is agentic development. And the other is simpler for like opening a file and inspecting it and interacting with it. You can probably guess which is which right now.
28:19Andrew:But like that, that's how I think about like, what's the simplest thing that gets that job done? What is the gap there? What are things that were assumptions that were made three months ago that I should basically toss out the window? Because three months ago, I wasn't writing 95 % of my code with AI today. Like I write a single line of code once in a while by hand. And it's just crazy how fast that shifts.
28:45Sahaj Garg:I love that. How really you're calling it more like a, you know, you should call some of these experiments or definitely reevaluate what they were doing in the first place. Like acknowledging that we're, you know, we're stepping on lily pads on islands like they're temporary and we're getting somewhere. And the way I keep thinking about it is I'm building things, but then I'm picking them up and I'm running with the stuff that I'm building as I'm building it. Like I'm not digging a mode. I'm not throwing down anchors. I'm not like putting anything I'm running. Right. So I think that's a really important call to action for folks.
29:16Sahaj Garg:If you find yourself with a bunch of like workflows, like maybe under maybe reevaluate if they were crutches, if there's a better way to solve it today and go for simplicity. I think in the world of SaaS software and people just vibe coding a replacement to something instead of renewing a vendor, I'd much rather in that world be something like grep or git. Like something that is so simple and solves such an effective problem that it can't be replaced effectively or efficiently or it would cost you too much in your time, efforts, tokens, compute, whatever. Right? So that's what I think leaders should optimize for in their workflows too.
29:55Sahaj Garg:And I like that you think about it that way.
29:56Andrew:I'll give you one example for us, which is actually on the customer support and customer success side. So I think a lot of the prevailing wisdom is like buy, not build for some of these kinds of things. And we've tried a lot of different solutions. And actually, what we've kind of come back to is the models have gotten so good now, that the most important thing is us as a company defining the kind of customer experience we want our users to have. And so much of that is actually about figuring out how to give these systems all of the context so that when anybody runs into any kind of problem, we know exactly what to suggest them to do first to try and also can automatically have that bug hooked up into not just a linear ticket, but like actually just kick off a PR.
30:43Andrew:Like that's actually a possible workflow now. And the hard parts of that are actually from everything we've tried less the tooling and less everything around it, but more so defining what you want your like 11 star customer experience to look like and how you get there and how you deliver like founder level customer support at scale. and that's an example of like you know i tried building this four months ago as well when i was frustrated that the solutions that i bought weren't kind of good enough and like didn't really work and like today it most definitely has started working most definitely and we're like investing very heavily into improving that internally because it relies on things that are unique to our core competencies as a company right understanding these kinds of challenges getting
31:34Sahaj Garg:aligned on what the decision should be you're right to call out that all of the work now turns inward just in the same way that like the proliferation of ai generated code exposed all of the problems in the stlc as what they were which is like communication and context lagging and baggage and you get so you get like this um bottleneck phenomenon right and in that same way like when like leaders don't like operate with the technology uh or rather like I guess what I'm trying to say is that if people don't acknowledge that they need to understand the problem they're solving as a business leader before approaching the build-buy scenario, then they're going to be much better equipped if they can make that realization.
32:20Sahaj Garg:Because a lot of times, historically, what software leaders have had to do is go out and shop and buy and get the closest to what they need and then conform it to what the company's goals are. And it's almost like a backwards process. so now people have to flip that around and really understand from a top to bottom level what am i solving with this piece of software that's what the software needs to do is solve something and we can be really precise i'm experiencing the same thing with models too where like if you know exactly what you want you can get there in a surprisingly short amount of time and energy so definitely changes the like the unit economics too of like how people buy i the one thing that i
32:59Andrew:often say to people is like Silicon Valley talks a lot about high agency. And I think the first step of high agency is actually just knowing what you want and knowing what problem you want to solve. And like people often skip that step, right? They like start with step two, which is actually solving a thing without really, really clearly identifying, hey, what is it that I actually want? And like, how do we get there? And this is more and more important right now, right? Because if a leader or an engineer or anybody can correctly and precisely articulate what they want a system to do, it's like so much easier to solve that, right?
33:34Andrew:Just like you said, than ever before.
33:37Sahaj Garg:This has been such an insightful conversation, but I want to ask before we kind of wrap things up, to that leader who does want to operate that way, you know, any final advice that you would give them? Because we're in a rapidly transforming industry and everyone is throwing away expectations and working with technology in a new way. Is there any kind of advice that you think is kind of pointing you forward for 2026 as we kind of start to figure out these new challenges?
34:03Andrew:For me in 2026, it's about reinventing yourself every three months, like properly and truly reinventing yourself and your organization every three months. And that's deeply uncomfortable and deeply unsettling because it is very hard for people to change at that speed. Super hard, super, like, literally uncomfortable to actually go through change at that pace. But things that don't work today will work tomorrow in a way and speed that is like hard to kind of fathom because of right around the inflection point where we are. And so like, maybe some things that I think about is like, nobody knows the answers right now.
34:46Andrew:We're all kind of learning very, very quickly to figure out what the playbooks ought to look like. And so not being scared of that and holding on to what's uniquely you, which is the ideas in your head and the way that you express them into the world. That's why also we're building what we're building out here at Whisper.
35:08Sahaj Garg:that's amazing i i can't agree more that it's about embracing and and building your taste and like you've called out very rightly in this conversation you know tearing down the barriers between you and expressing your taste and if you're listening to this and you're still using that dusty old thing called a keyboard uh this is definitely your call to action to try out some voice-to-text technology this is something that i have been using for well over a year uh especially in combination with agentic coding. I've tried a lot of tools and I think Whisper is something that's a truly delight to use.
35:41Sahaj Garg:And this is not at all a sponsored podcast. I just definitely wanted to throw that in there for people just because I can't get enough of this tool. So if you're a software leader, I really encourage you to pick up simple but delightful tools like this and figure out how they're going to take your team into the future. And Sahej, just before we wrap, where can our audience go to learn more about you and the cult of Whisper?
36:02Andrew:yeah you can head over to whisperflow.ai and the best thing that you can honestly do is just download the product and use it like people have been told for 20 years that this kind of stuff works and the only time you're ever going to believe that it actually does and that it actually is delightful to use is by actually trying it out so you can head over to our website give the product a try read some of our research and engineering blogs to learn more amazing so we'll
36:29Sahaj Garg:include those notes in the show notes. So please also be sure if you've listened, especially this far, to come check us out on LinkedIn and Substack, where the full newsletter along with this podcast is distributed, as well as reach out to us. We would love to hear your thoughts on our conversation today. Pick up and continue anything that we've talked about here. And thanks for joining us again. That's it for this week's Dev Interrupted. And Sahej, thanks again for chatting with me today. It's been a blast.
36:53Andrew:Thank you so much for having me on the show.
37:01Sahaj Garg:AI helps your developers write more code faster. But here's the problem. Your review process hasn't sped up. The queue grows, reviewers get burnt out, cycle time stalls. Linear B changes that. Our AI reviews every PR the moment it's created, catching bugs, security gaps, and performance issues before humans get involved. It even writes the PR description automatically. Your reviewers spend less time on first-pass problems and more time on architecture and business logic. Break the bottleneck, see how Linear B accelerates your workflow.
From the publisher
Your keyboard is the biggest bottleneck in your engineering workflow. This week, Andrew sits down with Wispr co-founder and CTO Sahaj Garg to discuss why traditional voice dictation failed us, and how his team is rebuilding trust by using contextual models to capture a developer's raw intent rather than treating speech models as "dumb" tools that just produce literal transcripts. Together, they explore the engineering hurdles of translating a messy stream of consciousness into perfectly formatted, zero-edit artifacts that can be instantly understood by both AI coding agents and human coworkers. Finally, Sahaj shares his framework for experimenting with new tools and why surviving this era of software development requires completely reinventing yourself and your organization every three months.
Follow the show:
- Subscribe to our Substack
- Follow us on LinkedIn
- Subscribe to our YouTube Channel
- Leave us a Review
Follow the hosts:
Follow today's guest:
- Try Wispr Flow
- Now on Android
- Connect with Sahaj on LinkedIn
OFFERS
- Start Free Trial: Get started with LinearB's AI productivity platform for free.
- Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.
LEARN ABOUT LINEARB
- AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
- AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
- AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
- MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.
