#285 Real-Time Speech AI and Accent Translation with Sanas CEO Sharath Narayana

19 May 2026 · 45 min · 16 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Sanas (CEO Sharath Narayana) explains on-device real-time speech AI for accent translation, speech enhancement, and language translation—aimed at making human-to-human (and human-to-agent) communication clearer without changing the speaker’s accent.

Guest background

Sharath is engineer by trade; previously worked in solution consulting at Akamai (Boston). He founded Observe.ai (speech/conversation intelligence), which raised capital and exited in 2021. He co-founded Sanas in 2019 after backing a Stanford white paper on reconstructing speech from phonetic sounds in sub-200ms.

Key claims

People initially paid for “accent” more than “enhancement.” Sanas uses phonetic-sound matching on-device (sub-100ms algorithmic latency) to harmonize accents (output: standard American or standard British). Language translation needs more phoneme pairs, so it has ~1–2s latency plus network delay; accuracy is prioritized (96–98% for launched languages). Human-in-the-loop supports healthcare/banking trust.

Notable examples

Windham Resorts (WFH sales calls with background noise) and Alorica contact centers (reducing mistrust in first 30 seconds; improving respect/attrition). Language translation used by healthcare/banking customers; human agents can switch to Mandarin/Arabic with 7–12 minutes setup. Roadmap: any-to-any accent in ~90 days; 5 languages mainstream in ~90 days, 30 by year-end; SDK release for accent/enhancement/language algorithms in 90–180 days.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Sharath Narayana's Journey to Saunus

0:45 to 3:29

Sharath shares his professional background and journey to co-founding Saunus.

“So look, very, very interesting company.”

The Evolution and Concept of Accent Translation

3:29 to 12:12

Discussion on the concept of accent translation and its significance in speech AI.

“And one thing that's fascinated me was something around speech and speech data.”

Customer Conversations and Market Approach

12:12 to 14:02

Exploring initial customer interactions and how accent translation is marketed.

“We genuinely thought people will want to pay for enhancement and clarity, but people wanted to pay for accent.”

Understanding Accent Translation Goals

14:02 to 14:36

Learn about the two-way accent translation technology in development.

“But eventually the goal, which should become a reality in the next 90 days, should be any accent to any accent, right?”

Customer Conversations and Use Cases

14:36 to 20:13

Discover how accent translation impacts customer interactions in various industries.

“I would just kind of double click on how does the customer side of this.”

Clarity in Communication and Sales

20:13 to 24:47

Explore the importance of clarity in communication for improving sales and customer satisfaction.

“first call resolutions and call handle time.”

Challenges in Language Translation Technology

24:47 to 27:35

Understand the challenges faced in developing accurate language translation solutions.

“Especially in healthcare, you have to be doubly, triply sure.”

Technical Aspects of Latency in Translation

27:35 to 28:00

Learn about the technical and philosophical aspects of latency in language translation.

“element in that call, things go a lot better.”

Technical Challenges of Speech AI

28:00 to 29:19

Explore the technical hurdles in speech AI development, particularly latency issues.

“So it takes a bit of time to process it.”

On-Device Processing vs. Cloud

29:20 to 32:57

Learn about the advantages of on-device processing in speech AI technology.

“We are like, we'll compromise the latency a little.”
Show all 16 chapters

Matching Algorithm for Speech Translation

32:58 to 36:29

Understand how a phonetic matching algorithm enhances language translation.

“of the 14 algorithms we have built, 13 of them work on device.”

Funding Journey and Market Insights

36:30 to 41:41

Gain insights into the funding journey of Sanas and the evolving AI market.

“which is a technical mode for us, and that's what is helping us do it better.”

Future Roadmap for Speech Technology

41:42 to 42:06

Discover what's next for Sanas in the realm of enterprise voice technology.

“We've taken a lot of contrarian approaches to building.”

Building a Profitable Business

42:06 to 42:30

Sharath discusses his approach to building a sustainable business.

“Otherwise, I'm very happy building this from here.”

Exciting Upcoming Roadmap for Sanas

42:30 to 43:31

Sanas is set to launch significant advancements in language translation.

“You know, you kind of teased there that things are going to be launched in the next quarter or something.”

Democratizing Speech Technology

43:31 to 44:38

Discussing the plans for SDK launches and the future of speech tech.

“But I think Sarnas as a speech company can go on a network, can go within a device.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We genuinely thought people will want to pay for enhancement and clarity, right? But people wanted to pay for accent.

0:12Sharath Narayana:Hey everyone, and welcome back. It's been two weeks. Welcome back to SlaterPod. So today on the podcast, we welcome Sharath Narayana, and he is the CEO and co-founder of real-time speech AI platform, Sanas. Really happy to have Sharath on the podcast. Thank you so much for joining. Florian and team, thank you so much for inviting me. I'm excited to chat. So we always ask, where in the world are you recording this from? What part, what city, what country? I'm at our corporate headquarters in Palo Alto in California, right next to Stanford, that's where it all started. That's where a lot of things started.

0:50Sharath Narayana:So, well, good morning then to you. Thank you. Cool. All right. So look, very, very interesting company. I've done some of the research, listened to a couple of podcasts you recorded before, where you kind of lay out the origin story of Saunus. And this isn't your first startup. You had a different company before called Observe.ai, which still operates. So tell us a bit more about your kind of professional background, the journey, and then co-founding Saunus. Engineer by trade. Graduated back in India. Now, gosh, it's been over 20 years. Did what most people do, did computer science and engineering and wanted to be a coder.

1:34And then in the first two jobs, I realized I can't sit in front of a machine and keep coding all my time. Wanted to be in front of people and talk to people. And my first full-time, authentic, real role was in Akamai, a Boston-based company. They started off MIT labs and they gave me a break into doing solution consulting. They put me in front of customers and I realized it was decent and talking to customers came across as authentic, genuinely loved technology and that kind of rubbed off the right way. Back in 2012 when I thought everything was going perfect for a 28 year old, I thought something was missing and when you feel something is missing in your life you do something crazy.

2:21I packed my bags, went back to India. I called it a golden handcuff at that moment in time. I think everything is going too much right. So I went and started up. So that's what you do when you want to do something crazy. First startup, learned a lot of hard lessons. I hit the ground as hard as I could, and I fell flat on my face, as bad as you can think. But lots of life lessons learned. And then it took me a few years to start up again. But this time when I wanted to start up again, I decided not to chase what was the evolving trend. At that point in time, I truly wanted to do something that spoke to me.

3:06And then it was early days of machine learning. The possibilities of computation was endless. And as an engineer, it spoke to me that, hey, if you can actually process lots of data and make some sense out of it and provide that intelligence to either people or businesses, it could be a wonderful thing. And I started looking at domains at which I could process large amounts of data. And one thing that's fascinated me was something around speech and speech data. And probably as a background, I'm probably one of the 600 people on this planet who speak Sanskrit as their mother tongue, Like Latin, it's a dead language.

3:51So I think language and speech has a particular affinity to me. And everybody told me that, hey, speech is very hard, but it's expensive. Speech-to-text has a lot of latency. And the more people told me it was hard, the more it attracted me towards solving it. and the first try for solving for speech or breaking down speech in real time was at Observe AI. And at that point in time in 2017, it was all about breaking speech to text in the fastest possible fashion. And we built a conversation intelligence system primarily based on speech data. And that's the journey at Observe. And Observe, I would say, always right place, right time, went to IC, raised a lot of capital, and then eventually exited in 2021 when software did our CDC.

4:46And though I made a bunch of money and exited, it was always feeling like I've still not truly solved real-time speech because speech to text is still 500 millisecond latency. You can call it as near real-time, but it is still not real-time. And I was looking for what can I do next and I meet this 19 year old whiskey out of Stanford who said I've written a white paper where I can break down speech by breaking it down to phonetic sounds and think about your name Florian the first phonetic sound is fur right that phonetic sound has no language right it's just a spoken sound right and he said when you break it down speech to phonetic sounds and if you collect enough data, right, and you build a lot of phonetic pairs, it's possible to reconstruct speech in sub 200 milliseconds.

5:39And that was fascinating for me to hear, right? And yes, it was a 19-year-old kid saying it. You can think about this as some dystopian reality, or you can believe in it and back it, right? I initially backed Sonus as an investor, right? One of the first few people who wrote the check, there was this person called Steve Schlunker from DN Capital who wrote a check. There was Varsha from Village Global who wrote a check. And I wrote a check too. And it was just about I call it as you pay money to hang out with smart people. That's what I did. So I paid the tax and I started hanging out with Max and Sean.

6:19And I think in a few months, we all were feeling excited that this could be something big. And then they made the offer that, hey, we are thinking of dropping out of college. We will do it and convince our parents to do it if you come in and do this full-time with us. And I took some time, thought about it, but this whole fascination of solving for speech in real time, even though it was theoretical, it was a white paper, it was so fascinating to me. I'm like, you know what? I don't want to regret later. Let me go in and try and attempt to crack it. And that's how we started Saunus. So it's been four and a half years since then.

7:02I've loved every single day post that.

7:05Sharath Narayana:So that was the original kind of use case you tried to solve. Like what would be the current elevator pitch? Like if somebody asks you, what is on us today? Because it's, you know, it's been quite a journey those past four or five years. I think that elevator pitch changes every 90 days. But I think the latest one is, to be honest. But I think what has stayed true is we are an on-device speech AI company, right? I think on-device and speech haven't changed in the last four and a half years, right? What we do with being an on-device speech company could be different. We were an application company before.

7:43Today, we see ourselves as an on-device speech AI infrastructure company because we believe that with the world of AI and everybody wanting voice to be their favorite interface. We believe both for human-to-human communications and human-to-agent communication, there is the fundamental acoustic layer that is needed to process raw speech, and we think Sonus could be that acoustic layer, right? So that probably part of the elevator pitch has evolved, but yeah, Sonus is an on-device speech AI infrastructure company. That's the elevator pitch.

8:20Sharath Narayana:And one of the key differentiators of one of the early use cases was also accent translation, right? Which is, if I'm very honest, is something I really did not have on my radar maybe like 12 months ago. This is something that kind of only popped up on my radar recently. So can we just talk about what's the definition? How would you define an accent? I mean, I have an accent. I'm a native Swiss German speaker. or I'm giving my best to sound, I don't know, somewhere between the Atlantic, maybe somewhere over Greenland. I'm doing my best here. But what is an accent? And then maybe on there, like what's the difference between an accent and maybe a dialect?

8:59Sharath Narayana:Is there, in your view? Generally, they both mean the same for me, right? And the whole goal was, like you just said, we don't want anybody to change their accent or dialect, right? Just like how phonetics is the fundamental of speech, dialect is a fundamental ingredient of what accent. Accent is how it is perceived, dialect is what it is. It's the source and accent is how people hear it. That's the difference between them. But our goal was there are 8 billion people on this planet and there are 8 billion different accents. Everybody has their own unique accents. And the goal was with speech AI, you can make that accent a little more homogeneous, something that people can understand each other very, very clearly without having to change the way they speak.

9:56And that's why we built, in our own scientific paper, we called it a dialect harmonizer. But then nobody will go on Google and search for a dialect harmonizer. So we called it accent translation, which is a more easily understood word. I realized it later. It was not the right phrase to use. But the intent was very, very, very simple that, hey, we are building a speech AI platform. We think we can break down speech in real time in a very unique way by breaking it down to phonetic sounds. And we think we can understand raw speech really well. And we can influence understanding of speech by doing this.

10:37And then we asked ourselves a fundamental question, what comes in between understanding each other?

10:44Sharath Narayana:Right? There was background noise. There was a lot of foreground voice. There was a lot of crosstalk. There was our dialects being very different. And there was infrastructure that we used to speak, like today, I'm wearing a headphone, you're wearing a headphone, right? That could induce its own infrastructure noises. and then there could be language, right? These are potentially six, seven things that could come in between me and you understanding each other. It could be either when we are in person and it gets accentuated more when you're remote, right? And we started Sonus as more like a speech lab and our goal was, hey, let's build a bunch of algorithms that could solve for human-to-human understanding first, right?

11:32and we built about 14 different algorithms in the process of two years and we bucketed them into four families right there was a noise family there was an enhancement family there was an accent family and there was a language family right and then when we went to the market for the first time in 2023 and our genuine question to the first three enterprises we went to is hey, we are a speech lab, we've built these four families of algorithms, what will you pay for? Two of the three of them said, we pay for accent. And we're like, okay, that should be our first product. And that's how it came. It was purely accidental.

12:15We genuinely thought people will want to pay for enhancement and clarity, but people wanted to pay for accent. And then a lot of them were fascinated by language, but it had an inherent latency then. because you just can't have a tri-phoneme pair to do language translation. You need to have a bunch of those pairs to do language translation in real time, because syntax of multiple languages are different, right? So we said, okay, let's commercialize this product, because we don't want to remain a lab. We want to make some money. And that's how we went to market, with Accent as our first offer.

12:51Sharath Narayana:And then what's the Accent that gets... Um, is there like one kind of, I don't know, mid-Atlantic, slightly Americanized accent that it goes into from, you know, maybe a call center overseas or what, or is the accent that it gets kind of translated into also variable? Like you could have any accent to any accent or is the target accent like some type of, I don't know, for lack of a better term, mid-Atlantic American version? So as of today, uh, we have something called a source accent and target accent. The source access is universal. It could be anybody from any part of the world speaking into the sinus microphone, right?

13:32The output for now is either standard American speech, which is the mid-Atlantic standard American speech as defined by the dictionary would sound like, and a similar standard British speech, right? because we initially built it for enterprises and we figured at least in the CX space, the largest outsourcers were North American companies or UK companies, right? So that's where we started this way. But eventually the goal, which should become a reality in the next 90 days, should be any accent to any accent, right? That's the ultimate goal and also two-way. We just didn't want to make it one way because just like somebody might have a problem in understanding my Indian accent, I also have a problem understanding Texan speech.

14:25I don't understand half the things they say any which way. So it has to be two-way. That's where it's truly homogeneous. And that's what we're building towards.

14:36Sharath Narayana:I would just kind of double click on how does the customer side of this. And I read in one of your LinkedIn posts that you have, I don't know, 106 flights a year, 500 customer conversations. So tell me a bit more about how would one of these customer conversations go? Like with a lead, let's maybe take one bucket, maybe for the accent. What's an initial lead conversation and maybe a customer conversation? How would it go? What would you talk about? I think the first two customers I'll tell you that we spoke to, one was a hospitality chain called Windham Resort. right the other one was a large contact center called Alorica right and luckily I spoke to the C-level folks in those companies right with Alorica I actually spoke to Andy who's the founder right and he could understand me as a fellow founder and when when he heard all the algorithms that we've built he said there were two things that jumped out to him one was hey I have a large global team right I have teams in over 20 countries right there are internal use cases where your accent product could be very helpful right because many a time I've seen we have engineers all over the planet also even though we are a call center we also build a lot of technology stack and people sometimes are not very open and comfortable talking in team meetings because sometimes they feel hey you don't understand my Italian accent really well.

16:11You don't understand my Korean accent really well,

16:13Sharath Narayana:right? So I want people to open up and be comfortable under their own skin and not have to worry about it. And the second thing he said is I also run a large CX operations, I have 100 ,000 people, a lot of them in the Philippines and India and Latin America. And one of the things that I struggle with as a large contact center owner is attrition, right? And when we do exit interviews, the number one thing that people say, yes, there is pay, there is sometimes crazy schedules, night shifts and all of that. But the thing that stuck out to me is I don't feel respected in my job. Right? And see, I have to serve American customers.

16:57So somebody sitting in India, somebody sitting in Philippines will have to work nights. I can't change that. right we are a margin sensitive industry i can't pay people whatever they want so there are two

17:09Sharath Narayana:things i can't change but this i think i can change right uh and and he started going down to why they don't feel respected in the job they do because it's not because people they're talking to are bad people right just think about a typical contacts in the queue right people would have waited 15 to 30 minutes. Enterprises will do everything possible for it to be solved digitally and not you having to talk to another person to solve your problem. So people are already frustrated by the time they come and talk to another human being, right? And the first 30 seconds is the impact 30 seconds. And there, if they hear somebody who has a lot of background noise, who has a scratchy microphone and has a different dialect, which is not what they're used to, there is instant mistrust, right?

17:59And then it can only go downhill. That's the only direction the call will go, right? And he's like, hey, if you can make that first 30 seconds impactful shot, right, where it should sound like I'm sitting in a studio, right? If the other person feels that I know what I'm doing, and I'm the expert here to solve your problem, then that person will drop their guards and be the conversation will be free-flowing and it'll be a real conversation if it's a real conversation both people feel respected right and that's the impact that they wanted to hear and then on the Wyndham side they were talking about their sales teams and it was just after the pandemic most of them were still working from home and they're like hey we have people working from home and I can't control the environment that my sales teams are working for.

18:46Here we are trying to sell an$8 ,000 cruise or a holiday package and there is babies crying in the background, right? And my customer is like, is this a scam or like, what is this, right? So how do you fix that? So those were like the two use cases that came in. So yes, Accent was the product, but the use case for me was clarity, Hey, can you make my people sound like they're sitting in a studio? And it's crystal clear that they sound confident. The fidelity is high. There is zero background noise. And the accent is homogeneous. If you do that, I think commerce can happen on that call. And with almost every conversation after that, it's always started with that.

19:33hey, what does this clarity mean to your business? You could be selling, you could be collecting, you could be supporting, either business support or technical support. What does this clarity mean to you? And in different use cases, in one case it meant sales going up, conversion rates going up. In collections it could mean people collecting more money on the same call. Or in customer support it could mean about AHT, handle times going down, resolution scores going up, and ultimately the CSAT going higher, right? And that's how we lead it, right? What we've seen in our sales pitch so far is in BPO environments, it's always about first call resolutions and call handle time.

20:18Those are the two metrics where they want to start a conversation, so naturally it becomes accent first, and then it goes to enhancement second, and then it ultimately leads to language. as the final frontier. In enterprises, it's usually always enhancement first because they just want their people to sound clear, right? And then sometimes it's accent and sometimes it's language, right? So that's usually how the sales cycles go for us.

20:45Sharath Narayana:Fascinating. You mentioned language being the final frontier. So this is not mass-deployed, basically AI, live speech translation, right? I would assume. Yes. In these scenarios that you described. I mean, it's still, you know, hundreds of thousands of agents. They speak one version of an accent in English. You're helping to create that connection, which you just described. But it's not like hundreds of thousands of people speaking different languages or a different language than the customer speaks. That's not yet correct, in a sense. Not hundreds of thousands, but it's getting there. I think by the end of the year, you'll probably have the same number.

21:30Today, I think Spanish, French, German are the three most used languages in our language translation solution. We have still not announced a public release yet. It's still limited availability. We've gone to some of our existing customers, some of our large healthcare customers and banking customers, and they've started using the product. It's been magical so far, but there are still certain latency nuances that we're trying to fix. It will be mainstream by the end of this quarter, but we have 37 implementations right now on language translation, but I would not still say it's mainstream yet.

22:13Sharath Narayana:Can we maybe isolate two different buckets of challenges? One is like the technical kind of the machine or AI translation bucket, the latency and the technical aspects, But I think there's probably also a kind of a cultural and general behavior on calls aspect here, right? If I call my local bank and I talk to my Swiss German speaking bank associate, that would be very different than if I call somebody in the UK or if I call somebody in Texas, right? Just generally from the behavior on a call. Is that something that is a factor also? Or people that are using this technology, like they, I don't know, they're aware of this.

22:52Sharath Narayana:Or maybe the person that speaks German and then speaks to somebody in English would be somewhat briefed on the behavior that would be expected on a call. I think that education most enterprises have given to their people, right? and today agents are switching between using an accent product. Somebody sitting in the Philippines is speaking to a caller in the US using Sanus' accent harmonization product, right, to sound more homogeneous. A lot of them actually declared that they're actually using a product like this. So nobody's trying to hide anything. And if you've tried our accent product, and I can play a few samples for you as well, it still sounds like you.

23:38If you think about me moving to the US 18 years ago when I landed in Boston for the first time, I had a very regular Indian accent. This is how I grew up. And now for the last 18 years having spent time here my dialect has tweaked a little. I'm now more easily understood by a native Caucasian speaker here. So what Sanus is doing is make you sound like that. It's still you, but think about if you spent a decade in that country or in that region, how would your typical accent would have gotten influenced? That's the only tweak we are doing. We're not making you or trying to make you sound American or British.

24:23That's not the intent at all. It still preserves your original timber. like if you do a bioprint test you'll still pass the test because it's still you right but just that you sound a little more homogeneous so that the other person understands you better and we've maintained the same thing with language as well like if you use our language translation product we actually have it on iStore and Android Play Store as well as a B2C version it's free it'll still sound like you but speaking Mandarin right i wish i wish i could speak my no that's what the product does and the product has a human in the loop because see uh like we've seen enough instances now where they're speaking to somebody they've enabled sonas but then the person tells them that hey uh i'm a native turkish speaker or an arabic speaker or a mandarin speaker i'm more comfortable especially we've seen this in health care where the length of the call is usually 25-30 minutes they're describing what's going on and they're most comfortable talking in their mother tongue yeah right and and the agent said that hey I have to have this product where I can speak your language right there will be a small latency I can enable that or I can get you a live translator it usually takes between 7 to 12 minutes to get that most people have said hey why don't we just try because I trust you right so then the person enables it and then they start speaking the feedback has been fantastic most of them think the agent has magically learned mandarin or arabic and they're like no your mandarin is not bad you speak pretty well it's a little formal but you speak pretty well and and the person sometimes tells them no i'm using this product which is helping me translate I can see the translation on my screen so I know what I'm seeing is translated the right way.

26:23Especially in healthcare, you have to be doubly, triply sure. But people just love it. And there is about a second to a two-second latency, but you only feel it in the first few seconds. But once the conversation starts rolling, it becomes natural. And people love it, right? The only reason we are not jumping in and waving the flag and going full throttle in it is because, for me, accuracy is the most important thing in language translation because what we are doing with our speech AI technologies is actually building trust between people. The moment you are not accurate, you lose that trust. And you lose that trust, you lose everything.

27:03So the three languages that we have launched today has between 96 % to 98 % accuracy. right and that's my benchmark like I will only launch the next language like Arabic or anything else when I'm sure I can be at that levels of accuracy latency people give you that benefit of doubt it's fine it'll take a second or two but if you say hepatitis C versus B then then you're playing with somebody's life so that's why we're being a little careful and because of that human element in that call, things go a lot better.

Read the full transcript

27:39Sharath Narayana:The latency is fascinating. I mean, is the latency still a technical issue or is it basically almost like a philosophical, like you kind of need to wait for that other person to finish parts of the sentence before you can actually get cracking on the translation? Like in German, it's the, you know, the verb, it's only clear at the very end if it's like I have done it or have not done it. Right. So it takes a bit of time to process it. So it's actually both, but it's still, I would say, technical more than philosophical, right? Because in accent, in enhancement, you only need one triphoning pair to start dissecting speech.

28:23In language, you at least need between three to six phoneme pairs, right? Every phoneme pair takes about 50 milliseconds, right? So if you need between three and six phoneme fares, you need between three to 500 milliseconds. And then that's one way. And then it's the other way around. So it takes about a second, even technically. And then you add another second for the network latency. That's something that I can't change. Though I do a lot of my inferencing on device, but the person is sitting 6 ,000 miles away or 7 ,000 miles away or 1 ,000 miles away. So there's the typical network hops involved.

29:03That will take you another second. So it's still technical. I don't know from satellite if you can beam it faster and that latency can come down to 100 milliseconds, which Mr. Musk is doing. It can make it better. But there is a technical challenge to latency, which is not very easy to come around. We've tried doing prediction there. see I can build a very massive model on the cloud so then if I do a prediction engine then I can't turn it locally on the device then I have to go to the cloud as soon as I go to the cloud then there's this whole data sovereignty challenges that comes in especially in healthcare and banking they're not very comfortable sending their traffic to the cloud

29:51but in prediction my philosophy is most ai systems are still tokenization engines by definition they say they can never be 100 accurate they can be 98 accurate they can be 99 accurate it's never 100 accurate so i'm like do you want to take the chance uh of a prediction engine where you might be wrong once in a thousand conversation but the time you were wrong something catastrophic happened so which is why we have decided not to do prediction. We are like, we'll compromise the latency a little. But generally, at least so far, people get around latency pretty quickly because they usually start the conversation in English and then they switch to a different language.

30:41So by then, they already know the person. They know this person is a healthcare expert or a financial expert. They are talking in their language to help them. So the leverage is on the agent's side. down on the customer side. So since there is that circle of trust built, it usually works better.

30:58Sharath Narayana:Can we just go back to that on-device? I don't think I fully understand. So usually, you know, companies in this space are like, it's all cloud and you're connected by an API. So you guys, when you say on-device, like this is a massive enterprise deployment. So how does that work on-device? So we have an app that you download on your desktop, laptop, mobile phone, right? And you use Sarnus on your device. And the on-device decision came in with a lot of painful experiences I had with my last company. In the last company, we built everything on the cloud. We were a cloud-native company. We grew very quickly.

31:35Our first$15-20 million of revenue happened very quickly because we offered something which was very innovative and noble. All the early adopters took it. But then if you have to build a$100-200 million company, you have to go to large enterprises. and we were stuck in an endless CIO loop of data sovereignty, people asking us to sign indefinite indemnity, like equitable measures if something goes wrong. And I'm like, hey, I'm not selling my company to you to win a 100K deal, right? And that's why when we started Sonus, I'm like, hey, the world is going towards building the largest model there is.

32:11Like we were hearing 10 billion parameter models, then we heard 100 billion parameter model, now we are hearing a trillion parameter model. right? But the bigger the model, the more money you spend on compute, right? If you look at all the hyperscaler stocks going through the roof, it's because everybody's compute bill is going through the roof. And I'm like, hey, at some sense, at some time, look at Apple. They are an on-device company, right? While everybody is spending money on building data centers, Apple is sitting idle on cash because they know they have an on-device play, they have a locked-in ecosystem and ultimately they'll win with the watches they have.

32:49So I'm like, why don't we take a very contrarian approach and try and build as much as we can on device. So of the 14 algorithms we have built, 13 of them work on device. Language translation has one small component which is outside the device, and we're trying to get that into the device as well, right? That's on-device for me because, one, you protect data sovereignty. You build a very high gross margin business. You're not transferring the cost of compute to your consumers, right? And then, finally, I think when you're local, your latency is very controllable and manageable. If you're not on the device, then you're dependent on the network latency.

33:35Sharath Narayana:And there's been quite a lot of progress in small models, which I would assume that you guys are building, right? I mean, it's quite exciting. We spoke to a company recently called Newphonic. We did like a startup meetup in London. They're working on it. So there's been quite some breakthroughs, right, in the past 12 to 18 months, which I guess would go in your favor. if the capabilities of these small models work better. Today, my mobile phone can actually run a 20 plus billion parameter model, right? So the on-devices are becoming powerful by the day. The hardwares are getting more sophisticated by the day.

34:21So I think our bet, I can see that paying off. But the way we've built our models, It works on a$100 laptop, like 1.6 gigahertz processor Pentium machine.

34:39Sharath Narayana:Our model will work on that, right? So that's what we built it for because we built it for the lowest common denominator, right? Yeah, you give me a MacBook Pro. Yes, my model will work even better. I can put a more enhanced model on it. But I built it for the lowest common denominator because the logic always was, hey, don't ask your customer to change anything. If somebody has a Pentium machine and they like it for some reason, they should be able to use Sanus on it, right? It should work on a$20 headphone if they decide to put Sanus on their headphone, which one of the headset manufacturers are actually trying now, trying to put Sanus on their headphone, right?

35:19It should work on a mobile device. It should work on an old laptop. That has been the goal. So even before this concept of SLM became popular, what we actually do on the device is a matching algorithm. So we've built a library now of 68 billion phonetic pairs, which covers every syntax of every language on this planet. And what it is doing on the device, it's essentially matching phonetic sounds. You basically choose an input language, you choose an output language, and it's just matching, or input accent, output accent. It is just matching the phonetic sounds in real time. It is almost like a stateless system where it is taking an input phonetic sound, it is matching it with a phonetic sound that's there in the dictionary, and outputting a similar phonetic sound in a different accent or a different language.

36:15That's all this matching algorithm does. So it's even one step below a typical SLM. You can call it an SLM if you like to call it an SLM, but it's actually a very simple matching algorithm that's there on the device. That's why the algorithmic latency is sub-100 milliseconds, which is a technical mode for us, and that's what is helping us do it better.

36:40Sharath Narayana:Can we also touch on some of the financing? I think one of the most recent was a Series B you guys raised. about a year ago. So you having written the first check for Sonos, how did it evolve from there? And what were some of the conversations like and how have they changed in that AI boom since you guys started? I've been grateful. I think the first seed round in the series A happened with Insight coming in with a large investment during our series A. This was early 2022. That was all on the art of possible, right? We said we want to build a speech lab. We had this white paper. We had built some small MVPs.

37:23I had some bit of credibility of having built Observe. We raised it on that. I think we benefited from the 2021 SaaS boom. This was towards the end of 2021, early 2022. We raised a decent chunk of capital, about$35 million, just as the company was starting, right? That gave us the war chest to not have to chase revenue from day one. We could take 18 months, go deep, build all these algorithms, fundamentally build it right. And then about 18 months ago, we announced it about 12 months ago, but the round happened in October 2024. Once we felt we've had that product market fit moment, we had over$10 million of revenue.

38:08That's when we raised our Series B because now I felt, okay, now this is repeatable, scalable, right I don't have to take 106 flights every year so now I think I can have a team who can have a playbook now and they can build it that's when we did our series B our lead investor inside with Quadral came together and then a lot of our customers wanted to come and invest like three of our customers have invested in our business which is generally a good sign of you doing something right so it came together really well but doing that series B was the first time I started hearing all these whispers where there was a camp which felt we are doing something magical and wanted to invest.

38:51There was also a camp who was saying, hey, but a majority of your business is in the human-to-human communication space, but that might all go away. And I think one of the investors had said, in 12 months, call centers will be dead. It'll be only agentic conversations and nothing else. And my logic at that time to him was, I get it. I think agentic companies coming in and more and more agentic interfaces being native will only be good for the business. It will actually not reduce the number of conversations. It will actually increase the number of conversations. I gave him the analogy of Uber.

39:27It was not that we were not taking taxis before. We were taking it maybe once a month, once a quarter. But after Uber came in and made it democratized, we take it every day, right? With more and more agentic systems coming in, if voice becomes the new keyboard and voice becomes the primary interface for enterprises versus digital, there'll be an explosion in conversations, right? Yes, a lot could be handled by an agentic interface, but the human-to-human communication will not go away because at the end of the day, people buy from people, people want to talk to people. I don't think that's going away anywhere, right?

40:01and 18 months out I still see the same moments I think there's an explosion in AI investment people still say there was an article very recently by a hedge fund from Apollo where they said with all this AI boom why are jobs increasing in the Philippines it should have been the other way around that all the jobs should have gone away right and and I feel validated with that thought process coming right where I always think 18 months ago maybe 100 % human to human, zero human to machine. Today it's probably 5-10 % and still 90%. That number might flip around in the next decade. That 10-20 % would be human to human communication.

40:44Everything else will be human to machine. But that 10-20 % will be a part of a much bigger pie. It will still be relevant. and I've always told investors that we are a speech company because even a machine is talking to a human, right? Machine can be perfect, but a bot might call a human when he's driving in a car at 50 miles an hour with his kids yelling in the background. It still needs to understand a human and a human could be from Bangladesh, from India, from Latin America. He might have an accent or she might have an accent or they might prefer speaking in a different language. So a machine to human understanding is also where Sonus can play a big role.

41:27So I don't get phased by it. But what that realization taught me was, I think in Sonus, if you realize, going on device, right, building for customer experience, prioritizing human first. We've taken a lot of contrarian approaches to building. I'm like hey after that round we had a big war chest which we even now have I said let me not depend on investor capital to build my business let me build a great business a profitable business I have a 96 % gross margin have insane demand we've grown from 14 million to 62 million in the last 18 months right let me just build a profitable business if people like to come and invest if I like the terms, I'll take it.

42:13Otherwise, I'm very happy building this from here. So that's been our approach.

42:18Sharath Narayana:I guess that's why it took me a second to kind of wrap my head around it. I understand the contrarian approach because I kept thinking, you know, yeah, cloud, SaaS, et cetera, but very different, very, very, very different, fascinating. Let's close on the roadmap, 2026. You know, you kind of teased there that things are going to be launched in the next quarter or something. Is there anything you can share with us, anything that our listeners should watch out for? Absolutely. I think still the vector we're talking about is, I call it enterprise voice. CX has been a focus, but now we're seeing it expanded across the enterprise.

42:52So language translation will become mainstream in the next 90 days. We're targeting five languages where we'll have more than 95 % accuracy in the next 90 days. And by the end of the year, we'll take that up to 30 languages. So we think language translation will become mainstream. The universal accent translation should be launched where it will be any to any. Today it is any to few. That's a big one coming up in the next 90 days. And then what you'll also see is go after other vectors. We are thinking about Telco as the next big horizon. What we're seeing is, hey, what we're building is still B2B.

43:31But I think Sarnas as a speech company can go on a network, can go within a device. and it can impact millions of people. I think we are attempting to go into a device. We are attempting to go on a network. And I think those will be two big horizon bets for us. And we see that being mainstream in the next 12 months. And from a productization standpoint, today, Sanus is still a black box. I think all the algorithms that we've built, we've never made it publicly available for people to build. as more and more people are building voice-first applications. We think Sanus can be that infralayer. We released our noise cancellation SDKs into the public domain about a month ago.

44:14The usage is through the roof on that. But I think in the next 90 to 180 days, you will see us launching our accent algorithms, our enhancement algorithms, our language algorithms, all as an SDK. We'll put it up on whatever LifeKit or PipeCat so that more and more developers can build speech applications on top of Sanus. I think the more and more speech gets democratized, I think this will be a beautiful world. More commerce can be enabled and unlocked because of that. And I want Sanus to play a small role in all of it.

44:46Sharath Narayana:Great closing worlds. Thank you so much, Sherrod, for taking the time today. This was fascinating. Looking forward to sharing this with the audience. Thank you so much for the opportunity, Ken.

From the publisher

Sharath Narayana, CEO and Co-Founder of Sanas, joins SlatorPod to talk about the evolution of real-time speech AI, the rise of accent harmonization, and why voice will become the next major enterprise interface.

Sharath traces his journey from engineer to entrepreneur, including the founding of Observe.AI before launching Sanas alongside Stanford researchers focused on solving low-latency speech processing.

The CEO explains that Sanas initially operated as a speech lab focused on improving human understanding in conversations. The company developed multiple algorithms covering noise cancellation, speech enhancement, accent harmonization, and language translation before discovering that enterprises were most willing to pay for accent-related technology.

Sharath emphasizes that Sanas’ accent technology is not designed to erase identity, but to improve clarity and reduce friction in customer interactions. He says enterprises, especially contact centers, adopted the technology to improve first-call experiences and reduce mistrust between agents and customers.

He also discusses Sanas’ focus on on-device AI infrastructure rather than cloud-only deployments, where running speech AI locally improves latency, protects data sovereignty, and lowers compute costs.

Looking ahead, Sharath says Sanas is preparing broader launches for real-time language translation, universal accent translation, and developer SDKs that will allow third parties to build voice applications on top of the Sanas platform.

More from SlatorPod

All 39 episodes
#285 Real-Time Speech AI and Accent Translation with Sanas CEO Sharath NarayanaSlatorPod · 45 min
Listen in VO