How Hume AI Is Teaching Voice Systems to Understand More Than Words

30 Sep 2026 · 59 min · 26 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Hume AI argues voice AI must be evaluated and improved using speech signals and real-world outcomes, not just transcript text. They build an evaluation/benchmarking infrastructure that captures emotion, turn-by-turn dynamics, and acoustic/language variability so voice agents are reliable in production.

Guests

Andrew Edinger, CEO of Hume AI. Background: leads a company built on a decade of research by co-founder Alan Cohen on semantic space theory; Hume previously pioneered text-to-speech and speech-language modeling. Hume co-founder Alan Cohen and engineers joined Google DeepMind earlier this year under a non-exclusive technology licensing agreement; Andrew took over as CEO and shifted focus to evaluation/training infrastructure.

Key claims

Voice has a “listening problem” when judged by text; word error rate misses meaning, emotion, and safety-critical differences. Voice is multidimensional (emotion changes per turn), so benchmarks must score audio/behavior in realistic scenarios. They built a framework (“CLEAR”) using millions of human judgments and 15,000 samples across 60 models.

Notable examples

“I’m fine” can be confused/scared despite correct words; a healthcare agent saying “fine” could be dangerous. Penicillin analogy: optimizing for measurable text can mislead real-world decisions. They also discuss vocal bursts, dialects, interruptions, and noisy environments (e.g., two people speaking, background screaming).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Voice Systems and Their Limitations

0:00 to 1:00

Explore the inherent challenges of evaluating voice systems primarily based on text.

“Voice actually has a listening problem because it just reads the transcript.”

Talking with Andrew Edinger

1:14 to 2:15

Core discussion with Andrew Edinger on advancements in voice AI technology.

“We'll hear more from them in just a few.”

Understanding Emotions in Voice AI

2:15 to 4:00

Discussion on how voice technology must align with human emotional understanding.

“This is I was just as I mentioned before we went on, I remember the first time Grant and I tried Hume's models and being really impressed.”

The Complexity of Real-World Voice Interaction

4:00 to 5:56

Exploring the real-world challenges and expectations of voice systems in various environments.

“At the same time, it's so much easier to speak and get the work done that you would like as opposed to typing.”

The Complexity of Real-World Voice Interaction

6:14 to 6:55

Exploring the real-world challenges and expectations of voice systems in various environments.

“AI pilot projects are everywhere, but real business value takes more than experimentation.”

Testing and Improving Voice AI Models

7:09 to 9:32

Insight into the processes for testing and enhancing voice AI models based on real scenarios.

“How do you go about, like, setting, like, basically testing all of that stuff, right?”

The Future of Voice as an Interface

9:32 to 14:01

Discussing the role of voice as a primary interface and the integration of multimodal communication.

“And it first starts with just being able to measure, right, you know, kind of where you're at, then it's how do I actually evaluate, right, where I should go fix things and get all of the analytics.”

Hume AI's Data Processing Pipeline

14:01 to 15:44

Learn how Hume AI utilizes a robust data processing pipeline to enhance voice systems.

“How is Hume actually detecting that information then?”

Understanding Emotions in Voice Interactions

16:10 to 20:39

Explore the multidimensional approach to recognizing emotions in voice conversations.

“In talking about emotion, we tend to think of this as all being in like neat buckets, like happy, sad, angry.”

The Ethics and Applications of Emotion Detection

20:43 to 24:13

Discuss the ethical implications and potential applications of real-time emotion detection technology.

“So all of this has me just thinking like this kind of technology would not just be helpful for agents.”
Show all 26 chapters

Challenges and Opportunities in Voice AI

24:13 to 28:00

Understand the challenges of creating reliable voice AI models and the potential for specialization.

“Like, I might take it, but, like, I'm not going to feel comfortable coming back to this, like, automated thing from my doctor.”

Harnessing Voice Data for Competitive Advantage

28:00 to 29:09

Explore how private evaluation infrastructures can leverage institutional knowledge.

“And so it's now about the abstraction level above that and how do you apply your intellectual capital and your horsepower and your domain-specific knowledge against that to drive true competitive differentiation.”

Addressing the Cold Start Problem in Voice AI

29:10 to 30:56

Learn about the challenges of cold start in voice projects and how to overcome them.

“We're able to run that through our system and deliver back to them a very comprehensive analysis and set of analytics that says, hey, look, these are the types of things that you need to really understand.”

Achieving Consistency in Voice AI

30:57 to 33:26

Understand the importance of maintaining performance consistency over time.

“So now how do I go back and see where I was at and start that process all over again to take sort of the one plus one, right, equals three approach?”

The Future and Challenges of Voice AI

33:27 to 35:37

Discuss the current state and future potential of voice AI technologies.

“You just got very, very clarity of focus from so many years of doing it.”

Depth of Conversation in Voice AI

35:38 to 37:51

Examine the depth and complexity needed in voice interactions and customer service.

“Are other languages as far along as English or farther?”

Unlocking Voice Data for Better AI

37:52 to 40:13

Discover the untapped potential of voice data in enhancing AI performance.

“And that's where people are working and investing really heavily right now.”

Evaluating Naturalness in Voice Models

40:14 to 42:01

Learn about the balance between natural sound and functionality in voice models.

“You guys really saw an opportunity and took a risk.”

Evaluating Voice AI Models

42:01 to 43:30

Learn how to effectively evaluate voice AI models beyond traditional metrics.

“are actually doing the job to be done correctly.”

Safety in Voice AI

43:30 to 45:08

Understand the importance of safety measures in voice AI applications.

“I'll leave it up to you if you want to get into that.”

Voice Identity Verification Challenges

45:08 to 46:32

Explore the challenges and advancements in voice impersonation detection.

“So that's why we're very excited about the position that we're in from an infrastructure layer and continually to add value for folks over time.”

Voice as an Interface

46:32 to 49:10

Discover the complexities and potential of voice technology as an interface.

“Like, I keep thinking of this as voice, as a gateway to real agentic work pretty quickly.”

Personalization in Voice AI

49:10 to 51:01

Learn how personalization is key to enhancing voice AI systems.

“Like just if you think about it in terms of let's say I had a perfect system.”

Future of Voice Technology

51:01 to 56:00

Gain insights into the anticipated developments in voice technology over the next few years.

“Where do you feel like the next two or three years take voice tech?”

The Social Contract of Voice Interaction

56:00 to 57:28

Explore how voice technology is changing social interactions and cooking.

“But in voice, you should just be able to speak very naturally and very freely.”

Closing Thoughts and Future of Hume AI

57:28 to 57:58

Reflect on the discussion with Andrew and where to follow Hume AI's work.

“Yeah, I think we can all relate to that.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Voice actually has a listening problem because it just reads the transcript. And when you fundamentally evaluate voice systems based on text, you cannot perform. In the wrong setting, right, that actually could be the difference between something very catastrophic happening or not. Fundamentally, we just do not believe voice systems should be evaluated to start based on text. I think you need to look at this as a multidimensional problem, and voice fundamentally is multidimensional. Text is rather flat, right? So to start with, you have to be able to score it and leverage it, but also do that on a turn by turn basis.

0:39Our social contracts actually will have to change a little bit, right? But the job of the AI systems, right, and the ones that will win are the ones that don't force us to do that, right? And that's the whole point of being natural and consistent and reliable and acting and aligning like we are, is you should just be able to talk freely and not have to change your patterns in order for the AI systems to work. Welcome, humans, to the Neuron AI Explained. I'm Corey, and I'm here with Grant, as always. How are you, man? I'm doing good, Corey. How about yourself? Doing well. Doing well. Before we dive in here, today's episode is brought to you by Seth.

1:14We'll hear more from them in just a few. But today's episode, we're going to talk about how AI voices have gotten remarkably condensing. They can clone voices, switch personalities, respond almost instantly, and increasingly hold conversations that don't feel like talking to an old school voice assistant. It's much more human, and understanding humans are two very different people.

1:37Corey Noles:And to have that conversation, today we're talking with Andrew Edinger, CEO of Hume AI, a company building exactly that, voice technology, around that second problem. Understanding all the information carried in speech that disappears when you turn a conversation into text. And the timing here is fascinating. Earlier this year, Hume co-founder Alan Cohen and several engineers joined Google DeepMind under a non-exclusive technology licensing agreement, while Andrew took over as CEO of Hume. Since then, the company has increasingly focused on the data evaluation and training infrastructure behind Advanced Voice AI.

2:11Corey Noles:And we're really excited to have a conversation all about that. Yes. Before we get started, please take just a quick second to like and subscribe to the channel so you never miss an interview or one of our wild live streams. Andrew, welcome to the Neuron. Thanks, Corey. Thanks, Grant. Great to be here. Look forward to the chat. Same. We're really glad to have you. This is I was just as I mentioned before we went on, I remember the first time Grant and I tried Hume's models and being really impressed. So when when the opportunity came along to to have you on, I was like, yeah, let's see what's what's going on there.

2:42And it sounds like you guys have been doing some some really amazing things. Yeah, it's been a great time. It's been a lot of fun. Hume's got a decade's worth of research in semantic space theory, which, you know, as we talked about earlier, we're complex humans. And that means we have many emotions that are present at once. And as voice takes over as the primary interface in AI, it's got to really align with how humans act and expect it to perform. And we're clearly very well positioned to help that and look forward to chatting with you about it. Was there a moment in voice AI that kind of convinced you this wasn't just another interface layer, but maybe a fundamental part of the AI stack moving forward?

3:23Yeah, I mean, I think when you look at all of the different modalities that are happening inside AI at the top level, right, multimodality will be present for a long time to come, right, across a wide variety of spectrums. And then you start to see them merge, right, with humanoids and wearables and, you know, a lot of those aspects. And really with voice, it became very easy to create a realistic sounding voice. But it became very hard to align that voice to the real world outcomes that us as consumers or practitioners expected that to perform. At the same time, it's so much easier to speak and get the work done that you would like as opposed to typing.

4:07And as agents start to proliferate inside of the enterprise and voice agents now start to speak with other voice agents, you have this explosion and this massive opportunity. And so it really seemed like a perfect time to sit there and say, look, we need to make sure that everything that happens with a voice on a global scale in every single language is accounted for and aligned to the real world outcomes that the humans that are using that expected to perform. So that's real. It's natural. It's accurate. And you have a very pleasurable experience.

4:40Corey Noles:Yeah. Is that a monitorability issue, a benchmarking issue? How do you solve that? Well, Grant, certainly, you know, once you're in production, you've got to be able to monitor things to make sure that they stay consistent. But we're a long way from ensuring that every use case can go from ideation and testing and fine tuning and the RL loop to get into production. Right. That's that's that's easy to say. And that loop is like easy to draw. It's really hard to do when you start to think about languages, dialects, vocal bursts, noisy backgrounds. I was just with someone the other day and they're like, yes, we need our model to be able to have two people speaking to it and have one person get interrupted by a significant other screaming at them in the background.

5:24Say, wait, I'll be right there. Right. And that continue. And when you do those things, models shut down because they weren't architected to be able to handle that. So there's just so many of these different permutations that are coming up. Not the least of which is like, look, in the end, you generally aren't calling customer support when you're in a studio like we are today. Right. You're generally in a car. You're generally multitasking. You could be calling your bank while you're walking the streets in New York or San Francisco, which means you have background noises and you have like less than ideal acoustic conditions.

5:55But you still expect the other side to perform appropriately. These are all very hard variables that you have to be aligned to what we call the real world scenarios by which humans expect them to operate. Yeah. I'm not great at understanding under those conditions either. Yeah, exactly. So you're right. Right. So think about AI, right? AI pilot projects are everywhere, but real business value takes more than experimentation. SaaS helps businesses move from scattered AI activity to measurable results by focusing on the foundations that matter most, like strengthening data readiness and governance, aligning AI efforts to clear business outcomes, defining a practical AI strategy with meaningful KPIs.

6:39SaaS delivers guidance for practical steps to bring structure, focus, and confidence to your AI investments. SaaS can help you scale what works and turn AI momentum into business impact. Grow faster and scale confidently with trusted AI solutions. Visit www.sas.com slash smb. That's s-a-s dot com front slash s-m-b. We also have additional resources for SaaS in the description below. Now, back to today's video.

7:10Corey Noles:Oh, my gosh. How do you go about, like, setting, like, basically testing all of that stuff, right? Do you have to come up with these intricate scenarios that are difficult to test the models in and perform them? Like, what's the process there? Yeah, it's a great question, Grant. It's a lot of fun, actually, because every single day we get to work on these use cases. And, look, fundamentally, here's the way I would think about it, right, is that voice actually has a listening problem. because it just reads the transcript. And when you fundamentally evaluate voice systems based on text, you cannot perform.

7:47And let me give you an example. I'm fine. If you hear the words or see the words, I'm fine in a transcript, it could mean you really are fine. It could mean you're confused. It could mean you're scared. It could mean a multitude of things. And in the wrong setting, right, that actually could be the difference between something very catastrophic happening or not. right as one example to polarize it right and so fundamentally we just do not believe voice systems should be evaluated to start based on text because measuring word error rate really doesn't account for those types of things when you don't understand the way in which humans actually feel the way they actually sound the way they actually want to converse the way they actually want this to be reliable consistently over time and so we built a system based on our decade of research and the way in which we built our models that fundamentally said, look, it should be evaluated on the voice and it should be aligned to real world outcomes and it should be aligned to the way in which humans think, act and believe this to be.

8:50And so our benchmarks and our framework with a million different human judgments and 15 ,000 samples across 60 models was a really big piece of research for us to just fundamentally say, we believe this is the new way that things should be measured and it's gone extremely well since then makes sense so does that mean that's a true

9:10Corey Noles:like uh voice to voice uh like system this in this case or or how how exactly is it different well grant i mean you can measure this across all modalities right but yes you truly are using the voice and humans to evaluate the outcomes of all of those very specific scenarios and that's exactly what we do. And so we've built the platform and the infrastructure to help everyone improve all of their voice AI systems. And it first starts with just being able to measure, right, you know, kind of where you're at, then it's how do I actually evaluate, right, where I should go fix things and get all of the analytics.

9:49And then how do I have a process to actually then improve those so that I could affect the outcomes in my model, right? And so if you reverse engineer that and you say, hey, look, I'm going to work with folks and focus on the North Star metric of how fast I can help improve your model output and outcomes to be more well aligned to real world scenarios, you're in a really interesting spot because everyone's looking for outcomes in their model improvement. And that's exclusively where we're focused. Got it. Nice. You know, I always just just a quick step back. I always hear people and I think it as well, that voice will one day be the primary interface.

10:31But I never really hear why. I would love to hear your thoughts on what it is about voice that you think makes it better. Yeah, well, it's not necessarily if it's better or worse. It's just going to be a primary modality, right? And so, you know, and it's going to be pervasive. And you have a multitude of languages, a multitude of dialects, and a multitude of ways in which we as humans are just very comfortable communicating. Some might do better via text, right? Some might do better via voice. But voice is a primary interface, and it's not going anywhere. However, it's complicated, right? And trust is a really important thing, right?

11:13If you use a text-based system and there's something that's inaccurate in there, A, you might not know. And B, if you do, we're kind of now trained to be like, okay, well, 90 % of this was awesome. I wouldn't say those three things. So that thing doesn't seem right. I want to go check that. You start using voice and it gives you the right word said the wrong way. You don't trust it. It gives you the wrong word in the right voice. You still don't trust it, right? If it doesn't handle the conversation in the way in which we're chatting right now, you're going to say this doesn't feel natural, right?

11:48And yet this is what we want. Like you don't actually have a problem calling your bank or your cable company to perform a very specific customer service action talking to their voice agent. The problem you have is that the voice agent doesn't actually understand and can't interpret and can't relate to you. Right. It doesn't hold the tools in the right way. It doesn't have that conversation. And if you could do that and make it be very efficient and get faster outcomes, far more reliable and cover a far more multitude of use cases, you would love it. If an agent caused an incident at your company tomorrow, what evidence could you actually produce?

12:25Most teams can see the prompt and the final answer, but the commands it ran, files it changed, or systems it reached in between, that can disappear the moment a session ends. Origin gives companies endpoint AI observability by recording an agent's work as a trace from who started it and what they asked to what the agent touched and changed. See what that looks like at originhq.com slash AI explain. And now back to today's video.

12:52Corey Noles:Well, I think about, so there's voice, but then I also think, you know, when you and I are having a conversation, there's also gesture. So I think at a certain level, it has to be not only a voice system, but it also has to have vision capabilities as well. Because I might be saying, no, no, no, it's the thing over here. Or, you know, because we have multiple, you know, senses with which we can communicate with one another. And so I think that the most aligned voice system would probably have some vision capabilities as well. Like an avatar element of some sort at least or something or something.

13:24Well, it's a really good point. And so our models take into account in real time facial expressions as well as the audio expressions. And you can look at them in separate channels and in separate dimensions. or you can marry them and ultimately derive sort of the implied score, the implied emotion as a result of that, right? And we do that across 48 different emotions and thousands of different variables and emotional states on that. And so folks like what love working with us because they take raw audio in and you get 600 different pieces of metadata out, that really now starts to deliver to you the type of signal that you would want as really golden set of ground truth that you can use to go inform how you would train these systems, how you would evaluate them, how you would ultimately fine-tune them and help them achieve what you would want, and Grant, to your point, both from a facial as well as a raw audio perspective.

14:20Interesting.

14:21Corey Noles:How is Hume actually detecting that information then? That's cool. Yeah. Well, I mean, Grant and Corey, you had discussed that a while back you had listened to some Hume models, which were very early pioneers in the text-to-speech game and in the speech-language model game. And this really goes back to Alan's research from 10 years ago around semantic space theory. But I think what's really interesting that I always like to say is that our research is not just a set of research that's very powerful. It's manifested into a very robust data processing pipeline with an ensemble of a dozen different models that's worked at hundreds of millions of hours at scale with human judgments infused into that.

15:04And so we're able to take that raw audio in through this data processing pipeline and deliver these types of rich insights out and then build the systems and the tooling around that to go happen. And so this isn't just research. This is actually working code in models in a data processing pipeline that scales across thousands and thousands of concurrent GPUs running together. Wild. Yeah, it's a lot of fun. It's a lot of fun. It's a lot of fun. We're very privileged to be able to do this every single day, as many others are in this space. It's a really fun problem set. It is. So your agents can talk to each other, but they still can't think together.

15:41That's the gap OutShift by Cisco is closing. OutShift is Cisco's incubation engine, building the Internet of Cognition, shared intent, shared memory, and guardrails that allow agents to work together from anywhere. Open source protocols and code, a system for intelligence that scales. Read the paper and experience the demo at outshift.com. That's O-U-T-S-H-I-F-T dot com. Check out the links in the description. Now back to our video. In talking about emotion, we tend to think of this as all being in like neat buckets, like happy, sad, angry. But your research describes it as more of a high dimensional spectrum.

16:25And I think that makes sense because there's often a blend there. And I'm wondering how do you pull that out and recognize the differences that that makes? Well, Corey, not only is there a blend, but it changes per turn. So you need to be able to not just identify it, but you and I might be having a very friendly conversation right now. Let's say the emotion was friendly, right? You might say something. I might get aggravated and all of a sudden get pissed or angry or confused or agitated, right? And so now what has to happen? And so I think you need to look at this as a multidimensional problem.

17:03And voice fundamentally is multidimensional. And to the earlier point, text is rather flat. So to start with, you have to be able to score it and leverage it, but also do that on a turn-by-turn basis. Because fundamentally, that's really where the power happens. So it's okay if there's a voice agent that works very well where you start to have a real conversation. That's great. We're having a real conversation. This is fantastic. But, like, are you now listening to me? Are you understanding, right, exactly what I'm saying and how I'm saying it? Are you then understanding my expression so that you can adapt and change and be accurate for a minute or two minutes or three minutes or 32 minutes if you're, you know, used to kind of dealing with some customer service, you know, experience?

17:51Which I won't name. And then lastly, can you be reliable and do it time over time again? And so what I just walked through is a framework that we developed called Clear, right? And that Clear framework really enables you to look at all of that in totality, but also on a turn by turn basis so that you can understand how all of that works, how that affects the emotions, which then affects the outcomes that the AI systems deliver and what you need to fix and align and tune for it. How does that happen in real time, I think, is what blows me away. And I'm not asking for, like, trade secrets. It blows me away that when you mention turn by turn, I mean, like, you need to be able to respond.

18:35I expect you to respond in for actions of a second, maybe a couple seconds if you, hmm, you know, if you bide your time a little. For that to be able to happen just on the spot is wild. Yeah. So our processing pipeline works at scale on raw audio. So if you're a call center, by example, and you have 100, 200 million hours of this audio stored up over time, or you're in a regulated industry where you have to have this stored for seven years or nine years, that's very easy to go through and kind of deliver these insights. And I have that ground truth to be able to go and create your agents off of and really mimic how people have truly interacted in your system.

19:17If you are some of these healthcare apps, right, or you're doing coaching assistance or you're deploying these voice agents, right, you first need to be able to just fundamentally understand from the recorded audio that you have, like, what happened? How did that perform? How was Corey? How was Grant? How was Andrew? What did all of that look like so that I can understand that? And you could do that by frame, by five-second intervals. You could do that in real time. You could do it at the end of the conversation, right? And so now we're able to inject this into applications and do this on a streaming basis in real time.

19:49However, the next step of this is now can the application provider, can the agent, can the models now take that and actually affect how Corey or Grant would respond to Andrew in real time? Because I said, Corey, you just pissed them off. You should go to this next step versus just respond. Thank you. I'll look into that. Right. Or Grant being able to say, okay, you know what? We better get this person to an agent. And this is a worthwhile use of a human's time to speak to this very angry customer who's of high value, who's been with us for 12 or 18 different years. And you now would have kind of all of those characteristics.

20:28And that's where this is going. And that's what's really exciting when you start to think about this. But it also first just starts with having to have the infrastructure to understand this before you could measure, evaluate, and improve it. And we're coming at this in the entire spectrum on a global basis. Wow.

20:43Corey Noles:Okay. So all of this has me just thinking like this kind of technology would not just be helpful for agents. I feel like people could use this. If you had a real-time feed telling you like the emotions of the person you're talking to at any given moment, you're like, ooh, they didn't like that. Shut up, Corey. She's angry. I mean, does this product not exist? What are the ethics of this? I don't know. It's very interesting. Yeah. Yeah. No comment there because let's just say there's probably a lot of people in my close circle of my life that would love, right, to be able to be like, hey, I'm telling you to be quiet.

21:16I'm telling you to do this. I'm telling you to do that. Yeah. That might be a little scary, but it's a very good idea, Grant.

21:23Corey Noles:There you go. Anyone can go make that. Maybe they'll use their technology. Yeah. Yeah. Yeah, exactly. Exactly. But for now, we'll just start with the interfaces that we're all, you know, interacting with. Yeah. And it's, you know, what's also interesting, though, is you start to get into like vocal bursts and the dialects and the nuances per language. Right. And some languages, which you have to remember, certain ones, they speak very loud. And you might think from the outside they're screaming at each other, but they actually love each other. Right. Just based on the way in which certain people, you know, communicate with each other in certain in certain different regions of the world.

22:02And so these are all the different variables that you have to account for.

22:06Corey Noles:What can a model that's trained to understand vocal expression do that, let's say, if you're just talking to, let's say, like ChatGPT voice mode or whatever, what's missed between those two things? Like if I'm just chatting with, you know, be it Whisperflow or ChatGPT or something like that, what is that vocal expression AI really, really missing? The real thing that's missing is that our research proved this out in the benchmarks that we did and the research is archived and it's all available to review. And as I mentioned before, largely voice has this listening problem because it's really just, you know, reading the text.

22:47But fundamentally, when you then say, how does that apply in the real world against real world scenarios across all of these models, across all of these modalities? The thing that was so striking about our research is that there was no one model to rule them all. And in fact, you did have to decide between tradeoffs on naturalness and reliability and accuracy and all of these different dimensions. And so you really did have to decide between all those. So what it basically proved was that it's really hard to get all of this right. So Grant, to your point, you could get something that sounds very natural, but it may not be reliable across turns or for a two-minute or a five-minute segment or in a particular use case.

Read the full transcript

23:30The inverse might be true where you've solved a very great use case for a customer service agent talking about billing for one particular product. But the minute you start to introduce a new product or something that you have another family member tied to your plan, right, or you start to ask another question, things now start to get confused and it doesn't hold up, right? And so that's really where we've got a lot of work to do in the industry and where we're very well positioned working with many of the large model providers out there on this and the agent providers to ensure that as they start to look at these use cases that they want to optimize for, they fundamentally don't have to expect users or customers or enterprises to have to figure out what the tradeoffs are.

24:12I often joke, right, if penicillin was the right drug for you to be told, but they said, Grant, we want you to take penicillin, you're just like, that's weird, right? Like, I might take it, but, like, I'm not going to feel comfortable coming back to this, like, automated thing from my doctor. And you're pretty educated and, you know, pretty forward-leaning with that, right? That's actually a real story that we had with a real customer. And even they were kind of laughing. They'll be named nameless. They're just like, you're not going to listen to that, right? And you won't see that in text because it's flat.

24:45And you won't be able to measure that in a transcript because it's probably spelled the right way. And there is no word error rate assigned with, like, you should take penicillin. Make sure it's not on an empty stomach. It's like, ah, cool. It's the right number of words and it said the right thing. But that's not how it performed.

25:02Corey Noles:It reminds me of the problem with generalized models where you push the capabilities really far in one way. We've kind of seen this with the Frontier lately. Like we push really far in coding, but then all of a sudden they are really bad at writing and like, you know, Opus 5. Something over here is less good. Yeah. Yeah. Opus 5 would famously tell you like this is most insane stuff like, oh, like this load. This is the load bearing feature over and over. And you're just like, I can't read what you're saying because it's so like jargony. It makes me wonder, is this bullish for, you know, everyone making specialized voice models and actually just like leaning into, hey, maybe it is better to specialize as opposed to try to make one.

25:39Corey Noles:perfect general model that's incredibly difficult. Like a medical care model, like a legal model. Yeah. Or even down to your company, like a Cori company model. Or to use the medical example, this hospital's model versus that hospital's model. What do you think, Andrew? Yeah, I think these are great ideas, and I fundamentally agree with you all. I think obviously the large hyperscalers are pouring billions of dollars to make sure that they get this right at massive scale across all of those languages and provide the tooling and the infrastructure for people to be able to go build on top of that and not have to find other things.

26:18And they're doing a very nice job at that. Don't compete with scale is what you're saying. Well, but the flip side, right, is if you think about it, right, in all of the vertically enabled AI solutions that are out there that primarily took back office legacy human workflows and put those into LLMs just to make those things easier, you're seeing the same thing happen with respect to voice. And you're seeing that happen on a language by language and country by country basis. I mean, one of the most fun things that I have is waking up in the morning and going to my Slack channel to see who wrote into our website, because invariably it is someone in a country that you cannot believe that they are building something so specialized and so powerful and so impactful for their particular use case that it really is just such an exciting time, right, to be alive.

27:07And you are seeing a lot of this, whether it is in Africa or South Africa or in the Far East or in Australia or, you know, in the Netherlands or even Canadian access. I mean, you're seeing it literally all over the place and it's really exciting. And I think in the end, right, there's not going to be one thing that rules at all and there's going to be a multitude of options. So I guess to your point, there is an opportunity just like there is in these vertical AI solutions with those workflows. There's opportunities where you have an advantage by understanding how humans should perform with respect to voice in a particular function.

27:43Providing the nuance that scale can't, I guess. Absolutely. Absolutely. Absolutely. And look, at the end of the day, access to the network, compute, storage, data, and underlying math is getting commoditized, right? And so it's now about the abstraction level above that and how do you apply your intellectual capital and your horsepower and your domain-specific knowledge against that to drive true competitive differentiation. And we believe this harness and this infrastructure we've built is going to allow people to have this private eval infrastructure and effectively compound all of the institutional knowledge and capital that they've got across all of this with respect to voice to deliver those outcomes in an outsized manner.

28:31And that's where we believe it's really fun because, you know, if you just are publicly releasing all of the data for benchmarks, people are just going to go benchmarks and get models to climb the leaderboard. and you're just going to go rely on those and it's not really accurate. So private evals is really where this is all at, which also supports the folks that have this institutional knowledge are going to be able to create very impactful models on a global scale by language, by use case, by vertical. Wow. So let's say company A calls you and says, we're getting into voice or we're in voice and we don't like where we are.

29:07what does that look like when when you come into the picture yeah well the first question is why don't you like it sometimes they know sometimes they don't know uh sometimes there's a cold start problem where they just can't get a voice to even sound real before they even kind of get to how does it perform in real life so we help with the cold start problem given our heritage in these things. We also help them identify, assess, and understand where the model is or is not performing with respect to their peers, with respect to their use cases, and with respect to how they're fundamentally expecting it to perform.

29:45We're able to run that through our system and deliver back to them a very comprehensive analysis and set of analytics that says, hey, look, these are the types of things that you need to really understand. Here is how you would evaluate them. Here's how you would fix them. And most importantly, here is what you would do to understand regressions and how this performs continually time over time, because it's not enough to go through all that and get it into production. It has to stay that way consistently. And you have to be able to measure and account for that. Right. And so we work with people to design that on a holistic basis so that they can ensure it stays that way.

30:25Because then what happens? You get a new customer, you get a new use case, you get new feedback, right? You realize that that was great, but now people are asking for other things because you got it to be so good. So at some point, you can be a victim of your own success where it's like, I nailed that. Corey and Grant can chat for five minutes and this thing sounds literally like it was on their podcast, right? But now you start to go to the next topic, the next thing, the next language, and all of a sudden, right, I don't want to pick on one of you, But one of you starts to veer off and you're like, oh, wow, I didn't account for the fact that when I tried to go do that.

30:58So now how do I go back and see where I was at and start that process all over again to take sort of the one plus one, right, equals three approach? Yeah.

31:07Corey Noles:That is interesting how like your frame of reference always changes. This is like a classic thing in AI where we get used to a certain level of capability and we think, wow, this is amazing. And then we, you know, we get bored of it in a month because we're trying to push it to the next capability frontier. And I've heard that, you know, a lot of software engineers in particular are very susceptible to this because, you know, when you don't have a specific like, okay, this is the end date. this is the feature we're shipping and then we're moving on. You just get in this loop where you're like, well, I can make it even better and I can make it even better.

31:42Corey Noles:And you have to be really careful with that in the age of AI when everything is so easy. How do you handle that? How does your team handle that at Hume and how do you keep control of like what's most important and don't get like featureitis and go off on all these different tangents? Featureitis. I like that. Featureitis. It's a bad disease. Yeah, it's a bad disease. Look, we're in a really fortunate position because we've built this infrastructure for ourselves over many years. And we've delivered models against this. I mean, we're no longer competing or delivering those models on a consumer basis.

32:19So our infrastructure is very hardened across hundreds of millions of hours at large scale. And we've been doing this for the last year with large model builders and many others. And so for us, it's really about repeatability, scalability, and consistency, right? And I think if you just go back to the workflows that we talked about earlier and making sure that those workflows are best represented into vertical AI solutions so that they work for the end users there, the same thing is true with respect to voice, but it's harder because it is multidimensional. And so how do I make those workflows very easy to repeat consistently?

32:54And so a couple of good examples is you've got large model providers that are doing checkpoints on their model every couple of weeks. Right. And every couple of weeks, they want to run the regressions and make sure that the regressions didn't change with respect to where they were last time. So they didn't make one change and break another. At the same time, they rolled out five, 10, 20, 30 new voices, new use cases, new languages. Right. Any combination in there. Right. And so you really just need to provide a consistent, reliable and scalable framework for them to go do that. And that's what we perfected because we had that for ourselves.

33:24And now we're just opening it up to effectively be the Switzerland of voice AI so that everyone could use this and provide this type of research infrastructure for folks so that they don't have to go build it themselves.

33:35Corey Noles:Yeah. You just got very, very clarity of focus from so many years of doing it. That's cool. Yeah. Yeah. It was a lot of fun. How far do you feel like voice AI has come and how much farther do you feel like there still is at the horizon to go? it's come a long way but we're still in the early days there's no question about it right for all of the reasons that i had mentioned and i think primarily right you have to look at this on a global scale and you have to look at as i mentioned earlier the different dialects the different vocal bursts the different mannerisms and the way in which everyone interacts by the way last time i checked let's just take you know what everyone calls english or the united states of America or where you have English speaking countries.

34:21We have plenty of folks that live here that like, you know, weren't born here and speak many other languages, right? Or married significant others. And so they grew their family up, you know, speaking in a different language and different culture, even though they live in Texas. Right. And so you have to account for all of those things that were very early, right. And voice performing right at that level with all of these, you know, multi-dimensional factors, right? That even seemingly in the English language, you would think are easy. But when you sort of start to double click on all of this, it's a very hard problem.

34:53And there are people investing a lot of money and a lot of time and a lot of effort to go fix that. And so we have a long way to go, which is really exciting because at the end of that, you're going to see a better customer experience for all of us. You're going to see new paradigms exist. You're going to see, you know, the intersection between humans and robots and wearables actually manifest in a way in which it's a core and central part of our life. And you're going to see voice really, truly become that primary interface as it starts to really understand us and work with us as we expect it to.

35:24And we got a lot of work to go get there. So it's early days and early innings, but the promise is absolutely there and people are working really hard to get there. So, you know, over the coming years, you're going to continually be surprised month over month as we all are with everything that we see in AI. Yeah. Are other languages as far along as English or farther? Or I just think when I think of the planet and then the hundreds of languages or thousands even maybe that it encompasses, I'm curious, are we just lucky to be in the U.S. and at the forefront of this? Or are there other languages that are doing that are performing really strong as well already, too?

35:59Yeah, I mean, I think, look, I think the large majority of CapEx that's been deployed into AI systems holistically, like U.S. has the lion's share of that. So, you know, we certainly are doing a fantastic job there and are certainly ahead. I think when you start to think about other languages, like obviously if you go to India, right, and you think about the number of dialects, right, it's actually exponentially more complex, right? And then you have dialects speaking to other dialects, and then you have dialects and you have villages even inside of the dialects. So you really just have this complex set of dimensions in many other countries.

36:36And that's not to say India is not doing a good job here. They actually are. And they're investing heavily, and there's a lot of innovation happening there. My point there is there is a set of complexity across the world that many are dealing with and many are trying to tackle. And sort of holistically, there's just a long way to go. But U.S. obviously is doing a fantastic job here. But voice is still a multidimensional problem, as we had discussed, and you got to be able to measure it appropriately and solve for it in the right ways.

37:02Corey Noles:What do you think we're the most behind in voice? Like, is it long running duration? Is it the dialect problem you just talked about? Like, where would you want to see the most improvement in the near term? Like, where's the weakness, kind of? Yeah, yeah, from your perspective. Yeah, it's in the depth of the conversation, right? So, again, you know, one, two, three, five, seven turns, like, okay, fine. But like, you know, how many times have you talked to a customer service agent and looked up at your phone? You're like, oh, wow, I was just on with I don't want to pick on an industry for 32 minutes.

37:32Right. And whether they were helpful or they weren't helpful, it was a 32 minute phone call. Right. You know, we've got to look up your information.

37:39Corey Noles:It just goes on forever. You're on hold for 10 minutes. Read through a guide. Correct. Exactly. So you've got duplex models with tool calling and reasoning, right, is kind of where that technology stack is going to address those problems. And that's where people are working and investing really heavily right now. And we definitely have a long way to go because of a lot of the ways in which these systems fundamentally just have complexities. But that's what people are going to be solving. And I think it's going to be fantastic once you start to see that come to the forefront. And a lot of people are working on that in a very deep way right now.

38:14And so when you start to look at things like that, there's a lot of work to do, but it's going to be very powerful when it comes back. Absolutely. Absolutely. I've always felt like, not specifically, well, yeah, call center industry and places like that definitely have a lot of motivation to see a technology like this reach point where it is reliable and performable. Because like you said, there's this whole element of just calling customer service. It's inconvenient. You're on hold. You're waiting. They're literally looking through a script to figure out where your issue lines up. To be able to make that a five-minute issue would sure be a better experience, I would think.

39:00There is. And, Corey, I've done the research. there are tens of billions of hours of audio in that industry locked up that are sitting in storage bins, S3 or otherwise, that folks are fundamentally not unlocking human emotions, human understanding, and all the metadata associated with that as the real ground truth to go build their voice agent experience on top of. Have they not seen the billions of dollars that they could make if they licensed that data?

39:32Corey Noles:Why are they not making licensing deals? Like, what's wrong? Then there's that. Then there's that for sure. Like to the extent that they actually have the consent to go do that, of course. Yeah, sure, sure. Yeah, if they don't have the consent, yeah. But even then, right, the inverse is also true. Many are deploying voice agent solutions, right, without trying to mimic the human behavior and understanding that they have from being in business for 30 or 40 or 50 years and all the recordings that they have. And that just is unfair to us as consumers and as loyal customers of a lot of these companies.

40:07And that's what we believe is a massive opportunity to help in that space. And we're the only folks in the world that do that. And we're really excited with a lot of the work that we're doing in that space right now. It's really cool. You guys really saw an opportunity and took a risk. I mean, the fact of the matter is you had great models. as when we were testing them the thing that stood out to me the most about the human models was that i felt like like the human element was the most there of a lot of other ones i was listening to so so i uh i had i admire the the guts it takes to shift so fundamentally and and go this direction but it makes all of the sense in the world yeah well thank you uh we got a lot of work to do but we, to your earlier point, saw the fact that we truly believe voice will be the primary interface inside of AI and that in order for that to happen, it has to align with humans and that we just so happen to have a research infrastructure built that scale to go serve that and we thought that would be a much larger market opportunity particularly at the time in which the models were going to get commoditized over time.

41:17It meant more people were going to build more systems than we had discussed and that we were probably better situated in the infrastructure business than the raw model business. And so far it's played out really well and we're excited about the future. Yeah, because in the model business, you're very much more competition every day, a race to the bottom in price. There's a lot of reasons to shift. It makes perfect sense.

41:41Corey Noles:So you have this voice replication leaderboard, right? Is that correct? Yeah. I think that's really interesting. So a model there can sound extremely natural without actually sounding much like the person it was supposed to clone. So my question is, is the industry optimizing too much for like, oh, wow, that sounds real, instead of measuring whether these systems are actually doing the job to be done correctly. So does it actually sound like the person? Yeah. You hit the nail on the head, Grant, and fundamentally what we have built is this scenario builder, which is, hey, look, when you want to evaluate a model, you shouldn't worry and optimize for what's measurable, which is what everyone does today, i.e.

42:23the transcript. You should optimize for what's most important. And to your point, Grant, right, it's not most important that you just sound real, right? Like, do you sound real in this context? Can you do that across a multitude of questions? Can you do this in a multitude of environments of acoustic quality? Can you do this across languages? And you should be able to just define those scenarios. And we have an orchestration layer that says, got it. Thank you for defining that scenario. We are now going to programmatically give you the plan by which you can most effectively and efficiently get that measured, get the analytics back and understand where that was at.

43:01And that is exactly one example, Grant, and one leaderboard. And there's a multitude more that we will continue to publish and continue to publish more evals. And then all of these become composable templates and workflows in our library for people to be able to go run and be able to go measure and choose what models to bench against and choose what scenarios and what languages so that over time this becomes very easy and you're never going to be guessing how you're going to perform in a real-life production scenario. Makes sense.

43:29Corey Noles:AI safety, right, has been a big buzzword the past two weeks for all sorts of reasons, which we won't go into. Boy, it hasn't. Unless you want to, Andrew. I'll leave it up to you if you want to get into that. But I do wonder, do you have any benchmarks around voice AI safety? Like how do you think about that? Because like Voight-Navis is a whole can of worms, right? It does. It does. I'm not taking the bait, Grant. Good try. I'm not taking the bait. I know. I know. It might be good for ratings, depending on what I said. I don't know. Or distribution. But I'm not going to do that. Look, I think safety is really important, right?

44:05And that's fundamentally why you have to measure this against real-world scenarios with humans, right? And back to my silly, very simple example, I'm fine, is actually a safety hazard, right? If a healthcare voice agent thinks you're fine and doesn't help you and you actually really needed something and you actually were scared, you just didn't know, right? That's a big problem, right? And that could be a human safety issue, right, just on one end of the spectrum. And so we're trying to work with all of the right guardrails in place to be able to measure those effectively. We're actually also doing a lot of research right now into the safety aspect and partnering with a number of the world-leading organizations on this.

44:47So stay tuned for some more research to be published from us in that respect because we think it's very important and we'll continue to invest heavily in that area. Because now that you sound natural, now that you can do the right things, now that you can do it at the right time in the right place, we need to make sure, right, that you are who you said you are and you're performing in the way in which you are. And the right things aren't linking and all of those things. So that's why we're very excited about the position that we're in from an infrastructure layer and continually to add value for folks over time.

45:16Corey Noles:I was wondering about this. Is there like a synth ID equivalent for a voice? Like if someone was to impersonate, I don't know, to pick a person, Joe Biden, right, on the phone and say like, hey, I'm calling my bank. Can you wire me$5 ,000? This is Joe Biden. Is there any way for the industry to have any automated tools to say, oh, that's not him. Like we know that that's built into the model. Yeah, it's very challenging, but there are a lot of folks working on that. and it's an area of massive investment in the startup community and it's a really hard problem. A lot of really smart people working on it and I do expect many people to solve that at large scale and it's going to be a fantastic win for the voice community and I think you'll continually see a lot of advancement come out of that.

46:04So absolutely, yes, it's well on the path. I think many people will tell you they've already solved it and many might have. I'm not saying they haven't, right? But you're going to continue to see additional innovation there. Nice. Awesome. Yeah. Nice. Something I'd like to touch back on is we talked a little bit about voice as an interface. And that makes me wonder, do you see voice as a model in the traditional sense or as more of an orchestrator of other models, like maybe a large model, maybe specialty models or various tools? Like, I keep thinking of this as voice, as a gateway to real agentic work pretty quickly.

46:48It's all the above, Corey. And I think let's take the call center example, right, where you would have a duplex model with a whole bunch of reasoning, right, which basically means there's two of us having a large set of conversations. And action has to happen. So I have to go look at the billing system. The billing system gives me one thing. Unfortunately, the systems in a lot of these companies aren't architected perfectly. so they might have to go to a different billing system for the second piece because you came as a customer through an acquisition right uh or that product is in another system and then you might actually need to understand where the field technician was which is in your you know uh maintenance repair and operation system which might be a different vendor and you might be migrating all of those things right so now if you just think about that and i'm doing this from voice right i'm having to have the right set of conversations in the right way now you have a bunch of tools that are being called, right?

47:40And there's going to be reasoning and chain of thought reasoning to go through a multitude of those where you're going to have LLMs in the middle of all that. You're going to have perhaps other agents that are going to do other things, stitch all that together to deliver it back so that they could basically just tell you that a technician will be there tomorrow to go fix your issue. And they're going to credit you for your service that was down yesterday. Right? And that was a lot, but you just want that answer. Like, when will my cable be back up? Which I know is not life or death, but like, When you start to think about it and that and blow that example out, right, to your point, Corey, you're going to have a multitude of complexities that go into these systems at production scale to be able to account for this.

48:18And absolutely, you're going to have large models. You're going to have LLMs in the middle of that. You're going to have the internet in the middle of that, right? You're going to have buffering. You're going to have all of these other factors that go into quality and packet loss and background noise and speaker identification. And when I stop talking, am I actually done or am I confused or is this like a healthy pause? Right. Am I thinking? Right. Correct. Correct. Correct. And again, it's because this is multidimensional. It's a lot easier in text, which I'm not suggesting there aren't really hard problems in text.

48:51There are. Right. As evidenced by all the folks that are doing great work in coding and the legal tech startups that are doing a fantastic job and all of these other things. But like when you start to look at that domain specificity, right, that folks are solving in text and you start to apply that to voice, it's just, you know, a lot more variables are in play there, right?

49:09Corey Noles:When I think of voice as an interface, there's like multiple layers of interaction, too. Like just if you think about it in terms of let's say I had a perfect system. At some level, that perfect system, it would be me talking to my computer and I don't have to open another application. It's like my computer just has some sort of voice model built in. But then every other application that I use would also have some sort of voice interface. So how am I now navigating the voice on my computer, the voice, like to your point, like Google Chrome or whatever I'm using to access the internet, and then the voice model on the website that I'm actually going to?

49:47And your phone, your earbuds, your hardware as well. Right.

49:50Corey Noles:It's actually like a Russian nesting doll of complexity, the more that you zoom out on it. It is. And say what you want about the algorithms that target you on certain social media platforms that display weirdly relevant, right, sets of information for you to take actions that you weren't thinking about. But at the same time, right, there is a level of personalization that we do enjoy, right? Whether it's ad targeting, whether it's searching, whether it's context that remembers our last conversation and where we picked off, whether it are the documents that we uploaded to have the context for how my business runs, right?

50:28And all of those things. And so the summary here is that folks want the voice AI systems to be very well aligned to that same level of personalization. And again, I just keep coming back to everything that we spoke about and how difficult that is in voice and how we have to have the right level of systems and infrastructure to support the testing, evaluating and improving of those systems. And that's where we're focused because personalization has to be at the core and the forefront of how voice AI systems work because that's what we're used to in the rest of our life. And that's where other systems are advancing in dramatic fashion.

51:00Let's say everything you're working on is right and you're hitting on all cylinders. Where do you feel like the next two or three years take voice tech? Again, the primary interface in voice AI where that just becomes true, right? Where a large portion of what we do is with voice and we feel comfortable, we feel confident, it feels natural. And we feel like we actually have the personalization to deliver us the speed and the convenience and the level of accuracy by which we've come to enjoy text-based LLMs for a variety of things that we do today, whether it's planning a trip, whether it's running a simple web search, whether it's trying to find a document, whether it's writing an email or a multitude of ways in which we use all that today.

51:44You would have the same way and the same thing happening with voice. Look, you saw some of the large coding agents deliver voice mode and coding agents and a multitude of other things. And so that is absolutely where you're going to see it. And I spoke a little bit earlier about these duplex models with reasoning, and that's where you're going to start to see these things deliver far more complex, personalized interactions for you that are going to become so seamless. It's going to be that aha moment, like the first time you went to chat GPT and you're like, oh, my God, this is incredible. Right.

52:19You know, you're going to have those same type of feelings across, you know, a large multitude of voice systems very shortly. Man, and in a coding agent, it's something special. The exact same thing.

52:30Corey Noles:Yeah. Like I told Corey, like this when GPT-6 came out, it was the first time I really gave voice mode in Codex like a legit shot. And just the fact that you could be like, yeah, can you go add this feature for me? and then you're, you know, let's look this up on the web. And then, oh, yeah, can we start up this new Greenfield project that we're going to make from scratch? Like, it's just one chat the whole time. I'm sitting on my couch talking to my computer over here while I'm watching television. Yeah. And you've got it managing multiple different tasks at once. Like, I ran into a RAM problem.

53:04I had a RAM issue because I didn't have enough machine that I was trying to do for the number of projects. And but but what a what a way to deal with that. And it's so good at like, hey, check on all my projects. How are these going? And just coming back with a detailed report from all of those. And it's like, man, this is people don't know what they don't know right now. But that's the power. I code every single day. I'm not a coder. I've grown up more of the business operations. But yet I can't call a customer service desk. Right. Of my cable company. Right. And find out right when service will be restored.

53:42Right. Using the voice agent and any level of reliability. Right. But yet I'm rather sophisticated in how I think about these AI systems. Right. Yet I'm very unsophisticated when it becomes decoding. But I could just type and actually get a website spun up right in a matter of minutes. It's a big disconnect. What's the barrier there?

54:02Corey Noles:Why haven't more companies adopted this, do you think? Is it because they don't have the benchmarks like you've been doing? Is it that it's just they don't know they can do it? Where do you think is the barrier? Well, the benchmarks just allow you to keep score, but it's underneath of the benchmarks that is the research and the process, the optimization and the workflows that have been codified into a platform. And this decade worth of research at very large scale that allows all of this to happen. And again, fundamentally, right, you know, voice is multidimensional, as we had discussed, and text is flat.

54:38And so it just is far more complicated. There's a lot more factors that have to come into play. And that's why we believe an infrastructure designed to help improve voice AI at global scale is something that really was necessary in the world. And, Corey, to your point, that's why we made the pivot and are exclusively focused there and really excited about the future. Love it.

54:59Corey Noles:Andrew. One last question. Oh, go ahead. One last question. It's a fun one. We can keep it short. Do you think that – so, like, you know, we talked about talking to voice interfaces a lot here. On a personal level, do you think that at some point we're going to have to change the social contract a little bit to be more comfortable with people talking to themselves? How do you feel about this personally? Or do you want to know this? You know, Grant, it is the social contract. I hope you didn't trademark that because I might start using it. That's a great one. Go for it. No, I didn't come up with that.

55:31Yeah, look, our social contracts actually will have to change a little bit, right? But the job of the AI systems and the ones that will win are the ones that don't force us to do that. And that's the whole point of being natural and consistent and reliable and acting and aligning like we are is you should just be able to talk freely and not have to change your patterns in order for the AI systems to work. Like, sure, I can optimize and write a better prompt and follow the right skills and tell it to think like a McKinsey consultant and all of those things that you do where you don't get better results to go do that.

56:03But in voice, you should just be able to speak very naturally and very freely. But I think the social contract that will exist is because you're actually now going to be able to do far more things in a multitude of ways. And you might have to start to think a little bit differently, just like we had to think differently with how all of these things work in the LLM world. Right. And so I think that's where the real promise is going to come into play, because, look, you know, when you go to Thanksgiving, right, in six or eight weeks or whatever it is, And I'm pretty sure people are going to have, you know, their aunts or uncles or grandparents that are there are going to use the word chat.

56:38Right. Oh, I was using chat the other day and X, Y, Z. It's how I made this apple pie. Right. I mean, just think about that for a second. I have an 80 year old that just use chat to update like their pie recipe that's been in the family for generations. Right. You know, like that's wild. Like that's a new social contract. Right. of basically how I would go interact and kind of make something that I've made for 30 or 40 years. Right. It's kind of wild. Right. And so fundamentally, you will see the same thing in voice. It's a hell of a chef, by the way. I'll be honest. It is. It is. It is. I use it a lot in the kitchen and I am constantly amazed with with how well it knows for me to reheat a steak or something.

57:22Yeah. The problem is execution. Execution is tough. At least if you're challenged like I am in the kitchen. Yeah, I think we can all relate to that. I definitely relate to that. Andrew, thank you so much for joining us, man. It's been a lot of fun. Yeah, it's been great. Thanks for having me, Corey. Thanks for having me, Grant. It was a lot of fun. And keep up the great work. Huge fan of all the work you guys do. Thank you. Where can people go to keep up with Hume and what you all are working on and learning along the way? Hume.ai. Feel free to reach out anytime or feel free to reach out direct to me on LinkedIn.

57:56Very active there and happy to chat with anyone. Excellent. Well, I'd like to give one more quick shout out to our sponsors at SAS for making today's video possible. If you like today's video, please take just a moment to like and subscribe. And don't forget to check out our daily newsletter read by more than 700 ,000 people just like you over at the Neuron.ai. And that's all we have for today, everyone. So farewell for now, humans. Meow.

From the publisher

AI voices can already sound startlingly human. The harder problem is whether they actually understand how humans speak.


Andrew Ettinger, CEO of Hume AI, joins Corey Noles and Grant Harvey to explain why voice AI still has a “listening problem.” A transcript can capture the words someone says while stripping away tone, hesitation, frustration, background noise, dialect, speaker identity, and the other signals that shape what those words actually mean.


The conversation explores how Hume evaluates voice systems against real-world conditions, why emotion is better understood as a high-dimensional spectrum than a handful of neat labels, what breaks during long-running voice interactions, how speech-to-speech systems and AI voice agents could become a primary interface for computing, and what new safety and personalization challenges appear when AI starts talking back.


If you want to understand what has to improve before voice AI feels truly natural—not just realistic—this episode is a practical look at the infrastructure, evaluation, and human feedback behind the next generation of spoken AI.


Subscribe to The Neuron at https://theneuron.ai for more AI news and analysis.


Sponsored by SAS: Grow faster and scale confidently with trusted AI solutions. Visit www.sas.com/smb


Sponsored by Origin: Origin is endpoint AI observability. See what a trace looks like at https://www.originhq.com/aiexplained


Sponsored by Outshift: “Scaling Out Superintelligence” Vijoy Pandey, January 2026. The technical whitepaper detailing the Internet of Cognition architecture, three-layer stack, and cognition state protocols. https://outshift.cisco.com/internet-of-cognition/whitepaper?utm_campaign=fy27q1_outshift_ww_paid_ioc-neuron-wp_podcast&utm_channel=podcast&utm_source=podcastInternet of Cognition Interactive Demo Clickable walkthrough showing per-agent activity, intent, context, and collective reasoning across a multi-agent SRE system.https://outshift.cisco.com/internet-of-cognition/explore?utm_campaign=fy27q1_outshift_ww_paid_ioc-neuron-explore_podcast&utm_channel=podcast&utm_source=podcast

More from The Neuron: AI Explained

All 106 episodes
How Hume AI Is Teaching Voice Systems to Understand More Than WordsThe Neuron: AI Explained · 59 min
Listen in VO