In short
Voice AI’s “uncanny valley” problem (latency, interruptions, wrong tone/emotion) and why moving from text/voice to audiovisual avatars requires new real-time engineering, audiovisual tokenization, and human-experience benchmarks. Smola argues voice is an intermediate step toward AV agents that look/feel human, but natural interaction is still far away.
Guest
Alexander (Alex) Smola, co-founder and CEO of Boson AI; professor at Carnegie Mellon University.
Key claims
- Natural conversation depends on interruptibility around ~150 ms (linked to human audiovisual perception).
- Single-model end-to-end audio becomes unaffordable if made “smart”; Boson uses a hierarchical/agentic setup with background tool calls and “thinking” turned off to avoid awkward pauses.
- Audio understanding (reasoning over audio + prompts) differs from ASR (audio-to-text without deep semantic grounding).
- Boson prioritizes affordability: better benchmark results at ~1/10 cost (with thinking off).
Notable examples
- Dinner at KDD conference in a noisy bar: system handled Slovenian and Hindi switches; iOS noise cancellation helped.
- Apple boot screen as an example of “fake progress” that feels natural.
- Hitchhiker’s Guide computer cheerfully announcing imminent death as an EQ/voice mismatch example.
- “Foot rule”/trim-mean analogy for learning from noisy internet labels at scale.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe State of Voice AI
1:24 to 2:18
Discussion on the limitations and advancements of voice AI technology.
“For over a decade, I've been exploring the ideas and innovations shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters.”
Real-world Challenges with Voice AI
2:18 to 3:36
Sharing experiences and challenges in using voice AI in noisy environments and different languages.
“But in general, I tend to think that, you know, they've gone through several iterations of it.”
The Importance of Advanced Microphone Systems
3:36 to 4:12
Exploration of the role of microphone arrays in enhancing voice AI performance.
“I would say a non-trivial amount of the credit definitely goes to the iOS team doing really good noise cancellation and audio separation on the mobile phone.”
Future of Video Avatars in AI
4:12 to 5:29
Discussion on the timeline and challenges of developing animated video avatars for AI interaction.
“I mean, well, I've got a huge microphone array sitting in my in my laptop here.”
Engineering Challenges in Voice AI
5:29 to 10:48
Deep dive into the engineering and research challenges faced when developing voice AI systems.
“I think probably in a year, audio will become pretty much bulletproof.”
The Economics of AI Model Development
10:48 to 14:00
Insight into the cost considerations and development strategies for AI models.
“And then, of course, once you have all of that running, you need to also take care of the engineering implementations.”
Evolving from Text to Voice Models
14:00 to 20:11
Explore the transition from text-based AI to voice models and the steps taken.
“Can we take a step back and maybe have you talk a little bit about your kind of arc or trajectory or path?”
Challenges in Audio Data Processing
20:11 to 22:20
Discuss the challenges faced in processing large audio datasets.
“So I think one of the things that are a meaningful differentiator is that we can process and have a lot of audio data.”
Sourcing and Annotating Audio Data
22:20 to 24:26
Learn about data sourcing and the complexities of audio annotation.
“Is it kind of internet style videos and maybe you extract audio from video, that kind of thing?”
Leveraging Noisy Data for Machine Learning
24:26 to 28:00
Understand how to extract valuable insights from noisy datasets.
“First of all, I mean, we know that it's possible to extract meaningful information, even from noisy data.”
Show all 24 chapters
Challenges of Training AI Models
28:00 to 29:33
Discusses the complexities of training AI models and the necessary resources involved.
“and suddenly you get a much harder data set where you have effectively labels that are with very high likelihood, very good.”
Model Architecture and Economic Decisions
29:33 to 31:39
Explores different model architectures and the economic factors affecting their development.
“So this basically, whatever you can get on public models is just a good prior, and then you optimize from there.”
Maintaining Conversation and Intelligence in AI
31:39 to 33:38
Examines how AI can maintain conversation while processing other tasks, similar to human behavior.
“So what goes on there may be, you know, quite variable, but humans are pretty good at that.”
Hierarchical Models in AI Systems
33:38 to 36:48
Discusses the benefits of hierarchical models and their role in creating fluid AI interactions.
“The problem is if you were to try and make them really smart, that would become unaffordable in terms of compute costs.”
Cost Efficiency in AI Development
36:48 to 39:48
Highlights the importance of cost efficiency in building AI models and the engineering behind it.
“So right now, on benchmarks, we are better than, let's say, GPT and Gemini and CROC at a fraction of the cost.”
Utilizing Off-the-Shelf Servers for AI Solutions
39:48 to 42:00
Explores the use of off-the-shelf servers in AI development and their impact on performance.
“But the quantity of servers enabled at the same time right now is a little bit limited.”
Building Efficient AI Systems
42:00 to 43:00
Learn how to create fluid AI applications without reinventing the wheel.
“Like I'm imagining the way you would build a weather MCP server is a lot more efficient than, you know, in an ideal world.”
User Tolerance for Delays
43:00 to 45:45
Understand how humans perceive and tolerate delays in technology.
“And it would feel totally normal for me, right?”
Importance of User-Centric Design
45:45 to 47:39
Explore how user happiness is crucial in AI voice interactions.
“Like, you know, is the model properly interruptible?”
Benchmarking User Interaction
47:39 to 48:36
Discuss the benchmarks for evaluating voice model interactions.
“that's, you know, what AI for humans really means to optimize in a way that it's pleasant for humans.”
Exploring Emotional Intelligence in AI
48:36 to 53:06
Delve into how emotional intelligence impacts AI interactions.
“I mean, we do use third party audio models in the defined benchmark.”
Personalization and Recursive Self-Improvement
53:06 to 56:07
Learn about continuous learning and personalization in AI models.
“But, you know, this is, I think, a really exciting time, at least not unless you do this appropriately.”
Personalization in AI Interactions
56:07 to 1:01:55
Explore how AI can personalize interactions based on user preferences and cultural nuances.
“But, okay, so there's the, you know, given that those models by now have a decent idea of how humans work, you can use that for, you know, RSI to some extent.”
Learning from Human Interaction
1:01:55 to 1:04:04
Understand the importance of learning from interactions to improve AI's handling of human behavior.
“Well, there's a bit of a paradox in that if most humans don't have them, they're not in the data, and therefore we won't be able to easily train models to follow those patterns.”
Transcript
Automatic transcript. May contain errors.0:00Sam Charrington:Voice AI has gotten very good over the past few years, but it still suffers from a bit of an uncanny valley problem. Humans are remarkably sensitive to the subtle things that make a conversation feel natural or not. A little too much latency, an interruption handled awkwardly, the wrong tone or emotional response, and suddenly the illusion breaks. And the bar only gets higher as AI systems evolve to incorporate more of our senses, Moving beyond text and voice to systems that can see and be seen introduces a whole new set of challenges around natural interaction. My guest today is Alex Mola, co-founder and CEO of Boson AI and a professor at Carnegie Mellon University.
0:41Sam Charrington:Our conversation explores what it takes to build these systems, from the models and inference infrastructure needed to make voice work in real time, to emotional intelligence, and ultimately, audiovisual avatars. Here's Alex on where he sees all of this heading. I would argue that voice is an intermediate stepping stone. You might think, well, you know, what's next? It's clear that eventually we will have avatars. And so I think what this is converging to is that you'll be talking to an AV agent that looks and feels like a human. And we're still very, very far away from making this really natural.
1:21So it's quite exciting, actually.
1:24Sam Charrington:I'm Sam Sherrington, and this is the TwiML AI Podcast. For over a decade, I've been exploring the ideas and innovations shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in.
1:47Sam Charrington:When I think about kind of the frontier of voice AI that most folks probably have access to, I'm imagining it's something like chat GPT advanced voice mode. And actually, I want you to react to this. Do you think that actually that's, you know, crap and there are much better systems and you should point me to, you know, X, Y, and Z? Or are there, you know, are there better systems behind closed doors and labs and enterprises or what? But in general, I tend to think that, you know, they've gone through several iterations of it. It does keep getting better, but it's still very infuriating. And in my experience, it really only works in perfect conditions, meaning a silent room, no background noise, being very conscientious about interrupting.
2:35Sam Charrington:If you interrupt, either you have to be committed to plowing forward or you have to stop and let the thing, you know, catch up to you. You can't like, you know, try to have a natural engagement. It's very brittle, I think. So you can definitely do better than that situation. And so fun situation from last week. So I was in Jeju in South Korea for the KDD conference. And we were out at dinner for drinks in a bar. And basically I was showing off our system and, you know, just, you know, me foolishly saying, hey, watch this. Let's see what happens. And so what I can confirm is that our model was able to handle people switching to Slovenian and then somebody else to Hindi fairly well, even though the bar was very noisy.
3:36I would say a non-trivial amount of the credit definitely goes to the iOS team doing really good noise cancellation and audio separation on the mobile phone. So I don't think that this would have happened just by waving a microphone somewhere. But I think we're on our way. And so basically, you know, having a single microphone will probably never really solve this. You need microphone arrays to really do proper noise cancellation, especially if you have many sources. But that's a solved problem, right?
4:16Sam Charrington:I mean, those things ship. I mean, well, I've got a huge microphone array sitting in my in my laptop here. I don't know if Codex was using it, but like I've also had like kind of experiences with the relatively new Codex voice. It just that that interaction just wasn't very fluid. Yeah. Okay. So that's I mean, sometimes voice input is good too. If you if you're feeling too lazy to type, I find typing way more precise for technical work. But for, so as in, you know, hey, let's configure our storage array. Well, no, I don't want to talk to you. I want to type. Because here's a very specific serial number and a specific URL, and I want to get this exactly right.
5:04So that's where I still prefer text, but maybe this is old school. But for, you know, if you're out in a bar and you're asking for information or advice or something, I think we're getting there. I will give it another year, and this will actually become quite ubiquitous, and I'm pretty happy where we are at with our models. I think probably in a year, audio will become pretty much bulletproof. Video feeds, as in avatars and so on, are going to start making some appearances. probably year, year and a half. We'll see that quite widely deployed on robots, basically with a face fully animated. So we're working actually with the startup on some of those problems related to that.
6:00It's a very talented team. So, and there, I think the lead times are a little bit longer because if you want to do this in the animated head or whatever, you actually need to make hardware, and hardware is hard. Right? Software is easy comparatively.
6:22Sam Charrington:When you think about the challenge of voice AI, how do you break it up in terms of primarily engineering problems, primarily research problems? If you try that, it's not going to go well. You need to do both at the same time. So let me give you a simple example. So let's start with the very simple basics. Namely, it takes about 150 milliseconds for basically a photon hitting your retina to your cortex actually doing something with it. and that number is reasonably stable that you can actually use that as a non-invasive diagnostic to find out whether you have a neurodegenerative disease because in that case the signal takes many detours and it takes longer so what people will do is they will flash a checkerboard pattern in front of you and then wait how long it takes for that stimulus to hit your cortex from the ears it's actually a little bit shorter because it's a well half the way right now Now, what that means is that basically humans kind of operate at around maybe 6 to 10 hertz, really, when it comes to audiovisual perception.
7:41I mean, of course, we feel like the world is very fluid. And I mean, we have our fancy 30 or 60 hertz displays, but the brain kind of, you know, within that time frame, it's kind of okay. So what we did correspondingly is make sure that our model is interruptible within about a similar time frame. So it does interruption handling within about 150-ish milliseconds or so. Now, the next thing is you need at least the current paradigm of how to deal with audio and other things is at some point you convert everything into tokens. and then those tokens will go to whatever LLM style backbone that you have.
8:31I mean, the architectures may differ and maybe at some point we'll get to diffusion-based models. I would think it's probably going to be more like speculative decoding with diffusion and the more traditional sequential models for just overall the statistical modeling. But, you know, that's a completely different argument and Stefano and one may have very different opinions on that. But basically, you know, you need to turn the audio into tokens. Then your model does something and then you need to turn those tokens back into audio. Now, you can get very high fidelity by having a very high token frequency.
9:19The problem is a high token frequency means that you need to ingest, thus pre-fill, and also generate many, many tokens per second. And we all know that this is expensive. So now you have this rather unpleasant dilemma where many tokens per second means your model cannot have too many parameters. Whereas if you have a smaller number of tokens per second, you can afford more parameters. text is the ultimate compressed format in that sense it's about 3 to 5 tokens per second that humans want whereas for audio you can easily have 10 plus and tokens per second also means the temporal granularity, let's say I have 10 tokens per second, then that means that each token covers about 100 milliseconds so this is why you actually get biology user experience, engineering, because you need to make this cost effective, and then science, namely, how do we actually represent this?
10:21How do we send it into a model, maybe doing something else more cleverly in the entire pipeline? So this is why all those things really need to come together. It may not necessarily be one engineer doing everything. That would be pretty amazing if you could find somebody like that. I mean, there are very few people but that's okay. But, you know, it's a whole systems challenge. And then, of course, once you have all of that running, you need to also take care of the engineering implementations. You probably need to buffer a little bit. So remember when I mentioned you want to be interruptible, but you probably want to have longer buffers.
11:04So now we're talking about essentially, you know, real-time type AV streaming, right? You also have a video feed. The video feed may come in at a different frame rate. So, yeah, basically, it's a really nice range of problems there. And for the video feed, I mean, we all know if you use WAN or some other models or Flux, they will happily produce, you know, video segments of 10-ish seconds. But if you want to have a continuous feed that, you know, is visually consistent for an hour, you need to modify those models a little bit and if you then want to make those models effective such that you can have multiple real-time factors there is yet another design optimization to be made.
11:57The TLDR is if you are doing video feeds for avatars you know that it's a video feed for an avatar so there are not going to be race cars driving in the background. In other words, most of the video feed is pretty boring. So for instance, you could easily compress the AV for this interview into a fairly effective stream. Of course, if I start moving my hands like crazy, then the bit rate will immediately go up by a lot. But most humans don't do weird things like this. And so it's a perfectly acceptable feed quality. And again, there's a trade-off between most beautiful quality and building something that actually people can afford.
12:46And that's probably also the other point where we are maybe taking a slightly different operating point from some of the, well, very large trophy models that are, you know, trillion parameters just to do chit-chat.
13:01Sam Charrington:And which point in particular, the affordability or something about the trade-off or which? So the affordability, right? So basically, you can always make your model smarter by making it bigger. Absent of the real-time criteria that you mentioned with regards to voice, there are some laws of physics that come into play here. Yeah. And so the problem is basically, you know, how much compute do you need to stream, you know, a conversation? And if you need to use, let's say, you know, a full Blackwell server GPU just for a single conversation, then that may not be the most economically viable model.
13:48I mean, this produces gorgeous demos, but your customers can't afford it. And that's, I think, where we went in with a price first and then worked backwards to, you know, how can you build something that actually people can afford?
14:05Sam Charrington:Can we take a step back and maybe have you talk a little bit about your kind of arc or trajectory or path? Like you started the company in 2023. You were not initially focused on voice. You eventually shifted direction to voice. When you started really focusing on voice, kind of where did you start and what were the steps you took to kind of evolve to where you are today? So we started off with text, just like, I guess, others as well, maybe with a slightly stronger focus on AI for humans. I mean, that's been with us since day one. And as mentioned, one of the key issues was that, you know, the text interface felt always a little bit awkward.
14:56And so we then started looking at, okay, what are good audio models? we also realized that probably building your own LLM was not the smartest idea if you just wanted to have really good audio but instead can you use the intelligence that's already baked into a high quality LLM and then make sure that it acquires effectively one extra modality namely in this case audio. Again, other people had done similar things, for instance, for video. So for instance, there are vision at least. So there are BLM, so vision, large language models. People are doing similar things for robotics and world models.
15:41So the overall pattern of using the intelligence that you kind of get for free from reasoning over large amounts of text, you then combine that with audio. Now, one of the problems is if you teach this model a new modality and you're not careful, it forgets everything that it knew before. Probably the easiest way to imagine that is if you, and I've actually seen that with a friend, so she adopted a kid from Latin America, so this was in Germany, and she spoke only German to her, and within a matter of months, the girl had forgotten every single word of Spanish, right? And a similar thing effectively happens if you take an LLM and you only get it to work with audio, then it will very quickly forget about all the reasoning and language itself.
16:35So you need to still maintain a somewhat competent mid and post training LLM pipeline while also having the same capabilities for audio. And you need to then also define tasks that nicely marry audio and text, so to ground things into each other. I mean, just like if you have a multilingual LLM, at some point you need to make sure that, you know, dog means xia or kane or hund. And if you don't have that, then it becomes a little bit tricky for the model to reason across languages and to get that strong generalization. but anyway so we did this we then you know released I think a fairly competent TTS model that was Hicks Audio V2 last year and I think we've significantly accelerated our release pipeline this year so we've been putting out TTS and audio understanding in ASR models And so in case you wonder what's the difference between, so TTS means text-to-speech, so basically, you know, text in, sound out, voice cloning, all of that.
17:54But what's the difference between audio understanding and ASR? So speech recognition. Speech recognition, it's audio in, text out. But the model doesn't really understand very much of what's going on. I mean, it understands a little bit, but these are fairly lightweight models, but they will not be able to resolve whether to recognize speech means to recognize speech or whether it means to wreck a nice beach, right? They both sound the same. And obviously, one is nonsense unless you are in some environmental conservation event where they probably mean the latter. but you don't know the ASR isn't going to be able to handle that but an audio understanding model that can reason over the audio plus maybe a text prompt plus maybe other audio references can.
18:58So these are then models that are much more similar to your favorite LLM just that they can now take different modalities as input and you could then also consider having video as another input in addition to audio and text. And you could have, you know, role model parameters or your robot or your self-driving car and other things also as inputs and then correspondingly all of that coming back out. So if you think about it, it's basically the tokens are like your bus on the back where all the information gets sent through and then comes back out again. so it's this is really the glue that ties everything together so we started releasing those models and I think by now we have something that has very good latency then we started having to really do performance tuning to make those models really interactive that's again extra work and I think by now we're in a in a decent position.
20:08So pretty happy and proud about what the team built. And so along that path, what was the, when you think about kind of the significant technical
20:17Sam Charrington:challenges that you ran into that, you know, the team really had to, you know, go heads down and you think came up with a clever solution, like, you know, talk about some of those key technical challenges. So I think one of the things that are a meaningful differentiator is that we can process and have a lot of audio data. And that's a meaningful mode. Meaning your training data set that you've collected or throughput or something else? No, it's the training data set, really. So we have in the order of 100 million hours of audio. And it's about 200 human lifetimes. If you live for 75 years in a very noisy environment, you would get about that amount of audio, not necessarily spoken, but that amount of audio.
21:19But, you know, you can actually, you know, go forth and scrape and crawl a fair amount of that data online. But then you need to process it. You need to extract, tag, normalize, transcribe. there's a lot of extra processing that's needed and having our own data center really helped us there. If you were to store this data on a new cloud, the storage bill would eat you alive. And yeah, if you were on one of the big three cloud providers, it would become even more problematic. so at least until recently when hard drives started becoming really expensive again this was a very nice situation i mean by now the hard drives are about three times the price of what they were a year ago so okay we'll have to be a little bit more prudent now and so talk a little
22:20Sam Charrington:more about this dataset? How is it sourced? Is it kind of internet style videos and maybe you extract audio from video, that kind of thing? This and many other things. The one thing we didn't do is we did not spend unreasonable amounts of money on an annotation company. So every once in a while I get emails from a company saying, hey, we can annotate 10 ,000 hours of audio for you. And at that point, I'm like, okay, that's good for you. Right? We are at between 10 ,000 times that scale. And so is that because the existing technology, like the existing tools are sufficient enough that you can transcribe?
23:18Sam Charrington:There was a lot of engineering. there's a lot of engineering that went into this. That's, I think, part of, I think, what's a meaningful asset of what we have. So my apologies for maybe being a little bit vague here. But yeah, that's where a lot of work went into. And that's... But speaking in generalities, you start with... Is more data is more better. And yes, it's basically, let's put it this way, it's sourced on the internet. And, you know, there are different types of data with different types of metadata that you can get. Then you need to be a good engineer and recognize what you can find.
24:01Sam Charrington:It prompts for me a question about the existing tools we have aren't perfect. So if you're starting with Internet data and processing those with imperfect tools, you have a noisy, a fairly noisy label set. and you're using that for training, like, you know, is that better? Is it worse? Has it created its own unique challenges? Like talk about the relationship between, you know, that. Okay, so there's a couple of things. First of all, I mean, we know that it's possible to extract meaningful information, even from noisy data. I mean, the simplest analogy is, let's say you have a really awful voltmeter and you want to find out what the voltage in your outlets in the houses, so you go around and plug it in many times, you get the numbers out and you average in the end and in the end you get something that's better than what an individual measurement of the voltmeter will do.
24:58And that, of course, only works if your voltmeter is unbiased. If it has bias, then that procedure is useless. And that procedure has been around for ages. in the Middle Ages there's something called the foot rule where you estimate the foot by just having people walk out of church and you grab the first 12 men you send the two with the shortest and two with the longest feet away probably due to deformity, take the other eight average their length and you've got a pretty good estimate of one foot Right. Now, okay, sorry for maybe giving a very crude statistics explanation, but that's literally where the footroll and trim mean estimators come from, right?
25:45So robust regression and estimation is half a millennium old.
25:54Now, with audio, right, I mean, you can, first of all, It's not unreasonable to assume that by having a lot of data, you can estimate a better model. It's also not unreasonable to assume that once you have a better model, you can get better annotation. So basically all of the learning from weak supervision, all those ideas are applicable. In some cases, you also have context and the context helps you more. And so with that context or privilege or side information, again, you can do a little bit better. So for instance, knowing that this podcast is between two people and maybe I can even see that in the annotation or whatever or having seen that other podcasts on TwiML are usually Sam and one other person.
26:55so then you can go and use that to figure out if there are only two people. It's very easy to separate those speakers and then you get large amounts of single speaker audio. It's also reasonably easy to tell, I guess, the two of our voices apart, which again helps you to then annotate a little bit better. And so basically having a lot of this stuff allows you to build better models and then you get that flywheel. I mean, if you think about it, people have used similar tricks for face recognition where usually matched impaired faces are not that easy to come by. But for instance, if you have a movie, you know that the actor may look very differently throughout the movie, but it's still the same actor and there are not that many actors.
27:51So what you can do is it's not that hard to identify the same actor within the movie, but then you can go and shuffle all the actors together and suddenly you get a much harder data set where you have effectively labels that are with very high likelihood, very good. So this is a lot of classical machine learning and statistics, and of course it will help. So you do the right thing and it's work, but it helps. I mean, this is about, I think, as specific as I can be, but I think most statisticians can do that, but you still need to do it at scale. And so, you know, you need to invest compute at scale.
28:42You need to invest into storage at scale for that. So a non-trivial amount of our resources was invested in data, not just training.
28:53Sam Charrington:With regards to training, you mentioned earlier on that it was clear that you didn't want to build an LLM. Does that mean that your models are primarily, you know, fine-tuned or somehow based on other LLMs? Or how do you describe the model approach that you took or training, you know, approach that you took? It would be very nice if we could just go and take a model and fine-tune it. But, you know, modifying those models is fairly major modification. so there's a full pre, mid and post training stage like in an LLM and so what goes in
29:47is substantially different from what you get in the end so they may share a non-trivial part of the same architecture but there's a lot of training that happens even on the LLM side again in order to make sure that the models retain their ability. Right. So this basically, whatever you can get on public models is just a good prior, and then you optimize from there.
Read the full transcript
30:15Sam Charrington:Have you published or said like what your base model or models are, or is that proprietary? It depends. So in some cases, it's proprietary. So for instance, one model that we released last year was built on top of LAMA and we acknowledged them appropriately in our model card.
30:38So basically some of those design decisions are a little bit variable for, you know, whether we just build an audio app model or an ASR model or whether it's an audio understanding model. So you do need the full pipeline, but we can either, you know, build and train on top of an existing model or we can build our own. But that's then often an economic decision and also, you know, which features you need. So, for instance, if you need particular performance in a particular language, then we may very well decide to invest also significantly more still on the language ability there. So, it really depends on the use case.
31:29But, you know, obviously, you know, if somebody gives you free steel, you don't build a steel mill, you build a car, right? and you know as long as you know those models are available it makes sense to take advantage of it it would be economically foolish not to but I would say the work that's required in building the audio models is somewhat commensurate with the language model itself maybe it's half or one third but it's not one tenth so the other thing is for instance if you want to have models that are able to then have a high amount of intelligence in the background you need to specifically make sure that they can do this while they're still maintaining a conversation in the foreground and I mean humans are pretty good at that so for instance you may idly chat with somebody and in the back of your head you're thinking about something else like well did I add some coins to the parking meter or I really need to go or you're actually trying to solve a difficult technical problem so for instance in an exam you may be stalling for time while in the back of your head you're feverishly thinking about the solution, right?
33:08So what goes on there may be, you know, quite variable, but humans are pretty good at that. Machines are still improving, and that's actually also one of the really exciting frontiers there. How do you design architectures?
33:26Sam Charrington:Talk a little bit more about that. But that's a great segue to the question that I wanted to ask, which is really around how you do that. As you were talking about the challenges of audio AI earlier, one of the questions that arose for me is, you know, are we able to do this all with a single model? does getting into hierarchical models you know your system one system two kind of bolted together to create a fluid experience does that ever make sense like how have you come to think about you know overall architecture for these types of systems I mean there are some systems that are just you know end to end audio and you know they have their place for something that's you know super responsive fairly small fairly low latency, but also fairly dumb.
34:25The problem is if you were to try and make them really smart, that would become unaffordable in terms of compute costs. So you really don't want to do that. Now, what you can do is you can have something that's maybe a little bit more of a two-stage architecture where the understanding and reasoning happens in one end with appropriate tool calls being fired off in the back end And maybe then the answer is being received. So it's really just like you would also do in multi-threaded programming, right? Where you have a main thread and it may fire off other threads in parallel that may do things and then get the answer back such that you can get a nice data flow.
35:12And I think a lot of tool calls, agentic programming and so on, have actually been quite helpful there. The other thing is it also allows us to build models that you can instruct just like you would a language agent. The prompts that you write for our model look and feel the same as if you were just instructing a regular LLM, just that our model also talks. So that makes it much, much easier for humans to work with that. And of course, in the back, there is a lot of engineering going on, deciding when to do a tool call, when to answer it directly. And then you also need to figure out when is the question easy enough that the model can answer it?
36:05And when is it sufficiently hard that you probably need something else? And again, that depends on the model size, depends on the information that's required. so do you need a database query do you need to talk to an MCP server right and yeah that is actually something where we I think spent a lot of time and the reason why we spent a lot of time is because we want to make sure that the model is affordable so you can easily build a very big model and that then uses a massive GPU and unfortunately do the user in the end pays for that massive GPU. So right now, on benchmarks, we are better than, let's say, GPT and Gemini and CROC at a fraction of the cost.
37:05So we are better than OpenAI's models at one-tenth the cost. Caveat with that.
37:12Sam Charrington:On what metrics? Intelligence. Speech accuracy, all of the above. So this is, for instance, for, you know, Big Bench audio or complex function bench and so on. So there's a couple of corresponding benchmarks. Now, the little footnote is this is with thinking turned off in these models. Now you might wonder, you know, why on earth would you turn off thinking? Well, because you don't want to have the big pause where the model thinks and then it responds, right? Because that feels very unnatural. So humans don't do that either. I mean, they may say, hey, let me think, but that's now, you know, you need to correspondingly integrate that.
38:01Sam Charrington:I think that's what, that was actually the thought that led to kind of this hierarchical thing. Like I'm envisioning a model whose primary function is maintaining the conversation and it might say, oh, that's a really interesting question. I'll have to think about that for a second. And kind of like you alluded to, kind of stalling, you know, while the tool is completing to retrieve the information. So, for instance, you know, if you test our, you know, you can test it out actually on our demo live afterwards. So basically, for the Higgs live demo, this performs web search in the background. And depending on how quickly it gets the result back, it will just answer or it will actually tell you, hey, let me look for that.
38:59and the challenge is now to make this all feel very organic such that the user doesn't feel any breakage of the entire interaction because you really want to maintain that illusion of a properly engaged other party. So you don't want to break that sense that the model is able to search and do all of those things in the background.
39:32And that's where, yeah, a fair amount of the engineering comes in.
39:38Sam Charrington:And for speech applications, are you able to use off-the-shelf MCP servers for most of the things that you might want to use? Yes and no. So, yes, we can. But the quantity of servers enabled at the same time right now is a little bit limited. the next iteration is going to support significantly larger numbers of them at the same time. Basically, this is all about context window sizes and so on. And again, cost of inference. So there's a trade-off because if you do quite a massive pre-fill and then you have a large KVCage that you need to log around with you, that, of course, makes the token generation more expensive.
40:29and this is really a little bit the trade-off. But what you can do is you can then have tool calls and the tools themselves have MCP configured and so on. So there are plenty of nice ways how you can do this efficiently. Of course, you can also do things where you build, you know, hundreds of billion parameters front-end audio model and that makes for gorgeous demos, but not so gorgeous, you know, cost as soon as, yeah. This is really, I think, where we took a slightly contrarian approach where we went for affordability first.
41:15Sam Charrington:That's very interesting. It's also a little different from what I was asking. And the experience that kind of that was the motivation for my question is, you know, as I build agentic systems with just kind of off the shelf MCP servers, internet, you know, based services, you know, personal data, they're slow, really slow. and it's annoying even in text. It's unimaginable in a voice scenario. And so really the question I was asking was almost, do you have to build everything from the ground up in order to make it work for voice? Like I'm imagining the way you would build a weather MCP server is a lot more efficient than, you know, in an ideal world.
42:09As a matter of fact, you can test out, you know, basically search news and weather and it feels very fluid in our application. So you can go to, you know, Pose on AI.
42:22Sam Charrington:And presumably you didn't have to build those yourself. You're just using... So you use a good server for that. And being able to use these, I mean, I think it's both a feature and a necessity. It's a necessity because you don't want to have to reinvent all the tooling again. it's also necessity because our customers don't want to have to relearn new techniques right and that that's that then also becomes a feature because it makes it easy to or at least easier to integrate within an existing system right but you know if you have a service that takes a long time I mean humans are pretty good if I ask you a question that you have to look up like if i were to ask you hey when was i last time on twimble and i don't think you remember the exact date you would probably tell me something like hey alex let me look it up and then you'll you'd open your laptop and do you know your search and maybe five or ten seconds or maybe a minute or two later you'll come back and say oh this was at that and that date, right?
43:40And it would feel totally normal for me, right? I would not be offended with you spending that time because you told me before, hey, this is going to take some time. So humans are very tolerant to delays if they are told that there's a delay. My favorite example is actually the boot screen from Apple. So Apple does something brilliantly sneaky there. And I guess we all have seen the boot screens on an Apple device, which goes nicely linearly. And then typically before it's even completely done, then suddenly it switches to, okay, it's on, right? And we've also seen the infuriating boot screen on a Microsoft device where it goes and goes and goes, and then it gets stuck at 99 % and you wait for two minutes for the last percent to complete, right?
44:41So you might wonder, you know, how does Apple do that, right? They, you know, after all, you know, is there any secret magic? No, actually, they lie to you. The Microsoft boot screen is the truth. What Apple does is very cleverly, they measure the boot time that it took last time. and they add a tiny amount to it and then they have a fake boot progress bar that's timed to go all the way to the end within the time that it took last time to boot. So in other words, you get this very nice progress bar that is eye candy. It means absolutely nothing. It just indicates how long it took last time. humans are perfectly fine with it right so i don't know whether that is still the very implementation now but it was the implementation for a long time and this is a brilliant slate of hand where i can create a very good user experience without having to solve an impossible technical problem right the other thing that you also need to do is you need to start actually finding benchmarks, like what makes for a good user experience for interaction for voice.
46:05Like, you know, is the model properly interruptible? Can the model recover properly from that? Right. Does it know that, for instance, in, you know, in Japanese, it's very common to say, hey, so basically you're back channeling the other person and saying, hey, I got it. Yeah. Oh, that. Right.
46:32Sam Charrington:But if the agent stops for every one of those, that's going to be infuriating. Exactly. On the other hand, if that person were to say, well, I don't understand, you want the model to stop. Right. So what that means is you need to make sure that the interruptibility is really scene dependent. So for instance, we released a benchmark exactly on that. That paper I think we put it up on Archive a month ago. We looked at the level of productivity in interaction. So basically whether the model actually can pick up if there's something missing and whether the model should also step in. So basically you do need to do extra user experience engineering and benchmarks and measurements and optimization to really make the model work really well in this context.
47:27So that's something that I think is quite different and new relative to what we had in text. And, you know, that's, I mean, to some extent, that's, you know, what AI for humans really means to optimize in a way that it's pleasant for humans. So this is a slightly different optimization track rather than models for code generation, right? they worry about task completion. In our case, we worry about human happiness. You know, you still need task completion and all of that, but you also want to make it enjoyable.
48:06Sam Charrington:What was the name of that benchmark? So ProactBench is the one for productivity. And then there's another one also by the same team. This is basically my Toronto team that has released it. If you look at my blog, there is a fairly detailed analysis and discussion. So IHBench, which explains exactly, you know, the various metrics that we use, how you then also go and make this nicely reproducible TLDR. I mean, we do use third party audio models in the defined benchmark. So what happens is you actually need to measure appropriately all the interruptibility and responsiveness and whether, for instance, the audio matches what's being said.
49:01My favorite example is from the Hitchhiker's Guide to the Galaxy. I guess I'm showing my age here. There's the BBC series. and they have this notoriously cheerful computer on a spaceship. And there's a scene where the entire crew is about to die in two minutes because the spacecraft is going to crash into something. And the computer very cheerfully announces,
49:32Sam Charrington:you will crash in two minutes, we will all die. And of course, this is for comedic relief and spoiler alert, no, they don't die in two minutes, obviously. That's an EQ point that you were mentioning earlier. Exactly. So what I'm trying to say is that for voice and then also for video, you need to really care about how humans feel rather than just doing text only. And I mean, yeah, you want models that are not as dumb as a brick, But there is a difference between IQ and EQ and the latter is also important for humans. I was going to ask, there's quite a bit of research on EQ from a generic AI perspective, you know, probably text focused to, you know, but I'm imagining part of the kind of the through line in our conversation is that things are just different for text.
50:35Sam Charrington:the challenges tend to be more systematic as opposed to we're just going to solve this model and put a text front end there are a lot of moving pieces i'm imagining that same thing is true for eq solving eq from a text perspective doesn't necessarily get you you know eq voice ai it will get you to parts of it because so from a text transcript for instance for interruptability, you can figure some things out, right? But then in some cases, it's also a matter of, you know, what's a good point of reference. So let me give you examples of terrible points of reference. So you could think about, you know, maybe movies are a really great source of how humans should interact with each other.
51:29Well, take the romantic comedies and usually, you know, the obsessed guy who in the end gets the girl essentially stalks her, right? If you did that in reality, the police would show up and lock you up, right? likewise in some other movies people are quite liberal with their fists or kicks or whatever and in the end you know they will you know kiss and make up so to say or you know be best friends again again in reality if you do this you end up in jail so it's very clear that a lot of social behavior in movies isn't quite so realistic. I mean, there are more realistic things in terms of talk shows and so on, but even there, I mean, I sincerely hope that nobody will train a model on Dr.
52:27Phil and assume that this is normal human behavior. I mean, you know, it's entertaining TV, But to make it very clear that you do need to, you know, be more careful in terms of, you know, how interaction should really work. The good thing is that LLMs by now have a decent theory of the mind. So that gets you somewhere. And then, yeah, I mean, we're now entering, you know, the exciting new world where we can actually go and, you know, explore how humans really interact. We have to be very careful about it because, you know, ethics matter and you don't want to, you know, experiment on humans. But, you know, this is, I think, a really exciting time, at least not unless you do this appropriately.
53:23So I think it's a really exciting time of where this is going.
53:26Sam Charrington:And by that, you mean the technology is sufficiently far along enough that we can start putting in front of real people and engaging their reactions and that kind of thing? Is that where you are going? That is effectively, I think, what will happen where those models will get better by learning how to interact with humans. And yeah, I mean, this is an exciting new world. Yeah. And that makes me think of, you know, there's a lot of conversation right now about recursive self-improvement, right? And, you know, we, it's commonly discussed from the perspective of, you know, as the models that we use to build tools or to build things, you know, get better, you know, we can build more things, the models, you know, build better models, et cetera, et cetera.
54:15Sam Charrington:But is there an extent in which like the learning loop can become kind of continuous and the model can, you know, in a conversation, you know, learn what works for a particular person like that? I think that brings in a lot of things, this kind of self-improvement, like personalization. But it's something that I don't think we see very, you know, we don't see that in interactions today. This is starting, I would say, and you have a number of avenues. The closest thing I think we see is like memory in, you know, your traditional LLM. Okay, so let's actually go over a couple of pieces there. So the recursive self-improvement, it's cheap if you can do it in, you know, so to say fully on a computer by just interacting, let's say, with an LLM that has a decent theory of the mind.
55:19So for instance, if I ask an LLM, well, what happened? Yeah, so like, and they're a user simulator. So for instance, NVIDIA did a great job at releasing some digital personas and some scenarios and so on. And we use that. And we also wrote papers about it. So it's no secret. And what that does is it also allows us to build models that work well for humans that are not always cooperative, that are maybe a little bit abrasive, that are a little bit unpleasant to deal with, right? Or people who might want to break the system, right? So no human is perfect and people will try to have fun with those things.
56:07But, okay, so there's the, you know, given that those models by now have a decent idea of how humans work, you can use that for, you know, RSI to some extent. Another extent and direction, though, is then personalizing things. So, for instance, knowing that Alex likes, you know, equations and facts and technical details and maybe has a little bit of rough edges on the social side or whatever. However, you know, these are things that I can use for personalization such that, you know, next time I interact with Alex, I can probably, you know, personalize better for him. Right. Or, you know, knowing which devices I have, what my preferences are, what my language background is and all of that.
57:04this all can be used to improve direct personalization by just having models reason over it. In a similar way how, for instance, if you look at Hermes agent, which then looks at past interactions and past things that it's done and it reasons and it commits that into memory. So think of this as a glorified CRM, but now for everybody. But at the same time, you also want to learn across all the interactions. So, for instance, knowing that if I insult a user, the user will not take kindly to that. I mean, okay, sure, it's an egregious example. And of course, nobody would actually do that. But the point being, you know, learning such things across is something that you can then overall use for an improved model.
57:57So you basically get global improvements. So it feels very similar to recommender systems just now in 2026, where you have an overall behavioral improvement, but you also have personalized improvement. The latter improving as the model gets to know more about, you know, let's say Sam or about Alex. But then overall, the model getting better at dealing with those pesky humans or maybe dealing with those pesky humans in North America versus, you know, there's certain appropriate behavior in some parts of the world that's perfectly inappropriate elsewhere. so like eating with your left hand in India is seriously frowned upon because you use the right hand for eating and the left hand goes behind your back right so that or if somebody sees the bottom of your feet in Thailand it's also not a good thing but basically there's a lot of cultural weirdness and you know weirdness that really matters for the people that you know live in that culture and learning that will allow probably also to personalize a little bit better for you know specific regions specific cultures but then also for individuals within that culture and, you know, at the other end, you know, behavioral improvements for, you know, agents interacting with, you know, just those squishy humans.
59:38So that I think is a really exciting future. And by now we will be able to do this at a scale where we can actually build things rather than just positive theories, right? There's, you know, glorious work by, So for instance, if you look at Maya Mataric, so she builds those robots that interact with babies and get babies to move their legs and whatever. And there's a lot of engineering and science brilliance in building this stuff. Brilliance is good, but brilliance is not repeatable and automatable. what I think the future is going to be is going to be systems that automatically learn this stuff such that in the future you don't need quite as much brilliance anymore and you know you can just let data speak you know for movie recommendations this is by now a solved problem mostly but for overall human behavior interactions and improvement, I think even humans are not particularly good at it.
1:00:55If I was trying to aggravate somebody and to just harass them, a lot of people might at some point lose their cool and respond in kind, right? Because not all humans are nice. but for the purpose of you know achieving a given task for an agent you actually want the agent to keep his cool and to to handle this right and i mean this is one thing when you you know watch you know school teachers and how great they are with kids and to get them to behave even though those kids are very uncooperative maybe initially, right? There's a skill. Most humans don't have that. But, you know, we can learn those skills eventually with these organic systems.
1:01:54I think this is an exciting world.
1:01:56Sam Charrington:Well, there's a bit of a paradox in that if most humans don't have them, they're not in the data, and therefore we won't be able to easily train models to follow those patterns. Well, you get it from interactions, right? And basically, you can think of each interaction as another new experiment, a new data point. And so eventually, you know, by some randomness of exploration and so on, you'll be able to learn what works and what doesn't work.
1:02:35I mean, the question then is, you know, how you use that appropriately. But for instance, if you need to have a difficult conversation with somebody, I mean, the fun example is the millennial shit sandwich, right? where you have praise, critique, and praise. And by doing that, you know that you can deliver the critique without people getting too upset, right? And this is a strategy of communication, right? Or when you learn how to become a manager, you take all those training courses. How do you deal with people who are happy, who are unhappy? How do you, you know, basically deal with people in a way that, you know, they can be productive, that they can enjoy?
1:03:41I mean, there are techniques. And, you know, these can be learned. They can be learned from instructions. You can take Cialdini's books and they probably contain some good instructions. but we can also then eventually go and learn from data. This is, I think, really going to be quite an exciting revolution for the next maybe two to three years. I think it's going to move fast.
1:04:08Sam Charrington:Well, Alex, it's been wonderful to catch up with you and get into a little bit of what you're seeing and working on with regards to voice AI. Thanks. Thanks for having me. Awesome. Thank you.
1:04:24Thank you.
From the publisher
Voice AI has gotten remarkably good, but natural conversation remains a high bar. Small delays, awkward interruptions, or the wrong tone can quickly break the illusion—and adding vision and visual presence only raises the stakes.
In this episode, Alex Smola, co-founder and CEO of Boson AI and professor at Carnegie Mellon University, explores the path from today’s voice agents to audiovisual agents and AI avatars. We discuss the technical tradeoffs behind real-time voice, including audio tokenization, latency, model size, and inference cost, as well as what changes when these systems can both see and be seen.
We also explore the role of emotional intelligence in AI, how agents can learn from human interactions, and what it will take to move beyond impressive demos toward interactions that actually feel natural.
🗒️ Full show notes: https://twimlai.com/go/777.




