Voice AI’s Big Moment: Why Everything Is Changing Now (ft. Neil Zeghidour, Gradium AI)

19 Feb 2026 · 1 h 23 min · 34 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

```markdown The MAD Podcast with Matt Turck

Episode Summary

Voice AI’s Big Moment: Why Everything Is Changing Now (ft. Neil Zeghidour, Gradium AI)

Overview In this episode of The MAD Podcast, host Matt Turck engages with Neil Zeghidour, CEO of Gradium AI and a leading expert in Voice AI. They discuss the evolution of voice technology, the current developments in AI-driven voice interactions, and the implications for future interfaces and applications in AI.

Key Themes and Topics

  1. Voice AI's Current Landscape
  2. Voice AI is experiencing a significant surge in interest and innovation.
  3. Historically viewed as an underdeveloped modality compared to text, image, and video AI.
  4. Recent advancements have led to improvements in naturalness, latency, and accuracy of voice interactions.
  1. The Cascaded Voice Stack
  2. Neil describes the dominant "cascaded" voice stack, which involves:
  3. Speech recognition converting audio to text.
  4. A text model processing the information.
  5. Text-to-speech converting text back to audio.
  6. Downsides:
  7. Increased latency due to multiple processing steps.
  8. Loss of paralinguistic signals (e.g., tone, emotion).
  1. The Future of Voice Interaction
  2. The next wave involves integrating speech-to-speech models that bypass the text bottleneck.
  3. Full-duplex interaction enables simultaneous speaking, making conversations feel more natural and fluid.
  1. Technical Challenges
  2. Addressing real-world complexities such as:
  3. Noisy environments.
  4. Multiple speakers in conversation.
  5. On-device voice processing for efficiency and privacy, using compact models.
  1. Voice Cloning and Privacy Concerns
  2. Voice cloning has practical applications but raises ethical concerns, particularly regarding deepfakes and privacy.
  3. Neil states that traditional watermarking methods for audio security are ineffective and calls for more vigilance from users.

Neil Zeghidour's Background

  • Transitioned from finance to machine learning after discovering AI during an internship.
  • Worked at prestigious organizations including Google, DeepMind, and Meta, contributing to significant advancements in voice technology.
  • Co-founded Gradium AI to focus on real-world applications of voice AI.

Industry Insights

  • Gradium AI operates as a nimble, focused startup, contrasting with larger organizations by specializing in compact, efficient models.
  • Emphasizes the importance of building blocks for developers to create diverse voice-related applications, rather than attempting to build comprehensive solutions themselves.

Conclusion The podcast highlights that while voice AI has made significant strides, challenges remain in real-world applications and ethical implications. The conversation encapsulates the optimism for future advancements in voice technology and the belief in the potential for small teams to innovate in a competitive landscape.

Additional Resources

  • Neil Zeghidour
  • [LinkedIn](https://www.linkedin.com/in/neil-zeghidour-a838aaa7/)
  • [Twitter](https://x.com/neilzegh)
  • Gradium AI
  • [Website](https://gradium.ai)
  • [Twitter](https://x.com/GradiumAI)
  • Matt Turck
  • [Blog](https://mattturck.com)
  • [LinkedIn](https://www.linkedin.com/in/turck/)
  • [Twitter](https://twitter.com/mattturck)
  • FirstMark Capital
  • [Website](https://firstmark.com)
  • [Twitter](https://twitter.com/FirstMarkCap)

Episode Timeline Highlights

  • 00:00 - Intro
  • 01:21 - Voice AI’s big moment — and why we’re still early
  • 03:34 - Why voice lagged behind other modalities
  • 11:01 - Voice vs text: where voice fits (even for coding)
  • 34:01 - On-device voice: where it works, why compact models matter
  • 54:05 - Hardest frontier: noisy rooms, multi-speaker chaos
  • 1:12:30 - Voice + vision: multimodality in interaction
  • 1:16:32 - Paris/European AI: talent density and the future

This episode serves as a comprehensive exploration for those interested in the dynamic field of voice AI, highlighting both its challenges and exciting prospects. ```

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Evolution of Voice AI

0:46 to 1:48

Discussing the significant changes in voice AI technology over the last 18 months.

“Neil is one of the very top AI researchers in the field and a key architect of the rapid evolution of voice AI towards real-time native audio intelligence.”

Neil Zeghidour's Insights on Voice AI

1:49 to 4:13

Neil shares his perspective on why voice AI is gaining prominence and its current state.

“and in voice, for example, the progress in latency, naturalness, accuracy have been really, really huge in the past years, in particular in the two last years.”

Challenges in Voice AI Development

4:14 to 6:33

Exploring the historical challenges and underdevelopment in voice AI compared to other AI modalities.

“you had to have an application either in computer vision, like image classification or NLP.”

Expertise in Voice AI

6:34 to 8:00

Neil discusses the rarity of expertise in voice AI and the impact of individual contributions.

“You know, so all of this, when you bring them together, you can make competitive models.”

Vision for Future Voice AI

8:01 to 9:39

Diving into the future of voice AI and the potential integration of voice technologies.

“And in that context, there is no real latency anymore.”

Neil's Journey in AI

9:40 to 13:59

Neil shares his personal journey into the AI field and his experiences at various companies.

“and it's even interesting how relevant it still is, despite the fact that there's so much work around voice.”

Early Journey into AI and Speech Recognition

14:02 to 16:57

Learn about Neil's initial experiences in AI and his work in speech recognition.

“So back then there were no LLMs and it was not really about generative AI.”

Researching Efficient Learning in Speech

16:57 to 19:45

Discover insights on the efficiency of language learning and data usage in speech technology.

“It's like Silicon Valley or the HBO show.”

Innovating with Neural Audio Codec

19:45 to 22:22

Explore the development of Soundstream and its implications for audio compression.

“that allow to compress audio while making it as transparent as possible for the human ears, basically.”

Founding Gradium AI and Shaping Voice Technology

22:22 to 24:10

Understand the motivations behind founding Gradium AI and its focus on voice technology.

“At the time where Gemini started at Google, that's when I left.”
Show all 34 chapters

Creating Conversational AI with Moshi

24:10 to 27:10

Learn about the development of the Moshi model and its significance in conversational AI.

“Because we had done our PhD together at Facebook.”

Commercializing AI Models: The Gradium Approach

27:10 to 28:00

Discover how Gradium transitioned from research to commercial applications in AI.

“So as I said, our open source models were very successful.”

The Journey from Academia to Impactful AI

28:00 to 29:14

Discover the shift from academic success to real-world applications in AI.

“to be multilingual, higher quality, you know, all the things that make an actual product.”

Navigating the Competitive Landscape of Voice AI

29:14 to 31:31

Learn why smaller companies can compete in the voice AI market against giants.

“So I think in the end, even in terms of achievement, academic success is one thing, but nothing is near in terms of impact to the fact of having people using your models for the real world.”

Building Blocks for Voice AI Products

31:31 to 34:01

Understand how specialized models can advance voice AI applications.

“and obviously you are a new entrant in the field of AI where there's tremendous amounts of competition.”

Challenges of On-Device Voice AI

34:01 to 36:20

Explore the complexities and advancements in on-device voice AI capabilities.

“And then there's an additional aspect to this, which is that voice can also, and should be pretty often on device versus an API call to the cloud.”

Open Source vs. Closed Source in AI

36:20 to 38:32

Examine the relationship between open source development and competitive product creation.

“two weeks ago is also based on our framework.”

The Importance of Engineering in AI Technology

38:32 to 41:28

Learn how engineering expertise contributes to the success of AI applications.

“And I think if we want to stay at the cutting edge, we should be a product company, but also be a frontier lab.”

Evaluating Voice AI Quality: The Role of Human Judgment

41:28 to 42:00

Discover how human perception is prioritized over metrics in voice AI evaluation.

“If you want a model to run on all Android and iPhone and so on, you know, that's a big engineering challenge.”

Challenges of Objective Proxies in Audio Evaluation

42:00 to 45:00

Explore the difficulties in creating objective measures of audio quality compared to human judgment.

“People have tried to make objective proxies of human judgment.”

The Evolution of Voice AI Models

45:00 to 48:20

Understand the evolution from traditional text-to-speech systems to more advanced AI models.

“I think then people talk about TTS already, right?”

Challenges in Speech Recognition and Interaction

48:20 to 55:00

Learn about the current challenges in speech recognition, including turn-taking and environmental noise issues.

“You know, a lot of information is conveyed that is not in what to say.”

Data Requirements for Voice AI vs. Text AI

55:00 to 56:00

Discuss the differences in data requirements and training methods between voice AI and text AI.

“Again, it's both a hardware and software issue.”

Understanding Data in Voice AI

56:00 to 58:00

Explore the differences in data requirements for training voice AI compared to text AI.

“For this kind of hard understanding problem, I hear people seeing the exact same stuff as they did 10 years ago.”

Quality vs. Quantity in Speech Data

58:00 to 1:00:20

Learn about the importance of high-quality speech data and its implications for AI models.

“So for the diversity of voice, the amount of data is extremely useful.”

Challenges in Language Data Collection

1:00:20 to 1:02:50

Understand the complexities of gathering data for less common languages and dialects.

“And the accuracy, for example, of annotations, it's paramount and it's a lot of human labor to be able to annotate precisely.”

The Business of Voice AI Products

1:02:50 to 1:08:52

Discover how voice AI products are being developed and the role of voice cloning.

“And then they want to talk to it three hours a day because it's a very good game.”

Addressing Privacy Concerns in Voice Technology

1:08:52 to 1:10:00

Examine the implications of voice cloning technology and the challenges of ensuring privacy.

“Let's make a voice that represents your brand.”

Understanding Deepfake Detection

1:10:00 to 1:11:16

Learn about the challenges of detecting deepfakes and the importance of vigilance in communication.

“It's probably not the case if someone gives you a phone call.”

Voice Control and Ownership

1:11:17 to 1:12:08

Explore how voice cloning works and the potential for community sharing and compensation.

“So I think this one has a lot of value because it's a good way of sourcing a lot of voices.”

The Future of Voice Technology in Media

1:12:09 to 1:14:12

Discover the integration of audio and video in technology and its applications in media.

“Because then, again, people typically are going to clone the voice of someone, but what they wanted is someone from a specific gender, specific demographics, age, accent, and so on.”

Innovations in Voice AI

1:14:13 to 1:16:28

Learn about new developments in voice AI and the creative potential of audio transformations.

“You can record a small, short love message.”

Overview of the French AI Scene

1:16:29 to 1:19:48

Gain insights into the current state of AI in France and its potential on the global stage.

“And last theme or question, you're building this company out of Paris.”

The Strength of French Talent in AI

1:19:49 to 1:22:22

Examine the talent pool in France and how it contributes to advancements in AI technology.

“Because in a way, you know, when you see like the most successful French companies, it's luxury or that kind of stuff.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00For the first time, it actually can be enjoyable and even more convenient to talk to an AI on the phone than talking to a human. I don't want to be mean to my people, the speech scientists, but historically, for some reason, voice did not attract the visionaries in machine learning. All the new hardware companies have voice at the heart of the product. All of these devices, they got rid of keyboards. They don't really have a screen or an interface and voice is going to be the main one. Hi, I'm Matt Turck from FirstMark. Welcome to the Matt Podcast. Voice AI is having a big moment. For years, the field was stuck in the uncanny valley, lagging well behind other AI modalities, robotic, slow, and frustrating.

0:38But in the last 18 months, everything has started to change. My guest today is Neil Zegidor, CEO of Gradium AI and formerly of DeepMind and Meta. Neil is one of the very top AI researchers in the field and a key architect of the rapid evolution of voice AI towards real-time native audio intelligence. This conversation is a deep dive into everything you need to know about voice AI, where we explore many key concepts in a very accessible way and discuss plenty of fun stuff, including why voice AI has so few experts, the massive challenge of building native audio models and the rise of autonomous voice agents.

1:12Please enjoy this terrific and very educational conversation with Niels Aguidor. Hey Niel, welcome. Hey, thanks for having me. So a lot of people in the industry are saying that voice AI is having its big moment. There's certainly a lot of activity. There's a lot of funding rounds. From your perspective, so you've been in this field for many years now, DeepMind, Meta, Nagradium. Is voice AI indeed having its big moment or are we still early? I think it's both having a big moment and we're still early. It's having a big moment because there is progress all around AI models and in voice, for example, the progress in latency, naturalness, accuracy have been really, really huge in the past years, in particular in the two last years.

2:01And at the same time, text models have evolved into what we now call agents, which are not only text models, but they can actually make actions and manipulate data, access information and so on and so forth. And now when you bring both together, you can have voice interfaces that at the same time are going to solve complex problems. And so I think there is a moment now because for the first time, it actually can be enjoyable and even more convenient to talk to an AI on the phone than talking to a human because you can call any time of the day or night. And the interaction is working pretty well and it sounds really nice and the latency is low and so on and so forth.

2:42So it's definitely having a moment because I think in a way, now it can be used in much more use cases than it used to but it's still early because it's still quite experimental so anybody who is using even the most advanced voice agents and compares that to the Her movie from 12 years ago it's obvious the gap that is still remaining and there are so many topics that are completely undressed at the moment in particular every time you watch a voice agent demo, just realize that it's someone talking to a phone in a quiet room. So the day where you will have someone shouting to a robot in the middle of a factory and having the robot understanding what's happening and who's talking to them, that will be, you know, like we'll be there and we are not there at all.

3:28So we'll get into some of the technical details in a minute, but at a high level, why has voice AI being, I guess, the most underdeveloped modality? There's been obviously extraordinary progress on text AI and then image AI and then video AI, but it seems that voice has been a little bit the poor parent in terms of progress. Why is that? I don't want to be mean to my people, the speech scientists, but historically, for some reason, voice did not attract the visionaries in machine learning, right? So even if you looked at the dynamics in conferences, if you propose a new method, like fundamental algorithm, and you wanted it to be accepted in a prestigious venue, you had to have an application either in computer vision, like image classification or NLP.

4:19If you did it in speech, you would get rejected because it was like two speech. And at the same time, the prestige of speech conferences used to be much lower than that of computer vision or NLP. So honestly, I don't really know why because when you look at the details, the first big success of deep learning, everybody knows the AlexNet model in 2012 where for the first time, you know, you had a deep learning model outperforming every single alternative on image classification. But actually, the really first big success in deep learning was speech recognition. We've worked from Geoffrey Hinton himself.

4:58And when was that? It was in, I think, 2007 or 2008. So way before. Yes. And so I think it's, you know, it was just not as prestigious. and so it wouldn't attract a lot of people that could have made significant contributions, I would say. And then what happened and was very nice is there was kind of a convergence of algorithms around transformers and LLMs so that now pretty much regardless of the task or modality you are looking at, you're always looking at the same technology. And there started to be much more progress also thanks to now the similarity between different modalities because in particular what I contributed to as my team was to take a lot of inspiration from successes in vision and NLP and apply them directly to speech.

5:48But it's still interesting because in a way, there are way fewer people who can train a competitive speech model than in text or in vision. So, I mean, it's a good position to be in because it's, you know, very few people have really gone into the depth of this topic. And it's one that is very challenging because it's bringing, ideally, if you want to solve the problems, you need to understand machine learning, signal processing that is much more, you know, completely different literature around telecommunication, audio compression and so on, along with psychology and not like cognitive psychology, but psychoacoustics.

6:30How does the human hearing work? How does speech production work in humans? You know, so all of this, when you bring them together, you can make competitive models. But it requires kind of a very wide scope of expertise in very different domains. Fascinating. To put it in numbers, how many would you say people there are in the world with that expertise? Are we talking about 100, 500, 10? Between 10 and 100? No, I would say. 50? I don't know. So it's tiny. It's hard to say. But yeah, I think it's very few and really meaningful contributions that have pushed the field forward have been made by very small groups of people.

7:09And I think that's also what's nice. So AI, I think in general, is one field where individuals can have a disproportionate impact because, you know, the amount of things you can do by yourself, you have access to compute and data sets is huge. and in voice in particular, since the required compute is much lower and is that the same for data? Really, a few individuals can make stuff that is completely, you know, just changing applications at very large scales. Great. So you mentioned her a minute ago, which is the individual reference for any conversation about voice. What is the ultimate success in voice that the field is working towards?

7:50Is that super low latency, expressiveness? what is great? So latency latency is already something that only makes sense if you are in a turn-based conversation because latency and the definition of latency is how much time there is between two turns one of the things we contributed doing is getting rid of speaker turns completely with what we call full duplex conversation so the idea that it's always listening always speaking and when it's not speaking it's just that it's producing silence, but you know, it's always on. And in that context, there is no real latency anymore. So because the model is just basically can talk at any time and it can talk over you and you can talk over it.

8:33And that makes the conversation really natural. Then naturalness, it's not only these dynamics in terms of tempo, you know, like when the model can jump in the conversation, when they should remain quiet, there is this dynamics question. And then there is emotion. And so there is emotion in what the AI expresses. Its emotion is natural, but also appropriate. That if you start feeling confident enough to start sharing about stuff that makes you unhappy or sad or feeling miserable, it's not saying, oh, I'm so sorry for you. Let's talk about it. And it will also understand when you're getting upset, when you're getting confused and so on and so forth.

9:13this will make already voice AI in terms of interaction extremely natural and as close as possible to human which is basically what is one of the things we see in the Her movie which I hate mentioning as a reference because it's so overused it's annoying but at the same time everybody understands you know the gap between where we are right now and the movie so that's I think it's still a relevant one and it's even interesting how relevant it still is, despite the fact that there's so much work around voice. But then there will be other questions about how voice is integrated into our life. So, you know, there are paradigms in voice AI, such as wake word detection.

10:01So, you know, when you use Google Home or Alexa or whatever, you have a wake word that is going to turn the speech to text on. So now let's say you want to work with your assistant that is always listening to you. So in a way, you would have something that is just running constantly without having even necessarily a wake word. So all of this, I think, is going to be both technical challenges and product challenges around where do they sit, how are we interacting with them, the link to the hardware as well. So I think what is a good sign for voice as a field as well is that in my perception, all the new hardware companies have voice at the heart of the product.

10:39all the prototypes that we see, whether it's glasses or pendants or, you know, like the new stuff that Johnny Hive and Sam Atman are working on. Voice is at the heart of the product and will be the main way of interacting. So all of these devices, they got rid of keyboards. They don't really have a, like a screen or an interface rather on, you know, and voice is going to be the one that is the main one. What's your vision of the future? Where does voice fit in? Is that voice and text? Is that primarily voice for certain use cases? There's certainly an argument that you'll hear a lot of people saying voice is great, but most of the time I'm at the office, the last thing I want is for people to hear my conversation and therefore I don't want to talk to a machine.

11:21So where does voice fit in that vision of the future? For example, I used to think that one obvious application where voice was kind of irrelevant was coding because it's fundamentally, you're not going to read code out loud, right? Yet now, since coding is going more and more towards vibe coding, which is natural language, it makes a lot of sense to do it by speaking. And now people are developing products that allow you to dispatch orders to coding agents in a way that is much more efficient than if you had to type in each different window to each of them. Even prompting LLMs now is doing by voice is much more convenient rather than typing.

12:01I still agree that there is one part which is more social about what the office's environment will look like. I don't know, maybe we will just also rethink the way we just structure office environments. What is sure is now people have AI assistants that are almost colleagues, right? I mean, you talk to any software engineer, the anthropomorphization of cloud code is, I find it extremely funny. you know even the verb coding I mean it's going to be clothing pretty soon and so these people you know they will if it's more convenient to interact with their main tool through voice that will justify also rethinking office spaces I guess so yeah I think there will be work around and we will naturally find them if voice becomes the main way I mean if it's more practical to interact with AI through voice super interesting before we go further, let's talk about you a little bit and your journey and the company.

13:01So I mentioned DeepMind and Meta and now Gradium and Utah in the middle. Just walk us through your life story and your work. So I studied mathematics and I started my career with a short internship in quantitative finance. That was in Paris, right? Yes, in Paris. And I was born and raised there. And what was interesting during my internship, I had access to Bloomberg, you know, terminal. And so I will see the news, like the constant news below the screen. You know, I was thinking, what if I could have an algorithm that just reads these news and take positions on the market faster than anyone because it was just able to analyze the news live.

13:42And I was looking, but I had no keywords about that, right? So I literally Googled how can I analyze text automatically or whatever. And I found machine learning. You know, it was an epiphany. decided to completely stop. Started studying again. My goal was to go back to finance with AI, you know, and machine learning. But I got passionate about all the possibilities there was around. So back then there were no LLMs and it was not really about generative AI. It was about medical imaging, text understanding, a lot of things around audio, speech recognition, obviously. And I looked for an internship which was about unsupervised learning which was already pretty cool and I just wanted to do it and so I pretended that I was passionate about language and I got the internship and then Yann Lequin opened Facebook Paris and I was able to interview so for the anecdote I did my coding interview with Sumit Chintala who then invented PyTorch and now I think he's the CTO of Thinking Machines and so I didn't know how to code because I had only studied mathematics And so he asked me to implement K-Means, basic algorithm.

14:53And I asked if I could do it in MATLAB because I didn't even know Python. I knew like very basic Python. And he was kind enough to let me do my coding interview in MATLAB. And I got the job, which when I think about it, it's so cringe because, oh my God, thank you, Sumit. And yeah, I did my PhD there. It was very interesting. I was already around speech and I was spending half my time at Facebook and half my time at Ecole Normale Superieur in Paris. in a lab that was studying language acquisition in babies. In particular, the main observation on which the lab was built was that humans learn language from mostly two speakers, their parents, with a few hundreds to 1 ,000 hours overall in the first four years, with huge variance between social backgrounds and without annotation, right, because you learn to speak before you learn how to read.

15:43and that still makes us already pretty okay for conversation when we are kids. And, you know, speech recognition back then was trained with already hundreds of thousands of hours of annotated data. Now it's millions of hours of annotated data. So the topic was more around efficient learning, which is interesting because it was 10 years ago, but now it's still as relevant as, I think there was a new company that raised a large scale round recently to make learning more efficient. So it's still as relevant as it's used to. And then I joined Google. At that time, it was interesting because, so I joined working on speech in Google Brain and there were almost nobody working on speech in Google Brain.

16:24It was not considered vibrant research topic. It was like a product topic. A lot of people were saying, oh, but it's solved, you know. Oh no, it's solved. It just works. So already back then. And what year was that? 2019. 2019. And so I found someone to work with me and we did like a lot of work around speech. And then I got excited about a specific topic around compression, audio compression. So it was just out of patience. I wanted to do like a new compression format that will not be MP3. It's like Silicon Valley or the HBO show. Exactly. Absolutely. And I wanted to do it with neural networks.

17:02The idea was that it would be computationally more expensive to compress and decompress the audio, but then you could complex it much more efficiently. And that's something we worked on for Google Meet. And it was called Soundstream. That was the first what we call neural audio codec. And I had no plan of doing generative modeling back then. We didn't care about that. But I was very lonely in a way. And I just wanted, I was trying to lure some people around me who are working in reinforcement learning. I wanted to get them to work on speech with me. I started a project around diarization, which is a task of you're listening to a conversation and you have to tell who said what, which is probably the less sexy research topic out there.

17:47I'm sorry. I think it's fascinating, but you cannot get people excited. I mean, it's very hard to get people excited about that. So within speech, which is not very sexy, this is the less sexy part. Like the monk project, you know, like very lonely and very, yeah. I was not very successful with that to get people to work with me. And I thought, okay, generative models, the nice thing is that if we generate speech, people will listen to it. And they'll be like, oh, that's cool. My thing, you know, my method has generated speech. So honestly, it was very opportunistic for me. I thought that it would be a good way to get people to work with me.

18:20And so we started a project. And the idea was that we started to see success around language models. So it was 2021, way before ChatGPT. But, you know, internally at Google, there were already quite a few projects that were successful around language models. And the idea was that just after the work we had done on the neural codec, so now if instead of using your codec for real-time communication, but you just use it to compress audio, now you have... Do you want to define quickly what a codec is? Yeah, a codec is just a compression decompression. So you have an audio, right? And you want to send it over the network when you're having a Zoom meeting.

18:56And you're not going to send the uncompressed WAV file because it's too heavy. So you're going to compress it in a much lighter file that you will send over the network, and then the receiver can decompress it back into audio. And the secret is, based on a lot of science and knowledge around human hearing, we know what kind of information we can remove from audio so that it won't create a perceived degradation, basically. So there is a lot of science around what specific information you can remove from an audio. that will make it almost as good for a human as the original one, despite the fact that you removed a lot of information, which allows you to compress.

19:39And the main idea was that instead of using hard-coded rules to do that, we would learn from data what are the transformations that allow to compress audio while making it as transparent as possible for the human ears, basically. And so now we had this way of compressing extremely efficiently, much more than MP3 or Opus on audio. And in a way, you could consider that it was so compressed that it was almost like text. And so we just, simple thing we did is just train LLM to predict this compressed audio instead of predicting text. And then you could do the exact same thing you can do with text.

20:16You could prompt, I take three seconds audio, compress it, I pass it to LLM and I let it predict the next compressed audio and realize we had, in one week, we had invented instant voice cloning. so we could replicate any voice with a few seconds of audio. And yeah, this became extremely successful because it was all the advantages we had with LLMs we could benefit. So LLMs are great at modeling long context. They scale very well. So if you want to have a large model, you just scale the small model. It sounds obvious like that, but for a lot of architectures, it's very difficult to go from a 100 million parameter to a billion parameter model.

20:51With transformers and LLMs, it's obvious. And I could go into more details, but in a... What was that project called? Audio LM. And then it gave Music LM. And then it gave Notebook LM, that was an automated podcast. And that became the standard framework for audio generation. There were two families that were kind of fighting during some time. It was diffusion models, which I think what Eleven Labs was based on early on. And we were the audio language model family. I think today virtually everything is audio language models because since they are auto-aggressive, so they run in a streaming fashion, they are naturally compatible with real-time inference, which is kind of the main topic around voice right now.

21:35And so everybody is using this technology today. And yeah, so it was very, very successful and extremely easy to apply to new tasks. So, you know, we did it on speech first. And so then we collected a data set of piano performances and then we had a model for piano and then we did more general music. and then we could do pretty much anything. It's even used by a non-profit lab working on animal vocalizations to try to decode the language of animals called the Earth Species Project. So yeah, it's as flexible as LLM Vortex, basically. Amazing, amazing. You played an incredibly important pioneering role in the current state of voice areas.

22:20And then what was the next stop after that? At the time where Gemini started at Google, that's when I left. So I wanted to create a small research environment that reminded me of the early days of Fair or Google Brain. So very small team, elite, no distraction. Sorry to say that. No product manager, no, you know, like just research scientist. No emails, just, you know, locked in a room with the machines and focused on science. And in particular, the goal for me was really to keep working on fundamental research and keep pushing the field forward and training students and so on. Because I felt very grateful to have been able to do research in such an open environment.

23:09It was also obvious for me that, and for this, I think I agree 100 % with Yann Lequin, the fact that what made AI dynamic and get from ImageNet in 2012 to where we are right now today is open research. That's because it's kind of a worldwide collaboration and everybody benefits from the progress of everyone. So that's for me, it was important for the field itself to keep this going. And so we decided to create a non-profit with the help of Eric Schmidt, Xavier Nen, and Rodolf Sadeh. So for the anecdote, the codename was Sphere because that's the name of the restaurant where we discussed the project.

23:46And we then understood that we could never trademark the name Sphere, obviously. So we just asked ChatGPT for a sphere in a few languages and a sphere in Japanese is Qtai. And there was AI in it. So I'm like, OK, that's the name of the lab now. And so that's how we created Qtai. The first person I reached out to is Alex Defossé, who is also now co-founder of Gradium. He's our chief science officer. Because we had done our PhD together at Facebook. And then when I joined Google, we kind of became rivals because, you know, we are working on the same stuff at the same time. And every time we meet one another, we'll just not talk about anything.

24:24Like, what are you working on? Oh, nothing. And yeah, and so, yeah, we had a small team, but with big expertise in speech. And we decided to work on the stuff that, again, it was an opportunistic decision that we made. So I looked at, you know, the kind of stuff we could do around voice. and for a lab, okay, we had 1 ,000 GPUs. That seemed a lot in 2023. It was still already at least one order of magnitude lower than biggest labs. So we wanted a project where we could do a difference despite the fact that we were four people. It shouldn't need too much compute and it should be very innovative so that just by being smart about what we did, we could make a difference.

25:07And so we focused on conversational AI and real-time conversation because I had seen from the inside at Google that nobody was daring to touch this topic because it was so challenging. And it really seemed extremely far to be able to cast the task of dialogue into an LLM, right? People are just working on like TTS and speech to text. The interactive stuff was really not there yet. So we thought, okay, we're going to work on it and we're going to make it full duplex. Might as well do something really innovative. What we didn't know at the time is that OpenAI was already working on a speech-to-speech conversation for a while.

25:44But in six months, we have a core team of four people and then six. We were able to ship a model trained from scratch called Moshi that is still to this day the only full duplex model. You can talk to it. It's a bit dumb because it's archaic in terms of intelligence, you know. But the latency is still one year and a half after. It's still the best in the world by far. And I think what was very interesting was that a very small team could do would make such a difference. Because then we shipped the first speech-to-speech translation system and then streaming speech-to-text and then streaming text-to-speech.

26:17And our models have been used across all industry. We are always proud to hear that a lot of big companies are using it, a lot of small companies are using it. The Maya and Miles demo from Sesami was built around our open source models. So I think what I really like about voice and what makes Gradium a project that I deeply believe in is it's one of the modalities in AI where a very small team can make a difference. You don't, I don't really see benefits from having extremely large organization with a lot of people and resources because you don't need that many resources. You're not, you don't need 10 ,000 GPUs to train a speech model.

26:57You don't need 1 ,000 people. The ability to go fast, iterate fast with the right people to me is far superior to the advantage of having a big organization. And then let's talk about Gradium. So Gradium is a commercial spinoff? Yes. So as I said, our open source models were very successful. They are still downloaded millions of times a month. And we started having companies reaching out from all sizes, small, very large. They saw the potential. Sometimes it was a bit weird because I was talking to people leading extremely large teams and I was explaining how I could train by myself, streaming speech to text.

27:37that was better than all alternatives. There was something we were doing very well, but at the same time, our open source models remained limited. They were fundamentally prototypes. For us, they were not even the actual contribution. For us, the contribution was the invention and the related research publication. For us, it was kind of like an artifact accompanying the paper, these models. But people wanted such models to be multilingual, higher quality, you know, all the things that make an actual product. And so we considered kind of outsourcing this part, you know, working with another team that will lead the product and so on.

Read the full transcript

28:11And honestly, after a few interactions, I realized nobody could carry such a project except us. Nobody can believe in something they have not developed it from scratch. We were in conversations with companies that wanted to create partnerships to improve our open source models. and I would look at the specifications they wanted. And I realized it's something we could do in a few days and they had been struggling for months. So, yeah, I mean, it seems that we are also the best team to do that. And from a personal point of view, you know, I was kind of an addict to academic prestige, having best paper awards, all this stuff.

28:52It was always nice, but it was never enough. Every time I did a paper, I was happy for like two hours and then I was already thinking about the next version. And at the same time, it felt a bit weird eventually that I could not imagine that at the end of my career, I would have worked on something so applied and so close to actual real life applications and not doing them myself, you know, not contributing to them directly. So I think in the end, even in terms of achievement, academic success is one thing, but nothing is near in terms of impact to the fact of having people using your models for the real world.

29:27And in a way, I think it's something that is kind of generalized now in the industry. And what's also something that made me think is I looked at all the people I respected the most. The main one being, to me, a living legend, Aaron Vandenord from Google DeepMind. All these guys, they were making so nice contributions scientifically. And then they decided to focus on products. And I think it's, I want to think it's because they also realize that academic prestige is one thing, but making a real impact is having your models being used in the real world. And so for me, it's the ultimate impact we can have.

30:06We still do science. In particular, Qtai keeps doing open source and open science and so on. For me, the upside about it is mostly to be able to train the next generation of AI researchers and keep the field alive, as I said, because I think it's to have a healthy and vibrant AI field, you need to have scientific dissemination so scientific exchange between institutions and they're also you know like the Chinese lab are making a remarkable work and it's kind of are forcing everyone to stay open to some extent because otherwise it also hurts the ego I think of the people who are in the lab that don't publish that was also something where I was very opportunistic about so my strategy was like if we publish in a world where the authors don't publish they will get so pissed you know of us claiming all the inventions that it will make them join us eventually and it's true that you know it's kind of far because some people they want to be in the place that is the cool place where the cool stuff happens and it's not only about compensation and so on now it's I would say everybody working in AI and doing a good job in AI is going to get good economic outcomes so then what can you you know glory is also very important and it can be scientific or it can be just being proud that you are making the best products.

31:22But I think it's also an important part. Great. Now Gradium is an actual company that was launched a few months ago and obviously you are a new entrant in the field of AI where there's tremendous amounts of competition. So I think for voice in particular, like the obvious question is why has OpenAI or Google or Meta not already won voice AI? and I think you probably alluded to some of the reasons up front, but why is that? Why can a small company hope to become the leader? So one thing I mentioned was if you have the right team, it can be extremely small and still make a significant impact. Other arguments I think is one is focus.

32:07So for example, if you look at large multimodal models, right? Like these generic models that understand images and can generate text and can produce code and so on. You have like a limited budget, which is the number of parameters and data you're going to feed to your model. When you want to add speech to them, you're fighting with coding and image understanding and so on. So you are playing with a lot of trade-offs that are irrelevant to the tasks that you want to solve. And at the same time, these models are so large, they cannot run at scale because they will just make everyone lose money in the process.

32:43The only format that makes sense for speech models to run at scale is to be extremely compact, which also means that the training resources you need to train them are much smaller than what you need to train other kinds of models. So the resources are not as challenging as for text models. I think also another aspect is, in a way, not trying to just make a conversational product. So really making building blocks so that people can build the product. So we could make the Gradium conversational assistant and think a lot about its capabilities and what it can do and what it cannot do and what will be the use cases for people and so on and hope that like the voice mode of OpenAI it's used for this task and this task and so on.

33:30But it's impossible to cover everything with that. And so now if we want to cover NPCs, fake sport commenters in video game and a language learning app and an annoying character in a cartoon and the customer center agent and so on. Then we just make the building blocks. And this, again, is not really, I think, in the DNA of big companies to do this kind of very specific models that are targeted towards developers rather than trying to solve a lot of things at the same time. And then there's an additional aspect to this, which is that voice can also, and should be pretty often on device versus an API call to the cloud.

34:11Is that fair? I think what is very challenging right now is if you want to have the full intelligence on device, like your full conversational AI on device. Honestly, I would say at this point, if you want such a model to be useful, we are not there yet, right? You can have something that can chit-chat a bit and it will be decent. Or we also have shipped models on device, but they are much more constrained in terms of applications. So, for example, we started a year ago with on-device speech-to-speech translation, which is something that makes a lot of sense. because when you're traveling, maybe, you know, you don't have a data plan that is going in every country.

34:45So it makes sense to have something that works on your phone if you want to order at a restaurant, something like that. I think it's a particularly adapted use case. But now we also, we released two weeks ago a model called Pocket ETS that not only is on device, but CPU only. So there are already mods for AAA video games where the NPCs can be powered through these voice models. And now you unlock a completely new kind of applications because on-device models allow to do very large-scale personalized content that will be economically not realistic with an API. So again, these kind of things is, you know, if you want to make meaningful progress in that direction, making small models in voice is much more difficult than making large models in voice.

35:31So keeping the quality while reducing the size of the model, that's where the big challenge is. And our CPU model, in terms of algorithms it's like the cutting edge of what we know. You know, it's really the later generation of everything we've been doing so far. So just a few days ago, there was this big announcement by Alibaba slash Qwenn that they were open sourcing the Qwenn 3 TTS family for voice design, clone, and generation. How do you think about open source in your world as Gradium? Is that a friend? Is that a foe? So if you read the paper, you will find our names in several pages. It's mostly inspired from the Moshi architecture, like pretty much every model right now.

36:18Even the VoxTral model that was released by Mistral two weeks ago is also based on our framework. I think that's really interesting because this proves that there are things that we do right because everybody is building on them. At the same time, I would say it's quite an advantage because I would say there are two kinds of research papers. There are research papers that are meant to be as explicit and reproducible as possible, which is what we try to do when we do one. And there are some that are more about, I would say, marketing in the sense that they are mostly focused on the results and the performance rather than explaining the underlying mechanism and the data and so on and so forth.

37:02The nice thing of people building around the frameworks we introduce is that even when they don't give details, we can infer the details. So in a way, in a competitive landscape, I think it's quite an advantage because in a way it would be more challenging if people were transitioning to something completely different from what we've been doing because, you know, then we could not infer anything when reading their papers. Now all of this is very familiar when we read those. And at the same time, we have all the issues that people are facing. We have been facing them for a while and so we have already workarounds and new versions and so on.

37:35Open source, at Qtai, that's like the end goal of the lab. At Gradium, that's not the end goal of the lab. The end goal of the lab is to make, of the company is to make competitive products that outperform every alternatives. But open source is, I think, is a good way of allowing developers to prototype stuff, understanding what people expect, what they want. You know, it's also a way to train talent. A few days ago, we released Hibiki Zero, so that Hibiki was our speech-to-speech translation system. Now it's a new algorithm that makes it even lower latency, better voice cloning, multilingual, and so on and so forth.

38:11And the PhD student who is working on that project at Qtai is joining Gradient for a few months to do a visiting PhD, and then he will start his PhD again. So I think there are very healthy relations between open source and closed source that can be done. At the same time, I don't think it hurts any defensibility quite the opposite, honestly. And so I think it's also for talent, it's very attractive because people, when they join us, they know that they can work on really competitive products, but at the same time, they can keep sharing models and more exploratory research and do really frontier stuff.

38:51And I think if we want to stay at the cutting edge, we should be a product company, but also be a frontier lab. and being a frontier lab, you need to do fundamental research. What's the gap between open source and closed source in voice? Obviously in text, there is like this cat and mouse game and it seems that the commercial labs are constantly like pretty far ahead of the open source. Is that the same thing in voice AI? So what is interesting in voice AI is now I think a high school student could make something that is decent. But then what people want is the last mile. And the last mile is extremely difficult.

39:33And the last mile is pronouncing well all the difficult cases. He's having a latency that not only is low, but is robust and almost zero downtime. And, you know, being able to clone voices regardless of accents and so on. So all these hard cases, for me, the only way to find the energy to solve them is because that's your business. because otherwise if you look at benchmarks like the benchmarks for TTS it's a Libri speech so it's books a lot of them from Dostoevsky so it's more about whether you are going to miss an I or a Y in a Russian name or it doesn't evaluate can you pronounce phone numbers and email addresses and URLs and all this stuff so when you optimize for these benchmarks you lose a lot of the actual real world cases so I think the main difference is the incentive to do things really with the finest details only makes sense for when you want to be competitive in product.

40:33And it's not only about the models, right? The infrastructure to run these models at scale is extremely important. I think particularly in an API business model, your margins mostly depend from the efficiency of your inference. And so in particular, I'm very lucky to have in my team one of my co-founders, Laurent, who did several years in, you know, he did all his career in quantitative finance, mostly at Jane Street. And now we have more people coming from quantitative finance. And these people, they are really passionate about efficient inference. That's their bread and butter, you know, because in financial, like in high frequency trading, that's kind of the only way to exist.

41:11And the engineering challenges are really significant. And I think a team in engineering as well, and not just like training models, makes a huge difference. And now, if we're talking then about on-device inference, that's even more complex, right? If you want a model to run on all Android and iPhone and so on, you know, that's a big engineering challenge. You mentioned benchmarks a minute ago, and that seems to be a really interesting question for voice and video as well and images. How do you make a case that your technology is better than the next provider? Because some of it seems to be a little bit around vibes, right?

41:53Like how you feel when you're on the receiving end of a voice AI. One thing that is clear is that you can only trust human judgment. People have tried to make objective proxies of human judgment. Like that would be a neural network that listens to an audio and gives it a grade. it sucks like so many people try and it works on their constrained setting and on real audio it doesn't work at all and completely breaks so we don't trust anything about you know but our ears so we do a lot of blind tests internally we do a lot of blind tests externally so we are working with human judgment constantly so every single decision we make is based on human listening we don't trust metrics at all so it's fundamentally subjective experience the quality of audio, but there are some things that are going to be widely shared.

42:46Reserves like the prosody, so the tone and the rhythm are natural or not. A lot of people would agree on that. Is a voice nice or not? Nobody agrees on that. And so then the only way for me to claim to have the best solution is to have the largest catalog and most diverse set of voices that people can pick. because then, you know, it's the kind of what voice are people going to like. This really depends on between people. I had faced that in the past when we did Music LM at Google. So for the first time, we made, you know, text to music. So you could type like death metal with marimba or whatever, and you get death metal with marimba.

43:28And so we put like a website online where people could do this stuff. And at the same time, we were already planning to do for the first time RLHF, So reinforcement learning from human feedback on music. And so what we had designed is people could give, so there will be two generations every time and people could give a trophy. I had planned that because, you know, I wanted for the first time to have integrate humans in the loop of music generation, because I think, you know, scientifically that will be really huge. And what was interesting is that, so we made a paper about that called Music RL and it was quite nice, but the results were not extremely convincing.

44:08And then what we did is that we just did judgment ourselves. So, you know, with people, with my colleagues, we took like, I don't know, 10 or 20 pairs of audio and each time we choose our preferred one and there was zero agreement between people. So there is no way, you know, like an algorithm is going to learn human preferences when there is so much subjectivity. So the only thing you can make is make it more steerable for each user to be able to customize it because there is nothing that is going to please every user consistently. The obvious next question is if nobody can agree on whether this model is better than that model, then aren't all the models more or less the same and therefore the entire voice AI model industry is sort of like commoditized?

44:56All the models being the same. I think then people talk about TTS already, right? which is much more constrained than voice AI. So in text-to-speech, I mean, there are factual metrics about accuracy and latency. And then there are more subjective things around expressivity and so on. But you can, again, you can really make a difference by making it more controllable, more customizable, and so on and so forth. So I don't think that it's clear that the best TTS, the most controllable, the most robust, the most smart in terms of expression is in front of us. Nothing is close to it yet. And then there is everything that is not TTS.

45:38You look at transcription. Now, again, what if a lot of people are speaking? So I was talking about diarization as the least sexy problem ever. At the same time, it's an extremely useful problem. And you look at the error rates and they are very bad. I mean, it's just not working. In difficult cases where you have a podcast with a lot of people talking at the same time, it just completely breaks. Full duplex, we did Moshi a year and a half ago. Still nobody has made it into a product. There are so many things that are just not existing today. I think the communitization, maybe it will happen someday, but we are very, very far from it, honestly.

46:16And the gap in, you know, there is already bridging the gap in quality and abilities, you know, to make that as powerful as human production of speech and understanding of speech in complex environments and so on. And then even if you reach this point, there will always be challenging about getting the same performance with smaller models. Because then, again, like a full duplex, smart voice AI that can run on a gamer GPU or on an iPhone. Good luck, you know, like that. Okay, awesome. I'd love to spend a little bit of time on the technical aspects of voice AI. So we talked about some of them, alluded to others, but I think it'd be nice to bring everything together.

47:02So in terms of the fundamental innovation that you described around speech-to-speech models, can you for us compare and contrast like the old way, which I believe is the cascade versus this new generation of models? How does that work? The cascaded system, which is the old way, but is actually pretty new. So like the old way is without LLMs, you know. So we talk, you know, the old LLMs. So the old way is like Alex, Nancy, and so on. Back two years ago, back in the day, two years ago. The old way is natural language understanding with parsing and with very limited vocabularies and so on. But the cascaded idea is you basically just take an LLM and you wrap it into a speech input and speech output.

47:46So you have like streaming transcription and streaming text to speech. The nice thing about that is you can plug any LLM and you can customize its behavior through its prompt, like you would do with a text model. A main limitation is around first latency of the whole thing because you're running three models in a cascade and each of them is going to add to the overall latency of the interaction. And by going through the bottleneck of text, you lose what we call paralinguistic information, which is all the information we convey when we speak on top of what we say. So that's emotional states, irony, lying is one.

48:20You know, a lot of information is conveyed that is not in what to say. In particular, I encourage people to look at every few years, InterSpeech, one of the main speech conferences, organizes a paralinguistic challenge with challenges that are more random every time. So lying detection is one, but then you have trying to recognize the origin of the parent of someone from them speaking. or there was one about people are speaking while eating and you had to recognize what they are eating. So, you know, anyway, when we speak, there are a lot of information that comes about us and this is lost through...

49:00So what you are eating, obviously, is maybe not the most relevant one, but understanding in the customer care context, understanding that you're getting lost or annoyed and so on, it's extremely important to be able to recover the conversation into a good state. Sometimes it's obvious from the text, like if the AI says go to hell yeah probably they're upset but sometimes it's not that obvious and that kind of cues can be helpful and so that's the only way I mean one way to address all of that all together is speech to speech and full duplex on top of that addresses what I think today is the worst part of voice AI I hate it from the bottom of my heart it's turn taking in any introductory course to conventional networks or image classification there is always this analogy that is made that you can make rules to tell whether an image is shot in the morning or night just based on colors, right?

49:53You can just look at colors, values, and make an handmade rule. But now, if you want to recognize whether it's a cat or a dog, you cannot make rules that recognize all the shapes of cats and dogs and all the angles and so on and so forth. So that's why we do machine learning, right? You just learn from the data when you cannot make the rules. With turn-taking, we are back at the archaic era of handmade rules, which is ridiculous. It's like you have an algorithm called the voice activity detection algorithm that just says whether it's silent or not. And if it's silent more than X hundred of milliseconds and then this rule and this rule and this rule, then it's an interruption.

50:30But if this happens, it's not an interruption. And so we have rules on top of rules on top of rules to decide whether the model should talk or not. And that makes it such that, I mean, you know that everybody who is talking with cascaded systems right now, you need discipline when you talk to AI. You need to adapt to its flow. Otherwise, it gets lost and gets confused and interrupts and so on. And this is extremely annoying. Moshi, again, people can talk to it today. It's not really powerful, but the quality of the interaction is unmatched in the sense that you can really take five seconds to think about what you're going to say and then talk.

51:07And the model won't be lost. they will just be extremely flexible. And what we did is it was an extremely simple thing. So we took like the audio language model. Instead of having it modeling one stream of tokens, we called it multi-stream, just two streams of tokens. One for the user, one for the AI. And both can be active at the same time. And there is no turn-taking anymore. And you just train it on stereo data. People talking on the phone and you have like one person on the left channel and one on the right. And your model is both at the same time. And then you play the role of left channel and the AI plays the role of right channel.

51:42And the model will learn how to handle a conversation like humans on the phone. And so when we did the motion announcement, in particular, there was a fun early demo that we did. At some point, we trained the model on the Fisher dataset. So it's a phone conversation that we are recording in the US in the late 90s. And so we could literally give a phone call to the late 90s. And it will mention like Saddam Hussein and Jacques Chirac and talk about, you know, like a lot of political stuff. It was very weird because you're talking to a guy and say, hey, I'm Bob from Arizona. And every time it's a new guy and you can talk about whatever and they tell you about their job and so on.

52:16So it's a kind of paranormal experience. Very fun. I think eventually it's impossible to think that we will stick to turn-taking in the future. So obviously, for us, the end goal is full duplex. That was always our motivation. At the moment, what we do is cascaded systems because that's where the market is right now. It's people that are still iterating a lot on the underlying text models that they want to use, on the tool use and so on. And so there is so much progress on the text side. One drawback of speech-to-speech models is that since everything is integrated, when you go from a text model to the speech-to-speech model, you need to fine-tune it on speech data.

52:57So now it's the cost to switch the underlying text model is extremely high because you will need to refine-tune everything from scratch. People want something that is modular plug-and-play. What will solve everything is providing the same flexibility as cascaded systems so that you can change the backend on demand, basically, and you get the same customizability and customization that you have with text models, but with the full duplex. And this, I think, is going to be a convincing solution to all the current limitations. So that's the frontier of Voice AI? one of the frontiers I would say again one where I would say a frontier is I could like bet to every single speech team in the world that they don't solve it in the next year or so it's a robot in the model in the factory and there is a lot of noise from machines and you have a lot of people talking to the robot and the robot has to figure out what the hell is going on this I can you know I can I can sign that it's this is extremely challenging if if I had to pick the worst topic I mean the one that will be the most challenging I think that will be this one.

54:02Even more than full duplex. Much more difficult, I think. And what happens behind the scenes when you have that noisy environment? What does the model actually do? And how do you solve that problem? If you're calling a receptionist in a restaurant and he or she picks up and there's super light in the back, how do you solve that problem? One thing that is really important is to have several microphones. So we do. We have two ears and it's extremely useful because that's the lowest, for example, to localize an audio source. So the reason why we can localize where our sound comes from, because we have a small time difference between the time when it arrives in this year and this year.

54:46So if it arrives here before, my brain will understand it will come this and you get like a longer. And then the acceleration of the phase gives you the other dimension. And so that's how you can locate in 3D. So having the ability of doing this specialization, understanding where the sources are coming from and which one is saying what and so on. Again, it's both a hardware and software issue. Creating training data for that is extremely challenging. Honestly, the level of robustness that we have as humans on that is extremely high. Also, we are all lipreading. We don't know that we are doing it, but everybody is lipreading all the time.

55:25That helps understanding a lot. in noisy environments in particular. We do it unconsciously, but we all do it. I think interactions on the phone are quite okay because, you know, the mouth is close to the phone and the phone can do a lot of work to enhance the quality. But now, you know, people want to have a robot that can be even a static robot that is just in a room and you want to shout to it from your living room. It's called far-filled speech recognition. It's really broken. Every company who has made these home assistants know how difficult it is to have them work in environments where there are several people.

55:59The reason I say I think it's one of the biggest challenges is in TTS, we see a lot of progress. For this kind of hard understanding problem, I hear people seeing the exact same stuff as they did 10 years ago. Like, there was zero progress. You mentioned data a second ago. How does that work for voice AI? And how does that compare to text AI? Obviously, text AI, the LLMs are training on the whole internet, but presumably there's a lot less audio and speech data to train on. How does that work? So if you do the math, basically like training on a few trillions of tokens, which is what you will do for a basic text model, that will amount to hundreds of millions of hours of speech or something like that, which is kind of amounts that are very hard to get.

56:46I think it's a very interesting question that comes up in a lot of discussions. And everybody has their theories. In particular, one impact, one, let's say, one attribute of speech data was that if you train a conversational model on speech data, it's going to be much less intelligent than a model trained on text. And I think it's because when you listen to speech data, the density of information is much lower than you will have in text. So you don't have Wikipedia like, you know, in speech data. You don't have Stack Overflow, Reddit and so on. Getting your model to learn about the world from speech, I think it's a terrible idea.

57:21I think you should start. I mean, you know, we have text models. So what we did for Moshi is we started from the text model. And then we took this text model and trained it on speech while trying to prevent as much as possible a loss of intelligence. So all the time we recompute the text metrics and they will degrade, but we are trying, you know, like to keep it a bit contained. But indeed, the quality of speech data, extreme variable. Honestly, I think we are training on way too much data. So for Moshi, we train on 7 million hours of speech. It's ridiculous. I mean, we could probably do that with 10 ,000 hours if we had the right method.

57:55We do not care because we had the data. So like, yes, we just trained on that, which is great to capture all the very specific idiosyncrasies in voice, the French accent, like the things that make a voice unique. So for the diversity of voice, the amount of data is extremely useful. But at the same time, you could think, okay, if my model is intelligent as a text model, it could learn dynamics of conversation from a few thousands or dozens of thousands of hours of data and not millions of hours of data. That's one source of data, the largest volume. It's just, you know, unstructured conversation, like publicly available, audiobooks and this kind of stuff.

58:30And then you get towards more specific data that you make for your applications. For example, for TTS, you want to have expressive data of high quality. You don't want to have something that is recorded with arbitrary conditions. You want to read studio recording, very low level of noise, professional or semi-professional actors. And then there is what kind of text do you give them to read? And so it's a very interesting problem because if you want to generate hundreds of thousands of hours of scripts for voice actors, it's not that easy. I mean, you cannot ask Claude or ChatGPT, write 100 ,000 hours of scripts and make them as diverse as possible.

59:15So it doesn't work. It's going just to be in a loop and collapse on a few topics, you know. Even if you ask for phone numbers, they are not really random number generators. They will produce eventually like always the same phone numbers. So Alex and my team spent a lot of time on making very complex machines for script generation with like a taxonomy of all possible topics and subtopics and subsubtopics and subsubtopics and sampling constantly. And every time we generate a phone number, it's generated with an actual random number generator so that we actually cover the whole scope of phone number.

59:51So I think it's a very interesting problem. It's painful. It's annoying. It's really about details. And I love that because I think that's always where we can differentiate. It's just because we always, as gamers say, we tank the painful stuff. And this pays off a lot because as I said, in terms of computational resources, we don't need as much as to train large reasoning models or video models. But for speech, it's not really about the volume of data, but having high quality data is extremely useful and very important. And the accuracy, for example, of annotations, it's paramount and it's a lot of human labor to be able to annotate precisely.

1:00:27There is a labeling aspect to do this. Yes, and the labeling aspect is interesting because automatic annotation works very well. And where humans are useful is really for the few mistakes that the best automatic annotations will do right now. And so the only way to make it worth the time of humans is if they have perfect annotation. If they have slightly imperfect annotation, that will be worth it. How does it work if you want to add a new language? Especially, not to pick on any, but like, I don't know, or do like Afrikaans, like something where there's less data on the internet, presumably. Do you have to hire actors to, you know, create scripts specifically for them?

1:01:10How does it work? What is interesting is that we see transfer between languages, in particular languages from the same family. For example, we have Hibiki Zero that released last week. So it was trained on 50 ,000 hours or maybe 100 ,000 hours for, for example, Portuguese and Spanish. and then to do Italian translation to English, 1000 hours was enough. I would say for languages that belong to a completely different family, that's more challenging. What is nice is that fundamentally the whole pipeline is the same. Right now, we have been focusing on a few languages to find the recipe that works and then it can be applied to most languages.

1:01:52What is very difficult is unwritten languages. So languages that don't have an official writing system. That's a lot of dialects where people are mixing languages and putting like a bit of French and then a bit of Creole and then a bit of something else and so on. This is very challenging because getting data is difficult. Getting annotated data is even more difficult because you have much fewer annotators that can do that. But these ones are the most challenging, I would say. But it has been a challenge for a long time. and there are a lot of projects around collecting data in such language. Interestingly, there is one.

1:02:31So, for example, the content that you can find in the most languages is the Bible. And so, like, the main source for, like, the most widely available content in the world is the Bible in all languages. Because a lot of people spend time into translating it and so on. So, that's a very valuable source of data. Obviously, you won't have a mention of modern vocabulary in it. and what about a hardware so you mentioned earlier this is not a big hardware gpu play compared to the big lms is that is that correct if you want to have voice models to run at scale and make sense from a you know in terms of economics you need to have this model compact anyway if you think about having npcs in the game where people are going to pay i don't know 70 bucks or 90 bucks or whatever for a game.

1:03:21And then they want to talk to it three hours a day because it's a very good game. So you spend a lot of time on it. If you have a large model, it's just impossible. Not only it doesn't fit on the GPU, but even for APIs, it just wouldn't make sense economically. So fundamentally, I think these models need to be small, but sometimes they require access to large models because they are solving a complex task. I'm quite, in a way, stating the obvious, but selective and adaptive compute usage based on the context and the difficulty of the task at hand, the task being as precise as like the next few words.

1:03:52Can I just answer like that? Because somebody say, hey, and I just say, hey, they ask for like the next flight to SF and I have to look up on internet to find the time. So being very selective about when to use compute, I think that's the only way for all of it to make sense economically. Let's talk a little bit about the product and business aspect of Voice AI you mentioned that you were building a model company, but also a product company. What is a product in voice AI? And the two things I'm thinking of are like one, cloning and then agents. That's the two things that people are talking about a lot.

1:04:33Take whichever one you want. Yeah. So I think for this one, so far as the product is the underlying models, right? Because what we see is that people are building voice agents. So voice agents, it sounds, I think, a lot like a business agent, like a customer care agent or something like that. But an NPC in a way is an agent, right? Because they have a voice interface, they have an underlying text model. Eventually in video games, they will be able to do function calling to launch a quest, to decide that you solved something or to control an action and so on and so forth. What we focus on is providing the best technology to the people who want to build the agent.

1:05:10So we don't build agents ourselves. We want to focus on the quality of the model because that's our specialty. That's where our talent is. And at the same time, I think it will be a bit unrealistic to think that we can address all the needs of the market with our agents. Because again, I think what is very exciting is when we look at our customers, every time there is a new one, it's completely unpredictable what it will be about. And it's about learning and customer care and video games and media. and personalized press. And it's so wide that we'd rather, you know, just make it easy for people to build voice agents and provide them with a reliable models and infrastructure.

1:05:57And that also works pretty well because in terms of staff, we are less than 15 employees today at Gradium. When we partner with a company that deploys voice agents in banks or hospitals or different kinds of businesses, it requires a lot of what we could call today a forward deployment engineer. So there is a lot of staff, human labor that is involved for these deployments into various complex systems and infrastructures. And we can just focus on the model. So that also allows us to remain pretty small and really focus on the science and the engineering. And voice cloning, is that a use case? Actually, at Gradium, we have the best of the industry.

1:06:38and the best means not only replicating like the specific characteristic of someone but I mean also the accent, some unusual recording condition. And so in a lot of contexts, for example, if you want to create a vintage sounding character with old radio effects and it's something that we do pretty well. If you want to have a robot voice, we can do that pretty well as well. If you want to have any kind of accent or speaking style like posh or more laid back or more urban. Any kind of social aspects and individual aspects of voice is something that we replicate pretty well. I think what is interesting is that voice learning in itself, where I see the most potential, is about creating interactive experiences around licenses.

1:07:27I tried to pitch, for example, who wants to be a millionaire? There is a video game and the questions can be generated on the fly. You would like the voice of the host to pull on them all the time. So this one makes a lot of sense with cloning a specific voice because you want to replicate the voice of a character, of a person, of an athlete, of a K-pop star, whatever. I don't know, you know, any kind of experience where people want to engage with a voice that they know, but they want to go through an interactive experience. Then in a lot of cases, for example, in customer care, people don't want a specific voice.

1:07:56They want a voice with specific characteristics, and that is nice for the use case. And so typically they clone the voice of a colleague that has a good voice. For these customers, I think the solution that makes the most sense is not cloning, it's voice design. So being able to create voices from a natural language description and having your voice generation being really faithful to the prompt so that it's really worth iterating on it. Because there are a few solutions that exist today around voice design, but as far as I know, they are not very popular and people still stick to the existing voice catalog because they cannot steer it precisely enough.

1:08:33But when this works, this will allow people to design the voices that they want for their use case, which I think is very interesting. Even though sometimes I'm a bit surprised because we have some customers that just take a random stock voice from our catalog, which is one of ours, like me or my colleague, and they don't care. And I'm like, I try to put them like, let's design a voice together. Let's make a voice that represents your brand. But some people are like, no, it's fine. Fascinating. Or maybe you just have such a great voice. That's what the market wants. inevitable question around privacy and then deep fakes, obviously in a context where one can clone somebody else's voice pretty easily.

1:09:12What does that mean and how do you protect all of us? So first thing I want to say, watermarking is a scam. I'm sorry, I have to say it. It just doesn't work. I worked on it. We have an appendix in the Moshi paper around how we could break so easily any watermarking that was supposed to be state of the art. So people should not rely on that. Also, watermarking, who verifies the watermark, right? You will need a platform to do the verification and remove the audio. But if someone, you know... Maybe explain what watermarking is in the context of audio. Yeah, in audio, that's the idea of either finding a way to...

1:09:48So when you generate an audio, you put a hidden stamp in it, like you would do in an image, but now it's in audio, so you cannot hear it. But it makes it very recognizable that it's not only fake, but it's fake from this specific model. and also there is deep fake detection which is just recognizing that an audio is synthetic even if there is no watermark in it just telling true from fake these are very difficult to do unfortunately you know if you get a phone call from a scammer unless you have inside your phone's automatic detection it's useless right you cannot upload it to a website that says it's true or fake honestly I would say at this point I don't know who has their grandma who lost their credit card and it has the airport and it's$1 ,000 today.

1:10:32Because it's always a story I hear. It's probably not the case if someone gives you a phone call. So I would say just in that context, honestly, I think there is nothing as safe as just asking personal questions that only the person that they pretend to be could answer. So you think it's going to be a fact of life and therefore we need it all to be just more vigilant? Yeah. I mean, you know, it was already the case with emails, right? Yeah. And people just need to be much more vigilant and probably will find ways to have authentication, for example, on the other phone side that it's an actual person that is calling and so on and so forth.

1:11:13But on the privacy front, if I clone my voice with Gradium, then you keep control of it? Only you can use it. You need to own it. but nobody will be able to use it. And so it's only for your own usage. And I think what eventually we could also do, which was done before, is allow people, if they want, to opt in to share their voice with the community so that they can also get financial compensation if their voice is used by other users. So I think this one has a lot of value because it's a good way of sourcing a lot of voices. Eventually, again, if we omit the case of replicating familiar voices from licenses or existing people who give their voice knowingly to create specific content, I think voice design is going to just remove this issue.

1:12:11Because then, again, people typically are going to clone the voice of someone, but what they wanted is someone from a specific gender, specific demographics, age, accent, and so on. And so they could just fill this information and get a few propositions for voices that will fit their need. And that will remove the need for voice cloning. So as we get towards the end of this conversation, just like a few sort of quick ones. One, there seems to be an emerging discussion about the intersection between voice and screens or image. Is that something that you focus on? There are several things in particular.

1:12:48So if you want to do speech understanding, having access to the video can help a lot. Again, for diarization, for example. If you're filming, like if you're listening to presidential debates with several candidates, it's awful to understand. If you watch, you can see who is speaking at each time and so on. So audiovisual understanding is much stronger than audio understanding in itself. I think now also what is interesting is, so we see that in a video generation, like with VO3 from Google, Now there is native audio that is included. And that's why I think it's for video generation, I think the most natural thing is integration of video and audio.

1:13:28I don't think it really makes sense to do video to audio generation separately because typically the data exists as a multimodal signal, right? When you have a video, you have the audio track as well. So you might as well just exploit both to train your models. So, however, now, so for example, we released for Valentine's Day a small app called Bridger Clone. Now, if I put my face in a photo, I put to video, I say make a video. It's going to make up a voice that it thinks sounds like me based on my appearance, my age and so on. Where do people find that? It's called Bridger Clone. Bridger Clone, like Bridgerton.

1:14:08Yeah, exactly. Bridger Clone.app. Yeah. And so what it allows you to do is now you can clone your voice. You can record a small, short love message. And now you get a video of you in your voice, you know. And this again, I think that's why it's useful to have a voice that is treated as kind of a separate component. Because now you can have much more control on the actual voice of the virtual character, rather than, you know, just having a likely voice that sounds like it could be yours. Yep. Okay. So we're talking about cloning. Obviously, cloning is just one of the many apps. Ultimately, just to play it back and drive it home, you provide building blocks and models to create all sorts of different products based on voice, whether that's a customer service use case or any kind of interaction, translation.

1:15:02and what is interesting is so at Qtai in two years we were able to do conversation, translation TTS, speech to text, always competing with much bigger and much more mature teams. It's still the same at Gradium and one of the big strengths is that we have kind of this fundamental framework for audio generation around audio language models and it's extremely flexible so I think one of our strengths as well is our ability to do like a new task and when people come up with new needs, whether it's about annotating various stuff or generating various stuff, all of this can be cast pretty easily in our framework.

1:15:38So then if we see that there is huge demand for speech separation, it's pretty easy for us to do it for voice transformation, for accent transformation, whatever, you know, that's also why we are always interacting with the developers to understand what they want. We are also now giving access to alpha models that can do, you know, stuff that are still experimental but our world premieres. And yeah, I think that's something I'm pretty excited because we can be much more creative than just speech to text and TTS. Obviously, these are kind of the master tasks where we want to be the best in the world because that's where most of the opportunities are.

1:16:16But at the same time, we can do a lot of fun stuff that is completely orthogonal to that. In particular around transformation, audio effects, and so on. I think that can be... There are a lot of things too that can be very cool. Good. And last theme or question, you're building this company out of Paris. Yes. Obviously, you're building the company very much in a global way and you have multiple customers in different geographies, including very much the U.S. We're recording this today in New York. You're flying out to San Francisco in a few hours. Any thoughts on the current state of the French AI scene and the European, I guess, AI scene?

1:16:59You know, there's always this fun back and forth, you know, scene from the U.S., combination of like occasionally admiration, pretty often mockery. mockery and the fact that Macron, whoever runs his social account, sort of misfired the other day by saying that he was going to allocate 30 million euros to AI when in reality he meant a specific program to attract a few academics to France. But so what on the ground there, what is your sense of the current state and the strengths and weaknesses of France? A lot of things to say about that. I was born and raised in Paris. I did all my career there.

1:17:40Facebook arrived when I started my PhD and then Google Brain moved to Paris and Google DeepMind afterwards. I'm what we can call terminally online in the sense that I love the very mean memes against Europe. So there's this guy, I don't remember his name, it's like a fake Swedish name. And he keeps posting about how, you know, he has, after only 20 meetings, he has contributed a 10 ,000 euros check for a C-ground. And it's a compliance first company and everything. But I love it. It's very mean. And honestly, I love it because it's very mean. I'm like, I love when people have so much time to spend just to be mean.

1:18:14I mean, I think it's quite, you know, I respect that. But at the same time, it's so far from reality. So, you know, French AI and European AI. So European AI is mostly French AI, to be fair. There is also Germany, but a lot of it is in France. And before French companies, as there was French talent in American companies. So again, a lot of the current audio generative models of Google, a global company, obviously, were developed between Paris and Zurich. Actually, most of it. Lama was started in Paris. Dino, which is the most groundbreaking vision work from Facebook, was developed in Paris. A lot of things have been developed in Paris that are not seen as Parisian because they were made in American companies.

1:19:06And now I think the field in Paris is, like the talent is so dense and the people are extremely strong and extremely committed. And the best signal that proves it is we used to have Facebook and Google in Paris. Now there is OpenAI, there is Anthropic, there is Coher. Pretty much everybody is opening an AI lab in Paris. And the reason why they do is for the talent. So I think we have everything in France to develop global companies. In particular, you know, in AI, which is an economy really built around talent, that's a perfect deal for it. Also, as a French guy, I really don't want us to screw it up, you know, in France.

1:19:51Because in a way, you know, when you see like the most successful French companies, it's luxury or that kind of stuff. and so yeah I really want AI to be you know that's a field where France can make a big difference and that can become one of the most biggest drivers in the European economy if it doesn't work then Europe will really have to look itself in the mirror and be like how could we screw it up with so many strong people because you know the people are here and you know capital we can get capital in Europe as well I think all the conditions are there for it to be competitive in a way the more people are mocking Europe, the more it can make the people who mock overconfident about themselves and everyone who is overconfident eventually get displaced by the underdog.

1:20:37So, you know, I mean, people should always, you know, like when I was the 996 like LARPing on Twitter, it's ridiculous. Like people coding at the gym, like what the hell, like code and then go to the gym. You're not doing a good gym and you're not doing good code. So you're just, it's just pretending. So I think we don't have a culture of pretending and that's whole Europe from the most Western to the most Eastern part. We tend to be a bit more pessimistic maybe and a bit more down to earth about our impact and the challenges and so on. We don't try to sugarcoat things. We rarely look enthusiastic or happy about our work, but you know, it's a good discipline.

1:21:18So I think the results kind of speak for themselves. What I think, you know, so Mistral, for example, was, you know, sometimes it's mocked because it's not considered the frontier lab like OpenAI, Entropic, whatever. Okay, but look at the staff and the resources and so on. I know the best people from Mistral, they are, you know, they can be compared to the top of the top of the biggest labs. So then there's a question of scale and scale of resources, indeed. So what you, the unfortunate tweet you mentioned about the 30 million euros, you know, is one of them. But yeah, I mean, other than that, people are, there are a lot of very strong people.

1:21:56And, you know, they have done also a lot of great things for American companies. I mean, Jan Leukien is one of them, obviously, but it's not only him. Sammy Bengio, who used to lead Brain, is now leading Apple MLR. The amount of French people in the leadership of big tech AI research is very large. Wonderful. Well, I love that. Yeah, I get carried a bit when I speak about France. I love the fighting spirit and this is a wonderful way to end it. So, Neil, thank you so much. This was terrific. Really appreciate it. Thanks. Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast.

1:22:33If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you on the next episode.

From the publisher

Voice used to be AI’s forgotten modality — awkward, slow, and fragile. Now it’s everywhere. In this reference episode on all things Voice AI, Matt Turck sits down with Neil Zeghidour, a top AI researcher and CEO of Gradium AI (ex-DeepMind/Google, Meta, Kyutai), to cover voice agents, speech-to-speech models, full-duplex conversation, on-device voice, and voice cloning.

We unpack what actually changed under the hood — why voice is finally starting to feel natural, and why it may become the default interface for a new generation of AI assistants and devices.

Neil breaks down today’s dominant “cascaded” voice stack — speech recognition into a text model, then text-to-speech back out — and why it’s popular: it’s modular and easy to customize. But he argues it has two key downsides: chaining models adds latency, and forcing everything through text strips out paralinguistic signals like tone, stress, and emotion. The next wave, he suggests, is combining cascade-like flexibility with the more natural feel of speech-to-speech and full-duplex conversation.

We go deep on full-duplex interaction (ending awkward turn-taking), the hardest unsolved problems (noisy real-world environments and multi-speaker chaos), and the realities of deploying voice at scale — including why models must be compact and when on-device voice is the right approach.

Finally, we tackle voice cloning: where it’s genuinely useful, what it means for deepfakes and privacy, and why watermarking isn’t a silver bullet.

If you care about voice agents, real-time AI, and the next generation of human-computer interaction, this is the episode to bookmark.


Neil Zeghidour

LinkedIn - https://www.linkedin.com/in/neil-zeghidour-a838aaa7/

X/Twitter - https://x.com/neilzegh


Gradium

Website - https://gradium.ai

X/Twitter - https://x.com/GradiumAI


Matt Turck (Managing Director)

Blog - https://mattturck.com

LinkedIn - https://www.linkedin.com/in/turck/

X/Twitter - https://twitter.com/mattturck


FirstMark

Website - https://firstmark.com

X/Twitter - https://twitter.com/FirstMarkCap


(00:00) Intro

(01:21) Voice AI’s big moment — and why we’re still early

(03:34) Why voice lagged behind text/image/video

(06:06) The convergence era: transformers for every modality

(07:40) Beyond Her: always-on assistants, wake words, voice-first devices

(11:01) Voice vs text: where voice fits (even for coding)

(12:56) Neil’s origin story: from finance to machine learning

(18:35) Neural codecs (SoundStream): compression as the unlock

(22:30) Kyutai: open research, small elite teams, moving fast

(31:32) Why big labs haven’t “won” voice AI4

(34:01) On-device voice: where it works, why compact models matter

(46:37) The last mile: real-world robustness, pronunciation, uptime

(41:35) Benchmarking voice: why metrics fail, how they actually test

(47:03) Cascades vs speech-to-speech: trade-offs + what’s next

(54:05) Hardest frontier: noisy rooms, factories, multi-speaker chaos

(1:00:50) New languages + dialects: what transfers, what doesn’t

(1:02:54 Hardware & compute: why voice isn’t a 10,000-GPU game

(1:07:27) What data do you need to train voice models?

(1:09:02) Deepfakes + privacy: why watermarking isn’t a solution

(1:12:30) Voice + vision: multimodality, screen awareness, video+audio

(1:14:43) Voice cloning vs voice design: where the market goes

(1:16:32) Paris/Europe AI: talent density, underdog energy, what’s next


More from The MAD Podcast with Matt Turck

All 44 episodes
Voice AI’s Big Moment: Why Everything Is Changing Now (ft. Neil Zeghidour, Gradium AI)The MAD Podcast with Matt Turck · 1 h 23 min
Listen in VO