In short
Practical AI Podcast Episode Notes
Episode Overview
- Title: Full-duplex, real-time dialogue with Kyutai
- Description: Discussion with Alexandre Défossez from Kyutai about their real-time speech-to-speech AI assistant, innovations in AI, and the AI ecosystem in France.
Key Guests
- Alexandre Défossez: Co-founder of Kyutai
- Chris Benson: Principal AI Research Engineer
- Daniel Whitenack: CEO at Prediction Guard
Main Topics Discussed
- Introduction to Kyutai
- Foundation: A non-profit lab launched in Paris with funding from notable donors, including Xavier Niel and Eric Schmidt.
- Mission: Focus on open science and open-source research to foster innovation in AI, particularly amid competition from major labs.
- AI Ecosystem in France
- Engineering Culture: Strong emphasis on mathematics and engineering education, leading to a rich talent pool for AI research.
- Independence and Growth: Growing independence from major American tech firms, with a diversification of startups and research initiatives.
- The Moshi Model
- Definition: A speech-based foundation model designed for real-time, full-duplex dialogue.
- Features:
- Full Duplex: Can listen and speak simultaneously, unlike traditional turn-based models.
- Low Latency: Approximately 200 milliseconds response time.
- Historical Context of Speech Models
- Advancements: Progress in real-time speech recognition has become more feasible with the advent of deep learning, particularly with transformer models.
- Transformers: Innovations in audio processing through models that can handle multiple audio streams, allowing for better speech recognition and generation.
- Data and Training Strategies
- Training Data: Combination of proprietary and public datasets, including the Fisher dataset of phone calls.
- Instruct Datasets: Development of specific instructions that cater to conversational dynamics, emphasizing shorter, more interactive responses.
- Future Directions
- Model Development: Interest in creating smaller, more efficient models that maintain high performance while being more accessible for devices.
- Open Science: Commitment to making research methods transparent and accessible for the greater community.
- Competitive Edge
- Agility: As a non-profit, Kyutai can make quick decisions without the constraints often faced by larger corporations.
- On-device Applications: Focus on developing models that can run on devices, providing more privacy and control over AI applications.
Key Takeaways
- Real-time Dialogue: The Moshi model represents a significant leap forward in AI-assisted communication, enabling more natural interactions.
- Open Science Commitment: Kyutai's approach emphasizes transparency in research, promoting collaboration and innovation within the AI community.
- Evolving AI Ecosystem: France is emerging as a notable player in AI research, bolstered by strong educational foundations and a shift towards independence from larger tech ecosystems.
Conclusion The episode provides a deep dive into the innovations from Kyutai, particularly with their Moshi model, while also exploring the broader implications of open science and the evolving AI landscape in France. The insights shared by Alexandre Défossez highlight the importance of collaboration and transparency in advancing AI technologies.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:03Welcome to Practical AI, the podcast that makes artificial intelligence practical, productive, productive, and accessible to all. If you like this show, you will love The Change Log. It's news on Mondays, deep technical interviews on Wednesdays, and on Fridays, an awesome talk show for your weekend enjoyment. Find us by searching for The Change Log wherever you get your podcasts. Thanks to our partners at fly.io. Launch your AI apps in five minutes or less. Learn how at fly.io. What's up, friends? I'm here with Kurt Mackey, co-founder and CEO of Fly. As you know, we love Fly. That is the home of changelog.com.
0:44But Kurt, I want to know how you explain fly to developers. Do you tell them a story first? How do you do it? I kind of change how I explain it based on almost like the generation of developer I'm talking to. So like for me, I built and shipped apps on Heroku, which if you've never used Heroku, is roughly like building and shipping an app on Vercel today. It's just it's 2024 instead of 2008 or whatever. And what frustrated me about doing that was I didn't, I got stuck. You can build and ship a Rails app with a Postgres on Heroku, the same way you can build and ship a Next.js app on Vercell. But as soon as you want to do something interesting, like as soon as you want to, at the time, I think one of the things I ran into is like I wanted to add what used to be like kind of the basis for Elasticsearch.
1:24I want to do full text search in my applications. You kind of hit this wall with something like Heroku where you can't really do that. I think lately we've seen it with people wanting to add LLM's inference stuff to their applications. On Vercell or Heroku or Cloudflare or whoever these days, they've started releasing abstractions that sort of let you do this. But I can't just run the model I'd run locally on these black box platforms that are very specialized. For the people my age, it's always like, oh, Heroku was great, but I outgrew it. And one of the things that I felt like I should be able to do when I was using Heroku was run my app close to people in Tokyo for users that were in Tokyo.
1:59and that was never possible. For modern generation devs, it's a lot more Vercel based. It's a lot like Vercel is great right up until you hit one of their hard line boundaries and then you're kind of stuck. There's the other one, we've had someone within the company, I can't remember the name of this game, but the tagline was like five minutes to start forever to master. It's sort of how we're pitching Fly. It's like you can get an app going in five minutes, but there's so much depth to the platform that you're never going to run out of things you can do with it. So unlike AWS or Heroku or Vercel, which are all great platforms.
2:31The cool thing we love here at ChangeLog most about Fly is that no matter what we want to do on the platform, we have primitives, we have abilities, and we as developers can charge our own mission on Fly. It is a no limits platform built for developers, and we think you should try it out. Go to fly.io to learn more. Launch your app in five minutes. Too easy. Once again, fly.io.
3:15Welcome to another episode of the Practical AI Podcast. This is Daniel Whitenack. I'm CEO at Prediction Guard and joined as always by my co-host, Chris Benson, who is a principal AI research engineer at Joaquin Martin. How are you doing, Chris? I'm doing very well today, Daniel. How's it going? It's going great. I think we talked about this a little bit on the last show, but now we're officially up against the Thanksgiving break. So a couple days off here in the US, which will be nice. Maybe I can catch up on some of the cool AI stuff that I've been meaning to play around with in my spare time.
3:54But one of those cool AI things that definitely made its rounds over here at Prediction Guard and we were talking about was the recent kind of advances in real-time speech assistance. And in particular, this sort of, you know, what OpenAI was doing, but then also what a lab in France called Qtai released. And today I'm really excited because we've We finally got the chance to have Alexandre Defoussé, who is a scientist and co-founder at Qtai with us. Welcome, Alex. Thank you, Daniel. Thank you, Chris, for the invitation. Looking forward to discuss about the details about Moshi. Yeah, yeah, we're excited about it.
4:38And maybe before we do that, if you could give us a little bit of a background on kind of what Qtai is, how it came about. Yes. So Qtai is a non-profit lab that we launched a year ago in Paris. We have funding from three donors. So Xavier Niel, Rodolfo Saadet, and Eric Schmidt. Eric Schmidt is probably the one you know the best. And Xavier Niel is a tech entrepreneur. I mean, entrepreneur now is a successful one. And then Rodolfo Saadet works in logistics. So they gather together to try to fund this effort to bring a kind of independent lab with a mission to do open source research at a time where the open source is maybe suffering a bit from the competition between some of the major labs.
5:35So that's, I think, a big motivation for everyone on the team. and basically we have sufficient capacity to be kind of competitive with big labs. We can't really fight every battles, but as we show with Moshe, we can definitely bring interesting ideas and innovation to the table. Yeah, and I find it, I mean, maybe for those here in the US AI ecosystem, we do see a lot of kind of innovation and interesting things happening in France France and in Paris. I'm wondering, just out of curiosity, what is the ecosystem like there? And how would you, I mean, you seem to be kind of formed out of part of that.
6:21So how has that sort of shaped you? And what is the ecosystem like there? I think the ecosystem kind of starts with the studies with France. There's a very strong engineering culture, also very strong emphasis on mathematics, which I think was like giving a good soil that initially attracted a number of big American players like Facebook that opened. So I think at the time, the Facebook AI lab in Paris was probably the second largest after the Californian one about tie with New York. So I think that kind of says how attractive the city can be because it's not so easy to compete with the attractivity of America.
7:09So now I think what has changed in the recent years is really the kind of independence that's growing from this kind of initial seeding. I think for many years, there weren't a number of truly French organization where you could have access to a sufficient number of GPU, like large enough clusters, so as to develop machine learning model for a number of applications. And that's especially the case with large language models. But there's been a number of events that has kind of led to this diversification of the ecosystem in France. Yes, and so now I guess there's like a number of big startups.
7:50There's like Qtai, And I think that's only going to grow. Also, there's one specificity in France, which I think is very nice, especially for deep learning. And it's the fact that we can do a PhD as a resident in a private company. So for instance, or even like a nonprofit. So at Qtai, we're going to have PhD students at Facebook, where I partially did my PhD. There was also a number of PhD students. And I think it's such a great opportunity to get to use graphic cards so early during our kind of career and even as students. And I think that's very specific to France. And that's also part of the success we're seeing at the moment.
8:35And that I think can only be growing as we train more and more people in such a way. I'm curious as you were describing the ecosystem there in France and how strong it is, what was the specific dynamic that brought about with all these for-profit organizations around you that brought about the desire to have the nonprofit? And how did you find yourself in the middle of that as you were in the formative stages? I think for me, there was a growing wheel to become a bit more independent. I think even though at Meta, for instance, there was a lot of value put on the Paris office, at the same time, an American company always takes decision in its center.
9:20So that would be California. And satellite's office always have to kind of bear the consequences of those, no matter the contribution they will make to the overall value of the lab. So that was kind of the initial desire to be a bit more independent. in terms of the decision making, the ability to lead the research. I got the opportunity. So I was contacted by Nelzeguidur, who was doing his PhD with me at Facebook, at Meta, and then had been at Google doing very successful research there. So he was part of the first team that was contacted, I think, by Xavier Niel. And I think the project was initially very appealing because it's like, you know, same business as usual.
10:09So doing research, what I love the most, having sufficient resources to do it in a completely independent and French environment. So that was, of course, very appealing. I didn't hesitate very long. I guess even at first it seems a little bit too good to be true. But so far, so good. So, yeah. Qtai kind of promotes this idea of open science and, you know, democratization of AI or artificial general intelligence through open science. Some people in our listeners might be familiar with sort of open source, open source AI or even like open access models. How would you define and think about open science as a thing and in particular how that connects to kind of the way in which you envision the building of AI or AGI?
11:04Yes. So I think the two are quite related. Usually the open science comes really around explaining how you arrived at the final results and kind of what are the mistakes you made, what are the things you tried, what was important and what not. So I would say that's like a first part that we've been doing really well with Moshi. We released like a preprint technical report with a lot of details that actually took us a bit of time. And that's something that's not necessarily, I don't think if we were not with this kind of nonprofit mindset, we would dedicate as much time. But I think on the long run, it's kind of important.
11:48and then there are several aspects the open sourcing can go from just the weight to like full training pipelines so releasing more code around the training of touch models is also on our roadmap we didn't get a chance to do it yet because yeah the paper already took us a bit of time and we have other things we're working on but i think that's also part of it like explaining exactly how you got to the final results and not just having a set of weights for one specific task but being kind of stuck with it if you need to adapt it to something else. That's kind of the, I think, the vision of open science.
12:27Could you talk a little bit about kind of what you're able to do with that model that maybe the commercial labs that you have in the same ecosystem aren't able to do? And maybe also kind of is it more standard within other nonprofits around the world that are doing similar things? Or do you guys, you know, is there something very, very distinctive compared to you that maybe other nonprofits that you've seen or maybe even modeled after don't have? Yes, that's a good question. So I'm not necessarily familiar with all the nonprofits in the AI ecosystem. I know the Allen Institute, for instance, is one of them.
13:09I think it's very, there's also the Falcon team, TIA. Yeah, I think we're kind of serving a similar mission. I don't think there is necessarily a big difference. Some of them might be more around like contribution to science, for instance, like general science or core deep learning. I think for us, we are mostly focused on core deep learning. We don't necessarily want to compete, for instance, on the purely text-based LLM space. So there's differences in terms of the choices of the research we're doing. But yeah, fundamentally, I don't think there is a big difference. And then your other question was with respect to like other for-profits.
13:51What do you feel is really in your sweet spot, to put it in another way, you know, compared to these competitors, it's very easy to kind of say, recognizing that all the resources that some of the largest companies in the world have, and they'll put into their labs. But there's definitely a place for others out there. And I think that gets missed a lot by the public. And so given the fact that you have this space that you're playing in, what sets you apart from those commercial in terms of maybe advantages that just having the mass number of GPUs available to them, what are some of those distinct things?
14:29Compared to some of the for-profit, if we take the biggest labs, obviously i guess we have agility that that is not really possible in a like super large company where every action will have consequences in the stock market for instance so the decision process can be really fast that was the case for the release of the model for instance we were able to release it under a commercially friendly license which would be a bit harder in larger structure then i think we have a strong for instance we have a desire to go more and more towards on-device models i think so moshi is kind of barely on device we we demoed it on a macbook pro but it was like top tier macbook pro so it's kind of like proof of concept runs on device not every device but i think we definitely have a value there because a number of for-profit are not going to develop really powerful on-device model because that would be a potential threat to their, like, it's harder to protect in terms of intellectual property.
15:32And I think in general, between the bigger players, there is kind of the race to the very top, very best numbers on like the benchmarks, MMLU and everything. And so, you know, if it takes like 10 times more inference time to beat the other on the benchmarks, they are going to do it because it's either beating the other on the benchmarks or kind of leaving the arena. So we're not really in this mindset. were more like the on-device, I think, could have a very large number of applications. It definitely cannot solve all issues. But I think as a non-profit, we won't have the kind of reservation other for-profit might have for on-device model.
16:26Okay, friends, I'm with a good friend of mine, Avthar Suwathan from Timescale. They're positioning Postgres for everything from IoT, sensors, AI, dev tools, crypto, and finance apps. So after I helped me understand why Timescale feels Postgres is most well positioned to be the database for AI applications. It's the most popular database according to the Stack Overflow Developer Survey. And Postgres, one of the distinguishing characteristics is that it's extensible. And so you can extend it for use cases beyond just relational and transactional data for use cases like time series and analytics.
17:01That's kind of where Timescale the company started, as well as now more recently, Vector Search and Vector Storage, which are super impactful for applications like RAG, recommendation systems, and even AI agents, which we're seeing, you know, more and more of those things today. Yeah, Postgres is super powerful. It's well-loved by developers. I feel like more devs, because they know it, it can enable more developers to become AI developers, AI engineers, and build AI apps. From our side, we think Postgres is really the no-brainer choice. You don't have to manage a different database. You don't have to deal with data synchronization and data isolation because you have like three different systems and three different sources of truth.
17:40And one area where we've done work in is around the performance and scalability. So we've built an extension called PG Vector Scale that enhances the performance and scalability of Postgres so that you can use it with confidence for large scale AI applications like RAG and agents and such. And then also another area is coming back to something that you said, enabling more and more developers to make the jump into building AI applications and become AI engineers using the expertise that they already have. And so that's where we built the PGAI extension that brings LLMs to Postgres to enable things like LLM reasoning on your Postgres data, as well as embedding creation.
18:15And for all those reasons, I think, you know, when you're building an AI application, you don't have to use something new. You can just use Postgres. Well, friends, learn how Timescale is making Postgres powerful. Over 3 million Timescale databases power IoT, sensors, AI, dev tools, crypto, and finance applications. And they do it all on Postgres. Timescale uses Postgres for everything, and now you can too. Learn more at timescale.com. Again, timescale.com.
19:03so alex you you've mentioned moshi a few times now maybe if if you could just give those that haven't heard of this an idea of first what is moshi and then maybe if you could then after that step back and describe, well, how did the lab, how did Qtai start thinking about that sort of model or that sort of research direction as a research direction of the lab? Yes. So Moshi is a speech-based foundation model that also integrates text as a modality. So it's especially built for speech-to-speech dialogue and especially real-time dialogue. So we put a real emphasis on the model being able to act in a way that's the most fluid as possible, like a real conversation with a human being.
19:56And so one of its characteristics is that it's completely full duplex, meaning that the model can both listen and speak at any time. So it's not turn-based, like walkie-talkies, which I think is an important feature when us, we communicate. So we wanted the model to be able to do the same thing. We also, yeah, as I mentioned, And that allows us also to have a very low latency. So we have like around 200 milliseconds between the time the audio leaves your microphone and the time you get a reply that has accounted for that audio. And yeah, at the moment, it's kind of like mostly we designed it as a speech agent with which you can discuss, ask questions, ask for advice that could potentially serve as a basis for a much larger use case.
20:46That's why we also mentioned it as a kind of foundation model and also a framework for a number of tasks that would require kind of reacting to your speech and beyond just being like kind of assistant. And then the second part of the question was, how did we start working on that? So we were two people on the initial team. So Nell and I, two have done most of our research on audio modeling. And then, uh, Eduard Grave had been on the, on the, like a core member of the initial team of LAMA, uh, the very first LAMA at, uh, at Meta. So we kind of had the right tools. So I guess the first reason is like, basically we, we sat together and we're like, what can we do?
21:33And where do we have an edge on the competition? And I think on this aspect of like combining the text knowledge and the, like top of the line audio modeling techniques, we had a real edge compared to other labs. So that was important. And also there was a sense that like speech was becoming an important modality and what had been done in a number of other modalities was still completely lacking. So I was back in November at the time, like OpenAI hadn't made any announcements. So it was still pretty much a new area to cover. So we kind of immediately started working on that. We actually started both on Helium and in parallel, we worked on the MIMI, the codex that we use, with the goal of having a really highly compressed representation at 12.5 hertz to get as close as possible to the text, which would be around like 3 hertz.
22:32Of course, it's not regularly spaced with respect to audio. Yes. And then once we were happy with Mimi, we immediately moved on to the kind of aspect of how do we model the speech? How do we handle the full duplex? How do we instruct the model? Like a number of challenging questions that arised all the way to the kind of first demo, public demo in July. That's great. And just one more kind of background question for those. some people might have seen, I guess, non-real-time agents. So agents that would take in audio, transcribe that, you know, maybe transcribe that with one model, use a language model to generate an answer, and then use a third model maybe to generate speech.
23:19So that's one kind of way to process this pipeline. You're talking about something different here, particularly for these speech-to-speech models or the kind of multiplex models that you're talking about. Could you give a little bit of a background? How long have people sort of been studying this, researching this type of model? And has it really only been possible in sort of recent times to make this kind of real-time speech a reality? Because I think some people are, at least public-wise, they may have seen things like Alexa in the past, right? The process of speech in certain ways, but this sort of demos, at least that they're seeing from OpenAI, demos that they're seeing from Qtai, this is a different type of interaction.
24:07So how long has this sort of been possible? And what is the kind of history of research? I know that's a hard question because there's probably a million things that have been done, but from an overall perspective, how would you view it? So I guess just to put in perspective, so I'm not necessarily entirely familiar with how Alexa works, but it's more, I mean, anything that's kind of pre-GPT model would be kind of rule-based, based on like automatic speech recognition, which is actually a fairly old field. And even real-time speech recognition has been successful for a while, not necessarily with the amount of success we see with deep learning.
24:45I mean, it was already using some of them deep learning before. But then it's kind of rule-based. So if you don't formulate a request in quite the right way, it's quickly going to say, I don't know, or just do a Google search. Then what brought a chain of paradigm was all the GPT model and chat GPT in particular with this ability to perfectly understand human requests, no matter how it is formulated. Then to bring that to the audio domain, what you need is the ability for a kind of language model, like a transformer to process the audio streams. Ideally, you would think it's very easy for a GPT model.
25:21You have text tokens in and you have text token, you predict the next token, and then you just need some special characters to differentiate between the request and the reply. And you want to be able to do something similar with audio, but things are not quite as easy with audio. Audio is not as dense in terms of information. You can think of words as being like really the almost information from an information theory point of view, like optimal way of transmitting information while audio as recorded by a microphone is just a wave that's oscillating like maybe 40 ,000 times per second. And if you just look at it with your naked eye, it will make no sense.
26:00So you need the right representation to be able to fit that into like a transformer model, have the transformer understand it and be able to produce the output. and that has been quite a challenging task. Like just if we talk about audio, like the first few success were, for instance, WaveNet. And on top of WaveNet, there was Jukebox by OpenAI that I think was the first, like, let's use a transformer language model to try to model audio. But I think I record from their paper that kind of processing one minute of audio would take eight hours on a top-of-the-line GPU at the time. So obviously, the technology has progressed a lot.
26:42And I think some of this progress was especially done by Nel Zegidur, for instance, who's another co-founder at Qtai, at Google with Soundstream in particular that provided this kind of discrete representations at a relatively low sample rate, low frame rate. And then already very quickly, Nel and his team showed that this could be fed into a transformer. At the time, they were kind of using a technique where you would still have many more, like for one second of audio, you would need to do maybe like a few hundred autoregressive steps, which is very costly. One second with a transformer of like equivalent information would be maybe three autoregressive steps.
27:24So that naturally put a constraint of both your context and the kind of lens of the sequence you can generate and completely rules out the real time aspect. then when i was at meta i also worked on similar topic especially on how to kind of not do as many autoregressive steps but try to predict some of the information in parallel and how to organize it in a way that you would have kind of minimal dependency between the different aspects you need to predict that maybe i guess it's a bit hard to say orally but basically it's like for each time step instead of having just one token like you would have in text now you have maybe four or eight or 16 tokens and yeah you need to make sense of that you cannot just flatten everything because that's that's just not going to work in terms of of performance and then there was a number of work i think one we use for moshi the rq transformer that kind of models the dependency between those tokens for a given time step with a smaller transformer i guess was a pretty important algorithmic contribution from i'm trying to find back who did that but i don't have it under my eyes but yeah so we kind of build so both on this expertise the work that now i've been doing the work that i've been doing and this kind of rq transformer paper and that's to solve the aspect of like being able to run a big language model so let's say seven billion parameter to like take audio as input and take and then output audio sufficiently fast for real-time processing and yes then the other aspects i guess the one where we kind of brought a lot of innovation was the full duplex aspect of kind of having multiple audio streams so one audio stream for the user one audio stream for moshi and i think that's kind of it's not something you would naturally do with with text because you already have one stream so going to two stream you know it's kind of a hassle but if you think of it for audio it's like all those kind of tokens in parallel they already form like up to 16 streams that we already had to handle so it was just like okay let's just double the number of streams then now we have two of them that are clearly separated we do actually the model is trained for instance during pre-training to also generate some of the users reply even if at that stage of the training there's no real like it's just kind of a participants in the conversation that sample randomly then obviously with the model we release there's no it only tries to model its own own stream but yeah so that that's kind of like the rough line of work that led to moshi then of course in audio modeling there's many other techniques that i didn't mention in particular diffusion is very popular so there's many models doing diffusion for music generation, for instance, for TTS, for a number of things.
30:25And obviously that's not compatible or like that's much harder to make compatible with the real-time aspect where autoregressive language model is still kind of the more natural and dominating paradigm. That was really fascinating in terms of like understanding. And I definitely learned as you were kind of describing it. I don't think I've heard such an excellent, you know, kind of not just from OSI, but just how to get there on that. What I'm wondering in my head is like, what are some of the I can I can imagine as you're talking so many cool things to do with this technology? What are some of the cool things that you've seen already that are that you guys have tried specifically that maybe wasn't possible before or that maybe people could only do at some level with something like like a chat GPT 4.0 kind of, you know, know through the API that way but you know this is open source it's open science they have a lot more capability there must be some pretty awesome stuff out there I mean there's like a few things that we've done that were really really funny for instance just training on this old data set from the 90s and like early 2000 of phone calls and then it's it was not really like a assistant anymore so it's just like you end up on the phone with someone random and they will tell you their name they will tell you what they think about u.s politics at the time and it's really it's kind of a different thing that we try to keep with the final moshi but obviously with the phase of instruct tuning we we lost a bit of this uh i mean it still quickly falls back to the helpful ai assistant personality uh that's maybe not as nice uh but that was a funny thing like basically we can train it on anything.
32:09And then this is going to act like a kind of actor that would pretend to be a certain person in a very realistic way. There's a number of things that we're exploring with this kind of approach, anything that would be like speech to speech or text to speech or vice versa. Some of them we kind of mentioned in the paper or with just this framework because we also have a text stream that's basically we use only for the model to be able to like output its own words. We don't actually represent the word from the user, but the model output its own words. And this kind of aspect, by making the text late or early on the audio, we can turn the model from being like a text-to-speech engine because if the text is early, then the audio is just going to follow it.
Read the full transcript
32:56But if the text is late and you kind of force the audio to some value and you only sample the text tokens, that now becomes automatic speech recognition. So I think that kind of shows how versatile this multi-stream approach is and all of those applications are really streaming. So we could actually, something we did for the synthetic data was using this kind of approach to generate long scripts and you could imagine like generating maybe 15 minutes or whatever. That's our things that we're working now more independently. And yes, in terms of more the general community, I'm not aware of anything in particular I think one thing we want to do though is to release code to allow fine tuning maybe with LoRa and also make it really easy obviously the pipeline is a bit more complex because you need audio ideally you need transcript you need separation between the agent you want to train and the users so we want to help with that regard and try to make it easier to adapt it to a new use case.
34:15Well, friends, I'm here with a friend of mine, Michael Greenwich, co-founder and CEO of WorkOS. We're big fans of WorkOS here. Michael, tell me about AuthKit. What is this? How's it work? Why'd you make it? WorkOS has been building stuff in authentication for a long time, since the very beginning. But we really focused initially on just enterprise auth, single sign on SAML authentication. But a year or two into that, we heard from more people that they wanted all the auth stuff covered. Two-factor auth, password auth, you know, with blocking passwords that have been reused. They wanted auth with, you know, other third-party systems.
34:50And they wanted really WorkOS to handle all the business logic around tying together identities, provisioning users, and even more advanced things like role-based access control and permissions. So we started thinking about that more how we could offer it as an API. And then we realized we had this amazing experience with Radix with this API, really the component system for building front end experiences for developers. Radix is downloaded tens of millions of times every month for doing exactly this. So we glued those two things together and we built AuthKit. So AuthKit is the easiest way to add auth to any app, not just Next.js.
35:25If you're building a Rails app or a Django app or just straight up Express app or something. It comes with a hosted login box. So you can customize that. You can style it. You can build your own login experience too. It's extremely modular. You can just use the backend APIs in a headless fashion. But out of the box, it gives you everything you need to be able to serve customers. And it's tied into the WorkOS platform. So you can really, really quickly add any enterprise features you need. So we have a lot of companies that start using it because they anticipate they're going to grow up market and want to serve enterprise.
35:54And they don't want to have to re-architect their auth stack when they do that. So it's kind of a way to like future proof your auth system for your future growth. And we have people that have done that. People that started off and they're like, oh, I'm just kicking the tires. I'm just doing this. And then poof, their app gets a bunch of traction, starts growing. It's awesome. And they go close Coinbase or Disney or United Airlines or, you know, it's like a major customer. And instead of saying, oh, no, sorry, we don't have any of these enterprise things and we're going to have to rebuild everything.
36:22Just go into the WorkOS dashboard and check a box and you're done. Aside from the fact that OffKit is just awesome, the real awesome thing is that it is free for up to 1 million users. Yes, 1 million monthly active users are included in this out of the gate. So use it from day one. And when you need to scale to enterprise, you're already ready. Too easy. You can learn more at offkit.com or, of course, workos.com. Big fans. Check it out. 1 million users for free. Wow. WorkOS.com or OffKit.com.
37:12So, Alex, you touched a little bit on the data side of this and also kind of hopeful future fine-tuning opportunities. But I'm wondering if you could go into a little bit in particular, because we're able to talk about this sort of thing, which sometimes we're not able to talk about, given the nature of the models that we're talking about on the podcast. What was the sort of data situation that you had to put together in terms of the specific training data sets or fine tuning data sets that you put together and curated for the model that you've publicly released as kind of model builder? obviously we had to put kind of a both like a pre-training data set in audio and in text initially we had to put the text data set together also because there wasn't necessarily at the time an alternative that we could use in terms of license and also we wanted to be able to keep training both on text and audio so as not to have a kind of catastrophic forgetting of the knowledge that will come from the text.
38:18One thing we realized that basically it's much easier to have very wide coverage of human knowledge with text than with audio. And then there were a number of other difficulties, in particular the fact that for the last stage of the training, we needed audio with clearly separated speakers. And also we needed some kind of instruct data set. So for the separation, we bootstrapped things from the Fisher dataset, which is the dataset I mentioned earlier of phone calls that kind of gave us a good enough base to then be able to train TTS model with separate speakers, also in combination with some recordings.
38:57So actually, as I talk about like, you know, taking faster decisions and in larger organizations, at one point we're like, okay, we need like really good studio quality recordings of people on separate microphones. So then we got in contact with a studio in London. Then the next day we were on the Eurostar and just like recording a few people, which I think was really fun. It's good to have a break from just launching jobs and crunching numbers now and then. And yeah, leveraging that plus the Fisher dataset, then we could train a TTS model that we could have follow like specific emotions and have two separate streams as output.
39:34So for the two speakers. And then we use that to bootstrap an Instruct dataset. Initially, we tried to convert to audio existing Instruct dataset for text, but we quickly realized that a few scripts that were specifically tuned for audio would give much better results. And one of the reasons is that if you look into some of those existing Instruct datasets, it's very geared toward first the way we're going to use text models. So maybe like some people copy-paste a markdown table and they ask to comment on it. there's a number of entries that are specifically done for kind of benchmark types of questions so it's going to be multiple choice question and the model just answers b but that's not something you're going to do orally you're not going to like give for choice and the model just answers b we needed like a lot more multi-turn like also shorter reply you don't want the model to spit out like an entire paragraph for a reply so with that in mind we had to kind of rebuild everything So that that Edward did a lot about that.
40:36So some of it was kind of pinging existing LLM being like, OK, what are like 100 different tasks we could do with a speech assistant? And then for each task, give me give me like 100 possible scenario. And then we had another model that we had fine tuned specifically to follow kind of the oral style. So shorter answers, like maybe short change of turns, we would like randomly sample topics and have discussions around them. So we tried to cover different aspects like that. And then we synthesized everything. So at the end, the dataset was fairly large. I think a few tens of thousands hours. And it was kind of sufficient to get to the state for the demo.
41:19So even though that was kind of cool that we could bootstrap this entire modality, basically from this one or 2 ,000 hours recordings from the early 2000s and a few hundred hours that we had recorded in the studio, one thing we noticed is that there is still what we call the modality gap. So there is still a gap in knowledge between the text model that we started from. And even actually, as we train the model, we can, we still train it on text. So we can always switch it to text mode and ask it the question in pure text. And the model would be like, get much better replies on trivia QA than it would get with the audio.
41:55and that's I think a really fascinating question of how to make the model understand that it's the same thing at the same time it's very easy for it to think oh it's two different modalities especially with the pre-training on audio where it gets like kind of random audio not necessarily a focus on like giving the right answers all the time we could recover some of that with the inch truck, but I think there's still work to do to be kind of as simple, efficient as a text model into really becoming super useful and factual. I'm curious, and you may have mentioned it, you mentioned 7 billion parameters earlier, but is that the size of the model?
42:35Is it a 7 billion parameter size model? Yes, it is 7 billion parameter. So as I mentioned, so it's a RQ transformer architecture. Actually, I found back the author. So it's Doyouk Lee and the collaborators that published first this model, which has the main backbone transformer and the small transformers that just tries to predict the different acoustic tokens. This one is kind of smaller. I don't have the weight, the size exactly, but in terms of runtime, inference time, it's negligible. Most of the knowledge and decision is done in the big 7 billion transformer. How did you pick the model being that size?
43:19And also as an addendum to that, what is your perspective on the relatively smaller models versus the relatively larger models? How do you see that? Yeah, I guess when we started, 7 billion was kind of the minimum size for large language models. Now, I guess 2 billion and 3 billion down to 1 billion, especially with the advance of distillation techniques from bigger transformers have become very efficient. They are now as efficient as tech's model at 7 billion parameters from like a year and a half ago. But yeah, at the time when we started, we're like, OK, we don't know exactly how much compute, how much capacity it's going to take to solve the task.
44:00So we don't want to take too many risks. 7 billion was a well-charted territory and at the time, like a pretty good balance between the two. Now that we know that we can solve the task with 7 billion, obviously we want to try to go lower than 7 billion. And that's something we are exploring because like the way we see things, it's going to be very hard and probably not super useful to try to put all the like thinking capacity and problem-solving capacity into Moshi, but we want it to be smart enough to have a direct conversation, understand what the user wants, and potentially then access other source for getting more complex answers that would also allow more plug-and-play aspects.
44:45And now you have a new text language model you don't necessarily want to retrain from scratch the audio part. So the way we see it is going towards a smaller model for managing this direct low-latency interaction, delegating some of the works to a larger model when needed. So for sure, now that we know it works with 7 billion, we would try smaller so that we can run on a much larger number of device. I guess you already started kind of talking towards additional things that you want to try with respect to Moshi and these types of models in the future. But maybe stepping back a little bit as we get close to the end of the episode here, when you as a researcher in this area look towards the future, it could be work that you all are planning to do internally or just things going on more broadly.
45:36But what kind of is some of the most exciting things for you as you think about the kind of next year of your work and the things that you're following, the things that you're looking at? What's on your radar and what are you excited to kind of participate in and see happen in the coming months? Okay, in the coming months, that's a good question. I mean, I think one topic that I'm interested at the moment is the question of whether we're going to be one day in the post-Transformer era. Like, I love Transformer and I love not having to wander anymore. I mean, if we look at the set of hyper parameters to train those models there, I've been frozen for maybe two years, two years and a half.
46:20You know, the architecture is frozen, which is good because now we mostly focus on just like making the right data to solve problems. And there's a lot we can do. At the same time, I think I would be really excited to see advancements that could happen either on the optimization side or the architecture side. We've seen a lot of like interesting work in this area. but at the moment we're more on a parity thing okay so we found other ways of doing kind of the same thing but there is not really something that has won a decisive on a decisive aspect or feature that would be potentially not sufficiently well done with transformers there's been like tons of engineering going into it so you know each time you think oh maybe quadratic cost is bad but then people are like no you can just hardcore optimize your CUDA kernel and now it's no longer your problem.
47:13But yeah, that's, I think in terms of the, just the scientific excitement, that's, that's one thing that I want to keep my eye on. Obviously, as at the same time, there is a lot of competition going on, just applying the current model. It's not necessarily easy to free time and mental space to try to think about those issues. So that's one aspect. And yeah, then, then I'm also curious about how all the the framework uh aspect is going to evolve uh working uh day to day with those technologies really feels like you're back in the 70s of like pre-c era where you have to think about the cuda like the code is different for each architecture there's a lot of leakage abstraction leakage it's not like you're not going to write a nice function you you need to write kind of dirty things you need to do the equivalent of pointer arithmetic all the time so that's another thing so maybe i'm not replying your question of what's coming in the next few months but longer term sometimes i just think of myself in 10 years and you know you can just write your attention kernel in a few lines of code in a dedicated language and you get like almost perfect code and i think that would be amazing to just explore more things more easily uh but we'll see so two yeah two big uh big potential changes but i think something's gonna happen in the coming years Yeah, well, thank you very much for sharing your perspectives with us.
48:37And also thank you for the way that you and the Qtai team are inspiring many people out there that are working on open models, open source, open science, and kind of just generally collaborating in this space. Really appreciate kind of what you're doing as part of that. And thank you for taking time to chat with us. It's been great. Thank you very much for the invitation, opportunity to present, and hopefully we'll have some other in the future. Definitely. Enjoy Thanksgiving.
49:15All right. That is our show for this week. If you haven't checked out our ChangeLog newsletter, head to changelog.com slash news. There you'll find 29 reasons, yes, 29 reasons why you should subscribe. I'll tell you reason number 17, you might actually start looking forward to Mondays. Sounds like somebody's got a case of the Mondays. 28 more reasons are waiting for you at changelog.com slash news. Thanks again to our partners at fly.io, to Breakmaster Cylinder for the Beats, and to you for listening. That is all for now, but we'll talk to you again next time.
50:00Game on!
From the publisher
Kyutai, an open science research lab, made headlines over the summer when they released their real-time speech-to-speech AI assistant (beating OpenAI to market with their teased GPT-driven speech-to-speech functionality). Alex from Kyutai joins us in this episode to discuss the research lab, their recent Moshi models, and what might be coming next from the lab. Along the way we discuss small models and the AI ecosystem in France.
Changelog++ members save 10 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
- Fly.io – The home of Changelog.com — Deploy your apps close to your users — global Anycast load-balancing, zero-configuration private networking, hardware isolation, and instant WireGuard VPN connections. Push-button deployments that scale to thousands of instances. Check out the speedrun to get started in minutes.
- Timescale – Purpose-built performance for AI Build RAG, search, and AI agents on the cloud and with PostgreSQL and purpose-built extensions for AI: pgvector, pgvectorscale, and pgai.
- WorkOS – AuthKit offers 1,000,000 monthly active users (MAU) free — The world’s best login box, powered by WorkOS + Radix. Learn more and get started at WorkOS.com and AuthKit.com
Featuring:
- Alexandre Défossez – Website, GitHub, X
- Chris Benson – Website, GitHub, LinkedIn, X
- Daniel Whitenack – Website, GitHub, X
Show Notes:
Something missing or broken? PRs welcome!




