In short
Podcast Summary: No Priors - State Space Models and Real-time Intelligence with Karan Goel and Albert Gu from Cartesia
Episode Overview In this episode, co-hosts Sarah Guo and Elad Gil speak with Karan Goel and Albert Gu, co-founders of Cartesia. They discuss the innovations in state-space models (SSMs) and their implications for real-time intelligence, particularly in text-to-speech technology. The episode highlights the significant advancements made by Cartesia in creating Sonic, an ultra-fast text-to-speech engine.
---
Key Topics and Discussions
- Introduction to Cartesia and Sonic
- Cartesia's Mission: Focused on building real-time intelligence for devices.
- Sonic: A text-to-speech engine offering low latency (135ms), specifically designed for interactive applications such as gaming and voice agents.
- Background of Karan Goel and Albert Gu
- Both co-founders have a shared background from Stanford AI Lab, where they developed foundational models.
- Their journey involved working on sequence modeling and exploring alternatives to transformer-based architectures.
- State Space Models (SSMs) vs. Transformer Architectures
- Key Differences:
- SSMs provide linear scaling efficiency, while transformers exhibit quadratic scaling.
- SSMs operate by updating information in a compressed state, modeling sequences more efficiently for certain types of data.
- Applications of SSMs
- Primarily effective for raw waveforms, like audio and video, but less so for text where transformers excel.
- Hybrid models combining SSMs with transformers are becoming increasingly popular, demonstrating enhanced performance.
- Advances in Text-to-Speech Technology
- Discussion on the challenges of current text-to-speech systems, emphasizing the importance of emotional engagement and naturalness.
- Sonic aims to improve the conversational quality of synthesized speech, addressing the limitations of existing systems.
- The Future of Real-time Intelligence
- Emphasis on developing models that can run efficiently on smaller hardware, promoting edge computing.
- The potential for reducing costs of advanced models from millions to just dollars for on-device applications.
- Hiring and Growth at Cartesia
- Cartesia is expanding, with ongoing recruitment for model and engineering roles.
- The company is focused on assembling a strong team to drive their innovative technology forward.
---
Key Takeaways
- Real-time Applications: Sonic represents a significant leap in text-to-speech technology, paving the way for interactive and immersive voice experiences.
- Efficiency in Modeling: The move from transformer-based architectures to SSMs offers potential advantages in efficiency and processing speed, especially in real-time environments.
- Multimodal Capabilities: The integration of audio understanding and processing is critical for advancing AI interactions in the future.
- Industry Impact: The evolution towards edge computing and efficient models could democratize access to advanced AI capabilities across various domains.
---
Conclusion The dialogue between Sarah Guo, Elad Gil, Karan Goel, and Albert Gu offers a deep dive into the state of AI technology, particularly focusing on the advancements in state-space models and their practical applications. Cartesia's work exemplifies the shift towards more efficient, real-time intelligence that could transform interactions across multiple industries.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Welcome back to KnowPriors. We're excited to talk to Karen Gohl and Albert Gu, the co-founders of Cartesia, and authors behind such revolutionary models as S4 and Mamba. They're leading a rebellion against the dominant architecture of Transformers. So we're excited to talk to them about that and their company today. Welcome Karen, Albert. Thank you. Nice to be here. And Kate, tell us a little bit more about Cartesia, the product, what people can do with it today, some of the use cases. Yeah, definitely. We launched Sonic. Sonic is a really fast text-to-speech engine. So some of the places I think that we've seen people be really excited about using Sonic is where they want to do interactive, low latency voice generation.
0:45So I think the two places we've really kind of had a lot of excitement is one in gaming, where folks are really just interested in powering characters and roles and NPCs. The dream is to have a game where you have millions of players and they're able to just interact with these models and get back responses on the fly. And I think that's sort of where we've seen a lot of excitement and uptake. And then the other end is voice agents and being able to power them. And again, low latency there matters. And even with what we've done with Sonic, we're already kind of shaving off like 150 milliseconds off of what they typically use.
1:23And so the roadmap is let's get to the next 600 milliseconds and try to shave those off over the course of the year. That's been the place where it's been pretty exciting. Love to talk a little bit just about backgrounds and how you ended up starting Cartesia. Maybe you can start with the research journey and like what kinds of problems you were both working on. Cart and I both came from the same PhD group at Stanford. I did a pretty long PhD and I worked on a bunch of problems, but I ended up sort of working on a bunch of problems around sequence modeling. It came out of kind of these problems I started working on actually at DeepMind during an internship.
1:54And then I started working on sequence modeling. Around the same time actually that Transformers got popular, I actually, instead of working on them, I got really interested in these alternate kind of recurrent models, which I thought were really elegant for other reasons. And it kind of felt like fundamental in a sense. And so I was just really interested in them and I worked on them for a few years. A couple of years ago, me and Kern worked together on this model called S4, which kind of got popular for showing that some form of recurrent model called a state-space model was really effective in some applications.
2:21And I've continuing to be pushing on that direction. Recently, I proposed a model called Mamba, which was kind of brought these to language modeling and showed really good results there. And so people have been really interested. We've been using them for applications and other sorts of domains and so on. So yeah, it's really exciting. Personally, I also, I just started as a professor at CMU this year. My research lab there is kind of working on the academic side of these questions while at Cartesia we're kind of putting them into production. Yeah, I guess my story was that I grew up in India, so I came from an engineering family.
2:56You know, all my ancestors were engineers. So I actually was trying to be a doctor in high school, but my aptitude for biology was very low. So I abandoned it and instead became an engineer. So, you know, I kind of took a fairly typical path, went to an IIT, came to grad school, and then ended up at Stanford. Actually started out working on reinforcement learning back in 2017, 18. And then once I got into Stanford, I started working with Chris, who was somewhat skeptical about reinforcement learning as a field. And so... This is Chris Ray. Yes, Chris Ray, who is our PhD advisor. So I had a very interesting sort of transition period where I started a PhD because I had no idea what I was working on.
3:40And so I was just exploring. And then ended up actually, we did our first project together too. Oh, yeah. It's like good times. Actually, we knew each other before that. And I think then we started working together on that first project. and we would hang out socially and then start working together. The only memory I have of that project was I kept filling up this disk on G Cloud and expanding it by one terabyte every time. And then it would keep filling up and I would insist on only adding a terabyte to it, which he was very mad about for a while. Well, by the end of the project, it was like running a bunch of experiments and logs would get filled up faster than...
4:18Yeah, then you would... Basically, I would be there tracking the experiments, and Connor would be there deleting logs in real time so that our runs didn't crash. It was a really interesting way to start working together. Yeah, so we started working together then, and then I eventually started working with Albert on the S4 push when he was pushing for NeurIPS. And I think he was working on it alone and then needed help. I got recruited in to help out because I was just not doing anything for that NeurIPS deadline. So ended up spending about three weeks on that, two or three weeks, something like that.
4:51And then we really pushed hard. And that's kind of how I got interested in it because, you know, he had been working on this stuff for a while and, you know, nobody really knew what he was doing. To be honest, in the lab, it was just like over in the corner, scribbling away, talking to himself. We don't really know what's going on. Could you actually tell us more about SSMs and, you know, how is it different from transformer-based architectures and what are some of the main areas that people are applying them right now? because I think it's really interesting is sort of another approach. They really kind of got started from work on RNNs or current neural networks that I was working on before as an intern in 2019.
5:23It kind of felt like the right thing to do for sequential modeling because the basic premise of this is that if you want to model a sequence of data, you want to kind of process the sequence one at a time. If you think about the way that you will kind of process information, you're taking in sequentially and kind of encoding it into like your representation of the information that you know, right? And then you get new information and you update your belief or your state or whatever with the information that new information that you have. You can basically say almost any model actually is doing this.
5:52And then there were some connections to other like dynamical systems and other things that I found really interesting mathematically. And I just thought this kind of felt like a fundamental way to do this. It just felt right in some ways. You can kind of think of these models as doing something, there's like some loose inspiration from the brain even where you kind of think of the model as encoding all the information it's seen into a compressed state. It could be kind of fuzzy compression, but that's actually powerful in some ways because it's a way of kind of stripping out unnecessary information and just trying to focus on the things that matter and code those and process those and then work with that.
6:28We can get more than technical details, but kind of like at a high level is just this thing. It's just representing this idea of this fuzzy compression and fast updating. So you're just keeping this the state in memory that's just always updating as we see new information. Is it better architecture for certain types of data or did you have applications in mind besides the general architectural concept? Yeah, so it really can be applied to pretty much everything. So just like Transformers, these are applied to everything, so can these models. Over the course of research over a few years, we realized that there are different advantages for different types of data and lots of different variants of these models are better at different types of data or others.
7:11So the first type of model we worked on were really good at modeling kind of perceptual signals. So you can think of text data as kind of a representation that's already been really compressed and like tokenized, right? Pre-cooked. Yeah, sure. And it's kind of like very dense. Like every token in text already has a meaning. It's kind of just dense information. Now, if you look at like a video or an audio signal, it's highly compressible. It's, for example, if you sample at a really high rate, it's basically like, it's very continuous. And so that means it's compressible. And it turns out that, yeah, different types of models just have different inductive biases or like strengths at modeling these things.
7:48The first types of models we were looking at were really good actually at modeling kind of these raw waveforms, raw pixels, things like that, but not as good at modeling text and transformers are way better there. Newer versions of these models like Mamba, which was the most recent one that's been out for a few months, that's a lot better at modeling the same types of data as transformers. Even there, there's subtler kind of trade-offs. But yeah, so one thing we kind of learned is that in general, there's no free lunch there. So people think that like you can throw a transformer at like anything and it just works.
8:19Actually, it doesn't really. Like if you try to throw it at like the raw pixel level or the raw sample level in audio waveforms, I think it doesn't work nearly as well. And so you have to be a little more deliberate about this. They really evolved hand in hand with the whole ecosystem of the whole training pipeline. So it's like the places that people use Transformers, the data has already kind of been processed in a way that helps the model. For example, people have been talking a lot about tokenization and how it's both extremely important, but also like very counterintuitive, unnatural, and has its own issues.
8:53That's an example of something that's kind of developed hand in hand with the Transformer architecture. And then when you kind of break away from these assumptions, then some of your modeling assumptions no longer hold. And then some of these other models actually work better. Do you think of the advantages like natural fit that translates to quality for certain data types? At least if we think about like, let's say perceptual data or I don't know, richer raw pre-cooked, not pre-cooked data. Or like, you know, how do you think about efficiency or the other dimensions of like comparing the architectures?
9:24Yeah. So I guess so far we talked kind about the inductive bias or the fit for the data. Now, the other reason why we really cared about these is because of efficiency. So yeah, maybe we should have led with that even. So people have yelled for a long time about this like quadratic scaling of transformers. One of the big advantages of these alternatives is the linear scaling. So it just means that basically the time it takes to process any new token is basically constant time for a current model. But for a transformer, it scales with the history that you've seen. This is obviously a huge advantage when you're really scaling to like lots of data, but it is actually something that's sort of a little bit of like a no free lunch thing.
10:03The fact that the transformer is processing, is taking longer to process things also means that there are things that it's better at modeling. So this is kind of what I was talking about. There's some like subtleties when you're talking about the trade-offs there. One way that we've sort of been thinking about it more is kind of thinking of, so as I mentioned in the beginning, we think of these state-space models as kind of being fuzzy compressors. And I think maybe kind of like the bulk of the processing should be done there. But at the same time, it benefits from having some sort of like exact retrieval or some cache.
10:32And that's exactly what a transformer is. So one way to think about the transformer is that you're processing all this data and it's just memorizing every single token it's seen basically. I mean, some kind of representation of it, but it's literally remembering every single thing you've seen. And you're allowed to look back over all of it. So that's why it's a lot slower, but that could be useful. But probably that shouldn't be what the bulk of your model is doing. Kind of the same way that like, I mean, again, using like very rough and probably not accurate analogies, but like the way a human brain is probably most of the intelligence is in this, you know, it's the statefulness, this real time processing unit.
11:08But it is helpful to augment it with some sort of scratch pad or like lookup ability retrieval. Right. And so these these ideas are actually quite synergistic. And what people have recently been doing is finding that combining them into hybrid models tends to work really, really well. Seems like it's better than either of them individually. So, and interestingly, kind of also maybe in line with that intuition, people have found that the optimal ratio tends to be mostly SSM layers with a little bit of attention. So maybe a ratio of like 10 to 1. I know of at least probably like five groups that have independently verified that this is kind of the optimal ratio of things.
11:47So, yeah, I think it makes intuitive sense. Are there specific domains that you're seeing initial applications of these hybrid approaches? People are mostly using this on text because that's what everyone cares about. I think they've been investigated a bit on some other things. So I actually just heard from some collaborators today that they applied a mama-based model on DNA modeling. They're basically bringing this idea of foundation models to DNA, which is kind of this new idea. Sure. I'm just wondering, because DNA is just, do you mean translation of DNA into proteins? Is it a protein folding model?
12:23No. So what you do is you can like pre-train a model on long DNA sequences and then fine tune it or use it downstream on things such as even just - Sure. But DNA itself just encodes proteins and RNA that fold into certain molecular shapes. So that's why I was wondering what the problem set is. Yeah. There's a bunch of them and I'm honestly not familiar with a lot of the exact details. Yeah. I used to be a biologist. That's my background. Well, just like current, my brain can't handle biology. Okay. Yeah, yeah, yeah. I think I get it. I was just curious, like, what the specific application area was.
12:56Oh, this is not protein folding, per se, but probably more like classification tasks. Like, one thing that people are interested in is, like, detecting whether, like, point mutations in DNA, like, what downstream effects that can have and stuff. But I'm not sure the exact classification setting. Yeah. And then one of the areas that you folks really started focusing on from a company perspective is tech to speech and voice. How did the research lead into that domain? We were kind of interested in showing the versatility and the actual use case of these models. So previously it was done mostly in academic context and at CMU my students are still kind of carrying that fundamental research forward.
13:30But we were pretty sure this would just work in a lot of places that are interesting. And so we think it will really work on all sorts of data, but kind of audio seemed like a pretty natural fit at first for some of the benefits. Like what we talked about is like much faster inference and so on. And so doing like streaming settings and so on or natural fit, we just thought this would kind of be like cool first application. Maybe Karan can kind of say more about. Yeah, I mean, there's so many applications that are interesting for these models because they're so generically useful. And I think part of the challenge is, you know, sort of picking the ones that are most interesting and impactful long term.
14:08Obviously DNA is an interesting one but you know we don't personally have it doesn't personally motivate us as much or we would have worked on DNA. But I think like to me like the things that are interesting about multimodal data are really the the places where SSMs have the most advantage right which is that you have data that's very very information sparse. Compression is actually an advantage because you can stream data through the system really fast and then process it very quickly so you update this sort of memory um and and so being able to handle very large context is kind of something that you want by design the other thing that i think really interesting about audio is that i think commercially there's a lot of very interesting applications where audio is starting to be important i think both on the voice agent side and like being able to kind of interact with your system in a more you know natural way like you would with a human is something that a lot of people want to be able to do is because there's a lot of places where you don't want to type into your computer, you actually want to talk to a human.
15:08Even on things like gaming and stuff, I think it's really interesting to think about how in the future, you will be essentially replacing graphics and rendering with essentially models that are outputting streams of data in real time. So I think the real time aspect of it is really core to signals and sensor data, and of which audio and video are both very important. And I think audio in particular for us felt like a very natural place to start because I did some work on audio in my PhD, which was also something that was helpful. And I think there's just so many applications in audio that are really emerging right now that require these types of capabilities to exist.
15:47So I think that's particularly exciting. The other piece I think that's really interesting about SSMs and that we're very excited about and are trying to do is the fact that like the model is so efficient that you can hope to put it on smaller hardware and actually push inference closer to on device and to the edge. And I think that so far the theme in a lot of the models that people use has been data center, very big model, lots of compute, lots of GPUs being burned. I think that in the long run, like what we would hope is that you're pushing this closer and closer to the edge. you're actually using much less compute to do inference.
16:27And you're actually able to basically reproduce a capability that maybe today costs a million dollars in the data center in$10 on a commodity GPU or accelerator at the edge. So I think that will be a very powerful shift because essentially what it means is that instead of running batch-oriented workloads in the cloud, you're basically pushing the processing and the intelligence to closer to where the data is being acquired and where the sensors are. And that's kind of what you want because if you think about security cameras or any kind of really sensor that's deployed, you really do want to be able to kind of sift through the information very quickly, discard what's not useful because most of it isn't, and then really kind of remember all the stuff that is and use that to do prediction and generation and understanding problems.
17:15So I think that's the theme. I think in general for what we're trying to build is sort of the infrastructure to be able to train these models, make them run fast, and then bring them closer and closer to kind of be, you know, very edge-oriented rather than cloud-oriented. That's super interesting. I guess as part of the recent Apple announcement, they mentioned that a lot of the models that are running on device are 3 billion parameters. Yeah. So in size, and so you really have to focus on small models. Yeah, so it is partly like, in my head, like it's always like two waves, right? Like there's like the first wave of companies that came was sort of really about like, how do we figure out if we can do something interesting, right?
17:49Nobody knew that scaling to this amount of data and compute would be interesting. So somebody took a bet there and did that. And that was great because now we have all these great models. I think the second wave is always about efficiency. And that's been the case in computing as well. We're like, you know, now we have phones that can do so much powerful work. And similarly, I think for models, what you would want is the smartest model ever, but run it real cheap. So you can run it repeatedly at scale. You can run it 100 times where you might run it once today. So a lot of that needs to happen.
18:21So I think that's what's interesting. So the 3D models are interesting, I think, but they're still small and not very capable. I think the question is, how do you make the capabilities just very, very good, but then have that low footprint on device, all of that? So I think that's where the technology that we've been playing with for the last few years and then building with now has really huge potential to actually be the default to run these workloads. Because I think that part of the challenge with Transformers is the fact that, you know, if you try to take a LLM 7B and try to run it on your Mac and you open your profiler, you will notice that the, you know, tokens per second goes down and the memory goes up.
18:59So I think that's, you know, obviously not great. And power and all these things aren't something that people have like, even I think in the data center, people are talking about it now in the last year or so. But now, you know, you'll start to see more of that conversation shift. And so I think that's kind of where we want to be, which is like, you know, the future will be more intelligence everywhere. And how do you kind of enable that piece, I think is kind of what we're excited about. Yeah, I think we get really different applications if people start making that assumption. As you said, we see it in the data center first, where even as an investor betting on applications or full stack companies that say, it costs a great deal to do 1 ,000 calls per query right now, but we're just going to assume we can make it cheaper and we can focus on quality first.
19:43I think when you assume that you can run the model or if you make it possible to run the model on hardware everybody has already, then you just get very different applications continuously and that quality without that cost and ongoing compute being a problem. Yeah, just the set of things you would want to be able to do will change because, you know, in the same dollars, you will be able to just do way more intelligent computation. And so I think that's cool. I think the way that, you know, you run games on your computer and the games are like very powerful, very, very rich and interesting. How do you kind of bring models to that place where if you think about on device, I shouldn't be able to have a music model on device and that should be my personal musician that I can talk to and get it to play whatever I want and I don't need to go to the cloud to do it.
20:27So I think these are all things that should be possible and just require this type of infrastructure and work. Yeah, and Cartasia recently launched its initial sort of text-to-speech product and it's really impressive in terms of performance and how fast you've gotten to ship something really is that performing. Can you tell us a little bit more about that launch and that product? Yeah, I think it was sort of a natural transition for us to kind of now start thinking about how to put the technology to work because there was a lot of pre-work that happened and Albert continues to do the pre-work for the next set of things.
20:59But I think it's sort of like how do you kind of build an efficient system that will allow you to do, for example, in this case, voice and audio generation. So I think the way we're thinking about it is we're building these fairly general models inside the company that allow us to kind of do fairly generic tasks very efficiently. So in this case, it's audio generation and then being able to condition on things like text transcripts. The philosophy is like, oh, audio generation is a problem, needs to be very efficient, needs to be very real time. And so we need to kind of work on the groundwork there to build the sort of the model stack.
21:32And then we need to have great training stacks so that we can actually train a model that's high quality that people want to use and that actually has a really great experience. So when we were putting together the Sonic demo, it was sort of like we wanted to show that the tech that we were using really can kind of give you something that's really interesting. And text-to-speech is very interesting to me because, you know, people have been building text-to-speech systems for the last, you know, probably 30, 40 years. There's constantly improvements happening. And yet we're not at ceiling, right?
21:59Like there's still so much more you can do in this area. Can you actually talk about that? Because I think a lot of people would say, like, that feels a lot more solved in the last year. Just text-to-audio generation. I feel like... Like, what's left between here and the ceiling in terms of thinking about the application experience? Yeah, I think, like, the way I think about it is, like, would I want to talk to this thing for more than 30 seconds? And if the answer is no, then it's not solved. And if the answer is yes, then it is solved. And I think most text-to-speech systems... Kern's audio touring test, yeah.
22:28...are not that interesting yet. Yeah. You don't feel as engaged as you do when you're talking to a human. I know there's obviously other reasons you talk to humans, which is, you know, sorry, I don't want to come across as crazy here. But yeah, there's a society that we live in. So we want to talk to people for that reason, obviously. But I do think the engagement that you have with these systems is not that high. When you're trying to build these things, you really kind of get so into the weeds on like, oh, I can't say this thing this way. and it's like so boring when it says it that way and how do I control this part of it to say it like this?
23:01You know, the intonation. Are there specific dimensions that you look at from an eval perspective that you think are most important in terms of how you think about? Yeah, evals for, you know, generation are generally challenging because they're qualitative and based on sort of, you know, the general perception of someone who looks at something and says, this is more interesting than this. And so there is some dimension to that. But I think for speech, like, you know, emotion is something that matters a lot because you want to be able to kind of control, you know, the way in which things are said.
23:30And I think the other piece that's really interesting is how speech is used to embody kind of the roles people play in society. So like different people speak in different ways because they have, you know, different jobs or work in different, you know, areas or live in different parts of the world. And that's sort of the nuance that I don't think any models really capture well, which is like, you know, if you're a nurse, you need to talk in a different way than if you're a lawyer or if you're a judge or if you're a venture capitalist. you know, very different forms of speech. The highest form of voice.
Read the full transcript
24:00Yeah. So those are all very challenging, I would say. So it's not solved, is my claim. There's also an interesting point, which is kind of like, even just for your basic evaluations of like, can your ASR system like recognize these words or can your generation, can your TTS system say this word? Even that is actually not quite a local problem. And for a lot of hard things, you actually need to really have the language understanding in order to process and figure out what is the right way of pronouncing this and so on. And so actually to really get like perfect, even just TTS or like speech to speech, you actually really need to have like a model that has more understanding, like at least of the language, but kind of like, it's not really an isolated component anymore.
24:41And so you have to start getting into these multimodal models just to even do one modality well. And so that's kind of like somewhere where that we were kind of eyeing from the beginning as well. And we were kind of using this as an entry point into building out the stack toward all of that and hopefully that's all gonna help the audio as well, but also start getting other modalities into that. That's really cool. I mean, I guess you've done so much pioneering key work on the SSM side. How is multimodality or speech really impacted how you've thought about the broader problem or has it, and it's more just the generic solutions are the ones that make sense?
25:13I don't think multimodality by itself has been kind of a driving motivation for this work because I kind of think of these space models I've been working on as like basic generic building blocks that can be used anywhere. So they certainly can be used in multimodal systems to good effect, I think. Different modalities have presented different challenges, which has influenced the design of these. But I always look for kind of like the most general purpose, fundamental kind of like building block that can be used everywhere. And so that's like multimodality is more of like a sort of a different set of challenges in terms of like, how are you applying the building blocks to that?
25:50But like you still use the kind of the same techniques and they mostly work. Given that versatility of model architecture, generality of the building block, what do you do next for Cartesia? You focus on the headroom for Sonic and audio. You work on other modalities. I'll take that one. We're obviously really excited about the Sonic work because I think it kind of shows the first example of something that we're excited about, which is it's a real-time model. You can run it really, really fast at low latencies, and it's capturing this idea that you want to generate a signal of some kind. So we're going to continue to obviously improve that piece.
26:28Also, you know, just generally things that folks want out of speech systems that need to get built that are orthogonal to the technology piece, which is, you know, being able to support lots of languages and just generally providing more controls. orthogonal access that's really important for generative models, which is how do you add more controllability in general to the system so you can get the desired output that you want. So that's obviously one focus for us is how to put that piece in. A few things that we're doing that I think in the short term are really interesting. One is bringing Sonic more on device.
27:01You can run the model real time in the cloud. Wouldn't it be cool if you could run it on your MacBook and it ran real time and it was just as good? I actually have a demo I can show there that I think is super cool. Over time, what we want to do is what Albert said, which is that, you know, audio benefits from text reasoning and, you know, the ability to kind of converse with these models and actually have them understand what you're saying beyond just, you know, superficial understanding is very important. So what we want to do is enable that piece next, which is, you know, you should be able to have a conversation with this thing and actually be able to have it respond to you intelligently and reason over data and context in order to do that.
27:39And so Sonic, I think of as sort of the output piece of that in some sense, which is like, what does the response for that model look like? And then there's the input piece, which is ingesting audio natively into these models and doing that kind of thing. Is the intention then to train a large scale multimodal language model on your side as well? Yes, but we have our own sort of set of techniques that we're developing in order to be able to do that effectively. I think that I will maybe leave for another podcast. But I think, yeah, I think that is the intention at the end of the day is build a great multimodal model, but then make it really, really easy to run on device and make it really cheap to run and really focus on kind of the audio piece and making that as good as possible.
28:21because I think that's sort of where, you know, the fidelity and the quality that you get from SSMs is just very different than what you've been able to see. That's pretty amazing because it seems like a lot of the limitations right now in terms of different application areas or use cases where text-to-speech is basically the extra latency or run-trip associated with pinging a language model in the middle. Yeah. As you go from speech to text, the text, and then out. And so if you do have multimodality, then obviously that shrinks the time on the inference side dramatically and that has a huge impact in terms of that.
28:47Yeah, I think the latency is going to be a big theme there because it's obviously like quite painful to orchestrate multiple models to do this piece. And then I think also just the orchestration itself adds so much overhead. It turns what is, in my mind, something that the model should do into an engineering problem that requires so much, you know, orchestration and just engineering work to... It feels almost inelegant. Yeah. From a computer science perspective. Yeah, maybe that's, you know, some of the thematically the general bias here, which is, you know, the inelegant things we're trying to chip away at.
29:24In the end, all the systems go away and it's just one model. And then we also go away, apparently. People ask me, like, how do I treat my research problems? And I can't explain. My answer is just aesthetic. It's just like there's something that I find elegant and we're aesthetically pleasing about things. And to me, that's almost the most important thing. And that's kind of driven a lot of these things, too. So like I said, like for how did SSMs come about in the first place, is just like, I just felt like there was something like really like nice about it, like elegant about it. And I just want to keep working on it.
29:56And I'm continuing to try to like do that, like find like the simple, nice solutions to hard problems. But it's not always possible. So at Cartesia, we of course need to solve the actual like the engineering challenges and there's always going to be hairy things. But as much as I can, I'm always trying to strive to kind of like make everything simple unified as possible that's great yeah i remember i can't fit is it uh erdos or somebody uh used to talk about uh certain theorems coming out of like god's book oh yeah or so elegant yeah i very much adhere to that uh the idea so the it's called uh proofs from the book is what he would say yeah and that's actually kind of thing that kind of um guides a lot of the way that i like picking choosing problems and what you're referring to is of course in like in pure math, sometimes you see like proofs or ideas that just feel like this is obviously just the right way of doing things.
30:49It's so elegant. It's so correct. Things are not, in machine learning world, things are often not nearly that clean, but you still can have, still the same kind of concept, just, you know, maybe a different level of abstraction, but sometimes certain approaches or something just seems like the right way of doing things. Unfortunately, this thing is also kind of like, it can be subjective. Yeah, sometimes I tell people this is just the right way of doing it, and I can't explain why. But maybe we should kind of have like one of our pillars should be about the book so I can start saying this. Let's see the demo.
31:24Yeah, I'd love to show you. Cool. Yeah, I have our model running on our standard issue Mac here. Basically, this is, you know, our text-to-speech model, Sonic, on our playground is running in the cloud. And so, you know, part of what I talked about earlier was how do you kind of bring this closer to on-device and edge? And I think the first place to start is your laptop and then hopefully bring it, shrink it down and bring it closer and closer to a smaller footprint. So let me try running this. It's great to be on the No Priors podcast today. You know, we have the same feature set that's in the cloud, but running on this and...
32:02Prove it's real-time and not Coup. say, you don't have to believe in God, but you have to believe in the book. I think that's the Erdos quote. Was that the quote? Let me grab an interesting voice for this one. Erdos is, where is Erdos from? Hungary. I mean, that's a default guess for any mathematician. Oh yeah, sure. He's just assuming. All right, I'm going to press enter. You don't have to believe in God. You have to believe in the good. That's pretty good. Lain C is pretty good. Yeah, it works really fast. And I think that's part of what I think gets me really excited, which is like, you know, it streams out audio instantly.
32:38I would talk to Ardosh on my laptop. Yeah, me too. That would be a great way to get inspired every morning. Yeah, I know. Yeah. Yeah, that'd be great. Your team is now how many people? We are 15 people now. And eight interns. Sarah always gives me shit for this. It's a big intern class. Yeah, we have a lot of interns. I really like interns. They're great. You know, they're excited. They want to do cool things. And are there specific roles that you're currently hiring for? Yeah, we are hiring for, you know, model roles specifically. We're hiring across the engineering stack, but really want to kind of build out our modeling team deeper.
33:17So always looking for, you know, great folks to come to Team SSM and help us build the future. The rebellion. Yeah, the rebellion. Yeah, we used to actually call it. Yeah, yeah, yeah. What do you call it? Overthrowing the empire. Yeah, yeah, yeah. That was the theme during our PhDs. And yeah, I would love to continue to, you know, have folks inbound us and chat with us if they're excited about this technology and the use cases. A lot of exciting work to do, both research and bringing it to people. Yep. Find us on Twitter at NoPriorsPod. Subscribe to our YouTube channel if you want to see our faces.
33:54follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-priors.com.
From the publisher
This week on No Priors, Sarah Guo and Elad Gil sit down with Karan Goel and Albert Gu from Cartesia. Karan and Albert first met as Stanford AI Lab PhDs, where their lab invented Space Models or SSMs, a fundamental new primitive for training large-scale foundation models. In 2023, they Founded Cartesia to build real-time intelligence for every device. One year later, Cartesia released Sonic which generates high quality and lifelike speech with a model latency of 135ms—the fastest for a model of this class.
Sign up for new podcasts every week. Email feedback to show@no-priors.com
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @krandiash | @_albertgu
Show Notes:
(0:00) Introduction
(0:28) Use Cases for Cartesia and Sonic
(1:32) Karan Goel & Albert Gu’s professional backgrounds
(5:06) State Space Models (SSMs) versus Transformer Based Architectures
(11:51) Domain Applications for Hybrid Approaches
(13:10) Text to Speech and Voice
(17:29) Data, Size of Models and Efficiency
(20:34) Recent Launch of Text to Speech Product
(25:01) Multimodality & Building Blocks
(25:54) What’s Next at Cartesia?
(28:28) Latency in Text to Speech
(29:30) Choosing Research Problems Based on Aesthetic
(31:23) Product Demo
(32:48) Cartesia Team & Hiring




