In short
Mirelo AI (CJ Simon-Gabriel) announces a ~$41–$42M seed round led by Andreessen Horowitz and Index, to build foundational European audio models for generating synchronized sound for video and games.
Guests
CJ Simon-Gabriel, CEO and co-founder of Mirelo AI; previously did ~10 years of AI research and worked on AI at AWS Lab/AI research, and is also a trained musician (Conservatoire in Strasbourg). Co-founders mentioned: Florian (electro musician background in Berlin; co-founder) and team members including other musicians.
Key claims
Audio models can be “best in class” with far less compute (audio models ~1–10B params vs LLMs trillion-scale; ~50x less compute). In audio, competition shifts from capital/data-center size to model quality. Audio should be treated as a separate “second layer” (sound drives ambience/emotion; “50% of the movie going experience”).
Notable examples
Dog barking, seagulls, car driving-by sound effects generated from a video and auto-synchronized in seconds; George Lucas quote about sound; roadmap from basic generation to professional editing capabilities.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOMirelo AI Overview
0:45 to 2:15
CJ explains Mirelo AI's focus on audio for video content and games.
“the sound effects mostly and possibly also add music if you want that.”
Founders' Background and Motivation
2:15 to 4:05
The founders' journey from AI research to starting Mirelo and their musical backgrounds.
“and it's a very big point for a lot of people.”
Model Development and Use Cases
4:05 to 6:10
Discussion on building their own audio models and current use cases in AI video.
“And you've built your own model here and you've built it with a very lean team.”
Advantages of Lean Development
6:10 to 8:06
CJ describes the efficiency of their lightweight models compared to larger models.
“and not in the trillion-ish parameters, essentially.”
Building Talent in Europe
8:06 to 10:00
CJ discusses the benefits of building a tech team in Europe versus Silicon Valley.
“You'd think also in terms of representations, you know, like, you know, how music is being represented, you know, like in music scores.”
Funding and Investor Confidence
10:00 to 12:10
CJ explains the conviction of their investors based on technology and team.
“I don't exclude to have a team there, but fundamentally I don't see the very big advantage.”
Future of Audio Models
12:10 to 13:40
CJ shares insights on the potential future of audio models in comparison to large language models.
“because the size of your data center defines the size of your model, and the size of your model defines how good it is, right?”
Market Strategy and Target Audience
13:40 to 14:00
Discussion on Mirelo's approach to market their product to consumers and professionals.
The Evolution of Audio in AI Video Generation
14:00 to 17:06
Learn about the advancements in audio editing capabilities in AI video generation.
“to work with audio and without you like the pain of of having to synchronize everything manually and keeping all the fun about iterating on how you want the audio to sound.”
Mirelo's Focus on Sound Effects and Music
17:06 to 19:15
Discover Mirelo's strategy to prioritize sound effects and music in their offerings.
“because the audio layer is going to define the ambience that you want.”
Show all 12 chapters
Hiring Plans and Company Growth
19:15 to 23:28
Understand Mirelo's hiring strategies and the importance of audio in video production.
“And when you look forward to expanding the team from 10 to 20, 30, 40, 50, whatever, who are you going to be hiring first?”
The Economic Value of Quality Audio
23:28 to 24:12
Learn how quality audio can significantly impact video success and revenue.
“but really as a second aspect of video creation, then we've won.”
Transcript
Automatic transcript. May contain errors.0:00CJ Simon-Gabriel:Hello and welcome back to the Scaling Europe show. I'm Seb Johnson. I'm here with CJ. CJ is one of the co-founders of Morello. Morello has just announced a whopping$41 million seed round led by Andreessen and Index. So we have a big round led by some real top tier VCs. And I think what's super interesting is that you're building a foundational model here in Europe. So for those who don't know, can you just give a quick introduction to Morello? Yeah, thank you. Well, thank you for having me here. Yeah, Mirelo AI is, we are focusing on audio for video content and games, essentially, right? So we are doing right now music and sound effects mostly.
0:43And the idea is really you show me your video and we tell you what sound to use, where, and we generate you essentially the audio for it. the sound effects mostly and possibly also add music if you want that. Why was this the business that you decided to build? I mean, both co-founders, my co-founder and I, Florian and I, we've spent roughly 10 years in AI research before. So it was clear that we wanted to do something there. But we are also both musicians. and I spent probably more time at the Conservatoire in Strasbourg than at school playing piano, organ and doing a bit of composition as well and Florian, he was very deep in the electro seat in Berlin and so combining both those things seemed like almost a no-brainer so we decided when we worked at AWS Lablet and we were working on AI in general and first large vision models and later large language models.
1:50And we saw everyone after ChatGPT starting to work on large language models. We were like, hey, why don't we instead focus on something else and focus on music and sound? And that's how we thought about, hey, let's go start Mirelo and do something in that area. And very quickly, this morphed into, hey, let's not just focus on just music, but actually let's do all the audio for videos because there's a lot of need there and it's a very big point for a lot of people.
2:18CJ Simon-Gabriel:And so you launched, I think you started one of the company two years ago. What kind of use cases have you seen people kind of using it for already? So first of all, we've been in stets for quite a while. So we've actually, because we train our models, our own models, it actually required quite a while until we got our models ready. First of all, because at the beginning, was just three people, two founders and a founding engineer. So we first had to build our team, then train those models. And we have two models now. We have a music model and we have a video to sound effect model. And we are very happy that they are actually extremely good on their evaluations.
3:02Actually, despite having much, much less capital to do all of that, they are actually best in class, especially the video to sound effect model. even though we're competing with very big other labs on that area. So we first had to build all those things, and we've just started basically the productization of it and creating our Mereo studio. Now the current use case mostly is to create, to add sound effects and a soundtrack and possibly also a music track for AI videos, right? That's what we're seeing it being used most for. But our goal in the long term is also to address all the audio in any kind of visual and video content, meaning all videos, games.
3:51And while so far it's mostly AI videos and AI creators that actually use our software, in future we definitely want also to create a tool that can be used by professionals as well.
4:05CJ Simon-Gabriel:Amazing. And you've built your own model here and you've built it with a very lean team. Why did you decide that you needed to build a model as opposed to sort of using like a multi-modal stack? Two years ago, there were almost no audio models around. So you had to, there was almost no choice. And for us, this was actually a very, very good thing because it meant like that by focusing on audio, so that we could actually, we can and we could and still can actually build a real mode because there's just so much less research happening there and so much less actually other labs focusing on audio especially if you think about sound effects and music specifically you know now there's a bit more a bit more uh speech or so is is taking off a little bit but but even sound effects is very very small and two years ago it was even smaller so it was i mean first of all there was almost no choice at that point like if we wanted a good sound effect models we had and good music models we had to build them ourselves and the other thing is that's our mode you know it's actually a very big opportunity yeah that makes sense i read that your your models are really lightweight right and they're able to they require 50 less compute than like a typical large language model how have you able to build that well that's actually the other big advantage of of audio is that in general those models are much smaller so we've of course also invested a lot in making them even more efficient and that depends a lot on you know like on on the codecs that you use because meaning how you the codec is if you think of it is sort of the tokenizer for for for for audio so basically how do you represent audio so that it's readable by the machine so if you if you can be more efficient on how you how you encode your your music so your model can become more and more efficient but I mean the bulk of it is just that most audio models are just so much smaller than large language models if you look at everything that Most, for example, text-to-speech models as well, the number of parameters is typically between 1 and 10 billion parameters and not in the trillion-ish parameters, essentially.
6:15So that's another big reason why it makes totally sense to work on that because you will not have those crazy compute expenses that you see with large language models.
6:28CJ Simon-Gabriel:And what role did your co-founder's musical background, Do you think play and how you train the model and you're kind of developing audio and sound? Well, first of all, it's a huge motivation. Like you need to be passionate about this, especially when you build your own company. It is quite, it is not as easy as you would think sometimes. You really, it is, it is, it is like people say it's a roller coaster. It definitely is a roller coaster. So if you don't have this intrinsic motivation of, hey, that's what I want to do and that's where my passion is, you know, that I think it's going to be very, very hard.
7:08So just as a motivation and as a driver, it has helped a lot. This also helps a lot actually for hiring because it turns out many AI scientists are super excited about music in general. they are either musicians themselves or they listen to a lot of music when they work or whatever. Many scientists in general are very much into music. And so when you tell them, hey, you can come here in Europe where there's not so many companies training for their own foundational models. So you can come in Europe, in a company, train your own models and combine it with your other passion, which is music. Many people just love that, right?
7:49And especially here in Europe, there's not so many other opportunities where you can have news where you can do that. So that also helps to basically get very good, excellent people right into it. And finally, maybe the third part of my answer is when you've worked a lot on music, obviously, you have also a certain perspective. You think also in terms of harmonies. You'd think also in terms of representations, you know, like, you know, how music is being represented, you know, like in music scores. And that can certainly also have an impact on how you want to, how you think about building your architecture and representing that music, which is a core part, actually, of the IP that you build when you train those models.
8:40when you talk about uh hiring talent and building in europe were you tempted at all to to relocate
8:47CJ Simon-Gabriel:or build from somewhere else where there's there's a greater degree of or people say there's a greater degree of technical talent oh there's there there's no no such place i think europe for that i i literally think like like like people when they say that they think of san francisco or the west coast you know uh typically uh i think it's probably the building like putting your tech team in the west coast why would you do that when you can put it in in europe you know i think that the people are as good the scientists are as good in europe maybe it's a little bit less dense but they have so much less avenues where they can go and and deploy the talent you're like so when you one of the few it's it's it's uh uh like actually you have more choice i think here than than you would have in san francisco also they get much much less poached right if you if i talk when i talk with founders in san francisco like like constantly you have you have poaching stories it's it's totally crazy in europe it's you have that too uh but but on a much much lesser scale um so no i mean there's there's especially for the tech part you know i don't see any reason why why we would put it in San Francisco or have it in San Francisco rather than in Europe.
10:02I don't exclude to have a team there, but fundamentally I don't see the very big advantage. The answer is different if you start thinking about go-to-market and so on. Obviously, there's a lot of concentration of startups that are there, and that's very interesting from the go-to-market aspect. But for the tech team, I think we have all we need in Europe. So the only thing that's often missing is capital.
10:29CJ Simon-Gabriel:Well, let's talk about that, because you've raised$41 million, co-led by Index, who are a very multinational global firm, but who have their roots here in Europe. And of course, Andreessen, who are an amazing US firm. You know, two tier ones, co-leading a really large seed round. What gave them the conviction to invest? The tech. Like, I think it's essentially the tech and the team. because we've really managed with a ridiculously small amount for training a foundational model to train something that's leading benchmarks by far even when compared to large companies such as Tencent or Sony models, etc.
11:20That's amazing.
11:21CJ Simon-Gabriel:Do you think this is the future? You know, we're seeing a lot of, you know, there's the big news about the Red Alert and OpenAI and how all the other big models were catching up to them. And what you've managed to build is something that is leading, but with a much smaller, much leaner team with, I imagine, a lot less capital. Do you think that's going to be the future of all kind of frontier models? I don't know if it's the future for all frontier models, but at least in audio, there is a good chance that it's going to stay like that. Because if you look at it, typically all those audio models that you look at, the size is not exploding, right?
11:55So basically means you don't have a real gain in increasing the size of your model, right? And so that means that it doesn't help you. It's not like in large language models where essentially it's mostly the only question about is how big your data center is because the size of your data center defines the size of your model, and the size of your model defines how good it is, right? And so the competition is mostly about who can raise most money to create the biggest data center and then to create the biggest model. That's not at all the case in audio, and that's very, very good because it means suddenly you're competing not on the amount of capital you have, but rather on how good you are at developing those models, right?
12:42And that's something that's much less capital intensive and where startups have a much better chance to hold strong, even against very big labs. In a sense, I like to say big tech loses its main advantage, which is, let's say, have hundreds of billions of cash flow that they could put in the interest. But that just doesn't really help here.
13:06CJ Simon-Gabriel:Yeah, interesting. And looking forward, you mentioned that you're starting to do kind of like take the product to market. who are you looking to partner with on this product are you going to go direct to consumers and have them you know add audio to some videos that they themselves are generating or making are you looking to partner with other large companies who are building their own videos what does that look like um both both so we have and we are here we are very open to both we have on the one hand we have miralo studio which is which is direct to consumer as i said the goal currently is it's it's designed currently for ai creators and prosumers and people you know like that that that want to to to that are not sound professionals but need sound for the videos and want to get very good quality sound sound for their videos but in the long term sort of we also want to we want it to be a tool that also addresses professionals right and just gives them new ways to work with audio and without you like the pain of of having to synchronize everything manually and keeping all the fun about iterating on how you want the audio to sound.
14:11And right now, I would say this will still need some development. We will need more editing capabilities in our models, maybe also increase the audio quality yet, because we are not yet a Dolby Digital kind of quality. But that's all going to come. And I think it mirrors a little bit also what happens with AI video generation models, which is at the beginning it was very basic it was sort of text to video maybe one image to a video but with very little control actually about about sort of the camera angle and and changing whatever has been rendered whereas now we see stuff with those new generation of models you have more and more editing capabilities that sort of you can change sort of certain aspects of certain persons like change one object replace it with another one then all that is all that is also going to come for audio and as we move from this basic capabilities which is just you show me the video and we create you all the sound up to all those editing capabilities adding all those editing capabilities it's going to address also more and more people from currently more like ai creators or amateurs up to sort of professional studios once like we have all those editing capabilities so that's my studio sorry and then we have the api our api so we're completely happy also to sell our model to other platforms, video generation platforms, all AI video generation platforms, I think they should all think about audio as something separate, not just in the context of, not just as an afterthought for the video, because it is not.
15:45Audio is 50 % of your video. And that's exactly George Lucas said this, like sound is 50 % of the movie going experience, at least. And that's absolutely true because sound drives the ambience and the emotions you feel in a movie, right? And so if you get it wrong, people will feel the wrong emotions. You can take the same video. If you change the sound, you can completely change the ambience, right? And so it's something you really need to think of as separate. And for sure, now the new generation of AI video models, they start to get a bit of sound. For us, that doesn't change anything because sound is always, you need to think of it as in a second layer afterwards and you will want to iterate it, you will want to edit it, you will want to change it because also our ear is so sensitive to it just because it's the way we communicate as well, right?
16:43And so historically, it has always been like that in the video and the cinema industry. It's like first you shoot the video and you try to basically get as little sound as possible except the dialogue. And then you start, once those images are in place, then you use a completely different stack of software and foliar artists, and you add the audio layer, because the audio layer is going to define the ambience that you want. And that, I think, will not change. That will continue. And we are the ones that want to own that part of the tech stack. But obviously, if video generation companies want to use us and want to build that second part also into the platform, we are more than happy here to also give them access to our models.
17:32CJ Simon-Gabriel:And are you focusing equally on sound, kind of like sound effects and music? Or are you focusing more on the sound effects initially and then you'll be developing music on the side? Look, we are roughly a 10-people company. So we need to focus. We've started a bit with music at the beginning because it was sort of what we were most passionate about. And that's where we were coming from. I mean, the founders at least and sort of the few founding employees. Most of us are musicians actually too. So that's where we started. But then very quickly, we saw also that need for sound effects because no one else was doing that at all.
18:14And it turned out that there we also had the most traction, probably also because it differentiates us most from people. So just maybe for what we have there, it currently is a model where you give me the video and we generate you in a few matter of seconds. It's faster than real time. We're going to generate you all the sound effects for that video. So basically the dog that's barking, the seagulls in the sky, the car that's driving by, et cetera. And we automatically synchronize also all of that. you know um and uh that's basically where also we suddenly now saw some more traction so that's why we decided for now we focus on that but with a new capital now we have the means to we finally have the means basically to to also hire new people and hopefully work on all those different different uh works um parts of the technology and and certainly for we want to own all the audio so So music is a part of it.
19:16Sound effects is another part of it. So we work on both.
19:19CJ Simon-Gabriel:And when you look forward to expanding the team from 10 to 20, 30, 40, 50, whatever, who are you going to be hiring first? Oh, we need to hire on all fronts. So we are hiring. The core part of the core mode of our company so far is the technology. The technology we've developed and our know-how in audio specifically. so definitely we want to increase the amount of research scientists at least double the size of the model team if not triple it but obviously now that we have the technology we also need to make sure that we are building the cool Mirelo studio and improving Mirelo studio and possibly build other products so we are also developing the product team currently it's only two people we want to get at least very quickly to six people and then see if we need more and then the third aspect is a go-to-market go-to-market aspect so we need we want we want people to help us with marketing size with the growth size possibly also with the sales size because as i explained we have two things we're selling two things right mirale studio and our api and the kind of like one is b2c or the other one is much more b2b so we really have both aspects um yeah and then everything around it that you need in a company as well when you when you start scanning when you look forward you know what what needs to go right what do you need to achieve for the next 18 24 months to be a success um we actually want to see more and more people using using uh using a studio but i think we we will win if people understand how important audio is for you for a video and most people simply just don't realize it today it always comes as an afterthought and it comes as an afterthought both for the people who create the videos the AI video creators or in general even the YouTubers typically you start by focusing on the content on the stories what you want to film etc and then you have your film and at some point suddenly at the end, relatively at the end of the process, you are, oh, and now I need audio, or now I need a music track, now I need that, and it's always like, oh, the thing that's really annoying that you don't know how to do, but that you have to solve within the next two days, because in two days, you wanted to launch your video, right, and it's actually totally banana when you think about, when you think about the fact, about really the sentence that it's 50 % of you, of the success of your video, is the quality of the sound, is the quality of the music track, right?
22:08That's why the very big Hollywood movies or the AAA games, they actually start thinking about the audio almost before the game or at the very beginning of the game development and not afterwards. For Hollywood movies, it still comes afterwards, but still they spend a lot of time in sort of creating the right soundtrack, the right music track for the video. And that's if people, like in a year or two, if we see more and more people sort of starting to value that and starting to understand like that the importance of the audio then i think we've won we've won also because you know um it also means that that the you sort of you can recognize also much more the economical value like of of audio because what i'm saying is like if 50 of the success of your video is is the sound It also means that if you get the right sound, you are going to get more clicks and you're going to get more revenue in the end.
23:07So there is a very big economic aspect to this. And if we manage to really make people understand both the creators, but also the video gen platforms, that audio is super important and that they need to have a proper thinking about how they integrate it. and not as an afterthought, but really as a second aspect of video creation, then we've won. Because then many people will need, even more people will need audio, will need good quality audio, and we'll be there basically with the best possible models to serve them.
23:45CJ Simon-Gabriel:Amazing. It's super interesting. Yeah, your job is not just to build and sell a product, but it's to educate the market about the value of how important audio is. Well, look, CJ, thank you so much for joining me. I think this is an amazing testament to what we must have built over the last couple of years. I think it's an amazing story for Europe that we can build amazing models here in Europe and access the capital required to scale them. So, look, thank you so much and best of luck with it. Thank you so much for the interview. And go and try me out.
From the publisher
The team at Mirelo have built the market leading foundation model for developing sound for videos.
The founders have a unique background of being musicians, academics, and having experience working in the big tech labs.
They've just announced raising $41m from Index Ventures and a16z which we get into to.
