20VC: Mistral's Arthur Mensch: Are Foundation Models Commoditising | How Do We Solve the Problem of Compute | Is There Value in the Application Layer | Open vs Closed: Who Wins and Mistral's Position

29 Apr 2024 · 50 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: The Twenty Minute VC (20VC) - Episode with Arthur Mensch

Episode Overview Title: 20VC: Mistral's Arthur Mensch: Are Foundation Models Commoditising | How Do We Solve the Problem of Compute | Is There Value in the Application Layer | Open vs Closed: Who Wins and Mistral's Position Host: Harry Stebbings Guest: Arthur Mensch, Co-Founder and CEO of Mistral AI Date: [Insert Date Here] Funding Raised: Over $520 million Valuation: $2 billion

In this episode, Harry Stebbings interviews Arthur Mensch about his experiences, insights, and the future of artificial intelligence, particularly regarding foundation models and Mistral's positioning within the competitive landscape.

Key Topics Discussed

  1. From Models to Team Building: Lessons from DeepMind
  2. Key Learnings:
  3. Arthur learned the importance of small, agile teams; a team of five can often be more productive than a larger team of 50.
  4. Discussed the need for sufficient uncoupling within teams to avoid inefficiency while maintaining some level of shared infrastructure.
  5. Impact of DeepMind:
  6. DeepMind's culture and practices heavily influenced how Mistral was established, particularly in optimizing team performance.
  1. Scaling Mistral to a $2 Billion Valuation
  2. Mistral 7B Model:
  3. The success of Mistral 7B highlighted the importance of efficiency in AI models.
  4. Key lessons included the significance of targeting developer needs and maintaining a balance between research and sales.
  5. Barriers to Growth:
  6. Current bottlenecks include limited compute resources, with Mistral possessing only a fraction of competitors' compute power.
  1. Winning in AI: Open Source vs. Closed Models
  2. Open-Sourcing Decisions:
  3. Discussed the strategic reasons for both open-sourcing and closing certain models.
  4. Arthur emphasized that open-sourced LLMs may shift the focus from just the model to the application layer and customization.
  5. Cost of Compute:
  6. Arthur projected that while compute costs would decrease, they would not reach zero, highlighting the ongoing challenge of balancing model quality and operational costs.
  1. The Future of LLMs
  2. Model Quality Bottlenecks:
  3. Data quality was identified as a significant challenge impacting the performance of AI models.
  4. Discussions on whether future models will be more generalized or focused on specific applications, suggesting a trend toward vertical-specific models.
  5. Profitability of the Application Layer:
  6. There is optimism about the potential for monetizing the application layer through developer tools and platforms that facilitate AI model customization.
  1. The European AI Landscape
  2. Challenges and Opportunities:
  3. Arthur highlighted the slower pace of AI adoption in Europe compared to the US but noted a growing interest and support from executives.
  4. Emphasized the need for European talent retention and the development of a robust venture capital ecosystem to support AI startups.

Key Takeaways

  • Team Efficiency: Smaller teams can lead to faster innovation in AI.
  • Navigating Growth: Mistral's focus on efficiency and developer needs is crucial for scaling and maintaining relevance in the competitive AI landscape.
  • AI's Future: Anticipation of increased demand for vertical-specific models and the importance of open-source platforms for fostering innovation.
  • Enterprise Readiness: Enterprises are ready for open-source AI, but require better tools for customization and deployment.

Conclusion Arthur Mensch's insights reflect the rapid changes occurring in the AI landscape, particularly with foundation models and the dynamics between open and closed systems. Mistral's strategy emphasizes efficiency, developer engagement, and a keen understanding of the evolving enterprise demands. The episode provides a comprehensive look at both the challenges and opportunities facing AI startups today, particularly within Europe.

---

This summary provides a clear structure to understand the discussed topics, key insights, and takeaways from the podcast episode, making it accessible for readers interested in venture capital, AI development, and startup growth.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Do you feel like you have enough cash now? I guess the start -up is always fundraising. Do you think enterprises are ready for open source? The most technical saving enterprises are definitely ready for it. We know that to widen the adoption, there's definitely some tooling to be brought to the market. What are the biggest barriers to mistrial today? We are still bottlenecked by compute for sure, but that's because we don't have many of it. We have 1 .5 THN, which is a few percent of our competitors. Can I ask, Juanley, was it a mistake for you to not scale that quicker? You can't really scale that much quicker.

0:27You can't raise like 2 billion on the seed round. I mean, at least you could have done it in 2023. Which competitive do you most respect an admire? We are surprised by the **** recently. It's a competitive year on SK. We are always mixing it up with new intro styles. Let me know what you think of that new intro style and what a show we have for you today on 20VC. Mral is one of the most exciting AI companies today at the forefront of the foundation model charge. Joining us is Arthur Mensch, co -founder and CEO of Mral, where he's raised over $520 million dollars in funding from the likes of Andres and horror as general catalyst, light speed venture partners and Microsoft, and before founding Mr.

1:07All Arthur was a research scientist at DeepMind, one of the leading AI institutions in the world. But before we dive into the show with Arthur's day, we're all trying to grow our businesses here. So let's be real for a second. We all know that your website shouldn't be this static asset. It should be a dynamic part of your strategy that really drives conversions. That's marketing 101. But here's a number for you, 54 % of leaders say web updates take too long. That's over half of you listening right now, and that's where Webflow comes in. Their visual first platform allows you to build, launch and manage web experiences fast.

1:43That means you can set ambitious marketing goals, and your site can rise to the challenge. Plus, Webflow allows your marketing team to scale without relying on engineering, freeing your dev team to focus on more fulfilling work. Learn where teams like Dropbox, IDO and Orange Theory, Trust Webflow to achieve their most ambitious goals today at webflow .com. And speaking of incredible products that allows your team to do more, we need to talk about Secure Frame. Secure Frame provides incredible levels of trust to your customers through automation. Secure Frame empowers businesses to build trust with customers by simplifying information security and compliance through AI and automation.

2:24Thousands of fast -growing businesses, including NASDAQ, Angel List, Doodle and Coda, trust Secure Frame, to expedite their compliance journey for global security and privacy standards, such as SOC2, ISO2701, HIPER, GDPR and more. Back by top tier investors and corporations, such as Google, Client of Perkins, the company is among the Forbes list of top 100 startup employers for 2023, and business insiders list of the 34 most promising AI startups of 2023. Learn more today at Secureframe .com, it really is a must. And finally, a company is nothing without its people, and that's why you need Remote .com.

3:02Remote is the best choice for companies, expanding their global footprint where they don't already have legal entities. So you can effortlessly hire, manage and pay employees from around the world, or from one easy -to -use self -serve platform. Plus, you can streamline global employee management and cut HR costs with remote free HRIS. And hey, even if you are not looking for full time employees, remote has you covered, with contractor management, ensuring compliant contracts and on time payments for global contractors. There's a reason companies like GitLab and DoorDash trust your remote to handle their employees worldwide, go to remote .com now to get started and use the promo code 20VC to get 20 % off during your first year.

3:46Remote opportunity is wherever you are. You have now arrived at your destination. Arthur, I am so excited for this. JC introduced us quite a long time ago now. I've known you for a while. I've been wanting to make this happen for a while. So thank you so much for joining me today. Thank you for having me. It's a pleasure. The pleasure is mine, my friend, but I want to start. What would your parents or teachers have described the young Arthur? I'm just always intrigued by the characteristics and traits of the best founders. How would they have described a 9 -10 -year -old author? I guess I was a bit curious and a bit stubborn, not very nice to my brothers.

4:23I was the eldest of them also. I don't know, you should ask them. I think they have good memories, hopefully. Do you know what Sandy or Mother wasn't in our reference list? So we missed that one out. But, you know, JC provided some great commentary. So I do want to start there also, you know, growing up, what was your first exposure to AI? You're a kid in France. How did you first get exposed to AI machine learning and what was that first Passion point that was Andrew Eng flying an helicopter helicopter backwards. It's a control problem Which is not easy to solve and I'm not sure if it was really AI related I think he was saying that he was using a normal network to control over this But that's the one of the first memory of me being shown what you could do with Machine learning at the time that was in 2020 2013 I think Most recently, you spent two and a half years, three years at DeepMind.

5:14What are the biggest takeaways for you from that experience? How did they impact how you think about building Miss Rale? A team of five is faster than a team of 50. Exceptive, sugar organized, the team of 50 should be 10 teams of five that are sufficiently uncoupled. One finding that I learned the hard way at DeepMind and the reason why we created the company in a slightly different way in terms of organization of the science team. And also the reason why we knew we had a chance to do interesting things with a smaller team. Can I just ask sufficiently uncoupled, do you not lose efficiency or is there not leakage between those silos and it creates actually inefficiency by having such silos?

5:54You have to share some things. So you share the infrastructure, you share the code base, you share findings. We are doing general purpose models and general purpose models. You need to evolve them in different directions. So you need to make them speak different languages. You need to make them be able to code, be able to do mathematics, be able to reason, you need to add multi -modality to them. All of these things are loosely coupled. It's useful if you use the same framework for optimization, for data, for training, but you don't want to have your team spend their entire day in meetings for coordination.

6:25It's actually pretty hard to figure out. I think so far we've managed to scale it relatively well. The team is only 25 people, so that's actually not super challenging. It will become more and more of a challenge. That's what I remember from DeepMind. It worked very well at the beginning. JimiNay was a bit too slow and I think they recovered sufficiently well since. We have optimized the team to be as fast as possible and to ship as fast as possible. Was it an easy decision to leave to start, Mistral? You're at DeepMind to one of the best institutions in the World Fair. I was some incredible talent around you.

6:56Was it an easy decision and just take me to that moment when you decided to leave to found or co -found Mistral? So it's not a zero to one decision. It's not a binary decision. You start to think like I'm 10 % leaning on living and then it grows and at some point to cross the threshold then you say, okay, well, I guess, no, that I'm sufficiently decided there's no way in which I stay more than a few days. Otherwise, I wouldn't be candid with my colleagues and so that's how you get started. You say, there's no turning back any tuning. What was that point for you? That point for me was probably around much end of March at last year where I decided to leave on Friday and I resigned on Monday.

7:35You can't stay if you've decided to resign all the way to Stodbury Fair. I totally agree with you. I do want to run this with some chronology. I spoke to so many of your advisors and investors and I want to start with actually the first model, MIR Mischar 7B, being one of the most popular released a while ago now. Why do you think it was What did you learn from that? I think it served two purposes. So the first was to show that there was a lot of slack in compressing models. And so from a scientific perspective, it was a good finding and a good learning from the community. It also filled the gap in the efficiency to performance to the space of models where there was definitely something missing.

8:217B is the size that allows to run efficiently a model on your MacBook or on your SmartSaven. And we made it sufficiently smart so that it was still useful. So there was already seven B models before, but there weren't good enough to do interesting applications. And so by targeting this specific space, we talked to the developers immediately, because developers, like the casual developers, running on a gaming GPU or on its MacBook. So it created a lot of curiosity and adoption because it was a missing spot in the performance to efficiency space. When you look at like lessons from that and how it impacts future releases, any that really stand out for you?

9:03I guess it taught us that there was a lot of interest for efficiency other than scale. And so that's why we continued of targeting very efficient models with the mixed trial at 87B and more recently mixed trial at 8622B, ensuring that for certain then a cost and for a certain size we were reaching the top performance of the market. That has been our major motivation for us to target efficiency. Well, simultaneously scaling towards the larger and larger. I spoke to Sarah Groyd before the show and she said the core question that I think is, you know, with the focus on efficiency and the efficiency frontier, does scale matter?

9:38Well, scale matters in the sense that you should spend more training on computer to make the models more compressed. So you do need to have some computer to compress models. No, scale isn't the only only ingredients to the recipe you need to scale, but you also need to have proper data. Otherwise you reach some data quality limit. You need to have proper techniques for training. I mean people call it computer computer player, I guess. How do you actually make some efficiency gain that are not costing you computer because computer is expensive? And so one of the things that we do at Mralist try and harvest is computer computer players.

10:10Can I ask, in that chasm of efficiency gain without costing more compute, is there much more efficiency we can e -count? Is there a lot for us to e -count, or are we working really at marginal improvements already? I think it's an open question. I believe there is. I believe we can make models that are much better for a certain size. But it's as open a question as can you make a much better model on the same kind of data by making it bigger and training it for longer. Things you need to discover them also on the way. You can try and predict the kind of performance you reach out. You will achieve at the end of the day You need to try it out.

10:44So that's the I mean, it's really much a research field You need to do the research and you need to try things. So I asked Sam Altman this question What is the end state for the model landscape? Most people say ah? It will become commoditized and actually there will be 12 players and it will be erased to the bottom What is the end state for models in your mind and how do you think about the commoditization question? I think the end state is to have more features on developer platforms that allows to do customization, that allows to make low latency models that serve a certain purpose, that allows to evaluate them and to improve them over time.

11:18And so the model is only like a tiny part. I mean, it's a central part, but it remains a tiny part of an application. And what you want to do across time, and when you deploy an application that you expose to users, you want to ensure that it works and ensure that it's latency reduces over time and ensure that its quality increases over time. And so I think that the end state is, models are effectively going to be a starting point for any AI application developer. They need to be surrounded by tools, by a lifecycle management platform, basically, and that's the one thing that we started to build.

11:50Like general purpose models are a bit indifferentiated, but the differentiation that you need to create for your application comes from the data you put into it, the user feedback that you gather and the intelligence that you have to figure out what the application should be doing. And that is not committed to all this. No recipe that allows to go from a general purpose model to model that is super good and better than all of the others at your specific task. This is a missing piece in the puzzle and that's one of the aspects where we're putting our strength on the product side. Sam and Brad are the other day that models just aren't actually that good at any, like yet, and they need to improve a lot in quality.

12:27What are the largest constraints or bottlenecks on model quality today and what needs to change for them to improve? I think the data quality is a constraint. How do you leverage the entire world knowledge and ensure that the model follows a certain path toward learning more and more complex things? That's a very important part and I think it has been a neglected part. There's obviously compute, but given the amount of data you have, we have a tent. Compute is already running into, it's no longer the bottleneck. The bottleneck is more the data at that point. You should get text to text models.

13:01And so the question is, how do you refine the data and how do you feed very high quality data to the model itself in order to improve it over time? And I think in that setting, it becomes a bit, one bottleneck that is associated to bringing better model performance is the question of how do you evaluate these performances. You need to have very good evaluation that targets very specific topics. You want the model to be good at helping diagnosis in hospital but in French. And oftentimes you're a bit out of domain compared to the data you have. And that's where you should identify a gap and you should try and fill it out.

13:37The pushing the model capabilities become also a question of mapping where they're failing and fearing out ways of improving it. For instance, they are failing at mathematics, how do you improve their mathematics thinking, how do you improve the way that demonstrate theorems. The answer to this is very different from the way you answer the question to how do you improve the medical diagnoses in French violence. Well, we see large scale generalized models that are able to answer huge swathes of very complex problems. What do you think we'll see much more vertically specific, smaller, more specific models that are much more vertically aligned.

14:14Yeah, we believe that. And actually these vertical models are not going to be out there. They're going to be built by the application makers because the only way you can make a low latency model that is super good at the specifics is to get rid of the general purpose aspect because a general purpose model is bit bloated, you can think about everything. But if you want your model to think thoroughly about a specific topic, so that you can call it in your AI application while maintaining a good user experience with low latency. What role do you play in that world? If it's actually in the application where you have that specific model creation, where that kind of value occurs, where do you play in that?

14:53It's a very hard job to make a specialized model. So it's actually very tight to the way you create a pre -trained model. And so bringing the tools that allow to do it in a full proofway, so allowing developers to create customized model that are performing very well at that task, but that doesn't require expert AI knowledge, which is hard to find, is definitely something where we're insisting. So I'm an investor today, and I'm pleased that you just said that there will be value accrued at the application layer, because I look and I worry that Blondney everything is going to get steamrolled by some of the players that we mentioned.

15:32How do you answer the question of will value occur at the application layer? And for me, as an ambassador, say Arthur, you know me. How would you advise me? There's two opposing directions. The first is that the models are getting better and better. So it means that creating a verticalized application, as long as you have the data for it and a good understanding of the use case you're facing, is going to be easier and easier if you have access as to the tools that facilitated. So that's the first aspect, which would make me think that the application layer is going to grow thinner and thinner.

16:01But then there's also the fact that the models are getting cheaper and cheaper because we managed to compress them, because we make a lot of improvement on their efficiency. And so that means that effectively, vispeless the competitive pressure there is on the model layer means that the price around the model, the dollar -parent intelligence unit, let's say, is definitely going to reduce. So there's two aspects of growing ability, compressed price, which on one side says that the application layer is going to grow thin and on the other side says that the model part is going to grow thin. So for us, the project that we are taking is that the model part is still going to be big enough and that we need to build this platform on top of that.

16:41Because that's where we are going to enable all of the vertical applications that will be interesting for humanity. How do you think about that positioning in brand? Because there are other players who are much more direct in saying, hey, we're going to dominate a lot of different verticals and kind of be afraid. How do you think about that, enable a two vertical applications or not in that position? We are not a verticalised company. We started to bring value to developers and to bring freedom to developers. So when we started, there was basically one API out there, soon too. And the field of generative AI was starting to look like it would be very centralised around a couple of players.

17:20And we took this platform approach where the model that we're making and the technology that we are making, we are allowing developers to own it, to modify it. And so bringing freedom to developers and AI application makers is I think the best way in distributing generative AI as widely as possible, which is our objective as a company. Making AI ubiquitous, bringing frontier AI into everyone's head is the reason why we started. We did a good job at it, but this open source part was, I believe, a good enabler for the community and made people realize that they could build very interesting technology by modifying the models themselves, instead of depending on the APIs of a couple of providers.

17:57Dude, what do AI developers care about? Everyone kind of gets on Twitter and goes, oh, did you see Axe's performance this week is better than Y's performance last week? What do they care about? Efficiency, scale, cost, what drives that usage and decision -making? They care about cost for sure. They care about customization, being able to modify the models that will. And on that aspect, I think we are only scratching the surface of what can be done. Like the fine tuning aspect that has been like the go -to solution. It's probably a little too low level from what we should be doing. They care about being able to deploy anywhere.

18:35So they operate in a certain space, in a certain cloud. They might be operating on prem. They might have some edge devices to deploy to and they want to be able to put that technology there. And so they also care about portability, which in turn offers data control. Usually LLMs AI becomes very useful when you connect it to knowledge bases or to anything that is related to certain business. In that respect, it becomes a very sensitive part of your application because it sees everything. It sees all of the data you have. And so enterprises, for instance, do care about ensuring that the property data they have is accessed in something that they can completely secure.

19:15And that's the reason why we deployed our platform on Azure in AWS, for instance, that is bringing the security layer that they need. We're going to get to enterprise. Can I just also just brand matter in this segment? You know, when we think about building brand, both in terms of developer adoption brand, corporate brand, is brand a large determinant of adoption in this segment? Brand seems to be critical and this is something that we have learned on the way. We use certain models because they are known to be good. You can't afford to evaluate everything out there. And so having some form of community watching is super important.

19:53The part we took with APHD distributed models has contributed to what I think has become, Well, at least a known brand and we believe that it's definitely going to be important. Brand is important because trust is important in that domain. And open source brings trust in terms that provide some trusted brand. You mentioned the word open source, then. I'm going to get to that. I do just want to touch on that you mentioned cost also. I want to touch on cost. How when and who will make marginal revenue that exceeds marginal cost in LLM based products? You should be telling me you are the investor.

20:30You have been your own company. That means I know nothing. I know who is doing the most margin at the moment. It's probably going to be over a bit more time. Who is doing the most margin at the moment? Nvidia is at that point. The club providers are pretty much at cost. And the club providers, we are not at cost. Hopefully, but the margin that are known to be lower than the typical software margins. AI application makers. Some of them, the one that almost used, seems to be doing a pretty good margin. I think it's going to be quite a moving space. As I've said, the capacity of models makes the cost of making an application lower and lower.

21:06I don't think there's any way in which the margin at cost and the margin of the most important part of that technology, which is really the foundation layer, become zero because otherwise there's definitely going to be a furnace problem. What do you mean by the furnace problem? Talk to me about that. Usually the value tends to accrue where most of the difficult part is and most of the defensibility is. It was for a while it has been on foundation on what is. I think it's obviously evolving with time and there's no mode that isn't disappearing or evolving with time that will remain the part where most of the innovation will be made and where most of the, well, it is the significant part of the crude value will there.

21:43Well, the value will accrue. Is there actually much of a barrage creating a foundational model company today? I know that's a really broad, stupid question in many respects, but you have so many different players now and new ones popping up every day. Is the barrier just reducing day by day? I don't think it is. To be relevant in that space is a very hard topic. You need to be dominating on the cost efficiency performance by ToFront and there isn't, there's only a few companies that are currently well positioned. So you can try and do something but if it's not relevant, if it's strictly dominated by another model or another technology, then you have a problem.

22:18There's a few barriers that are pretty hard to face. You need to raise sufficient capital to have enough compute and be relevant. You need to have people that knows how to train models, which is still the scarce resource. And then you need to have a good brand, because as you've said, it's highly competitive. And this is not something that comes out of thin air. So I think there's still a lot of defensibility on the market. Although there is a lot of noise, which is different. How quickly does the cost of compute go down, do you think? Because if you look at those things, actually, you said cost of compute, access to talent and brand, if we drastically bring down cost of compute, like many think we will very quickly, you've got access to talent and brand.

22:54Two of those are more doable. The cost of compute reduces over time, just based on hardware costs. It reduces around 30 % every two years if you show low Nvidia on map. The other thing that creates is the efficiency of algorithm. So if you look at where we train models from three years, three years ago, and the way we train today, I think we have probably made something around the 100 times algorithmic improvement. That's probably where most of the gains were actually made in the last three years. Obviously the cost of compute does reduce, but it doesn't reduce faster than the more low. So our bet is more on efficiency where I think there's a lot of improvement that can still be made.

23:32Given Nvidia's prominence there and Nvidia being the one where the gains are, As you mentioned, is bluntly one of these single most important things, not simply the quality of your relationship with the core provider, being as you're or being in video or being one of these players. Is that not the core determinant of success today? I guess it's an important aspect. There is a strategic dependency from the AI layer on the club providers and on Nvidia. The competition is heating up as well, but it's effectively important. It's effectively useful when you develop a software to also know the hardware provider because they can help you out In optimizing for the hardware issues when you're selling your developer platform to enterprises to bring that platform through their usual provider Which happens to be a cloud provider.

24:18So there's definitely some important collaboration to be made there When my Amazon invest like two billion dollars in anthropic or whatever it was Is that not just like a trade where like anthropic then spend one point eight billion dollars on Amazon and return it back to them. Do you see what I mean? Is it not a bit of a misnomer? It looks like you're on flipping, yes. I don't know about that deal particularly, but it makes sense from both perfectly. Can I ask, how does the unlimited availability of open -source LLM's impact the answer to the above being marginal cost and marginal revenue? Does it change much?

Read the full transcript

24:54It moves the value a little higher than the model itself. It moves the value to the platform in customization part, which is really, I guess, something that we're expecting. And it accelerated that process. You started off completely like very open source, very much open to the community. Now you have small models open and then larger ones closed. Am I right? We also have large models that are open, though. Depends on the threshold for small and large, but eight times twenty to be is actually relatively large by any standard. What was behind the decision then to close some models? Is it just a business case where you need to make money?

25:29Opportunity to grow the business using that asset as something that we are selling. It's still the case that we're growing that our business on top of commercial models in particular. There's also a good way of cementing some strategic relationships with cloud providers. And it's going to be to continue to be the case. We still intend to be a leader in the open source part and to have some unique assets that we can license and to have some unique platform that developers can use. It's hard when you suddenly have some closed and you start building an enterprise team. For you as a founder now, how do you think about our balance between a research team and a sales team and making sure that the two cultures come together well?

26:13I think some important thing is to create empathy. So ensure that the science team also understand the problems that the users are facing. It improves the science because at the end of the day, the general corpus technology we are making is only general corpus if you identify the use cases so that comes back to the earlier discussion We had so ensuring that the science team has some relatively direct exposure to the product and to the business team is Actually important to make them understand what where the model is failing and how it could be improved significantly And on the other side the go -to -market team has to understand it.

26:46It's a very technical sales Sales motion because you you're selling not the product but you're selling something that is going to pull out the product So you need to tell the customer how these things should be used to actually make something that brings value to the business and that only goes through strong enablement of the go to market team. So it's I think it's a challenge. They'll not operate on the same scale. The science team as cycles of several months to go to market team goes faster, shorter cycles, let's say. But I think so far we've managed to go to market people that have some Technical interest and and technical people that have some business interest and I think that's how that's how you ensure that You don't have silos at the end of the day One of my worries with bluntly this space as we move into enterprise is that brand Matters so much in terms of enterprise actually and they already have existing agreements with Microsoft And I worry that actually product or model quality doesn't matter as much as distribution Microsoft just tack on existing clients with new products.

27:46How do you think about that as a core challenge? Am I wrong to be worried about it? I think it's true. Distribution is very important. Short cut distribution is to create demand through open source models. Do you think open source is ready for enterprise? What do you think enterprise is ready for open source? And do they care about it enough? It depends on the enterprises, but some have been early adopters and are using a lot of these trouble models into production. So for sure they're ready enough in order to bring them to the next level of putting things into large scale production, etc. I think they're still lacking some product around like managing correctly load balancing Customizing the models because you can do it with the AY solutions But if you want to make it robust enough and scalable enough, it's actually not easy And if you want to actually increase the quality of the models the custom models the recipe are a bit hard to set So the most technical saving enterprises are definitely ready for it and there's a few there's actually many use cases at hand production using using open source models no in order to widen the adoption.

28:49There's definitely some tooling to be brought to the market. Obviously every enterprise today is sitting in a boardroom going what's our AI strategy? What do you advise them and what questions should they be asking? Start thinking about how they are going to change all of their product using AI as the premise. You think the existence of like very clever agents because you can build very clever agents today, assuming that presence and working backwards to understand the consequences of organization. Not thinking about generative AI as a way to, as a way of increasing productivity in a world processing, but rather as a way to change completely the way you operate your core business, which usually involve taking models and customizing them pretty heavily to create the differentiation that you will need in like five years time when everybody will have adopted the technology in its company.

29:36So my question to you is your in France, I'm in London. We both know that European enterprises do not move very fast. Most do not even have slack today. My concern is that we drastically overestimate adoption in the near future and maybe underestimated in the 10 year 20 year future. Do you think that's the case here and do you worry about the lack of G of a lot of enterprises, especially in Europe in adoption? I mean, it's a general phenomenon in the tech that you always overestimate the speed but under estimate the impact. I think it's probably occurring today. It's slightly different in the sense that there's some executive support for pushing generative AI solutions even in Europe.

30:16So there's some delay compared to the US market for sure. I wouldn't say it's very significant. It's one year maximum in terms of delay. The challenge here is that it's a technology that can take many forms. And so trying to focus on some specific thing that you can bring to the market that have AI in it is a prioritization challenge. And so you need to be very strategic around that. I don't think this is super easy for enterprises generally. It will become easier once they try up of the shelf solutions a bit more. Once they realize that there are some developer platforms that are low to do it without hiring very expensive and hard to find AI sentist in house.

30:55And so we expect that this is going to accelerate in the coming years. Do you think that we're still just playing in the experimental budget game? Or do you think that we're moving into core budgets as well? It depends. It's moving into core budget for customer support, for instance, where like areas where the application of AI is pretty obvious. It's definitely moving into core budgets. It's also at the experimental stage in many other functions. and for core applications in the industry, the telecom industry and healthcare, this is still in the playground, but I think it's going to evolve in the next year.

31:30Kamelowski, as you build our enterprise, it's another expensive thing to build out on top of compute and talent. It costs real money. And I spoke to Paul at light speed before and he was mentioning to me bluntly how much less capital you've raised compared to a lot of your competitors, most obviously open -air, and anthropic. a set in a world where capital equals compute equals quality of model. How does Mischraul keep up and stay relevant? So the good thing is that capital is correlated to compute. The then compute is correlated with quality. It's not completely dependent on it. And as I've said, there's some strong opportunity for providing models that are the best of their class.

32:07They might be sufficient to actually solve certain use cases. That's where we're playing. In addition to playing on the scaling part, because obviously you do need to keep stay relevant. You need to keep your technical team motivated and to keep your technical team motivated. You need to give them the experimental bed they need to make new discoveries and to progress science. And that is where you need compute in addition to growing the model across time. I mean, we're growing our compute like every company. We are convinced that we don't need to grow at the same rate because there's a lot of barriers that are not computer -related that are appearing on the way that we're already seeing.

32:45We think we can scale and we are also convinced that we on the efficiency front we are already very well positioned and we are strengthening that position What are the biggest bias to mistral today? We've had a few delays with our compute providers for sure that has been in barrier So the last answer to your question is to be taken with a grain of salt We are still bottleneck by compute for sure But that's because we don't have many of it. We have 1 .5 kH 100 which is a few percent I think of the capacity of our competitors And so that's definitely a bottleneck that is going to improve significantly in the coming months.

33:18Can I ask, Juan Lee, was it a mistake for you to not scale that quicker with the benefit of the lines right now, which you wish you'd scale that quicker? You can't really scale that much quicker because you can't raise like two billion on the seed round. I mean, at least you couldn't in 2023, but you can, today, but you can only hire that fast. You can only scale your infrastructure to manage more GPUs that fast, then you can only raise capital that fast. There's some acceleration constraints that are pretty hard to fight and that are pretty much the first principles of starting a business. You mentioned about the scaling constraints and cash.

33:52Does it matter where your cash comes from? Does it matter if you have European funded, Saudi funded, US funded? Did you think that matters? I guess governance matters. The way it is important for a young company like us is to be under the control of the founders because there's a lot of things to be invented and vision can only be carried by them. We have very good governance term, a very simple and clean governance that makes us a for -profit company, growing a business to actually push the science frontier. This is something that we're very attached to, being able to control the company, leverage our funding partners appropriately to grow in different parts of the world, in the US, in the EU, it has been critical as well.

34:34So it does matter in the sense that we want to have partners that are supportive and long term because we are in the field that is fast -moving where we don't know yet exactly where the value will accrue and so being flexible and being smart is definitely a requirement when you raise money. Would you take money from Saudi or China? Good question. It depends on the term. China is a bit hard. For us it's even hard to operate in China. I mean, we don't operate in China because you can't really operate in US and China without being like a very very large corporation And so you need to make some choices.

35:10What chance do you think that Europe has an AI? I know it sounds deterministic and defeatist and so you might be going to fuck Harry shut up. But it's like what chance do you think Europe has an AI? And what does it take for us to stand up as a serious AI industry with Europe? I guess the chance it has is that it's a revolution. It's changing the way we do software. And so as every revolution it opens a lot of opportunity for new actors and there's no reason why there shouldn't be an actor That is that was created in Europe that could grow pretty fast and that's the mission that we gave ourselves We have the talent Capital can cause oceans without without too much problem We have the market the market is more fragmented than in the US for sure the ecosystem the digital native ecosystem is definitely smaller But it exists and it's growing there's local opportunity for business development on the talent side, we can hire 23, 24 years old people that we can onboard informants and they operate as well as any software engineer in the valley.

36:09So people are quite talented here. And so if we manage to keep them and to convince them not to go to the US, we have a lot of opportunities. When we look at computers, mobile, cloud, the kind of core technology shifts, the way that it's worked is Europe has kind of seeded control to the US and then just taxed US companies for access to our citizens if one's being defeatist. Is it different now? I mean, Europe is paying the price of not setting up a VC system in the 60s, but setting it up like 40 years later or even 50 years later. And so, I'm by the way, the dirty secret is that the VC ecosystem in Europe is US funded.

36:49Yeah, it was. I think it, is it still? I want to honestly in large part yes there's government institutions which are backfilling it but largely backfilling it with bad players who aren't very good but the best providers in Europe largely US funded by top US institutions. Okay I think yes I've said it takes time for an ecosystem to build so you have layers of entrepreneurs and investors that stack on top of each other. The US has 60 or 70 years of venture capital investments. I think Europe has only 20 years. I mean, it takes time, it takes an incompressible time to build an ecosystem. It takes also some wheel power.

37:24And I think now we're seeing that wheel power. We are seeing entrepreneurs creating companies, we're seeing this is like you, not going to the US. Everything is positive, it just takes time. I'm adamant that we'll manage to do something interesting. On the engineering side, you feel that you have the depths of talent pool to hire from. As you scale now, you do. On the engineering side, on the AI side, we do. we have a team in the US though which is Working on like specific topics like for senior AI scientists You find them more in the value than than in France for junior AI scientists There's a wealth of talent in France in Poland in the UK.

37:59I think one of the strength of the area When you were raising money was it very different speaking to European investors versus US investors? I guess in the seed round. No, it wasn't that different because it was a seed round for the series A which was a bigger round, it was European funds were unstructured to do the kind of deal that we were proposing. We didn't even have a lot of conversation because they just couldn't get their head around the investment that needed to be made, whereas we were approved in your company. Yeah, I think what is lacking and it's related to the ecosystem part in Europe of our growth funds that are able to take huge bets with lots of conviction and that in turn should improve over time, especially if we manage to use European wealth and channel it more into that growth firms than it is today.

38:47I think you have more hope than me on that one. That is not going to happen. We are not going to see many more European growth funds be built in the next few years for sure, not in the next three to five. Yeah, it hinges on a few political decisions. I think it hinges on supply of capital and belief in a future European ecosystem that can contend with other large ecosystems. It's a chicken and egg problem. This could be nudged into the right direction if politics wants to do it. If a couple of companies show that you can actually have companies that grow fast in Europe and that's what we're trying to do.

39:17I'm not too pessimistic. I find you too pessimistic. You should come to France. I think you will get more optimistic. Do you know what ever Prision was telling me that I'm too pessimistic then shit, I really need to be more optimistic. My question to you is, you know, you just mentioned that the speed of scaling. Hardest thing, dude, is scaling with your company at the same speed. What was the hardest thing about your self -scaling as CEO with such speed of scaling of the company? I mean, we are learning on the job. It's effectively, you have organizational challenges. How do you ensure that 45 people communicate well together?

39:52How do you manage your time in terms of representation time? In terms of business development time, because we're still at the stage of the company where we get involved a lot in the deal making aspect. And how do you ensure that you set proper directions and maintain the team in the state of tranquility, despite the amount of noise that there is on the competitive side. The fact that direction is obviously going to be changing over time because there's a lot of uncertainty in that field. So this is, I think, the hard part. I don't think I'm doing it properly, that we are actively trying to find sources of information to learn new things.

40:31If you could call yourself up to the night before you became CEO and founded Mistral and give yourself some advice with the now knowledge that you have, what would you say to yourself, Arthur? Maybe stage a bit more of the product development and go -to -market development. We did start the go -to -market motion at the time where we had absolutely nothing to sell. It did work out. It sense of anything. I think it might have been slightly simpler to save things maybe a little more developing the product a little before developing the go -to market. But since it's such a fast -moving field that we did start everything a bit together with some organization that was a bit lacking and now we are solidifying it on the fly.

41:13It has worked out, it hasn't been optimal for sure. And so in hindsight you can always give me like a few tactical advices on who to I or when. Generally, I think the strategy we had one year ago hasn't changed much. We did realize that we need more capital and that we couldn't have to create only from Europe and that we need to go to the US very quickly. Those were findings that we did on the way. I don't think they would have helped. It would have helped that much to know it a year ago. One do you feel that you have enough cash now? I guess the start -up is always fundraising. It's a field where for the years to come, the investment are going to exceed the revenue by design because you do need to scale and you do need to stay relevant as the Fertia R company.

41:57So effectively, there need to be some investment. The revenue is ramping up, so there will be some revenue to reinvest. But today and for the years to come, the speed for developing research should be faster than the speed at which you can develop your your good market before we do a quick fire when you look at the landscape today which compass did you do most respect and admire? I mean they all delivered we were surprised by Co here recently came up with new good models and I think that was a surprise for us and obviously OpenAI and Entropique and my friends at Google are also doing a good job so it's a it's a competitive landscape and we respect all of them we also all work in the same direction and eventually with the same higher goals.

42:42So it's great to have respect for one another. Is it too late to start one now? We see like hella sticks starting now. Is it too late? I'll stick and know them well. Is it too late? I wouldn't recommend going into the foundation layer business. I know some didn't recommend to do that one year ago and we did and it seemed to have so far a change of you think. So I think it would be arrogant for me to say that there's a new chance for a new competitor to arise and beat us. Listen, my friend, I want to move into a final thing, which is just a quick fire. So I say a short statement, you give me your immediate thoughts.

43:15Does that sound okay? Yeah. Let's do that. So what worries you most in the world today? Global warming. There's a race of the planet hitting up and as finding solutions for it, I think AI is part of the solution, brings more control. It brings potentially more efficiency in some of the processes, but there's a effectively a race for survival. So I think this is something that we should be a bit more aware of. What if you change your mind on most in the last 12 months? I think I've changed my mind on a lot of management premises that I had and that I had never tested the in -roll. What was the biggest one?

43:51Transparent feedback is actually a super useful for a company and so operating in a almost really transparent manner has helped out growing with our breaking. What element has been the most unexpectedly challenging in the scaling of MISTERAL? The amount of demand that we had to manage, which is too high for what we can handle. The brand success, the fact that people know us, was a bit unexpected. We knew that it would be noticed. We had no idea that people would start using us that fast. What do you do to calm down? You have a lot going on now Arthur and you have a lot of expectation and cash on your shoulders.

44:32What do you do to just? I run, I cycle, I think my partner will yell at me but I try to take care of my daughter. Okay, you've recently become a father. What do you know now that you wish you'd known when you first had your daughter? You know very recently. I had no idea that you needed so much energy to care for a small children. What do you think AI will take the world in the next 10 years? Like what does the future of society look like in a world where AI is embedded into everything? Well, it's changing the way people work significantly in the sense that it requires to be more creative and to bring more value being what can be automated.

45:14So it's a very structural change on the job market, which means that there should be some adaptation that are taken pretty quickly in training in education, that people can get a sense of what is going to be expected from them in their daily job, assuming that there's some AI out there. Do you think the fears of job replacement are grossly over -exaggerated? I think they are, I mean, depends on what you're speaking to. I think the job are going to be displaced, for sure. Some will be replaced, some will open up. We're just trying to move humanity to a higher level of abstraction. So we can no talk to machines, and machines can understand, and answer in a human -like fashion.

45:53This is not so much of a paradigm change compared to what we're doing with computers. I think what's happening right now is that probably the speed in our elevation throughout the higher abstraction level is probably occurring at an unmatched rate in history. Though that means that the society adaptation is going to be more challenging, it needs to be anticipated. Final one for you. We do a show in 2034, 10 years time. If everything goes right, where's Mr. Allen? Mr. Allen has some very relevant models, commercial and open source, and it has a very strong developer platform that allows to do everything that you need to create your AI application.

46:33That would be a good achievement. Ah, but listen, I've so enjoyed doing this. Thank you for putting up with me going in many different fast -moving directions. You've been incredibly patient and a brilliant guest, so thank you so much, my friend. Thank you for hosting me. What a fantastic guest table in the show, I want to say a huge thank you to Arthur for being so patient with me there and for being so open with some of those answers. If you'd like to see more, you can of course find it on YouTube by searching for 20 VC, but before we leave you today, we're all trying to grow our businesses here.

47:02So let's be real for a second. We all know that your website shouldn't be this static asset. It should be a dynamic part of your strategy that really drives conversions. That's marketing 101. But here's a number for you. 54 % of leaders say web updates take too long. That's over half of you listening right now. And that's where webflow comes in. Their visual first platform allows you to build, launch, and manage web experiences fast. That means you can set ambitious marketing goals and your site can rise to the challenge. Plus, webflow allows your marketing team to scale without relying on engineering, freeing your dev team to focus on more fulfilling work.

47:41Learn where teams like Dropbox, IDO, and Orange Theory trust Webflow to achieve their most ambitious goals today at webflow .com. And speaking of incredible products that allows your team to do more, we need to talk about SecureFrame. SecureFrame provides incredible levels of trust to your customers through automation. SecureFrame empowers businesses to build trust with customers by simplifying information security and compliance through AI and automation. Thousands of fast -growing businesses including NASDAQ, ANGEL list, DUDELE and CODA trust secure frame to expedite their compliance journey for global security and privacy standards such as SOC2, ISO 2701, HIPER, GDPR and more.

48:24But by top tier investors and corporations such as Google, Client and Perkins, the company is among the Forbes list of top 100 startup employers for 2023 and business this insider's list of the 34 most promising AI startups of 2023. Learn more today at secureframe .com. It really is a must. And finally, a company is nothing without its people. And that's why you need remote .com. Remote is the best choice for companies, expanding their global footprint where they don't already have legal entities. So you can effortlessly hire, manage and pay employees from around the world, all from one easy to use self -serve platform.

49:00Plus, you can streamline global employee management and cut HR costs with remote free HRIS. And hey, even if you are not looking for full time employees, remote has you covered, with contractor management, ensuring compliant contracts and on time payments for global contractors. There's a reason companies like GitLab and DoorDash trust remote to handle their employees worldwide, go to remote .com now to get started and use the promo code 20VC to get 20 % off during your first year. Remote opportunity is wherever you are. I so hope you enjoyed that show. As always, it means the world to me that you listen.

49:37You can check it out again on YouTube by searching for 20 VC and stay tuned for an incredible episode this coming Wednesday with an OG of the Venture space, the one and only Mark Sustet at Upfront Ventures.

From the publisher

Arthur Mensch is the Co-Founder and CEO of Mistral AI. Since its inception in May 2023, Mistral has raised over $520M in funding from investors like Andreeseen Horowitz, General Catalyst, Lightspeed Venture Partners, and Microsoft with a current valuation of $2 billion. Before founding Mistral, Arthur was a research scientist at DeepMind, one of the leading AI institutions in the world.

In Today’s Episode with Arthur Mensch We Discuss:

  1. From Models to Team Building: Arthur’s Greatest Lessons at DeepMind

  • What were Arthur’s biggest lessons from his time at DeepMind?

  • How did DeepMind shape how Arthur built Mistral?

  • Why does Arthur believe smaller teams are better for AI?

  • Why did Arthur decide to leave DeepMind and start Mistral?

  1. Scaling Mistral to $2 Billion Valuation Within a Year

  • What made Mistral 7B so successful? What did Arthur learn from the model release?

  • What are the biggest barriers at Mistral today?

  • How does Arthur balance the sales and research teams at Mistral?

  • What does Arthur know now that he wishes he had known when he started Mistral?

  1. How to Win in AI: Open Source, Cost, & Adoption

  • Why did Arthur open-source some models? Why did he close some?

  • How quickly will the cost of compute go down? Why does Arthur believe marginal costs will not go to zero?

  • How will open-sourcing LLMs affect the marginal cost?

  • Does Arthur think open source is ready for enterprise adoption?

  • What questions should enterprises be asking about AI adoption today?

  • What are the biggest challenges to AI adoption today?

  1. The Future of LLMs

  • What does Arthur think are the largest bottlenecks of model quality today?

  • Does Arthur think future models will be more generalized or vertical-focused?

  • What does Arthur think about the future of commoditization in models?

  • Why is Arthur optimistic about the profitability of the application layer of AI?

  • How should models differentiate themselves today?

 

More from The Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch

All 521 episodes
20VC: Mistral's Arthur Mensch: Are Foundation Models CommoditisingThe Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch · 50 min
Listen in VO