In short
Podcast Episode Notes: Databricks Founder Ion Stoica: Turning Academic Open Source into Startup Success
Episode Overview
- Podcast Title: Training Data
- Episode Title: Databricks Founder Ion Stoica: Turning Academic Open Source into Startup Success
- Episode Hosts: Stephanie Zhan and Sonya Huang, Sequoia Capital
- Guest: Ion Stoica, Professor of Computer Science at UC Berkeley, Co-founder of Databricks and Anyscale.
- Main Topics: Transformation of academic open-source projects into successful startups, Databricks' growth strategies, AI infrastructure, partnerships, and the future of AI technology.
Key Concepts & Discussions
Databricks’ Foundation
- Background: Databricks was founded on the idea of enhancing Spark, an open-source data processing engine, optimizing it for AI applications.
- Vision: The initial focus was to solve classical machine learning problems and scale them up using Spark.
Importance of Partnerships
- Strategic Alliances: Emphasis on building partnerships, notably the partnership with Microsoft, which significantly accelerated Databricks' growth.
- Competitive Edge: Stoica emphasizes the confidence in their ability to create superior products for Spark, fostering a competitive yet collaborative ecosystem.
Current AI Landscape
- Growing Complexity: The evolving AI ecosystem is increasingly complex, with numerous techniques needing to integrate effectively.
- Production vs. Demo: There is a substantial gap between impressive AI demos and deploying AI solutions effectively in production settings.
Open Source Models
- Control and Security: The decision to create and open-source their models was motivated by the needs of enterprise customers who prioritize data privacy and control.
- Enterprise Adoption: Enterprises are gravitating towards open-source solutions for enhanced control, especially in light of recent data privacy concerns.
Databricks’ Competitive Advantages
- Mosaic AI: A key component of Databricks’ strategy, providing infrastructure for training and fine-tuning models tailored to enterprise needs.
- Focus on Value: Databricks aims to help customers identify high-value use cases for AI, guiding them through implementation.
Compound AI Systems
- Definition: A compound AI system consists of multiple models and components working together, akin to programming where different functions collaborate to create a cohesive application.
- Future of AI Applications: Transitioning from human-in-the-loop systems to fully autonomous AI systems, enhancing reliability and accuracy.
Entrepreneurial Insights
- Focus on Problem-Solving: Stoica advises aspiring founders to concentrate on meaningful problems rather than succumbing to hype.
- Research and Industry Connection: Strong ties between academia and industry foster innovation; however, funding constraints can hinder academic research.
Key Takeaways
- Understand Market Needs: Identify and solve pressing problems that will grow in importance over time.
- Balance Between Control and Innovation: Open-source models provide enterprises with the autonomy they desire while balancing the need for cutting-edge technology.
- Partnerships are Crucial: Strategic partnerships can catalyze growth and expand market reach.
- Focus on Reliability in AI: As AI continues to evolve, ensuring the reliability and accuracy of models will be critical for widespread adoption.
Closing Thoughts Ion Stoica's insights blend academic rigor with practical entrepreneurship, illustrating the path from research to commercialization. His emphasis on problem-solving, strategic partnerships, and the importance of understanding enterprise needs offers valuable lessons for current and aspiring AI entrepreneurs.
---
Disclaimer: This summary is intended for informational purposes only and does not constitute investment advice.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00In general, this was our approach. We are going to be aggressive also about partnerships, even though the partners could compete and overlap. Because you have to trust yourself that at least when it comes to Spark, you can build the best products. You're saying internally, well, you know, if someone else is building a better product for Spark, then we deserve to lose, right? So that's kind of always was the confidence that we can build the best products for Spark and eventually, if Spark we are going to win.
0:51Hi everyone, welcome to Training Data. Today we're excited to welcome Jan Stoika, Professor of Computer Science at UC Berkeley. and co -founder of both Databricks and AnyScale. He is a uniquely exceptional career as a leading professor and founder of companies at truly legendary scale. Today, we dig into questions like Databricks positioning in AI, how research projects like Spark and Ray have led to the founding of Databricks and AnyScale, how he ties his research projects closely to industry from day one, new projects out of his lab like VLLM, MGMGBT, LMSIS and Vekunah and what research fields he's thinking about next.
1:31John, thank you so much for joining us today. We're really excited to have you on the pod. To kick things off, we'd love to hear a little bit about where DataBricks aspires to fit into the overall ecosystem, especially with some of the recent launches. What are you personally most excited about? First, thanks for having me here. So, I think with DataBricks, alloys we wanted to provide the platform, an intent platform, which helps our customers get most of the value out of their data. And one of the best ways to get value out of the data is using this kind of new developments in AI, including Larry Language Model and everything else.
2:22And the one thing is that to note that this was all always our vision from day one. Actually Spark was created one of the main reasons to create it, but when we build it, it was to solve, to speed up classical machine learning algorithms, right? and to scale them up. Yeah. And so in some sense, for us, it's full circle. We started with AI with machine learning. It's a classic machine learning. And right now, we are back and doing more and more AI because to take advantage and to create value out of the data. So you mentioned that you were founded around enabling classical machine learning. How do you see the current AI moment as same and how do you see it as different?
3:15And I'm curious, what are the specific things you're doing for this moment in time, like the Mosaic acquisition, things like that? Yeah, so certainly the momentum around AI, it's on a different level today. Right? You just look at the investments out there, right? I should tell the big part of the story. So I think that obviously what the way we are looking at is that taking advantage and being successful is AI is not easy. AI ecosystem is growing in complexity. It's not just a simple call with a model. Right? You have so many techniques now like rag and raft and right. to improve the accuracy of your application using AI.
4:12Obviously, everyone is excited about AI because it solves so many problems, and there are so many, you know, generally so many headlines. But still, it's not what you want. You want a product of AI. A lot of things still we are seeing today are fantastic demos. Right? And demos are inspirational. And when people see demos, whatever charge -up it is, all this kind of problems, all impacts math problems, it's very easy to think that, wow, it's doing that, it's going to do everything. So, but going from demo to production, it's like I said, it's a big step. A demo means that you need to find an instance, at least an instance, to be really impressive.
5:00In production, is there exists such instance? But when you go from the demo to the product, when you have a product, it has to work for all cases. So that's kind of the big gap. So that's why the effort is to improve the accuracy, improve reliability, of course, eliminate hallucinations as much as you can. And really also to find where it provides the best value, you know, because you know, you can apply AI to 1 ,000 use cases, but which other use cases which are going to provide you the biggest value. I think that's what we are also trying to help our customers with. So to navigate that, how to successfully apply AI to their product, to their, to improve their services, to their business.
5:49Yes. In that vein, one of the most interesting things that I thought came with the Databricks AI launch was the new Open General Purpose LLAM created by Databricks. What was the reason behind training your own models, open sourcing them, and what do you think are some of the best use cases for that model? Yeah. So I think the main, if you look at like our main market, its enterprise and enterprise customers. They do have a lot of concerns about data privacy confidentiality and obviously the kind of one control is not only that, but our disability. Right? You want to be able to audit what data you use, what the results have been used and what the decision have been made, which is what data, the decision have been made.
6:37And so that's why in general the enterprises, everything being equal, migrate towards the open source models, we can host on their machines or in their VPC and so forth. So it's one of the reasons for releasing the outbreaks is to help our customers. And then the customers say, you know, in many of these cases, they start from this model and then find you on with their own data and To optimize for their particular use case, right? So and it's again, so our enterprise customers, they want, you know, open source model, they want kind of control as much visibility as they can. They want privacy and confidentiality, especially in the recent light of the breaches, which are widely published.
7:38The other thing is that the DNA of Databricks has been open source. It's not only Spark, but Delta and MFLO and many others. The other thing I thought was fascinating about deep -brex is its excellent programming abilities. What do you think has made it such a capable code model, especially even compared to something like CodeLama 70B? Look, I think it's about, obviously it's about data and how you train it. And it's one thing, I think, you know, one of the major Databricks advantage is Mozayk, right? We have also the entire infrastructure for training. or so we have not only about the data, but the entire infrastructure of training and fine tuning.
8:27And that is making it much easier and most cost effective to optimize models for different use cases. And obviously the co -pilot programmability, it's one of the very important use cases in this enterprises because software engineers are still very expensive. Yes, that's interesting. So making the people you have, and it's not only about that, it's like hiring top software engineering is very difficult. If you are a large company, like maybe, I don't know, 4G or a thing like that. And so making those people productive, it's extremely important and critical for their business. You mentioned Mosaic as a key part of the strategy.
9:16I guess what do you think are your most important chess pieces in this kind of AI battleground? I imagine Mosaic is one of them. And our most of your enterprise customers looking to train their own models. And how does Mosaic and your other oppositions fit in with your customer's needs? Yeah. So I think there are a few enterprises who would like, so it's a kind of a pre -training and then it's fine tuning and using the model on your own hardware, on your own VPC, on your own machines, you know, you rent it, you own it. So, and you may expect there are a few which are going for training, but there are still a few who are doing going to pre -training ones that data have enough data.
10:09A lot of them, they want to do fine tuning, right? Because it's a look, if you are an enterprise, right, and a company, right? And you want to improve your business, right? What do you have? What do you have and which other don't? The, the, the, we think you have, which other don't, It's the data, the data about your business, about your users. Right? Right? So, therefore, because this is something you have and other do not, you want to take advantage of that. Right? So, how do you take advantage of that? Again, there are many ways and you try different ways to do it. Right? One is about fine tuning, you use that data and you have an open source model, you find you know that.
10:52Right? the other one to use RAC, right, and seeing like that. And so you are going to do any of those. But then you want to do, like I said, in your own BPC, to preserve security, to have, you know, 10 -4 security perimeter, perimeter, you want to do it in order to protect confidentiality of your users. And obviously right now, GDPR, California, consumer privacy act and so forth, is there also a regulation? And the number of regulation is going to increase. So the fact that you can have having an open source models, you can find you and you can use your swag in your own VPC. It's very compelling value proposition.
11:42And are you finding that most of your customers want to go that approach versus, you know, OpenAI and the very powerful close source models in the market? I think enterprises and still we are still early on. And there are many enterprises obviously using OpenAI for different use cases, different applications. and I'm sure OpenAI and Microsoft Azure will come with new products to provide better confidentiality, better security. But at the end of the day, what I was saying is that everything being equal, right? As an enterprise, I will prefer more control. Security in strategic, right? It's like it's less kind of locking or thing like that, right?
12:39So I think that's what we are saying. So if as open source models are going to catch up with the proprietary models, in particular in the use cases which matter for the enterprises. It doesn't need to be perfect for all the particular use cases. It just need to be very competitive or it matters. Then enterprises will prefer these solutions on which they have more control and more secure. And how far away do you think we are from that moment? like do you think we're there today where all else is equal or when do you think we cross that that? You tell us of the open source, that's the proprietor.
13:24Yeah, open source being on par, all else being equal for the core use cases. So you have a lot of use cases, you know, it's again is the open source plus data. And right now the application are more complex, it's not just a call to a language model. You have this kind of, you know, it's very cold, you know, you compose, you have application which build from many components and it's not called compound AI. So actually it turns out that if you can build an applications for doing you know, recommendation or single agvato or, you know, for programming for a particular copilot for a particular tool, you can actually do better than even something like OpenAI on the latest chat GPT because you have more data.
14:32And the other things about, you know, data waste, which I think they're also announcement about that is that is a kind of unity catalog and and see like that you have access and you know also about the structure of the data which helps you tremendously to improve the accuracy of your applications right so is not only the model right it's everything else you have around it And the quality of the data you feed it in. So I think that's, it's a data stupid like you'd say, right? Like at the end of the day. It sounds like control and security are two primary areas of things that enterprises really care about that you've noticed.
15:23And Databricks obviously has a tremendous advantage with having access to data as well to help these companies use models. Control and security for the customers. Exactly. What about some other factors have you noticed that they also care about is how much does cost matter, how much does diversity of other models matter. I think obviously cost is important. And it's like this, right? First of all, you know, you want to, you know, it's initially the important thing is about the value. Right? Can you provide the value? Right? It is like, so that's the first thing. And at that stage, cost is not as important.
16:03Right? And in here actually, in some of these early stages where people also try to use the most powerful model like OpenAI and so on and so on. But once you cross that and you have a use case which you conclude that is good at value to your business. Now you want to scale it up, right? And now you are talking about having more control and more security and all of these things, protect confidentiality of the data, privacy of your users, and so forth. And now basically people consider how they are going to deploy it. And this is where having multiple choices and having, like you said, more control, security, it matters.
16:49And where open source models and platforms like beta bricks are very valuable. And again, it's also all the other components we share together in the beta brick, like I mentioned, you need to catalog and every single to increase the value of your applications. Yeah, super interesting. I'd love to talk about compound AI systems. I think you guys probably coined or popularized the term. And it seems like that's a lot of what the industry is latching on to now. Maybe for audience, can you explain what is a Compound AI system and what our enterprise is thinking about when they're building these? Yeah.
17:28So a Compound AI system is basically consists of multiple components, multiple calls to language models or agents. And it's very much you can think about when you write a program, you have multiple components. You have different function procedures to do different things. And then you then put them together to create the program. The same is very similar here. You can use maybe one model to parse the data to extract the data. You can use, and then you can use, for instance, depending of the prompt, you may use a model. You know, if the prompt is about mass problems, right, it's like you can use one model.
18:19If the prompt is about, maybe about programming, you may use a different model, right? And then you may use, for instance, about formatting the result, right? That's another one. You may use that, for instance, now more and more, you are talking about agents and with agents. you have to, you know, you call external services or functions like search, or you can use a calculator and things like that. So now you have different models, you can do a better job, and there are small models actually to take the prompt and convert it to a function call. So that's kind of what it is. But the way to think about it is like the conceptually, it's like your writer program has different components and make it easier to develop and deploy and manage the same thing you want to apply to applications.
19:19So a collection of smaller components that work together where some of the parts is greater than the one monolith that you're placing. I'd love to dive a little bit into the Databricks story quickly, which I think is an incredible legendary journey over the last decade. But with a lot of nuances that I think maybe many folks don't yet understand, from today it looks like you were building the right company, the right at the right place, at the right time. But the actual nuances that Databricks was originally started originally for data scientists. It happened to cater well to machine learning workloads because of all the work that you were doing in data.
20:00And over time, you made the right strategic decisions to actually really grow with the AI market. Can you share a little bit of some of the learnings and journey that got you to where Databricks was today? Yeah, I do think that very happy to. I do think that a lot of it's also about being the right time that I place and things like that. It's being lucky, or I think that's all of this is true. And there are many things, needs to go right to be successful. Some of the things you control, some of the things you do not control. When we started, indeed, we are focusing to build a cloud product, hostage product for data scientists.
20:46Like we have this kind of notebook, book and we have provided hosties Spark and we targeted data scientists. And we targeted data scientists. It was one of the reasons we targeted data scientists because again, we all have, like I mentioned, Spark early on, also targeting machine learning applications, machine learning workloads. And at that time there are not many data scientists. It was like 2013. However, when you look around, most of the universities already have data science programs, like degrees, you're starting to offer data science degrees and you're saying, okay, we are, it seems a good, you know, pass forward.
21:37A good market. And I remember you're looking at LinkedIn to see how many data centers there because it's not our users. Not initially, they're not as many. Thousands. Especially when you convert the data with database analysis, analysis, and so forth. Who started the build? And I think it was a reasonable good product. And then we started to grow quickly in terms of the number of customers. Initially we had small customers.
22:14But then, and still, the interactive analysis, the data science has been for a long time. It's actually one of the biggest workloads, in particular in terms of revenue. because the interactive workloads are price higher than batch workloads. But I remember that I sold to these companies to use data bricks and they are aspirationally buying it to the data science AI. And after a few months we look up to them because we didn't go much earlier, because we saw that they were doing well, they're usually just growing and everything seems good. So no reason to worry. So we went back to them and said, you know, to see what they were doing, maybe we can write a blog post or whatever they were doing, right, it's like marketing and everything.
23:15The surprises are very few are doing actually machine learning at that time. And we asked them what happens, right, or still. Well, it turns out that obviously, that in order to the machine learning, you need data, like we discussed early on. And they realize that for the particular applications they want is they don't have the data they need. So they need to now put maybe to add new logs, you know, to call like new logs in their products and things like that. And also they need to clean up the data, you know, curate the data and so forth. So now we are lucky because Spark was very good also for data engineering, right?
23:57For data processing, right? It's a data processing tool at the end of the day. And so they are using Spark for data engineering. And then obviously we start focusing also on serving data engineers much more than before. and just how we started. And then obviously when later now, you know, God's, you know, data engineering, right? And still, you know, have all these data scientists exploring and starting to build models. And then it is a very natural extension to start to add more products for our user, for our customers, like I mentioned, to get more value out of their data and that meant building machine learning models and using the open source models.
24:57Very interesting. I want to talk about the Databricks Microsoft partnership. I think that was the stuff of Legends. I think still probably one of the only case studies of a successful, like truly transformative partnership. Maybe walk us through What that partnership was? Was it a bet the company moment at the time? Do you think Databricks would have become what it became? If you hadn't struck that partnership, maybe just talk about that? Look, obviously, the partnership with Microsoft was a great partnership for us. The one thing I want to say is that, and this is most visible, we are very, from the day one, we are actually very focused on the partnerships.
25:42Our idea was always, you know, like we now we make Spark successful, you know, hopefully the fact of standard for data processing and then we make Databricks the best place to run Spark. So early on actually, you know, we started, you know, a few months in the life of the company we had this kind of partnership with Cloud Data. And then we had these Hortonworks. And these partnerships was mainly because you are, you know, to advance Spark, right? Because Spark, it was created in Hadoop ecosystem, right? And these are the Hadoop companies, right? And so we had this partnership, or data stacks and so forth.
26:31So this was like in the first year or two. So we had all this partnership. Despite the fact that some of these companies, we knew that they could become our competitors, right? Because, you know, just making them, is making and helping them to deploy, to manage and to sell Spark based services, right? It's like, so in some sense, you know, the Microsoft is like, it's, it's, it fits our, kind of, approach of trying to, you know, strike as much, you know, the partnerships with other, being aggressive, doing partnership with other, you know, organizations in the ecosystem. Even though, again, in some cases, it was not a clear cut whether we are, you know, they are going to compete with us or not.
27:32Right, so you are very just growing the ecosystem and growing spark that was the priority. We even had a point in the partnership with Snowflake. So it requires a lot of heavy lifting. So we always look for partnership which are meaningful. And I think great negotiation, Ali and so forth. Did a fantastic job there. But at the end of the day, we also need to commit to it. It's like it took to build your data bricks. So we are an IWS before that. It took tens of engineers one year, and you have a small company at that point. So it was a huge commitment and a huge bet also from our perspective to do that.
28:28And yeah, and we, you know, I think engineering and everyone executed very well and it was a successful product. Right. And, you know, Microsoft was great partners, you know. And yeah, this is what happens. We are obviously a bit lucky, but in general, this was our approach. we are going to be aggressive also about partnerships even though the partners could compete and overlap because you have to trust yourself that at least when it comes to Spark you can build the best products. You're kind of saying internally, well you know if someone else is building a better product for Spark then we deserve to lose right.
29:18So that's kind of always was the confidence that we can build the best products for Spark, and eventually if Spark wins, we are going to win. Do you think that Databricks will have become the company that is today without that Microsoft partnership? I think so. It maybe could have taken a little bit longer. Yeah. But yeah, I think so. Right, we would have still have a good, very good offering on Azure. Like we have with GCP, you'd have taken a bit longer. But I don't see any fundamental change in dynamics, right? Because one of the advantages of Databricks of course is one spark won, and we could provide the best product for spark.
30:04We are in a very strong position. And compare with other clouds, remember that one of our advantage, like every one else's advantage, like confluent and so forth, versus that, you can provide a service on multiple clouds, right? And multiple clouds is, you know, multi -cloud has been more and more very strategic for, especially for large enterprises. We do not want necessary to be locked in or to one, they want a choice. So yeah. I love the confidence and conviction that you have in Spark and your own execution ability is internally, But that married with the practicality and aggressiveness of winning as a business and pursuing the right partnerships and doing whatever it takes to win.
30:53Yeah, I mean, you try to simplify things. That's what I was saying. You know, you like, initially when you said, like, look, you know, like, we have to make Spark Queen. You know, there are many, I remember we look at all these combinations. It's like spark queens, the product fails, right? Or spark loses, but we have a product which is successful or boss fails. No, that's not very interesting or boss are successful, right? And we convince ourselves that, you know, we need to bet on spark to win because that's the most likely way also for the product. Again, for better or worse, right? But sometimes, there could be many ways to success, right?
31:39And in retrospect, you cannot go back and try other alternatives. Maybe there are better alternatives at that point. But it's important to commit to un -sync, which hopefully it's a reasonably good solution, a good password, right? Okay, and there are many passes to the peak, right? The most important is to go meet one which list is a peak, right? It's like not the shortest one or the easiest one may not be, but it has to be one together. And to the highest point. And that's how you said, okay, Spark has to, has to end, right? And we have to build, you know, the best, need to be the best place for Spark.
32:15And then we are saying, you know, to be the best place for data and AI. We need to eventually we knew and we assume that if you are to be hugely successful, you are going to go beyond Spark. Right? It's like that's why the name of the company, the Databricks is not Spark Labs or something like that. Yes. So that's kind of you try to simplify. So you do that and you start to execute. OK, so you want to make very successful as an open source. So you want everyone to use it. So that's how you do Cloud era and try Horton also and so forth to do this Spark. Because there's a time there are other solutions.
Read the full transcript
32:54You know, people knew that Hadoop MapRedews, you know, these times have passed, so to speak. So they were talking about new systems. They were all like, tells. It's actually Hortonworks has this project and so forth. So that was very important. And then it was also for data science. It was kind of when we bet it's like it's an issue. and we can, we saw that we can build a best product for it. So, you know, so that's kind of, you need, and you need to have ultimately confidence in what you bet on, right? You have to bet, right? Right? Because you are a small company, right? If you don't bet how you are going to win, right?
33:41Everyone, right? And then you need to have some level of conviction, right? to do it. And yeah. I'd love to kind of pull in that thread and switch gears a little bit into tying a lot of the entrepreneurial path that you have with a lot of the academic and research background that were the roots for the beginnings of these companies. You have a very unique career as both a leading professor and a founder of multiple unicorn and deca corn companies. I don't think there's anyone who comes close to pursuing both disciplines at the scale of success that you have. Maybe specifically, you know, Ray with any scale, Ray to any scale, Spark to Databricks, take us into your head.
34:26What is the process in which these research fields start to ruminate in your mind? When do you kind of continue to give them resources to develop and then when do you know that it's the time to then start a company to pursue that in a better, more open and fast way. That's a good, that hard question. So I think that, and by the way, I like to preface, to preface, to say that, to start with saying that, it's obviously also a lot of like in the middle of the world. And being in a place at Berkland having fantastic students and colleagues around you, you couldn't do without that. It's like, is there married probably more than mine?
35:19But one thing I think is that I've been always trying to focus on the problem. And actually, when, even to my students, I'm telling them, one of the most important things you need to do is to figure out what problems you are going to work on. Because every one of us comes to Berklar, stop schools. One thing they have in common, they are good problem solvers. right? They have good grades, you know, good scores, you know, bright papers, papers about solving a problem. So therefore, if all of them are good problem, problem solvers, the differentiator is a problem you are working on. Yes. So you start with that.
36:08And also, I think it's like, especially at Berkley, you get exposed of kind of not only new ideas, but willingness to take risk and get into new areas. In some sense, and this is what I like about Berkeley, if you look traditionally, Berkeley, among top schools, the first to open new areas is like, of course, a risk processor. It was a stand for this well. but databases, networking, a sense of networks, even in open source, so is the UnixBSD, TCPIP, right, part of the SV. So there are always kind of more, you know, a little bit trying to experiment. So I think that's kind of, you know, that's kind of culture.
37:09I really resonate with it. And then the other thing that happened is very clearly we have these slabs, which are like five years slabs, which basically each lab has kind of a vision. And you know, it's a group of faculty coming together, which believe in that vision and with their students and try to make it happen over, you know, five years. And this, you know, has a lot of great, you know, impact, you know, it's like all the way. It started, this tradition started 40, 50 years ago, which they passed on the land, the cuts and others. And they built, you know, risk rate, redundant array of inexpensive disk now, network of first station, as commodity, you know, this is everyone is building now, this huge cluster of commodity machines, servers.
37:59And again, and many more. So there are these kind of elements, and these labs are a very strong relation with industry. Connections are funded. Since I came, when I came actually, this was one change happened. Before these labs are also supported by government, in particular DARPA. But it was a point in which that kind of, at least that particular DARPA funding dried. So, when I came to Berklaizoro, this kind of now you need to go and remember getting more money from industry. Remember from Google first time we got, and it was unheard by then because you were asking for 500 ,000 per year for four years.
38:49But you know we got, and this is we got. So now you have also this kind of very tight connection with the industry. It's a very good environment to see the problems, right, to understand the problems. And then you can see that. And you try to think about also about, obviously, about trends, because trends are important. You have to be aligned with these secular trends. You need to bet on the right trends, because these are things you cannot change. Or it's very hard to change. So you are not aligned. It's not good. And these trends actually, there are multiple trends. And the multiple trends kind of open gaps between them.
39:36And these are kind of opportunities for problems. Like for instance, for big data. It was clear you have more and more data. And the amount of data people collecting were just growing. It was pretty clear. It's like Google has seen that year before. And they built all these systems. But now everyone wants to emulate that. That's why Hadoop was created. And then you start to see again, you look at that and you are working in that area. And we have all these Hadoop people coming to our retreats on this kind of labs and we are friends with them. And we start to see problems. and then you try to use them.
40:22And then you just like for instance, there are two things which happen. Isadop, one thing it's about a group in our lab. Oh, by the way, and the other thing what happens in these labs, they are interdisciplinary. Right? Because people, you know, we are in this kind of lab, was that it was people from machine learning, systems, databases, networking. Right? So there are these groups of Michael Jordan students, which wanted to compete to this Netflix challenge. The Netflix released some data basically and asked for people to provide recommendations, come up with recommendations, to build recommendations, algorithm systems.
41:05So that to beat their own recommendation. So they come to us and, okay, it's a lot of data, what can we do about it? Right, you want to, and you know, tells them to do yes, hard up. And but hard up was very slow, right? Because, you know, and then, you know, but they put together something quickly for solving this problem in which the data was kept in memory. The other thing I've seen is about, like, I have a previous company, Conviva, and it's about its own analytics company. And kind of it was very slow and we try to do it, you know, it's like, you, for our queries, and there is no way to do it.
41:47And again, keeping the data in memory was a solution. It's one solution. That's kind of how we start it. That's one thing. And you look at the trends, yeah, it's obvious, right? It's like on one hand, you have more and more data, growing faster than the more slow. So you need to have, it's not going to fit on on machine. And therefore, you need to do multiple machines. And then the only other question is that, are you going to have data sets that are going to fit, important data sets are going to fit in memory, right? It's like that's kind of the first question. And they are because people, even for, when they were doing, for instance, are, who query you notice that?
42:25And they look at the data from different cluster, how to cluster from Yahoo and Microsoft and others. And they notice that in a lot of cases, actually, when you look queries and you're doing analytics, they are doing not only very rare on all the data, You do say for the most recent data. You want to see what happens yesterday, what happens last week, something like that. So once you get that and you have a lot of cases, you know, the data is fitting in memory and memory is still growing quite quickly at that time. You know, that's kind of your pretty, you know, you connect the dots, right? And then it's about solving the problem.
43:06And the other thing is happens and why they are related. I'm talking about academia and industry because I'm telling people, some people, you know, there are people in academia push back. You say, you know, it's a lot of engineering here. This is not you should do in academia. But one thing is what I've always found it is very satisfying. If you build a system, right, in a new area, and that system, it's used by other people. then you are in the best position to understand the new problems in that area. Because people are going to use your system in different ways, right? And then you understand that, right?
43:52And going back, if you know the problem, you are also in a good position to solve it, right? So actually, it directly helps you with your research, right, to be ahead. Because otherwise, what is a choice? Well, you are finding the problems, right? Of course, there are problems, very good problems in theory, which are not solved by decades and so forth. But other things, what people do, they go to Google and Microsoft and so forth, spend time and to understand what the problems they have, right? Because to solve. But that's kind of a little bit unsatisfying, right? Because you go to someone to learn about their problems.
44:29But the question is why don't people solve those problems? Maybe they don't solve the problems, because maybe they're not as important as they're given time and maybe for the right reasons, they are too much in the future to that. But this is a thing, right? It's you have to focus about the problem, on the problem, and you have to focus on the trends. And the way they're connected, you want to solve a problem, ideally, which is going to be more important to the more of them today. What are the problems that you're most excited about right now? Like Spark, Ray, what's gonna be the next data breaks or any scale?
45:08So a few things. So I still think that it will going to be a lot of work. It's right now what happens if that need to rethink most of the software stack. Why? Because it's again, if going back to the trends is that the demands of this application, in particular, air application and so forth, growing much quicker than capabilities of a single processor, single node, even if you consider accelerators. Right? So one on hand, this happens. On the other hand, the infrastructure becomes much more complex. You need to run that application not only on one node, but on many nodes. It distributes it. But it's not only that, it's becoming very heterogeneous, right?
45:57Because in order to breeze a gap between the demand and capabilities of hardware, people build accelerators. Like, that's why NVIDIA is a trillion dollar company, right? And but now the infrastructure becomes even more complex, right? It's heterogeneous. is not only distributed heterogeneous. When we start this spark, it was homogeneous. All the nodes are the same, some storage, some CPUs, that's all. But right now, look at the heterogeneity. You have Nvidia, and you have many others. You have TPS from Google, every cloud. It's having their own chip. Now you have MD and Intel. say that it's everything, it's about, yeah, yeah, right.
46:45So that's kind of what happens. So now you have a huge gap between this application and it's very complex infrastructure, just growing in complexity. And then it's not only about the CPU, it's a compute. It's about networking, you have infinite band. And you have all of this, you know, RDMA and so forth, right? So huge heterogeneity. And the software stack has to abstract the way that complexity for the developers. There is no way around. You want a single machine or you have this operating system to have structures of complexity, right? That's why that makes it easy to develop this application.
47:21Now it's extremely hard. So something is going to happen there. I think the other one is about this building application. You're talking about compound AI, you're talking about compound AI and things like that. Right now, this application, application in particular, large language models, every one is talking about, large language models. You know, the application are like assistance for humans. So, right, humans are in the loop. Right, if you are thinking about customer support, if you are talking about copilot, if you are thinking about Q &A question and answering, even summarization, you have a human helping a human to be much more productive, which is a fantastic application.
48:04But they cannot be autonomous. They are not yet autonomous. And to go from having the human is a loop to being autonomous, it's a huge gap, because it's autonomy. To have something autonomous, you need to have someone which is running more deterministic, it's more reliable. It's like far more accurate. Right? And you need to get there because if you don't get there, you are still limited to having the system or the human, it seems to look. And the human is a bottleneck, it becomes a bottleneck, it's just certain number of people on the planet. And I think that there is a lot of work to be about how to make, I would say at least some aspects of this lag language, building the lag language model application, more like an engineering discipline, right?
49:05Where you can build much easier systems for smaller components. Okay, so the two next data bricks will be distributed computer across, whether a genius hardware and autonomous compound AI systems. No, they... I'd love to tell them another thread that you mentioned just now of kind of how funding constraints drive what you're working on. I'm curious what you think right now of there's been a lot spoken about kind of the almost the brain drain in AI right now because the universities don't just don't have the funding that you could get if you went to work at one of the big research labs. How do you think about that?
49:44What do you think is ideal? Does it does operating under constraints force creativity for you? Like what do you make of all that? Yeah, that's a great question. And it's true. I mean, it's challenging. It's very challenging. And when I came to United States, I came to do my PhD, and I graduated from Kennedy, even on university. I'm originally from Romania. So one thing was admired, people admired about United States, It's like, how well this kind of, you know, the collaboration and the partnership, the three -way partnership between academia and government and industry are working. And prime examples at that time was obviously the Internet, right, which was DARPA project and of course academia had a huge impact and also industry was, you know, it was whatever, third industrial revolution people were saying.
50:55And that when it comes to AI today is that kind of partnership is broken. It's like industries, every company is doing this, it's a rich in silos. right even you know they don't talk each other as much. Academia like you said doesn't have resources and the government doesn't invest as much right so I think that's that's something to be very concern about right and that's why I am also a big proponent of open source models and And US and California kind of lose in long term, if this doesn't, is not fixed. So what happens now, in unfortunate academy, what I think it happens is that, of course, there are some bigger universities and labs which can still afford to maybe spend one to millions, maybe to train some models.
51:54It still there, not in the same league, like open AI, it's like open AI and Microsoft, of the world, they are talking about building the data centers about $100 billion. And I think the danger is that you are going to, some of the academics are going to give up and try to innovate around the ages. Now you can still innovate in this application and sing like that. and I think there's a lot of innovations there. But clearly, it will be harder to come up with new model architecture and wait.
52:41It will be harder. It's not impossible. Nothing is impossible. So yes, it's a challenge right now. I mean, there are people like, you know, and which are a better, more fortunate position, maybe I'm one of them, in which have access to resources outside the academy, right? But you know, it's, you do want to level the playing field, you know, there's to maximize the innovations, the innovation comes from everywhere. Now, this being said is through that scarcity and always in the past, scarcity, spurs innovation right as well. But the concerns about not having access to resources, to play the same game like in industry, it's a concern.
53:44I'd like to switch gears into some rapid fire questions. If you're ready for it. Sure. Well, anyone take meaningful market share from Nvidia over the next five years? I think they it will be. They will be at the minimum because they are going to, it's probably that it's because Nvidia will not like to be accused about monopolytic, monopolytic behavior. So, their market share has to decrease under some percentage, whatever, 70 -80%. So, that one will be one of the reasons. But I think that if I have to name one company to... Of course, there are clouds and for strategic reasons, they are going to push their agenda to build their chips, like still probably the biggest competitor right now in terms of the market share, it's Google, it's TPs.
54:41Yeah, yeah. Probably that will continue for a while. What's one project or students in your lab right now that you'd want to highlight? I'm going to check here because I think that both VLLM and Charbot Arena has been tremendous. It's like, I'm not talking about Skype pilot because Skype pilot is like, they started a company. So I'm not talking about it. But I think Vialalam has been amazing. It's like it's one year old project. And I haven't ever seen such a rapid growth. And of course, it's also part of the AI. It's kind of compresses the time. So it's as there is something about that. And I think the other one is Charbot Harina because he just is fascinating to see the development in the space and to see how these different models was they are strong or they are weaker.
55:36And I think that having a front seat to see that kind of development in the ecosystem and in the space, it's fascinating. Do you think the foundation models will commoditize? The foundation models are like a GPT -4 or a cloud. Do you think there's a market to be made and providing these models over time or do you think it commoditizes? I think people will continue to build larger and larger models. I think the words of when it comes to serving is it looks like model distillations works quite well. By the way, the model is where you're trained as smaller models on the outputs from the bigger model.
56:20It has been a lot of success as that. By the way, this in some way it says that it shows how important is the data, right? for training a model, right, going back early on in our conversation, right, because you have higher quality data from the big model and use that to train the small model and it's working very well. So I think using multiple distillation models to reduce the cost of inference is going to be a way forward. but yes, I think for advancing and pushing the frontiers, so to speak, right? Pan intended, you are still going to see a lot of effort on bigger models. What are you most excited to see in the world of AI in one, five, and ten years?
57:06What I'm actually talking about AI. Look, there is no question is transformational, right? I think that it will change a lot of things, everything maybe. I think the most excited I am about it, it's about how do you make this AI, more C systems, more predictable, accurate, very viable, how you can debug the systems. All of this is kind of in the realm of software engineering like techniques. This is what I think it's exciting. What advice do you have for founders building an AI? It's the same thing, focus on the problem. Don't focus on the hype. The hype is emotional. It's not reliable. Just look at the facts.
58:04It's like, and look at the problem. Try to understand the problem. and try to be truthful to yourself. It's about, if you build an application, it's about production. It's not about the demo. Right? It's beyond the demo. Of course, the demo are important, don't get me wrong, are very important. But the production is the way the mindset you have to have. And, yeah, and it's a dangerous thing, there's so much hype. And you can solve everything, and you can do everything, probably in certain number of years. But now just focus on exactly what problems you are going to solve. Convease yourself, it's a good problem.
58:45Convease yourself that you can solve it. At least you have, you can have an MVP right. You can solve a smaller, a smaller version of that problem, which is still very valuable for your customers. That's what I would say. Nothing, no silver bullet. Amazing. Thank you so much, Yon, for joining us today. I've loved hearing a lot about your own thinking and reasoning behind your own journey. A lot of the thought process behind finding the right problem to solve, building the right systems to actually be in a position to understand the best problems, and then applying that even to many of the bold decisions that you've had to make in founding multiple companies from research into commercialization and the incredible success of Databricks today.
59:31Thank you. Thank you for having me.
From the publisher
Berkeley professor Ion Stoica, co-founder of Databricks and Anyscale, transformed the open source projects Spark and Ray into successful AI infrastructure companies. He talks about what mattered most for Databricks' success -- the focus on making Spark win and making Databricks the best place to run Spark. He highlights the importance of striking key partnerships -- the Microsoft partnership in particular that accelerated Databricks' growth and contributed to Spark's dominance among data scientists and AI engineers. He also shares his perspective on finding new problems to work on, which holds lessons for aspiring founders and builders: 1) building systems in new areas that, if widely adopted, put you in the best position to understand the new problem space, and 2) focusing on a problem that is more important tomorrow than today.
Hosted by: Stephanie Zhan and Sonya Huang, Sequoia Capital
Mentioned in this episode:
Spark: The open source platform for data engineering that Databricks was originally based on.
Ray: Open source framework to manage, executes and optimizes compute needs across AI workloads, now productized through Anyscale
MosaicML: Generative AI startups founded by Naveen Rao that Databricks acquired in 2023.
Unity Catalog: Data and AI governance solution from Databricks.
CIB Berkeley: Multi-strategy hedge fund at UC Berkeley that commercializes research in the UC system.
Hadoop: A long-time leading platform for large scale distributed computing.
VLLM and Chatbot Arena: Two of Ion’s students’ projects that he wanted to highlight.




