E53: Tomasz Tunguz on the Evolution of the Data Ecosystem

27 Aug 2024 · 43 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Turpentine VC - Episode 53: Tomasz Tunguz on the Evolution of the Data Ecosystem

Episode Overview In this episode, Erik Torenberg interacts with Tomasz Tunguz, General Partner of Theory Ventures, to analyze the current investment landscape in data businesses, with a particular focus on Databricks, its growth trajectory, competition with Snowflake, and the implications of AI on the data ecosystem.

Key Themes and Discussions

Introduction

  • Welcome back to the podcast; the discussion centers on Databricks and the data ecosystem.
  • Tomasz shares his initial thoughts on Databricks and its evolution.

Databricks

Company Insights

  • History & Growth: Databricks was founded to commercialize Apache Spark, initially challenging traditional data processing models.
  • Market Positioning: Tomasz discusses how Databricks is positioned against its main competitor, Snowflake, emphasizing their strategic differences.

Understanding Data Ecosystem

  • Unstructured vs. Structured Data:
  • Traditional data infrastructure focused on structured data, primarily for business intelligence.
  • Emergence of unstructured data (text, video, images) necessitates more robust systems for processing large datasets.
  • Modern Data Stack: Databricks facilitates the handling of unstructured data, allowing for more sophisticated data pipelines.
  • Comparison with Historical Models: Transition from MapReduce and Hadoop to modern data solutions is highlighted.

Investment Perspectives

  • Bull & Bear Cases for Databricks:
  • Bull Case: Databricks gains dominance in AI infrastructure, increasing its market share from Snowflake.
  • Bear Case: Competitors like Snowflake effectively retain their customer base and advance their own AI capabilities.
  • Historical Skepticism: Initial doubts among investors regarding the value of unstructured data have shifted, particularly with the rise of AI.

Competitive Landscape

  • The rivalry between Databricks and Snowflake is intensifying, with both companies aiming to capture larger portions of the market.
  • Challenges of transitioning from Snowflake to Databricks are discussed, emphasizing customer inertia.

Future Directions

  • AI's Impact on Investment Opportunities: The growing need for AI capabilities is reshaping the data landscape.
  • Emerging Trends:
  • The growing importance of data catalogs as repositories of company data and lineage.
  • The unbundling of databases and the rise of multiple computing engines.

Open vs. Closed Source Dynamics

  • Discussion on how Databricks (open source) and Snowflake (originally closed source) are positioning themselves strategically in the marketplace.
  • The importance of managing open-source communities as companies grow and evolve.

Looking Ahead

  • The potential for further consolidation in the data ecosystem is anticipated, with increasing competition expected among data processing tools.
  • Tomasz shares insights into investor sentiments and the evolving landscape of data business strategies.

Key Takeaways

  • The investment landscape for data companies is rapidly changing, influenced heavily by advancements in AI.
  • Databricks has become a significant player, challenging traditional models and focusing on the unstructured data market.
  • The competitive dynamics between Databricks and Snowflake will shape the future of data infrastructure.
  • Understanding the significance of data catalogs will be crucial for companies aiming to optimize their data strategies.

Additional Resources

  • Tomasz Tunguz's Blog: [TomTunguz.com](https://tomtunguz.com/)
  • Theory Ventures: [Theory.ventures](https://theory.ventures/)
  • Databricks: [Databricks.com](https://www.databricks.com/)
  • Turpentine Network: Join over 400 founders and executives in the Turpentine Network [here](https://hmplogxqz0y.typeform.com/to/JCkphVqj).

Conclusion The episode emphasizes the significance of Databricks in the evolving data ecosystem, underlining the importance of understanding both the technical and market dynamics as companies navigate this rapidly changing landscape.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Welcome back to Turpentine VC, a podcast where we discuss the art and science of building successful venture firms, VC to VC. In the episode ahead, Tom Tenguz, GP of Theory Ventures, returned to the podcast to talk about the past, present, and future of data businesses, specifically through the lens of Databricks. Tom offers a bull and bear case for Databricks, an analysis of the market and main competitors, and what's next for the data landscape in the AI era. Let's dive in. Tom, welcome back to the podcast. Thanks for joining. Pleasure to be here. Thanks for having me on, Eric. so we're going to cover databricks today maybe let's start with your journey with databricks and when did you think hey this is going to be a you know pretty spectacular company i remember meeting them at the a and the b i mean meeting ali and um uh i mean he's a force of nature uh and clearly like i mean he was coming out of the the berkeley labs and they were commercializing spark spark was a very difficult technology to run by herself and databricks was building software to manage it.

1:07Maybe it was the B and the C, not the A and the B. And, you know, many of the leading data scientists were using this technology as a way of manipulating enormous data sets to get them to, from whatever unstructured raw place they were into whatever place and packaged output they needed to be. And I think he's just done a phenomenal job building a really aggressive company. I think the story, one of our founders is building within the data ecosystem and he went to an event with Ali. And I think the story really sort of encapsulates his business aggression, which in my view is an extremely good thing.

1:46He said, he got on stage talking to all these different startups and he said, we would love to partner with you and we will be competing with you. So if you win, great, we'll work with you. And if we win, well, we win. I don't know if it's exactly what he said, but that was his sentiment, which is, this is an open, competitive playing field. We are playing to win. Let's be aggressive. Yeah. And talk about the company itself. What do you think was sort of the insight or what do you think about like what makes Databricks so special or what makes it work and situate it within a larger kind of context of the market?

2:23Yeah, great question. Okay. So let's say you're building within the data ecosystem. Then the classic data infrastructure has been, the output has been BI, right? Business intelligence. The CEO wants a dashboard of how many widgets am I selling per region per month? And the way that that's built is there are transactional systems. So let's say you have like a point of sale system that measures how many widgets you sell. That data is moved from the point of sale database, which is called the transactional database using ETL. It's basically just extraction, transformation and loading pipes like 5Tran into a cloud.

2:53data warehouse like a snowflake. It's then a snowflake gets manipulated, reconfigured, and then goes into a dashboard. That's what we would call sort of the modern data stack. And that works really well for relatively small data volumes. And it works really well for structured data numbers, things that you would find in an Excel spreadsheet. There's this entire other world of what's called unstructured data, which are things like raw text or video or images, and even really, really large files where that doesn't work because it's really expensive. And because they're so large, the latency doesn't matter so much.

3:27You may not need down to the second latency. And so what you want are bigger systems that are burlier, that can handle much larger data volumes and can parallelize them and manage them over large numbers of servers. And before the world of Databricks, there was a technology that was called MapReduce that Google had built that did this and which became commercialized in Hadoop in the form of Cloudera and Hortonworks and MapR. Those were the three companies that were sort of building this technology. And the way that it's been described to me as sort of the most tangible difference between these two technologies is Snowflake is really to deal with like data frames.

4:05And then what the Hadoop vendors, those were built by people who were file systems people. In other words, like how do I manage logs files or video files? and even when i was at google i was trying to play around with map reduce and i had candidly i had a really hard time it was very difficult for me to understand how do i set up a map and the reason it's called map reduces like a map this is how all these fields map together and then here all the here how all the files kind of compute and go out and get you to a reduction which is the that particular nomenclature for an output and so what spark said is okay file system people built this really sophisticated thing that can manage unstructured data and process it.

4:47But a lot of data people are having a hard time. So why don't we build this really sophisticated pipelining infrastructure at scale in a way that data people can understand? And so that's a very simple sort of reductionist view of why Databricks has had so much success is because data people now can take advantage of absolutely massive computing resources to build these data pipelines. So what are some of the applications of a big unstructured data pipeline. Let's say you want to train a machine learning algorithm to identify license plates on, I don't know, at the scale of New York City Police Department.

5:21Well, you use Spark. You'll use Databricks as core open source technology, and you'll piece together a constellation of different open source projects that take those video files, put them into a folder, effectively process them with code, and then spit out an output. And a lot of the times that output is just either refined files, but now you're seeing Databricks actually starting to compete with Snowflake. You can look at the data warehousing revenue has 4xed in the last year. It's gone from about 100 million to 400 million in a single year. And two years ago, we would ask enterprise buyers, are you using Databricks for BI?

5:57And the answer broadly was no. It was very difficult to find anybody. And now we're hearing about them more and more. So these two worlds of the structured data and unstructured data are starting to converge. And that's why there's increasing competition between Databricks and Snowflake. Why didn't the solution exist before Databricks? Like what took so long? Was it sort of this had to be the right time for the company or was it just too hard to do? Or talk about the time, the why now? Well, most companies were not dealing in huge volumes of data, right? Google was the first one. And, you know, the apocryphal story is that Google, when they started, they were going to use very fancy Oracle databases, but because they were crawling the web, it was just too expensive.

6:34So they needed to build a new technology to build MapReduce. And that's what started. But Google, there were only a handful of companies that were really operating at web scale. Yahoo also was operating at web scale. Yahoo created Hadoop. And so Hadoop and MapReduce kind of came to be. Then over time, everybody, once we saw the business model of Google, everybody in the world said, okay, now this is the era of data. We need a lot of data in order to compete. And so data volume started to increase and the market size, the addressable market of people needing massive data infrastructure solutions started to grow.

7:07And as that started to grow, the way to capture the market was to improve the user experience so that the data people could actually use these technologies in a much more sophisticated way than they could with the MapReduce Sadoop world, which was targeted to basically a different user. Did Databricks have, what were the major competitors along the way, or is it really just such a project that they didn't have any there it's hard to point to a competitor um it really is i i don't know i think at this point you have snowflakes uh snow park which spark spark is databricks's technology snowflake came along a little bit later and created a technology called snow park very deliberately which is um snowflakes equivalent there's been some adoption but it hasn't sort of derailed the locomotive that is Databricks, that's probably the most viable competitor of any sort of scale.

8:05Yeah. And was Databricks obviously investable? At what point did it become obvious that it was going to be a pretty big company? At the beginning, I would say the venture ecosystem as a whole believed more in structured data than they believed in unstructured data. And that's because, remember, we were coming off this wave of Cloudera and Hortonworks and MapR. And there had been massive commoditization in that market. When Cloudera started, list price for a node, in other words, running Cloudera on a single server was about$4 ,000. And then at the end of like five or six years, it was about$1 ,000 a node.

8:39And there weren't that many players, but they're all trying to win relative share. And so they just kept cutting price and price and price. So there was a fair amount of skepticism, I think, in the venture community overall that in this part of the ecosystem, somebody could charge a premium. There was also, I think, a fair amount of skepticism that was a lot of value within that unstructured data. So you think about, okay, what is unstructured data? Well, maybe it's logs, right? The value per gigabyte of a log relative to the value of a gigabyte of a point of sale system, massively different. Log is not that valuable.

9:09Very few people actually look at it. The value of one frame in a video is not that valuable. And so you're basically pressing these giant lemons, not getting any juice in unstructured data, or at least that was the perception. And whereas the converse was true within structured data. And now with this AI wave, the reality is the transformer model, well, that's totally enabled by extremely large compute and extremely large data sets. So you could argue Spark and that technology is, and the related technologies are basically a necessary or requisite to getting to where we have been in AI. AI is just massive tailwinds for Databricks, basically.

9:48Like as the market has matured or developed, it's only just better for Databricks. Yeah, that's right. I mean, I think the reason that there's so much interest in the company in the late stage markets and the reason why I think when they go public, you'll see a significant amount of demand is it's thought of as the AI company, the AI infrastructure company, not in the way that OpenAI is thought of as the model provider. But if this is like the data movement company, if you're moving data, if you're manipulating data, if you want to set up pipelines to train machine learning systems or modify them, you will be using Databricks predominantly.

10:20uh and so that's why there's there's just so much demand for it it's interesting so let's go to your your post a bit so you were starting to explain basically how databricks and snowflakes are intersecting but let's compare and contrast the businesses a a bit more well so i was at both the databricks summit and the snowflake summit of databricks event which was two two two and three weeks ago respect well snowflake summit first and the databricks summit san francisco and you can get a sense of the difference in the buyer universe snowflake summit probably director level and above a lot of the topics covered are about transformation of a company from wherever they were to the future of data data bricks is attended by engineers a completely different prototype completely different phenotype of individual coming to those two events and five years ago they didn't really overlap why was that well we talked about the modern data stack and how most of data movement, particularly for structured data, is the creation of reports.

11:18Well, it turns out that there are three categories there. There's business intelligence, Omni, and Looker. And then there's data exploration where I'd put Hex. And then the last is machine learning systems. Historically, machine learning systems have all been post hoc offline. In other words, CEO asks, can you cluster the customers? Who are the customers most likely to churn? What is the revenue prediction for next month? Those are not part of a product that is delivered to customers. that's something that's presented to a ceo as a report and then when we had the launch of the transformer everything changed and now all everybody wants ai and in the old world the people building the reports for the ceo were not part of the engineering organization you had a vp of eng and then you had a vp of data and those are two parallel structures where they might use the engineering teams were producing some of the data and the data teams were consuming it but That was it.

12:05That was the end of the data pipeline or the data supply chain. And now what's happening is the engineers are saying, actually, you know that data that we're producing around points, whatever, people buying widgets? We want you to send it to the data team. We want you to make a machine learning model out of it or an AI system. And then we want to put it into the path of production. We want to make a recommendation system or a chat bot for customer support. And so instead of the data dead ending within the data team, it's now basically a circle. The two teams are working together and the two teams are converging.

12:37And so the Snowflake people are now starting to talk to the engineers. And the engineers are saying, well, we need these production grade systems. And so these two teams are coming together. And as a result, Databricks and Snowflake users are starting to see each other more and more and more. The big question strategically is if you're a large company, do you standardize on one? Do you standardize on Databricks? Do you standardize on Snowflake? Do you have some teams running on Databricks, some teams running on Snowflake? And the rationalization of that infrastructure, I think will be sort of the defining, the battle of the giants here over the next five years.

13:10Hey, we'll continue our interview in a moment after a word from our sponsors. How deep do you go to seek out an answer to a question? Maybe you've spent hours clicking the source links on an obscure Wikipedia page, or maybe you're even the type of person who checked out the entire shelf on the topic at your library. If you're nodding along, then check out GiveWell. an organization that researches questions about global health and philanthropy, even if a satisfying answer might require years of reviewing studies, talking to experts, and over 300 footnotes. GiveWell has now spent over 17 years researching charitable organizations and only directs funding to a few of the highest impact opportunities they've found.

13:45Over 125 ,000 donors have used GiveWell to donate more than$2 billion. Rigorous evidence suggests that these donations will save over 200 ,000 lives and improve the lives of millions more. GiveWell wants as many donors as possible to make informed decisions about high-impact giving. You can find all of their research and recommendations on their site for free. You can make tax-deductible donations to their recommended funds or charities. And GiveWell doesn't take a cut. If you've never used GiveWell to donate, you can have your donation matched up to$100 before the end of the year, or as long as matching funds last.

14:17To claim your match, go to givewell.org and pick podcast and enter econ102 with Noah Smith and Eric Torrenberg at checkout. Make sure they know that you heard about GiveWell from econ102 with Noah Smith and Eric Torrenberg to get your donation matched. Again, that's givewell.org to donate or find out more. Present kind of the bull case and the bear case for Databricks. Like what's the, if everything goes well, what does that look like? And if things don't go according to expectations, what does that look like? Well, the bull case, so the company grew 50 % last year and then is on a 60 % growth rate at scale.

14:57So there's a meaningful inflection in growth. Very rare to see a company actually accelerate, primarily driven by AI. The bull case is that Databricks becomes synonymous with AI infrastructure. Anybody building AI systems will be building on Databricks. combine that with starting to win some significant share away from existing Snowflake structured data customers, and it becomes a dominant platform. I think the counterpoint to that would be Sridhar now as a CEO of Snowflake is able to make pretty meaningful inroads and retain and continue to grow the existing Snowflake customers by presenting an extremely compelling AI story.

15:33And you've seen that, right? At the announcement when Slootman appointed SweetR and announced it, there's been significant recruitment of people with material AI backgrounds. You have the NVIDIA partnership. You look at what they've launched with Cortex, which is their AI product, and the announcements at the Snowflake Summit. In addition, just the cadence of PR has been completely different, much more aggressive, trying to position themselves within the world of AI, developing their own model that competes with Databricks' model and others. So I would have to say the pace of innovation from Snowflake is pretty significant.

16:10The question is, will customers move? And the answer so far, at least according to our diligence, is that it's very hard for a company to move from Snowflake to a Databricks. You're talking about pretty significant change management. And so at least today, we haven't seen much of that. And then, if I could just go on for a little bit longer, one of the accelerants of that transition and one of the justifications for the tabular acquisition, which was Databricks buying a company that commercialized a technology called Iceberg, is the ability to move data and particularly compute. So, okay, what do I mean by that?

16:49Well, a database is, there's what's called a query engine. So you ask it a question, it does a bunch of math, and it comes back to you with an answer. But when it asks those questions, it reads a bunch of files. And a database is basically the combination of those two things with a bunch of other stuff, security caching. But it's a query engine and a bunch of files. And Snowflake is a database that's optimized for a particular thing, structured to data analysis. And Spark's Databricks, or Databricks' Spark, is the same thing. Well, what's happened is we're starting to see the unbundling of the database.

17:21Well, now there are many different kinds of compute engines. There's Snowflake, there's Spark, there's MotherDuckDuckDB, there's PrestoTrino, Dremio, and graph databases. And what customers are asking for is, I want to pull my data out of the database. And why do big customers care about that? Well, it's expensive. It's expensive to store data two or three times in two or three different database centers. It's expensive to be compliant with international law if your data is in a whole bunch of different places. And now with these AI workloads, these companies actually want all their data in a simple place because if they decide to build some cool AI feature, it's just much easier to have it in one repo, one repository.

17:58And so Databricks bought a company called Tabular, which commercializes this data format called Iceberg, which allows you to basically move data across different workloads. So if you decide, I need an dashboard for the CEO and it has to have 60-second latency probably going on Snowflake. But if there's something that's a little bit slower, a little bit much larger, you might move it on to Databricks Spark. And the more this trend happens where the customers are in control of their data, the more compute workloads will start to work around, the greater the competition at the query engine layer.

18:29But you're much more bullish, right? Like of the two cases that you outlined, you think it's much more likely the bull case or say more about that? Oh, I think it's, yeah, yeah, for sure. I mean, look, to see a public scale company growing at accelerating their growth rate, I think they'll do just fine. If you, we ran a linear regression based on five key variables, which predicts about half, you know, has about prediction actually about 50%. That model basically predicts that Databricks would trade at about a 14x forward multiple, snowflakes roughly around 12. but I think if it were to trade today, if it were to list, it would probably trade significantly higher than that because of the public market appetite for AI stocks, which doesn't exist.

19:13I mean, the only way you can get AI exposure today is Microsoft, which is a massive conglomerate or NVIDIA, which is a fantastic chip company, but chip company nonetheless. And so there's a void there. And so I think when it lists, it will trade incredibly well. Yeah, it's really interesting. If you had to put like a basket of companies that if, you know, you just that similar to Databricks, you're making a similar kind of bet. What other companies would you put there? Well, public is pretty hard. I mean, I think, okay, so let's say you had a dollar to spend,$100 ,000 to buy any stock and you wanted to have maximum exposure to AI.

19:50Databricks would be in there. Snowflake would be in there. OpenAI would be in there. Microsoft would definitely have to be in there. They have a business that went from zero to 5 billion run rate in about 18 months. And that's just, I think, the beginning of their ability to cross-sell. I think deeper than that, it's hard. It's hard because you have commoditization at the model layer. And we can sort of talk about that. But if you look at like the large language models, they range in sizes from about 2 billion parameters, let's say, to about 175 billion parameters. And there's a benchmark that's called the MMLU, which is high school equivalency.

20:21So how good is this particular LLM relative to a high school grad? and they're all basically converging to somewhere between 70 to 80 percent of the performance even it's for the very very big ones so i think it's it's hard to kind of invest at the model layer particularly for and the capital intensity so then you probably have to go application by application okay let's take a look at code automation let's take a look at legal automation let's look at accounting automation let's look at security automation there the companies are much younger so i'll probably stick with my my mega cap basket for now that makes sense what do you think people most misunderstand or underappreciate about the company?

20:57What do they not fully understand or what do they get wrong? I don't think the technology and what it does is broadly understood. You know, a lot of people think about databases and an easy mental construct for it is just a really big Excel spreadsheet. And what Databricks does is imagine you had like a million Excel spreadsheets and they all had images in them. And what you wanted to do was, and all of those images had license plates on them. And let's say you wanted to count the numbers of A's and B's and C's and 1's and 2's and 3's, you would use Databricks to do that. And that's not a very sort of like grounded example, but just to give people a sense of scale about what the ultimate technology does, that's a really hard technical problem.

21:42And it's a hard technical problem not because the optical character recognition, like looking at the 1's and 0's and A's and B's and C's are hard. It's a hard technical problem because you're talking about thousands of computers all working together at the same time to achieve a single result. And so it's like I was a rower in college. And so it's like a rowing boat. But instead of eight people in it, you have a thousand people. How do you coordinate that? Yeah, that's really interesting. What else about the company do you feel like we haven't gotten into that is maybe interesting for people who are either students of the company or students of business more generally?

22:17I think a big push is the data catalog. And the data catalog was a strategic announcement for both Snowflake and Databricks at their most recent summit. Data catalog is the address book of the data ecosystem. So let's say, Eric, you create a table about number of podcast listens by episode by date. You might own that table, right? And I would create one about venture capital investments. I would own that table all within a company. Each of these data fields are related in some way. They all start from the same source tables and they all kind of evolve in a big company to the ultimate product that they'll be.

22:59And there's this strategic push to control that data catalog. Okay, that data catalog can be used for lots of different reasons. One is I need to look up who is the subject matter expert in a 10 ,000 person company who owns the podcasts table. Okay, it's Eric. There's another reason to understand the data catalog, which is lineage. How do all these fields actually come together? Lineage is important for compliance. So if there's an external data set that touches social security numbers, I really need to know that to make sure it's super secure. there's data quality which is is the data that's being fed through all these pipelines actually accurate are the distributions remaining the same are the data volumes constant so that the people who are just using deciding using data can be confident in it there's um bi anyway so that if you can think about like the data catalog it's basically the nexus it's an understanding or it's a roadmap of how all the data within a company fits and if you can control that and understand that, you can build a whole bunch of products around it.

24:02The challenge with the data catalog historically is nobody wants to update. It's like a wiki. Somebody creates a page once and then 10 years later, somebody comes back and visits it, but it is super valuable to a company. It's like an internet. And so now there's a strategic push to control that data catalog for both Snowflake and Databricks. There's an open versus closed source dynamic here, which also we haven't talked about, but maybe I'll just pause there on the data catalog. Yeah. Let's talk about the open-close dynamic. Sure. So Databricks is an open source project. I mean, Spark is an open source project, came from a lab, very difficult to implement.

24:37The commercialization motion for Databricks is the infrastructure to be able to manage tens or hundreds of thousands of machines operating this particular open source project, a little bit akin to Kubernetes, which was a technology that came out of Google. The company has done, Databricks has done a really nice job commercializing it. When they started to move into structured data workloads, they created a technology called Delta Lake, which is effectively a cloud data warehouse. And that was closed. That was closed source, which was a bit of a departure. And I think, and I don't know this for, you know, this is just my conjecture.

25:13I think that was basically a way of ensuring that other players within the ecosystem couldn't leverage the existing technology that Databricks might have built. And it was also a way of keeping people, their existing customers within a closed universe. What's interesting is Snowflake started closed source and is now pushing more into open source. And so they're basically like complementing each other. One was closed source moving into open Snowflake. Databricks was open. It's now moving more into closed. We'll see what they do with the iceberg acquisition. And so there's this like, it's like two sumo fighters trying to figure out, okay, where am I weak?

Read the full transcript

25:47How do I, how do I buttress my weakness when I'm facing my opponent? And so we'll see how this all sort of evolves. And I think one big question is Iceberg, which is the core technology of this company, Tabular, is open source today. It's an Apache license. The CEO from that company is on the steering committee for that project. And so there was a lot of questions about, does that remain open source? Is it closed source? How does this impact Databricks' strategy? my sense is they'll probably open it I think that's sort of the illusions that they've made in the market or to keep it open and open other things but it's still completely TBD and commercially how will that impact a business is also unclear do people actually care, does the data but in terms of the hearts and minds of developers open source matters a lot and what ends up happening with lots of open source companies as they grow, as they change the licenses as they shift to more commercialization And it's very easy to alienate those communities.

26:42You look like HashiCorp when they change their license from whatever open source license they had to a business source license to commercialize it more. There tends to be a lot of flack and a lot of PR management around that evolution that whichever decision Snowflake decides to go with their open versus closed source and Databricks with their open versus closed source, whichever path they decide to go, they will definitely need to manage. What else about the data catalog sort of line of discussion is worth mentioning or making sure that it's being kind of fully understood here about the significance of it?

27:14It's a strategic component to the infrastructure that historically has been very difficult to commercialize. So there are some companies, some great companies building data catalogs. But because it's not self-updating and because nobody really owns it, no one is basically promoted for managing a beautiful data catalog. it's probably a complement to one of those key products, business intelligence, compliance, data quality, security, rather than a standalone product. And so now all of a sudden, and I think we'll start to see this more, Snowflake and Databricks will start to compete or at least step on the toes of some of their key partners, particularly because they're vying for bigger and bigger markets.

27:53They want to show higher and higher growth rates. And so they'll probably start to acquire or compete as they push out. The most valuable commodity in this space is a compute workload. It's not storage, but they're related. But what Snowflake and Databricks are both competing for is the opportunity to process data. They want compute. Storage is important because if the storage is with Snowflake, then it's very easy. There's no friction associated with the compute. If it's outside of Snowflake, then you have to pay. Amazon charges you based on ingress and egress fees, moving data in and out. So there's friction there.

28:31So this is why these companies are really, they care about the storage, but the crown jewels, the most valuable part, the filet mignon of this market is the compute. And so that's why compute heavy workloads, ETL will be really important. Materializing all these different views, like the DBTs, these, the five trends, I think they'll play a really key part here in the next five years. And it'll be interesting to see, uh, how the, what, as the elephants dance, how all these businesses need to respond. Yeah. And can you explain those businesses as well? You mentioned the DBT five trend, can you give a little bit of insight into, into those businesses you just mentioned and how they work?

29:09Sure. Yeah. So let's go back to the point of sale. So, you know, I started a business for selling widgets, all the data is stored within the point of sale, like a Square or whatever. And we need to move that data into Snowflake. And so we need a connector from Square into Snowflake and then a pipeline, some computers that calculate, make sure the data is correct. If the connection drops, we try. That's 5Tran. And this is super high level. And then once it's in the database, it will be in one particular format, but let's say I need to change it in a particular format. So let's say the transactions are listed by microseconds.

29:44And I really want to see them aggregated by the month. Well, you need to transform the data. How do you do that in an intelligent way? There's a technology that's called DBT. Looker had a technology called LookML that does this too. So you can create metrics, define them once, and then reuse them in multiple places. So 5Trend are the pipes that move the data. And then DBT is once it's gotten to its particular place, how do you reconfigure the data to be in the format that you want for the right output? Do you have a request for startups along these lines? Like, has everything kind of been picked over or where are we?

30:17No, I think, well, a lot is changing within the sort of the modern data stack. A lot is changing when you have lots of AI enabled BI companies that allow you to ask questions of a BI system. And I think that there's definitely a role for that. I think what will happen is you'll see a pretty significant wave of consolidation here over the next two years. and once the dust settles, then it will be easier to pick your spots. But right now it's tough. And one of the reasons is the modern data stack has been very heavily financed, particularly through 21 through 23. And there has not been a lot of consolidation aside from the very, very largest players.

30:58That's really interesting. How much as an investor, are you sort of like, hey, there's this great opportunity, like tops down versus bottoms up of like, I think this thing should exist. Let me go look at every, everyone who wants to do something here or is doing something here versus just kind of like responding to the, like, yeah, how are you, you're sort of what you're looking for evolving over time? Well, we're, we, we're extremely thematic. Um, and so we'll spend lots of time researching spaces, trying to find what's the angle, where is there a pain point? Where is there a top three priority that's meaningfully underserved?

31:33And so we've been spending time in the modern data stack. The thing that we hear from buyers is, one, don't sell me another tool. I don't need another tool. The second is cost reduction is really important. And then the third is either my board or my CEO is really pressing me for an AI strategy. And I need an answer, but I'm afraid because I've never shipped production systems in the past. Technically, it's probably okay, but I really don't want to be fired for losing a bunch of data or exposing somebody's social security number. And so how do I make sure that those things are secure? We have looked at the large language model security space, and there are many different approaches there.

32:11So we definitely are spending time to try to understand data loss prevention and prompt injection attacks and model poisoning data toxicity, those kinds of things. But still, it's also really early. Our estimate is there's probably less than 5 million of ARR in that space today. And when you did Looker at the A or the B? At the A. And when you got excited about Looker, was BI still a very nascent space? Or what did the market look like? So in the late 1990s and early 2000s, there were four, basically the four original, at least from my view, four original BI companies. There's just Cognos, Hyperion, Business Objects, and MicroStrategy.

32:53All those companies were worth about$2 to$3 billion, which is about$8 to$12 billion today. They were three-layer cakes. So they had databases in them, they had caching layers, and they had visualizations on top. And they were super tightly controlled by IT. When I was working at Google, there was a MicroStrategy instance. And there was a very small team, and it took forever to get the metrics. Tableau then came out of Stanford and said, these systems are really locked down. And they're also very complicated. Let's just take the top layer of the cake, visualization, and let's have anybody be able to make a pretty chart.

33:27So that business grew to be$16 billion over the course of about 14 years, grew bottoms up$5 ,000 to$10 ,000 at a time and a spectacular outcome. And so the pendulum swung from super centralized, very structured to very decentralized, no structure, bottoms up. That worked for a long time. Then what happened, then the cloud data warehouses started to come around. And there was Redshift, which was the fastest growing product at AWS. And there was BigQuery, which was Google's database. And then Snowflake came around. And so the classic sort of centralized IT systems were not architected to handle that kind of a scale because the databases they had inside were for much smaller data sets.

34:07And so the idea behind Looker was try to strike a balance between centralized and decentralized control, but really build a BI system that could work natively with the cloud data warehouses. And so originally, Looker would sell and bring Snowflake in because the buyer wanted a complete package. And then the Snowflake started to grow and grow and grow. It became the other way around where Snowflake had actually many more AEs than we did. Not to say that Looker wasn't growing fast, but their ability to spend and their ability to raise capital was huge. And so people wanted this bundled solution.

34:36Now the next wave of BI, we're investors in a company called Omni, which is the Xlooker team. And they're again, striking the balance there between centralized and decentralized IT in a really beautiful way. Is there anything you think we didn't go into in full depth that you think we should cover more? Otherwise happy to either on Databricks or any other company we've discussed or happy to wrap on that. You know, I think one question that we are actively researching is whether these databases like Snowflake and Databricks will ultimately be part of applications and will they serve AI models? And I think they're trying to push in that direction.

35:13They were never architected for this, but if they can push into the path of production, it opens up a new market for them. Because I think the fastest growing part of software today in terms of venture-backed software are inference companies. So what does that mean? Well, you have a large language model. You do two things with models. You can train them, teach them how to predict, or you can ask them to predict. And so the first one is training. The second one is called inference. For a long time, training was by far the largest part of the market, 80 to 90 % plus all the GPUs were there just to train and train and train.

35:46Like Facebook's buying 20 billion of GPUs this year. Well, just to kind of give you a sense, and then Google and Microsoft and Amazon spending 12 to 15 billion a quarter. Half of that is on GPUs just to build out data centers. And now what we're starting to see is inference. So it's actual usage when a user asks a question, well, you have to ask for a prediction from a large language model where NVIDIA is trying to push into the inference market. And I can imagine Snowflake and Databricks also starting to push meaningfully into this market because the growth rates on these companies are enormous, like 1 to 20, 1 to 30, 1 to 40, 50 to 150 in a year.

36:18And that's just because the demand is there. It's a commodity or relatively commodity market, but the total top line growth is staggering. So I could imagine acquisitions there for both of those businesses or meaningful product pushes and then tying the training systems ultimately to the inference is a way of keeping all the data within an ecosystem and capturing more and more customer spend, particularly for the very largest businesses. Yeah, that's a good note. I have one more question, which is when you talk to your peers, the people whose opinion you really respect and who know the most about some of the topics that we've talked today, and you have disagreements with them about where something is or where something is going or the state of something, what is an example or what are some of those disagreements?

37:06Value of NVIDIA in three years.

37:12I think that's really interesting. If you were hosting a dinner, I mean, two, three years ago, the question was, what is the price of Bitcoin in 18 months? That was sort of the catalyst for a great debate. I think there's an equivalent question here, which is what is the total market cap of NVIDIA in three years? And there are bulls who believe that the demand for GPUs is basically unquenchable here for the next five years as a result of small fabrication capacity, none and effectively very little in the US. That, you know, Tesla is buying 10, I mean, whatever, 15 billion of GPUs, Facebook's buying 20 billion of GPUs.

37:48And this kind of goes to the moon. and you believe that larger and larger models become more and more important. And they go from 150K a run to 10 million to train to a billion to train, which is what Amazon has had recently. And so you just have scale. Scale dominates, GPUs rip, and NVIDIA continues to go to the moon. It's already the most valuable company in the world. And then there's another view, which is Amazon and Google have their own GPUs. AMD released the chip set that's about 30 % more powerful than the current generation of A100s. And this market ultimately reaches commoditization.

38:24I mean, you look at NVIDIA's revenue. Revenue has grown 5x in the last four years. Profitability has gone from$4 billion in profits to about$46 billion in four years. You really have this beautiful engine of a business. And so how sustainable is that? How does that grow as more GPUs enter the market? What does the dynamics look like? And what do you think is the steel man of the other side or like the best argument or something that could almost convince you, but not quite. Of what? That NVIDIA would be worth less? Yeah. I think the biggest argument is small language models start to really dominate.

39:01So you look at Lama's 8 billion parameter model that was over-trained on data. It's really good. And it compares favorably to 70 billion parameter models and even some 140 billion parameter models. and then a large number of the enterprise use cases don't need parameters in the billions they need much much much smaller models that can be deployed on a computer or a mobile phone for privacy reasons or latency reasons and so these big fancy gpus become less important i think another dynamic there is just the competition from like google and amazon on the google's tpus and amazon's inference chips where the bottom falls out of the market because you have there's this awesome paper we studied in business school called the bullwhip effect the bullwhip is you know you crack a whip and you can see a wave going through it whole idea so in business school we had this belgian professor and he made us play this beer game sort of illustrate this and there were four students one was a beer maker another one was a a wholesaler another one was a distributor the first distributor then a wholesaler then a retailer and the professor would modify the demand just a little bit at the retail level it's like one or two units a week.

40:07Then it would go back up to the producer and the producer would see these massive swings in demand. And there's this great paper from the Sloan School if you want to read it. And then that massive swings in demand would go all the way down. And by the end of like four or five cycles, maybe 10 cycles, everybody in the supply chain is bankrupt. And the producers turn to the retailer and say, what is going on? And it seems like all of a sudden all these people want this beer and then the next week nobody wants any beer and you can look at nvidia's uh any one of the multiples like evita or pe multiple and over the last 15 years you see this huge spike when gaming became really big and the next generation mmos and the first person shooters really needed beefy gpus and then it surged and then fell and then there was the next generation which was the GPUs used for crypto mining.

40:58And so the PE multiple went from about like 15, 16, all the way to 50 or 60, and then back down to around 25. And so you have this like, particularly at the chip layer, you have this incredible amount of volatility that exists because it's not a recurring business. It's not a subscription business. And at some point people will have enough GPUs. And this is a big reason, and I completely agree with the strategic direction, that NVIDIA is pushing more and more into software. Again, I would not be surprised if they start getting in the business of inference, but they've started to push into weather prediction and lots of different other applications as a way of diversifying their revenue and stabilizing the value of the business.

41:33I think that's a good explanation of both sides. How confident are you in your position? I don't know how to answer that. I mean, I think it depends over what timeframe. I was looking at leaps, which are long-dated options for NVIDIA. So I think this is the question of the day, which is how long does this GPU demand last? And I don't know if I can answer that for you. So 50-50, you know, coin flip.

42:09Yeah, makes sense. So let's wrap on that. For people who want to go deeper on your work, if they enjoyed the conversation here, I want to point them to your blog, to your Twitter. any other plugs that people should know about listening? No, there's three places. There's TomTungus.com, there's Twitter, and then there's the LinkedIn. So you can find the content there as well. Cool. Awesome. Well, Tom, thanks so much for coming on the podcast. And until next time. Thanks for having me, Eric. Such a pleasure. Talk to you soon. Turpentine VC is a podcast from Turpentine, the network behind Moment of Zen and Econ 102.

42:48too. If you liked the episode, please leave a review in the Apple Store or rate us on Spotify.

From the publisher

In this episode of Turpentine VC, Erik Torenberg welcomes back Tomasz Tunguz, GP of Theory Ventures. They dissect the investment landscape of data businesses through the lens of Databricks. They analyze the company's growth trajectory, competition with Snowflake, and its potential in the AI era. Tom presents bull and bear investment cases, discusses valuation metrics, and explores the strategic importance of data catalogs. They also cover the emerging trends in the modern data stack, potential market consolidation, and the impact of AI on investment opportunities.


🔥 Apply to join over 400 founders and Execs in the Turpentine Network: https://hmplogxqz0y.typeform.com/to/JCkphVqj


—

RECOMMENDED PODCASTS:


🎙️ This Won't Last - Eavesdrop on Keith Rabois, Kevin Ryan, Logan Bartlett, and Zach Weinberg's monthly backchannel. They unpack their hottest takes on the future of tech, business, venture, investing, and politics.

Apple Podcasts: https://podcasts.apple.com/us/podcast/id1765665937

Spotify: https://open.spotify.com/show/2HwSNeVLL1MXy0RjFPyOSz

YouTube: https://www.youtube.com/@ThisWontLastpodcast


—

SPONSORS:


🛠️ Building an enterprise-ready SaaS app? WorkOS has got you covered with easy-to-integrate APIs for SAML, SCIM, and more. Join top startups like Vercel, Perplexity, Jasper & Webflow in powering your app with WorkOS. Enjoy a free tier for up to 1M users! Start now at https://bit.ly/WorkOS-Turpentine-Network


💥 Head to Squad to access global engineering without the headache and at a fraction of the cost: head to https://choosesquad.com/ and mention “Turpentine” to skip the waitlist.


—

LINKS:


Tomasz’ website: https://tomtunguz.com/

Theory Ventures: https://theory.ventures/ 

Databricks: https://www.databricks.com/ 

Databricks' Accelerating Growth: https://tomtunguz.com/databricks-growth-2024/ 


—

FOLLOW ON X:


@eriktorenberg (Erik)

@ttunguz (Tomasz)

@turpentinemedia (Turpentine)


—

TIMESTAMPS:


(00:00) Introduction

(00:40) Tomasz' initial thoughts on Databricks

(02:09) Understanding Databricks' technology 

(05:00) Unstructured data pipelines & applications

(06:11) Databricks' market opportunity & competitors

(08:05) Investability of Databricks

(10:24) Databricks vs. Snowflake: investment insights and market predictions

(13:10) Sponsors: WorkOS | Squad

(15:36) Bull & bear case for Databricks

(14:12) Unbundling of the database

(17:23) Databricks bullish outlook & investment considerations

(21:51) Databricks: a difficult technical problem

(23:15) The power of the data catalog

(25:25) Open vs. closed source: the Databricks & Snowflake dynamic

(28:58) The importance of compute workloads

(30:01) The modern data stack and the AI strategy

(31:31) Investor perspective and the evolution of the market

(33:30) The evolution of business intelligence 

(35:55) Databases & Al models

(37:46) The GPU market and NVIDIA's future

(43:09) Wrap 

More from "Turpentine VC" | Venture Capital and Investing

All 87 episodes
E53: Tomasz Tunguz on the Evolution of the Data Ecosystem"Turpentine VC" | Venture Capital and Investing · 43 min
Listen in VO