Trino, Iceberg and the Battle for the Lakehouse | Justin Borgman, CEO, Starburst

30 Jan 2025 · 1 h 6 min · 38 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Starburst’s approach to the “lakehouse” data layer—Trino query engine, Apache Iceberg open table format, and the “Ice House” end-to-end lifecycle—plus hybrid/on-prem vs cloud strategy and product architecture.

Guest backgrounds

Justin Borgman is CEO of Starburst and a serial founder in data infrastructure. He previously co-founded HadoopDB (a Yale spinout) in the early lakehouse era; it was acquired by Teradata. At Teradata (starting 2017), he worked on next-gen data warehousing and helped discover and contribute to Presto/Trino (originally from Facebook).

Key claims

Lakehouse performance gaps vs warehouses are now “de minimis,” making lakehouses the future. Starburst differentiates via hybrid deployment, open formats, and end-to-end Iceberg management (“Ice House”). Iceberg won due to adoption by large internet companies; Starburst pairs Trino/Presto with Iceberg and supports streaming ingest and governance. Galaxy was rebuilt after an initial SaaS architecture failed to meet the “Apple approach.”

Notable examples

Trino used by Netflix, Airbnb, LinkedIn; Iceberg adoption by “super scaled” internet companies; bank CEO quote: “we’re never going to move everything to the cloud” (four clouds, incl. on-prem). Dell Lakehouse is powered by Starburst for AI on-prem.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Starburst and the Data Landscape

0:45 to 2:27

Justin Borgman explains the role of Starburst in the data analytics field.

“And while there's been widespread embrace around open formats and Iceberg in particular, we've been doing this forever, though.”

Analyzing Databases: Transactional vs. Analytical

2:27 to 5:03

Discussion on the differences between transactional and analytical databases.

“And then we can go into all sorts of technical details.”

Justin’s Journey: From HadoopDB to Starburst

5:03 to 7:18

Justin shares the story of his first company and its evolution into Starburst.

“So you and I, we've done this, a variation of this a couple of times.”

Challenges Faced While Building Starburst

7:18 to 9:57

Justin discusses the challenges during Starburst's early days and the pivot from Teradata.

“And this was at a particular point in Teradata's history where they were starting to feel the pain of both open source and cloud.”

Transitioning from Presto to Trino

9:57 to 12:23

The evolution of Presto to Trino and the challenges of brand recognition.

“And that was, I guess, serendipitous for me and my team because it allowed us to then create a business around this independently.”

Comparing Data Lakes, Warehouses, and Lakehouses

12:23 to 14:00

Comparison of data lakes, warehouses, and lakehouses and their functionalities.

“And yes, a lot of old fashioned, you know, developer relations and meetups.”

The Evolution of Data Storage

14:00 to 15:00

Learn about the transition from traditional data warehouses to lake houses.

“You could store infinite amounts of information.”

Performance Improvements in Data Warehousing

15:00 to 16:00

Discover how file formats and query engines have improved analytics performance.

“You know, Teradata was just dramatically faster than what you could do on Hadoop at the time.”

The Lake vs. Data Warehouses

16:00 to 17:10

Understand the differences and advantages between data lakes and warehouses.

“But you couldn't really run analytics because it's not super structured.”

The Rise of Hybrid Models

17:10 to 18:20

Explore how hybrid data solutions are integrating cloud and on-prem systems.

“So, data lakes on one end of the spectrum, cloud data warehouses, and then lake house is meant to be best of both worlds in the middle, combining the advantages of both.”
Show all 38 chapters

Trino and Iceberg as a Reference Architecture

18:20 to 19:20

Learn about the partnership between Trino and Iceberg in data querying.

“It's like just leave your data anywhere and we'll find it wherever it is for analytical purposes.”

Starburst's Product Offerings

19:20 to 20:40

Get an overview of Starburst's products and their unique features.

“I think part of it is the industry is starting to mature.”

The Importance of Open Source in Data

20:40 to 21:40

Discuss the role of open source in data architecture and innovation.

“So we're an open engine querying open formats.”

Challenges and Opportunities in Open Source Business Models

21:40 to 23:00

Understand the balance between community contributions and commercial needs.

“And there are some use cases where being able to join a table that lives in another database system with the table that maybe sits in the lake can be very valuable in getting a fast response time.”

Enterprise Solutions and Customer Needs

23:00 to 24:20

Explore how enterprise solutions meet complex customer demands.

“These technologies get really tested to the limit that way and can be a good indication of sort of where things are going.”

Cloud vs. On-Premises Debate

24:20 to 25:40

Examine why some companies are resistant to moving entirely to the cloud.

“And then around it, we've built all the enterprise functionality that you would need.”

The Future of Data Infrastructure

25:40 to 28:00

Discuss future trends in data infrastructure and the role of AI.

“And presumably as people, as customers build stuff, they presumably contribute some of it to the community.”

Future of Multi-Cloud and On-Prem Solutions

28:00 to 29:10

Discussing the trend of multi-cloud strategies and on-prem solutions in the context of AI.

“I won't name his name just to protect his innocence, I guess.”

Starburst Enterprise and Galaxy Overview

29:10 to 30:50

An overview of Starburst's Enterprise and Galaxy offerings, focusing on their functionalities and adoption.

“Yeah, so that's a classic, you know, SaaS product that's hosted and managed by us.”

Connecting Data Sources and Query Execution

30:50 to 33:20

Explaining how Starburst connects various data sources and performs query execution efficiently.

“And so those things all really matter when you're, you know, ultimately fitting into their margin.”

Real-Time Analytics and Data Types

33:20 to 35:10

Exploring the capabilities of real-time analytics and the types of data Starburst can process.

“Or you can use a BI tool like Tableau or ThoughtSpot or others.”

Governance and Performance Optimization

35:10 to 37:05

Discussing governance features and performance optimization techniques within Starburst's architecture.

“The data itself can be structured versus unstructured.”

Data Applications and Customer Use Cases

37:05 to 39:35

Describing how customers utilize Starburst to build data applications, including specific use cases.

“So you have, joke aside, you have, so ingestion we talked about, you have table maintenance.”

Challenges in Building Galaxy

39:35 to 41:05

Justin shares insights on the challenges faced during the development of the Galaxy platform.

“But so for the internal use case, like I'm a bank and I want an AML and I'm an online drink.”

Re-architecting for Seamless Experience

41:05 to 42:00

Detailing the decision to re-architect Galaxy for a better user experience after initial failures.

“You know, I would say building Galaxy, the SaaS platform, you know, there's a story there where we actually built it twice.”

Navigating Nerve-Wracking Decisions

42:00 to 43:04

Learn about the challenges of re-platforming and the importance of investor patience during development stages.

“And, you know, that basically took another year, you know.”

Managing Product Complexity

43:04 to 44:16

Discover how to manage multiple products with shared components while addressing unique customer needs.

“I'm sure that must have been made for good board meetings.”

The Tug-of-War: On-Prem vs. Cloud

44:16 to 46:17

Explore the complexities of offering both on-prem and cloud solutions in today’s market.

“You know, if you're a big multinational bank with an on-prem footprint, enterprise is probably your best bet.”

The Rise of Iceberg in Data Management

46:17 to 47:37

Understand the impact of major acquisitions on the adoption of data formats like Iceberg.

“And I think part of that was also our bootstrapped orientation where we were very customer obsessed because we depended on that revenue to like pay the next month's payroll.”

Strategic Insight on Acquisitions

47:37 to 49:39

Analyze the implications of acquiring companies built around open-source projects.

“Was that the acquisition of Tabular by Databricks?”

The Evolution of Open Data Formats

49:39 to 52:12

Learn why open data formats are gaining preference in enterprise software over proprietary systems.

“Then you spend$2 billion, you buy a company.”

The Current State of Data Mesh

52:12 to 55:26

Evaluate the evolving concept of data mesh and its implications for organizations.

“technology, wherever there are two things that are mostly the same and one is open and one is not, the open is going to win over time just because economics will eventually drive long-term decision-making.”

Integrating AI into Data Strategies

55:26 to 56:00

Explore the role of AI in data access and model training within organizations.

“Of course, the topic that we need to talk about is AI.”

Data Access and Transformation in AI

56:00 to 57:15

Learn about the importance of data access and transformation for AI models.

“Your models are only as good as the data that you train them on, and, you know, access to more data or better data is going to influence that.”

Lessons Learned in Selling Enterprise Software

57:15 to 59:17

Discover key insights on selling enterprise software and the importance of partnerships.

“Are you mostly outbound, sales driven, presumably given the type of customers?”

Building a Services Network for Startups

59:17 to 1:01:19

Understand how to build a services network and the importance of quality delivery.

“So, actually, a very interesting topic, services for, you know, any founder listening to this.”

Navigating Partnerships with Larger Companies

1:01:19 to 1:03:48

Learn how to effectively partner with larger companies and the importance of internal motivation.

“which is more of a, you know, ISV kind of, which is the other sort of side of the partnership world.”

The Future of AI in Production Use Cases

1:03:48 to 1:05:55

Explore the future of AI and its transition from experimentation to real-world applications.

“So people should, yeah, should view those as very long term projects.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hi, I'm Matt Turck from FirstMark. Welcome to the Matt Podcast. As AI takes over the world, the battle for domination of the data layer is more intense than ever. My guest today is Justin Borgman, the CEO of Starburst, one of the key players shaping how enterprises handle their exploding data needs. We covered a lot of ground in this conversation, including in particular Starburst's key strategic moves over the years. Building Galaxy, we actually built it twice. We wanted to take a more, I'll call it like a more of an Apple approach. That basically took another year. The critical moment they had to restart their open source project pretty much from scratch.

0:34It was tricky because you're sort of starting from zero in terms of name recognition. The early bet on the lake house architecture and Apache Iceberg. You really need fast read performance. Lake houses are the future. And while there's been widespread embrace around open formats and Iceberg in particular, we've been doing this forever, though. We have that sort of end-to-end lifecycle around Iceberg, and we call that the Ice House. The decision to support on-prem deployments when many others bet everything on the cloud. I had dinner with the CEO of one of the largest banks in the country. He told me point blank, we're never going to move everything to the cloud.

1:06And lessons learned from building their big partnership with Dell. Where are we? What are the negotiation points? What are we working through? It really needs to be that kind of priority on both sides to make it work. Now, Justin is a very thoughtful serial founder in the data infrastructure space. And he's particularly great at explaining complex technical concepts in simple terms. Please enjoy this great conversation with Justin. Hey, Justin. Welcome. Thank you for having me. You're the CEO of Starburst. Let's just start with the one-minute elevator pitch just to frame the conversation. What does Starburst do?

1:37Sure. So Starburst is a data platform for analytics, building data apps, and now increasingly incorporating AI in the applications that you build. We're the creators of an open-source project called Trino, which is a pretty popular project used by a lot of the big internet companies like Netflix and Airbnb and LinkedIn and so forth. And it essentially allows you to run fast SQL queries on data in your lake, as well as connecting to other data sources as well. So you can really run analytics across all the data that you have, be it in a traditional database, a data lake, on-prem, in the cloud, you name it.

2:16So I'd love to start the conversation with a little bit of a sort of broad overview of the space to make this interesting to, you know, anyone that may be curious about the world of data infrastructure. And then we can go into all sorts of technical details. But like to start with, what world do you operate in if we think of all of this as, you know, databases? So databases, data warehouses, how would you sort of compare and contrast the various databases of the world? So I think of the world, the database world, as being really divided into two halves, the analytical side and the transactional side.

2:55The analytical side are systems that are built to be very read optimized, so reading data as fast as possible. The transactional world being more oriented towards writing data fast and consistently. So when you think of a transactional system that's maybe powering an application, you know, a canonical example would be like an ATM that needs to record a debit and a credit very quickly and has to be consistent every single time. An analytical application would be something like how many customers bought product X last year and slicing and dicing the demographic profile of your customers and understanding their journey, you know, through the website all the way to a transaction at the end.

3:36And so those are broadly speaking the two worlds. Yeah. And when you say writes, what matters most is to like never lose any data and the data needs to be like 100 % correct 100 % of the time. Whereas analytical is what do you optimize for in the read? Yeah. So analytical is maybe a little less – I'll call it maybe mission critical in the sense of, you know, losing a byte is not going to be the end of the world. That's the transactional side, which is very focused on ensuring that consistency. But on the analytical side, you have a different challenge, which is how do I process massive amounts of information, do very complex joins of information across different tables to get to an answer as quickly as I can?

4:21And that requires a lot of, I'll call it rocket science, in the optimization of how those queries actually get executed. And that's what we focus on on the analytical side. And maybe to just like drive it home, some examples of famous company names in the first bucket and in the second bucket. Yeah. So on the transactional side, you'll sometimes hear people also call them operational database systems. So I think of operational, transactional as sort of, you know, interchangeable words to some degree. You'd think of Oracle or MongoDB as being two very prominent examples on that side of the house.

4:54On the analytical side, you'd think of Teradata, Snowflake, Databricks, and Starburst. And Teradata acquired your first company. Yep. Right. So you and I, we've done this, a variation of this a couple of times. Like the first time I was looking this up was in 2013 when you were running HADAP. So tell us a story, like your first company, what led you to where you are today? Yeah, my first company was actually a spin out from Yale University. And it was really commercializing the research of my co-founder, a guy named Daniel Abadi, and his PhD student, Kamil Bayda-Pavlikovsky. The research was called HadoopDB, which was really like the first imaginations, I would say, of a lake house architecture, which is what's become pretty popular today.

5:43But back then, we were sort of the first to really do that. And this was like 2010 timeframe. Hadoop was just gaining momentum as the data lake of the day. And we were some of the first to think about using this for analytical data warehousing purposes. And so we built that business over four years, raised venture capital, and eventually were acquired by Teradata, as you said. And it's funny how history repeats itself in a lot of ways because the concepts of a data lake, the concepts of doing analytical processing on a data lake were really things that were being played with and pioneered 15 years ago.

6:18But today have now become really the status quo or what everyone is doing today. So it's sort of interesting to see how the space has evolved over that time frame. And I'd say the biggest thing that has changed, you know, over 15 years is that Lakehouse technologies have just improved dramatically in terms of performance and functionality. And I think that just goes to show that building databases is hard and it takes a long time to get quite there. So 15 years later, we finally delivered on the vision from, you know, 2010. Yes. And so when did you leave Teradata to start Starburst? So that was 2017.

6:56Yep. Yep. And yeah, maybe tell us about the early days and that worked. Yeah. So when I was at Teradata, my job, so Hedapt had been acquired. I became a vice president general manager of a small business unit at Teradata that was really focused on next generation technologies and thinking about the future of data warehousing. And this was at a particular point in Teradata's history where they were starting to feel the pain of both open source and cloud. Snowflake was still a young company. It wasn't materially impactful to Teradata yet, but you could see them on the rise. And companies like Cloudera were certainly taking workloads away and causing some pain there.

7:39And so my job was sort of to solve the innovator's dilemma, if you will, and figure out the future for Teradata. And so in that vein, I actually discovered a new open source project, which is today known as Trino. Back then it was called Presto and it was coming out of Facebook. So Presto later became Trino. And it was really how Facebook was running all of their data warehousing analytics. It's very interesting because Teradata was always trying to sell their data warehouse to Facebook and never really getting anywhere. And in fact, I had a lot of meetings with Facebook as a Teradata person at the time.

8:14And, you know, it was like, oh, wow. Okay. They actually do, you know, a lot of what Teradata can do, but they built it themselves with this open source project. And they've run it at ridiculous scale, hundreds of petabytes of data, thousands of queries. And that got me really turned on to this idea that maybe this could be the future. And I'd already seen the power of open source just watching Hadoop's rise and the viral adoption of that technology that I thought, you know what? Open source is a really powerful distribution mechanism. So, you know, this could be something. And that began a, I'll call it maybe a awkward marriage of sorts, where my team from Teradata actually started working with the team from Facebook to contribute towards this open source technology and advance it.

8:58And, you know, we built the query optimizer for Presto or Trino now. You know, we added access control. So, while at Teradata? While at Teradata. As a like service offering of some sort? Yes, yes. So the idea, at least my business plan, if you will, for Teradata was to turn this into an actual product line that would, you know, have the virality of an open source project and that we could build an enterprise offering around it. Unfortunately, during my time there, Teradata had a number of changes at the executive level, the CEO, multiple CEO changes actually during my brief period there. And nobody could really get on board with what we were doing.

9:35I think they were all concerned that it might be competitive or cannibalistic towards the existing business, which I can understand those concerns. But that's also kind of goes to the innovator's dilemma. Sometimes you have to eat your own lunch if you want to still have lunch. And so long story short, we parted ways in 2017. And that was, I guess, serendipitous for me and my team because it allowed us to then create a business around this independently. And that's really how Starburst was born. And then you co-founded the company with some of those top Presto contributors from Facebook. That's exactly right.

10:14Yep. So you convinced them to just leave the mothership and strike out on their own. Okay. Exactly. And I think, you know, the project itself has just grown tremendously since that time frame. So they've been able to realize, I think, a couple of benefits there. A, you know, the rise of a commercial business that's quite successful today, but also seeing the open source project just become more vibrant and users continue to grow. Yeah. And the whole Presto to Trino kind of evolution was partly because at some point Facebook said, no, this is our open source project. Yes. That was a tricky moment in our history just because it's so hard to create awareness of any brand name.

10:56And essentially because Presto had been created at Facebook, they own the trademark, just sort of the way intellectual property laws work. And so when my co-founders left Facebook and joined Starburst, you know, they continued development of the community and of the open source technology, but we're still calling it Presto. And that led to a sort of legal challenge where they had to rename it. And that's why Presto SQL became known as Trino now today. And there was a rebranding. But it was tricky because you're sort of starting from zero in terms of name recognition at that point. Even today, there are a lot of people who still know.

11:33So Trino at some point had zero stars on GitHub or whatever. Exactly. Yes, and you had to rebuild from scratch. Rebuild from scratch. So that was, yes, definitely. If there's a book later about Starburst, that will at least have a chapter for sure. Sounds like a fun time. So how did you build it from there? Was it just like good old, I mean, clearly I'm sure the community knew what Trino was, but then it was like good old developer relations and meetups and content to build the open source project? That's exactly right. I mean, the hardcore, you know, Presto users immediately became hardcore Trino users and followed, you know, my co-founders in the transition.

12:13But outside of that group of like, you know, the people in the know, the broader market had no idea what Trino was. And so we had to recreate that awareness, which is very challenging. And yes, a lot of old fashioned, you know, developer relations and meetups. Of course, we also had a pandemic right around this time. Oh, yeah. So the in-person element was completely taken out. Nice. Which was, yeah, which was a major bummer because I think in-person events are so great. That was a big challenge for our marketing team, our developer relations team, our community team. But, you know, fortunately today, I think, you know, the awareness of Chino has grown.

12:50And I think more and more people are finding out about it. And is Presto still active or is it completely different at this stage? Completely different at this point. That's right. They started as identical copies, but the code bases have diverged quite a bit. And today, Presto is really just used by Facebook. So it's sort of like their own private branch in a way used by a small number of people. And Trino has become the mainstream community branch. You know, Apple, LinkedIn, all the big guys are contributing there. So let's talk about, you know, lake houses and all the things. So again, in an sort of effort to make this interesting for a broad group of people and, you know, the hardcore data engineering people can fast forward through this.

13:34But can you compare and contrast data lakes versus data warehouses versus lake houses to make it super clear? Sure. So, you know, the term data lake really originated in the, you know, call it maybe 2011, 12 timeframe when Hadoop was gaining momentum. It was described that way because you could basically store anything in it. It was just this distributed file system that was infinitely scalable. You could store infinite amounts of information. And from that, people started to say, hey, you know, we want to do real analytics. We've got to actually optimize the way we store the data, the way that we lay it out in this lake, in this file system, and started to build new file formats that allow you to basically store the data in the same type of way that you would a traditional data warehouse.

14:26So there were early files that are still very popular today called like Parquet files or ORC files or Avro files, which were basically columnar representations of the data. And without getting into too much technical detail, storing the data in that way delivers faster read performance. Back to the opening point about analytical systems, you really need fast read performance. And so how you store and lay out the data has an impact on that. And so people kind of like recreated the storage of traditional data warehouse platforms, but did it in a lake. And that's really where the lake house evolution took place, which was doing data warehousing activities now in a lake and calling that a lake house.

15:06And I think, you know, that is what has really evolved over a 15-year period was in the early days, again, when I was working on my first company, the gap between doing analytics in a data lake and doing analytics in a data warehouse was pretty significant. You know, Teradata was just dramatically faster than what you could do on Hadoop at the time. But as a result of the evolution of those file formats and the query engines themselves getting faster and faster and faster, 15 years later, that performance gap is de minimis at this point. And that really has changed the game. And I think now lake houses are the future.

15:45You even see the very traditional data warehouse companies embracing open formats and data lakes and, you know, the rise of iceberg, which I'm sure we'll talk about at some point, really playing an important role in that as well. So to play it back, you had the lake, which was like super scalable, but you couldn't dump anything in it. But you couldn't really run analytics because it's not super structured. On the other end of the spectrum, you had data warehouses, which are very structured, sort of designed for analytics, but maybe less flexible, arguably, although. And, yeah, almost exclusively all proprietary, right?

16:21You know, that was always the biggest complaint, you know, when I was at Teradata. It's an amazing database. I will still say that now, you know, seven years later, it's an amazing database. There are things that that system can do that really nobody else on the market can do. But it's proprietary and it's expensive. And I think that the evolution of big data as a term and as a concept, like just the volumes of data growing so much in each of our businesses now becoming so data-driven, necessitates scalable architectures everywhere you go. And data warehousing is certainly one of those. If all my data is trapped in a proprietary system, that really inhibits my ability to scale and grow and do the types of analyses I need.

17:04And so, hence, I think open source plays such an important role in data today. Yeah. So, data lakes on one end of the spectrum, cloud data warehouses, and then lake house is meant to be best of both worlds in the middle, combining the advantages of both. That's right. And so, obviously, the most famous cloud data warehouses just, again, put names in the categories. Yep. And maybe I'm too obsessed with this as like the guy who does like all the industry landscapes. But I like categories and logos in the categories. But in the cloud data warehouses, obviously, you have Snowflake, Redshift, Google BigQuery.

17:43In the world of DataLex, originally Hadoop. At some point, early Databricks. Yep. Dreamio. Dreamio and then Starburst. And then Starburst. But then that group evolved to the leghouse architecture. Yes, that's right. Or, yeah, I would argue, yeah. Yeah, we've been sort of focused on the lake house from the beginning. And Databricks has evolved into that. That's right. Because you always have that vision that storing all your data in one place is a terrible idea. Yes. And you are always the federated query engine and the core value proposition. My word's not yours. It's like just leave your data anywhere and we'll find it wherever it is for analytical purposes.

18:29Yeah. Yeah, I think where our view on that has evolved slightly is we do think data lakes are where you're going to want to store as much data as you can just because the economics will drive that. But you'll never store everything, to your point. And I used to use, actually, when I was raising venture for Starburst, I used to use your landscape all the time because especially when you show the progression from, you know, 2014 or whatever to today, 2025, you see the Cambrian explosion of data sources. And so being able to federate out to those has advantage for certain use cases for sure. Yeah, yeah.

19:05It has become the claim to fame of that landscape is to show the absurdity. That's right. It's not exactly what it was meant to do, but that's what it's become. Okay. And is it fair to say, I think you just alluded to it, that between like those various buckets, the lines have started blurring as well because people have a federation as well on top of being the repositories and different things. Is that fair? Lines have definitely blurred. Yeah. I think part of it is the industry is starting to mature. you have a couple giant players in Databricks and Snowflake that are now looking for adjacent markets to continue to grow their TAM and justify their valuations and drive their revenue into the future.

19:48And that's leading to more overlaps into boxes that they weren't typically part of. And so how do you position today versus those and, you know, I guess Dreamio and also the hyperscalers, right? And like, how do you, how do you, yeah, how do you position? Yeah, I think it necessitates being even crisper on the differentiation between those things. And so for us, there are a few things that we point to with customers. Number one, we're one of the only hybrid players. So Databricks and Snowflake are cloud only. So if you happen to have data on-prem, we're pretty much your only bet. And, you know, it just so happens that that turns out to be most of the Fortune 500, almost the entirety of the financial services sector in particular.

20:32And so we do a lot of business in those industries as a result. The second thing that helps differentiate us is the openness of the platform. So we're an open engine querying open formats. And while there's been widespread embrace, I would say, especially last year in 2024, around open formats and Iceberg in particular really winning that format war, that's new. That's new for this industry. We've been doing this forever, though, and that's, you know, the first queries run on Iceberg were Trino or Presto queries. So that pairing in the open source community of Trino and Iceberg has been kind of a reference architecture for years at this point.

21:10And that gives us an advantage because we can help manage your Iceberg deployment holistically. holistically, everything from streaming, you know, ingest, loading that data into iceberg tables, maintaining it, doing things like compaction, data maintenance, data management, and then, of course, querying. We have that sort of end-to-end lifecycle around iceberg. And we call that the ice house. So that's our play on lake house is the ice house. And then the other area where we differentiate, which you touched on already, was federation and being able to query into other data sources. And there are some use cases where being able to join a table that lives in another database system with the table that maybe sits in the lake can be very valuable in getting a fast response time.

21:56So, you know, we serve anti-money laundering use cases, fraud detection use cases, where very often the patterns to detect bad behavior actually exist in more than one data source. So you are very early to that vision of the lake house, now ice house. Was it, I don't know, controversial at any point, the, you know, iceberg plus Trino as a reference architecture? Or was it hard to convince people that it was going to be the future? I would say yes. Until last year, there was a lot of debate over which format is going to win. You know, Databricks had created one called Delta. There was another open format called Hootie.

22:32And then there was iceberg. And if you're just evaluating those for the first time, you think, well, there are three formats. How do I know which one's going to win? I think the reason we had conviction that it was going to be iceberg was simply that it was the one that had already been adopted by a lot of the super scaled up Internet companies. And I think there's something to be said for saying that, you know, you can watch those companies and kind of see where technology is probably going to go because they're the ones running at the most ridiculous scale. These technologies get really tested to the limit that way and can be a good indication of sort of where things are going.

23:08So we saw them all adopting Iceberg along with Trino and, you know, felt like that was likely going to be a pattern that is adopted by the industry as a whole. Okay. We'll go back to Iceberg in a second. But since we started talking about the reality of the Starburst product, so you have three offerings, I guess two core ones and one that's kind of newer. So you have the, as you were saying, on-prem version, which I think is Starburst Enterprise. That's right. Then you have the cloud product, which is Starburst Galaxy. So maybe walk us through those and then just to mention it up front, you have a specific offering with Dell, which is the third product.

23:55Maybe walk us through those different products and let's start with Enterprise, which is the on-prem version. How do you think of, you know, a classic question for, you know, open source businesses, How do you think of that enterprise offering versus the underlying Trino functionality, open source functionality? Yeah, so I would call enterprise a classic open core model, which is to say that the core is open, meaning there's a Trino engine inside where the leading contributors to Trino. You get that as part of the offering. And then around it, we've built all the enterprise functionality that you would need.

24:31So fine-grained access controls, row level, column level, data masking, query auditing. We've also built in a number of extra performance features. We have something called Warp Speed. We have a lot of fun with the names of our sort of subproducts. That is smart caching, smart indexing that delivers, you know, 10x performance boost over. Just faster SQL? Faster SQL. Exactly. Exactly. Some techniques there that give you faster SQL. That's exactly right. You know, we have extra connectors. We have, you know, management capabilities and so forth. And so that's what Starburst Enterprise is. Because it is self-managed, what we mean by that is the customer is managing it, and that allows them the flexibility to deploy it anywhere.

25:15It could be deployed in an air gap facility. It could be deployed in a vehicle if it needed to be. Like, you could run it literally anywhere. And so, again, that works for customers who have on-prem environments, hybrid environments, very complex environments, and gives them a lot of flexibility to integrate with the particular technologies that they have. Yeah, and, you know, classic question, again, for that kind of business model is, you know, the underlying community, because you're the major Trino contributors, but there's contributors to Trino in lots of different places. And presumably as people, as customers build stuff, they presumably contribute some of it to the community.

25:52So it's like this natural tension in any open source business between, you know, the open source that keeps getting better and the commercial products and perhaps a narrowing gap. So what is your experience? Every company is going to be different. But what is your experience with that tension, I guess? Yeah, I think you're absolutely right to describe it as sort of nuanced in the sense that you're always trying to make the right tradeoffs between contributing to the community, which generally drives adoption. So that is a benefit versus keeping something proprietary, which drives conversion. And that's the way I think about it.

Read the full transcript

26:31It's like adoption or conversion. And you need both. You need to grow the pie and then you need to make sure that you're getting a good piece of it. And so I think it's more art than science. But certainly, you know, things around performance and security are very logical places to kind of draw some lines where you know that enterprise customers value that. And, you know, that's something that can be monetized. Do you have any kind of, you know, in your license, anything that forces companies above a certain size to pay or not? We don't. And that's something that at least philosophically my co-founders have been pretty, I guess, consistent on is a desire to continue to use the Apache license, which gives widespread flexibility and freedom to users of the technology.

27:23And so we haven't created any, you know, more, I guess, restrictive licenses like some open source companies have. And so you were saying the on-prem version of Starbucks enterprise is particularly popular with large companies and financial services in particular. Yes. It's like another super interesting topic of, you know, cloud versus not cloud and who will resist ever going to the cloud. So your experience is that those companies are not going to the cloud. Yeah. I mean, they're pretty emphatic about it. I had dinner with the CEO of one of the largest banks in the country. I won't name his name just to protect his innocence, I guess.

28:07But, you know, he told me point blank, like, we're never going to move everything to the cloud. We're going to have four clouds, in his words, and the fourth one being his own, you know, on-prem, essentially. And so I do think that's the future. I think also there's, you know, an interesting case to be made that there may be some repatriation with the AI emergence. And, you know, our partner Dell is very much betting on that as well. They're selling a ton of servers right now, AI servers, which are basically, you know, have a lot of GPUs in them, a lot of NVIDIA products in there. and are seeing customers that are trying to get economies of scale by deploying that infrastructure on-prem and running it themselves.

28:49So, yeah, I think this dichotomy is going to exist for as far as I can see. Yeah, it's been fascinating, actually, to see the Dell stock price. Everybody's talked about NVIDIA, you know, for all the right reasons, but it's amazing how Dell, as the quintessential on-prem player, has seen their fortunes accelerate as well last year. Okay, great. So that's Starburst Enterprise and Starburst Galaxy, the managed version. How does that work? Yeah, so that's a classic, you know, SaaS product that's hosted and managed by us. It is connecting to your storage. So it's your own S3 buckets, your own, you know, RDS, your own MySQL database.

29:32But the compute and the control plane is managed by us. And so we're able to offer a very seamless, easy to use, turnkey, push button type of approach while still giving you all the performance and functionality that you need and that you're looking for. And that product has actually evolved very, very quickly for us. We've been able to get a lot of new interesting features and functionality there. And one of the use cases where we've seen a lot of adoption and interest is actually customers who are building their own data applications and using Galaxy as the embedded engine. where, you know, there's a data analytics portion of the SaaS app that they provide to their customers.

30:14And the analytics are essentially powered by our engine. And, you know, we were digging into, like, you know, why are they choosing us for this? And I think it goes down to, you know, if you're going to be part of someone else's margin, essentially, their cogs, right, you need good TCO. I think what we're seeing is data application developers are choosing Iceberg for storage because that gives them a lot of flexibility. They want something that's based on open source so that they always have that optionality as well. And then, you know, we can deliver them pound for pound the best cost performance, we believe, in the industry.

30:50And so those things all really matter when you're, you know, ultimately fitting into their margin. Yeah, very interesting. And that's in contrast to like this whole series of companies that specialize in embedded analytics and BI. So the. Yes. Size sense and good data. Yes. Well, I'll clarify. We're not doing the visualization piece. Yeah. So some of those companies. But you help them do the analytics and then embedded math. Exactly. Yeah. We're doing the query execution behind the scenes. We're the engine for that. Exactly. Okay. All right. So that's Galaxy. And then we alluded to the Dell thing.

31:26That's more recent, right? Yes, that was almost 12 months ago that we released what's called the Dell Lakehouse. So it's their product powered by Starburst. And that's essentially, you know, Starburst inside with Dell's object storage, Dell's hardware. And it's become a real centerpiece of their AI strategy in selling this into large customers. Some of them themselves are CSPs that are building out their own infrastructure for analytics and AI. And I think at the end of the day, AI is only as good as the data that you're training the models on, only as good as the data that you're accessing through RAG workflows.

32:06And so having a lake house at the core of that architecture is important. Okay, great. So those are the three offerings. And still on the topic of product, maybe walk us through the overall architecture. And you have a chart somewhere on the website that maybe we'll put in the video where, you know, you have the classic sources on the left and then, you know, Starburst in the middle and magic on the right side. So walk us through that. So in terms of sources, how do the data go into or how does Starburst access the data? What kind of data? Sure. Yeah. So at the heart of Starburst is this notion of connectors.

32:50And we think of everything as a connector. So even if you're just accessing S3 and you're going to be querying iceberg tables, that's technically our S3 connector or data lake connector to access that. And so every connector is basically just connecting to the underlying catalog of the system that you're connecting to. And then as soon as you've connected, which is like a one-time, you know, setup thing, now you can run queries. And you can do that at the command line, like just start to write SQL queries, joining tables across different systems. Or you can use a BI tool like Tableau or ThoughtSpot or others.

33:27Or, and then this goes to the sort of data application side, we're seeing customers, you know, build more programmatic ways to interact with the data that we have access to. One of the features that we've built that's really nice, especially for internal purposes, is something called data products, which is basically allowing you to stitch together a view of your data across these different data sources. And that's where you really start to get some interesting optionality because you can decide to materialize that view or not materialize that view. And there are trade-offs to both. You know, if you're querying the data source directly, you're going to get very fresh access to the data directly where it lives, but you might trade some performance.

34:07And if you're really optimizing for performance to power some kind of dashboard, well, that might be a case where you want to materialize that view in the lake. And so we have a lot of flexibility on allowing you to do either one. The data can be batch or real-time? Yes. Yep. So meaning that you just have a Kafka endpoint in gesture? Yes. Is that a word? Ingestion. Yes. Of some sort. Yeah. Yeah. And actually, that's something we've really optimized recently is we call it streaming ingest. And that's connecting to a Kafka stream and landing it in iceberg tables automatically for you. And we're able to do that at incredible throughput.

34:48And that's an architecture I would say that we're seeing a lot of out there is Kafka as opposed to more traditional batch-oriented ETL. And so you can land your data, have it there, you know, it might be 15 seconds old, you know, from when it was created. And that enables, again, more real-time analytics as a result. The data itself can be structured versus unstructured. Does it matter? Yeah, typically it's structured or semi-structured for our SQL analytics offerings. But we're also, you know, doing some interesting things around vector search. We can connect to different vector databases. There's some things that we're working on there.

35:30But yes, today I would say primarily SQL analytics on structured or semi-structured data. Semi-structured would be something like JSON files or XML files. Yes. All right. So that's sort of the, for the most part, the left side of the magic chart. So the middle, so there's a SQL query engine, which is what we talked about, augmented by warp speed. That's what it's called. And then there's a governance component as well. Exactly. Yes. And that's very important for our large enterprise customers. You know, the access control piece in particular, having good auditing controls of being able to see who accessed what, who saw what, and being able to do both what's called RBAC, role-based access control, but also ABAC, attribute-based access control.

36:18And that's important because that allows you to tag certain data sets as being, you know, PII data. Maybe this one has certain data sovereignty constraints. You know, only people in Switzerland can see this data, for example. So it gives you a lot of flexibility. Is there other bits in that middle that I forget? You know, I'm quizzing you on one chart on your website. Yeah. Because as a CEO, of course, you know all the charts on your website. I've probably seen them all at some point. Yeah, I mean, I think the other is just obviously if you're using Galaxy, then we have a lot of automation of the cluster management itself.

36:54You know, you can scale it up, scale it down. You can have an auto scale on its own. and all those things are just going to save you money if you don't have a cluster running all the time. So I'm totally cheating because I'm bringing up the chart in front of me, but I'm not showing it to you. So you have, joke aside, you have, so ingestion we talked about, you have table maintenance. So governance we talked about, accelerated SQL analytics we talked about. The two things we didn't talk about is table maintenance and automatic capacity management. Do you want to go? Okay, yeah. Yeah, so the table maintenance piece, That's, you know, there's a lot of activities that you want to do to optimize for performance reasons how the tables are structured and laid out.

37:35If you've been around for a while, you know, this is kind of like defragging a disk drive from like a long time ago, right? Like you want to take a lot of small files and bring them together into larger files. And that's called compaction. And that's a particular thing that you want to do with iceberg tables to get better performance. And you could do that yourself. It can be tedious, time-consuming, error-prone. There's a lot of work involved. Or you can use Galaxy and we do it all for you, basically. So that's sort of the table maintenance side of things. And then on the capacity management, that's really those autoscaling capabilities where you can spin clusters up and down and have them automatically scale based on incoming compute.

38:13It's a more serverless type of experience behind the scenes. The data apps you just described, is that – so that's a materialized view that you were describing? Well, that's the data products piece. That's the data products. All right, great. So there's data products and there's data apps. So what are the data apps? So data apps is really what I would say is an emerging use case where customers are building their own applications, leveraging our platform, and doing that because, again, we become part of their COGS, part of their margin. So they want a platform that's going to be very cost-effective to them.

38:47And so we're powering the analytics behind some of the large SaaS vendors out there. As I'm reviewing my notes, I joked down things like AML, Anti-Money Laundering, Customer 360. Yep. So those are examples. So those are apps that are delivered by SaaS companies just to play it back, but powered by Starburst? Well, they can either be apps built internally. So like an AML use case, anti-money laundering use case might be something that a bank has built themselves to meet their own regulatory requirements. But, you know, for example, one of our customers is called Vectra, Vectra.ai. And they're a – Because that's Amr's company, right?

39:32Yes. Yes. And they're a cybersecurity-oriented company, and we're powering the analytics behind Vectra. So that's an example. Yep. Okay. But so for the internal use case, like I'm a bank and I want an AML and I'm an online drink. So what walk me between the difference between a data app, like a Starburst data app versus just a use case? So you formalize, you have like a building block, like a Lego block that helps them build the AML app? So not necessarily for AML specifically. I would say like, you know, this is a horizontal platform, but because of, you know, our REST API, they can start building their application directly against that.

40:20And so now access to all of the data that they need to, you know, either show or power dashboards or analytics within the app itself or deliver reports or whatever analysis is sort of taking place, we can be the engine behind that. And we've tried to make it as easy as possible for those developers to build against that. Okay. Very cool. So that's a lot of things you guys have built because you also have, you know, we were talking about connectors. Like I wrote down, you have like 50 plus data sources. That's a lot of building. What turned out to be the hardest to build in the history of Starburst?

41:02Ah, the hardest to build. You know, I would say building Galaxy, the SaaS platform, you know, there's a story there where we actually built it twice. And, you know, customers never really saw the first version because at the last minute we said, you know what, we can't deliver the seamless, easy to use, consistent experience that we're going for. Oh, wow. So you had to go back to the drawing board? and went with a totally different approach. And actually what's interesting is that first architecture was actually fairly similar to Databricks' architecture where we did not control the compute plane.

41:41We were just this sort of control plane and we were running everything in the customer's environment. And we saw potential benefits to that, but the biggest drawback is you don't have as much control over the experience. And, you know, we wanted to take a more, I'll call it like a more of an Apple approach, you know, to really delivering, you know, customer joy. And so we re-architected that. And, you know, that basically took another year, you know. So it was a big, expensive, you know, re-platforming. But ultimately, we're really happy with the way Galaxy turned out. That must have been a – was it an obvious decision or was it like a super nerve-wracking decision?

42:15It was a nerve-wracking decision. It was obvious that we had prototyped this different approach that became Galaxy as it is today. and it was obvious that that was going to be better. But even still, it was a nerve-wracking approach because you're sort of throwing away like two years of development and starting over again. And, you know, of course, we're venture-backed and we spent a lot of money to build that. And that was actually one of the reasons we raised venture in the first place. You know, we have a little unusual history that we were bootstrapped the first two years and we were running a nice little profitable business.

42:47It was great. But we thought, you know what, we can't build a SaaS solution. We can't build a cloud platform off our little bootstrapped, you know, small business. It's just, you know, too expensive. And so we raised venture, built this cloud platform, and then, you know, threw it out and built it again. So that was definitely stressful. Good times. I'm sure that must have been made for good board meetings. Yes, for sure. We're lucky we have very patient investors. Yes. And the flip side to having a lot of different products in addition to the effort and time to build them is that It's a lot of surface to manage, a lot of like product management that's involved.

43:26How does that work practically having, you know, I guess now three different products that need to be sort of synced in terms of functionality? Yeah, definitely there's complexity to it. I think the way that we try to simplify it is find as many common components between that. You know, if you think of like a car manufacturer, you know, they usually, the big ones, the VWs, you know, of the world, you know, they'll build on one platform and that platform will then get used by, you know, five different car models or maybe 10, you know, where they've built the chassis, you know, once. And they'll use the engine, you know, the same engine in, you know, five different cars.

44:03So we've basically taken that kind of approach where there are common elements that are used across all three of those. But even still, there are differences among those different products. And that requires a lot of, you know, intentionality, puts some added pressure on our product managers to really make sure that they have their ICP really nailed down, their ideal customer profile that that particular product caters to, you know. And we touched on some of those things. You know, if you're a big multinational bank with an on-prem footprint, enterprise is probably your best bet. If you're more digital native, then you probably want to use Galaxy, you know.

44:39And so each one has its own unique components. And how does that work from a team perspective? Do you have different teams working on the different versions? We do, yeah. Yep, different engineering teams, different PMs. And again, there's common elements that are shared, but yep. And then so they all come to you and you have to decide on the priority and who gets more money to go faster? Yes, we're actually doing that right now for, you know, our next fiscal year. Very cool. And, you know, it's really, you know, one of the very interesting parts of this conversation is precisely that kind of on-prem versus cloud because, you know, the dominant narrative has been in open source has been that, you know, people who did that kind of like hybrid approach kind of ended up not being, not really liking it.

45:37and then recommending to anyone that would listen that people should do cloud only. But you've embraced the complexity and you do both. Yes. For better or worse, that's right. That's exactly right. I mean, you know, we started the company with this idea of optionality. It was probably one of the words that I used, actually, the data-driven event that you hosted, geez, I don't know, five or six years ago, right? Yeah, 2019. Yes. That's the second one we did. The one we did for Starburst, yes. Yes, yes. You remember words you used in 2019? I do because it was a word that I used a lot, and I still do.

46:14You know, optionality is certainly a word that you may hear in finance context, but not necessarily used that much in technology architectures. And, you know, we sort of founded the company on this idea, which is to really give customers, you know, the power of optionality, the flexibility to build a future-proof architecture that you can change out components, you can work on-prem, you can connect to different data sources. And I think part of that was also our bootstrapped orientation where we were very customer obsessed because we depended on that revenue to like pay the next month's payroll.

46:48And so optionality became sort of a core ethos. Now, the consequence of that is complexity for us as a vendor. So you're right to touch on that. But it's one of the reasons customers choose us. Okay. So we talked about the different offerings. We talked about sort of the architecture under the hood, maybe to close on architecture. So we did talk about Iceberg. You support the other two as well? We do. Optionality. Optionality. There you go. So Hoody and Delta. And Delta, yep. But you see most of your customer use Iceberg? Yes. I feel like the summer of 2024, the world said, okay, it is Iceberg.

47:33And that was like VHS, Betamax, you know, decision made. What triggered it? Was that the acquisition of Tabular by Databricks? I think it was two things. I think that was a big, big piece of it. I think the other thing that happened just maybe two weeks before that was that Snowflake also decided to support Iceberg. Yeah. And of course, you know, maybe some of your viewers are aware there was a bidding war for Tabular, which is what led to such an incredible outcome for them. Yes. And to put this so. Actually, I don't know what the numbers were, but I heard that Snowflake was bidding like 300 and then 600.

48:11And then what was the final price for? It was like 2 billion, right? Databricks. So Databricks won over Snowflake with a$2 billion deal. Yep. And now... Which is fascinating, which is probably partly why Ali wants to... One of the reasons why Ali wants to keep the company private for a bit longer is because he can do that kind of thing, which is a very bold move and exciting move. Yes. But they can do that without the scrutiny of public markets. A hundred percent. Yes, I totally agree. And what do you make of the acquisition from a strategic standpoint while we're on the topic? Because so Tabular, for anybody that may or may not follow the space very quickly, was a young Series B company that was the commercial company on top of Iceberg.

48:59and, but ultimately Tabular is an open source project. Yep. Well, Iceberg, yes, is. What did I say? I'm sorry. Iceberg, yes. Iceberg is an open source project. And it's sort of unclear what you buy. If your ultimate strategic goal is to take control over an open source project, sort of unclear what you actually buy by acquiring the company that's a commercial company on top of the open source project. I mean, on there, you know, this, you know, certainly better than me. But aren't there like a bunch of like iceberg contributors like AWS? Yeah. Starburst as well. Starburst and Microsoft, right.

49:39So. Snowflake? Yeah. I mean. So what do you get? Then you spend$2 billion, you buy a company. Yeah. But like to which extent, in your opinion, does that help you take control over that project? Yeah. I think this is a really interesting question. I don't think it does allow them to take over the project, to be perfectly honest. And I think that the market is actually resolutely determined to ensure that it continues to be independent, which is important actually for Iceberg. And that's what made Iceberg popular in the first place over Delta, you know, which was Databricks' own format. The market wants an independent standard.

50:17That's what they want, independent of any vendor. And so fortunately, there's enough groundswell of people like ourselves, like Snowflake, like some of the others you mentioned, where it is – and it's also, by the way, an Apache Software Foundation governed project. So you have the Apache Software Foundation also ensuring independent governance, which is important and makes it truly, I think, independent. I mean, Ryan works for Databricks now, and he will for some period of time. And Ryan being Ryan Blue, the CEO of Tabular. Exactly. But Iceberg is bigger than Databricks, honestly. And it will continue to be, I think, the ubiquitous format.

50:58You know, so if they had hopes of like killing it, I don't think that worked. If they had hopes of controlling it, I don't really think that will work either. I think probably the biggest thing that they get out of the acquisition is the ability to tell the market that they can do iceberg two and not look like they had made a mistake with Delta. You know, that they can they get sort of a marketing win out of it. But as a practical enduring matter, I don't think that they get extra influence beyond that. I also don't think like, I think Ryan cares too much about the future of Iceberg, like even to like, even if Ali was like, do this thing.

51:36I don't know that Ryan would do it, you know? Yeah. And why do you think Snowflake embraced Iceberg before that? Because that's fundamentally against their interest, right? Snowflake is all about put all your data in our proprietary database and leave it here forever. Yeah. So was it just pressure? Customer pressure. I think that's exactly right. I think that the market is basically saying we want to store our data in open formats. That's better for us. And by the way, this was sort of one of my theses, if you will, for starting Starburst is I believe gradually over time, enterprise software technology, at least infrastructure technology, wherever there are two things that are mostly the same and one is open and one is not, the open is going to win over time just because economics will eventually drive long-term decision-making.

52:26And, you know, that's what I saw with Cloudera. Again, Teradata was the better system at the time, but Cloudera was the open, open-source system. And that got a lot of adoption. Oh, that's right, right. Because so Teradata as in HADAPT at Teradata because, yeah, because HADAPT was basically a competitor in Pala, right? That's exactly right. Yeah. So we're going back memory lane, but... That's right. Yes. That's right. But Impala was open source. It was. And Hedapp slash shared data was proprietary. That's a lesson learned there. Exactly. I am the product of many lessons learned. And that's making me feel old.

53:04But that's, no, that's exactly right. Yeah. Which is a good thing in this space of enterprise software. Like having learned the lessons and try different things is very much a strength. Okay. Another thing that seems to be a part of the overall Starburst story and positioning is the data mesh. I would say, I would ask, what's up with the data mesh? That seems to be, you know, a very powerful idea. And we had Jarmak on, you know, either a data-driven or on earlier version of this podcast. But that was, you know, like all the rage sort of like a year or two ago. And I guess what's the current state of it?

53:50Is that as vibrant and exciting as ever? What's your take? No, I think the phrase or the concept has lost a bit of momentum. I think that there are companies who are deploying data meshes for sure, but some of the attention has certainly waned. I think the lasting legacy of that, though, is this concept of data products, creating these sort of curated data sets from data that can live in multiple places and thinking about them from a product perspective with a product mindset, which is to say that there's a clear owner of the data. And so I think like any, you know, major sort of wave or, you know, phenomenon, there's a lasting impact of that.

54:36Yeah. I mean, Jamak has her own company now. She's been very focused on, she probably hasn't been quite as publicly vocal on the data mesh concept. And I think what we've seen with customers is there's a people and process element that's actually harder than the technology part in implementing a data mesh. So we're very capable of helping customers implement data meshes from an architecture perspective. And we do do that. But I think, you know, it does require some retraining of sort of how you do things and how you organize your people internally. Yeah, yeah. And I think we're saying the same thing, but I think the term itself is perhaps a little like past the initial excitement, but the fundamental concept remains as something powerful.

55:24Yes, yes, exactly. Of course, the topic that we need to talk about is AI. Ah, yes. I'm hearing good things about AI. Apparently it's going to be big. It's going to be big. Yes. To which extent is that part of your positioning? We've talked a lot about SQL search and structured or semi-structured data. Yep. Equally, you were saying that for Dell, you were part of the AI story. So, like, where do you fit in the sort of AI stack world? Yeah. So, two places, I would say. First and foremost, on the training of models for those who are actually building their own models. Your models are only as good as the data that you train them on, and, you know, access to more data or better data is going to influence that.

56:10And so we become this sort of access layer to all the data in our organization. There is no piece of data that we cannot get to, essentially, through our federated architecture and being able to work across on-prem and cloud. And so that's one element is, you know, you're going to access data, you're going to do some transformation of data, prepare data, train a model with that data. The other is the more practical aspects of putting AI into production, which, you know, we see RAG workflows as an essential part of that. You know, people are building these agents that need to access contextual information, pass that along to the LLM to get the appropriate response as part of the agent or part of the application that they're building.

56:53And we think we can play a central role in that, both in terms of, you know, the access to structured data, but also as performing vector search. And so we expect to be doing more, and you'll hear a lot more about us, I think, this year around RAG workflows and how Starburst plays a unique role in that. A few words on the go-to markets, you know, lessons learned selling enterprise software. What has worked? What has not worked? How do you currently sell? Are you mostly outbound, sales driven, presumably given the type of customers? Yeah. Anything you can talk to? Sure. Yeah. Yeah, we're mostly direct.

57:36And then, of course, we have an important partnership with Dell that we touched on as a channel for us. You know, things we've learned. I think the biggest thing that we learned was that we are, and this is going to sound silly now in retrospect, but it was an important lesson. We're hardcore enterprise software. You know, we're not a PLG motion. And I say that that feels silly in retrospect because, like, it's so clear to us now today. But there was a period where we thought we wanted to be PLG, you know. Well, I think just about any company in 2021 thought they were PLG, right? Yeah, for sure.

58:11For sure. And so, you know, we hired sales leaders who had PLG experience and we thought that would make us PLG. But the reality is like PLG starts with the P, which is the product itself. And we're pretty hardcore, complex, you know, data infrastructure stuff. And our solution architects are some of the superheroes of our go-to-market. Like they're the guys and gals who know how to, you know, deploy this and work within the specific unique data ecosystem that a particular customer has. And each one is different and unique. And that's part of what makes it hard to really create a true PLG motion around this is that, you know, every customer is a little bit different in terms of what they're trying to do and how they want to deploy it.

58:55And do you charge services? Is like service is a component? We do. Yes, yes. And we do that both ourselves. We have our own professional services organization, and then we've trained up some amazing partners. We work with one that's pretty prevalent here in New York City called Kubrick, which does amazing stuff. And we work with a number of SIs outside of that, too. But, yeah, services are definitely helpful, especially for large, complex enterprises that are deploying this. So, actually, a very interesting topic, services for, you know, any founder listening to this. How did you go about building that network, so Kubrick or others?

59:35When did you feel it was the right time to start working with SIs? That's a good question. I would say, you know, you should probably learn how to do it yourself first. So I would say not at the very, very beginning. But then my advice would be try to start training up one or two and probably boutique firms like don't don't go after Accenture on day one, but start with a smaller boutique firm. Because what's most important, especially in the early parts of developing the services piece, is the quality of the delivery. And you want to control the quality of that as much as possible. Obviously, you control it when you do it yourself, but that's difficult to scale.

1:00:17And it's not the margins that your investors probably care about. They don't want you to build a services company. So that's probably not going to be the dominant way you want to deliver ultimately. But by choosing just one or two and really focusing on them, there's a mutual benefit there. A, you're helping to bring them business and you're creating that positive reinforcement cycle that learning how to deploy your software is going to lead to more revenue for them. But also you're creating a reliable partner that you can count on, that you know when you recommend that firm to deliver services at customer X, you know they're not gonna make you look bad, you know, and that's really important.

1:00:57So I would say, you know, rather than going far and wide and signing up hundreds and thousands of SIs, which we did that too along the way, I would say go deep with one or two, probably on the smaller end, just because you'll be able to get more attention span and try to make them as good as delivering as you are. And then you can build on that relationship. And that can be very symbiotic. And then the Dell partnership, which is more of a, you know, ISV kind of, which is the other sort of side of the partnership world. To the extent you can talk about it, how did that come about? Because, you know, a lot of people want to do that.

1:01:31A lot of startups want to do that. And that actually happens very rarely. So what was the history and how long did he take and any lessons learned there? They actually came to us. And I think that's a really important element. Hard to recreate, of course, for any entrepreneurs who are listening to this. But the reason that's important is whenever a partner is coming to you, you know that there's some motivation behind it. Maybe that you don't even know initially on why that's so important. And that's going to just drive momentum and focus and interest because the biggest challenge working with somebody so much larger than you is how do they get focus on it, right?

1:02:05I mean, they have tens of thousands of employees. How do we know this is going to be important to them? Well, because, you know, a VP within Dell was like, this is important. We need to close this gap. Michael himself thinks we need to close this gap. And a VP of product. A VP of product. That's right. Not a VP of partnerships. That's actually a good point. Yeah, I think of the partnerships organization, at least in this context, as the facilitators of the relationship. But the real impetus, the real motivation has to, I think, come from product. It could come from sales. That's also not bad. You know, in some cases, it's a revenue leader who's saying, hey, we need partnership with this company because we know it helps us get into these customers or close these deals.

1:02:45I think that's great motivation as well. But I think you do need that stakeholder who's willing to stand up and sort of sponsor the initiative to get the right resources around it. So in our case, it was on the product side. They had come to us. They looked at all the players in the market. They did some bake-offs and decided, hey, we're the best technology for what they were trying to do. And that led to this OEM agreement. We started as a reseller. So we did about a year as a reseller agreement, meaning that they could resell our software. And then that emerged or evolved into a true OEM where we're now embedded in their product offering.

1:03:21It is a Dell product. You know, for all their sellers, it's a first-party product. They're getting comped in full. You know, it is treated as a Dell product in every way. So how long was the whole cycle from first call to where you are now, like a couple of years? Yeah, I would say probably, yeah, we're probably now starting year three. And I would say it was two years to get it really out the door as a skew that they could sell. So it is also long. Yeah. Yeah. So people should, yeah, should view those as very long term projects. And who drove that on your end? Like, how did you make it successful on the Starburst end?

1:04:01So I think that idea of focus needs to be present on both sides to make it work. And so in my case, I have a SVP of alliances. His name is Tony, who sits on my executive team. He works directly for me, reports to me, who shepherded that deal on our end. And that meant that it was always visible to me. You know, it was something we talked about in every one of our one-on-ones. How's it going? Where are we? What are the negotiation points? What are we working through? And so I think it really needs to be that kind of priority on both sides to make it work. Okay. Maybe to close, what's the next year going to look like for you?

1:04:40Or I guess we're at the beginning of 2025. So what's the next year going to be for you and then for the industry? So that's the prediction, 2025 prediction part of the question. Yeah, question. The thing I'm most excited about is really seeing AI start to get implemented in real production use cases. I think, you know, the past year or year and a half has been a lot of experimentation, a lot of prototyping. And I think we're now just starting to see AI actually matriculate into real production use cases. And that's going to get us over the hype cycle. Because I think even in my own organization, when we started to get into AI, you know, some people were like, is this just a fad?

1:05:22Is this, you know, like – Is this like crypto? Yeah, yeah, exactly. No, like we're going to quickly get into, I think, a real like, you know, plateau of productivity or whatever they call it in the hype curve here where we're going to actually see material results, which is only going to just reinforce, I think, the demand for it. And I think, you know, RAG in particular will be a really important element of making AI production ready. Okay. So that's both what you have on the docket that starburst and your prediction for the industry. Yes, exactly. Very cool. All right. Well, Justin, thank you so much.

1:05:59That's been wonderful. Really interesting chat. Appreciate it. Thank you for having me. Always a pleasure. Hi, it's Matt Turk again. Thanks for listening to this episode of the MAD podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.

From the publisher

In this episode, we explore the cutting-edge world of data infrastructure with Justin Borgman, CEO of Starburst — a company transforming data analytics through its open-source project, Trino, and empowering industry giants like Netflix, Airbnb, and LinkedIn.


Justin takes us through Starburst’s journey from a Yale University spin-out to a leading force in data innovation, discussing the shift from data lakes to lakehouses, the rise of open formats like Iceberg as the future of data storage, and the role of AI in modern data applications. We also dive into how Starburst is staying ahead by balancing on-prem and cloud offerings while emphasizing the value of optionality in a rapidly evolving, data-driven landscape.


Starburst Data

Website - https://www.starburst.io

X/Twitter - https://x.com/starburstdata


Justin Borgman

LinkedIn - https://www.linkedin.com/in/justinborgman

X/Twitter - https://x.com/justinborgman


FIRSTMARK

Website - https://firstmark.com

X/Twitter - https://twitter.com/FirstMarkCap


Matt Turck (Managing Director)

LinkedIn - https://www.linkedin.com/in/turck/

X/Twitter - https://twitter.com/mattturck


(00:00) Intro

(01:32) What is Starburst?

(02:32) Understanding the data layer

(05:06) Justin Borgman’s story before Starburst

(10:41) The evolution of Presto into Trino

(13:20) Lakehouse vs. data lake vs. data warehouse

(22:06) Why Starburst backed the lakehouse from the start

(23:20) Starburst Enterprise

(27:31) Cloud vs. on-prem

(29:10) Starburst Galaxy

(31:23) Dell Data Lakehouse

(32:13) Starburst’s data architecture explained

(38:30) The rise of data apps

(38:54) Starburst AML

(40:41) “We actually built the Galaxy twice”

(43:13) Managing multiple products at scale

(45:14) “We founded the company on the idea of optionality”

(47:20) Iceberg

(48:01) How open-source acquisitions work

(51:39) Why Snowflake embraced Iceberg

(53:15) Data mesh

(55:31) AI at Starburst

(57:16) Key takeaways from go-to-market strategies

(01:01:18) Lessons from the Dell partnership

(01:04:40) Predictions for 2025

More from The MAD Podcast with Matt Turck

All 44 episodes
Trino, Iceberg and the Battle for the LakehouseThe MAD Podcast with Matt Turck · 1 h 6 min
Listen in VO