914: Data Lakes 101 (and Why They’re Key for AI Models), with Oz Katz

15 Aug 2025 · 26 min · 15 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Explains data lakes vs data warehouses for AI, why multimodal AI makes data management messy, and how LakeFS uses Git-like versioning to keep data consistent and reproducible across training and deployment.

Guest backgrounds

Oz Katz is co-founder and CTO of LakeFS, a company providing data storage for AI applications. He has a systems engineering background.

Key claims

Data lakes are like shared folders where unstructured/messy data from multiple sources can be combined without rigid upfront schemas. AI pipelines create multiple “sources of truth” (raw files, embeddings, features, labels) that drift out of sync. LakeFS provides a unified workflow (branch/PR/merge) over data to prevent teams from stepping on each other and to enable reproducibility via versioned snapshots.

Notable examples

Image pipelines (raw images → embeddings in a vector DB → extracted features/labels in other systems); automated checks like verifying images meet quality thresholds before committing/merging. Mentions Apache Iceberg tables on object stores (e.g., S3) and vector database support.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Guest Introduction and Travel Chat

0:39 to 1:46

Oz Katz shares his travel plans and experiences with time zones.

“Oz Katz, welcome to the Super Data Science Podcast.”

Transition to Data Discussion

1:46 to 3:08

Hosts shift focus from travel to discussing data lakes and AI.

“14 hours or something like that, but I land before I take off.”

Understanding LakeFS

3:08 to 4:32

Oz explains what LakeFS stands for and its naming background.

“And so specifically, I mean, we will eventually get to LakeFS, your company that you co-founded and were your CTO.”

Defining Data Lakes

4:32 to 5:21

Oz provides a clear definition of what a data lake is.

“It's distinctive looking because our audio-only listeners won't be able to see this right now, but it's lowercase lake, L-A-K-E, and then FS in all caps without a space.”

Data Lakes vs. Data Warehouses

5:21 to 6:45

Oz contrasts data lakes with traditional data warehouses.

“These are often misused or just used differently by different companies or vendors or people with some interest in the field.”

Data Needs in AI Era

6:45 to 8:05

Discussion on the evolving data needs for AI applications.

“is that it, I think it maybe is even implied by what you just said, but it's that the data aren't necessarily very structured.”

Challenges of Multimodal Data

8:05 to 10:00

Oz discusses the complexities of managing multimodal data for AI.

“Oh, that's a good one because I think this is, the answer I'll give you two years ago is probably very different from the answer I'll give you now.”

LakeFS Solutions for Data Management

10:00 to 12:00

Oz outlines how LakeFS addresses challenges in data management.

“Am I correct that this is a tricky situation and how are we resolving it?”

LakeFS Functionality Explained

12:00 to 14:00

Overview of LakeFS functionalities and user interface.

“Yeah, so one of the reasons why we built LakeFS and started working on this problem to begin with was to try to kind of allow humans to better interact with that mess, right?”

Understanding LakeFS Features

14:00 to 15:10

Learn about the user-friendly features of LakeFS and its automated data management capabilities.

“You don't have to type in the commands if you don't want to use a CLI, although one exists for you.”
Show all 15 chapters

Future of Data Storage for AI

15:10 to 16:30

Discover the trends in data storage and how object stores are becoming central to AI models.

“It sounds like I could make use of that for sure.”

Introduction to Apache Iceberg

16:30 to 18:40

Get insights into Apache Iceberg and its role in creating table structures over object storage.

“So that's definitely something we're seeing and also something we're doubling down on.”

Using LakeFS in Organizations

18:40 to 21:31

Explore the user experience of implementing LakeFS in an organization and its benefits.

“Think about the structured data story, right?”

Book Recommendation on Data Systems

21:31 to 22:20

Learn about 'Designing Data Intensive Systems' and its relevance for data scientists.

“And now we just let you manage it that way.”

Connecting with Oz Katz

22:20 to 24:21

Find out how to connect with Oz Katz and join the LakeFS community.

“I always ask my guests for a book recommendation.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:Welcome to episode number 914 of the Super Data Science Podcast. I'm your host, Jon Krohn. In today's episode, we've got Oz Katz, who is co-founder and CTO of the company LakeFS. LakeFS provides data storage for AI applications. And so we talk about key data related things for AI, like what's the difference between a data warehouse and a data lake? And what are the kinds of things that a modern database needs to be able to do in order to support the data warehouse? support, training, and deployment of AI applications. It's an invaluable episode. I think you're going to enjoy it.

0:39Jon Krohn:Oz Katz, welcome to the Super Data Science Podcast. It's great to have you here. Where are you calling in from today? So I'm in New York. And thank you, John, for having me, by the way. As I'm usually in New York, that is the most likely place that you would find me. But today I'm in San Francisco. And I suspect that you make your way out to the Bay Area from time to time as well, Oz. I sometimes do, yeah. Are you there for a conference? Anything happening that's interesting? I'm actually just on a stopover on my way to Brisbane, Australia. Oh, wow. Okay, that's a journey. Yeah, so that's coming up.

1:17Jon Krohn:After we finish recording, I make my way to the airport. And yeah, I'm going to have a 14-hour flight. And yeah, I'm going to lose a day. It's kind of interesting. I fly on a Wednesday night and then land, it seems like much more than 14 hours later, because you land on Friday. But it's even trippier, I think, when I come back. I gain a day, so I'm in the air. Yeah, you get that day back somehow, right? Yeah, exactly. So I think on the way back, I'm in the air for, again, 14 hours or something like that, but I land before I take off. Somehow I always feel like I'm not smart enough to understand time zones fully.

1:59Jon Krohn:it just it doesn't make sense to me no i know exactly sometimes i think when i'm wrestling with it that i'll like i should get a globe out like you know kind of the globes that we i don't know i had growing up as a child this is like the full-size globe and just kind of spin it and kind of spend time thinking about how it's moving because it's probably possible with practice to figure it out but yeah it's definitely confusing so the trippiest thing about it that i've found so far like there's this thing called the international date line and you would expect it to be a straight line going down the ocean it's definitely not a straight line some countries are like some islands want to be on that side some want to be on that other side so there might be islands where it's like the next day but they're actually further down into the other the other direction than other places they can just travel across days one mile apart from each other yeah Yeah, that is trippy.

2:52Jon Krohn:I guess some countries might want to be kind of more aligned with Asia, I guess, and some more with the Americas, I guess is the idea. Interesting. I hadn't thought about that. I'll have to look into it. But anyway, we are not here for a lesson on time zones today. We're here to talk about data for AI. And so specifically, I mean, we will eventually get to LakeFS, your company that you co-founded and were your CTO. but uh we're and so we will get there at some point uh in terms of talking about it technically and how it fits in as a solution and and why it's so badly needed but first let's just talk about the name what is lake fs what does that stand for all right so i'll preface and i'll say that as the old saying goes there are two things that are hard in computer science one is cache and validation.

3:45The other one is naming things. I hope we got cache and validation right, at least that. So LakeFS stands for lake as in data lake, which is typically the architecture where you would put LakeFS into. FS stands for file system. I come from a systems engineering background. So for me, it makes perfect sense. I don't know that most of our target demographic or the people that get value from Elikifest necessarily understand that.

4:18Jon Krohn:So there's that. So it was named maybe by the tech team, not by the marketing team. Yeah, that's what happened. We liked it enough. We decided to keep it around. People tend to gravitate towards the name. So I think we probably did okay there. Well, it's distinctive. It's distinctive looking because our audio-only listeners won't be able to see this right now, but it's lowercase lake, L-A-K-E, and then FS in all caps without a space. So it's kind of like visually, it's a little bit like iOS or Mac OS. But yeah, I've never seen a name quite like it and that makes it memorable, which is a good thing.

4:56Jon Krohn:I guess so, yeah. So expanding out, you mentioned there how Lake FS fits into a broader category of solutions called data lakes. Explain to us what a data lake is because I also hear terms like data lake house and things like that, which that seems like a marketing gimmick to me. But maybe you can fill me in and give me the technical background on these different kinds of terms like data lake, lake house, that kind of thing. Sure. And that's a great question. These are often misused or just used differently by different companies or vendors or people with some interest in the field. I'll give my definition.

5:33It's not necessarily like AWS's definition or Snowflake's definition, but I'll try to give mine. So a data lake at its like the most basic form of it would be just a shared folder. Right. Some central location where John, a member of one team and Oz, a member of some other team, can collaborate on top of data. Right. Maybe you're on the marketing team and you have some data coming in from Salesforce or HubSpot. And maybe I'm doing inventory for the company and I have something coming from my ERP, CRM, whatever it is. And we all want to be able to pull that data into one location. That's the lake.

6:12All the streams go into that lake.

6:14Jon Krohn:All the streams go into the lake. I think that's where it's coming from. I see. And then we have one central location where now we can combine that data to make more out of it. One plus one equals three kind of thing. And once we have everything in one centralized location, all other teams can start using it, communicating with it, building models out of it, do all the nice things that data lets us do. So that's kind of the core idea. Correct me if I'm wrong about this, but I think another feature that I have, that I think about when I think about data lakes is that it, I think it maybe is even implied by what you just said, but it's that the data aren't necessarily very structured.

6:58Jon Krohn:You know, it's kind of like it's a place where all those streams can flow in and you don't necessarily worry about integrating those all into one kind of... So it's very much that, right? Maybe if I contrast it with what we typically had before, which is a data warehouse. In a data warehouse, everything has to be well-organized into tables that have this rigid structure behind them. Typically, there's a team in place that's kind of like the gatekeeper, if you will, because they're the only ones with the expertise. The technology is kind of complex. It's also expensive, right? Whereas with that big shared folder, first of all, you throw your data in, right?

7:35If anyone wants to get access to it and wants to answer some business question, they can do it even if the data is not well-formed or even if it's not like optimally designed to meet a specific query. So that's why I like the shared folder analogy. like there is some inherent mess with that right but sometimes you want that mess because it allows you to move faster right talk more about that mess later i i'm okay nice yeah i mean maybe maybe now's the time even to start getting into that so we in this era of ai what are the particular kinds of needs that data scientists, AI engineers, people who are fine-tuning, pre-training, running

8:21Jon Krohn:AI models, what are the particular needs that they have in terms of their data solution? Oh, that's a good one because I think this is, the answer I'll give you two years ago is probably very different from the answer I'll give you now. So it used to be that the data that we get value from and that we can actually derive like business impact out of was your, let's say, tabular data, right? Tables organized into a database. It's all very well structured. But that's not necessarily the case nowadays, right? Especially like with advances around AI models, what we can do with them. We can extract value from a lot of different other kinds of data as well, right?

9:06So it could be images. It could be embeddings out of those images, labels attached to them, right? All different kinds of modalities that are representative of the information we have at the company. Right? So it's a very broad definition and it's no longer just tabular data or one specific kind.

9:26Jon Krohn:Wow. That's interesting to hear. And something that occurs to me is that we're in this era now where we have more and more data types that we are throwing into, particularly like multimodal large language models. They might have vision data, text data. They can be outputting different modalities of data, images, video, also text. and it seems to me like that could be a particularly tricky situation for data management. Is that correct? Am I correct that this is a tricky situation and how are we resolving it? Yeah, it does make it a lot harder to manage, that's for sure. I think it's not just about the model itself being able to utilize or to output different modalities.

10:14It's also that our use cases around those models become more complex, right? So maybe I have a great model that's doing only text, but I have another model running just before it that I'm using as part of a pipeline that converts an image to text or that converts audio to text. And then the output of that might be converted into an image I'm generating. So it could be a mix of all of these things together. The place where this gets difficult is that even though the dream or the idea of a data lake is that everything streams into that one centralized location, Reality is a bit more dirty than that, right?

10:53If I have images, yeah, I might have those images on just the raw data lake, just as image files. But I'm probably also going to have an embedding of those images stored in a vector database. And maybe I have some features that were extracted from those images in a database. Maybe labels in some other third-party labeling solution. now instead of having that one source of truth that i was expecting to have i have multiple sources of lies right they get out of sync very easily i don't necessarily have access to all the different modalities that were extracted from that one thing it becomes a lot harder to manage

11:32Jon Krohn:right right right yeah tons of complexity and so what are the what are the kinds of solutions what are you guys doing at lake fs to make this new era that we're in with so many different types of data, object stores, feature stores, you know, the different, obviously, data formats that we're talking about is inputs and outputs, different kinds of data pipelines for handling training of these models. Even inference time could potentially be challenging. Yeah, so one of the reasons why we built LakeFS and started working on this problem to begin with was to try to kind of allow humans to better interact with that mess, right?

12:11And this comes up in several different ways. First, it's a very error-prone environment. Maybe I'm, OK, I looked at that images folder. I see 100 images. Great, I'll start working on my model to build something out of that. But here, John comes along from some other team and says, OK, these images are not great. Let's replace a few of them. And let's also add a few others to the mix. And he doesn't know that I'm working on some other thing that depends on those images. Right. So now we have people stepping on each other's toes. And then our CISO would come along and say, OK, I know that Oz needs access to those files, but he doesn't know that John also does and has a valid reason to do so.

12:52And now suddenly you're locked out. And of course, this multiplies by the amount of modalities that you have and the number of people in the organization. So part of the idea behind LakeFS is essentially to create like one facade, one way of working with the data. across all these different modalities, and now we all speak the same language. So just like Git and GitHub, the same idea that they brought to code, if I want to introduce a change, I'll open a pull request. And John's team, they have to sign off on that change before it gets introduced. I have a very structured workflow to how changes get implemented to bring that same concept, that same notion to the data itself, not just the code, but also the data itself, regardless of what modality it is or what type of actual business value represents.

13:44Jon Krohn:I gotcha, I gotcha. So when you talk about this unified facade that's like Git, is it literally like Git where we type commands similar to the kinds of Git commands where you talked about pull requests there, so I guess we're making commits and that kind of thing, that same kind of paradigm? Yeah, so it's very similar in concept. You don't have to type in the commands if you don't want to use a CLI, although one exists for you. And I know some people that's like, and they're religious about using their terminal for everything. I'm not one of those people. I like having a nice UI. So LakeFS also provides that.

14:18But there's also a bunch of SDKs in place, right? If you have this daily job that scrapes the internet and brings in external data, you might want that system to run those commits for you, right? So you'd have a new commit arriving every day at some hour that represents new data coming in, along with just like in GitHub where you have actions and you have those green checkboxes that say, okay, this was tested and is validated. You can do the same for that new data that you just introduced. These are all images. They're all above a certain, I don't know, sharpness level, all at least 500 pixels wide.

14:55You can run those checks automatically and then have the system commit and merge those changes for you. If you're a consumer of that data, you know that for sure if I look at that directory, it's going to contain images that adhere to that level of quality that we set for. Super cool. That sounds very helpful.

15:12Jon Krohn:It sounds like I could make use of that for sure. And so where do you, where do you see this all going? I mean, maybe that's something, maybe that's something even that's proprietary and maybe there's, there's parts about that that are like, you know, confidential roadmap. But to the extent that you can share, where are we going with data for AI models and what kinds of solutions might we need in the future? Yeah. So first of all, it's not confidential, right? We pride ourselves of like building Legafest out in the open. And the core of Legafest, the versioning engine at all scales is fully open source as well.

15:47Oh, really? So transparent that you can just go to GitHub and look at the code. I see. So where do we see this going? So first of all, if I look at like broader than just Legafest, looking at the industry as a whole, One thing we are seeing happening is that everything tends to converge around the object store. It's kind of the source of truth that's emerging. We see like S3, probably the most popular object store out there, release support for tabular data using Apache Iceberg recently, and also being a vector database as well. All of this is like from the last few months. So all of these modalities, instead of going to all these different locations, might be in different technologies and different systems, but they all converge to the object storage, the underlying storage.

16:38So that's definitely something we're seeing and also something we're doubling down on. Right. This is your kind of organizational source of truth. And that will let you manage all the different types of data you might want to store there in like one pane of glass, as they say. Right. One way to manage all of them.

16:53Jon Krohn:Supporting vector databases, like you mentioned there, that's definitely something very important. For our listeners who aren't aware, probably many of them are, the idea of a vector database is that you have some numeric representation of the similarity between typically some large number of documents or files. and it allows you to retrieve similar information or some particular type of information very rapidly across millions, billions of documents in fractions of a second. And so very powerful technology, something of course important to be handling with LakeFS or any kind of data solution that people come up with today.

17:38Jon Krohn:What is Apache Iceberg? You mentioned that there as well and that actually that's not one that I'm familiar with. Yeah so Apache Iceberg is an open source project it's been around about five six years now. What it does is essentially lets you represent a table just like a database table on top of an object store. All right so imagine you have a table on top of something like Amazon's S3 or whatever storage which appliance you're using, Apache Iceberg will let you turn that into an actual table that has a schema that you can modify. You can insert, delete, and update records on it. And it does so without actually forcing you to use one specific compute engine on top.

18:24Maybe you're using Pandas, and I'm using AWS's Athena or Snowflake, but the data is still centralized in one place. And regardless of the tool that's consuming it, They all see the same state of the table, the same abstraction. Nice. Cool. That does sound important. Think about the structured data story, right? Those labels or the metadata about those images, they can also get stored back to the object store as an Apache Iceberg table.

18:50Jon Krohn:Perfect. So to just kind of bring the idea home of how a system like LakeFS works and allows us to manage data in particularly a large organization, It sounds like the bigger the organization gets, the more helpful having a tool like LakeFS in place is. Could you walk us through what it's like if I'm a listener and I start using LakeFS at my organization or I go to the open source repo or maybe I do the commercial version, but for whatever reason, I get LakeFS up and running. What is my experience like day to day as a user of LakeFS? So imagine you had something like GitHub. But instead of showing you the code that you're using, you would see the data that you have in your organization.

19:43And so I can say, OK, maybe I want to experiment. I want to build or optimize an existing model. I'll go ahead. I'll look at the data that I need and I'll branch out of it. I click the big green button on my screen, and now I have, for all intents and purposes, my own isolated copy of everything. And Legifest does this intelligently without actually copying everything to another location. It does keep track of whatever changes you make on your branch. So now I have this entire big data lake that my organization has at my disposal. I can remove half of it to see what happens. I can add new stuff in to improve my results.

20:22And when I address that data from my code, I have to specify which version I'm using. So it's no longer just here's the data lake and let's hope for the best. I'm saying it's this set of data at this version, this specific kind of snapshot frozen in time. What this guarantees is kind of a side effect, is that whatever I'm building now is going to be reproducible later. As long as my code doesn't introduce any variability into it, if it's deterministic, same code, same input data would guarantee the same result. So that's kind of step one. If I'm happy with the change, if I modify the data, introduce new data, remove some stuff, I now can either merge it back so everyone else gets a view of the nice changes I just introduced.

21:10Or as we mentioned, I can open the pull request. I can have those tests that automatically run to ensure the quality. I can do all that. And at the end, the idea is that now I'm treating the data not as just this random thing that flows in and out without anyone actually controlling it, as something that's actually an asset of the company, which typically it is. And now we just let you manage it that way. So all the power that GitHub gives you in terms of collaboration and manageability just for the data itself.

21:41Jon Krohn:it's kind of the wow yeah that was crystal clear uh and i'm really glad that i asked that question because it you layered in this level of visualization things like pressing the big button to get my own branch and then now i have the the whole organization's data lake at my disposal but i also have the security the peace of mind that i'm not going to mess up anyone else's workflow uh or anything else that anyone else is working on i love it it makes a lot of sense Oz, thank you for taking the time today to explain how data are stored for modern AI systems, the kinds of things we need to be prepared for in the future.

22:18Jon Krohn:Before I let you go, I always ask my guests for a book recommendation. That's a good one. So I'll give one of my favorites. I think the typical target audience for it is not usually data scientists, although I I think they can benefit the most from it. It's called Designing Data Intensive Systems by Martin Kleppman. It walks you through pretty much everything involved in actually building a data system. Think of how is a database actually, how does it work? What's an index? How does a database represent its actual data on the storage layer? What happens if two people try to write to a database at the same time?

23:01How does it manage that? kind of all the basics of a large distributed system but also a very small-scale database gets covered by that book and it's very approachable right even though it's a very complex topic it's written in a way that's like digestible you can actually read perfect and that's an o'reilly book i think uh because i think it is yeah yeah i haven't actually read it but i

23:23Jon Krohn:did buy it at some point and it's on my bookshelf uh it ended up being one of those books where I haven't needed to get into the wheeze myself of designing a data intensive system. Someone else has always done it, but hopefully people working with me have benefited from that book being on the shelf. And I think that's why I recommend it because I think as a person that's like consuming those systems, like as a customer of those systems, it really helps understanding at least like one layer below where you are in the stack really helps. Perfect. Thanks for the recommendation. And then finally, this has been a really interesting episode.

23:58Jon Krohn:have loved chatting with you. I'm sure lots of audience members enjoyed learning from you as well. What's the best way to follow you after the episode? So I would say the LakeFS Slack is a great place, not just to talk to me, but also with other people kind of in the same boat trying to solve those same challenges. I'm also available on LinkedIn and Twitter, so happy to leave links to that below as well. Nice. And the LakeFS Slack, yeah, that is easy to find. I just Googled it quickly. And so I'll have the link to that in the show notes. thousands of people have already joined, I can see right now.

24:32Jon Krohn:Indeed, yeah. Nice. All right, Oz, thanks so much for taking the time and catch you again soon. Yeah, thanks for having me.

Read the full transcript

24:41What a nice informative episode.

24:43Jon Krohn:In it, OzCats covered how data lakes provide the flexibility to throw in unstructured rivers of data from multiple sources without rigid schemas, unlike traditional data warehouses that require careful organization upfront. He talked about how modern AI systems work with images, embeddings, audio, and text all at once, creating complex data management challenges where different modalities tend to get stored across multiple systems that quickly fall out of sync. And he talked about how his company, LakeFS, applies Git concepts to data management, allowing teams to branch entire data lakes, make isolated changes, and merge improvements back through pull requests with automated quality checks.

25:24Jon Krohn:Alright, that's it. I hope you enjoyed today's episode to be sure not to miss any of our exciting upcoming episodes. Subscribe to this podcast if you aren't already, but most importantly, I hope you'll just keep on listening. Until next time, keep on rocking it out there, and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.

From the publisher

In this Five-Minute Friday, Cofounder and CTO of lakeFS Oz Katz talks to Jon Krohn about data warehouses, data lakes, and how companies can handle increasingly complex data infrastructures and formats. Hear about lakeFS’s collaboration with Legofest, lakeFS’s approach to helping users collaborate on data lakes, and how to overcome the challenges of working with multimodal data.

Additional materials: ⁠www.superdatascience.com/914⁠

This episode is brought to you by the ⁠Dell AI Factory with NVIDIA⁠.

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
914: Data Lakes 101 (and Why They’re Key for AI Models), with Oz KatzSuper Data Science: ML & AI Podcast with Jon Krohn · 26 min
Listen in VO