In short
Episode Summary: State of AI Startups with Sarah Catanzaro
Podcast Overview Podcast Title: Latent Space: The AI Engineer Podcast Episode Title: [State of AI Startups] Memory/Learning, RL Envs & DBT-Fivetran — Sarah Catanzaro, Amplify Episode Description: Sarah Catanzaro discusses the evolving landscape of AI startups, investment trends, and critical insights from the modern data stack era to forefront AI infrastructures and applications.
Key Themes and Discussions
- DBT-Fivetran Merger
- Significance:
- Not an end for modern data stack but a strategic move toward IPO.
- Both companies were performing well financially.
- Implications:
- The merger signals readiness for larger revenue targets and liquidity options.
- Reinforces the symbiotic relationship between data and AI.
- Data Infrastructure in AI
- Usage of DBT and Fivetran:
- Essential for managing training data and analyzing interactions within AI systems.
- Frontier Labs:
- Relying on advanced data management tools for complex interactions in AI.
- Challenges with Data Catalogs
- Initial Expectations:
- Data catalogs were anticipated to be central to the data stack but have struggled as standalone products.
- Reasons for Struggles:
- Often designed for human use rather than for machine efficiency.
- Consolidation into larger platforms (e.g., Snowflake) has diminished their relevance.
- Funding Phenomena
- Trends in Seed Rounds:
- Increasingly common for startups to raise large seed rounds ($100M+) without clear near-term roadmaps.
- Raises concerns about sustainability and investor confidence.
- Investor Insights:
- Emphasis on understanding a startup's short-term plans and resources needed for milestones.
- World Models and Their Viability
- Current Sentiment:
- Mixed views on world models being overhyped and underspecified.
- Need for clearer definitions and applications across different use cases.
- Market Potential:
- Possible applications in video editing, autonomous driving, etc.
- Consumerization of AI
- Key Theme for 2026:
- Focus on personalization, memory management, and continual learning.
- Essential for retaining users and improving product engagement.
- Market Dynamics:
- Increasing need for AI applications to evolve with user preferences.
- Re-evaluating Reinforcement Learning (RL) Environments
- Criticism of Current Approaches:
- Sarah Catanzaro considers RL environments a temporary trend; real-world data offers richer, more valuable insights.
- Preference for Real-World Data:
- Emphasizes the effectiveness of using actual user interactions over simulated environments.
- Investment Thesis
- Ideal Startups:
- Companies that integrate challenging research problems with practical, impactful applications.
- Focus on areas like memory management, continual learning, and ensuring user retention.
Closing Thoughts Sarah Catanzaro encourages a focus on genuine value creation in AI startups, emphasizing the importance of solid groundwork in areas like memory management and personalization to ensure sustainable growth and user retention.
Links and Resources
- Sarah Catanzaro on X: [@sarahcat21](https://x.com/sarahcat21)
- Amplify Partners: [amplifypartners.com](https://amplifypartners.com/)
- Latent Space Podcast: [latent.space](https://latent.space)
- Full Video Episode: [Watch Here](https://www.youtube.com/watch?v=LdVldygYE6I)
Episode Timestamps
- 00:00:00 - Introduction: Sarah's Journey from Data to AI
- 00:01:02 - DBT-Fivetran Merger Discussion
- 00:05:26 - Data Catalogs and Their Challenges
- 00:08:16 - Insights on Data Infrastructure at AI Labs
- 00:10:13 - Funding Environment Overview for 2024-2025
- 00:17:18 - Discussion on World Models
- 00:18:59 - Memory Management and Continual Learning
- 00:23:27 - Critique of RL Environments
- 00:25:48 - The Ideal AI Startup Profile
- 00:28:02 - Closing Thoughts and Resources
This episode provides a comprehensive overview of the current trends and challenges in the AI startup ecosystem, making it essential listening for anyone involved in AI engineering and investments.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding the Data to AI Transition
0:45 to 2:07
Discussion on the symbiotic relationship between data and AI.
“That's actually what brought me into data.”
The Impact of DBT and Fivetran Merger
2:07 to 4:00
Analysis of the DBT and Fivetran merger and its implications.
“And they were the presumptive winners in their categories anyway.”
The Role of Data Catalogs in Modern Infrastructure
4:00 to 6:32
Exploration of the importance and challenges of data catalogs.
“But you're now saying that the dbt5 trend are being used for training data.”
Data Stacks in AI Companies
6:32 to 7:59
Observations on how AI companies manage their data stacks.
“So I think there were a couple of things.”
Trends in AI Startup Funding
7:59 to 10:00
Insights into current trends in funding for AI startups.
“So I do wonder at times if we built data catalogs for the wrong people and potentially even for the wrong use cases.”
Challenges in Evaluating Startup Viability
10:00 to 14:00
Discussion on the difficulties in assessing startup potential.
“Like upwards of$100 million in a seed round where you have a long-term vision, but not a near-term roadmap.”
Understanding Startup Valuations and Hiring Dynamics
14:00 to 16:42
Learn about the complexities of startup funding and how valuations affect hiring.
“So like, I understand why they need that funding.”
Exploring World Models in AI
16:42 to 18:54
Discover the current skepticism and potential of world models in AI technology.
“Okay, so obviously we can go about that forever.”
The Importance of Memory Management in AI
18:54 to 22:56
Understand memory management's role in AI applications and user retention.
“A theme that I've been spending a lot of time thinking about is memory management and continual learning.”
Evaluating AI Environments and Their Effectiveness
22:56 to 25:37
Examine the relevance and effectiveness of AI environments in real-world applications.
“That's also a fascinating infrastructure problem because you have to load and unload and, you know, cache and all the good stuff.”
Show all 12 chapters
Archetypes of Exciting Startups and Their Innovations
25:37 to 28:00
Explore the characteristics of innovative startups and the relationship between research and applications.
“We have maybe three minutes for any other stuff that you think about just the state of startups in general, state of funding.”
Final Thoughts and Farewell
28:00 to 28:29
The host and guest exchange final thoughts and farewells, sharing contact info.
“that I want to sort of dig into there, but we're short on time.”
Transcript
Automatic transcript. May contain errors.0:11Okay, we're here with Sarah Karenzaro from Amplify. Welcome. Thank you. First time on the pod. to be here. Too long. I know, I know. We've known each other for so long. Yeah, never made an appearance. It also made the transition from data to AI, I guess. I don't know if, I did, I don't know if you were always like as deep on AI, but I'll be, there's a lot of simpatico. Yeah, I've always actually kind of oscillated between data and AI. Sure. Like arguably, I started my career in quote unquote AI. It was just more like symbolic systems back then. But as you said, I think like there are, They're so symbiotic.
0:47It's almost hard to divorce them. That's actually what brought me into data. I was like, I want to better understand what happens when I write a SQL query. Yeah. Let's briefly touch on data because I think, obviously, that's a lot of where you and I first met. DBT 5.Tran. That was so cool. I mean, how do you think about the end of the modern data stack? Okay. So a lot of people look at the DBT-5Tran merger and talk about the end of the modern data stack. And I think that is a fundamentally wrong take. Both of these companies were growing very healthily. And you funded DBT? We funded DBT. So both of the companies were actually beating their revenue targets.
1:34I think what you're more seeing is, you know, IPO environment wherein companies are expected to have far more than, you know, like 100 million revenue. And so. What would you say the bar is now, 300? No, like above 600. Yeah, yeah. And the combined company is 400? I believe that they'll actually be close to 600. I don't have the exact number. But they're clearly just getting ready for IPO. So, you know, basically, like the merger was a way to accelerate that path to liquidity. As you might remember. And they were the presumptive winners in their categories anyway. Exactly. Exactly. You know, I think one of the things that has actually pleasantly surprised me, and this speaks to, again, the symbiotic relationship between data and AI.
2:23Many of the big frontier labs are actually using both DBT and 5TRAN. I recall talking to folks at Thinking Machines, like within weeks of the company's formation, and DBT was already an important part of their stack. It was certainly like training data sets need to be managed. We need insight into what users are doing on these platforms. And in fact, like the way in which you would analyze interactions with an agent or analyze interactions with an LLM is even more complicated. And so while I think perhaps like the demand for analytics engineers, the demand for data scientists didn't explode in the way that some people thought, like analytics engineers are not one third of personnel.
3:08That doesn't actually mean that the demand for the tools is not still like very prevalent. But you go where you want it. You want it to democratize things. You got it. Yeah, yeah. I mean, I guess we democratized things by perhaps reducing the need for the people. I don't know whether or not that is a good thing. But honestly, I do think that like the fact that it is easier than ever from a tooling standpoint for people to make data driven decisions is probably a step in the right direction. And I've become actually convinced that like, well, every company does need analytics engineers and does need data scientists.
3:46They probably don't need armies of them. And probably having like a moderately sized data and analytics team is a good thing. Yeah. So you touched on an interesting thing I wasn't planning to ask, but this is interesting. As I come from the data field, data was synonymous of analytics. Yeah. But you're now saying that the dbt5 trend are being used for training data. Is there any notable differences in the workloads or the requirements? Undoubtedly, there will be. I mean, I think one of the things that we saw with analytics that was surprising to some of the people in the data infrastructure space was that the workloads were actually quite predictable.
4:27They were quite predictable because like many of them were actually not being generated by humans, but rather by deterministic systems. So like a lot of it was like BI dashboards that are, you know, Tableau that is actually hitting your database or maybe not Tableau, but like Looker or, you know, Hacks or something like that. I think with analyzing, curating, preparing data sets, it's a bit more ad hoc. And so undoubtedly, it will be less predictable. I don't know if that really changes the way that we approach developing data infrastructure. I talked like some people are quite interested still in like things like learned indexes, learned optimizers.
5:12And it's a bit easier to build a learned optimizer if you have more predictable workloads. And so it could change the way that we approach things like that. Yeah. Data catalogs, do they become more important? Are they transferred? Oh, man, like straight to the gut. So that was something I got wrong. I'm sorry, I don't know the background. What did you? I just, I really believed that data catalogs were going to become an important part of, you know, the modern data stack. And the players are Atlin. She's Singaporean, so I. Yeah, yeah. There was. Data world. Girl, data world metaphor within our portfolio.
5:54They've all struggled as a category. They all have struggled a bit as a category. Many of them have been acquired subsequently, which suggests that this was not perhaps a standalone category. As a data scientist, I spent so much time working on data catalogs. And so I kind of felt like this was the thing I wanted. I didn't want to have to build the... More to the point also, pre-training data, you have a lot more heterogeneous data all over the place. Yeah. And like you need to keep on top of it and you need to make it discoverable, accessible and all that. So why didn't it work? So I think there were a couple of things.
6:35I think we have seen some consolidation in the modern data stack, particularly around, you know, some of the key components, whether it was, you know, Fivetran or DBT or, you know, Hex or, you know, Snowflake. Many of these products offered kind of like data cataloging capabilities as a feature. And I think for humans, that was good enough. Like the data catalog that you had available in Snowflake was good enough. The data cataloging capabilities available in DBT, like those were good enough. They did. DBT, like obviously as they didn't build the cloud, they were going to build it. Yeah. What else do you do?
7:15I mean, it's actually funny. In fact, my colleague Barr at Amplify was the like products lead on these kind of like metadata services. I think it's still not obvious to me, but I think one opportunity that might have existed and or could have been realized was the opportunity to build data catalogs, not for humans, but for machines. This would look a little bit more like metadata services. I don't just mean for agents, although I think that opportunity is arising more, but even like microservices and things like that. Okay. Yeah. So I do wonder at times if we built data catalogs for the wrong people and potentially even for the wrong use cases.
8:07I think a lot of data cataloging companies ended up focusing on discoverability when perhaps the real market opportunity was in governance. Governance is very important. Any other comments just about what you know so far about the data stacks of the large labs? I guess, obviously, a lot of data people who might be listening would want to sell into them. Yeah. I mean, a couple of observations. One is that they are actually paying careful attention to their data stacks. I think they're thinking about problems ranging from data discoverability to data preparation to even things like the efficiency of data loading.
8:47Like if you're unable to load data to a GPU efficiently, then the GPU is going to sit idle and that's going to be a kind of like a cost. Yeah, yeah, exactly. So what solution handles that? I don't actually... I mean, I get to talk about, yes, exactly. Plug my portfolio companies. We have a portfolio company called Spiral that has developed a file format called Vortex. and they make data loading like super efficient. Specifically to GPUs? Specifically to GPUs. Okay. Yeah, yeah. Good to know. One of the things that has surprised me though is actually that like so much data infrastructure has actually scaled quite elegantly to meet the AI use case.
9:32You would hope. You would, but like the scale of these AI companies, it's incredible. um and so it's not as big as ads maybe maybe yeah i i think that could change you know like as agents actually become kind of like more prevalent or and are interfacing with each other and therefore like perhaps like the number of transactions explodes i have a friend who works on transactional databases at open ai and i was like so you must be like building databases like this is like a paradigm shift in terms of like the scale that like databases are like going to need to handle and he's like no we use rocks at like it's the one they acquire right yes exactly yeah yeah um very cool okay let's just talk about funding around it because obviously that's like a big theme this year what comes to mind in terms of looking back at 2025 uh what stands out it was crazy um yeah you can give anonymized examples of like what what does crazy look like yeah i mean i I think crazy looks like raising upwards of$100 million.
10:39Seed. Like upwards of$100 million in a seed round where you have a long-term vision, but not a near-term roadmap. Yeah. This is something that I'm seeing happening not just occasionally, but quite frequently. Yes. And it definitely makes me anxious because firstly, when founders are asking me, how much should I raise? I'm typically saying like... Three, like five. Well, like, what do you need to do? Like, what are your milestones for the next, let's call it like 12 to 24 months? What resources do you need in terms of, you know, headcount, compute, equipment to unlock those milestones and then like maybe add like a 20 % buffer or something like that.
11:29But doing that analysis requires you to understand what you're going to build in the next zero to, let's call it, 24 months. I've talked to some companies and they're like, we're building a frontier lab for X. And I'm like, OK, cool. I get the long-term vision. There is an opportunity to make AI more secure, make AI more humane, make AI more data efficient, whatever it might be. So I'm bought into the long-term vision. And that, you know, for me as an investor is super important. Like, so let's talk about, like, what your team's going to work on in the next six months. They're like, maybe we might build a consumer app.
12:07Like, you know, we're definitely... I feel like I know exactly the company you're talking about. But like, I wish I was talking about like one specific company. I'm actually talking about like several companies. And look, like, I'd be a hypocrite to say that like, I've never done investments like that. But I've done investments like that when I really know the people and I'm like, they're going to figure it out. What is frightening about this funding environment is that you meet a founder. They're like, I'm raising$100 million. I'm raising like a billion dollars maybe at times. And you need to make a decision in seven days.
12:43And I can't tell you what I'm going to do for the next six months. And so like you have no way of even gaining conviction that they're going to figure it out because you only have like seven days to get to know them. I think what some of the founders are missing is like you only have seven days to get to know me. If you haven't figured it out, like you probably want a partner who's going to be working closely with you to help you figure it out. I mean, they're absolutely viewing it as transactional, right? Like, yeah, they don't care. No, they care about, you know, the most money at the highest valuation.
13:13I mean, the crazy thing is that they don't even seem to care about dilution. It's just like the most money at the highest valuation. Yeah. But, you know, it does send a signal that helps. So, I mean, yes, I think it does right now send a signal. Okay. I'll tell you how it affects me. And I hate it. I hate it. All right. Antithesis came out of stealth this week. Right. And it's like the only thing I know about them is they do something, something in AI testing. And Jane Street led a seed round of$100 million. We invested in it too. I can tell you what they do, but they do the permanistic simulation testing.
13:48The thing that is the lead is the money. Yeah. And then like, okay, well, who else uses it other than GeneStreet? Like, what do you do that's innovative? Palantir. Okay. Warp stream. Yeah. So, yeah. Okay. Anyway, so maybe, maybe, and so this is a bad example because they're actually legit, but like, you know, there's, there's a lot of similar examples where they just lead with the money and like, there's no, no much substation behind it. maybe it's just bad storytelling and that's why i as a podcaster get to talk to them i just talk to general intuition and like once you spend some time with them then you're like okay this is why they raised 100 million dollars but like without that context it's like really hard to understand anything well and and like i think there are some companies that are raising you know a hundred million dollars or more because they need it like a good example might be like periodic in addition to, you know, yeah, they need to build out a wet lab and like designing a wet lab that can support high throughput biology, which is absolutely critical to their goals.
14:48That's costly. So like, I understand why they need that funding. But again, there are others where like they don't have these near term milestones. I think the thing that is a little bit perturbing to me, many of them are doing it because it makes it easier for them to hire because you know there are all of these candidates who like want to be want to work at a company that is like a unicorn or a near unicorn they're pitching because the alternative is work at a big lab where you know it's yeah the prestige and the money is there yeah well or the alternative is like work at like an early stage startup uh but but but but like there's something about like the big valuation, that becomes enticing.
15:27They're also kind of pitching candidates. They have a compelling equity pitch where they're like, okay, maybe you're getting less than 0.1 % of the company, but given the valuation, the value of your equity is already like$10 million or something like that. And they also guaranteed a dollar value of the equity. You mean that they'll offer them a loan to pay? A buyback if you want to sell it. Yeah. Because they have so much cash. But the thing, though, is that the valuation is a made-up number. Valuation, until a company exits, it is an entirely made-up number. So I could just be like, you know what?
16:14The latent space pod, that is worth$5 billion. And we could agree. Like, I as an investor could say, like, that is the price. And now the company is worth$5 billion. Like, do you think that, like, if you were to... Yeah, it's not real. It's less actual, it's less active than any volume. And given the funding amounts that they're raising, too, like, if they spend that and they, you know, get acquired for less than that amount, then, like, their teams are getting nothing. I wish people were kind of like more sensitive to this dynamic and thinking more about like what is the upside associated with the company and you know more fundamentally like do I deeply believe in this vision because I think like joining companies because like they have a billion dollar valuation it's just it's not the right way to choose a job.
17:02I hear you. Okay, so obviously we can go about that forever. And there's also some stuff with cyclical funding and all that stuff. But I do want to be more relevant to engineers and researchers. What are the themes that are really strong? So one thing I'll point out is world models, just in general, are a really strong bet. I would say, so I have every near-ebs, I go to this group of researchers and we take a vote on the top themes of the year. Everyone's extremely skeptical about world models. I think it's a trailing indicator because LLMs have been so enormously successful. You're like, I don't need anything else.
17:42I don't know if you ever take on world models or any other top theme of the year. My take on world models is that we have not yet defined what a world model is. Oh yeah, there's like three definitions right now. Yeah, I think there's a lot of confusion about what a world model is and therefore what it should be used for. We're already seeing plenty of market potential for video models, including for things as perhaps banal as video editing. I think, you know, we're already seeing some applications of world models to things like autonomous driving and potentially even coding. But again, it really hinges upon like, how are you defining world models?
18:20And I think one challenge that people have seen is that like world models perhaps designed for one specific use case might not generalize to others. So as an example of this, like world models for like video game generation might not like generalize to like factory settings or robotics. I use the word might like strategically because I think like it is potentially a research problem that might be figured out. So that's part of the general intuition podcast that we did is that they have some evidence. Yeah, yeah. I think like it is possible. It's just we're not there yet today. A theme that I've been spending a lot of time thinking about is memory management and continual learning.
19:03I work with a lot of companies. Save startup. Okay. I think I know what startup you're thinking about as well. But I actually see a lot of market potential for memory management and continual learning. My interest in this is actually more driven by conversations with practitioners. Personalization is so important right now. I think what we're seeing is that like a lot of AI application companies, they're growing really quickly, but they suffer from relatively low retention, relatively high churn. So if you're developing an app like Cursor, how do you ensure that your users don't switch over to Windsor?
19:49Yes. or, you know, cloud code or cognition or, you know, whatever else when they release new features. Yeah, cursor rules isn't enough, right? Like, it's like the shittiest form of memory. Yeah, yeah. You know, and it's great. But yeah, I agree with that. But also it's like, I've publicly mused about this before, where, like, memory is very poorly implemented today in a lot of surfaces. Like, even chatGVT, I wouldn't say, like, people are particularly excited about it. Okay, all right. You feel stronger about it than I do. Yeah, yeah. I mean, I wish ChatGPT had, you know, much better. Yeah, this is mostly the leading one.
20:29I don't know. So, and then I think, like, just in general, it makes product management harder because what is the product? It's a combination of you plus memory. And, like, when you have a bug, is it the memory or is it something core? And that's, as a user, Especially if it's consumer, there's going to be zero patience for any of this. I agree. But that said, consumers seem to be tolerating products with no implementation of memory today. So I think better is still probably better than what exists now. Better is better than nothing, I guess. Would you agree with the statement that basically, let's say, a key theme of 2026 is this personalization?
21:12I would call it kind of like the consumerization of AI in the same way that consumerization of enterprise was a trend like 10 years ago. Yeah, I mean, I think that is a good way of putting it, too. Like, I don't, for what it's worth, think like this is just a consumer or prosumer phenomena. If you are an enterprise that is adopting, again, like Devon or Augment or something like that, you probably also want your models to kind of like learn. Totally. I'm not limited. Yeah. Like you start to like K factor. I had to explain what that is to so many founders. And, you know, like this, like if you're in normal SaaS, this is what you obsess over.
21:49And to AI founders, they're like, what do you mean growth? This doesn't just show up. Like, I mean, it has, though. But I think like it has because for a while, you know, AI has just felt magical. But now we're getting more accustomed to the magic and it's no longer enough. And I think we need to revert to some of the old tips and tricks for retaining people and bringing them in. Personalization is one of them. I always kind of intermingle like memory and continual learning because I think like one interesting element of personalization is not just learning facts about your or your preferences, but like actually learning new skills from interactions with you.
22:38And, you know, learning as the world changes, like there are new versions of languages and frameworks and, you know, other repos that are coming out all the time. The world is changing all the time. human intelligence is incredibly dynamic and yet like uh artificial intelligence is just so static today yeah but like so it must update weights yeah for you but but but that also means that like it's an interesting kind of like systems problem because like if you must update weights then like you know weights become stateful and today like inference is not stateful so so you know i think i think there's going to be like a lot of kind of fun gnarly problems to figure out as we figure out things like personalization and continual learning.
23:19That's also a fascinating infrastructure problem because you have to load and unload and, you know, cache and all the good stuff. Yeah, exactly. One more thing. I think we have time for one more take on our own environments. Huge topic. Is it just a Docker container with some custom software loaded and logging stuff out? What are the good ones like and what are the average ones like? So I know I'm going on record on this. and like I'm actually okay to be wrong but I think our own environments is just a fad. Oh God. Oh no. They're all fake? I mean like people are like okay. The thing that makes me take it seriously.
23:59The labs I know are paying seven, eight figures for our own environments. And they could build it in-house. They're not. And I don't understand why. I mean they were paying seven to eight figures for like piss poor data annotation too. Yeah. Uh, so like, and then data labeling before, like the labs have a lot of money. I think perhaps like oral environments could create some value in the short term, but I think to the point about like what makes a good oral environment, what makes a bad oral environment, I think the best oral environment is, is, you know, the, the real world. Um, why would I, you know, want to, uh, buy a DoorDash clone when like I can just use logs and traces from you know DoorDash itself it doesn't mean that we don't need to blend in parallel yeah I mean I think like using the real world using real apps as like our L environment is in fact like the best thing and this is what Cursor does like they actually do use uh you know real user activity on their platform to significantly significantly like improve both their coding agents as well as tab and i think that's one of the the approaches that has like made the platform so compelling it doesn't like you still need to figure out like the right rubrics you still need to figure out like the right set of tasks uh so there are some aspects of aural environment design you know at least as we're talking about it today uh that i think are going to remain incredibly relevant but like just building a clone of an app I think is not that useful.
25:36Yeah. Okay. That is all I'll take. We have maybe three minutes for any other stuff that you think about just the state of startups in general, state of funding. Yeah. So maybe I can talk about like just the archetype startup that is like most exciting to me. Yes. Press for startups. Yeah. Yeah. I love investing in, you know, infra, tools, platforms, et cetera. And as we talked about with continual learning, I think like there will be opportunities for like new tools, platforms and infra in the future. I've spent a lot of time thinking about like applications today and specifically like the relationship between research and applications.
26:15An example of this is like, I think there were a lot of advances in RAG. And the biggest beneficiaries of these advances were the application companies for whom retrieval was a critical unlock. So as an example of this, like Harvey, Habea. I knew you were going to say Harvey. Yeah. I mean, they have like really interesting RAG implementations. They have hired researchers, like really good researchers to kind of advance the state of the art. And that enables them to build a better product. I feel this way very much about like rule following and customer support. Rule following is like a hard research problem.
26:55But if you solve rule following, then you unlock better customer support. And I think a lot of Sierra's success can be attributed to like their focus on this. So I've been thinking about like even for something like continual learning or memory, what is like the killer use case where you can either offer a dramatically better experience by having a good memory implementation or you can do something that was just not possible today. I think you can also think about this in the inverse. like and often the best company is emerging this way they're like i'm trying to do this thing but in order to actually do it i need to solve this hard technical problem uh that that that's kind of like the story of runway uh i don't think they would have built models if they didn't have to uh but i love that that combination of like we're delivering something that is like better for consumers better for prosumers better for users uh but we're doing so by solving these like really gnarly research and engineering problems.
27:56Yeah. I don't want to... Yeah, got it. There's so much that I want to sort of dig into there, but we're short on time. But just thank you in general. I don't know if you have a general call to startups for a page somewhere that you want to point people to. Twitter. Facts. Whatever it's called. Yeah, you can find me. You can find me there. We're in South Park. With the one-eyed dog. I'm easy to spot. Okay. Well, thank you so much for your time. I know you got to go, but I appreciate it. Of course. It was great seeing you. And thanks for having me. Yeah. Thanks.
From the publisher
From investing through the modern data stack era (DBT, Fivetran, and the analytics explosion) to now investing at the frontier of AI infrastructure and applications at Amplify Partners, Sarah Catanzaro has spent years at the intersection of data, compute, and intelligence—watching categories emerge, merge, and occasionally disappoint. We caught up with Sarah live at NeurIPS 2025 to dig into the state of AI startups heading into 2026: why $100M+ seed rounds with no near-term roadmap are now the norm (and why that terrifies her), what the DBT-Fivetran merger really signals about the modern data stack (spoiler: it’s not dead, just ready for IPO), how frontier labs are using DBT and Fivetran to manage training data and agent analytics at scale, why data catalogs failed as standalone products but might succeed as metadata services for agents, the consumerization of AI and why personalization (memory, continual learning, K-factor) is the 2026 unlock for retention and growth, why she thinks RL environments are a fad and real-world logs beat synthetic clones every time, and her thesis for the most exciting AI startups: companies that marry hard research problems (RAG, rule-following, continual learning) with killer applications that were simply impossible before.
We discuss:
* The DBT-Fivetran merger: not the death of the modern data stack, but a path to IPO scale (targeting $600M+ combined revenue) and a signal that both companies were already winning their categories
* How frontier labs use data infrastructure: DBT and Fivetran for training data curation, agent analytics, and managing increasingly complex interactions—plus the rise of transactional databases (RocksDB) and efficient data loading (Vortex) for GPU-bound workloads
* Why data catalogs failed: built for humans when they should have been built for machines, focused on discoverability when the real opportunity was governance, and ultimately subsumed as features inside Snowflake, DBT, and Fivetran
* The $100M+ seed phenomenon: raising massive rounds at billion-dollar valuations with no 6-month roadmap, seven-day decision windows, and founders optimizing for signal (”we’re a unicorn”) over partnership or dilution discipline
* Why world models are overhyped but underspecified: three competing definitions, unclear generalization across use cases (video games ≠ robotics ≠ autonomous driving), and a research problem masquerading as a product category
* The 2026 theme: consumerization of AI via personalization—memory management, continual learning, and solving retention/churn by making products learn skills, preferences, and adapt as the world changes (not just storing facts in cursor rules)
* Why RL environments are a fad: labs are paying 7–8 figures for synthetic clones when real-world logs, traces, and user activity (à la Cursor) are richer, cheaper, and more generalizable
* Sarah’s investment thesis: research-driven applications that solve hard technical problems (RAG for Harvey, rule-following for Sierra, continual learning for the next killer app) and unlock experiences that were impossible before
* Infrastructure bets: memory, continual learning, stateful inference, and the systems challenges of loading/unloading personalized weights at scale
* Why K-factor and growth fundamentals matter again: AI felt magical in 2023–2024, but as the magic fades, retention and virality are back—and most AI founders have never heard of K-factor
—
Sarah Catanzaro
* Amplify Partners: https://amplifypartners.com/
Where to find Latent Space
* X: https://x.com/latentspacepod
Full Video Episode
Timestamps
00:00:00 Introduction: Sarah Catanzaro's Journey from Data to AI00:01:02 The DBT-Fivetran Merger: Not the End of the Modern Data Stack00:05:26 Data Catalogs and What Went Wrong00:08:16 Data Infrastructure at AI Labs: Surprising Insights00:10:13 The Crazy Funding Environment of 2024-202500:17:18 World Models: Hype, Confusion, and Market Potential00:18:59 Memory Management and Continual Learning: The Next Frontier00:23:27 Agent Environments: Just a Fad?00:25:48 The Perfect AI Startup: Research Meets Application00:28:02 Closing Thoughts and Where to Find Sarah
Get full access to Latent.Space at www.latent.space/subscribe




