In short
Tristan Handy (dbt Labs CEO) explains how dbt helped define “analytics engineering,” why it uses SQL and a DAG of small transformations, and how dbt’s semantic layer and “Fusion Engine” improve trustworthy analytics—especially for AI agents. He also discusses agentic migration via skill.md files and why dbt fits an agent-driven future. He covers open-source/community stewardship vs monetization, and the recent dbt Labs merger with Fivetran (Handy becoming president).
Guest backgrounds
Tristan Handy is CEO/founder of dbt Labs (originally Fishtown Analytics). He coined “analytics engineering,” built dbt (trusted by 100,000+ teams), and led the company’s open-source community and training/certifications.
Key claims
dbt enables software-engineering rigor for data transformation; semantic layer prevents agents from confidently returning wrong metrics; Fusion Engine adds “type safety” across SQL dialects; skill.md files can compress large dbt migrations into ~6 weeks.
Notable examples
study of ~100 companies shifting to the modern data stack; Siemens-scale semantic-layer need; migrations from stored procedures/Spark to dbt taking consultants a year/$millions vs agents doing it in six weeks/tens of thousands.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOFounding DBT Labs
0:55 to 2:30
Tristan discusses the origin of DBT Labs and its initial vision.
“This episode of Super Data Science is made possible by Anthropic, Notion, and Gouroubi.”
Analytics Engineering Insight
2:30 to 4:50
Tristan reflects on the role of analytics engineering and its evolution since 2016.
“The prior company that I worked for had a bunch of smart people that worked at it.”
Empowering Analysts
4:50 to 7:10
Exploring how DBT empowers analysts to take control without overwhelming them.
“There's been this power shift more and more towards analytics type people.”
DBT User Experience
7:10 to 9:30
Tristan explains how users can effectively interact with DBT.
“And Spark is exactly the wrong tool if you want data analysts to be able to do any of this work because right at the outset, it is just incredibly challenging to get it set up.”
The Role of SQL in DBT
9:30 to 12:20
Discussing how DBT leverages SQL for data transformation.
“And what you don't want to do is you don't want to have these super, super complicated monolithic transformations, but you want to kind of stage things out.”
The DBT Community
12:20 to 14:01
Tristan shares insights on the open-source community behind DBT.
“And that means that when you go to your BI tool or when you load up your analytics agent or wherever you want to consume this data, the data that these front ends have access to has all been really nicely modeled.”
Balancing Open Source and Capitalism
14:01 to 17:28
Learn how open source principles can coexist with commercial success.
“DBT Labs is built around an open source foundation and a community driven identity.”
The Dilemma of Staying Independent
17:39 to 21:19
Explore the factors influencing decisions against acquisitions in tech.
“Yeah, it sounds like you're striking a great balance.”
Understanding the Semantic Layer
21:19 to 24:20
Delve into how dbt acts as a semantic layer for data handling.
“I think that that kind of drive to be as open source as possible, to be community minded, to continue to grow in a way that allows everyone, all of your users to benefit is going to in the long run be great for you.”
The Role of Ontologies in Data Science
24:20 to 28:00
Examine the importance and challenges of ontologies in evolving domains.
“And it's particularly relevant today because absent these types of cues, how do you measure X thing, AI agents have to try to re-derive that for themselves at every turn.”
Show all 17 chapters
Understanding Ontologies in Data Science
28:00 to 30:10
Explore the challenges of maintaining ontologies and the role of LLMs.
“And I, you know, I, again, I'm not a super expert here, but you can imagine how a domain like genetic research could be really suitable for this type of work because, uh, it is very verifiable.”
The Impact of dbt's Fusion Engine
30:10 to 34:08
Learn about dbt's Fusion Engine and its potential to revolutionize analytics engineering.
“There's another dbt, I don't know if you'd call it a product or a feature.”
The Future of DBT in an AI-Driven World
35:34 to 41:22
Discuss how DBT adapts to AI developments and its relevance moving forward.
“I mean, I'm like, to me, there's no question if people are using cloud code, you're using open AI codecs, and you really should be because they're so magical.”
The Evolution of Analytics Engineering
41:22 to 42:08
Examine the transition of analytics engineering and the role of skill files.
“Honestly, one of the things that we've been doing for as long as we've had a commercial business is we've been replacing kind of old school ETL products, whatever.”
Understanding dbt Migrations and Agents
42:08 to 46:02
Learn how dbt migrations have evolved and the role of agents in this process.
“It could encode hundreds of hours of collective human experience and run a dbt migration flawlessly.”
Book Recommendation: Musashi
46:03 to 47:28
Tristan shares his insights on the book 'Musashi' and its cultural significance.
“So despite it having been such an interesting episode, we should start wrapping up.”
Closing Remarks and Future Engagement
48:05 to 51:45
Wrap-up of the episode with insights on Tristan's work and how to stay connected.
“You know, I don't spend enough time thinking about that great culture of samurais, which as a kid was somehow so fascinating.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:Today's exceptional guest coined the term analytics engineering. He built the tool 100 ,000 data teams rely on and now says agents can compress a million dollar year long data migration into just six weeks. Welcome to episode number 1019 of the Super Data Science Podcast. I'm your host, Jon Krohn. Today's guest is Tristan Handy, CEO and founder of DBT Labs, the company behind, you guessed it, DBT, which is the beloved open source tool that brought software engineering rigor to data transformation. A decade after founding DBT Labs, Tristan is now merging the business with Fivetran and becoming president of the combined company.
0:40Jon Krohn:In this episode, he explains why he turned down acquisition offers for years, how the semantic layer keeps AI agents from confidently getting your metrics wrong, and how skill files let agents do work it used to take humans weeks of training to learn. Enjoy. This episode of Super Data Science is made possible by Anthropic, Notion, and Gouroubi. Tristan, welcome to the Super Data Science podcast. A treat to have you here. Where are you calling it from? Oh, yeah, thanks for having me. I am in the great state of Pennsylvania. I live in the main line outside of Philadelphia. That's right. And your company was originally called Fishtown Analytics.
1:19Jon Krohn:Is that Philadelphia is also known as Fishtown, I guess? Well, Well, Fishtown is a neighborhood in Philadelphia. When I started the company, I lived there. Yeah. Gotcha. Yeah. That makes a lot of sense. I probably should have looked into the facts behind that before asking you. It's been a long time since, but originally it was the biggest cod fishing area on the East Coast. I came to learn. There you go. And so that was a decade ago that you founded Fishtown Analytics. Now, DBT Labs, as most people would know it. You've raised hundreds of millions of dollars over the year, and your platform is trusted by over 100 ,000 teams today.
2:02Jon Krohn:That's a wild journey. In 2016, you articulated a vision that data teams deserve the same rigor and tooling as software teams. And you coined the term analytics engineering. What led you to that insight in 2016? And what has the journey been like to today? It's a good question that I don't often get asked. Why did I think that that was true? It was not like some bolt of lightning insight. The prior company that I worked for had a bunch of smart people that worked at it. I ran marketing. The company was called RJ Metrics. and we were a kind of prior generation BI. We could, in our numbers, see the market shifting in the like 2014, 2015 time horizon.
2:50A lot of the really early adopting like tech forward companies were leaving. So we did a study, a bunch of us on the exec team talked with, I think the grand total was like something like 100 companies. and we found out that they were moving to what eventually became known as the modern data stack. Back then, it was people were storing their data in Amazon Redshift and they were using BI tools that were purpose-built to run on top of Amazon Redshift. But the funny thing is they were using this new technology, but they didn't, like a lot of the problems that they were experiencing were the same problems that they'd experienced before.
3:32So they like new technology, same problems. And you're like, well, what's going on here? Like, what have we failed to learn something important? And, you know, I could I could talk about this for about forever. But the conclusion that I came to, along with some of the other folks on this team, was that we were asking data people to build production systems. The cloud meant that everything was always on and there was an expectation that you could hit the button run in the top right of every dashboard and that it would produce new data and that data was always correct. And so this is really a production software system, but we were not equipping them with the tools to do that.
4:08And so that's kind of where everything flowed from that has happened in my little neck of the woods over the past 10 years.
4:16Jon Krohn:And more recently in 2020, you said that DBT's purpose is to empower analysts as first class owners of another transformation process. So, you know, in 2016, you saw this transformation process, this opportunity to be providing tooling to enable analytics engineering to form it all. And then in more recent years, there's this desire to empower analysts to be transforming whole organizations. It seems like power has shifted as data capabilities, analytics capabilities, AI capabilities have become so important in organizations. There's been this power shift more and more towards analytics type people.
4:59Jon Krohn:And yeah, so how are you balancing empowering us technically without overwhelming us with software engineering complexity? The thing that I think is the most magical about people who call themselves data analysts is that they're not afraid of complexity and they tend to dive in to kind of multifaceted problems. Now, that sounds kind of like high level and vague. But the thing is that data analysts are a role that spans two big disciplines. One is the actual quantitative discipline. You have to actually know something about math and statistics. And I would even put the software engineering part in here, like the kind of more technical mathy brain stuff.
5:49And then on the other side, you have to know about the actual business domains that you're working in. Because a data analyst that doesn't actually know anything about the domain that they're working in is really just an order taker. They're not adding any value beyond their technical capabilities. And especially today, when anyone can have technical capabilities with your agent, those kind of can't stand on their own as a source of value. You know, almost everybody finds their way to the data analyst role in some kind of weirdy to use in Chronic Pathway because you generally become an expert in the one side or the other.
6:29And then you learn the other side on the job. The way that I've always thought about building DBT and is this like idea that we should empower these magical unicorn humans to be able to live up to their potential. potential, but not overwhelm them with technical details at the outset. Now, that's not to say that data analysts are not capable technically, but that's not the first order concern. That's not why somebody becomes a data analyst. Back in 2016, the kind of cool kid way to do data transformations was Spark. And Spark is exactly the wrong tool if you want data analysts to be able to do any of this work because right at the outset, it is just incredibly challenging to get it set up.
7:22The syntax is very complicated. Nobody with a data analyst background knows how to write Spark code at the outset. And so one of the reasons that I chose SQL was to make it accessible to data practitioners, to data analysts. And it kind of goes from there, but the idea inside of dbt is progressive complexity you start out and everything's super simple and as you need it it turns out that there's features that you can take advantage of that help you scale up from there
7:52Jon Krohn:let's double click on that experience of using dbt for our listeners who haven't had it before and it's kind of interesting in a podcast format where we have to be able to have an explanation that's easy for people to understand and potentially an audio only format even if people are watching this video of us talking it's not like just visual cues as to how dbt works so this will maybe test your storytelling abilities, but I'm sure it's a narrative that's well-worn for you over the past decade of running dbt labs. Tell us about the experience of using dbt and how that might be different. Like walk through a user story or two of how it's different from SQL or how you recommend analysts getting into dbt.
8:37Jon Krohn:Just, yeah, fill our brains with what the dbt experience is like. The core idea behind dbt is not particularly new. In fact, in the first five years of the journey, we probably came across literally dozens of, I think maybe 50 plus tools that had been built like this before. So the core idea is not unique. The idea is that databases became really powerful for analytics with the launch of Amazon Redshift, and then successively with things like Google's BigQuery, and Snowflake and Databricks. And increasingly, you could rely on databases that primarily spoke SQL to do work that had previously been done by like Hadoop or Spark.
9:25And so that meant that you could express it in SQL, but data transformation or taking the raw data that comes from your systems of record, whether those are kind of business systems like Salesforce or whether these are like logging systems like Snowplow Analytics or something like that, that raw data being transformed into something that you might actually write a query against as a business user is a many, many stage process. And what you don't want to do is you don't want to have these super, super complicated monolithic transformations, but you want to kind of stage things out. You want to take your raw data and at first make sure the column names are reasonable and that you have basic data guarantees like foreign key integrity and not null constraints and these kinds of stuff.
10:17And so you clean it up just a little bit. And then maybe you start to kind of consolidate a couple of tables together. You join customers and orders so that you can get first purchase date or this kind of thing. And as you go, you maybe get to a place where you have a single table that has all of your customer data in it. But really, that customer data has to be sourced from like 50 different tables because it turns out that customers touch every system that you have. And this process in DBT is accomplished by a series of what are typically like fairly modest SQL select statements. Like ideally, you don't have any select statements in there that are over, I don't know, 100, 120 lines long.
11:03And each of them does something really discrete. And then they form a what's called a directed acyclic graph, a DAG, that goes from left to right. And it does all these subsequent stages of processing. And the magic of dbt is that it makes all of this stuff really simple to implement. Building your first table takes basically nothing. You then type dbt run and dbt manifests all of this stuff in your data platform.
11:32Jon Krohn:I gotcha. So it allows people to be doing kind of in a simple, this is probably going to be an oversimplification and you can correct me on what I get wrong here, but it's allowing people to do similar kinds of things as they might want to do in SQL, but easier and at a larger scale faster. In SQL, if you are writing a select statement, that is something you might do in your BI tool. You want to build a chart. So you write a select statement that does some group by and you get some numbers back and you put them on a chart. And that doesn't persist anywhere. But what DBT does is it actually takes all of these different tables that you have defined and it persists them in your database.
12:17And there's a lot of work involved in that. How do you do that well, et cetera. But it persists them in your database. And that means that when you go to your BI tool or when you load up your analytics agent or wherever you want to consume this data, the data that these front ends have access to has all been really nicely modeled. And it's all there in the database waiting to just be selected from.
12:43Jon Krohn:Cool. And all of this stuff is available open source, right? Mm-hmm. Mm-hmm. So listeners can right now make their way. You can very quickly, we'll obviously have a link in the show notes as well, but it's as quick as typing dbt, three characters into Google and you will make your way. What does dbt stand for, Tristan, actually? Originally, it stood for data build tool. We didn't really know what to call this thing. And so the first person to ever write a commit to dbt needed to give the repo a name. and he was just like i don't know he typed dbt on the repo and at the outset we thought it was kind of a bad name because it i don't know it didn't have any character but uh over the years the dbt community has grown huge and the logo and the brand have a lot of recognition and so we realized that we had kind of missed the boat to ever change the name and this is in fact how people thought of it.
13:45Jon Krohn:Oh yeah. No, I mean, I feel really lucky to have you on the show. To me, the DBT brand is like this kind of rockstar name. And so to have the founder and CEO of DBT Labs on the show, it's a great honor. Speaking of the community that you built, DBT Labs is built around an open source foundation and a community driven identity. Even as the company has expanded commercially through acquisitions like SDF Labs, more recently, you had an all-stock merger agreement with Fivetran that will see you take on the role of president of the combined company, according to our research. So how do you balance this vision of open source and community stewardship with practical realities of managing now kind of multiple large successful commercial organizations?
14:36Well, hopefully they're not distinct commercial organizations, I think they fit together pretty nicely. But there can be, on a kind of tactical level, tension between open source and capitalism. Open source wants things to be free, and free as in beer and free as in speech. But then capitalism wants things to be monetized. And so I think that That is a common way that people who are not deeply involved in open source see the world and they're like, I don't really get it. How does that work? In fact, companies do things all the time that they don't directly monetize. I mean, just for example, the large tech companies in the U.S.
15:26have written academic papers for, I don't know, forever, for as long as I'm aware of. The Google's technology that eventually became Hadoop, they wrote about it in a paper. Google also did a paper, some folks inside of Google wrote a paper called Attention is All You Need, which created the entire transformer revolution that we are living in today that's led to large language models. So this stuff happens all the time and there's very discrete business reasons to want to do it. The commercial justification for having people write open source software is that you get to train an entire industry of practitioners how to do their jobs.
16:15And because, like, trust me, people did not do data work 10 years ago the way that they did today. And we're a big part of how that transition happened. But then they learn how to do it on your tools that you can then, like, sell commercial versions of. So there's a real commercial justification there. But the nice thing that's almost like a side benefit is that it's just like highly motivating as humans to feel like you're making a much bigger impact than your P &L might otherwise show.
16:47Jon Krohn:Machine learning predicts and Gen.ai creates, but neither is built for complex, constrained decisions. That's where mathematical optimization comes in, giving you explainable, trustworthy decisions you can act on with confidence. Girobe is the fastest, most reliable solver organizations rely on for their high-stakes decisions. Want to see it in action? Join the 2026 Girobe Decision Intelligence Summit, September 22nd and 23rd in Las Vegas, for training, expert insights, and Gen.AI-enabled accessibility. Discover why 70 % of the world's leading enterprises trust Gurobi and start achieving optical outcomes yourself.
17:28Jon Krohn:Head to superdatascience.com slash Gurobi for the conference details. That's superdatascience.com slash G-U-R-O-B-I. Cool. Yeah, it sounds like you're striking a great balance. You're certainly enjoying a lot of success. something that we noticed in our research, something that you said a while ago, and I can't, I don't have the exact quote in front of me right now, but we read about how over the years, over this decade of growing DBT Labs into this large organization, you've had various opportunities over the years to be acquired and to basically experience a big wealth event for yourself and probably lots of other people in the business.
18:14Jon Krohn:How do you end up kind of deciding, like, it kind of reminds me of watching HBO's Silicon Valley TV show where they, you know, they're, they're just always, you're always cheering for them to stay independent, to not be acquired, to make their next big leap, um, you know, on their own. And you continue to do that despite presumably the temptation of, you know, the, you know, the capital inflow that could happen through an acquisition. You started off this conversation asking where I was calling in from. And, um, the, honestly, I think that this, this has at least something to do with the answer to this question.
18:55Uh, culturally the East coast and in, you know, specifically Philadelphia does not have the kind of, let's just say like the, the, the culture inside of Philadelphia and the culture inside of San Francisco are like wildly different. Nobody cares how much money I have here beyond some like relatively modest number. The housing prices are, you know, you can like for the prices, you can get a reasonable home in San Francisco. You can like buy a castle in the Philadelphia suburbs. And so I like, I just don't need that much money in my life and generally optimizing for it. And I started a consultancy.
19:36Like all of this is a surprising upside to me. The thing that I've chosen to optimize for over the last decade is creating an impact in the world and primarily an impact on users of our product and data analysts most especially. And so when thinking about any type of corporate transactions, the thing that I've thought about first is how would this impact the users of DBT? And the merger that you mentioned earlier with Fivetran was the first time that something hit the bar for me where I was like, oh, actually, this would be good for users of DBT. generally i think that people don't set out with a specific intent to transform data what they set out with an intent to do is like hey i have some data it's sitting over there i need to like bring it over here and do some stuff with it and so fivetran is the pipes for that and we are the meaning creation part of that and uh you know thousands and thousands of companies use these products together.
20:50I can't remember. I feel like our best estimate was something like, it was like over 10 ,000 companies used these products together already. And so it just made a tremendous amount of sense. Whereas conversations we've had like this in the past might've made sense financially, but they didn't make sense for users.
21:09Jon Krohn:Well, we all thank you for continuing to be concerned first and foremost about this community. And I think it will serve you well in the long run. I think that that kind of drive to be as open source as possible, to be community minded, to continue to grow in a way that allows everyone, all of your users to benefit is going to in the long run be great for you. Probably financially as well, but you don't need to think about it as like your primary goal. You mentioned there in your most recent response how dbt adds meaning. So the fivetrans like pipes and dbt adds meaning. There's a term that came up a lot in our research for dbt labs, which is semantic layer.
21:58Jon Krohn:Do you want to explain how dbt acts as a semantic layer for your data? This is a topic that is particularly hot right now as analytics agents are very in view. The problem that the semantic layer solves is not a problem for small organizations. So if you imagine that you're a part of a whatever, a 20-person company, a 50-person company, you probably don't need a semantic layer. But now imagine that you are Siemens, you're a global company, you have 2000 data engineers that are serving 300 ,000 employees globally, it is not possible to just know the answer to random questions that you might need to know the answer to, without getting meetings together of people that you search for.
22:55in your Outlook phone book. You get everyone together in a room and you say, how should we be measuring this thing? And there's a lot of conversation and everybody's got to figure it out and which table should we be using and all this stuff. And literally that's how big companies have for the past, whatever, 30 years, tried to answer questions like this. That's the process. And the semantic layer is tooling that allows those types of decisions to be made and then stored so that successive people, when they ask those questions, can confidently measure things in the same way twice. They don't have to reconvene the whole group.
23:44It is a technically complex problem area, but the problem that it solves is really an organizational problem. It is how do you scale knowledge to increasingly large groups of people? And that is honestly a tale as old as civilization. I mean, we don't have to go too deep on this, but like as long as people have been organizing together to like figure out how to do stuff, there's been this question of, well, how do we make sure that we know things and that everybody in this organization knows a consistent set of things? And so the semantic layer is an attempt to do that. And it's particularly relevant today because absent these types of cues, how do you measure X thing, AI agents have to try to re-derive that for themselves at every turn.
24:33And oftentimes, they do one of two things, or almost always they do one of two things. One is that they, in re-deriving how do you measure something, they just take a tremendous number of tokens to do that. They just have to look into a lot of stuff and think a lot, and that becomes slow and expensive. And then the other outcome is that they just get it wrong, or they come up with an answer that may be kind of reasonable, but it's actually not the way that your organization measures these things. So the semantic layer extends very, very nicely into a world of agents.
25:08Jon Krohn:Yeah, really cool. It seems like one of the key use cases. And if people aren't aware, that word semantic basically just means understanding, just means like the underlying meaning of some data. And so by having this common playing ground, there's this common lingua franca, this common agreement on what meaning is across data sets, across agents. There is efficiencies, especially across the large organizations that you were describing. They're like 200 ,000 person companies with 2 ,000 data engineers, that kind of thing. Yes, totally. And there are these successive layers of meaning where you start off with raw data, you go to modeled data, then you go to semantic layers, which are typically, they help you understand how to join these tables together and how to measure certain metrics.
26:06But then you can even go one step further. And this is not an area of particular expertise for me, but it's a very interesting conversation happening in the industry is ontologies. So ontologies are another layer of meaning making on top of data that are not just how do you measure a thing, which is typically how we think about the semantic layer. But it's also how do you model business processes? How do you model causation? And I think that most of us do not operate in an environment where it's appropriate to say we need ontologies for this because the world changes quickly and sometimes it's hard to keep up.
26:50But in very specific high value domains, think like genetic research or like famously, this is employed. Ontologies are employed a lot in defense. So I think about these four layers as kind of like the knowledge or meaning hierarchy.
27:08Jon Krohn:Yeah, the ontology thing is a bear and you definitely want to have a system that can be flexible or do it automatically, I suppose. for five years up until a couple of years ago, I was a co-founder and chief data scientist at a business that was working in HR tech. And in HR tech, you're constantly, there's lots of companies out there that make their whole bread and butter on creating these ontologies and updating ontologies so that you have like, okay, what are the skills that make up a data scientist? And then it's this constant, okay, so you have the skill ontology that you're needing to update on a constant basis.
27:44Jon Krohn:is like, oh, now there's PyTorch instead of just TensorFlow, and we've got to add that in. And then whole new disciplines emerge. I might argue that data science is kind of fragmented into lots of these different kinds of specialized careers like AI engineer, LLM engineer, context engineer, all of these kinds of very specific roles. And it's somehow some ontologist's job to figure out how to understand a domain so well that they can manually create these relationships between entities in that ontology, you know, skills to jobs and jobs to different job categories and difficult in a constantly evolving world.
Read the full transcript
28:26Jon Krohn:Yes. Yes. And I, you know, I, again, I'm not a super expert here, but you can imagine how a domain like genetic research could be really suitable for this type of work because, uh, it is very verifiable. It does not change. So yeah, maybe, maybe that makes a tremendous amount of sense. I, you know, trying to keep a, an ontology of the definition of a data scientist up to date seems like a very fruitless endeavor. For sure. It would be tough. It would be tough. Lots of people out there trying to do it. I would recommend you just use an LLM have the you know, that's, that was kind of my, when we would have these conversations about, okay, like, you know, I think it's time things become complex enough, our businesses become big enough, we should start maintaining our own ontologies.
29:17I was like, that's a really bad idea.
29:20Jon Krohn:You're really not going to have a good time. And yeah, it's pretty amazing. The power of, I think it was about a century ago, Ludwig Wittgenstein, this famous Austrian philosopher, he had this theory that a given word is on average, the average of the meaning of the words around it. And that kind of idea is what allows initially word vectors, word-to-vec algorithms to work, and now large language models where, you know, and yeah, it's pretty magical. And it means that we probably, you know, I think for listeners out there, avoid ontologies wherever you can and just rely on the magic of modern large language models to handle it for you.
30:08Jon Krohn:Anyway, don't need to go on ontology too long. There's another dbt, I don't know if you'd call it a product or a feature. There's something that you have called Fusion Engine that I want to make sure we don't miss talking about. So you've compared dbt's Fusion Engine to what TypeScript did for JavaScript and what React did for web development, promising to take SQL development from text manipulation toward compiler level understanding while aiming for a universal babblefish, to use a quote from you. So tell us about this Fusion Engine and how it can change analytics engineering itself. In our quest to bring software engineering best practices to data practitioners, one of the things that I think data practitioners, very much including myself, didn't even know to ask for was type safety in their language.
31:04So if most data people speak SQL, type safety is just like not a thing that SQL has ever really provided. And in fact, it's not referring to SQL as a language is kind of a non sequitur because there is no such thing as a, I don't believe that there is such a thing as a SQL independent of specific database implementation. Every database has its own dialect of SQL and they are not the same. There is a standards body that issues standards for like a baseline level of functionality in SQL, but every implementation is pretty different, not just in its kind of high level, like what's the function signature of some standard function, but like in its weird idiosyncratic implementation details such that you could run a syntactically valid SQL statement on database engine A and database engine B, and you could actually receive different answers.
32:09And that is bad for the ecosystem. I mean, if you are old enough, you can think back on how miserable it was to develop web apps in the early 2000s when browser compatibility was basically non-existent. It was a really thankless job going from this works in browser A to this works in all the other browsers. And so that ends up meaning that historically data people have just said like, I'm developing for this database. And then you have Oracle as a result, which is essentially a business that thrives because they've made it really hard to switch away from them, which is not good for anybody. What type safety does is it says, no, I really understand the language.
32:59I understand the code being written here. I understand the operations that it's being asked to perform. And I understand it at such a deep level that if I needed to, I could actually rewrite this to be executed on a different engine and be able to prove that both of these queries would produce the exact same output. And this deep understanding of the actual code that's being written allows us to do a bunch of things that software engineers are used to, but data people are not used to. I mean, as basic as when you are writing code in the VS Code extension that we build, now we can automatically highlight, you know, red squiggly underline, any syntax errors before you even run the code.
33:53And historically, the way that SQL developers got errors is they would run it, and then whatever the database said, you would then have to try to interpret that and track it down to what the specific error was. And now we'll just tell you, and tell you in real time as you type instead of execute code against the database. So that's like one small improvement. But the overall trajectory of languages in computer science is to provide more and more functionality inside the language so that developers have more and more leverage to build amazing stuff in it.
34:33Jon Krohn:Agents are getting smarter every day, but even the smartest agents get stuck without the right context and the right tools. That's where Notion comes in. With the recent launch of custom agents, Notion became the collaborative AI workspace where teams and agents work side by side. And now, their new developer platform is turning that workspace into infrastructure developers can build on. The piece I keep coming back to is how easy it is to ship something real. The CLI authenticates in one line, workers deploy without provisioning any infrastructure, you write your code, deploy, and you're done.
35:05Jon Krohn:For me, that unlocks building purpose-built tools for my custom agents with the predictability and custom logic I need. Think a guest prep agent that pulls a researcher's papers, recent talks, and citation graph on demand. Tools my agents can actually call with parallelism and predictable behavior, not just hope for. Learn more about Notion's developer platform today at notion.com slash superdata. That's all lowercase letters, notion.com slash super data to try notion's developer platform today and when you use our link you're supporting our show notion.com slash super data yeah makes a huge amount of sense and fusion the dbt labs fusion engine does sound like a big step forward it's really exciting when you talk about being able to do more i feel like i've got to get back to the agents that you were talking about maybe 10, 15 minutes ago in this conversation, because there's still so much more for us to talk about with respect to dbtlabs and the future of the world, which is agentic.
36:05Jon Krohn:I mean, I'm like, to me, there's no question if people are using cloud code, you're using open AI codecs, and you really should be because they're so magical. It's mind blowing. And on an almost daily basis, I think of new ways that I'm like, I wonder if it could do this really hard thing. and you're like, boom, Fable 5 does it. At the time of recording, yeah, go ahead. This is a complete side point. But, you know, I have many of those experiences. We've all had many of those experiences. But I always still enjoy it when I ask it to do something and it just completely falls on its face. So yesterday, I've really gotten into cycling and I was like, well, maybe I could ask it to make me what's called a GPX file, like a root that you can import into by computer.
36:55and um it i asked it a fair you know like hey make me a 20 mile route starts at my house ends at my house you know in this neighborhood and it thought about this for a good solid 15 minutes it came back to me gave me a gpx file i imported into my bike computer and it it had me like riding through people's backyards and i'm just like this is you just like completely failed to do this thing. So, uh, yes, the, the overlords are very good at some things and not that good at others. For sure.
37:28Jon Krohn:Yeah. I, you know, I would have thought it could probably figure something like that out. So that is a, yeah, it is surprising where, where it falls down and, uh, yeah, you know, it used to feel because there was so many more places where it would fall down and And it was so easy to, like a few years ago, kind of GPT-3 level intelligence, when we were only using LLMs to replace us on tasks that took a few seconds long, it was easy to experiment and find out, okay, yeah, this doesn't work. That doesn't work. We're doing pretty well over here. It's great at this. But with GPT-4 kind of level intelligence, and then we're talking about many minute-long tasks, and it's quite competent at the seconds-long tasks, it becomes harder to test.
38:13Jon Krohn:And now that we're kind of in this era of, okay, sophisticated agent harnesses, lots of double checking, gigantic models with lots of different mixture of experts, nodes being pulled in for different tasks. So we're getting, you know, these really deep levels of expertise. And it allows us to be now replacing ourselves on tasks that would take us hours in some cases. And so it becomes harder to test and experiment and find what are these kinds of hour long tasks. So it's nice to, that's a good one. I wonder how soon that'll be figured out. The question, I think behind your question is, how does, and I spend all my time thinking about this, but how does dbt fit into an agentic future?
38:57And I remember somebody asking me at an onboarding session back in something like 2018 or something. And I was talking about how dbt was really born out of the paradigm shift towards the cloud. and this person asked me, what's a paradigm shift that could make DBT less relevant? And I was like, I mean, I literally don't know the answer to that because I don't know how the future will unfold. But, you know, essentially every founder in the world, when ChatGPT first came out, had to ask the question, is my product more or less relevant in an AI dominated world? And I think the answer for DBT is actually a little bit in the middle.
39:43We are not a product that, you know, because of AI, you're going to stop using DBT. That's not how that works. But at the same time, it's not yet clear that because of AI, you're going to use a lot more of DBT. It's funny, our business has been like very, very steadily growing over the last, whatever, six years. and you don't actually see the rise of AI show up in that yet. There's a couple of early indicators, but yet. And the reason is that we work really well in this future for a couple of reasons. One is that DBT is code first. And so coding agents are quite good at writing DBT models. And so there's more than ever of them.
40:26And especially, you mentioned the Fusion engine before. Fusion is very good paired with agentic coding because it leads to these really tight iteration loops where the agent can test its own code very quickly and cheaply. And then on the other end, with the semantic layer, it is a really important piece of infrastructural input into analytics agents where people throughout an organization can ask questions and feel good that they are getting trustworthy answers back. But we're still in the early stages of deploying analytics agents through an enterprise. So I'm excited about that S-curve coming up.
41:07Jon Krohn:Yeah, I think they're going to be creating a lot more data, a lot more data pipelines. It seems to me like LLMs, agents are a good thing for dbtlabs. I think so too. Honestly, one of the things that we've been doing for as long as we've had a commercial business is we've been replacing kind of old school ETL products, whatever. There's companies that have been around for 30 years that are deep in the bowels of Fortune 500 companies. And agents have done two things here. One is that it's made it very clear that like these old school products are not going to meet modern needs, but then also it really helps with the migration.
41:58And so it has never been easier to migrate from some legacy product to a modern thing with a good long running agent.
42:07Jon Krohn:Something that you wrote about recently in a LinkedIn post related to talking about agents and using tools like CloudCode is how you describe a 12 kilobyte skill file, very specific number, but relatively, not tiny, but relatively small. It could encode hundreds of hours of collective human experience and run a dbt migration flawlessly. So do you want to tell us a bit, just in case listeners aren't aware of skill.md files in cloud code or kind of related skill files and other agents? And so, yeah, what those are and how they're such a powerful tool for a specific task like a dbt migration. I mentioned before that one of the things that we've been doing over the past 10 years is essentially teaching a group of professionals how to do their work in a different way.
42:58We called this analytics engineering. We built tooling around it. We also built coursework around it. We have certifications. We have online learning. In order to work in this way, you need to know some stuff. And it's not an infinite amount of stuff. It's, you know, you can generally go, if you know SQL and you know data, you can become pretty expert at dbt if you spend about two solid weeks, like, you know, focused on it. But skills are this like magic way. I mean, they, to me, are the closest thing we have yet seen to the matrix where you plug that giant wire into the back of your skull and you download new skills into your brain.
43:48Here, you're downloading skills into your agent's brain. So we have authored a set of DBT-related skills that are essentially like the knowledge that we were training humans on via all of our online courses and certifications and everything. And we've packaged all that up and we've distributed it to agents. And it turns out that that just works really, really well. Now, I'm sure that we're going to continue to learn and evolve that skill, but it is right out of the gate. It was very impressive and is becoming pretty widely used.
44:27Jon Krohn:Cool. And this is probably kind of a dumb, simple question, but what is a dbt migration? A dpt migration is imagine that you were using some other product, or maybe you were just like, wrote a bunch of stored procedures, or you had a bunch of spark code or whatever. And all of that was doing data modeling in your organization. You know, in a small organization, that might be a thousand tables, but inside of like a big bank, I mean, that could be 10 to 50 ,000 tables that you're talking about here. And sometimes the people that wrote all of this code, not only are they not at the organization anymore, but like some of them are retired.
45:14It can feel very high risk and certainly very timely in like time intensive to take all of that code and migrate it into DBT. And so as a result, some companies choose not to do it. Other companies who do do it, they hire consultants that take a year and multiple million dollars to do that type of work. And now you can do big, big DBT migrations in six weeks and for tens of thousands of dollars instead of millions of dollars.
45:49Jon Krohn:Really cool. Thank you, agents. Thank you, DBT. Thank you, Tristan, for building such a powerful open source product and developing this analytics engineering community around it. I promised I would let you go early. So despite it having been such an interesting episode, we should start wrapping up. So my penultimate question that I always ask my guests is if they have a book recommendation for us. You got anything? I am a big audiobook reader. Here I have, I don't know, I've been an Audible subscriber for like 20 years now. My most recent book I thought was really fascinating. It's called Musashi.
46:29It is about Miyamoto Musashi, who was, I think, the most famous Japanese samurai to ever live. And it almost feels like stories of King Arthur. It like if you've ever read the once you've once a future king, it is almost like the Japanese version of that. And so I found it really like the story was interesting, but it was more interesting. Like, what does this culture, what are the stories that it tells itself? So I found it to be a really interesting read.
47:02Jon Krohn:regular listeners will already be aware that i'm obsessed with anthropic's fable 5 model and it has taken over my working life i'm writing a technical book that includes latex files mathematical notation python code examples and fable 5 and cloud code handles requests i make across whole chapters with accompanying jupiter notebooks end-to-end work that a few short months ago would have been dozens of separate requests with way more manual fiddling required with fable 5 it just works, essentially like magic, first time. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you.
47:39Jon Krohn:Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai slash superdata. That's claude.ai slash superdata. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Claude.ai slash superdata. That is really interesting. You know, I don't spend enough time thinking about that great culture of samurais, which as a kid was somehow so fascinating. And so I just, I looked up quickly here because I, you know, I was like, you know what, I've never known the name of a single samurai.
48:25Jon Krohn:this is my first Miyamoto Musashi and yeah he wrote a book I think called The Five Rings it's very digestible if you want to become a master in the art of the two sword technique read his book it's like less than 100 pages there you go and then you two like him can have an undefeated record in 62 duels yeah exactly nice final thing before we wrap up Tristan is how can people follow you after this episode or follow dbt labs i am um a a podcaster myself i write a newsletter um you can find all of that at uh the analytics engineering roundup and uh look forward to connecting with you there nice sounds great thank you so much for taking the time out of your super busy schedule with us we really appreciate it tristan and hope to catch you again in some years to come and see how dbt labs is coming along thanks a lot there's been a lot of fun extra informative episode today with an inspiring data entrepreneur.
49:24Jon Krohn:In today's episode, Tristan Handy detailed how DBT was born from a study of about 100 companies moving to the modern data stack, why he chose SQL over Spark back in 2016 so that data analysts, those rare people who span both quantitative skills and business domain knowledge, could start simple and layer on complexity only as they need it. He talked about how the semantic layer stores and organizations agreed upon metric definitions so nobody has to reconvene a room full of people to re-decide how to measure something and how skill.md files can package a decade of dbt training courses and certifications into a few kilobytes an agent can absorb.
50:00Jon Krohn:So migrations that once took consultants a year and millions of dollars can now be done in six weeks for tens of thousands. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Tristan's social media profiles, as well as my own. at superdatascience.com slash 1021. Yes, superdatascience.com slash 1021 for episode number 1021. Thanks, of course, to everyone on the Super Data Science podcast team, our podcast manager, Sonja Breivich, media editor, Mario Pombo, our partnerships manager, Natalie Jaiske, researcher, Serge Macis, and our founder, Kirill Aromenko.
50:43Jon Krohn:Thanks to all of them for producing another excellent episode for us today, for enabling that super team to create this free podcast for you. We are deeply grateful to our sponsors. You can support the show by checking out our sponsors' links in the show notes. And if you'd ever like to sponsor an episode yourself, you can get the details on how by making your way to johncrone.com slash podcast. Otherwise, share this episode with your favorite DBT fan, review the episode wherever you listen to podcast episodes or on YouTube. if you leave a text review on Apple Podcasts, that is especially helpful for us and I'll read it on air at some point.
51:23Jon Krohn:Subscribe, obviously, if you're not already a subscriber, but most importantly, I hope you'll just keep on tuning in. I'm so grateful to have you listening and I hope I can continue to make episodes you love for years and years to come. Till next time, keep on rocking it out there and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
51:45Thank you.
From the publisher
In Episode #1021, Tristan Handy (Founder and CEO of dbt Labs) joins Jon Krohn to explain how a study of about a hundred companies in 2016 became analytics engineering, and then became a tool that over a hundred thousand data teams rely on. Tristan coined the term, chose SQL when Spark was the fashionable answer, and spent a decade turning down acquisition offers because none of them were good for the people using dbt. He is now merging dbt Labs with Fivetran and taking on the presidency of the combined company, the first deal he says cleared that bar. In this episode, Tristan walks through what dbt does to your raw data, argues that the semantic layer matters more once analytics agents are asking the questions, explains the type safety behind the Fusion engine, and details how a 12-kilobyte skill file collapses a million-dollar migration into six weeks.
Additional materials: https://www.superdatascience.com/1021
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(00:07:22) Why Tristan chose SQL over Spark, and what progressive complexity means
(00:10:44) How a dbt project turns raw data into modeled tables
(00:17:50) Why a decade of acquisition offers kept failing his one test
(00:40:47) How 12-kilobyte skill files cut year-long migrations to six weeks




