In short
Agents are changing the data stack: pipelines and visualizations can be generated from prompts, but costs and failure modes shift downstream into review, monitoring, and fixing breaks. Jordan argues data engineering becomes “managing a fleet of agents,” and data artifacts must be treated like code with PR-style review, plus a shared “context layer” (guides) so agents understand business logic and schemas.
Guest
Jordan Tigani, co-founder and CEO of MotherDuck; founding engineer on Google BigQuery; previously worked on SingleStore’s SaaS offering.
Key claims
Most real queries are small (about 90% under 100MB); older distributed warehouse designs optimize for large hardware-era workloads, while modern single-node setups fit agent workloads better. Data pipelines/visualizations should be atomic, sandboxed, and reviewed via GitHub.
Notable examples
MotherDuck’s “Watertown” (data engineer managing agent fleet); “Flights” (scheduled Python/TypeScript scripts) and “Dives” (TypeScript visualizations); schema-change auto-chasing (e.g., int to float) and “trust but verify” via evals.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Evolution of Data Engineering
0:59 to 1:48
Discussion on the transformation in data engineering and the role of AI agents.
“Now, here's my conversation with Jordan Tagani.”
Big Data Insights from BigQuery
1:48 to 2:16
Jordan shares insights from his experience at BigQuery regarding data sizes and user behavior.
“through with our agents on, they need to transform to meet that new challenge.”
Challenges of Handling Queries
2:16 to 4:28
Exploration of the challenges in handling queries and user expectations in data processing.
“You built BigQuery and one of the biggest distributed warehouses like that does data, right?”
MotherDuck's Approach to Data Tools
4:28 to 6:49
Discussion on how MotherDuck is changing the approach to data tools and instances.
“Because even those giant, giant things are expensive.”
Democratization of Data Access
6:49 to 8:07
Jordan discusses the democratization of data and its implications for users.
“evolve and change and become more capable underneath us.”
Transitioning to New Data Workflows
8:07 to 12:31
Insights on transitioning workflows with the introduction of AI and agents.
“So like understanding how to throw away something that you took for granted before and like trust a new definition of good.”
Future of Data Pipelines with AI
12:31 to 14:01
Exploration of the potential future of data pipelines in light of AI advancements.
“have like, okay, you have these, these set of tools, you would just transform, analyze, visualize.”
The Evolving Role of Data Engineers
14:01 to 16:56
Explore how AI and agents are transforming the tasks of data engineers.
“On the other side of things where you're bringing data in, you're transforming the data, that also seems highly vibe-codable.”
Empowering Non-Technical Users
16:57 to 18:46
Learn how advancements in AI tools enable non-technical users to engage with data.
“some way to sort of handle, okay, there's a bunch of tasks to do.”
The Evolution of Data Interaction
18:47 to 19:51
Understand the shift from traditional SQL queries to conversational AI interactions.
“But that becomes the new barrier, right?”
Show all 20 chapters
Personal Data Management Stories
19:52 to 24:15
Hear a personal story illustrating how accessible data tools can enhance personal projects.
“Like a personal story for me is like I whenever I work out, I use an app called Heavy and it pushes a web hook for after my workout of all of the different things of what I did.”
Building Data Pipelines with AI
24:16 to 26:54
Discover how AI can streamline creating and managing data pipelines.
“rather than having the visualizations to compute that, I wanted to sort of pre-compute for our historical data.”
Evolving Expectations in Tech
28:41 to 29:50
Discover how to adapt to evolving technology and expectations.
“Yeah, you got to take stuff away, the shape is going to change and you have to just be comfortable reconfiguring it over and over again.”
The Challenges of Data Management
29:51 to 31:25
Understand the problems of data management and collaboration in tech.
“And sometimes that does mean simplifying.”
Building Context for AI
31:26 to 35:02
Explore the need for context layers in AI data understanding.
“And, you know, it turns out it can actually start poking around and looking at things and like it can figure out like, you know, is this a is this milliseconds or seconds just by looking at it?”
Trust and Verify in AI Usage
35:03 to 37:18
Learn about the balance of trust and verification in AI applications.
“And you can detect whether the LLM or your agents can, you know, are generating that right number.”
Navigating AI Costs and Productivity
37:19 to 42:00
Discuss the implications of AI costs and productivity in engineering.
“where agents and humans can understand them.”
The Challenges of Scaling Engineering Practices
42:00 to 43:14
Explore the complexities of scaling engineering practices and the financial implications involved.
“out of some engineering orgs and sometimes when you zoom in it's one team or one engineer or It's one swarm of agents and it's something that's cracked the code.”
The Future of Human-Agent Collaboration
43:14 to 43:56
Discuss the evolving relationship between humans and agents in tech environments and the tools facilitating this collaboration.
“And we have to have better guardrails and understanding of like how those things are working.”
Staying Connected with Mother Duck
43:56 to 44:31
Learn how to follow and engage with Mother Duck and its community resources.
“And it's been really cool to kind of dive into that.”
Transcript
Automatic transcript. May contain errors.0:02Welcome back to Dev Interrupted, brought to you by Linear B. My guest today is Jordan Tigani, co-founder and CEO of MotherDuck and founding engineer on Google BigQuery. He used to write more SQL than anyone at his company, and now, like so many of us, he can't remember the last time he wrote a query. We covered Steve Yege's Gastown here on the show, and Jordan wrote the data version. He calls it Watertown, where a data engineer manages a fleet of agents. Those agents write the pipelines and the visualizations from a prompt, and Jordan's point is that the cost lands downstream, in review and in catching what breaks.
0:41His answers that treat all of it like code, which adds to the review bottleneck, a key problem we're set to solve here at Linear B, with tooling like AI code review and merge automations. Learn how the mother ducking genius behind Mother Duck thinks about the data equivalents for his own gates. Now, here's my conversation with Jordan Tagani. But I'm really excited to have this chat with you, Jordan. You know, you're the CEO and co-founder of a pretty cool company right now, in my opinion. There's a lot of really innovative things coming out of Mother Duck. And your background as a founding engineer at Google BigQuery is, you know, just like a driving force, I think, in all of this change in narrative.
1:21And it makes so many folks look to you to understand like how data science is transforming. So we're really excited to have this conversation today. Talk about some of like the ancient history, even back in like the early 2020s. I know that's like a long, long, long time ago. But even right now, how Mother Duck, I think, sits at the center of an argument that the practitioners of the data stack are AI agents now. and data scientists and the way that we work with data and the tools that we use and communicate through with our agents on, they need to transform to meet that new challenge. So we're super excited to dive into all of this today.
1:56Jordan, welcome to Dev Interrupted. Yeah, thanks. Thanks for being here. These are exciting times and it's good to be in the middle of things. It really is. And you couldn't really be more in the middle of things. I want to back up for a second and talk about you know, Mother Duck and how we got to be here because your background was in big data. You built BigQuery and one of the biggest distributed warehouses like that does data, right? And so we're talking about a fundamentally different approach from how Mother Duck works now, which I want to dive into with you, no pun intended, but also understand like how that change came to be in your mind.
2:34Like what did you understand about what big data was and what it couldn't be and what it needed to change into that, you know, you were maybe one step ahead of on the insights. So, you know, we used, when I worked on BigQuery, we often used BigQuery to understand what was happening in BigQuery. And I remember doing some like analysis, you know, trying to figure out, okay, well, what, if somebody was using Snowflake, what size Snowflake instance would they, would they need? And so that we could kind of do some comparisons because the way BigQuery worked under the hood was very different than Snowflake.
3:06And so I was looking at kind of the size of data that people had and the size of queries that people were running. And the thing I found was it was dramatically less than I had expected. And something like 90 % of queries were under 100 megabytes. And megabytes is something that at Google people sneeze at. Actually, they don't even bother to sneeze at it. They sneeze at terabytes. But like, you know, and just something you wouldn't think you'd need, you'd need sort of a massively parallel engine to do. And like kind of as I, you know, I was looking like, wow, like most of our customers are actually using tiny data.
3:45And then even the even the giant customers like, you know, we had like HSBC and Home Depot and Walmart and like Equifax, like some of the biggest companies in the world, like the stuff that they were doing wasn't that big. And I think internally, there's a lot of energy around, okay, well, how do we make bigger stuff work and make it faster to do tens of terabytes and hundreds of terabyte queries? But the things that people were complaining about were the things that their 100 megabyte query wasn't any faster than their 100 gigabyte query. And so actually, the things that people cared about was like, hey, I've got a human waiting for this result.
4:26Like if you add a second to every query, you know, because of all this distributed mechanism, then that's actually going to make the experience of your users worse than being able to handle some giant, giant thing faster. Because even those giant, giant things are expensive. And so people tend to not do them very often. So kind of like I filed that, you know, away and it took me a couple of years. you know, actually I worked at Single Store for a couple of years working on their SaaS offering. And people kept asking for smaller instances. And we, you know, we had one that was half the size of Snowflake and the Snowflake smallest instance.
5:07And people were like, what if we did smaller one? And so it was like, it seemed like there really was a need for, you know, scale to zero serverless kind of smaller instances as well as large. And then the other thing that, you know, you should realize is that, you know, these, a lot of these technologies that were, we kind of built the data infrastructure around were designed when machines were small. Like, you know, if you had a hundred gigabyte query, well, a hundred gigabyte wasn't going to fit in memory. It wasn't going to like, like you needed, you know, if you wanted a hundred cores, like, you know, you, you had to, you had to combine a whole bunch of machines to get a hundred cores.
5:49And, and nowadays is like the just basic machines that are on ec2 like the physical machines are you know um hundreds of cores terabytes of ram and so like you don't actually need these distributed distributed systems and if you don't use the distributed systems you can just do things you know much more easily you can be faster lower latency like a whole bunch of a whole bunch of important important things. And so, yeah, it was like, hey, if we're going to do something, if you were going to build these things now, you would do them differently. And to me, that was an opportunity. There's a really interesting, I think, parallel in there to unpack of the idea that you even said the hardware was smaller, it was older, and you couldn't even get some of those mid-sized queries, those things could live in the RAM of the machine.
6:42It was always thought that we were moving towards having these bigger monolithic kind of like services that could handle all of that for us because it was just so hard to really visualize and understand how quickly the tools and the technology we were using were going to evolve and change and become more capable underneath us. And so like we make bets on like what we want to build and what we need for the future. But then sometimes we have to make the bets to take things away and to simplify them. And there's a lot of things I think we're learning about that even right now with, you know, how everyone's workflows are transforming with agents.
7:15You know, there's a lot of temptation to add a lot of complexity, a lot of stuff on top because you think you can orchestrate it in this kind of way. But there's a much bigger opportunity sometimes in taking away stuff you took for granted or that isn't necessary anymore. And you can only discover by starting to hack at it, you know. And so there's like, I think, a lot of lessons there that have parallels about how we work with technology and how our expectations change. And it's really important to call out, too, that data, too, became more widespread and just more accessible to more people. Like it was it kind of moved out of the domain of just being solely something that the data scientists or the data researcher would do at scale on these huge amounts of data.
7:56It became something that was more approachable and embedded into more tools and delivered reports and all sorts of ways to where working with data became more democratized. And because of that, we needed smaller, more customizable bite size on your device ways of working with that data to meet the demands. Right. So like understanding how to throw away something that you took for granted before and like trust a new definition of good. Like part of doing that, too, is understanding what you think the next level of good will look like. So for you, was that a vision of the access of these kinds of tools just being more widespread?
8:37And what are the opportunities that you still see for the tool to continue to evolve for now? It's like more agentic consumers. You know, I think a lot of things in technology, you know, you can have two things that solve a similar problem. They have a different architecture. They're going to be able to grow in different ways. And I think that kind of some of these older systems, the ways that they can grow are reasonably constrained and they move very slowly. I mean, it's just, I remember some of the optimizations that we worked on in BigQuery would just take months and months for something relatively simple.
9:15And, you know, working on this single node system, you know, we can just, those things can be done in, you know, in a week or weekend. And you can just move much, much faster. And so because kind of the system has been shaped a little bit differently, the way it can expand and extend and react to new changes is different. And I think, you know, you mentioned sort of, you know, all the changes that are going on with AI. I think if you have these sort of giant warehouses or distributed warehouses, just from a shape perspective, they feel the right shape for AI agents. We might have a lot of agents that are sort of hammering on different things and like, and having a single node system where you basically, you know, one, one warehouse, one agent, you know, for as many agents as you, as you have, like those, that seems to sort of fit the problem better than, than some of these, these older, older systems.
10:20Yeah, absolutely. I feel like we're trying to make all of the bits that an agent touches as atomic as possible so that we can isolate them and put them in these like more sandbox things. The idea of having agents at scale and at real time, all touching a same kind of data space is I think scary for most people who would want to protect their data. So I do think that becomes like even like a strategic wedge and like how the tool is adopted and who its consumers are. And I want to talk a little bit about who those new users are, because, you know, it is the agents, but it is also still the engineer, the data scientist, the data engineer, you know, figuring out how to orchestrate them.
11:01And that's where Mother Duck, I think, is really leading in terms of teaching the future of how the data scientists and the data engineers are going to work with their tools. Like, we really loved your entire coverage on Watertown. We've talked a lot about Gastown here on the show, and we've covered it since the top of the year. All of the many things from Steve Yage and how those different narratives have evolved. We've actually adopted a lot of them here on our engineering team and on the show. And so like when they see the same parallel applied towards data engineers and like the ways that they work with their data and how it will transform, I think was really powerful and really smart, too, because it gave it gave us language to talk about the levels of AI fluency as you got better using those tools.
11:44I wanted to dive into your head again, no pun intended about that and learn a little bit about how long it took to develop that idea after thinking about, you know, what Gastown was doing for engineering and like what, how, how innately did all of that come to you? So I think that we're in, you know, we're in this like transitional moment right now. Like I think I had a college professor and, you know, who like he was the guy who invented punctuate or named punctuated equilibrium, you know, in evolution. And I feel like software technology, like you often have something similar with meeting, like you're sort of at this equilibrium state and then some change will happen.
12:27And then there's sort of rapid transition to a new equilibrium. And so I think we had the modern data stack was this sort of equilibrium state where you have like, okay, you have these, these set of tools, you would just transform, analyze, visualize. And those are each different, different companies. And everybody had their swim names, swim lanes, and everybody was happy. And then along comes like AI and agents. And all of a sudden, like that, you know, the world is going to change. And so trying to predict exactly the way that it's going to change is sort of, it's a recipe for, you know, spectacularly bad predictions, but it doesn't mean, you know, you shouldn't, you shouldn't try.
13:09So, you know, it's also when it's sort of super exciting because that's when the biggest, the biggest opportunities happen when you are in these sort of transitional times. So we started to notice that like, hey, text-to-SQL is getting, you know, it works. It's not necessarily the way we thought it was going to work, you know, but, you know, with an agent, you basically can get very high fidelity. So it's like, okay, so that's happening. And then you realize, okay, this Claude and that agents can really do very good visualizations. and those are only going to get better. So if you can run a query and then you can visualize it, that's a lot of what you have in a BI tool.
13:53So it's like, okay, well, BI, there's really an opportunity to just sort of change how you're doing BI. On the other side of things where you're bringing data in, you're transforming the data, that also seems highly vibe-codable. If you ask Claude, if you point Claude at Salesforce force and say like, Hey, you know, pull my data in. It's going to be able to figure out how to do that. And, and it's going to get better. It's going to, you know, like the, the, the, how long it takes it to do today. It's less than it'll take in the future. And so you kind of think about the data stack and what people are doing and what data engineers do and what data analysts do and sort of like, okay, how does that job change in, you know, when building pipelines can be done from a prompt building, you know, visualizations can be done from a prompt, you know, transformations can be done from a prompt.
14:49Well, what, what you do that, what are the things that are going to go wrong? And, you know, I think, you know, well, people are, you know, the agent's going to build a bad, bad schema, or they're not going to have the right context or they're not going to have the right, you know, then there becomes like this, these other jobs that, you know, humans are going to have to do to, but they're almost like, you know, I mean, just like, you know, they say software engineers become managers of agents. Like I think a data engineer is really going to be sort of a manager of a fleet of agents that are going to be doing, you know, doing data work.
15:25And really the things that they have to do are going to be the, you know, reacting to changes. And I think one of the ways that data engineering is different than software engineering is that it's so much easier for things to break through no fault of your own. It's like, because you have this data that's coming in, data may be changing. You know, there's schema changes or, you know, something may be stalled upstream, or there may be a bug upstream or the format or the distribution of the data changes. And that, you know, changes how you have to visualize it. There's just all these things that can, sort of go wrong, even if you've built the right pipeline, that I think humans are going to have to, you know, somebody's going to have to handle that.
16:09And then you think, well, okay, well, but can some of that be done automatically? It's like, hey, if the schema changed from a, you know, big int to a float, you know, can Claude, you know, go and chase that through, or if a field gets renamed, can Claude chase that through all of your visualizations? And chances are probably yes, Um, so then you, you think about, okay, what do those agents look like? And anyway, so that was sort of the idea behind, you know, I, I called it Watertown after, after, you know, Steve Yaghi's, uh, you know, gas, you know, gas towns a little bit, a little bit tongue in cheek.
16:43And so exactly what are those agents going to be and what are they going to do? And what are the jobs going to be? Who knows? I mean, just sort of like, I think when, you know, in gas town, you know, is, do you need the deacon and the, like all of these specific roles? Like, you know, maybe not, you know, maybe, you know, but I think you're going to need some way to sort of handle, okay, there's a bunch of tasks to do. There's a queue. There's somebody, you know, and then there's like a mechanism to surface things to humans as well. And then again, exciting times, stuff is changing, stuff is moving, stuff is moving quickly.
17:20But I think we're already starting to see some of these things, you know, some of these things happen and something's changed. Like I can't remember the last time I've written SQL and I used to write the most SQL of anybody at the company just because I love to poke at stuff and to sort of be like, okay, well, this is happening, but like what's actually happening? And, you know, but now it's just a, you know, Claude session and so I think that's going to be happening through the, you know, through the rest of the data stack. And it's exciting. and it's not just, I think it's not just like going to make people less, you know, useful.
18:00Like I think you're still going to need people, but it also brings, you know, we call them at Mother Duck, we call them NTDs for non-technical ducks, which is like the people who would otherwise feel like they can't do this, this stuff. They're like, I'm not smart enough. I'm not good enough. I don't have the knowledge or the background to be able to, to do these things, to be able to, answer these questions about what's going on in the business or their sales book or their, you know, their marketing campaigns. And now all of a sudden they can do those. And I think that's really powerful and liberating.
18:34Yeah, exactly. It goes again to like the whole democratizing of the tool. Now all they need to do is know what to ask, which is its own challenge, by the way, understanding what you're looking for and what the query needs to be, what you're really asking for. But that becomes the new barrier, right? Which is exciting because text to SQL, like you said, just really transformed. It's gotten amazing and it's gotten a lot better. Like I remember back in October of last year, I was at a hackathon. It was a data hackathon actually for data agents. Brian Bischoff of Theory Ventures hosted this really wacky hackathon called America's Next Top Modeler, where we had to basically decide or figure out how we wanted to sort over like 10 ,000 unsorted documents and park it files and all sorts of just like unstructured and unstructured data for a fictional company that had several mergers.
19:28And then like adding cruelty to it, he even had like a dusty binder with like old things in it that can contradicted what was in our digital files. And we'd have to consult it and like the real world. So it was like mind boggling. And as part of that, I was struggling with, I remember back then with, with Claude trying to get some basic interactions with SQL where I was comfortable with, where I was like able to actually understand what was going on in the data. And now I just feel like whenever I do work with, you know, modern models to work with data, it's really, really simple and conversational.
20:01Like a personal story for me is like I whenever I work out, I use an app called Heavy and it pushes a web hook for after my workout of all of the different things of what I did. And I actually have that hit a database where it might have an agent that then just looks over it and helps give me like a daily understanding of my workout. and how I'm trending. And it even gives me nudges of like, hey, you haven't like increased that weight in a while. I see you're like, you know, holding back or something. And I've found that really fun and interesting to experiment with the data. But that's just something that's become really accessible to me is like, I guess, a non-technical duck or a non-data engineered duck.
20:39You know, I'm like technical, but I'm not of the data science world. But before I would have never really even thought about trying to do that. But as agents get better at making those SQL queries, like all of a sudden now you just have a rise and these huge or new or ad hoc SQL queries, you get the needs for having these like pipelines that can run it and create it. You know, I understand that y 'all are tackling that problem too with flights. And that's what pipelines are where agents build and they schedule and deploy things because now like you, they can build up that big SQL query and push it.
21:13Right. So like you're understanding like the platform that they need to do their distributed work on. I have a question for you, though, in that world where you have agents deciding what SQL to run and then to put it up into a pipeline to run it. Where do the new gates and the checks fall? Because agents can make so many queries or even mutate the data. And so in engineering and in the Linear B world where we talk about the bottlenecks and the places where the humans put down the gates, that's like code review. That's like a PR, right? But like, there's no PR for like a data or there's not one that's like immediately easy to visualize.
21:50So like, what is that? What is that like on your side? I think at some, to some extent, these things have to be code. These have to be like, and treated like code. I mean, obviously they're code because they, you know, they run, they run stuff, but like, but they have to be treated like code and they have to be, you know, I think there's going to be, you know, review processes and, you know, check, you know, people can submit PRs against, against them. Like, you know, whether it's going to be, you know, AI agents or humans, you know, writing the PRs or reviewing the PRs, like that's certainly, you know, an open question and things are going to be, you know, and that part is going to be changing.
22:31You mentioned that we built this thing called flights at MotherDuck. And so what flights are is really they're just a Python script. And we'll be adding Node as well, but TypeScript that runs on a schedule. And we have some hints around here's how to do good things with MotherDuck. Here's how to connect to certain data sources. These sort of set of guides that make that easy. We set our templates that basically the AI will be able to pick the template or be able to write these scripts. But at the end of the day, it's Python scripts. And I mentioned when we saw that a lot of the data pipelines and data ingestions is sort of heavily, highly bi-codable.
23:18We said, well, how do we – our goal really is to make it easy for people to get data in. because people would start using MotherDuck and we say to them, okay, well, what you need to do is you need to go and sign up for Fivetran and set MotherDuck as your destination and pull in a bunch of data and then you can start using MotherDuck. That's a lot of steps. There's a lot of things you have to get right. And so one of the things this lets us do is we just let you sort of start from in medias res and everybody's watching the Odyssey. It starts in the middle of things. The middle of things is, hey, I have a thing I want to do with my data.
23:58And then, okay, well, then we're going to go back and we're going to figure out where that data is and we're going to figure out how to get it in. I just did that this morning, actually. I'm working on a new workload intensity metric with Claw trying to figure it out. It's quite a complicated query. And so in order to get, you know, rather than having the visualizations to compute that, I wanted to sort of pre-compute for our historical data. And so, you know, it basically kicked off one of these flights and creates one of these flights and did a backfill. And that also is going to be running every hour.
24:33So they'll be able to sort of keep that up to date. But I didn't start out to sort of build a data pipeline. The data pipeline was sort of pulled in by, OK, well, there's this thing that I need to do. this like there's maybe this data that's missing or this is like this computation i i uh i need to do that you know uses some external data or is uh is uh is highly intensive and i think that tends to be how people how people work is that you know the the the end goal isn't the data pipeline the data pipeline is sort of this necessary thing in order to achieve this other this other goal you know i think getting back to the you know getting back to the gates i do think that you know so do visual on the visualization sides we have this you know you know our visualization tool is called dives which is which is similar it's basically you know it's a TypeScript thing that can be created by by you know your favorite LLM you know I use Claude as sort of the generic LLM monitor so you use Claude create this TypeScript file which is a data visualization and Claude is really good at it as you know as a gemini you know and um etc and so you kind of have these like two things you have this this these dives which are type script visualizations you have these uh flights which are python uh you know python scripts you know but those are both code those are both you know things that like you know the way we use them internally is we have a github github repo and we um we set prs against the github repo and then we basically a commit hook that like that synchronizes that to MotherDuck.
26:10And so for us, GitHub is the source of truth. But then that makes sure that we can do reviews. Anybody can see what they are. We can clone them. We can tweak them. But then also those get pushed into the tool so that it's easy to use from the tool. And I feel like that's a good sort of stable way of sort of making that stuff work. It's all code and, you know, treated like code. And so you can use the same mechanisms that you would use for anything, which are, you know, also being geared towards, you know, agents and AI and those things are getting better. I think just the last thing I'll say on this topic, which I realized being a little long-winded, one of the things that we're trying to do is make it so that as these agents and these models get better, the features also get better.
27:11And that which is tricky because often people work on something and as the agent gets better or the LL gets better, it's like, oh, I don't even need to do that. Like that whole thing goes away. But I think if you can build things in the right shape, then as the models get better, these things just get better. So for example, the visualizations are dives. Like they just get better as Claude gets better at building visualizations and gets better at writing SQL. And the flights, the same thing is, you know, get better at like pulling in data from different sources, you know, building, writing Python.
27:44Like those just get better. And synchronizing it to GitHub, like, well, you know, I would get up the whole like developer tools are only going to get better at working with these agents. And it can be tempting to just like, oh, well, the model doesn't do this. And so I'm going to fill this gap. But, you know, chances are the model is going to, as it improves, that gap is going to go away pretty quickly. Your AI software factory is shipping more code, but is it actually delivering more value? Linear B's new guide, Your Software Factory Needs a Context Layer, shows you how to find out. It breaks down why AI adoption alone isn't enough, how to measure effective PR yields, and why cost per effective PR may be the metric that your executives want right now.
Read the full transcript
28:31You'll also get practical ways to connect source control, issues, CI, deployments, and AI signals into one feedback loop. Download the free guide from Linear B today. Yeah, you got to take stuff away, the shape is going to change and you have to just be comfortable reconfiguring it over and over again. It goes back to like the beginning of like, just the paradigm of like, expectations change, the technology evolves really quickly faster than you think. And so you have to be comfortable with, you know, changing and evolving that way. And there's a lot of smart things in there to unpack, like how ultimately boiling it down to a GitForge kind of source of truth that's shared, and everyone can have visualizations on this also allows you to take advantage of those same kinds of gates and checks like AI review on the types of queries and things that are getting shared out with folks and you get the collaboration, right?
29:19And that's a really important part of creating and sharing the data too, because I think the temptation in now is like with all of these capabilities and flights is like certainly the right shape because it's compatible with the atomic idea of the small and the versatile and the sandbox, right? But it also allows for the sudden scaling of the thing you need. Like when you said you didn't set out to make a pipeline, you just got a pipeline for free just by nature of how the tooling is now shaped. And so that becomes like the new opportunities to look for. And sometimes that does mean simplifying.
29:53Another part of this too that keeps bouncing around in my head is when you have a lot of people with a lot of data needs and abilities and you have a lot of agents that are doing these things at scale, it becomes really easy for people to just reinvent things over and over again for collaboration or communication to break down, for people to reinvent the same stuff. And that's not economical for a company if they have 10 really agentic engineers and they're all reinventing the 10 same things. So what do you think has to evolve to help those organizations actually distribute the gains of everyone being able to work that way?
30:30What has to evolve? I think that's a good question. There's sort of the one way of thinking about it is a content management problem. Like how do you make sure that if somebody has done something that becomes visible to other people so that they don't, you know, they can either, they can either use it directly, they can, they can, or they can riff off of it versus like, versus generating it directly. And so we are building something, you know, but building that stuff into, uh, into mother duck. Then there's the other, the whole idea of, uh, of context and, um, you know, context meaning, you know, in the, you know, the data world, like everybody's got a definition, different definition for how they compute revenue.
31:14Like different companies have, you know, like different ways of names for their regions or like they, you know, there's just a bunch of, there's a bunch of sort of business logic that is, you know, tends to be in people's hands or tends to be at a dock somewhere. And I think one of the reasons that people have been skeptical of, of sort of some of the AI and data is because they're like, well, how are they going to, how's the AI going to learn about, you know, the stuff that only I know about my data. And, you know, it turns out it can actually start poking around and looking at things and like it can figure out like, you know, is this a is this milliseconds or seconds just by looking at it?
31:53And it's like, is this reasonable the same way a human would think that it's reasonable? But I think there's also a need for a context layer, semantic layer, something that teaches the LLM about your business, about the specific schemas and things that you care about. Sometimes it isn't necessary, but it's a way of creating shortcuts. You know, one thing you find is that like, you know, tokens are expensive. And so while the, you know, maybe, maybe an LLM can, can figure out, you know, what your fiscal quarter is, you know, like if it doesn't have to figure that out, if you just tell it, like it's, it's a way of getting, you know, getting to the answers you want much more quickly and less, and less expensively.
32:42So we just want actually today, yesterday, we call it guides, which is all they are is sort of markdown documents that you can create that describe, you know, aspects of, you know, your data, your database, your organization, your schemas. They can be, they can even be like style guides. Like they just sort of describe how this is, these are kind of like skills for your data. And it's pretty simple. It took us a long time to build because we were starting out with something much more complicated. We wanted, like my belief was it should be sort of self-driving. You should be able to, we should be able to glean this from what individual users are doing and then be able to combine them across users and be able to sort of like, and that turns out to be, you know, very hard to do well.
33:36and turns out what people actually just wanted was they oh i just i know what i want can i just write it and like and so like now we're letting you write you know write these documents i think the the next step of that is the uh is is sort of like okay do you want like because i think actually you can get very very far from just a markdown doc that describes that describes something yeah maybe you have little sequel snippets in it that describe how to do do some some computation but um you know llms are very very good at understanding english um and or understanding whatever you know other human natural languages uh so i think that they're going to continue to get better at you know like they're going to know that better than they are going to know like some you know metric flow or some some cement modeling language and uh and the the additional rigor involved just makes it harder to write.
34:29It doesn't actually make it, make the LLF do a better job. There is the argument that I've heard made that, well, the reason you need semantic modeling, semantic modeling meaning like actually something that effectively enforces every time we do this calculation, we do this calculation exactly the same way. and I think that those I don't know whether those are really going to be needed I think as the models get better that's going to be less important if you describe it in English or in some way of being clear about it I may be wrong but I think the other way of doing this is essentially the sort of trust but verify where you describe it and then you run evals I think evals are important in this world of agents and AI, where you say like, well, when I say, you know, tell me about the revenue in February, you know, 2025, the number should be this, like there is a right answer.
35:40And you can detect whether the LLM or your agents can, you know, are generating that right number. and you can use that versus having this sort of more fancy semantic model. But that's my belief. But I've been wrong about a lot of these things before and we'll see what happens with that one. Well, trust and verify is definitely the dual approach that's really important. I actually gave a talk called that recently at the Checkmarks AppSec subnet about that same as I think for like engineers that are using agents to produce code of like you can trust and understand you can have these guardrails, but then you also need to have these verification systems in place to look at the other end.
36:28It's about like measuring and understanding the inputs, but then also measuring and understanding those outputs on the other end as well. And I think that's been a big challenge for engineering teams. It's just like in the last year, there's been like a big mandate, just use AI, if I'm just possible and just like, you know, use your tokens, pick up tools, we'll try every tool on the market. And then there's been a lot of shifts recently with people trying to pull back on their inference budgets and companies burning through all of their tokens that are available to them just in the first few months of the year.
36:57And it speaks to an inability to pick the right tool for the right problem. Engineers maybe are, you know, they always want to use the best for everything. And so there's still a challenge ahead of us as like engineering leaders of optimizing and those costs and reducing those kind of like duplicated work. And I think that's going to be a big challenge. And it sounds like guides are like one step for kind of putting that those kinds of roadmaps in a shared place where agents and humans can understand them. There's also things like understanding, like just to get a number instead of having to crunch it every time and us having to really get the agents into that kind of motion because they are just, you know, apt to crunch.
37:41They love to use their tools. And so there's another part of this too about like in that world where people are just using a lot more tokens. We've been talking even about like the token maxing leaderboards and stuff. Like what are the signals that you look for as an engineering leader that tells you that like your AI usage and your AI productivity is actually giving you value to your engineering team or to your data sciences? I think that's an unsolved question. You're right. Everybody wants to use the latest model. It's just like, if you've got a Ferrari in your garage, you're going to want to drive up.
38:19Maybe you won't want to drive it in traffic, but you're going to want to use it. And if you've got Fable available to you and it's going to get things right more than you know, Gemini Flash, then you're going to want to use that instead. And so, so far we have not put any sort of limits on people. I mean, we do have per user limits in Claude that we just, we sort of, when people run into them, we increase them and we just sort of want to have some sort of, you know, visibility into what's happening. And the goal isn't to sort of try to restrict what people are doing. The goal is just it can cost money and we want to prevent a runaway Yeah, you don't want a runaway bill.
39:07Also, it helps you and the employee have the check-in about the AI usage and what you're doing with it. And it helps you know what are you using the tokens for. We'd love to, let's talk about that productivity. It's an opportunity even to share it with others. I think that's actually a really strategic idea for distributing. you know we have like ability for you know one of our engineers like i mean i just you know said hey can you bump my club code again and i was like and it was it was a thousand dollars thousand dollars a month and uh but it's like you know this person is incredibly productive engineer like one of our top top people like making him you know 10 more productive is certainly worth um worth a thousand dollars a month um or you know i we bumped it to something higher um And so that was an easy call to make.
39:55At some point, it won't be, though. At some point, I'm nervous for those days where it's sort of like it starts to get – people are spending – you were talking about Gastown. People are running dozens of agents and they're just like – yeah, if you're running dozens of agents and you have – if you're running frontier models, then yeah, that's probably going to be more than you – are going to be able to or want to or want to spend. When that world arrives, I don't know how we're going to deal with it. I've heard other founders talk about like, oh yeah, when you sign up, when you take a new job, part of your, you're going to get a token budget.
40:41A token budget. Yeah. And that's going to decide whether or not you're going to take the job because it's like this this company lets me drive a Ferrari to work and this, this one, I have to drive, you know, Ikea. Those are different, different experiences. And I transition, transitional states are, are, are exciting because I know we're certainly in one for now. Do you think this is one of those where it'll be like a quick equilibrium, like where it gets up to that new equilibrium, like going back to what you said earlier? I think, I mean, there will, there certainly will be an equilibrium about like, okay, this is just the way people do software engineering.
41:16And like, cause, cause right now, like the token budgets and the token usage and that AI usage and agent usage is like, it's just sort of rocketing, rocketing up. And we're just sort of figuring out what are the levers to, you know, levers to constrain it. And when does it make sense to do it? And when does it make sense to not do it? And it does impact how many engineers you can hire and impacts like, you know, a whole bunch of things. so I think we will need some more time to get to the next equilibrium but I think a lot of things are going to break before then which is you know absolutely it's like there's so much that buckles right now I think in modern engineering orgs underneath like the workflows and honestly just the raw outputs that are coming out of some engineering orgs and sometimes when you zoom in it's one team or one engineer or It's one swarm of agents and it's something that's cracked the code.
42:11And then the leadership and the teams, they want to emulate and replicate this. But then it quickly gets into like a huge ballooning amount of cost. And then there's a huge drive to prove the value. You know, the CFO is in these rooms now for all of these engineering teams because those token budgets are real. They're parts of compensation packages at some of like the biggest tech companies now. And you're so right that modern engineers, like they want to work where they have the most what they consider like intellectual commoditization available to them to do their job at scale, because that's that is how engineering is done.
42:48I think it's still a lot for us to learn. And there's definitely going to be more developments on the scene for sure. I think that in this conversation, we've talked a lot about how our expectations of how technology is shifts really dramatically underneath us. And we have to be challenged to create new versions and reinvent things that we took advantage of or took for granted the day before. And I think that we're in a world now where we have to manage the outputs of things that are running atomically and at scale and 24-7. And we have to have better guardrails and understanding of like how those things are working.
43:28But also too, critically, we need a fluency layer of data and intent between workers and humans and their agents. And, you know, things like Mother Duck, things like Duck TV, those are part of that language, right? They allow the agent and the human to work more closely together and in a way that leaves artifacts that are shareable and can be distributed and everyone can benefit from. So I think that's like the future of how people will work with their knowledge. And it's been really cool to kind of dive into that. Again, no pun intended with you, Jordan, today. But as we wrap up, I want to know, where can our listeners go to keep up with you and Mother Duck and everything that's going on in your world?
44:12I think probably the best place is follow Mother Duck on LinkedIn or me. That's probably where we're most active. And we do have a Mother Duck blog where we talk about all things Mother Duck. We have a Mother Duck community Slack for people who are using community Slack. There's also DuckDB Discord if you're a DuckDB fan. That's also quite an active place. Amazing. We'll get all of those links in our show notes. And to you listening, if you've made it this far, then you're obviously a data nerd and you obviously love today's conversation. So please give it a like or a comment wherever you're listening or watching this and join us on Substack or LinkedIn as well to read the whole newsletter that's accompanying this episode as it comes out.
44:57And if you have any thoughts about today's episode or the things that Jordan and I talked about, come find us on LinkedIn and let us know. You can drop us a comment. We'd love to continue the conversation with you there. and Jordan, thanks again for coming on the show. It was a pleasure talking to you. Thank you. This was fun.
From the publisher
Your data warehouse still thinks a human is on the other end of the query. That's a problem MotherDuck CEO and co-founder Jordan Tigani knows from the inside, having spent years as a founding engineer on Google BigQuery. This week on Dev Interrupted, he joins Andrew to make the case that one warehouse per agent beats the distributed systems he used to work on. He lays out Watertown, his riff on Steve Yegge's Gastown that recasts the data engineer as the manager of a fleet of agents, and walks through how MotherDuck's Flights, Dives, and just-launched Guides keep pipelines, visualizations, and business context as code reviewed in GitHub.
Get the guide: Your software factory needs a context layer
Follow the show:
- Subscribe to our Substack
- Follow us on LinkedIn
- Subscribe to our YouTube Channel
Follow the hosts:
Follow today's guest:
- MotherDuck: The DuckDB-powered data warehouse built for small data and agents at motherduck.com
- Water-Town: Read Jordan's take on the agent swarm data stack, his riff on Gastown, at motherduck.com/blog
- Flights: Learn how MotherDuck turns data ingest into scheduled Python scripts your agents can write at motherduck.com/blog
- Dives: See how MotherDuck replaced its BI tool with LLM-written TypeScript visualizations at motherduck.com/blog
- Guides: Read the docs on MotherDuck's new Markdown context layer for agents at motherduck.com/docs
- MotherDuck Blog: Keep up with all things MotherDuck at motherduck.com/blog
- MotherDuck Community: Join the community Slack at community.motherduck.com
- Connect with Jordan: LinkedIn | X
OFFERS
- Start Free Trial: Get started with LinearB's AI productivity platform for free.
- Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.
LEARN ABOUT LINEARB
- AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
- AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
- AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
- MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.
