The Father of Data Contracts - AI’s New Best Friend

10 Dec 2024 · 36 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Talking AI - The Father of Data Contracts - AI’s New Best Friend

Episode Overview In this episode of the Talking AI podcast, host Matt Paige and co-host Omar Shanti speak with Andrew Jones, Principal Engineer at GoCardless. The conversation centers on the concept and practice of data contracts, exploring their significance in improving data reliability, governance, and collaboration within organizations.

Key Concepts

What are Data Contracts?

  • Definition: Data contracts are agreements that specify the terms around the use, quality, and handling of data between producers and consumers.
  • Purpose: To ensure that data pipelines deliver consistent, trustworthy insights without disruptions.

Emergence of Data Contracts

  • Origin: The need for data contracts arose from the challenges organizations faced regarding data reliability, especially when using data for critical applications.
  • Initial Problems Addressed: Organizations were facing issues with schema changes that disrupted downstream applications, leading to unreliable data for business intelligence (BI) applications.

Importance in AI

  • Data Quality: Data contracts play a crucial role in ensuring that the data feeding AI models is accurate and reliable.
  • Governance and Compliance: They help maintain governance standards and compliance with regulations, especially in regulated industries.

Key Discussions

Data Reliability Issues Prevented by Data Contracts

  • Schema Changes: Regular changes in data schemas that break downstream applications can be managed through data contracts.
  • Data Governance: Establishing clear standards for data quality, access, and categorization can prevent issues related to personal and sensitive data.

Decentralization and Collaboration

  • Platform Thinking: Data contracts promote a decentralized approach that allows teams to share responsibility for data while adhering to clear standards.
  • Self-Service: Empowering teams to create and manage their own data contracts fosters ownership and accountability.

Practical Steps for Implementation

  • Identify Problems: Start with a clear understanding of the specific problems you want to solve with data contracts.
  • Engage Stakeholders: Involve the right people—data producers and consumers—in the creation and management of data contracts.
  • Start Small: Implement data contracts in small, manageable use cases before scaling up to larger initiatives.

Key Takeaways

  • Data contracts are essential for ensuring reliable data flow and maintaining accountability between data producers and consumers.
  • They can facilitate better governance and compliance, especially in industries with strict data regulations.
  • Practical implementation is vital: Start with clear objectives and engage stakeholders to ensure the success of data contracts.

Upcoming Resources

  • Webinar: A session on bridging the governance gap with data contracts is scheduled for December 19, 2023.
  • Books and Articles: Andrew Jones has authored a book on data contracts and offers a free white paper (dc101.io) for those looking to gain deeper insights.

Key Links

  • [Connect with Andrew on LinkedIn](https://www.linkedin.com/in/andrewrhysjones/)
  • [Connect with Omar on LinkedIn](https://www.linkedin.com/in/omarshanti/)
  • [AI Opportunity Finder](https://hatchworks.com/ai-opportunity-finder/)

Final Thoughts This episode provides a comprehensive overview of how data contracts can enhance the reliability and governance of data, especially in the context of AI initiatives. By promoting a collaborative approach and clear standards, organizations can improve their data practices and drive better outcomes from their AI applications.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Data is the foundation of AI. Without quality data, without data, you don't have quality AI. Welcome to the Talking AI Podcast, where we talk AI with both experts in the field and early adopters. I'm your host, Matt Page, and we're here to demystify AI for you so you can get some value from it. Let's talk some AI.

0:20What if your data pipelines came with a set of promises, clear, measurable guarantees that protected quality, ensured consistency, accelerated decision-making? Our guest today pioneered that very concept and even wrote a book about it. Known as the father of data contracts, Andrew Jones, saw before most that aligned producers and consumers around these explicit agreements with transform how we trust and use data. Welcome to the Talking AI podcast, Andrew. Thanks for having me. And we also have special guest, returning guest, Omar Shanti, CTO of Hatchworks AI, joining us as well. Welcome, Omar.

1:00At this point, am I even a special guest? Yeah, I don't know. I don't think you are. We're going to drop the special. You're just a regular figure now. All right, guys. Well, I want to kick it off here because data contracts, this was a new thing to me. I know, Omar, you have gone deep on this with Andrew. I'm excited for a little bit of back and forth there. But I think for the audience, like take us back to when this first idea emerged. Like what problem were you seeing in organizations that sparked the need for this thing? What is it? Yeah, so the initial problem I was trying to solve. So we worked in an organization.

1:41We were using data, but we wanted to use our data for more and more important applications. So like we were going to use our data to drive our models, that would actually be drive product features that drive revenue. so we had this aspiration as part of our strategy it was a clear goal of company but our data wasn't really reliable enough to do that like we um our data regularly the scheme would change get breaking things downstream and this is just like bi applications and we felt if our bi applications are not reliable enough built on the same data then how would these applications be reliable enough to tell to customers if it's built on the same data so that was a real initial problem i was trying to solve and i started thinking about problem because there's various things like observability and alerting and things like that but that didn't really seem to self-problem what i felt was self-problem would be a bit more like um like using api like we do in software engineering like i thought the problem was really we're building on top of a database directly we're using change day capture to get data from database into the warehouse.

2:46So building on database directly, the same schemas, those schemas kept changing. Therefore, I felt if we keep building our same source, there's not much we can do to make it more reliable. We need some sort of interface there. So that's kind of where it came from. And then I thought, well, to create an interface, I need to be able to describe it a bit like an API spec. So I created one of those. And I thought, well, actually, if I've got this document describing the data and it's both human and machine readable, actually there's loads more I can do with that. I can start putting things in there to help with data governance.

3:20I can start putting things in there to help with access controls. I can start using that as a document where consumers and producers can discuss around and agree and become like a requirements document as well and documentation. I started thinking, well, actually, all of these things are possible now that it's describing our data in sufficient detail through this data contract. so that's not where it came from it started off as a simple it's a simple idea describe the data but once you realize you can do that and once you realize you put almost anything in there then it can become very powerful and it solved my interface problem to grab the interface and that will end up using itself many other problems as well and when you and just to like simplify it for the audience like are there a few like core tenets that you define around data contracts that take it a step further there?

4:14Yeah, so to simplify it, and because since data contracts have become more well-known, people will use it to solve different problems, and it can do that. It can solve many different problems, not just the interface problem-mastering itself. But basically, a data contract, it contains metadata about the data, so metadata. and it describes that in a human and machine readable format. So it's not YAML or JSON or something like that. It doesn't really matter. Quick break in the pod. If you're listening to this podcast, chances are you've been thinking about how to actually use AI inside your business.

4:49And that's exactly why we built the AI Opportunity Finder. It's a free tool that helps you uncover high impact, tailored AI use cases based on your business, your goals, your pain points, and your industry. No fluff, no generic use cases, is just real ideas that fit your business and they're ranked by ROI potential. It takes about three minutes to run and it's like having your own personal AI strategist for free. If you want to try it for free, check out the link in the show notes or go to hatchworks.com backslash AI dash opportunity dash finder. So that's all a data contract is. And then you can use it for various solutions.

5:25So the interface idea, you can put data quality checks in there and have them run by a platform. You can put documentation. you can put governance, you can put categorization of data and personal data and data retention, whatever problem you need to solve. Once you have this document that contains the metadata, you can use that to power any kind of solution. Yeah, I'm curious, like this is bringing me back to like some, I'll call it nightmares of like sitting in a conference room, going through this huge spreadsheet of all of this data and defining, okay, what is this? What is that? And the reference I'm thinking of is back in my days at like AT &T and in the telecom side.

6:06And not only did we not know what the data was, there was duplicate data all over the place. So I'm curious, like what's Omar, maybe we'll start with you. Like what's your biggest like nightmare moment from the past that like data contracts resolves? And then Andrew, I'm curious what a, I'm sure you have several, but like, or maybe even within a client, don't mention the name. that you've seen where it's just like this pain point that exists? Sure. I'll give a broader answer and then I'll drill in. I had the luxury to work at an organization that was a bit decentralized in its data acquisition practice.

6:47And what that meant is individual teams were autonomous to go out into the marketplace, purchase large amounts of data. These corpuses could be anywhere from the six figures upwards. And what we ended up seeing was a lot of these individual teams were purchasing the same data sets. So to sort of fictionalize the example, imagine eight different teams purchasing the Nielsen annual report for things like TV and household sort of consumption practices. Now, at a smaller scale, that's not too bad. But if you see that is the rule rather than the exception, that these teams are investing so much in purchasing, curating, maintaining these data sets in a silo, then you sort of realize, okay, we need more communication across them.

7:32Now, that's not necessarily the problem that data contracts needs to solve, but it's important to recognize that data contracts fits into the broader arena of data governance. Data contracts is sort of a very key and specific sort of solution in the space of data governance. And as Andrew mentioned earlier, others such as monitoring data observability, data quality, having a data catalog rich with metadata, all of those are basic hygiene things that could have helped with the problem that I just mentioned. So that one stands out as just a really sort of commonplace issue that many Fortune 500 companies might face, but something unique to data contracts has to come down to, I would say, schema changes.

8:17So So I love what Andrew was mentioning earlier that in his org, he sort of felt hamstrung in the sense that you can't increase the importance of the model without increasing the reliability of the data. The more important the model, the more valuable that data reliability is. And something like an upstream schema change breaks it. Something like changing the semantics so that height is now measured in feet rather than meters or how an address is represented. That risks breaking the whole application. So in practice, I've seen it happen where a citizen developer within a company was empowered to make a change to the schema and ended up tanking a production critical application, which was sort of a BI visualization layer.

9:04So ultimately, something like data contracts comes in there and it puts in a very clear sort of spec as to what the schema should look like. it's sort of a commitment, hence the word contract. And there's accountability sort of caked in there, as well as practices around how you can go about issuing a change. So what we sort of call change management. And I think that's the perfect example, because in that case, this production grade downtime would have been avoided. Yeah, good old shadow IT. I've never been guilty of that before in my life. Andrew, what about you? You got a horror story for us?

9:41yeah yeah i think most commonly it's change management like i'm saying this is problem to solve with day contracts but i've seen it used in other places uh governance is a particularly common one so in regulated industries you just need to have better data governance like what data is person data or not how do you handle data retention how do you make sure the data doesn't go across borders if it's not supposed to across territories um and those kind of things capture the limited contract and platform can then automate the policy like put those policies into practice through the platform so that's quite a good one another i've seen a bit more unusual but quite good actually is that there's um when you're i'm not saying like getting data from third parties well this was pick up company they had a lot of subcontractors and they need to get data from those subcontractors back to them so they could put together and create a report for for regulation again.

10:38And they were thinking about how they can create a data contract that crosses into becoming a real contract, part of a contractor agreement, where they provide data in a specific format to the company who's hiring these subcontractors, make sure they get the data they need collected in one single place in the same format from different subcontractors across different countries, so they can create a clean data that they provide to regulators, particularly around, this one was around carbon data for the government. So that's kind of like pulling data from other people. We're still using data contract to formalize that agreement and make sure you get data in the right format, the right quality.

11:22Yeah. Going back to Omar, a point you had, it kind of triggered in my head because I feel like one of these core tenets of data contracts is this, it's almost this decentralization approach to it. You're almost productizing the data in a sense to where people can rely on the data, go get the data accurately. But you also mentioned shadow IT. That's like the negative version of that where people were kind of spitting up their own things. But curious, like how do you avoid, even with data contracts, this like utopian view of like this decentralized, federated data contract world, but avoid it like drifting into kind of shadow IT and stuff still kind of going off the rails?

12:07I think this conversation dovetails nicely with broader conversations around self-service. Whether you're looking to enable like self-service data analytics, which people these days often associate with something like a data mesh, or whether you're talking about enabling more self-serve access to development, and so autonomous teams decentralize, what you end up wanting to do is create standards in a central location that bring together sort of the best minds in the organization to create sort of a single set of practices everyone ought to follow. These should be responsive, not overly prescriptive, but they should sort of do a lot of the work of keeping these teams and then you decentralize that out.

12:51So rather than have the team that understands the domain, do, sorry, rather than have the team that understands the tech stack do all of the technological work, the idea is let's have this team that understands the tech stack come up with a set of standards and a set of reusable tools that they can then send elsewhere so that others can sort of do that work. And, you know, you might think of this as, well, think of maybe a chef might know very well how to operate a fryer. They know how to bring the oil up to the perfect temperature. They know how to get those fries right every single time. But it's difficult to scale that, right?

13:28You can't have this one chef now serves 300 restaurants. So instead, that chef might work with someone to build a machine to fry products for you. And then you can decentralize that machine as a sort of reusable tool across the org. Now, if it's too prescriptive, then that opens up some challenges. If it's not prescriptive enough, then that also opens up some challenges. So figuring out how to really thread that needle is where that technology part really interlocks with the people and the processes within an organization. Andrew, that's the way I think about it. I'd love to hear your thoughts as well on that delicate balance.

14:05Yeah, no, I love that. I completely agree with that. I love the metaphor as well. That was better from the start of date contracts for me as well. When I was thinking of date contracts, I was thinking, It's the same sort of time the data management articles came out, original ones. And I was thinking, yeah, that explains my problem very well. And I like the way the solution is sort of shaped there. Although I didn't think I could change the whole organization. I wanted something a bit more lightweight, which I thought data contracts could be. But also I was thinking about ideas that came from like platform engineering and DevOps and this idea of I worked in a platform team.

14:39I was leading a platform team. So I understood how the platform could enable this decentralization, enable this accountability and responsibility, and even get software engineers to create their contracts and own their data and provide good quality data products to other parts of the business. I felt the platform could do a lot of that work for them. So we're not asking everyone to become data experts and GDPR experts and whatever else it might be. that everyone can create and manage their own data, provide it to our parts of our business using the platform, but with decentralizing the ownership and responsibility to those in the business, to those in the domains.

15:21Andrew, thank you for introducing that concept of platform thinking because ultimately I think that the very core of a platform is building that sort of reusable set of services that people can then build on. So we build the platform in order for people to build the product. That's sort of the same mindset that an enterprise arc has when they come up with the API standards for how to do authentication, pagination, and so on. For what a data architect has when they come up with standards around schema definition or how to sort of capture semantics and so on. And what data contracts effectively solves for.

15:56So having an organization come together to create these sort of reusable components is basically bringing platform thinking and maybe tech light, maybe tech heavy, but even so platform thinking way to the organization's data governance landscape. Yeah, exactly. And that's actually what we did when we first built our data contract implementation is we built it on our same platform, our software engineers use, our same developer platform. we use the same language to define data contract, something called JSON on it which is a bit weird and random if you haven't heard of it it's like JSON with extensions and functions not something you necessarily choose unless you already had it but we already had it and we usually define our APIs and define our infrastructure's code so your data contract just fit alongside that something the software engineers could use easily built into a platform and in that a date contract is you chose what components you wanted did you want backups how long you want them for did you want in our case it was on google cloud so did you want pubsub you want big creed you want both and if you want both as a service gets deployed that syncs and keeps them in sync archives data from pubsub to bigquery so and as a um software engineer creating this date contract you don't know what all you don't need to know what's going on behind the scenes you're just writing a contract and behind the scenes a platform is automating all that for you and both y 'all mentioned the term data mesh and i know omar you talked some about that you were talking about this at our uh ai dinner we had here in atlanta the other day what is the data mesh is it an alternative to data contracts is it complementary uh talk us through that all right i can offer some quick thoughts so um the data mesh we can think of is a architecture It is a sort of proposal for how an organization can build this sort of decentralized form of data analytics landscape that is in a way built around a few different core pillars, but whose implementations greatly vary.

18:05So probably at the 2018-2019 mark or so, this now famous author, of course, Jamar Deghani, published this book around the data mesh, laying out these four principles. One of the key principles is around bridging the gap between the people who understand the data best, sort of that application team who are building that data, And the business team will understand how to get the most value out of that data without introducing all of the sort of heaviness of the data lake and the data warehouse teams who are bringing data from A to B, losing business context, but maintaining sort of technical context.

18:44So the problem the data mesh effectively solved for is something like 80 % of data lakes fail. Similarly high number of data warehouses fail as well. And a number of AI programs as well as BI programs similarly failed. So to bridge those gaps, DeMarc recognized that one of the biggest bottlenecks is in this gap between the people who are creating the data and people who are sort of consuming the data analytically. So the goal of the data mesh is why don't we, in a way, decentralize by empowering these people who are producing that data to go one step further and also publish their data in ways that support maybe analytical workflows or operational querying, which is sort of real-time querying at a scale that transactional systems kind of can't support.

19:32um yeah it was that uh the databricks uh ai world event the other day and they had the head of ncr voix uh their ai he was uh the analogy he gave is going from like a data swamp uh which is probably more related to the 80 that fail is some analogy to that but then progressing data lake uh data marketplace and obviously they were talking databricks with their uh lake house But yeah, interesting. Yeah. I didn't realize that. That's a crazy stat. I'll kick this to Andrew in a second, but I think one thing to stress as well is a data mesh is only one particular architecture to enable the self-service analytics, but also a data mesh takes on many different forms.

20:15So it's one of those where you can sort of, you know it if you see it, but ultimately the implementations vary so much that it motivated John Josperin and Eric Broder to just write a book recently around all of the varying data mesh implementations that they've seen, because they're just so diverse. And a lot of the sort of core principles that Jamak lays out are great, but still a bit underdefined. So there's a big gap to go from the theory of the data mesh to the practice. And for that reason, it tends to be a bit of an elusive concept. Data contracts can play a role, but they're not necessary in a data mesh, and nor are data meshes necessary for data contracts.

20:53You can use data contracts in a different architecture and you can do data meshes without contracts. But I think Andrew and I would both agree that whatever your architecture is, a data contract is a great way to increase the reliability of your data and then build the more sort of important models and dashboards on top of. Andrew, your thoughts? Yeah, I definitely agree with all of that. Yeah, like I said, data contracts are just good practice. Good idea to be applying some discipline to how you create and manage your data. data, good idea to describe your data so that you can then create some automations, you can then populate data catalogs, you can then have agreements with your producers and consumers, whether you're doing data mesh or not, that's still a good idea to have.

21:36But if you're doing data mesh, I think you'll find it quite difficult to do data mesh without data contracts, particularly when you're trying to implement this self-serve data platform, the idea of data mesh, particularly been trying to implement the federated computational governance, I think it's called a data mesh, that can be achieved through data contracts quite easily. So I think it'd be hard to do those concepts without data contracts. But yeah, data contracts are useful without data mesh as well. Yeah. And Matt, you'd mentioned the dinner that we had recently. I won't go into it here, but listeners can, of course, reach out at any time.

22:16One of the visions that I'd laid out was for an agent mesh. Now, the concept of a mesh is not unique to data. You have something called a service mesh, which in a way prefigured the data mesh. And service mesh comes from that microservice domain-driven design lineage, where the idea is you can take these big APIs and then break them out into smaller reusable components, have individual teams manage them, and then scale up that way. Well, I won't go into too much over here, but I think what we're seeing in the world of Gen AI and LLMs is that as the base models sort of plateau, people are having a chance to sort of build up the stack.

22:55And what that enables is teams to build more sort of finely tuned or smaller models that are better at performing certain tasks and then sort of delegating responsibility across multiple of these agents. And in a way, what we end up creating is akin to an agent mesh, this set of agents, each akin to sort of a microservice, which does like one thing very well and is able to sort of route to other agents because it sort of understands where they are and how to interact with them. So very much similar to the service mesh side of it. There are parallels with the data mesh, but I think they're a little bit harder to see.

23:32But anyways, we can get more into that if anyone's interested. it's offline uh but yeah and i'm sure some listeners are thinking right now this is the talking ai podcast finally they're talking about some ai but i hope listeners are starting to uh cue into this that data is the foundation of ai without quality data without data you don't have quality ai and i think what's really interesting andrew is like your concept of data contracts came in 2021, which was before the advent of and proliferation of Gen AI with, you know, ChatGPT is kind of the seminal moment. But I'm curious, how has that changed or evolved data contracts?

24:12And maybe most specifically, AI is very good at taking a lot of unstructured data. How does that play into the concept of Databricks and just like unstructured as a general topic? yeah i think generally the sort of increased use of ai as a lot has made a lot of organizations a lot of people start thinking about quality of their data that feeds into their ai like you mentioned and also these applications of ai they tend to be deployed in more critical places in business so it could be product features could be internal processes things that just can't break i mean i reporting should be able to break all the time anyway but like this is another level of reliability that we need and that's why people are starting to think about things like data contracts things like data quality reliability to make sure that data is dependable enough to serve the models has some sort of business outcome but it's important after that so i think that's why data contracts continue to be more widely adopted it's to solve that problem because the use of data is more important than it has been before i think when it comes to the second part of the question was around unstructured data the day contracts it's always really been about structured data for me so it describes the structure of the data and what's in there and how you can use it but when i think about unstructured data it's still some metadata around what you want to capture so like where is it what kind of s3 bucket for example is it in um who owns it what kind of person data is in there, can I use it in this particular use case, particular AI use case or not?

25:53So those kind of, that kind of metadata about the data again, that kind of the descriptive policies around data, what you can use it for, what you can't use it for, those are very important when it comes to creating, taking more explainable AI models. And I think that's a good place where data contracts can help capture information in in a human and machine readable format, but it's owned by VDataOwner. It's got some virtual control around it, some change management around it, and it's well-defined and well-understood. Nice. Omar, any thoughts on that? Yeah, I love that. I mean, it's a great question because ultimately what the advent of Gen.AI has done is allow people to more easily take advantage of unstructured data that was either too primarily intended for a human audience.

26:42So PDF documents, manuals, scanned images, right? All of these data sources were, for the most part, a bit inaccessible unless you wanted to invest time in building either your own model or using an open source model to extract natural language and so on. All of that's sort of available with a lot lower of a barrier to entry. So that raises important questions as to how do we even talk about the quality, the lineage, and other sort of facets covering the sort of metadata of these unstructured documents. I've been keen to sort of have these conversations. I've been speaking with some folks, including briefly Shashank Adas of AcroData and the Data Hub.

27:27But it seems that there's just a gap in the market. Now, there's a gap in our vernacular, in our concepts, as well as like in the tooling around identifying things like the data quality of PDFs and videos and how we can talk about data readiness and things of that sort. Andrew, do you agree? Are there any sort of tools that you think are letting companies get ahead of that? Because I think that's a really exciting problem space now for enterprises looking to build enterprise-grade RAG systems, for example. Yeah, that's really interesting. I don't know of any tools personally. I think I agree with you that there's a gap here that needs to be filled by something.

28:07So, yeah, I think, yeah, it's really interesting. I think it's an interesting problem that somebody, I'm sure somebody's trying to work out how to sell and be interested to see how we sell that in the next few years. Yeah, that's the cool thing. I think all of this emerging technology and innovations, just creating all these new problem spaces for new solutions to be solved. Two more things for you, Andrew. I'm curious, you know, we were just talking about AI and the explosion of it. Do you see kind of like a meta idea here, AI being used as a tool to help create, maintain data contracts? When you think of as the Databricks event yesterday, they were talking about using AI to define like documentation and kind of metadata and things like that.

28:51Do you see that being a tool that can be used to make data contracts easier in the adoption of those? Yeah, I think so. So, well, I think I've kind of got two opinions here and they kind of contradictory. One is, but yeah, anything we can do to reduce the adoption effort of creating a date contract or modifying a date contract for use of AI, because these are structured documents that have a structured schema. So a bit like code, you can kind of generate it for AI and if it doesn't work, you'll be able to submit it to production. It will just fail the checks there. So anything we can do to reduce the effort there is great.

29:29one particular idea might be to automatically categorize your data and things like that so you don't have to categorize yourself manually and that is a bit basse bit in particular a bit conflicted because on the other hand that can save a lot of hours like mainly categorizing data is labor intensive on the other hand sometimes it's better to have a human bed checking the data because AI, if they're not perfect, they are making assumptions. They are categorizing best they can, but they're not perfect. And if you're in a regulatory industry, for example, you're storing personal data or sensitive data, would you want you to risk miscategorizing that because you left it to an AI?

Read the full transcript

30:12Would you want a human to categorize it once when you start creating that data and then have this platform then know about that? And from then on, that's definitely sensitive data. that's a bank account, that's a national insurance number, or whatever it might be, and make sure you store and process that correctly. Because I suspect if you didn't, you know, if you had some sort of data breach or some sort of incident where the data was miscategorized and you said, oh, the AI said it was fine, but it wasn't, I suspect that wouldn't really be an acceptable answer. So I think sometimes where I think it's okay, it's better for a human to be involved.

30:52But yeah, I don't understand. It's a lot of effort, potentially, if you're doing this from scratch. And AR could tell you help with that, at least maybe as a first pass or something like that. Yeah, liability still comes back on the human, right? That's where it's there a human in the loop aspect of this to where you kind of get best of both worlds. And I think that's something that, you know, you kind of have to wade your way through to figure out what's the ideal state. And then the last thing I got is probably people listening to this that are like, yes, give me data contracts. This sounds great.

31:21I'm waiting around through this murky data swamp with lots of fog like we were talking about earlier. What's the logical way to get started other than just reaching out to you directly and getting somebody to help walk them through that? Yeah, that's a really good question. I think, first of all, you need to be clear what problem you're trying to solve with data contracts. Myself and I spoke about many different problems you can solve date contracts. So what is the one you're trying to solve? And often people say, I want to solve date quality and I want something to enforce date quality on my producers.

31:56But I'm really starting to use words there like enforce and control, which I mean, contract is another word, but it's kind of words where you're not really encouraging much collaboration. You're already starting off with friction. So I think before you start deploying date contracts, think about probably trying to solve who you want to be creating these date contracts So is it you as data people, data engineers? Is it software engineers producing data? Is it your Salesforce admin and the data coming from Salesforce, for example? Think about who it is you want to be creating these data contracts and then work with them from the start to engage them in the data contract solution you want to build, making sure that you're building something that they are happy to use and they understand that they are taking this ownership, responsibility, accountability for the data contract and by extension of data that's come through that.

32:48So that's kind of where I start. We also tend to start small. Find a good small use case to start. Don't, we know, I mean, everyone knows that like these large IT projects, they tend, you know, don't deliver anything for 18 months. They tend to not be successful. So start small, solve a real problem for the organization. Do it with a day contract, build a bit of platform while you're doing that. Then do the next problem. and start to show the value early and often while building out your capability as well. No, that's great. Well, Andrew, thanks for being on the Talking AI podcast. So a couple of things here, and I'll let you kind of go through the spiel.

33:29But you wrote a book on this. It's on Amazon. Tell us where people can find you, where they can reach out. I know you have a great newsletter. Give us the spiel of how people can reach out to you. Yes, I wrote a book. It came out last summer. and it's available at Amazon and wherever else you get books. But if you go to data-contracts.com, you'll find all the links there and all the information about the book. If you're looking for something a bit easier to digest or your new state contracts, more and more of an overview, an introduction, then I wrote a free white paper on date contracts recently.

34:03You can get that from dc101.io. So that's dcwithdatecontracts101.io. And that's a summary of what date contracts are and how you can use them to solve problems. And then, yeah, as you mentioned, I've got my newsletter as well. I'm quite active on LinkedIn. And if you go to andrew-jones.com, you'll find all the links to my newsletter and my LinkedIn and wherever else I post on there. Nice, that's awesome. And we also have on the Hatchworks AI side, we have a lab coming up, effectively, a webinar on bridging the governance gap with data contracts on Databricks on December the 19th. If you're listening to this before that date, come join us, come check it out.

34:42We'll have Omar and a couple other folks in there. It's going to be a really interesting and fun chat. But thanks, guys, for being on the Talking AI podcast. Thanks for having me. I really enjoyed it. Thank you. Thank you. Thanks for listening to the Talking AI podcast. If you enjoyed the show, give us a follow or subscribe on your favorite podcast platform. And don't forget to leave us a review. We love those. For more info on Talking AI, visit TalkingAIPodcast.com. The single biggest mistake we see companies make with AI is they don't properly train their teams. We see it all the time. Companies roll out AI tools and expect people to just figure it out.

35:23But using AI effectively requires a totally different mindset and skillset. And that's exactly why we built training for every level of your org, from AI training for teams and executives to training engineering teams on our generative-driven development methodology. Or if you've already identified your AI use cases and want to just prioritize where to start, we offer an AI roadmap and ROI workshop to help you build a quick plan. It's all about going from we should use AI to actually driving real value with it. Head over to hatchworks.com to learn more.

From the publisher

How can organizations ensure that their data pipelines deliver consistent, trustworthy insights without frequent disruptions? In this episode of the Talking AI podcast, host Matt Paige co-host Omar Shanti, CTO of HatchWorks AI, speak with Andrew Jones, Principal Engineer at GoCardless, about the concept and practice of data contracts. Andrew explains what data contracts are, why they emerged, and how they improve data reliability and collaboration.

The discussion explores how data contracts relate to data mesh and platform thinking, helping teams share responsibility for data while maintaining clear standards. Andrew and Omar share examples of data contracts addressing issues like schema changes, governance, and data quality.

They also consider how data contracts support AI initiatives by ensuring that the data feeding into models is accurate and reliable. Practical steps are given, including starting with a clear problem and involving the right people.

Andrew provides resources for deeper learning, including his book and online materials. Matt and Omar also mention an upcoming webinar on data contracts, inviting listeners to learn more and move toward more structured, dependable data practices.

Key Moments:

  • Introduction to data contracts and their purpose
  • Highlighting the common reliability issues that data contracts can prevent
  • Linking data contracts to data mesh and other modern data architectures
  • Exploring how data contracts help maintain governance and compliance
  • Demonstrating how data contracts ensure quality for AI-driven applications
  • Discussing platform thinking and its role in scaling data contract implementations
  • Offering practical advice on how to begin using data contracts in your organization

Key Links:


Mentioned in this episode:

AI Opportunity Finder

Feeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you’ll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action. 👉 Try it now at https://hatchworks.com/ai-opportunity-finder/

More from Talking AI

All 84 episodes
The Father of Data Contracts - AI’s New Best FriendTalking AI · 36 min
Listen in VO