In short
Podcast Notes: Talking AI - Episode: Don’t Trust—Verify: Building a Proof of Quality for AI Data
Episode Overview In this episode, host Matt Paige interviews Rowan Stone, CEO of Sapien, highlighting the significance of data quality and provenance in AI. They delve into Sapien's innovative decentralized data protocol, focusing on the "don't trust, verify" principle. The conversation covers the challenges posed by bad data, the complexities of achieving accurate AI outputs, and the implications of synthetic data.
---
Key Themes
- Importance of Data Quality
- Quality and provenance of data are crucial for reliable AI outputs.
- AI systems are only as good as the data they are trained on, raising the question of trust in AI-generated information.
- Challenges in AI Development
- Misinformation and model collapse are significant concerns.
- Risks are heightened when AI systems, such as autonomous vehicles, interact with the real world.
- Sapien's Unique Approach
- Sapien builds a decentralized data protocol to ensure data quality through validation and peer review.
- Their model includes incentives for data validators and emphasizes transparency in the data supply chain.
- The Role of Incentives
- Current AI systems are driven by perverse incentives that can lead to unreliable outputs.
- There is a tendency for models to provide agreeable responses rather than seek objective truth.
- Consensus and Collaboration
- Drawing parallels between AI and blockchain/crypto, consensus models are discussed as a way to ensure data integrity.
- Collaboration across different AI systems is essential for overcoming fragmentation in the field.
- Synthetic Data
- The difference between human-created and AI-generated data becomes less critical if data validity is established.
- Synthetic data can enhance model performance if properly sanitized and validated.
- Future of AI Models
- Expectation of multiple specialized models rather than a single dominant AI.
- The debate between open-source versus closed-source models continues, with implications for speed of development and collaboration.
---
Key Moments
- 01:04: Importance of Data Quality in AI
- 03:32: Challenges and Risks in AI Development
- 07:08: Sapien's Approach to Data Validation
- 08:35: Incentives and Trust in AI Systems
- 13:30: Building a Decentralized Data Protocol
- 23:22: Consensus and Collaboration in AI and Crypto
- 30:55: The Role of Synthetic Data
- 36:17: Future of AI Models and Open Source
---
Key Takeaways
- Data Validation: The necessity of robust data validation mechanisms to ensure AI systems are trustworthy.
- Incentives Structure: Understanding the perverse incentives that can lead to model failures and hallucinations.
- Decentralization: The benefits of decentralizing data sourcing and validation to improve transparency and trustworthiness.
- Synthetic Data Use: Recognizing that synthetic data can be valid and beneficial, provided it is properly vetted.
- Collaboration: The need for collaborative efforts in the AI field to overcome siloed developments and enhance overall progress.
---
Additional Resources
- [Sapien](https://www.sapien.io/)
- [Rowan on LinkedIn](https://www.linkedin.com/in/rowan-stone/)
- [State of AI 2026 Report](https://hatchworks.com/state-of-ai-2026/)
- [AI Opportunity Finder](https://hatchworks.com/ai-opportunity-finder/)
---
Lightning Round Insights
- Trend: Social AI applications perceived as mostly noise.
- Infrastructure vs. Product: Building infrastructure is more complex due to the need for general application across various use cases.
- Human Responsibility: Oversight remains a critical human responsibility despite AI advancements.
- Constraints on AI Progress: Availability of high-quality training data and regulatory guidance are significant non-technical constraints.
- Hardest Domain to Validate: Fields requiring specific expertise are often the hardest to validate due to limited knowledge availability.
---
Conclusion The conversation with Rowan Stone at Sapien emphasizes the pivotal role of data quality and transparency in the AI landscape. As AI systems become more integrated into daily life, approaches like that of Sapien may offer solutions to ensure that trustworthiness and accuracy are prioritized in AI outputs.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding AI Trust Issues
0:45 to 2:29
Discussing the importance of data quality and provenance in AI systems.
“And we're going to get into what that means.”
Building a Protocol for Data Quality
2:29 to 3:47
How Sapien is creating a system to ensure data quality for AI training.
“And like you said, learning from the on-chain world and applying some of those principles to this AI data problem.”
Identifying Failure Modes in AI
3:47 to 4:17
Exploring the potential failure modes and risks of AI in real-world applications.
“What are these failure modes that you're most worried about?”
The Dangers of AI Models in Real Life
4:29 to 7:37
Discussing the implications of AI models in real-world scenarios and the need for accurate data.
“Yeah, I mean, I essentially view where we are as an industry as a giant test in production mode.”
Understanding AI Incentives
7:37 to 12:03
Examining the incentives driving AI companies and their impact on AI trust and performance.
“They did this earlier in the year, but they basically gave full autonomy to the model to run and manage a vending machine.”
The Current State of Data Training
12:03 to 14:01
Analyzing the existing landscape of data training for AI and the role of companies in this process.
“of having something that you can chat to and something that you can bounce ideas off of.”
Understanding Current Data Annotation Models
14:01 to 18:06
Learn about the traditional models of data annotation for AI and their limitations.
“I think, I believe y 'all started more on like the services side and it kind of came into this almost infrastructure protocol world, but it has like learnings from the crypto world in a sense.”
Towards a Trustworthy Data Standardization
18:06 to 23:40
Explore the need for open standards and trust in AI data creation and usage.
“But I'm also curious too, When you think of what is ground truth, how do you get to that answer, especially when things are subjective?”
Parallels Between Crypto and AI
23:40 to 28:00
Examine the similarities and differences between blockchain technology and AI data practices.
“And I'm curious, like, obviously, kind of gotten into some of this, but you've worked at Coinbase, you know, on chain and all of that stuff.”
Exploring the Role of AI in Data Verification
28:00 to 31:24
Discussion about the potential for AI to participate in data verification processes and the implications of synthetic data.
“And it feels very messy and out there, but it just kind of feels like it naturally happens.”
Show all 15 chapters
Synthetic Data and Model Reliability
31:24 to 36:03
Examining the impact of synthetic data on model performance and the importance of data validity.
“Because there's a lot of people that feel like, you know, at some point you may have model collapse where it's almost recursive in and of itself and it starts to degrade.”
The Future of AI Models: Size and Accessibility
36:03 to 38:45
Insight into the future landscape of AI models, addressing size, open vs closed source, and user preferences.
“And then make sure to your point, it's failing in a safe way or you have guards and safeguards around that.”
Lightning Round: AI Insights and Challenges
38:45 to 42:00
Quick-fire responses on current AI trends, responsibilities, and the biggest constraints in AI development.
“All right, I want to wrap us up with a few lightning round questions.”
The Importance of Essential Tools
42:00 to 43:39
The discussion revolves around useful tools and personal preferences.
“that you personally cannot live without right now.”
Introducing Rowan and Sapien
43:39 to 44:14
Rowan shares insights about Sapien and its upcoming initiatives.
“We're starting to speak more openly about where we're going with proof of quality.”
Transcript
Automatic transcript. May contain errors.0:00We're obviously living in the information age, but with unlimited information comes unlimited noise. Welcome to the Talking AI Podcast, where we talk AI with both experts in the field and early adopters. I'm your host, Matt Page, and we're here to demystify AI for you so you can get some value from it. Let's talk some AI. AI is everywhere, but most of us can't answer a simple question. Why should we trust what it tells us? A lot of that comes down to something no one wants to talk about, which is the quality and the providence of the data these systems learn from. Today, I'm joined by Rowan Stone, CEO of Sapien, a decentralized data foundry focused on turning collective human knowledge into verified enterprise-grade training data.
0:47And we're going to get into what that means. But Rowan has helped drive on-chain products at Coinbase, and now he's applying that don't trust, verify mindset to AI, building what Sapien calls proof of quality for the data supply chain. We'll dig into incentives, validation, synthetic data, and what it'll take to trust AI agents in the real world. But great to have you on, Rowan. Great to be here. Thanks for having me. Yeah, this is going to be a fun chat. The world that you've spent a lot of time in, I think, has a lot of translation into the world of AI that we're playing in now and is so emergent.
1:24But I first want to ground the conversation in what you're building at Sapient and the problem you're focused on solving. I think that'll be good context setting for this conversation around data, which is like the most important thing when it comes to AI. Yeah. I think the best way to frame this is we're obviously living in the information age, but with unlimited information comes unlimited noise. And so for us, what we're trying to do is build a system that will enable AI builders typically to be able to figure out what's fact and what's fiction to try and actually get to the point where they can understand that the data set that they're using to train their model is actually the right data and something that's going to give them a good trustworthy output.
2:15And so we're basically trying to answer two key questions for builders. It's like who created the data and can that data be trusted? We're doing it in a slightly unique way, approaching the problem from a kind of first principles perspective. And like you said, learning from the on-chain world and applying some of those principles to this AI data problem. And so building a protocol and essentially enabling the builder to rock up, source either data, or it could even be agent output evaluation or validation of an existing data set. And then essentially let the protocol take over perform the work, which we can get into detail in terms of how that works, and ultimately receive what they need, be that an eval or a set of data, and an attestation that shows that the work has been performed in line with the specification that they provided, and importantly, that that work has been peer-reviewed by people that have the relevant knowledge to ensure that the data they're going to use for training is actually decent and not just noise.
3:25Yeah, and I think it's an interesting thing for the audience. You live on the side of the training being done in the models, the data being used to train the models, data being used to fine tune models, not necessarily on the inference side and whatnot. But I'm curious, with this insane amount of information out there. What are these failure modes that you're most worried about? We have misinformation, we have model collapse, we have bad decisions being done by agents or something else. What sticks out to you? And they're all potentially bad, but where are the biggest problems going to emerge where bad data is going to rear its head, I suppose?
4:16Quick break in the pod. Our State of AI 2026 report just dropped, and it breaks down what actually is changing in AI, what's hype, and what leaders need to be paying attention to this year. You can grab it right now on our show notes or at hatchworks.com. Yeah, I mean, I essentially view where we are as an industry as a giant test in production mode. Like we're in the world's biggest public beta. And so most people who are kind of either building within the ecosystem or just tech forward and kind of understand this stuff, they're going to be using ChatGPT, Claude, insert favorite model here. And they're going to take its outputs with a pinch of salt because they know like these things are not perfect yet.
5:04And so I think less concerned about standard text-based chat model style outputs. Sure, you can have something that's factually incorrect. You can also potentially maliciously ingest data into a model and have it kind of steer opinion and kind of move away from ground truth into some sort of prerogative that you may have. Like all of these things are potentially dangerous. There was a case relatively recently where a model, and I forget which one, and it's not worth talking about which one, but a model had been asked by a kid who was experiencing some mental difficulties at the time for support.
5:45And ultimately, I think the way it panned out is the model almost egged the kid on to take really horrible action. And so these are clearly super dangerous. However, even though they sound super dangerous, I'm actually much more worried about models that we invite into the real world. And so those could be autonomous vehicles. You've got a two-ton car driving around on the streets with like not one person, but everybody and their pets and everything else. And then we are increasingly seeing humanoid robotics kind of coming to the forefront of media attention. There's a ton of companies building some really cool stuff in that space.
6:27I've always been a fan of Star Trek and things like this. And so for me, it's almost like we're getting towards the point where we have an ability to make a data, which is really cool. But with that coolness comes a ton of problems. Like now we're inviting a robot that will initially have similar strengths to a human, but very clearly will soon have improved strength and dexterity and movement and speed to a normal person. And we're inviting these into the real world to interact with your kids, with your pets, walking down the street. And so all of these cases where we're bringing an AI model into the real world is really where I see the risk as being the highest.
7:09And that's for me where I view the kind of tolerance, if you like, for bad data being the absolute least. We need complete surety that these models are trained on super accurate inputs that are properly peer reviewed so that we don't end up with a ton of different safety issues, be that autonomous car, the plane that flies you on your holiday, the boat that takes you across the river, or a robot walking down the street. Yeah, you mentioned the scary use case early on, a more mundane one, but I saw Anthropic Cloud. They did this earlier in the year, but they basically gave full autonomy to the model to run and manage a vending machine.
7:50I don't know if you've seen this story yet, but it basically could go out and purchase things and it had a goal to turn around a profit. But it was very easily manipulated by different users getting coupon codes, you know, saying, hey, I'm the CEO. And it had a lot of kind of hallucination going on. And some of that could be related to the underlying data. A lot of it also related to just, you know, the system prompts and how the system's architected. And it's leaning more towards kind of being a helpful assistant. But you mentioned a point earlier, which I think is critical. people just trust these models a lot, right?
8:30So if I'm in chat GPT, the answers sound so good, right? And that's, it's kind of what, and it gets to the incentives. They're, they're incented to provide a probable good sounding answer and response, right? And I've seen this come through in some data, uh, where, you know, thinking about people converting or doing things from the responses they get on chat gpt are much higher than you would see in more normal forms where there's a lot more work on the human to go find the information the data figure out a decision and whatnot uh but get to the incentive piece because i think that's a huge underlying aspect to all of this i think this whole problem you mentioned right it's the biggest largest you know beta in the world i think it's a great way of framing it i think it's just an under discussed topic.
9:23Yeah. I mean, the incentives are perverse and insane. Like there is a ton of money to be made for all of the companies that are trying to build these things. And so you will see companies doing what's best for them and their shareholders and not necessarily thinking about what's best for the planet, what's best for us, what's best for humanity as a species. And this kind of opens a bit of a Pandora's box of discussion. But if we pair it back to the incentives for them kind of to proliferate a model, what they're really trying to do is put a friend in a user's pocket and get to the point where the user is able to converse and trust said friend's outputs.
10:11Because then they have essentially an ability to kind of sit on the user's shoulder and recommend courses of action and monetize in many, many different ways. If you think about how Google typically operates, they answer your questions and they make suggestions for things that they think that you might like. And they make a ton of money just by understanding you and suggesting the right things. I see that as being a likely course that we're going to end up with. But I think the tricky thing that we have, particularly in kind of the current giant public beta is that it's not well understood how the models are behaving and how they are trying to gain your trust and to kind of be that trusted confidant or whatever.
10:57And the way that, let's just pick on ChatGPT because it's huge and basically everybody uses it, the way it typically operates, or at least the way the current models operate, is that they'll be overly agreeable. And so if you ask a question, and it's relatively open, and you kind of suggest an answer within your question, it's quite likely that the model will give you that reassurance of like, yeah, you're on the right track. That's super cool. Great, great thinking. If you're in a business scenario, and you're like, I don't really know how to deal with this weird thing that's going on, like, what should I do with my customer or whatever?
11:34It's probably going to reinforce whatever it is that you've partly suggested and the danger there is that you're basically just talking to yourself like you're creating a circle jerk of like group think basically between you and a model rather than having like a real objective ground truth seeking advisor who'll basically just say no Rowan like you're full of shit that doesn't make any sense do this instead. Which we all need. And so we absolutely, like to me, that's the real benefit of having something that you can chat to and something that you can bounce ideas off of. And so I know your audience are pretty savvy, but there's a really quick way to, I guess, patch in some way.
12:18All of the models at this stage have an ability to tweak how they answer you and to tweak how they respond to questions and prompts. And so you can literally just prompt within your settings like do not be agreeable and seek the truth and make sure that you are challenging my thinking and position yourself as an expert in yada yada like you can make your own prompt but it's super important to understand that without that type of setting you're gonna get a circle jerk of like yeah that's great do more of that and that's because they're incentivized to to make you enjoy chatting to them. Yeah, it's the whole sycophantic nature of it.
12:59I mean, when it makes its way to a South Park episode, you know it's reached a mass scale in a sense. But another tip on that side, what I like to do is tell, when I'm in a conversation, tell whatever model I'm working with to be critical of the response it just gave me. Like go and be critical of your own response. And that's another way to kind of break it of that pattern in a sense. But getting back to the data side, because like you said, the training side of models is very much feels like a black box in a sense. And people kind of have this understanding that, oh, it's trained on all the data on the internet.
13:39It's just scraped all the data. But there's nuance to that. And there are companies that are assisting these model makers. You mentioned like autonomous cars in terms of labeling data all these different things. What does that world look like today? And then where is the gap? And I think what's super interesting is you're introducing this protocol and this is not where your sapiens started, right? I think, I believe y 'all started more on like the services side and it kind of came into this almost infrastructure protocol world, but it has like learnings from the crypto world in a sense. So I'd love to kind of understand what's the difference between what exists today, what you're building and how that is potentially different and better.
14:28So what exists today is large BPOs, people operations, essentially, typically in the developing world who are asked to create data sets or evaluate existing data sets, annotate over the top of them, provide extra context. If we take the autonomous vehicle example, they're driving millions of miles all around the world in different countries. They're collecting video data, sensor data, whatever it might be. It depends on the company, what they've chosen to kind of work with. And whenever they encounter a funny scenario they don't know how to deal with, then that creates a kind of data request, if you like, in an annotation queue.
15:14And it goes out to one of these shops full of people. and then people will look at the data they'll try and figure out where the car is in space and time what it's looking at and importantly give it the extra context it needs so the model can better understand that scenario in the future and know how to safely navigate whatever that might have been and so this the current model is very much traditional service business and you're right it's it's a bit of a black box if for example a car crashes into something let's just take an extreme example the obvious question is going to be like who was at fault like was it the autonomous vehicle or was it like did the wall jump out at it whatever and in a world where it's the autonomous vehicle at fault we don't have visibility in terms of where did the data come from who created the data can the data actually be trusted how was the training done how was the inference done it's very much proprietary to each individual company that are building these models and in some ways that's right but in other ways when we're bringing real heavy dangerous things into the real world and having them interact with people that don't give permission for example to have these then we need to kind of be a little bit more open in our thinking and so you're right we started as a service business we were doing work with amazon zoeks and with toyota and mid journey and a whole bunch of different customers and ultimately we were sourcing people to come and give models context or to check the kind of outputs of different models.
16:50And we were doing it in a slightly different way to the kind of norm. We were posting tasks in a task dashboard, allowing anyone anywhere to participate and not really caring where people are from. To us, it doesn't make a difference. Everybody gets paid the same irrespective of kind of lines on a map. And so in some ways, trying to move the standard service model on, but quite quickly came to the realization that the real prize to go after here is to standardize not so much how work is done but standardize how we define what good looks like for the entire industry and so i'm not sitting here saying that we've got the answer but what we're working on is a credible way to understand who created your data and if that data can be trusted and importantly to get an output and an attestation that's completely open and easy to audit so that if something does go wrong you can quite quickly figure out why and you can make sure you don't do that again and importantly everybody else can quite quick quickly figure out why it went wrong and they can learn from that lesson rather than the kind of siloed approach that we're in today and so it's early i want to kind of caveat what i'm saying here we have a fully working version of the protocol in testnet and we're aiming to ship a production version and invite the first set of customers and builders to start using it in february there or thereabouts pending some audits making sure that security is up to scratch uh however that's where we're at today um and yeah but it's but it's an interesting approach though because i think you're kind of decoupling it away from companies who obviously have their own motivations, right.
18:37And incentives, whether that be the model makers or the companies doing this kind of data labeling in services today, which does bring a little bit more, I guess, democratization to it and almost, you know, it trust I think is like the key foundational thing between this is the vision in a sense. But I'm also curious too, When you think of what is ground truth, how do you get to that answer, especially when things are subjective? Because I feel like trust and truth in general is not necessarily a binary thing. It's either true or not. There's a spectrum. And so how do you account for that, whether this is specific to Sapien or just your views in general around trust and quality of data and truth and all of those big meaty things?
19:32If we solve this, we win, basically. Like this is the thing to go after. The way that we're approaching it and everything that we're building will be open source. And so I'm very happy to talk about the way that we're approaching this. but the way that we approach is that we have in fact let me take a step back and let me kind of double click a little bit on my answer for the last question the current model is very traditional service business so if you're an engineer you're building a model you knock on your procurement team's desk and you say hey i need x data set annotated or i'm missing a data set go and find me this obscure thing and then the procurement team will go out to market knock on everybody's desks, send out tenders.
20:15We go through this giant process of annoying bureaucracy for weeks and weeks and weeks. Eventually, the procurement team will select a vendor, the vendor will get a purchase order, and they'll go and start doing the work. And then months later, said MLOps engineer starts receiving data that he needed two months ago. What we're trying to do is get rid of that entire nonsense and really collapse the process to get to the point where we can quickly get new data and quickly iterate on new models. And so in this new system, MLOps literally just comes into POQ, our system, proof of quality, or POQ can be embedded exactly where they're doing work today.
20:56And so the first adapter that we've built to allow engineers to work where they are is a Langchain adapter. This is a set of tooling used for the agentic space. And so it means that anybody building with that can directly connect and enable flow of data quite quickly. So rather than this huge complicated process, MLOps engineer is able to build a specification of what they're looking for. That then, the protocol then creates the kind of advert, if you like, to workers that can do the work. Once the work has been done by anyone anywhere that has the right knowledge, credentials and experience he goes into a validation queue validation is again completely open anyone anywhere can participate and it's kind of working on a free market principle where anyone that has the right credential and the right experience is able to submit a validation as well as a confidence score in terms of how confident they are that the answer is correct or incorrect so they're grading themselves essentially with that confidence score effectively exactly they're they're saying like yes this is airpods i am super very very confident it's airpods so i'm going to essentially bet a ton of value that i'm right and then it's it's done on a comment and reveal basis what we need to avoid is all of the validators checking each other's answers and colluding because then no value is created and so they're all done in private they submit their answer they submit the bet essentially and you essentially find the correct answer is kind of mapped with a bunch of people saying yes we all agree this is correct and you'll have a whole bunch of others who are just like not doing the work cheating whatever but they're way off to the right somewhere we essentially take value from those that are cheating and found to be wrong and disperse that value to everybody that gave us what's found to be the kind of logically correct answer.
22:49There's a bit more detail here and a lot more math, but it's a consensus algorithm to enable us to get towards what a group of the right peers believe to be the right output. And importantly, without colluding and being properly incentivized for making sure that they give an honest answer. Yeah. And that was one of my questions too, is how do you avoid the collusion and whatnot. But it's kind of a beautifully simple approach, which I think the best things typically are, right? It's almost that you're leveraging the knowledge of the commons, the crowds of people to get to where truth is versus relying on one single source, which has points of failure, right?
23:36So it's kind of a very interesting approach there. And I'm curious, like, obviously, kind of gotten into some of this, but you've worked at Coinbase, you know, on chain and all of that stuff. What was the, what is the same about crypto and the AI world that we're in? And where is it different? Obviously we're getting into some of this on the data side, but are there any parallels that you see or where the future is going from your experience in this world? There's actually a ton of parallels here. If we're just speaking about the consensus piece, which we just mentioned before, this is a very well thought through problem in distributed computing systems.
24:20It's typically called the Byzantine general or consensus problem. And in that on-chain world, what you're trying to determine is how many of the nodes in a system need to agree on a particular transaction or state change before you can class that as being correct and accurate. and so compared to the problem that we're doing here it's actually more simple everybody has all of the knowledge that they need to make the decision because everybody holds a copy of the entire ledger and so everybody can literally just reference the ledger look at the transaction or state change is that possible can they do this yes or no commit and so it's more about kind of game theory of making sure people are not trying to maliciously extract value or change things that are happening.
25:06So it's kind of like one layer of the problem. The difficulty in the AI space is that we have that same layer at the top, similar layer, kind of different topic, but the problem is very much the same. And then we have this extra layer of complexity in that things could be subjective rather than black and white, and the participants don't necessarily have all the information they need to participate. And so we need to kind of layer in credentials to ensure that the right parties are able to kind of come in and so that we don't end up penalizing the one person that actually submitted the correct answer way up to the right while the entire rest of the ecosystem said no it's actually wrong yeah and so it's it's a more complex version of the same type of consensus problem the benefit that we have having built in the on-chain world for the past decade or so is that as a team we actually have quite a lot of experience thinking about this and So we can apply principles and learnings from peers and from other protocols to kind of attack this data problem in a different way.
26:11I think the other piece, and this is a bit of a topic switch, but the other big parallel that I would draw between on-chain world and AI world is that we're still in the phase where the incentives are such that you should or you want to create something completely new from scratch and ignore everything else that's already been done. And so the cheesy analogy here is like, imagine a world where when the internet was first created, the world was like, you know what? TCP IP, that sucks. Like, I don't want to use that. I'm going to go and create 500 different versions of it. And you end up with this like crazy fragmented mess.
26:52Like we wouldn't have the things that we have today if that had happened. That's exactly what's happening in the crypto world just now. You have thousands of different chains. That's not necessarily a bad thing. The bad thing is that none of these chains can talk to each other in a truly trustless way. And so again, you end up with disparate, weird, siloed stuff. In the AI world, we have exactly the same incentives set up. And so everybody is building their own model. Everybody is using proprietary algorithms. not necessarily proprietary compute, but most of the big ones use their own compute.
27:27And importantly, they're using proprietary data and proprietary annotation and inference and training. And so nobody is learning from each other. And we're ending up with tons of different models who are developing at different speeds and are good at different things. and uh yeah the downside here at least in my mind is that it's going to take significantly longer to get to a good outcome because nobody's working together it's like we've taken all of the smart people in the world and then given them each the same task and told them not to talk to each other doesn't make any sense yeah it is interesting how you see some of these patterns in nature in life where you have tons of different iterations and things happening and the best thing emerges at some point.
28:19And it feels very messy and out there, but it just kind of feels like it naturally happens. Looking back, it was just obvious. So it's kind of interesting how some of that looks like a base kind of principle, first principle type of thing. But I'm curious, I want to get into the topic of synthetic data. Before I jump into that, one thought that hit my head, like, obviously these are humans in this protocol that you're building that will be doing this verification and going through this process you discussed. But do you see a world where AI can actually be one of those participants in the review of data?
28:56And like, is that possible? Because you can give AI incentives just like you can give a human incentives. Like, is that something you see that's plausible at some point in the future? I mean, it's happening today. It's not plausible. It's like we're here. The way that we're designing our protocol is that we don't really care what type of data is needed. And we also don't really care who the participants are. And so it needs to be serviceable for a user, for a validator, for somebody that needs work done. And that somebody could be an MLOps person at whatever company, or it could be an agent that is on a mission and doesn't have everything they need to complete that mission.
29:42And so I think thinking in that way, even if we're not in a world where agents are quite at that level of autonomy yet, we're not going to be a million miles away. But to answer your question more directly, this comes into play most for the validation section. You could for sure create a model that would do the work, for example, but the the real prize and the real efficiency that's more likely to be gained is a bpo somewhere that might have a hundred people and they're setting up themselves as a business to validate specific types of data that they have expertise in and they layer ai agents over the top just to supercharge from a kind of efficiency perspective and provide an extra layer of intelligence before they submit said data and so i'd expect that's the way that it's going to go initially but there's nothing to say that it wouldn't be entirely replaced by very specific models in the future.
30:36And if you think about it, if we come back to this siloed analogy of everything not talking to each other, if ChatGPT, for example, builds a model that's incredible at dealing with health stuff, it's basically Dr. GPT, and Claude is terrible at it. And like, ignore what I'm saying here, I have no idea which one would be better. You could imagine a world in which Claude would welcome the input of ChatGPT's incredible medical knowledge when it's trying to upgrade its understanding of that field. And so there's no reason why you would exclude that type of knowledge transfer. But where we are today is that we're still in the phase of transferring knowledge from humans into these models.
31:17And so it's less in demand, but certainly something that I expect to start happening relatively soon. So let's get into the topic of synthetic data itself, right? Because there's a lot of people that feel like, you know, at some point you may have model collapse where it's almost recursive in and of itself and it starts to degrade. So how do you think about synthetic data and its impact on the future of models and training data and all of these things? Like, is there, I think this gets back to the point we made earlier where there may be many ways of doing this and there may be one logical one that emerges that's just obvious to us in the future.
31:59But what's your take on that? It doesn't matter to me as long as the data is accurate. That's the key thing. if a certain piece of data has been created by a person versus created by an agent or a model it doesn't make any difference providing said piece of data is actually true accurate high quality whatever metric you want to measure it by and so i think it's much more important to determine fact versus fiction than it is to think about where said input came from and so that's why we're not uh our protocol is essentially agnostic to who you are where you are what type of data is it's really just seeking is that true or not is that accurate or not yeah that's where sapien becomes a big part of that future in a sense but like there's uh you know a lot of people talking about oh this will lead to model collapse but it could be the exact opposite to where synthetic data is almost has this exponential positive benefits of the performance of the models to your points like it really doesn't and i've kind of shifted my thinking on this like i'm questioning does it matter who created the data right but i have not really thought about it in the sense that you just mentioned that's not the vector to look at it on it is the trust and accuracy and validity of the data not that's a better way to put it is the data valid that's a much better way to put it and that's i mean i'm using the word quality because everything that we're framing it's like we're basically creating a trust and quality layer for ai data but you're absolutely right it's like is that input valid or not and that should be the piece that matters and what happens or the reason that you have model collapse and hallucination and like bonkers outputs is because you're not sanitizing the synthetic data sets and removing the nonsense and so of course you're going to get a bad output all of these systems if you peel away the complexity they're garbage in garbage out and so you have to remove the garbage to get to a good quality output yeah and i've mentioned this example several times on this podcast but it ties back to what you mentioned earlier around incentives.
34:17So it was a study or research that OpenAI had done around hallucinations. And the problem leading to hallucinations deals with incentives because these models are incented to provide a plausible answer, right? And it's just like when you are taking an exam, if you don't know the answer, they tell you to answer C. ChatGBT and Claude and all these are working in a similar fashion versus the SAT where it's like, if you don't know it, don't answer, right and it actually works in your favor so if you change that incentive and then what's the alternative maybe the model asks you know uh for more context more clarification to provide a better response and this is more on the inference side but i think this incentives are such an important aspect of this entire system that's being built in a sense but i want to shift to i don't know if Do you have any thoughts on that point for us?
Read the full transcript
35:12Really quickly, I think there is one comment there. And I think the idea of a system failing safely, I think is really important. And it's kind of the missing piece in how many of the models currently behave. It is for sure part of the story for things like autonomous vehicles and humanoid robotics. When they don't know what to do, by default, they will just do nothing because that's the safest thing to do. But it's a really kind of common principle from the kind of non-software engineering world, whether it's infrastructure or aeronautical or mechanical. If a piece of equipment fails, it should fail in such a way that doesn't harm the equipment around it and importantly, the people around it.
35:55And so in case of a model, if it doesn't know, it should be very clear that it doesn't know because that's the safest thing to do. and for a car it should pull over and not do what the cars were doing in austin recently where they were driving straight past the school buses that had a little stop sign on the side and just tearing past like they weren't recognizing that they for sure were seeing it but they weren't successfully feeling safe they were continuing on that way yeah that's it's the uh the charlie bell um evp at microsoft talking about similar things on the security side it's like just assume whatever you're building will be breached or whatnot.
36:34And then make sure to your point, it's failing in a safe way or you have guards and safeguards around that. But I want to get your take on, what do you think the future is going to look like? We have several big frontier models. Do you think that's just going to be the future? We have several big models and people are fine tuning those specific to their use cases? or do you believe we will have more smaller models being used in niche ways? And then that also kind of bridges into the topic of open source versus closed. I think Meta just recently has kind of pivoted their position some on this, which is interesting, but curious your take on the small versus large, open versus closed, and then fine-tuning aspect.
37:23Yeah, awesome question. I think first off, anyone that gives a real answer here or a confident answer here is full of shit. Nobody has any idea. My best guess is that we'll end up with a ton of different models. And you'll basically use the one that's best at that particular thing. But I think more likely, each of us will find a model that we're just more comfortable and familiar with, that's really good in our native language. And then that will become the model that we individually use. And so I expect in different parts of the world, you'll have models that are more popular. But I don't expect that we're kind of coalesced towards one model to rule them all.
38:02I don't see that as the future at all. And then the open versus closed is super tricky. To me, the tech stack is too important to be closed source. Like it has to be open source. However, the incentives are so huge that I can't see it becoming truly open source. at least not soon. Hopefully, eventually it will become open source. But for now, I think it makes more sense for it to be closed. And it's certainly accelerating speed of development because we have competition and competition spurs speed. And so I guess in some ways it's good. In others, it's a black box. And that's kind of scary when you invite it into your everyday life.
38:45Yeah, totally. All right, I want to wrap us up with a few lightning round questions. So just first thing that kind of comes to mind. But what AI trend right now do you think is just mostly noise? Social. This is really hot in the kind of on-chain AI world, which, to be clear, is primarily grifty nonsense, unfortunately. Although there are a couple of really cool agentic builders out there who I have a ton of respect for, Wayfinder and Mamo and others, virtuals as well. But AI for social stuff is one of the most infuriating things ever. social media at least in my opinion was borderline unusable pre-ai and it is completely unusable post-ai the videos are just nonsense the random garbled replies horrible hate it yeah yeah i do think we will have this uh turn to human created human oriented as the different differentiated thing which is kind of awesome uh but what's what's the hardest thing about building infrastructure like the protocol versus building product any nuance there that you've learned along the way the difference between the two i think the difference in the right way to phrase this to me product is what much more pointed you're building something with a very specific set of kind of confines and so it's much easier to figure out what good looks like.
40:14When you're building a protocol and you're aiming for it to be as general purpose as possible, it's much harder to figure out what good looks like. And you need to speak to a ton more people to get towards that good. And that's the process that we're in today. Defining every single kind of touch point and ensuring that it is specific enough whilst also being general enough and simple enough that it's properly extensible and not just a nightmare to use. I think that's the hardest part on the protocol side. What should humans always stay responsible for no matter how good AI gets? And an acceptable answer here is nothing.
40:54It doesn't matter as well, if that's your view. Oversight. Oversight. Okay. Nice. A nice simple answer there. What's the biggest constraint on AI progress that is not technical? that's not energy related and whatnot? Availability of high quality training data. And secondary to that, regulation and guidance that enables forward development in a safe way. What's the hardest domain to validate from a data perspective and why? Or maybe is there a few domains that are just very messy or weird? It's weird that's always the hardest. If there's only a few people in the world that know the knowledge, you're doomed from the start and it's really hard to get confidence in the actual data set.
41:52And so anything obscure, really hard that requires specific expertise. Yeah. All right, last one for you. Give me your tool in the AI world that you personally cannot live without right now. I love this framing I heard the other day. If you have a good product, if you were to lose it or it got taken away, are you going tomorrow to go get the replacement to that thing, which is kind of an interesting way to think about product. But what's the one that you can't live without? I'd love to say something like anti-gravity or cursor or one of these like IDE cool builder things. Truth is, I'm not the right person to be vibing with those.
42:33For me, it's granola. I am terrible at taking notes and granola just saves my ass every single day. It's fantastic. If you haven't used it, download it. It's in the background. It listens. It creates really intelligent, concise notes and saves me literally every day. Love it. Yeah. That's funny. You mentioned granola. I'm a big user of that too, but the cool part, and this is like completely off on a tangent, but they have what they call recipes. So like mid meeting, you can talk to it to have it do something. Like I had an example where my daughter came in during one of my meetings. So I had to go step out and do something.
43:09I was like, catch me up. And it did. And it even has one that's like, make me sound smart. It'll give you relevant questions in the meeting. It's just, it's pretty cool. I should have used that before this. I know you could have been using it right now. Yeah. Right. Rowan, thank you for being on talking, talking a bit of AI with us. Where can people find you? where can they learn more about Sapien? And you all have some really interesting writings out there just around this proof of quality and whatnot that I think is worth checking out. We're starting to speak more openly about where we're going with proof of quality.
43:44Expect a white paper very soon and expect full documentation shortly afterwards. Websites are currently kind of migrating from service world to a much more dev-oriented landscape. sapien.io main website earn.sapien.io is our kind of legacy app if you like i don't think anybody wants to find me but should they rowan rk6 on x awesome thank you rowan thank you thanks for listening to the talking ai podcast if you enjoyed the show give us a follow or subscribe on your favorite podcast platform and don't forget to leave us a review we love those for more info on Talking AI, visit TalkingAiPodcast.com.
44:29Quick break in the pod. If you're listening to this podcast, chances are you've been thinking about how to actually use AI inside your business. And that's exactly why we built the AI Opportunity Finder. It's a free tool that helps you uncover high impact, tailored AI use cases based on your business, your goals, your pain points, and your industry. No fluff, no generic use cases, just real ideas that fit your business and they're ranked by ROI potential. It takes about three minutes to run and it's like having your own personal AI strategist for free. If you want to try it for free, check out the link in the show notes or go to hatchworks.com backslash AI dash opportunity dash finder.
From the publisher
In this episode, Matt Paige and Rowan Stone, CEO of Sapien, discuss the critical importance of data quality and provenance in AI.
Stone, who has experience with on-chain products at Coinbase, introduces Sapien's innovative approach to building a decentralized data protocol that emphasizes 'don't trust, verify' principles.
They explore avenues such as incentives, validation methods, and the peer review process used by Sapien to create high-quality datasets.
The discussion touches on the implications of bad data, the role of synthetic data, the complexities of achieving accurate AI outputs, and the parallels between the AI and crypto worlds.
Key insights are shared on how to ensure models perform safely, the hurdles in the industry, and the trajectory of AI development.
Additionally, Stone provides a glimpse into Sapien’s efforts to demystify data validation and enhance the transparency and trustworthiness of AI applications.
--
Key Moments:
- 01:04 The Importance of Data Quality in AI
- 03:32 Challenges and Risks in AI Development
- 07:08 Sapien's Approach to Data Validation
- 08:35 Incentives and Trust in AI Systems
- 13:30 Building a Decentralized Data Protocol
- 23:22 Consensus and Collaboration in AI and Crypto
- 30:55 The Role of Synthetic Data
- 36:17 Future of AI Models and Open Source
--
Key Links:
Mentioned in this episode:
Free report from HatchWorks AI — State of AI 2026
What’s real in AI this year, what’s hype, and what leaders should prioritize — including production lessons, designing for agents, and governance. https://hatchworks.com/state-of-ai-2026/
AI Opportunity Finder
Feeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you’ll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action. 👉 Try it now at https://hatchworks.com/ai-opportunity-finder/
