In short
Eye On A.I. Podcast Episode #227 Summary
Episode Title Sedarius Tekara Perrotta: The Importance of Data Quality in AI Systems
Host Craig S. Smith
Episode Overview In this episode, Craig speaks with Sedarius Perrotta, co-founder of Shelf, about the significance of data quality for AI systems. The discussion revolves around the challenges posed by unstructured data and highlights the innovative solutions provided by Shelf to enhance the performance of AI tools like Microsoft Copilot.
Key Themes and Concepts
- Unstructured Data Chaos
- Definition: Unstructured data refers to information that does not have a predefined data model or is not organized in a predefined manner.
- Importance: It acts as the "fuel" for AI systems, and its quality directly impacts the effectiveness of AI initiatives.
- Shelf's Mission
- Objective: To provide transparency into unstructured data, helping organizations manage and optimize their data quality.
- Focus: Shelf's solutions aim to tackle "garbage in, garbage out" problems, ensuring accurate and reliable data for AI applications.
- Data Observability and Monitoring
- Approach: Shelf implements data observability, enabling organizations to monitor the quality of their data in real-time.
- Benefits:
- Identifying inaccuracies, duplications, and outdated information.
- Facilitating data-driven decision-making and improving AI performance.
- Tackling AI Hallucinations
- Issue: AI hallucinations occur when AI systems generate incorrect or misleading information based on poor data quality.
- Solution: Shelf's technology helps in fixing these hallucinations by ensuring that the data feeding into AI systems is accurate and trustworthy.
- Generative AI Adoption Insights
- Trends: Rapid adoption of generative AI among organizations, with over 70% having active projects.
- Future Forecast: A significant increase in production-ready generative AI initiatives is expected by 2025.
- Data Management Strategies
- Recommended Practices: Organizations are encouraged to build a strategy for clean and trusted data as essential for scalable AI initiatives.
- Crawl, Walk, Run Approach: Start with manageable data projects to build confidence and scale as trust in the system grows.
Key Takeaways
- Transparency in Data: Organizations must focus on gaining transparency in their data assets to facilitate effective AI implementations.
- Quality Assurance and Monitoring: Continuous monitoring and quality assurance are critical to maintaining the integrity of data used in AI systems.
- Strategic Data Management: A proactive strategy for data management will help organizations leverage their unstructured data effectively and reduce reliance on outdated or inaccurate information.
- Future of AI: The intersection of data quality and AI technology is crucial for the success of future business initiatives, particularly as generative AI continues to evolve.
Conclusion Sedarius Perrotta emphasizes the critical role of data quality in AI systems, advocating for organizations to adopt robust data management practices. As AI technologies evolve, the need for clean, structured, and trustworthy data will become increasingly vital for achieving successful outcomes in AI-driven initiatives.
---
Stay Updated
- Craig Smith Twitter: [@craigss](https://twitter.com/craigss)
- Eye on A.I. Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
Episode Timestamps
- 00:00 Introduction and Shelf's Mission
- 03:01 Understanding SharePoint and Data Challenges
- 05:29 Tackling Data Entropy in AI Systems
- 08:13 Using AI to Solve Data Quality Issues
- 12:30 Fixing AI Hallucinations with Trusted Data
- 21:01 Gen AI Adoption Insights and Trends
- 28:44 Benefits of Curated Data for AI Training
- 37:38 Future of Unstructured Data Management
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00We did it out of need, out of pressure, out of pain of being able to deliver more value for our customers because out of the box that didn't exist. So we built this observability layer and a lot has happened since then. We were founded in 2017 or we went to market in 2017. This need for transparency into data, especially unstructured data or files, different types of documents. This has been needed whether or not Gen.AI has come to fruition. It was relevant. It has been relevant since digitalization. And really what has happened with Gen.AI is a problem that has been old, that a lot of people are very, very familiar with.
0:36of what is in my file system, what can I trust before I hook it into Gen.AI. This has come to the forefront, and this is where we're in the right place at the right time, solving a really pertinent problem for scaling Gen.AI in businesses globally. So my journey into this space all started at the US Peace Corps. I was in Transylvania, of all places, and I was working at a think tank. So I was responsible for knowledge management at a think tank. And I did that for several years in my service. And then I went out and became a consultant doing knowledge management projects around optimizing the value of intellectual capital and best practices, lessons learned, key strategies, all that fun stuff.
1:21and did projects for Harvard, MIT, Stanford University, Society of Organizational Learning, and a whole bunch of stuff for the World Bank. And that led me to spend time at MIT, honing my craft at the Sloan School, where I really got into, that was when I got into AI and NLP specifically, because when you do knowledge management, it's very manual and tedious, and was looking for ways to optimize and speed up that process of getting quality data, because it's always been the problem. The problem with knowledge management or any, let's call it any file system management, is the quality of what's in it.
2:04People create multiple versions. They create and they put it in different places. Things go out of date very quickly. Some things are the approved version. Some things are not. things are non-compliant in there and there's really no transparency into it. So I started a consulting company. It was called Neuron Global and that's what we did. We did knowledge management consulting. We served hundreds of companies on knowledge management consulting projects and we really got specialized in SharePoint. So got to experience SharePoint firsthand. At the Ignite, the Microsoft keynote, they actually said that 2 million SharePoint sites are created every day and 2 billion files are uploaded to these sites every day.
2:53It's just the scale is... Yeah. And for listeners who don't know SharePoint, just give a little thumbnail. Yeah, it's Microsoft. There's two major file systems in Microsoft. There's SharePoint, which is a little bit more customizable, which enables you to create portals and communities. And then there's OneDrive, which is more of a Google Drive, Dropbox, Box version of that. So we're specializing in that portal version, which is really knowledge management, kind of 101. And we were doing that for a number of years and ran up into this wall of data quality because everybody's very enthusiastic at the beginning of the projects.
3:36You get a lot of buy-in on the executive level all the way down. And then it gets launched. And then months later, stuff starts going out of date. People find things. It's not what they're looking for or it's the wrong thing. And then they stop trying to use it. And then this negative feedback loop continues until adoption becomes very difficult and hard to enforce. And that was really where we needed to found another company that sat as an observability layer on whatever file system or whatever knowledge system was being used, whether it was Dropbox, Drive, SharePoint, Salesforce knowledge, Zendesk knowledge, whatever it is, wherever people were putting data, we needed an observability layer so that you knew you would be able to identify what was accurate, what was up to date, and what was trusted and what isn't.
4:23So that was the founding of it. We did it out of need, out of pressure, out of pain, of being able to deliver more value for our customers because out of the box that didn't exist. So we built this observability layer and a lot has happened since then. We were founded in 2017 or we went to market in 2017. And then, you know, we were this need for transparency into data, especially unstructured data or files, files and different types of documents. This has been needed whether or not Gen AI has come to fruition. It was relevant. It has been relevant since digitalization. And really what has happened with Gen AI is a problem that has been old, that a lot of people are very, very familiar with, of what is in my file system.
5:15What can I trust before I hook it into Gen.ai, this has come to the forefront. And this is where we're in the right place at the right time, solving a really pertinent problem for scaling Gen.ai in businesses globally. We were talking about SharePoint and what SharePoint does. And the problem is when you get your data uploaded into a knowledge base, then things start going haywire because things get out of date. you find that they're duplicates and that sort of thing. So let's go on from there. Yeah, so I think there's just a natural data entropy that happens whenever you are trying to save documents anywhere, really, whether it's in OneDrive, SharePoint, Google Drive, Box, Dropbox.
6:07It's just this natural problem of trying to manage this chaos of everybody creating documents at different times and different versions and things going out of date and out of compliance. Anyway, this is a real problem for Gen AI, especially RAG initiatives, retrieval augmented generation, and the Microsoft Copilot because they need that data in order to retrieve answers. And that is essentially the major difference between web-facing content and internal company content because there's a veil across an organization's documents and files. and it hasn't been pierced because there's no transparency.
6:46There's no analytics. There's no way to know what's there. So this is a major impediment to Gen AI at scale, and this is what we work on solving. And so that's interesting about RAG. I was thinking it was more preparing data for fine-tuning or even basic training of models, But it's data in a knowledge base that an LLM or some other gen AI application is accessing. Is that right? Yeah. Generally speaking, and these are techniques that companies are using to improve the RAG responses as they're fine tuning the models. But they're trying to work around bad data because most organizations have kind of put their hands up and said, I don't know what to do about this.
7:39There's no transparency into this. I don't know how to fix it, but we need to find a way forward because we see the potential. We see the awesome benefits of Gen AI in our organization to automate more, increase productivity, increase efficiency, increase competitiveness, all those things. So they kind of work around it. But that's exactly the problem that we solve is the data quality issue, the garbage in, garbage out issue. And that is monitoring and keeping track of what is in those documents and files and what can you trust and what you need to filter out. And how do you do that? I mean, presumably you're using AI to do that as well.
8:19We're using AI, we're using chains of experts, NLP, ML. We have over 150 autonomous agents behind the scenes working for us. So it surfaces in 22 algorithms that are specifically designed to, in the sections of documents, so not, you can't compare a document to document that's not very helpful, because a question is being asked, and it's in a particular sentence or paragraph on a particular page, and that's the area that you need to dial into. So we're breaking it down into those sections, and then we're looking for, we're comparing those sections to one another, and identifying duplication, identifying conflicting, identifying areas that don't have context, identifying areas that are non-compliant or toxic or bias.
9:09And then we're surfacing that on an interface so that people can actually see for the first time the quality of their data and can be empowered to take action. And when you say it looks inside the document, if I have one document that says, I don't you know, Craig's birthday is in October and another document that says Craig was born in March. It'll find that? Correct. That's exactly what it would find. That would be in our system on conflict. So you would link it up to whatever sources you have, any source. And then it's this quality assurance layer that sits on top. And in addition to that, you are able to, in real time, view your Gen AI answers.
10:02And we're able to know when an answer was derived from a document that has our section in a document that is bad, that we've identified with our 22 different scores as a risk. So we can tell you exactly when an answer has been generated that is not trusted, that is not accurate, that is not compliant. So that in real time, you have the power to fix things and to improve your RAG and your Microsoft Co-Pilot initiatives. Yeah, that's interesting. But can't you put this, you know, let it run on your knowledge base and reconcile all of the problems before connecting it to an LLM? Or is that just too big of an ask and you really do it live?
10:56Yeah, I mean, it's crawl, walk, run. Certainly that's where we're going, where you can automate more and more and more and then just automatically fix things. Right now, what we do is we identify it so humans can determine whether they're going to fix it. But humans could also set up filters and let's just say plays that serve as gates of content going into the LLM. So you can automate the filtering out of content and you can automate the creation of labels and named entities as well. So that the information, the sections have more context for that LLM. So there's things that you do that you can do in real time, and there's things that you do that accumulate over time and improve accuracy day by day, month over month.
11:44And so if someone's building a knowledge base for a RAG system, would they first do a pass with – and is the product Shelf or what is the product name? Shelf. Shelf. Okay, they would first do a pass with shelf. They would correct things that they know how to correct. And then shelf would run kind of in the background while during inference. And then say, you know, give the user an alert that this may have a problem. them be careful or look more deeply into it. Yep, that's correct. There's really several use cases. That's the very proactive one. We see more of a use case where they've already implemented a RAG or a co-pilot that they're having hallucinations and they want to fix those hallucinations because they can't scale it.
12:48So that's the more common use case is people saying, we have an issue, Houston, there's an issue here and we need a solution. So a lot of times we come into projects where they've been trying to do this manually. And then when they see this opportunity to automate a whole bunch, to actually see what's going on, there's a lot of resonance with that. And we have very happy customers. But yeah, you can also do it before you try to even launch a Gen AI initiative. And that's called data readiness, Gen AI readiness, responsible AI initiatives. Those are what those hook into. And a lot of organizations are doing those as well.
13:30Yeah. You guys started this company in 2017. I mean, that's recent history. But in the world of Gen AI, that's ancient history. What were you focused on before Gen.ai came along? And when did you develop this solution for RAG? Interestingly enough, we've been doing the same thing since we've launched, which was data quality, which is transparency into unstructured data. Like, what can you trust? What's accurate, trusted, up to date? So we've been always doing that because of the story I told you of not having any transparency in other projects when we were doing it from a consulting standpoint.
14:17What has changed is that humans are no longer the consumers of the good data, of the algorithms, of the analytics that we're running on top of that data. It's now it's LLMs are the consumption. consumption. So it's just accelerated something that we've been always doing and has always existed. There's always been bad data and people have been searching and finding the wrong answers. That's a known problem. It's just everything accelerated in the age of Gen AI because people realized that this was kind of their goal. This was their foundation for automating more and going faster and it has accelerated our, let's say, adoption and accelerated the market need for something more general.
15:08Like we started in the contact center where there's ROI to be proved in shortening a call, in reducing the amount of escalations there are, in improving CSAT. So this is where we've been because there's ROI attached to it. But every other department has been having the exact same problem. It's just harder to prove the ROI. And now what Gen.ai does is it enables an ROI, a demonstrable ROI out of the gate for any of these because we can see exactly how many bad answers there are and show them fixing them in real time, how much impact that they've had on their hallucination and answer accuracy. You know, I had a conversation a couple of weeks ago with a guy about data ontologies and how important it is to, you know, to build an ontology so that particularly for AI's, you know, gen AI consumption.
16:09So there isn't confusion about what different things mean. And is that part of this? Does this help in building that ontology? Of course. An ontology on a knowledge graph is the architecture and backbone of, I think, any real accuracy initiative, any hallucination reduction initiative. You really need to build it on that because it shows you the relationship between all the parts, all the sections of your data. And then on top of that, the metadata creates relationships that can't exist otherwise. So, yeah, of course, we have a graph, we have named entities, we have ontologies, we have custom taxonomies, topic modeling, all these things that are, it was part of the Web 3.0 back in the day of Tim Berners-Lee, all these really cool technologies like the OWL framework and whatnot and RDF and all these other things that just didn't take root because there was an ROI connected to them.
17:12but they were always known to improve the quality and the retrievability of accurate answers. It's just the investment wasn't there, the appetite to invest because it was very hard to prove that ROI. Now, with the speed of everything and running these models in real time and having all those conversations, you're able to see exactly when something is correct and accurate, something is not that helpful. Maybe it's kind of accurate, but not exactly what you need. And then there's the human in the loop there. And then things that are just flat out either wrong because they're pulled from bad data or wrong because the model hallucinated an aspect of the question and the prompt wasn't clear enough.
17:56And how running this, the shelf system on your knowledge base, how is there some metric about how much you can reduce hallucinations? Yep. There's real time monitoring of answer quality and the improvements that you're making using the platform. So yeah, there's a real-time answer quality dashboard where you can see exactly the number of inaccurate answers that have been produced and the source of those, and that you could either fix, filter, or enrich a bad answer, essentially, the data that's connected to that answer. And you fix and enrich in real time, or does it get shunted off into a queue that somebody else is working through?
18:44Yeah, it just depends on the organization and how they're set up. So someone can fix it right there and then, especially if it's something as simple as there are two documents that are very similar, like two products. Think of health care, like a health care plan. They're very, very similar, except some numbers that kind of mean everything. And the titles are even very similar. So you would want to put a named entity on that, telling the LLM that these are actually two different products. And that if you don't know the answer or you're going to try to pull from either of these, ask a clarifying question or essentially create a rule.
19:25When this particular plan is used, it follow these steps. And if this particular plan is used, follow these. But it's a way to have more control over the outputs of your Gen AI answers. Yeah. How does somebody get started with Shelf.io? It's pretty easy. There's three steps. First, you would connect any of your data sources. So connect your SharePoint, connect your OneDrive, connect your Google Drive. We have dozens of connectors out of the box and a connector framework. And then two, it's going to start running and identifying issues. So you have that dashboard. And then the third step is you get to fix, filter, or enrich and start impacting the answers being generated.
20:09And is this, how do you charge for this? Is this a sort of a pay-as-you-go? Are there different tiers, I imagine? Or do you sell seat licenses? It's a consumption model based on different tiers. So it's the number of pages processed, essentially and uh our starting plan is at like 3 000 a month and then it goes up from there yeah uh you you recently did a survey of uh of companies on on where they are in their gen ai journey uh specifically around the issue of uh data uh can you give us some of the insights Yeah, sure. I mean, one of the things that is very interesting, and we heard this from a major implementer at the Microsoft Ignite event.
21:11And the quote is, six months ago, I wouldn't know I had a problem. Today, this is very relevant. So I think one of the things that the survey indicated was how fast the market is moving. The adoption curve is just, it's moving a standard deviation to the right in six months, which is crazy. See how fast things are going. And that another thing that was really interesting was the number of active Gen AI projects. So over 70 % of the organizations have active GNAI projects, not in planning, but active projects. And then of those, the vast majority are still in the POC stage. And I believe Satya Ndodal said at the Microsoft keynote that 10 % of the market right now is ready to go into production in 2024.
22:12But in 2025, that number is being forecast to be 80%. So the speed of this market is blazing. And the other key takeaway I was surprised was that a lot of organizations do realize that they have unstructured data quality issues. They understand that the fine tuning and trying to chain different technologies together isn't the long term solution. It's a solution to trying to get something work right out of the gate and very quickly, but it doesn't scale. So I thought that was another really interesting takeaway from the market. Another thing I can say was the pure scale of most organizations. So over 80 % of the organizations polled had more than a million documents and files.
23:07Yeah, I remember that. Over 51 % of them had over 10 million. So just the breadth of that problem of how to know what you need to address first and how do you even try to manage such a huge number? You can't do that manually. Yeah. And what was happening before solutions like Shelf? uh i mean was it you know you have a a million documents but maybe you know it's kind of like my my office you know there's stacks and stacks of documents it's really only the top inch that that is active the rest are you know virtual archives uh yeah how yeah how does uh How does that relate to what you're talking about?
24:03Well, I think it's a complex problem, right? Because there are some cases where, let's say, take customer service. Customer service has a manageable amount of documents that are serving inbound questions coming from the customer. So the agents are only using thousands of documents, not millions of documents to answer questions. And when it's thousands, you can actually curate and fix the problem. You can't get 100 % accuracy, but you can get darn well close. So in that particular use case, which that's where we were before Gen AI, you can actually quickly, relatively, get to a trusted knowledge base.
24:50You can get to a trusted, secure body of knowledge, corpus of knowledge to support a use case. Now, then there's technologies that are like federated search or semantic search, which go across an entire organization. And I think the users of these technologies have a sophistication because they know that they're using it. and when it pulls something, the document that they're not looking for, like whenever you're doing, let's say a Google Drive or Dropbox search, you're like, no, that's not it. You're mostly using your human memory to identify whether it's the right version or not, but they're not particularly accurate because they don't know what the right version is.
25:35So the larger the corpus of knowledge, the larger the document base that you're trying to work with, the greater the problem is and the harder it is to solve. And that actually goes into what we advise the market to do. And that's crawl, walk, run. You shouldn't try to consume the entire beast out of the gate. We have companies that are like, oh, we're going to give you 1.5 petabytes of data to quality assure. And it's like, we can do that. But what are you going to do? What are you going to do with the output? We're going to identify tens of millions of issues. Do you have people that are going to be able to fix tens of millions of issues?
26:11A better move is to identify what use case you're trying to address. So, for example, it's an onboarding use case. Okay, great. We have an onboarding use case. We're going to make it easier and faster for people to onboard, to self-serve, to learn all the core company cultural knowledge and processes and procedures. And, you know, it's very similar. Then you identify which documents, which folders, which files that you actually need to support that. You port it over, you put the correct permissioning around that, and then you manage that because that also changes over time. It's not like you make this effort to get clean and accurate data and then you're done.
26:52No, because time passes and new products are launched and new people join and new directors put in new policies and procedures. Things are constantly changing. So for it to be accurate over time, you need to be monitoring it in real time as well. But to start with something small, to make it work, to get that confidence, to get people trusting the system, and then to secure it, and then to scale into more and more use cases. This is what we really see working. And this is what we see a lot of the consulting partners that we work with. This is what they're advising to. Yeah. And beyond RAG, would this work for creating a clean training data set so that you can, because that's part of the problem with hallucinations is there's so much bad data in the training data.
27:49and the models are really looking at probabilities based on what's in the training data. But if you got rid of all of the anomalies in the training data, then presumably, and I've thought about this with the coding assistants, if they were trained, because the first ones were, it was almost an accident. right uh they they train up this massive llm i think it was gpd3 and they realize hey this thing can actually read code uh but yeah can you talk about that about you know sort of training new models from scratch with better curated training data and and how that would have would apply to shelf.
28:44Yeah, there's two other use cases that I think are really interesting for your audience. I mean, there's the actually curating and pruning and caretaking the data that you want to use. So if it was 1.5 petabytes, then that use case of identifying the worst offenders in that obviously will have an impact and will be valuable. There's no doubt about that. So that's certainly an option. Anything that you're doing as far as data quality management before you're trying to do the actual AI task is going to have downstream benefits. And then additionally, the autonomous agents. So if you're creating experts on certain tasks and chaining them together, the data that they're consuming also benefits greatly from being pruned and caretaken before it goes into being automated and expanded very quickly.
29:41So they're really fixing the data and investing in the data. It just has downstream benefits for so many different use cases. It's why the structured data management space is so robust. And there's so many hundreds of companies there and huge businesses like Snowflake and their ecosystem and Databricks and their ecosystem. Because people understand if you invest in that, you're going to be able to use it in whatever use case that you want downstream. So you're kind of fixing things at the very source. And then where you pipe that is wherever you pipe it to, you're going to have more value, you're going to have more accuracy, you're going to have more trustworthiness.
30:18So it's the same model for unstructured data. The more you invest in fixing it upfront before you use it for whatever different use cases, the more valuable those use cases are going to be out of the gate and the less work you're going to need to do later. It's kind of like in the software world, having a really good specification versus not, and then designing something and building something and then finding bugs later, and then having to fix the bugs, it's like 100 times more expensive than just having designed it in the front. In the data management space, it's the same thing. You just fix things upfront, you save a lot of downstream headaches.
31:00And it's easy for me to say, because people have urgency to get something out of the gate and it takes time to do that. And that's why we kind of want to address that at both places, at the responsible AI and the data cleansing. But also, hey, we have a Gen AI project right now. We're having hallucinations. We need to improve the answer accuracy. You need to be able to do both. You need to be able to straddle both worlds because different companies are at different stages. And even different Gen AI initiatives within a company can be at many different stages. And you need to be able to walk with the market and enable them at whatever stage they're in.
31:37Yeah. You know, we were talking, you were saying, you know, a million plus documents and to focus on a particular problem or use case first. Is there, and a lot of that, those documents would be stale at this point of the million or 10 million documents. is there still value to be had in that so-called dead data, data that hasn't been accessed for a long time? That's a great question because we've heard some people say, just throw away all your data. As a knowledge manager, as my background of where I came from, that like sends my, you know, shivers down my spine because there are best practices, lessons learned, and key strategies.
Read the full transcript
32:26There are nuggets of gold in all of that. So just throw it away means you're just dumping out all of the learning, all the organizational learning that you've accumulated over many, many years from different, different people with different skill sets and specializations. You're saying, oh, we're just going to flush it down the toilet and we're going to get no value out of that. And I think that's a big mistake. There's so much there that can be harnessed and used. Now, do you want to use the duplicates of those of a document that was created five years ago when there's 10 versions of it? You want to use all 10 versions?
32:59No, you want to use the one that's correct. Do you want to get rid of all the information on policies and procedures around product literature? No, you can use that for R &D. You can use that for system design in the future. There's some really interesting use cases. I believe it's from P &G and IKEA, where they're using old information as inputs to design new products. old information about products and different markets that were very successful at the time, that a human brain couldn't possibly compare all these different markets and what was good and what was bad and what was good about the design or bad about the design.
33:40But when you input it into a model and you turn the model around that particular output of designing new furniture or new product based upon what has worked in the past, you can have remarkable results. And I think it's the same for organizational knowledge, whatever it may be. You can learn from old sales literature. You can learn from old marketing position. You could learn from old CS information. Not all of it, not all of it is equal. But to throw it all away is really a big mistake and a huge lost opportunity. And I would imagine the organizations that just throw it away, they're going to lose a competitive advantage to the ones that actually put in the time to get value out.
34:18Yeah. Yeah. Yeah. You guys started before Gen.ai, but there are a lot of companies coming into this space. I heard, well, I won't name them, but the space is, I don't know if it's crowded, but I see companies talking about, you know, mining your unstructured data or cleaning your unstructured data uh how do you guys see the market is it such a massive market as your survey suggests that there's plenty of room for plenty of players uh or or do you think it's getting crowded or you know well at like i'll answer that on several levels um first of all um this is a needed and known problem for the entire market.
35:17So every organization has unstructured data. And as the survey indicated, what was it? Over 80 % identified that they have at least 50 % of all of their documents have inaccuracies in them. So there's no organization that doesn't have inaccurate unstructured data. That is not possible because we are humans and we create things and we are flawed by nature. So first of all, the unstructured data as a repository, it's a known issue there. It is not 100 % accurate for those 10 million documents. So that's the first thing I would say. Second thing I would say is it is a big dynamic market, and there are a lot of different technologies needed to solve the problem.
36:06There isn't a one-stop shop to fix everything. What you're going to want to do is you're going to want to construct a tech stack around the very specific problems that you're trying to solve, the use cases that you're trying to bring to market. And I see we are very well positioned because of our we've been we haven't built this technology in the last six months. We built this technology over the last seven years and we've invested in this and this we are practitioners in this space. So we have a point of view and a perspective on the unstructured data, and we have pattern match around it, which enables us to bring more value than you would imagine straight out of the gate, straight out of the box.
36:49And that's really where our position, that's where we think we're really well positioned in the market overall. The technology is continuing to evolve. I mean, Gen.ai, for production at least, or LLMs, I should say, have really only been in the economy for, what, three years now, four years. And the tools that you're bringing to bear, you were talking about increasing automation. information uh where do you see this this technology going i mean specifically technology to uh to manage unstructured data well i i think to answer that you first have to see where where is the technology going in general uh and in the where is it going in businesses where's the most value to be had and i and i i mean jensen hong and and all the the the ceos of the hyperscalers are are essentially saying the same thing on this, is that you're automating the easier tasks first.
38:01You're taking on work that is repetitive, that is manual, and you're bringing in slight automations using the autonomous agent technology or the expert technology. And you're taking more and more workload and offloading it to AI, to different sorts and streams of AI. And whatever the organization strategy is on that automation, which I think every organization has it on their map because it's I'm not saying anything new here. It's what enables that. What what enables our organization? What enables us to succeed the fastest? What enables us? First of all, what are the first things that we should start automating?
38:44Do we want to start in the customer service function? Do we want to start a marketing? Do we want to start in HR? Do we want to start in a horizontal function? Do we want to build experts that improve the efficiency of our SEO optimization? I don't know. The use cases are almost infinite in imagination. So you're identifying what you're trying to fix, knowing that in the future, more and more things are going to be automated and there's going to be orchestration layers of these agents, agent swarms, agent frameworks. Whether we go to general AI or not, I think it doesn't matter at this point. What is very clear is that there will be an increased amount of capability year after year towards more and more automation with more and more confidence.
39:43So 2025 is going to have more confident, more secure, more proficient, effective Gen AI initiatives in 2024. And likewise, 2026 will follow that same trend until it gets to a point where you can't forecast anymore. You don't know what is happening because so many different technologies are converging. That's another really interesting thing. You have the biotech, you have Internet of Things, you have fabrication and manufacturing, you have robotics. I mean, it is an interesting time to be alive. And I say all that because all of it's built on unstructured data, because LLMs consume unstructured data.
40:29They're fine tuned on unstructured data and they output unstructured data. So it's the oil. It's the fuel source of the machinery of the future. So it is a critical component. It isn't the only component. Like if you're building a jet, there's a lot of other parts and they all got to be right. But if you try to launch that jet without fuel, it's not going anywhere. Yeah. You mentioned at the beginning of Microsoft a few times. What's the relationship with Microsoft and with Copilot? And yeah, tell us what you're doing there. Yeah. So because of where we came from, because we were doing a consulting practice for SharePoint instances, essentially advising on how to build them and how to optimize them and how to make things more findable.
41:22We've had a long relationship with Microsoft and we've seen them grow and be extremely visionary in this space. I think they're one of the most visionary companies at the moment as far as where they're going. There's a couple others, but they're certainly one of the top ones. And their strategy, as shared last week at their major event, Ignite, at the keynote, their strategy is to invest in these co-pilots, to invest in faster, more automated workflows. whether you're creating a document, whether you're editing an image, whether you're creating a presentation or you're doing Excel sheet, or you're trying to query SharePoint or OneDrive, whether you're trying to use meetings or Teams, Microsoft Teams.
42:19They want to integrate and they want to be part of everywhere where people are working. and as I shared before, they have great ambitions, they have incredible products, but they all need unstructured data to be accurate up to date and trusted in order to function properly and for people to use them and to adopt them and to scale them. And that's where we come in. We're an enabler of that. We want to work with Microsoft and with these co-pilots and with companies implementing the co-pilot studio and AI Foundry of Microsoft to build these technologies. We want to be part of that conversation to enable them to deploy more accurate co-pilots at scale.
43:05Yeah. And are you then integrated into some of those products or are there extensions or plugins or something that are available? Yeah, we are integrated into Co-Pilot Studio, into AI Foundry, into SharePoint, into OneDrive. we can be a connective fabric between all these different technologies and the pipe work of unstructured data that connects them. The bigger point that I'd like to make in considering AI for whatever organization's roadmap for 2025, the greater point is that this space is moving very, very fast that the LLMs are evolving almost by the week. And there's like a mini war of the LLMs of who's going to have the best one.
43:57And there's mini models and there's custom models and all these things. There's frameworks being used to patch them all together. There's different use cases of the agents and the RAG retrieval augmented generation and building one's own models and so forth and so on. And I guess just one of the things that I would want people to leave with is the understanding that all of these are powered by the same fuel. And if that fuel is crude, if it is unstructured, if it's crude unstructured data, the engines aren't going to run properly. So no matter what the use case is, no matter what you're looking at implementing in 2025, I would just implore you to just think about what is your unstructured data strategy?
44:41What are you going to do to make sure that the information being piped into these initiatives is accurate, up-to-date, and trusted. And I think that if people have that thought and going into the planning and going into the new year, I think they'll be in a good position to succeed in 2025 and beyond.
From the publisher
In this episode of the Eye on AI podcast, we dive into the critical issue of data quality for AI systems with Sedarius Perrotta, co-founder of Shelf.
Sedarius takes us on a journey through his experience in knowledge management and how Shelf was built to solve one of AI’s most pressing challenges—unstructured data chaos. He shares how Shelf’s innovative solutions enhance retrieval-augmented generation (RAG) and ensure tools like Microsoft Copilot can perform at their best by tackling inaccuracies, duplications, and outdated information in real-time.
Throughout the episode, we explore how unstructured data acts as the "fuel" for AI systems and why its quality determines success. Sedarius explains Shelf's approach to data observability, transparency, and proactive monitoring to help organizations fix "garbage in, garbage out" issues, ensuring scalable and trusted AI initiatives.
We also discuss the accelerating adoption of generative AI, the future of data management, and why building a strategy for clean and trusted data is vital for 2025 and beyond. Learn how Shelf enables businesses to unlock the full potential of their unstructured data for AI-driven productivity and innovation.
Don’t forget to like, subscribe, and hit the notification bell to stay updated on the latest advancements in AI, data management, and next-gen automation!
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Introduction and Shelf's Mission
(03:01) Understanding SharePoint and Data Challenges
(05:29) Tackling Data Entropy in AI Systems
(08:13) Using AI to Solve Data Quality Issues
(12:30) Fixing AI Hallucinations with Trusted Data
(21:01) Gen AI Adoption Insights and Trends
(28:44) Benefits of Curated Data for AI Training
(37:38) Future of Unstructured Data Management




