In short
Talking AI Podcast Episode Notes
Episode Title
Transforming Data from 2D to 3D with Graphs and AI
Host
Matt Paige
Guests
- Jon Brewton, Founder & CEO at data²
- Daniel Bukowski, CTO at data²
Episode Summary In this season premiere, Matt Paige discusses with Jon Brewton and Daniel Bukowski the transformative power of graphs and AI in data analytics, particularly for high-reliability organizations. They explore how knowledge graphs can convert flat, 2D data into a 3D landscape of interconnected insights, enhancing data transparency and explainability.
---
Key Concepts
- Transformation from 2D to 3D Data
- Traditional Analytics: Traditionally, data is represented in flat, 2D spreadsheets.
- Knowledge Graphs: They allow for a more nuanced understanding of data relationships by visualizing data in a 3D format, showing how different entities are interconnected.
- High-Reliability Organizations
- Defined as industries with zero tolerance for failure (e.g., defense, intelligence, energy, finance, healthcare).
- Require tools that provide transparency and explainability in decision-making processes.
- Knowledge Graphs vs. Relational Databases
- Knowledge Graphs:
- Represent data as nodes (entities) and edges (relationships).
- Enable easier visualization and access to interconnected data.
- Relational Databases:
- Use tables and require complex joins to understand relationships.
- Often lead to inefficiencies and longer query times.
- AI Integration
- The combination of knowledge graphs with Large Language Models (LLMs) is discussed, particularly how it enhances the reliability of AI outputs and reduces hallucinations.
- Retrieval-Augmented Generation (RAG): A method of enhancing LLMs by providing them structured data to improve accuracy in responses.
- Real-World Applications
- Examples include complex scenarios in high-stakes industries (e.g., intelligence agencies) where understanding intricate data interrelations is critical.
- Data²'s work with intelligence data showcases how their platform can answer complex questions with higher accuracy and reliability.
---
Key Moments Highlighted
- The Journey of Data²: Transitioning from traditional analytics to a cloud-agnostic, zero-trust AI platform.
- Demystifying Knowledge Graphs: Providing clarity on what knowledge graphs are and their significance in data analytics.
- Application Use Cases: Real-world applications, especially in high-reliability sectors, demonstrating the value of contextualized insights.
- Future of AI-Driven Analytics: Speculations on how AI and graphs will shape data analysis in the future.
---
Key Takeaways
- The integration of knowledge graphs with AI technologies can significantly enhance data analysis by making insights more accessible and reliable.
- High-reliability organizations particularly benefit from systems that prioritize transparency and explainability.
- AI tools should augment human capabilities rather than replace them, creating more efficient workflows without sacrificing accuracy.
- The future of analytics will likely see an increase in personalized AI models as foundational models become commoditized.
---
Additional Resources
- [Data² Website](https://data2.ai/)
- [Connect with Jon Brewton on LinkedIn](https://www.linkedin.com/in/jon-brewton-datasquared/)
- [Connect with Daniel Bukowski on LinkedIn](https://www.linkedin.com/in/danieljbukowski/)
- [AI Opportunity Finder Tool](https://hatchworks.com/ai-opportunity-finder/)
---
Conclusion This episode of Talking AI explores how advanced data structures, particularly knowledge graphs, are revolutionizing the way high-reliability organizations analyze and utilize their data. By integrating AI with robust data models, organizations can achieve more accurate outcomes that support critical decision-making processes.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00We needed transparency, we needed explainability, we needed to understand the chain of cognition associated to how we operated with AI, if it was going to catch any sort of traction in the environments that we used to work in. Zero tolerance to failure, defense, intelligence, energy, finance, healthcare. Welcome to the Talking AI Podcast, where we talk AI with both experts in the field and early adopters. I'm your host, Matt Page, and we're here to demystify AI for you so you can get some value from it. Let's talk some AI.
0:32Hallucinations and trust issues continue to be a major roadblock for AI adoption. What if the results from AI could actually be explainable? And that's exactly what we're going to get into today with a couple of really smart guys from Data Squared. We're joined by John Burton, the founder and CEO, bringing deep expertise from energy and government sectors. And Daniel Bukowski, the CTO, who's impressive resume. You got like Facebook, Microsoft, Robinhood, Neo4j, which gets really deep into the graph side of things, but excited to have you guys on the podcast. Welcome. Thanks so much, man. I really appreciate it and look forward to the conversation today.
1:12Yeah, it's going to be fun. So, John, kick us off. Just give some context for who Datasquared is. Sure. Just to kind of set the stage and then leading into, Dan, we'll get into some of the, what is a graph database? How does it differ from relational databases, LLMs, all that kind of stuff? Sure. Well, look, we started Data Squared in sort of mid-2023. Leading up to September 2023, I was leading a company in Australia that was a bespoke analytics consultancy, really building solutions, machine learning, and AI solutions at scale for a lot of folks in the Australia and New Zealand region. I had left Chevron after several years in the oil and gas industry.
1:58And throughout that sort of time at Chevron, we spent a lot of time trying to optimize how we work and how efficient we could be as an organization with just the interactions that we had with the technology that we had on. and then trying to find new solutions that we can incorporate from a commercial perspective that would change our ability to sort of go further up the efficiency horizon and make sure that we can get as much blood out of the stone as we could. And I called one of my old buddies that I used to work with at Chevron, a guy named Jeff Dogley, and said, hey, man, it would be really interesting to see if we could figure out how to use this technology with the tools that we knew from the past and see if we can do anything novel and interesting with them.
2:39We found pretty quickly that it worked exceptionally well. And when we presented a graphical construct to a large language model, it not only understood it, it fundamentally could crawl the connective tissue between any one entity or any relationship within the structure of the data that we presented it. So we thought, well, we might be on to something. It brings a real commercial focus to what we're doing. We said, all right, let's start a company. Let's see what we can do with this new technology. But we started it within the context of high reliability organizations and high reliability operations, which meant we needed transparency.
3:16We needed explainability. We needed to understand the chain of cognition associated to how we operated with AI. just in general, if it was going to catch any sort of traction in the environments that we used to work in. Very specifically, high reliability industry can be defined as sort of a zero tolerance to failure, defense, intelligence, energy, finance, healthcare, all sort of sub segments of this environment. But the thing that really gave us an opportunity to figure this stuff out was an initial pilot test we did with some intelligence data. We got accepted into a a catalyst accelerator program with the U.S.
3:54government and the Air Force Research Labs. And we started looking at intelligence and all-source intelligence as a play within the structure of what we built. And at the end of the day, we built this modular, cloud-agnostic, zero-trust-secured platform. It's an AI solutions development platform that gives us the ability to step away from traditional black box applications and step into an environment where we can create traceable, transparent, and explainable outcomes. And that's really where sort of Dan's expertise starts to hit the environment because he comes from a background that understands the space really, really well.
4:33Yeah, and that's a great intro. I'm excited actually to get into some of your experience with some of those more regulated industries in a big, like that's a whole nother ballgame, I'm sure. But Dan, like let's... Quick break in the pod. If you're listening to this podcast, chances are you've been thinking about how to actually use AI inside your business. And that's exactly why we built the AI Opportunity Finder. It's a free tool that helps you uncover high impact, tailored AI use cases based on your business, your goals, your pain points, and your industry. No fluff, no generic use cases, just real ideas that fit your business and the ranked by ROI potential.
5:12It takes about three minutes to run and it's like having your own personal AI strategist for free. If you want to try it for free, check out the link in the show notes or go to hatchworks.com backslash AI dash opportunity dash finder. Awesome intro, really cool perspective, but I'm sure some people are like, okay, well, what, what makes y 'all unique in a sense? And what is this whole, you mentioned knowledge graph. What is that? And what are the components that are coming together? You mentioned with this like Gen AI chat GBT moment that's making this, this unique opportunity right now. No, definitely.
5:48And I've, I've been working with graphs for 20 years. I, I started my career as an accountant and an intelligence analyst 20 years ago in the federal government. So I was building graphs by hand to do follow the money investigations. So if you think about a graph, you have, uh, you have the entities, the nodes and the relationships, which is how they're connected. And a real easy way to think about it is the nodes are the nouns in your data. If you, John will bring up a graph here in a little bit and we'll, we'll take a look, but it's the people, the places, the things, the ideas, you know, Dan could be a node.
6:27John could be a node. Then you have the relationships. And so that's how that's how everything is connected. And that's really the secret and the essence of a knowledge graph is it tells you how the things in your data are related to each other. So if you have a spreadsheet of accounting information or financial transactions, which is what I used to work with, when you're looking at it in a Microsoft Excel and rows and columns, it might not look like that data is connected, but it actually is because you have money being sent from one account to another account. You just can't see it in the spreadsheet format.
7:02Similar to when I was working at, you mentioned Microsoft and Facebook doing IT security, security engineering, we would have millions, if not billions of rows of IT logs, which were really similar. You have one IP address connecting to another IP address, transmitting data or accessing a certain service. When you're looking at that as a spreadsheet, you really can't tell how things are connected, even though the data does contain connections. And so that's what a graph allows you to do. We built on the Neo4j Knowledge Graph. I used to work there. It's also the original graph database. And so we believe it's, you know, the industry leader at this point and why we're building on it.
7:44So when you bring that data together, what you get is, you know, at the database level, so written to the database, you have the things, the nodes in your data, as well as the relationships. of how they are connected. And this is really important when you're an intelligence analyst or any user of the database, because if you're dealing with, you know, relational, you might have three or four tables that you have to do a join on to bring things together to show how they're connected. That could take 30 seconds. That could take 30 hours, depending on the volume of the data. But when you're working with a knowledge graph, all of those connections already exist.
8:25They've been written to the database. And so what that practically means for an end user is instead of waiting 30 hours to do five different joins of the table, which as an end user, I have no idea what's going on under the hood. I probably don't even care. All I know is it's taking me three hours or 30 hours to get my data. But if I have that data in a graph, it's much easier to visualize how it's connected, but then also to get that information when I need it. And so that's why we're building on the foundation of a knowledge graph with, you know, some additional technologies that John mentioned and we can get into as well.
9:03And that's what it almost feels like it gives your data, like dimension, dimensionality, depth to it. In a sense, it's like you're going from this. It's like a super Mario brothers to, um, you know, name your, your video game with the, the 3d world and whatnot. Yeah. Yeah, I was going to say, there's this great old book that Jeff talks about all the time called Leaving Flatland, which was sort of a conceptual theoretical application of moving from sort of 2D systems and understanding of information to just a three dimensional representation of that information. I, you know, Dan's probably going to talk about this in a slightly different way, but, you know, as an engineer and looking at things, maybe he'll talk about it the exact same way I'm about to say it, but traditional systems and people see rows and columns.
9:52They see disconnected rows and columns. And what we're really gearing toward is building a vast, really contextualized three-dimensional landscape that we can present in that format to a system so that it inherently understands the connective tissue, not only between the pieces of this puzzle, but how it can transverse that information environment to get to all of the information components relative to a given query. I mean, this really is about taking things from 2D to 3D in concept. And so like, let's take it to the next step, because I'm sure there's some people that are going, that may be familiar with kind of a knowledge graph and this technology going, okay, well, is this like something I'm using instead of a large language model?
10:36Is this this new different thing? But it's not, they're, they're complimentary in nature. I like, I think this is where it starts to become differentiated, but like just a story from back in my day, I used to work with a team of data scientists and we'd get all these requests or if the business needed a request, they put something and they do a ticket. Uh, whoever's creating the thing, they'd write the SQL query, they'd return the results that would trigger another question. Uh, then you'd have to get back in the queue and it's just, it's like really painful back and forth. But I'm curious either Dan or John, whoever wants to take this, what happens when you put an LLM with this knowledge graph.
11:16No, and that's something that we've seen over the past few years. So when I joined Neo4j, my background is data science. My technical background is data science as well. And when I joined Neo, it was to focus on using algorithms, you know, help customers use page rank and degree centrality and these algorithms that were designed specifically towards graphs. Well, then about six months later, the, you know, chat GPT was released and, you know, kind of the world, the technology world was turned on its head with new capability. And what we were hypothesizing, and then has eventually been proven out by not just Neo, but also Microsoft, Google, BlackRock, all kinds of other, you know, research papers and kind of publications, is that the traditional way that people are currently working with language models right now for providing them with that additional information that they need, the retrieval augmented generation or RAG, is to take unstructured data, so PDFs, presentations, text, chunk it up, and then put it into a vector database, which is designed to handle that specific type of data.
12:25What we've seen quite a bit, I saw it when I was at Neo, we're seeing it here, is that that's a quick and easy way to get started providing the language models with that additional information they not have your your own your company's policies your company's data research or whatever but they really can't get beyond that 60 to 80 percent accuracy it's almost like that ceiling that you just can't break through and the challenge with that is it makes it unreliable to put that into production, especially high reliability, as John was mentioning, or customer facing production. You know, there are those use cases, what is it, a year, two years ago, where the state of New York had an issue with a lot of inaccuracy, some kind of embarrassing stuff with Air Canada.
13:12Airlines selling the car. Yeah. Yeah. A lot of good, you know, flights for free, you know, and that's, you know, I'm surprised we haven't seen, seen more of those crazy use cases to be out. I'm surprised there's not, I feel like with this, the AI agents and operator coming out, we've got to see more, but sorry to like break. Oh no. Oh no. I think honestly what it did was it gave everyone pause of like, I am not going to put this in production or customer facing production until I know sure it's accurate. Well, that's where graphs come into play. And what we've seen through the research and the experimentation and kind of really what What we're building around is graphs provide more explainable and more contextual information to the language model than just the, um, the vector store, like the chunks of data.
14:00And what that means is the traditional approach to rag where you're just, you know, creating a vector store, what it's doing, it's scattershotting the information to the LLM. Here's 10 pieces of information. use it to answer the question and have fun figuring out how it all goes together. It might be related. It might not be related. I'm just going to give you, I'm just going to throw some pages of a book at you and read it, figure it out, come up with an answer. Yeah. That kind of work to get you to that 60 or 80%, but it's, it's going to really struggle to get you to that, you know, upper 90 % reliability that you need to put it into production.
14:38what a graph does is it says you know it says here's the pieces of information but here's how the pieces of the puzzle fit together so think of it actually really i'm on the fly i can either be throwing a bunch of pieces of a puzzle at you and saying figure it out or i can put the pieces of the puzzle together and hand it to you on a platter and say here's the big picture that we're getting to read it and analyze it and use it to answer the question um and yes i did just come up with that on the fly. I'm picturing you like throwing a, a box of puzzle pieces. Yeah. Like ninja stars. Exactly. Yeah.
15:13I'm putting your ninja throwing at you. So with, uh, I'm going to take this opportunity to, to learn while I got some smart people who got their attention, but like when, when you get the vector database, the whole concept there is like, you talk about like nearest neighbor, um, in essence. So like, but the mathematical, I'm starting to, right. Yeah. Yeah. So what, and you're getting to this, um, Daniel, but like, where does it differ with the graph? And if it gets too technical, just tell me and I'll go, I'll go chat GBT after the high level. So it's interesting. Graphs are unique in that they don't really have a starting point in a sense.
15:52Looking at a spreadsheet, you have, you have cell A1, you kind of know where you're starting typically top to bottom, left to right with a graph, it's, you can really start anywhere. And so the key is, you know, the key with the, we'll call this the retriever, how I'm querying and retrieving the information is where do I want to start? Where am I going to anchor? And then how am I going to traverse the graph? So I can really start anywhere. And that's part of what we're building. And that's some of the technology that's built into Neo4j as well. Um, you, you mentioned the traditional vector search, uh, kind of vector, you know, nearest neighbors, approximate nearest neighbor similarity.
16:32That's a great place to start. Like that works. You can use that. We use that for some of the data in our platform. That's what you use. You know, we actually, we use vector stores as well because they, while they will only get you so far, they do have their purpose as part of the whole application. And so those typically work for unstructured information. So if you have text or PowerPoints or, you know, kind of that, the data that doesn't fit into rows and columns. But what we also do is translating the query into Cypher, which is the Neo4j query language, SQL, but built for graphs. And so that will help you actually find a better or maybe a different starting point than just simple similarity search.
17:15And so you might traverse to a different part of the graph and then query the local nodes and relationships or follow a specific path. And what we're doing is bringing all of that data together. You might have nearest neighbor search that brings you this piece of information. Then you might have a text to cipher query that's going to bring you other information from the graph that all have their place in generating the response. And that actually highlights another benefit of the graph in that graphs do really well combining structured and unstructured information versus a vector store alone or vector only rag, which is really, you can work with structured information, but it really does its best work with unstructured information.
17:58So text documents. That's right. So they complement each other in some ways. They're not necessarily alternatives. Okay. Yeah, it's not a new or it's, it's bring it together. And that's what we're building. We're building that application layer on top of it, that, you know, it does a lot of things, but one of them is bringing all of these pieces together in a solution on top of multiple database types you know yeah yeah you have to like from our perspective you have to understand when we're talking about target market approach we're looking at environments where 95 is actually not a good enough answer and so you have to understand the different things you can throw at this the fidelity of the information you need to store or have to store where traditionally resides as i said earlier Like everybody sees the world in rows and columns.
18:48But when you're using things like vectors, you're essentially coming up with a, I think these two things are related within this context. And what we're doing from a graphical perspective is saying, we know these two things are related and here's why. And so that's really the departure. So one thing works pretty well on getting a roundabout answer. The other one is imperative if you want clear, transparent and explainable answers. See, that's great. Y 'all aren't just engineers. You get some business acumen as well, because you're taking like this, okay, who need this very accurate approach when you're leveraging AI?
19:26Well, government, energy, the ones you mentioned. So it's a great kind of like niche to focus in on. And that's where like the value is. Like we hit on earlier, hallucinations are one of those key things that is stopping a lot of AI adoption. This helps you combat that. And it gets to this whole idea that I see so many people, they run into an issue with AI. They're like, oh, it's cool. It's just kind of a nice toy and they forget it. But it's, it's this evolution we're going to see of how solutions are architected, the applications that begin to emerge that are going to make it this, you know, just amazing kind of new world that we're, we're not ready for.
20:03I equate it back to like the beginning of the internet, you know, you couldn't be on the phone and use the internet at the same time and look where we are now. But talking about like the application, I'd be curious, like, are y 'all up for pulling it up? Just because I think what's so cool about the solution is actually visually seeing like how these nodes look and we don't have to go super deep. And for those that are listening only on the podcast, we will be very, very descriptive. I'll start as John's bringing that up. I'll just be our application is called review and there's, I guess I'll anchor it on.
20:36There's three things that we're building it to accomplish. um yeah as john brings this up the first is helping you get your data into that graph form because while graphs have been around for like several hundred years um and graph databases for i believe 15 going on almost 20 years it's still a little niche and so that process of getting your data whether it's rows and columns or unstructured data into the graph can be challenging and so we're accelerating that. The second is providing the user interface that John is showing here. So you can interact and get those graph-based insights without having to run Cypher queries or kind of, you mentioned that process, you submit a ticket, someone wrote the query, you get the data, you submit another ticket.
21:23We just want to avoid all of that. Let you interact with the graph and build upon what Neo has with some additional features and tools. And then the third part will be the, the natural language interface that John will show as well, where you can just query it and have it perform tasks and work with you through that natural language interface without having to do it yourself. And so yeah, John. And that's a big piece. And like, for those that are, that are listening, what we're looking at is like some stats in this like web of just like interconnected nodes and things that all of them connect.
21:56But, but John, give us like just a quick, like, yeah, overview, cuz this is really cool. Yeah. And so I'll table set this because it's a really interesting use case and one that kind of pointed us towards, oh, we might've built something that's hyperscalable. We had built a solution around an oil and gas problem, which I mentioned earlier, and we got it accepted in this catalyst accelerator program. We had to do something DOD and intelligence focused. We reached out to somebody in network and they gave us the capstone project that graduating analysts that go into the CIA, FBI, EHS do before they get into their full position within these organizations.
22:35Now, normally the team composition of six analysts takes six weeks. The average score, because this is really differentiated data, is around 70%. If you've ever seen that meme of Charlie from It's Always Sunny in Philadelphia on the whiteboard with all the yarn connecting everything. That's what these analysts are doing. That's a great show, by the way. Fantastic. This is completely off topic, but I've watched it a little bit back in the day, and my wife and I just complete show hole, no idea what to watch. So we started going back to Always Sunny in Philadelphia. It's just a really funny show.
23:09Yeah. It's an incredibly funny show. But that episode very specifically just is a perfect proxy for what these analysts were working through. They were trying to connect human intelligence, signals intelligence, financial intelligence. This is in reports. This is in transactional data. We're talking about unstructured and structured information across a wide geographic region in the Middle East to understand a couple of key things. One, who's responsible for the killing of our confidential informant? Who financed it? And ultimately, are they planning any other attacks? And is there any incremental evidence within the structure of this data environment that leads to a another risk being identified, one from a cyber perspective, which we unearthed.
23:52So we get this case study of information and we start looking at a total evidentiary environment of 1 ,629 documents and spreadsheets. It's a huge amount of information to crawl through. And so we have these human intelligence reports, these signal intelligence reports, the finite. And so we have all this transactional density. And what we do whenever we ingest this information is we create a graphical representation of it. We define the entities, we define the relationships within the structure of those entities. So what you're seeing from a graph model perspective on the screen right now is this hyper contextualized reality of this information environment.
24:30So it's not only the constituent parts of that environment, it's how they're connected and the strength and density of those connections. And whenever you can see this stuff, to Dan's earlier point, whenever we present it in this context to a large language model, the model's ability to say, okay, I understand fully what you're showing me, how it's connected, why it's connected together, reduces model drift and answer drift almost explicitly. Now we do some things that are proprietary to us to make sure that we can narrow that to almost zero, like traditional graph rag applications that gets you anywhere from like 80 to 92%.
25:09We're operating closer to that 98, 99 % space. But it's how we do some of the proprietary stuff we do to enrich these environments and contextualize these environments. So on the screen, you see this like, ah, here's this known entity. We can infer right away that this known entity is very, very important within the structure of this overall map because of how much transactional density you see coming to and from every other piece of this equation. And for the listening audience, this known entity has about a million arrows just like pointing right at us. But that's cool. That's the density component you're talking about, right?
25:49And this isn't even the actual data. What John's showing here is just the design of the model. That's right. You would think of it as the overall schema. And that's why it looks a little bit like a hairball because you design the schema of a graph to be functional. it doesn't always present well but it you can see conceptually how are things connected and then as the data gets loaded it fills it's in these different nodes or relationships in the in the graph yeah and then the next cool thing is you can actually ask it questions and chat with it not that we have to go into demo of that but that's the cool thing you don't have to necessarily write all these complex queries necessarily that's right that's the core thing that we built was to make this easy for the people that were using these tools.
26:34At the end of the day, we want to empower down in organizations to the people that are doing the work to be able to use these systems with ease. And the easiest way to do anything is just to ask a question in the language you speak of a tool that understands that language and can return very detailed, hyper-contextualized answers to you as a byproduct of just asking a question in normal language. But we do some things to set that up. And I think this is like a really important point. We work on sort of this T cubed framework, which gives us like full cycle chain of cognition explainability. And that is from a traceability perspective, we assign unique identifiers to every node, every edge and the relationships these systems have.
Read the full transcript
27:16And so we ID them, we semantically embed them, we enrich them, We understand temporal effects, when they start, when they stop, who was involved, where it happened, geographic constraints. We can interpret, you know, how, where, when. And as a byproduct of that, we can start to get to the whys and the hows. You know, we can start to define these things and query quite easily. And the other thing that we really try to do and be effective with is making sure that we can really adhere to general data provenance standards. So we source things, we cite things, we understand where they come from, why they come from these places, when we got it, who was the last person to update these records, what does that mean for how that information is used.
28:03And when we do all of these things, we can contextualize these environments the way that we do. And it allows us to start breaking down this information environment just holistically in slicing and dicing just different components of it. So we can start to say, hey, I want to look at all my known entities and all the banking activities going from those entities to different pieces of this overall information puzzle. We can start to look at evidence at scale and say, like, we want to slice this up on different evidence types. We want to use geographic boundary conditions. We want to set temporal designations on when we retrieve information.
28:43So all the stuff we do to prepare this environment allows us to contextualize this environment on query, which makes it easy for these systems to understand this information because we're not presenting it in rows and columns. We're presenting it in ways that is actually connective tissue sort of in a 3D format. And whenever we do that, it makes it very easy for these systems to understand what we're asking, the information we provide it, and how to answer those questions. So all the stuff that we do in the background, you know, labeling people, the connections between people, the connections between people and locations, the timelines associated to it, are to enable that sort of natural language interface with these things and to make it easy on the people that are executing the work within the systems that we're deploying.
29:26This is a bit more of a back-end view of the data model itself, the information environment, and just how these things are connected. But the user interface that we build, I'll pull up here quickly, is really intuitive as well. I mean, we're essentially building an information environment out to answer questions like this. This is a great question because if you know about the oil and gas industry, this question is insanely hard to answer correctly. This question was asked of us as a company. Can you help me understand how to optimize my produced water disposal activities, reduce my total expense associated to that, and then identify any profit increase opportunities I have within the structure of the network of information I provided you?
30:17Well, if you know anything about oil and gas, you need about 10 real different segments of the business to understand how to answer that question correctly. There's a lot of functional representation, whether it's from the Wells organization, the Wells organization in the midstream asset base within this company, how these things have performed in the past. What does that mean for what they're regulated to and from? What are they authorized to do? What are the contracts associated to each one of these things? What are our activities forecast moving forward? What does that mean for our produced water forecast?
30:57What does that mean for our historic putaway? Like, are we planning to something that we've never been able to accomplish before? What do our permits say about how we operate? You know, what are our costs of operating? To answer that question, you need all of those pieces of information, or you can't answer it correctly. And so what we really built is a system that allows us to answer really hard questions in really simple ways. Like asking that question in that way in the past is impossible to get a right answer on. And what we built is a system that cross-references all of this information, pulls it based on the way that we ask that question and creates really explainable outcomes.
31:38Does that make sense? Yeah, that is so freaking cool. This stuff people are creating now with this technology just amazes me every day. I got like three tangents I want to go down and I know we're running out of time, but I'm curious if we could hit them. All right. So first of all, the first thing that comes to mind is you just asked this question. It takes a whole team of people typically to do this, write this report. what what shall take on what happens to our place in the workforce the functions we are doing as humans versus what we can uh you know deploy to ai tools like this especially the more it gets integrated in with our data any any just thoughts more broadly on that yeah we both definitely have thoughts on this i'll give you sort of a general thing you know like our mission as a company is to discover a better way.
32:36Our vision for this company is that we can build a world where machines and people make smarter decisions together. Like we're not working on replacement theory here. We're trying to empower the people that are doing the work to do more work, more efficiently, more effectively every day. Now that's going to have some knock on effect as you start to proliferate this technology and these environments. But we're really trying to do it in a way that empowers people to do the jobs they already have more effectively. And so like, we're not trying to replace people. Our old professor at Harvard has this great quote that says AI is not gonna replace people, but people with AI are gonna replace people without it.
33:15And I think that's the truth, right? It's about how can you get smart in using these tools and building agile systems that allow you to sort of plug different information in at any point in time and sort of 10X your capacity as a doer within the structure of the organizations that you already operate in. And that's our goal at the end of the day. We're not trying to replace anybody. We're trying to make it easier for the people that are there to do the things that they need to do and a really high efficacy rate. Dan, I'll turn it over to you. I know you have your own thoughts on this stuff, but please.
33:46No, it's a really interesting question. And I have a similar perspective where, you know, technology is, it evolves what people do. If you go back to, I think a good example is kind of the horse and buggy versus the automobile. There was some change, but the automobile became this enormous industry employing hundreds of thousands, if not millions of people that wouldn't have existed if you didn't have that shift from the horse and buggy. And you think of like traditional Sherlock Holmes type Victorian era to what we have today with all the different automobiles and just all the industry that's built around them.
34:27What we're seeing here is with this type of technology, with AI, and everyone's talking about agents and all these different tools, it is shifting how you do work. Programming, you know, the actual software developer writing code is going to be a really interesting one to follow because if you're using a tool like a cursor or a GitHub Copilot, it is automating a lot of the steps of building software, writing software, debugging software. But, you know, the cursor team did a really good interview maybe three or four months ago where they said they want to make coding fun again. And so you're taking away a lot of the debugging and the searching stack overflow where instead I can just find the issue that's being debugged and have it fixed or with tab complete.
35:11Hey, I made one change on the code. It can populate that through 20 other instances in my code. And so I really do think for the, at least the near term, it's going to be a shift in work. And for many cases, it's going to make the work better and easier to do. That's not to say that we shouldn't be very aware and very supportive of, you know, roles that might get changed, but that's not what we're focusing on here. We're really trying to bring the capabilities of a graph and AI and reliable AI to end users, intelligence analysts, or other data analysts who are already doing this type of work. We're just making them more efficient and more capable in doing that specific work.
35:56Sorry, just one last follow-up there that I think is so important. We're working with the FBI now. We won't get into the details of what we're doing, but the change in how they query the system and the depth that they can get and the fidelity that they can get out of something that we're doing versus traditional systems like Palantir, Scale AI and some other companies is significant. It's non-trivial. The ability to understand the connective tissue at depth allows you to answer questions with more certainty and more predictability. And that's kind of the crux of the system that we built. It's like raising the level of interrogation with ease.
36:36It just gets you to a deeper level of understanding of the environment you're operating in, which should give you better results ultimately at the end of the day. Yeah. And I'm curious, and you were just sitting on cursor as an example there, Dan, I'm curious, like how, how y 'all are using AI in your business? We obviously see it in the product, but how are you using it in the business? I'm kind of, I'm on the side, like I'm not a software developer. I understand how it works, you know, um, databases front and back and all that kind of stuff. but I've now gone on this kind of like journey to building stuff.
37:09And it's like this hugely empowering, democratizing moment to where I can work with AI. And I've gotten to the stage where I'm having it work up front in the requirements phase. And then I'm getting into, okay, now let's plan this out. Where do we want to start? Because you don't want to zero shot it. You want to kind of do it iteratively. And then you're just building upon it from there. But curious, like any, like, you don't have to go through a laundry list, but is there like one or two use cases where you're using AI in the business specifically? Well, Dan's doing a lot more of it than I am.
37:41I can tell you the use case that I use it for most consistently is checking the things that I write because I'm a product of the Texas public school system. And I got to tell you, this thing is fantastic in helping me make sure I don't make mistakes in writing. So, you know, it's trivial stuff like that that is really, really useful. Dan will talk to you about sort of the technical merits of it and the things that you can use it for. But for me, it's about making my life a little bit easier for the things that I don't do great at. Yeah. And I think big picture, it's how AI should be used today as an assistant with a human as part of the process.
38:18And I think anything you do has that foundation. And so from a developer perspective, yeah, we have tools like Cursor or Copilot to, you know, help, especially with debugging and, you know, just finding issues or just accelerating the process of developing a software application. Our developers are using those tools. You know, I like using tools like Perplexity for research that is grounded, you know, with in facts. um and then yeah like i know you know i do a lot of writing on linkedin and kind of putting a lot of content out there i write all of my own content but i will use it for idea generation or image generation or even you know when we're working on a new um a new feature in the product or building out a new use excuse me a case using it as a brainstorming partner you know we're we're doing you know we're doing a lot of work right now with all source intelligence and cyber security another core graph use case that we're building a demo for is secure supply chain and logistics.
39:22And I'm not an expert in that area, but I can describe what I want to convey and how I'm going to use it. And it can help me brainstorm. What should the demo be? What should the narrative be? And what should we kind of build around to show the features of the graph as well as have an engaging story. So really anything, but as long as it is that, you know, as an assistant with a human in the process and, you know, making sure you have the reliable output and you can either double check the response or as we do with our platform, you know, have everything cited to a piece of data. And that's half the battle is knowing ways you can use the tool and there's no better way to learn than just starting to use it.
40:02And last question, I promise I'll let y 'all go. Like we've hit this crazy moment. And by the time this launches, who knows where the heck we'll be. But DeepSeek just got launched. Well, it's actually been out. The model's been out. But there was this huge reaction in the market where you got over a trillion dollars just eviscerated in a day, right? And it's their R1 model. It's, you know, the reports are they trained it in like under$6 million versus hundreds of millions you have with AI. Uh, inference cost is fractionally cheaper as well. Um, I'd be remiss if I didn't ask the guys with lots of government background and experience, just like, what's your, your thoughts on it?
40:46Either, you know, from a U S perspective, more broadly from the industry and AI, just curious what your thoughts are. I will have, I think, um, some different perspectives on this, just given our backgrounds. But like for me, from a business perspective, I've looked at large language models, at least foundational models, and how you interact with these models as a race to commoditization. Like we're getting to a point, it's a race to the bottom. Context windows are going to get bigger. Your ability to interact with these things is going to get wider and more prolific. And the cost associated is going to go down.
41:20Like that is just where we're headed right now. And I think Andrew Ng had a comment yesterday, which I thought was a pretty accurate representation of an understanding of what happened. And it said, today's DeepSeek sell-off in the stock market attributed to DeepSeek's version 3R1 disrupting the sort of norms within the tech ecosystem is another sign that the application layer is the right place to be. It's a great place to be because the foundation model layer is hyper competitive and it's full on racing towards commoditization. And that's great for people that are building applications. That's great for people that are building platforms.
42:03And so, you know, like that was always going to be a race to commoditization. It just so happens that something happened that influences it in a very, very tangible way quite quickly. The depth of differential between what was normal commoditized and commercialized models at this point versus what you could get, at least in terms of cost and performance, was such a departure from what we've seen from some of these companies that it really sent shockwaves through the industry. because, and I'll just say this in general, from a business perspective, we had Project Stargate announced a couple of days ago while I was in DC.
42:40And that's a$500 billion program of work to like build data centers around these large scale infrastructure compute facilities. And while I don't think that changes like super tangibly in a very short period of time, we were always going to get to the point where we were going to have personalized llms or small language models running at scale on pcs that people are going to use in a disconnected capacity this just might accelerate our path towards that that's my perspective dan any other thoughts it just just one point before i do want to get your take but i think the other piece is jevin's paradox at play kind of like economics here when stuff gets cheaper commoditized paradoxically increases consumption of the thing it does and it's not training isn't the only thing you where the energy and gpus are needed you need it on the inference side as well and i think once these ai use cases get um you can do more things more things become viable which means you're gonna need a lot of uh power on the inference side so like in my mind you know nvidia ceo is probably sitting there saying, no, I'm good.
43:52Oh, I'm with you. I got a call yesterday from the state of Louisiana saying, not saying who, but what does it mean for our future data center growth? And I said, nothing. As you commoditize exactly the point that you just made, it widens the user base, which means that the use stays roughly the same as you scale or grows more. And so we're still going to need data centers, but the personalization of this stuff is going to be pushed down quite quickly. And so it's going to be, I think you, you nailed that completely. Sorry, Dan, I don't mean to step all over you. Dan, you got the final word. Give us your, your last, uh, hot take before we wrap it up.
44:30I think you really hit on it and this really won't be a hot take, but you know, we are really two, three years in, I know that some of the early GPT technology has been around longer, but a year and change ago, we were talking about mixture of experts. And before that it was, you know, ChatGPT had just been released and we were dealing with 2 ,000 or 4 ,000 token context windows. And so I think in the early stages, this is just one more advancement in the technology. And it's, you know, it's going to be copied rapidly. You know, I'm sure Llama and Google and everyone else, they're going to have their own versions of this.
45:05They're going to be testing it out, improving it. And I think that actually is going to also emphasize the importance of open source in addition to the proprietary models. The fact that they released this, both the model, the weights, and the paper that goes with it will mean this approach, it looks very promising. It probably still needs to be proven out a little bit more. There's going to be areas where it gets improved. But a year from now, it's going to be, it could be the standard. It could not be the standard. We really don't know yet. But the big picture is, it's just one more step forward in the technology.
45:40and to your point, Jensen and NVIDIA, everyone's saying we're good because this is just going to make it more accessible and more usable more broadly. And we're still going to be going down this path. Nice. Yeah, you rounded it out perfectly, bringing in the open source take there as well. But John and Dan, thanks for jumping on and talking some AI. Where can folks find you? Where can they learn more about DataSquare? John, go first. Yeah, LinkedIn. our handle is at data2us and our website data2.ai those are the two places you can find us the most both Dan and I are pretty prolific on the LinkedIn front so you can look us up individually at John Bruton and at Daniel Bukowski that's I think Dan you probably have some other things you want to hit but that's the broader ones no that's it yeah the website LinkedIn and there's some great demos of the software on YouTube.
46:40And especially whether through a website or LinkedIn, if you want to reach out and learn more, just send us a message and we'd be happy to talk as well. Yeah, we'll drop some of those in the show notes. I got some cool demos out there on YouTube and whatnot. And your dog agrees in the background itself. Either the mail just showed up or it's dinner time or both. Yeah, nice. All right, guys, have a good one. Thanks for listening to the Talking AI Podcast. If you enjoyed the show, give us a follow or subscribe on your favorite podcast platform. And don't forget to leave us a review. We love those.
47:13For more info on Talking AI, visit TalkingAIPodcast.com. The single biggest mistake we see companies make with AI is they don't properly train their teams. We see it all the time. Companies roll out AI tools and expect people to just figure it out. But using AI effectively requires a totally different mindset and skill set. And that's exactly why we built training for every level of your org, from AI training for teams and executives to training engineering teams on our generative-driven development methodology. Or if you've already identified your AI use cases and want to just prioritize where to start, we offer an AI roadmap and ROI workshop to help you build a clear plan.
47:52It's all about going from we should use AI to actually driving real value with it. Head over to hatchworks.com to learn more.
From the publisher
How can high-reliability organizations transform their data from flat, 2D spreadsheets into a dynamic, 3D world of interconnected insights? In the Season 4 premiere of the Talking AI podcast, host Matt Paige dives deep with Jon Brewton, Founder & CEO at data², and Daniel Bukowski, CTO at data², to explore how leveraging knowledge graphs and AI is revolutionizing data transparency and explainability.
Jon and Daniel share their journey from traditional analytics to building a modular, cloud-agnostic, zero-trust AI platform that empowers industries such as defense, intelligence, energy, finance, and healthcare. They break down the core concepts of graph databases versus relational systems, discuss how graphs provide the connective tissue needed to support AI-driven decision making, and reveal how integrating LLMs with structured graph data can overcome challenges like hallucinations and unreliable outputs.
Learn how data² is transforming complex data into contextualized, three-dimensional insights that not only answer hard questions but also elevate the work of analysts and decision-makers. Whether you’re curious about how to improve data reliability or interested in the future of AI-augmented workflows, this episode is packed with practical insights and forward-thinking strategies.
Key Moments:
- The data² Journey: From Flat Data to Graphs
- Demystifying Knowledge Graphs & Their Value
- Graphs vs. Relational Databases: What’s the Difference?
- Integrating LLMs with Graph Data for Better Accuracy
- Transforming 2D Data into 3D Insights
- Real-World Use Cases in High-Reliability Industries
- Empowering Teams with Transparent, Explainable AI
- The Future of AI-Driven Analytics
Key Links:
Mentioned in this episode:
AI Opportunity Finder
Feeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you’ll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action. 👉 Try it now at https://hatchworks.com/ai-opportunity-finder/
