In short
The TWIML AI Podcast Episode #681: GraphRAG: Knowledge Graphs for AI Applications
Episode Overview In this episode of The TWIML AI Podcast, host Sam Charrington speaks with Kirk Marple, CEO and founder of Graphlit. They delve into the concept of "GraphRAG" (Graph Retrieval Augmented Generation), discussing its architecture, workflows, and applications in leveraging knowledge graphs with AI technologies like large language models (LLMs).
Key Discussion Points
Introduction to Graphlit and GraphRAG
- Graphlit Background: Founded three years ago, Graphlit focuses on creating an unstructured data platform capable of processing multimodal data including documents, audio, and video.
- Integration with RAG: Graphlit’s approach incorporates knowledge graph integration with Retrieval Augmented Generation (RAG) to enhance data retrieval and processing workflows.
The Role of Knowledge Graphs
- Entity Extraction: Discussion on the importance of extracting entities (people, places, things) to construct a knowledge graph that is essential for effective data retrieval.
- Comparison of Models: The use of various models (e.g., Azure AI text analytics, LLMs) for entity extraction, including insights on when to utilize traditional named entity recognition vs. LLMs.
Data Ingestion and Preparation Workflow
- Multi-Stage Workflow: Graphlit employs a systematic approach for ingesting data, which includes:
- Ingestion: Downloading data from various sources.
- Preparation: Transcribing audio and extracting text from documents.
- Extraction: Performing entity extraction during the ingestion pipeline.
- Enrichment: Optionally enriching the data with additional metadata from external sources (e.g., Wikipedia, Crunchbase).
Storage and Retrieval
- Hybrid Storage Model: Use of vector databases, graph databases, and document stores to create a comprehensive and searchable knowledge graph.
- Retrieval Strategies: Implementation of keyword, vector, and hybrid searches while emphasizing the importance of metadata filtering for improved data relevance.
Challenges in RAG and Evaluation
- Complexity of Retrieval: Addressing the challenges of merging different retrieval methods and the need for dynamic tuning in production systems.
- Evaluation Methods: The shift from subjective evaluation (vibe checks) to more automated evaluation methods to assess the relevance and quality of retrieval results.
Advancements in Prompting Techniques
- Dynamic Prompt Compilation: Development of a prompt compiler to enhance the interaction with LLMs, making it adaptable based on query context.
- Use of Structured Formats: Preference for XML-style structured prompts, which have shown to yield better responses from LLMs compared to other formats.
Future Directions and Use Cases
- Beyond Traditional Use Cases: Exploration of interesting use cases beyond chatbots, including content repurposing, audio summaries, and dynamic content generation (e.g., generating marketing materials).
- Agent-Based Applications: Discussion on future possibilities for incorporating agents in workflows to automate tasks like content editing and publishing.
Closing Thoughts
- Kirk concludes by expressing excitement for the future of GraphRAG and the possibilities for leveraging knowledge graphs in AI applications. The conversation emphasizes ongoing exploration and innovation in the field.
Key Takeaways
- GraphRAG Potential: GraphRAG presents a promising framework for integrating knowledge graphs with generative AI, facilitating advanced data retrieval and processing capabilities.
- Dynamic Workflows: The flexibility of Graphlit’s platform allows developers to configure workflows, making it accessible for various applications.
- Importance of Metadata: Robust metadata extraction and filtering are crucial for improving the effectiveness of retrieval systems in AI applications.
- Future Innovations: As advancements continue, the integration of RAG with knowledge graphs is expected to unlock new opportunities and enhance content generation strategies.
For more details, visit the complete show notes at [twimlai.com/go/681](https://twimlai.com/go/681).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington, and today I'm joined by Kirk Marple. Kirk is CEO and founder of Graphlet. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Kirk, welcome to the podcast. Yeah, thanks so much. I've been a longtime listener and glad to finally be part of this. I'm excited to have you on the show, and I'm really looking forward to our topic. we're going to be digging into what you're doing at Graphlet, but in particular, the broad space of GraphRAG.
0:42Tell us a little bit about Graphlet and kind of how you're approaching RAG as a space. Yeah, for sure. I mean, we've been around for about three years and started really building an unstructured data platform for everything. I mean, multimodal data, documents, audio, video, and really started getting interested in the knowledge graph side of this, of pulling all that data into a knowledge graph to make it explorable and then sort of saw the integration of that with RAG kind of come out over the last year and how we can benefit from that. We've known each other for a bit now. You attended our first TwiMLCon conference in San Francisco.
1:22That was in 2019. You started out doing, really trying to go after applying ml and ai to media talk a little bit about how that led you to the way you think about the problem today yeah i mean a lot of the problem in rag is is the r i mean the retrieval side but you have to have you have to have the data to retrieve in the first place and so we started focusing on i mean there wasn't at the time and still kind of isn't a five trend for unstructured data. I mean, that's kind of what we were looking at of where is the data pipelines to make the data available to AI models, ML models. And so we started at the ingest side, pulling in really all sorts of data.
2:12And in my background, I had a company in the broadcast video space. So dealt a lot with file-based workflows for that. And there's a lot of parallels of, I mean, pulling in data, running NLP on them, running computer vision. And so we kind of started with the content first of how do you get data into a system like this? And then search and retrieval, obviously, is a big step. And we start a lot with the metadata. I mean, how do you retrieve data? Well, how do you store the metadata first? And then how do you retrieve it? So like metadata filtering is common now in vector databases. So we were already doing a lot of that a couple of years ago.
2:49And as RAG kind of became the concept, we realized, I mean, we already have a great retrieval system to plug in to LLMs. And we already had solved the sort of media side of it that plugged into multimodal models very well. So that's when we kind of really focused on Graphlet as a platform that anybody could build an application on. And that's really what we've really been pushing for the last year. And so graph rag is a concept that grew out of a paper that Microsoft published some time ago. To what degree is your system like trying to implement the specific approaches from that paper as opposed to the general idea of applying – using graph – graphical relationships in a rag model?
3:47It was really interesting to see that paper. I mean, we'd been doing a lot of that already and just hadn't been talking about it as widely. But I mean, the first part of it is just how do you build the knowledge graph? And that's what we had focused on first of doing entity extraction, people, places, things and creating that from the content. And as we kind of built up our RAG system, pulling in that data from the knowledge graph. I mean, essentially, GraphRAG is really where we're at now. And that's what it was really. I mean, it was good to see it in a paper and kind of see other people kind of looking at that because it's something I've been thinking about for a couple of years.
4:24And really, honestly, a lot of this started with a podcast discovery platform that I was trying to build six or seven years ago. And it was a little too early. I mean, and now I know there's a there's a bunch of great project projects out there and transcription got cheaper. And it's just so much easier to build a platform like that. Let's talk a little bit about that entity extraction. Where did you start with that? How has it evolved? Is it a solve problem? If someone who's listening wanted to go about doing that, what are they going to find that's difficult? Let's dig deep into that part of the problem.
5:01What we found is there's a lot of overlap with the retrieval side of text extraction, text chunking, pulling those pieces together to do entity extraction. and really any NLP. So you kind of have to solve that part first and then use a model. I mean, we've used like Azure AI text analytics. I mean, we've actually used LLMs for this. And I mean, instructing the model to identify people, places and things. But what they're also really good at is identifying things like places, address extraction. We found LLMs are especially good at that, which was really difficult with NLP. And I mean, with with original algorithms.
5:43And so that I think that extraction side happens during the ingestion pipeline. So what we look at is. I've heard a little bit of the contrary, in particular, that large language models aren't great in a lot of cases for kind of traditional named entity extraction and that the traditional models are still better and maybe more controllable. Like, how do you think about where to apply what? We've actually seen that as well. I mean, we can use a mix. um we've seen the we've seen lms work great for specific cases like um for events like we were working with a community website where they were wanting to pull out like okay when when was the dance concert and when like where was it um we were able to guide the lm really well i mean to extract a an event um but i've seen it in other ways where kind of people and companies the classic sort of AI text analytics, I mean, like Azure, works way better.
6:49And so we can support both. So during the same ingested pipeline, we can actually run both models. And we can instruct it to say, okay, places use GPT-4, people and organizations use Azure, that kind of thing. You would think you can give an LLM a paragraph of text or a document and say, okay, you know, produce a JSON document that has all of the people identified and it just doesn't work as well as you might want it to work? There's cases. I mean, it definitely works in some situations, but what we found is, I mean, you get more noise in some cases. And so I think it's something you kind of want to try and look at.
7:29I mean, we've definitely seen a mix. I mean, there's some cases where, and I mean, honestly, the Azure type, I mean, or Amazon type models have similar noise problems where it'll identify a term as a company and it's not really a company. So there's a data quality issue that kind of feeds into both these sides. And I think that's where it does take a bit of testing and evaluation to get it working right. And have you been able to identify specific patterns that characterize the cases? One thing we've looked at, we haven't released yet, is sort of a chaining model using sort of an NLP style to identify the entities and then go in and refine that.
8:15We do support data enrichment today where when we identify an entity like a company, we can go call out to like Crunchbase or Wikipedia and enrich metadata for it. But one thing I really want to try is once you've identified the entity, go and re-identify and prompt the LLM to say, hey, I think this is Microsoft's in here. go, I don't know, grab everything you can find about Microsoft. That's an area that I want to explore definitely a bit more. So you've got these multiple methods to extract the entities, and you mentioned that that is part of the data ingestion. How is that used in the context of ingestion?
9:02So what we call our content workflow is kind of a multi-stage. So ingestion is really just like downloading it from a blob storage or a website. And then we have a stage which we call preparation. And that basically includes audio transcription or text extraction, PDF extraction. And what we get out of it is a canonical form of, I mean, essentially we store a JSON file of here's the transcript or here's the text with semantic chunking or page chunking. And then we have an extraction stage. And that's where this would happen. And so we kind of have this sort of state machine that the content goes through.
9:42And in our platform, you can configure each stage of the workflow. And yeah, at extraction, I mean, you can tell it which model or which API to use, what you want out of it. And then we basically take the results of that and feed that, connect everything up in the knowledge graph, basically. And then we have an enrichment phase after that, which is optional, where you can kind of just say, oh, I've identified, I mean, a company now go get its address, that kind of thing. And you recently joined our generative AI meetup. You've been involved in our community and kind of talked a little bit about this and did a demo.
10:22And one of the contexts or one of the questions that came up in that conversation was like RDF and these entity relationships, you know, from the traditional NLP world. Do you use those kinds of relationships in the graph that you construct or is it more ad hoc? That's a great question. I mean, the way we currently do it is we're basing all of our entities on schema.org. So JSON-LD kind of style. So we're not inventing our own data model for that. And I think that's important to have canonical kind of reusable data. And the LLMs actually know JSON-LD really well. So we found. And so that's kind of the first part.
11:09I mean, today we're not doing, we're leaning more on entity to content relationships rather than entity to entity relationships. That's most of what, I mean, we could do both, but I think that's mostly what we're trying to see. Because with the entity to content relationships with GraphRag, you can then say, oh, for this piece of content, what are all the entities related to it? And what are those entities related to in other pieces of content? um and so that our our graph leans leans a bit more on that on the entity to content relationships okay okay and so there's less of a need or desire at this point to kind of trend traverse entity to entity relationships to pull in this broader graph of content it's more kind of two-step as opposed to a fan out and that's what i think um like yohei has been working on some graph things you might seen on Twitter and there's been different kind of open source projects that are a bit more like, I mean, I work at X, that kind of relationship.
12:09And I mean, we could definitely do that. I mean, but our graph just because of what we're using it for is not leading that direction today. So you were talking through your workflow and you kind of got through the ingestion part of the pipeline. What's next? The preparation stage and the enrichment stage and all those, essentially end up with data in a vector database and in a graph and a document store and so what we do is have a sort of a hybrid data storage model where um and object storage we can leverage as well and so i mean if we're we're caching things essentially there and so we kind of use that all together where the graph is kind of replacing the relational index and that's the relationship between all the content um i mean we have collection kind of patterns we have the entity to content patterns.
13:02And so this is something I kind of came up with over the last several years and just kind of refined. And that's worked really well for us where we can then walk from one piece of content, either via similarity in a vector, sort of that vector angle or through the metadata and the graph. And so we can kind of have that hybrid approach and essentially a searchable knowledge graph. What's been your experience working with the vector databases Do you kind of support all of them or do you have preferences? Are there like key features that you're using in your system that are, you know, not universally supported or are you just using the basic capability of the vector databases?
13:48Yeah, that's a good point is, I mean, we had looked at, I mean, all the kind of major players and talked to a lot of folks. There's a lot of great ones out there. we were Azure native today. So we're leveraging a lot of kind of Azure data services like for our graph database. And we were actually already using Azure AI search. I mean, then cognitive search for our keyword search. And right at the time when I was evaluating vector databases, they came out with their vector index support. And I tried it and it worked. And it's a nice because it supports metadata filtering and keyword, vector and hybrid all in the same box.
14:30I mean, we're not plug and play with vector databases like some other things. We're kind of more, we're going to pick different things connected together and have a managed platform on top of it. And it works great. I mean, and they even just increased their performance and updated their storage and things like that. And we're actually going to be at Microsoft Build next month kind of showing this off. And I mean, in that Microsoft domain, but it has worked great for us. And we could, I mean, theoretically swap it out if we wanted to. But we've actually had some really good experiences so far with this one.
15:09And so we're kind of broadly seeing, you know, as RAG and vector databases become more popular, vector becoming kind of a layer on top of, you know, a lot of different data stores. So I just came back from Google Cloud Next and they did a similar thing where, you know, BigQuery now has a vector layer. AlloyDB, which is kind of their Postgres, now has a vector embeddings layer. My question is, do you see like the same thing, you know, happening with graph as well? Like, is it going to end up, do you see us moving towards a converged world where, you know, you put your data in and you have all these different kind of semantic abstractions on top of it, you know, where you're not replicating the data?
15:58Or do you think that, you know, graph is, you know, so distinct that they, you know, remain disparate systems? it's a really interesting point i mean it's something i've thought a bit about because when i was first looking for kind of the magical database that i really wanted like it didn't exist and i had to kind of create this this franken database from from a couple different things um the searchable side like keyword search in a graph i know that a couple vendors have have solved and in that world i would envision the vector comes next um i think there's it could I mean, you could integrate something like that with the graph.
16:40I mean, the big problem we had, honestly, with graph databases is the payload. You can't store a ton of data in a node. And that was a big reason that we kind of came up with this way that it's sort of a three-tier storage model where the graph is the index, the heavy metadata is in a JSON store, and then the really heavy content is in the file system, in object storage. And so we kind of used that sort of HSM model of, like, layered storage. And that's worked great. I mean, we only kind of pull it off desk if we really need to. And but to your point, I think I don't think it's a done deal that it'll go there.
17:17I mean, I think we'll have to see. But I wouldn't be surprised, really, if they start. I know they've started adding the keyword search, so I wouldn't be surprised on Vector. And along the same lines, like I think, you know, Vector databases have been around for, you know, quite a long time. And we were making progress in kind of shifting from keyword-based search to more semantic search and embedding-based search. And most of the Vector databases, vendors that we see now kind of started as these tools to support the search use case. You know, RAG came and kind of popularized that whole space and, you know, Vector then started become getting pulled into, you know, every data store.
18:06Do you think that GraphRag will have the same impact on graph? Graph databases have been around for a very long time, like longer than vector databases, I think, right? Neo4j, I've known those folks for quite a long time. um you know they've always had their place but they've never like i think had the the vector moment like vector databases recently have do you think that that you know based on what you've seen with graph rag like do you think graph uh graph rag is that killer app for graph databases i think it's definitely possible i mean graph databases i've i mean it has been one of those sort of background things where it's really useful for some specific use cases a lot of times the algorithmic side of like i mean running i mean algorithms heavy algorithms on a very large graph um i think the idea of more of a property graph and this sort of entity graph and graph rag i think could be a way that it's like what happened with vector and rag i mean there are several vendors that like made themselves like that that created huge companies out of just that that plug and play so i think we'll have to see i mean the the sort of value of graph rag is still a little TBD.
19:27I mean, we're still exploring it, but I think that's what makes it interesting. I think there's a lot of interesting paths you can take to see how to get the sort of squeeze of value. And I mean, I'm optimistic. I really think that, I mean, more people putting these relationships in the graph, just rather than just a typical relational database, gives them more opportunity to find the sort of explore their data in new ways. You go through this ingestion process that populates your kind of four data stores um you know interrelated data stores and is that the end of the the pre-processing uh step and then the um you know then we're to the next half which is what happens at the time of a query pretty much yeah i mean other than enrichment i mean which is which is optional i mean once the data is in the data stores in the vector index i mean that's like i mean if you call our api and say hey ingest you are like a url then the footprint becomes okay the extracted text and the and the data in the graph then it's all about retrieval and then it's you can make a query and that query um basically can be a mix it could be keyword only it could be vector hybrid um the metadata we'd already been doing a lot with metadata filtering because one of our key sort of thesis points was index everything in time and space because we were actually doing working in geospatial as well so a lot of our metadata is not just title author keyword it's i mean what lens on a camera are you using what's the gps location of the the the image or the video so like we have we have a really thorough set of metadata that we can extract from any media and that's when it becomes really interesting i think for i mean metadata a driven reg.
21:20I mean, asking questions about locations, asking questions about time ranges. And then the graph reg is really pulling out those entities and asking questions about, okay, well, where are these entities mentioned? And so that everything kind of drives towards that retrieval model. It strikes me that there is a lot of potential complexity in like the fusion of these different, you know, retrievals. You know, with, yeah, I think folks who listen to the podcast have heard me talk, you know, previously about like, it's easy to get a rag demo up and running, but to get something that's really, really good into production is difficult in part because there's a lot of tuning that goes into the retrieval.
22:07and that's just with kind of one you know data store with just your embeddings and and you know chunks and all that kind of tuning all that can be difficult now you're talking about layering in you know metadata and graph and you know concepts like re-ranking like what does re-ranking mean now across three different you know sources of of information? Like, how do you approach all that? I mean, we sort of take a layered approach. I mean, retrieval, the first is this kind of search that is, it definitely does touch. I mean, it touches the graph, touches the search index, the vector. And so the retrieval kind of step happens first, but it applies to kind of that triad of data stores.
22:52And then we actually just support added support for the new Cohere re-ranking model recently. And so what we do is the output of retrieval gives essentially a ranked list of sources. And but usually the sources has metadata that we pulled from the data stores. And then we provide that to the re-ranking model or do it ourselves. When you say a ranked list of sources, meaning source documents or sources like your three different metadata vector graph? content sources like document chunks or document sections, things like that, or chunks of an audio transcript and things like that. So a source is kind of an abstraction for like, hey, here's a piece of text that we found probably in some content.
23:37Got it. So you somehow execute, you know, a query across these multiple systems, you get back a bunch of chunks, and you treat them, You know, from a re-ranking perspective, you treat them equally like you're just doing the best you can to rank them based on the content of the chunks as opposed to, you know, the context in a graph or the, you know, the neighborhood in a embedding space or something like that. At least today we are. I mean, I think that's something we're exploring. I mean, one is time, sort of time relevance of sort of doing time clustering. I mean, we're looking at geo clustering.
24:21We're looking at some things like that. The graph can help with that as well. I mean, today, I mean, what's in production today is essentially just the content sources getting re-ranked and those content sources can come from retrieval. um i mean the the big thing we're we're doing today with um we we do support retrieval strategies for like expanding um the the chunk into a semantic chunk in which is like a section of a document or i mean chunks of a of a transcript and things like that so we we call those strategies that then that's where it's the knobs people could turn during during the retrieval step and then that's where we just added added the re-ranking strategy which initially supports cohere um and that's really helped i mean we can definitely see is the relevance you get out of your vector database or even the output of the search it's not i mean that's why these models exist i mean it's not exactly in the order you're expecting and lms can adapt to misordered data but it's the filtering that i've found is throwing away the irrelevant data um lets us kind of have a filter, kind of a low-pass filter on, I mean, okay, let's just get rid of stuff that doesn't make sense and that could confuse the LLM.
25:35Maybe this is a good time to introduce the topic of evaluation and how you think about relevance and quality of a retrieval. Yeah, I mean, a lot of it, I mean, everybody kind of goes through that ad hoc phase of, okay, it looks good. And originally, yeah, vibe check, exactly. I mean, and we started there. And we've actually just started working with a vendor in this space that has kind of more automated evals. And hopefully, I mean, we're looking maybe by next week, we'll have some details on that. But yeah, we're starting to kind of lean more into that. Because, I mean, for me, I'm not a data scientist.
26:15I came from a more traditional software background. A lot of this is like unit testing and integration testing. I mean, and you have to have suites of these tests to run when a new model comes out or, and I kind of look at it in that way. And so we are trying to automate that more, more than just VibeCheck. And I think there's some good vendors out there and good projects that really help with that. Are you thinking about dynamic optimization, DSPY and those kinds of things? You know, dig into that whole prompting space. Yeah, I mean, I would call it, I mean, what we do is more of a prompt compiler.
26:49and so it is dynamic prompting we're not just using i mean there are there are static phrases that we use or like i mean paragraphs that we we use for like some of the instructions and guidance um one of the things i came across and this is actually when i started implementing cohere months ago though how they like xml i mean the the cohere prompt really like xml and i was actually able to realize that um opening eye likes it as well and pretty much on most of the other models So we have sort of an XML template that we sort of compile to that has different sections. And it has a context section at the top and instructions and guidance.
27:28And so we've kind of broken it out into, okay, here's how I'm going to guide. Sort of guidance is like what not to do. And instructions is kind of what I want you to do. And then we compile the sources and then the user prompt piece at the bottom. um we've had to do some things like dynamic based on model where like i don't know like haiku um forgot the schema the json schema like it you had to remind it of it at the top and the bottom i think it was and i've seen that with a couple different models and um and those are really some things that it that it lets us be dynamic on the fly basic so there's a there's a structure but it is kind of a compiler that we dynamically output based on the query, the incoming query and all the retrieval strategies and all that kind of thing.
28:19Is it an LLM that's compiling this into a prompt? Is it, you know, a set of rules or heuristics or a combination like? Yeah, the compiler is just code. I mean, honestly, it's just, it's straightforward. We are using LLMs for things like, what is it, a prompt rewriting. So we do have a way that we have a couple strategies for optimizing for semantic search versus kind of rewriting the prompt. We actually just have an experimental one for multi-query now that'll kind of break it into multi-queries, which actually works really nicely. And I know that, I mean, Lomindex and Limechain, they've seen those kind of results as well.
28:59Those work really nicely. We use LLMs for like summarization of the conversation history. so we support like a windowed conversation as well as a summarized history so those kind of things but the the compiler itself is kind of it's just it's a code i mean it's just kind of looking at all the context that it has and then generating that essentially a strength it just gets gets put into the llm and when you you said that the llm's like xml
29:31uh you know how different how much stronger is that statement than they like structure like they like that particular kind of structure apparently like as opposed to you know uh headings and paragraphs yaml json that for whatever reason xml is just like some magic pixie dust that makes them work better i mean i think it's it's a bit of both but i think i learned it from i mean it's It's just how Coher documents, they kind of say, hey, here's our preferred prompting format with and using XML to kind of create sections within that. And I found that when I backported that, because I tried YAML, I tried JSON, I tried a couple other things.
30:14And I mean, I don't know if I was a little surprised, but I mean, it's probably trained on a lot of data like that. But I was I saw how it we didn't have to change our method depending on the model as much. It was a common thread that I found that pretty much every model I've tried listens to the prompt very well when it's structured like that. Now, could I do that as YAML instead of XML? I think OpenAI is definitely more resilient, and you can give it more options for the structure I found. But Cohere would then have a downside, at least with their original models, where it wouldn't listen to the structure.
30:53So what we were looking for was something that kind of worked across everything and could just be kind of a standard approach. For a given use case, are you using multiple models and kind of orchestrating them and trying to identify what's best slash cheapest for a given step? Are you defaulting to a single model that you like the best? How do you think about the model space? Yeah, I mean, we let developers pick. We have what's called a specification. It's kind of a preset that you can pick your model, pick your tokens, like token limit, system prompt, all that kind of stuff. And then depending on what conversation they're having or what, like if they can use a different specification for JSON extraction per se than for having a conversation.
Read the full transcript
31:40So we kind of hand that to the developers that use our platform to choose. But we, I mean, we're currently using actually GPT 3.5, 16K as our default. Like if nobody, if they don't pick anything else. But I am actually evaluating moving over to Haku as our default. I think we'll, I mean, we'll probably do that before the end of the month. And just because I've really liked that model just from a price performance and just quality standpoint. So I'm thinking that that's probably going to become our default. You've mentioned on a couple of occasions, you know, what has given me the picture of like building blocks that a developer at least, you know, might conceptually think about their application?
32:26Are you presenting them as building blocks? Is there a library of building blocks? Or are these just kind of the natural steps that someone goes through? Or even like, how does the product present itself? Is it, you know, client libraries, APIs, you know, GUI thingy? it is it's an api first platform we have a developer portal you can sign up for free get an api key start using it same immediately and we have a graphql api that is kind of the native api to it but we actually just released this week native sdks for python and typescript for like a node backend and so they kind of hide the complexity of graphql and it just it's it's a simpler um and it's a type uh type safe experience which we wanted um and so yeah we just i just updated we have a bunch of streamlit apps that use the sdks now and it's been great to see the developer experience that um and i on our home page now it's like two lines of code for ingest a website and prompt a conversation with a prompt and so we've we've abstracted it down to that it is an interesting point a lot of what we just spent the last 40 minutes talking about the developer doesn't have to think about.
33:42This is stuff you're doing under the covers to enable the developer to pass you a URL and then be able to prompt against it. Yeah. And I think this is a really big difference in how we approach this versus a lot of the open source projects that have been around and have really led the way for rag awareness. But we take a more content first approach where it's almost like a CMS, like you're just putting data into this content management system. And then we're saying anything you want to change is configuration and so we have this workflow object that you can create that says here's how i want you to adjust data here's how i want you to prepare it enrich it or extract it and enrich it and all you do is when you say put me at a website and use this workflow and what we're actually pulling from is more the configuration is code model like github actions and it's really like hey okay here's your workflow it's static and as i check in code it's just going to use that and build it and that's kind of the approach we took with those workflows predefined or are they defined in code by the developer the latter yeah so the developer can and we have built-in ones so like if you do nothing else we'll transcribe your audio we'll i mean put it into the vector database like you don't have to do anything um but you a lot of it is just like opt-in do you want to change your transcription model or do you want to use gpt4 instead of gpt 3.5 for summarization so the kind of hooks or tools or things that you can use to kind of insert into the process and you can just configure all that and they're reusable and so you really would probably have a developer would have a set of them for their application maybe for data extraction or for conversations and they're just reusable across the platform and then as they ingest data they'd say okay and use this workflow um and so that was really where what i was thinking about this last year is really the i mean the idea of not it's really not building blocks per se like we have the building blocks already connected but you could turn knobs on each of the building blocks and and that's an opinionated approach i mean we were a bit different in that way but i think it for us we feel it makes for a really nice developer experience where it's really simple to get started and then you could tune and you're not having to put, you're not having to think about cloud infra.
36:02You're not having to think about the LMs. I mean, it just works. And we jumped, you know, right in and started talking about RAG. And I think, uh, for a lot of people, the kind of the, the use case concept when you're talking about RAG is like, you know, there's a chat bot, I, you know, issue a prompt or a question around, uh, uh, you know, a document or some sort of content or something. And, uh, um, you know, I get a response that's based on, you know, some knowledge that an LLM has beyond the kind of pre-trained, you know, training corpus. Right. Um, but you feel strongly that there are, you know, use cases that are interesting beyond chatbots you know talk a little bit about the ones that you're excited about and you know how you've seen them play out for sure i mean i think i think the chatbot or copilot experience is kind of like the first order it's it's direct you're seeing a direct rag put on a page which is great i mean i think that's there's a lot of value for that um but what we're seeing is rag as a pattern almost I mean, it's like SQL for unstructured data, I mean, or like a way to format data in a way that you can use it for so many things like content repurposing, which we have something called publishing that you can point at the data.
37:30You can filter, create your subset of data, summarize it, and then publish it in a new format, like a blog post or a long form article or social media. and what we really see is that content publishing angle being that kind of second order of of really how can you use rag as just a function a sense of a piece of functionality to deliver more value um because it's using rag in multiple ways i mean it's it's finding all the data it's i mean using lms doing different things and then creating like audio summaries um like i have one set up that I have a feed on my email and every morning it actually reads through the email, summarizes it, creates an 11 labs audio summary, I post it to Slack.
38:16And you can do that with like, I don't know, four API calls or three API calls with us. And that's the kind of stuff that I really want to see developers, what they can build with those kinds of build. I mean, those are sort of building blocks. Um, and, and that's what excites me is like be able to repurpose that content into new ways. And we're actually even looking at dynamic image generation, using LLMs to create the text to actually put into an image template and generate marketing copy or marketing graphics and things like that. If you think about the content generation use cases that have played a big role in driving LLM popularity and the image generation popularity, it's like, it's static in the sense that you try to stuff, you know, a bunch of stuff into a prompt or use a prompt to do a single prompt to direct the creation of some content assets.
39:20So, you know, if I want to write something, I might give it a bunch of resources and stick it into a prompt. And what you're essentially describing is that, you know, just like we might use RAG to create context dynamically for an interactive system, we can also use RAG to create content or context dynamically for some of these generation use cases. So you've got now a prompt that dynamically drives the generation of a context that results in a document, an audio, a video, a blog post, an image, whatever. Yeah. Yeah. And looking at like, I mean, integrating with like, hey, Jen, for kind of video clips with avatars, like we're, I mean, it's, it's, that's the kind of stuff that I think becomes interesting.
40:18And the one thing, when you publish that, it becomes a piece of content in the system. Like we re-ingest it into the system. And so it becomes searchable automatically. And eventually, I mean, we've been waiting for feedback of like, we could auto publish, like put it in your Google Docs for you, or put it onto SharePoint or whatever people want. I mean, that part's not hard, but I could see that being the connected tissue of dynamic content generation. And then I'll just bring up agents for a quick sec. But once you have that piece of content, then you could run an editor agent workflow of like, oh, hey, go have your editor.
40:54I've got to publish this. Use this LL prompt. Edit it. Republish. Write that back into the system. and you could have all these kind of editor relationships or go generate me a marketing graphic based on this piece of content and then paste it back in. So that's why when I say we're kind of like a CMS with LLMs built in, I think that's the long-term value is it really becomes this ecosystem for content generation that may be interactive or may be offline and it can work both ways. And I'm taking that agent discussion as kind of a future direction thing, something that's possible based on the foundation and not something that you have gotten very far with today?
41:36I mean, we have one thing we call alerts. And so that's kind of like a very simplistic one-step agent that we work with. And it's like basically the Slack alert of my email. And so it's sort of a two-step, like summarize, publish, and then notify. And we're actually looking and kind of seeing where, I mean, there's a crew AI and Autogen and all these different things that are great. And do we integrate with them? Do we try and do something our own? I mean, I think there's so much innovation there. The one thing I will say that I really like is I've worked a lot with like actor models in the past, kind of dynamic kind of workers that are interacting.
42:19and i've started to hear this now on some more podcasts of there's a lot of learning that people can take from kind of the distributed systems actor model world and apply to agents and i think the the ones that i like the best are kind of learning from those it's kind of a solved problem in a lot of ways you know how to spin up a kind of like durable entities or that kind of thing um so i just hope that i mean we don't try to reinvent the wheel um on that and i mean there's That part of it should hopefully be the easier part, but it's really the memory for them and the workflows and things that'll be the more innovative part.
42:56I mean, there's, you know, it's making me think about message queues and message passing, all these, you know, this infrastructure stuff that we've already figured out for, you know, distributed systems. We don't necessarily have to reinvent all that stuff. And I don't know that a lot of the stuff that I've seen reinvented in a good way. Like it's all single process. Like let's not start there. Well, that's, I mean, today we're, we're an event-driven system. I mean, we're built in that way from scratch. Alerts and feeds are actually kind of agents. Like they're actually like a, a demon process that's running, sitting there doing things on a periodic and running different events.
43:36We actually use this thing called durable entities from, from Azure, from, from Microsoft. That is sort of this Azure model that, or sorry, the actor model that has state. and it's sort of like AWS step functions in a way where you can run different steps and define a workflow. And to us, I mean, that's kind of our agent model. I mean, we can construct those dynamically and let Azure run it. I mean, they'll handle all the retries and provisioning and it just works. So if we come out with something that's agent-like, it's probably just going to be a, hey, it's a layer on top of that. I mean, because I think it's a nice pattern already.
44:15Is it open source? Is it as a service? Is it on-prem, off-prem? Yes, we're cloud platforms and service. So it's closed source today. The SDKs are open. We just open sourced the SDKs on top of it. All the sample apps are open source. But we take advantage of a good number of Azure backend services. So we don't do on-prem today. We're going to release in the Azure Marketplace this quarter that you can essentially have your own sandbox ecosystem of Graphlet running in your own Azure subscription. And we do have ingest from all the clouds, but currently today it's an Azure native service. Is multi-cloud a priority or is it kind of a wait and see?
45:00It's a wait and see. I mean, I think it's, I mean, we could have gone the kind of Kubernetes really abstract multi-cloud model. And there are reasons that I kind of thought, okay, if we lean in on one cloud, that we can get a lot more benefit. I mean, the managed databases, especially, that we're leveraging solve a lot of problems for us. I mean, and I, I mean, seeing other companies that struggle with their Kubernetes infrastructure and struggle with just all that, that side of it, and we kind of build up, we've decided to build on the shoulders of a couple of those things. And so there's still, I mean, there's a lot we do from the ingest pipeline standpoint but i mean i'm not worrying about managing a database i mean right now and so um i don't know i mean we're we're still i mean somewhat small company and so i think it's just where do you want to put your eggs in what basket at this point it's been great to see the progress you've made and how it's kind of evolved over uh past two three years uh since you first spoke with me about it and looking forward to seeing how it continues to evolve.
46:08Thanks so much. No, it's, I mean, there's been so many interesting topics about RAG. I think GraphRAG is really, we're still in the early days and that's what's exciting about it. I think we, we hope we kind of have the underpinnings of it, but I think we're still going to learn so much more. I mean, over the rest of the year. Well, thanks so much, Kirk, for jumping on and sharing a bit about what you're working on. Yeah. Happy to be here. Thanks so much. Thank you.
From the publisher
Today we're joined by Kirk Marple, CEO and founder of Graphlit, to explore the emerging paradigm of "GraphRAG," or Graph Retrieval Augmented Generation. In our conversation, Kirk digs into the GraphRAG architecture and how Graphlit uses it to offer a multi-stage workflow for ingesting, processing, retrieving, and generating content using LLMs (like GPT-4) and other Generative AI tech. He shares how the system performs entity extraction to build a knowledge graph and how graph, vector, and object storage are integrated in the system. We dive into how the system uses “prompt compilation” to improve the results it gets from Large Language Models during generation. We conclude by discussing several use cases the approach supports, as well as future agent-based applications it enables.
The complete show notes for this episode can be found at twimlai.com/go/681.




