In short
Podcast Summary: Knowledge Graphs as Agentic Memory with Daniel Chalef
Overview Podcast Title: Software Engineering Daily Episode Title: Knowledge Graphs as Agentic Memory with Daniel Chalef Episode Description: This episode discusses the challenges of contextual memory in AI, focusing on how current models retain and recall information. The conversation includes insights into ZEP, a startup creating a memory layer for AI agents using temporal knowledge graphs.
Key Participants
- Daniel Chalef: Founder of ZEP, with a background in software engineering and startups.
- Kevin Ball (KBall): Vice President of Engineering at Mento, experienced in AI discussions and engineering coaching.
Key Concepts Contextual Memory in AI
- Current AI models face significant challenges in retaining and recalling relevant information over time.
- Unlike humans, AI systems often rely on fixed context windows, leading to loss of important past interactions.
ZEP and Temporal Knowledge Graphs
- ZEP is a startup working on a memory layer for AI agents using temporal knowledge graphs to retain long-term contextual information.
- Founded in 2023 and part of Y Combinator Batch Winter 2024, ZEP targets mid-sized enterprises.
Definition of AI Agents
- AI Agent: Defined by three attributes:
- Uses a large language model (LLM) as its brain.
- Autonomously interprets instructions and makes decisions.
- Takes actions to achieve predefined goals.
- Agents require tools for action (e.g., querying the web, interacting with applications) and a broad understanding of their environment.
Importance of Memory in AI
- Short-term Memory: Relates to current interactions (e.g., ongoing conversations).
- Long-term Memory: Involves storing experiences and semantic relationships over time, akin to human memory.
- Challenges include creating a semantic memory that draws connections between various experiences.
Knowledge Graphs
- Knowledge Graph: A data structure that semantically models complex relationships, typically represented through triples (two entities and a relationship).
- Knowledge graphs are advantageous for building dense, well-described semantic datasets, allowing for efficient data retrieval.
Ontology Development with LLMs
- Ontologies define the categories and relationships within knowledge graphs.
- ZEP's Graffiti Library can build ontologies dynamically based on real-time data inputs, improving flexibility and adaptability.
Graffiti
A Temporal Knowledge Graph
- Graffiti supports dynamic data and allows agents to manage changing memories effectively.
- Implements a bitemporal model, capturing both event timelines and validity periods, enabling agents to reason with evolving data.
Query Mechanism
- Developers can send both structured (JSON) and unstructured data to Graffiti.
- It features a simple API for graph creation and querying, avoiding complex query languages like Cypher.
- Uses semantic and full-text indexing for efficient data retrieval, allowing near-instantaneous access to relevant information.
Challenges and Future Directions
- ZEP aims to navigate the complexities of transitioning from experimental prototypes to deployment in real-world applications.
- The focus is on enhancing privacy, security, and compliance aspects while developing innovative agentic applications in various domains, including healthcare and B2B solutions.
Key Takeaways
- The shift from static RAG (retrieval-augmented generation) systems to dynamic agentic memory frameworks represents a significant evolution in AI capabilities.
- Understanding and managing semantic relationships and contextual memory is crucial for developing effective AI agents.
- Future advancements in AI will require integrating probabilistic reasoning and exploring how diverse human perspectives influence knowledge representation.
Closing Thoughts
- The discussion highlights the importance of memory in AI and the potential of knowledge graphs as a solution to contextual challenges in AI systems.
- Daniel Chalef emphasizes the ongoing journey of transforming AI capabilities into practical applications and the philosophical implications of how agents perceive reality.
For further details, check the full episode of Software Engineering Daily [here](https://softwareengineeringdaily.com/2025/03/25/knowledge-graphs-as-agentic-memory-with-daniel-chalef/).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Contextual memory in AI is a major challenge because current models struggle to retain and recall relevant information over time. While humans can build long-term semantic relationships, AI systems often rely on fixed context windows, leading to loss of important past interactions. ZEP is a startup that's developing a memory layer for AI agents using temporal knowledge graphs, enabling agents to retain long-term contextual information. It was founded in 2023 and was part of the Y Combinator Batch of Winter 2024. Daniel Shaleff is the founder of ZEP. He joins the show with Kevin Ball to talk about the challenge of contextual memory in AI, temporal knowledge graphs, ambient AI agents, and more.
0:44Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space. Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.
1:20Daniel, welcome to the show. Thanks for having me, Kevin. Yeah, excited to get into this. So let's maybe start. Do you want to introduce yourself and Zep and what you all are all about? Yeah, so I'm a software engineer turned founder, and this is my second startup and my first AI startup. Like many, we're just under two years old. And our focus is on enabling the agentic future. And the way we're going to do that, or the only way we can do that is to ensure that agents have the right information available to them at the right time. And so ZEP is a memory layer for agentic applications. And we focus on midsize companies in the enterprise.
2:02And it's a very exciting space to be in. So let's maybe define some terms here, because while some of our audience is familiar with all these different things, maybe not everyone is. So when you say the agentic future, what do you mean by that? I feel like agent is a term that different folks mean different things by that. So how do you define agent? Yeah, it is so nebulous, right? There is a definition for agents. It's just become pretty loose. And the way I think about it is that AI agents today have an LLM as their brain, but they also have the ability to autonomously interpret instructions and make decisions, which is a very important aspect of agents, and then take actions to achieve goals that have been set for them.
2:51So those are the three high-level attributes that I look at when I think about what agents are. When you dig a little bit deeper into it, though, there are several high-level components or several needed components for an agent to be able to actually do what I just described. And they have tools, for example, for taking action. And tools might be something like querying the web, so searching the web. For answers, a tool might be actually taking action in a line of business application. So generating an invoice. And the other important aspect is for an agent to be able to reason and understand what to do next and make a decision.
3:38They need to have broad understanding of the environment that they exist in. So what is the user's world if there's a human in the loop? or what is the business world if they're purely reacting to changes in business state or our personal state, like our home. And they're going to have to have a very broad understanding of that. And so memory is important. So being able to recall what they did in the past and being able to then plot a way forward based on this new, maybe new stimulus that they've received. And memory really enables planning. And that's the other big, important aspect. And there's something called a perception action loop when we talk about agents.
4:27And so agents look at stimulus, so maybe a human's request or some event that is coming from their environment. They look at their memory, they perceive this, and then they decide to take action based on that using the tools that they have. So that's just a very high level view of the way I kind of look at what agents are. That's really helpful. So let's maybe dig a little bit deeper on what those pieces are. Tools, I think, is not going to be the big focus of what we're talking about here, but maybe first defining a little bit, that's essentially function calls, right? Making action using a computer somewhere.
5:04Exactly, yeah. Okay. And then reasoning, LLMs, like we're all kind of familiar with what these things enable. Memory, that's where you are focused. And I feel like it might be worth us kind of defining that a little bit more carefully. Is this just context that you're dumping into your LLM prompt in some way? Like, what do you mean by memory? Yeah. So I like to think of memory as quite expansive. and there are multiple types of memory that we look at when we think about agents. There's a short-term memory. A short-term memory is merely what is happening in the current, if you're talking about a human-in-the-loop agent, what is happening in the current conversation?
5:49What has the user just asked? What did they say previously? So what do I know? You can think of it as a human's short-term memory. But we also need long-term memory. And if you think about how the human brain works, long-term memories are often processed in some way, and they're put into a big data bank or memory bank for recall. And there are various different types of long-term memory as well. We could remember how to do things, so there's procedural memory, but then there's also semantic memory, where we build relationships between different events in our lives or things that we've perceived and connect them all up.
6:35And those go into the databank as well. And so those are various types of long-term memory. There are more, but those are kind of the big high-level aspects of long-term memory. And if we double click on semantic memory, this is where things get really challenging because humans have an uncanny ability to draw connections between various parts of our lives and our experiences. Sometimes we get it wrong. I forget people's names all the time. I forget what I ate for breakfast. But most of the time, we're able to file it away and recall it over very long periods, which is a marvel. But agents are going to have to have something that approximates that to be able to work effectively.
7:25Not only that, but the promise of AI is the ability to process vast amounts of data and make sense of it far more than humans are able to conceive or understand in a very short period of time. Okay. So that makes sense. And I think I understand why you double clicked on that particular area, because that ties into this concept of a knowledge graph, which is, as I understand it, what you all are building. So let's maybe talk about knowledge graphs. What are they? Is that this is the implementation of those semantic links? Yeah. So knowledge graphs are data structures that allow you to semantically model complex relationships.
8:09And they typically contain something called a triple, where you have two entities or what are called nodes in the graph. And those are things. So you could say a person is an entity or even a concept could be an entity. So you have two of those. And then between them, you have a relationship. And this is called an edge in the graph. And the relationship describes what an entity-to-entity relationship is about. So that's what the edge does. So that's a knowledge graph in a nutshell. And why they're very useful from a memory perspective is that we can build very dense and well-described data sets, semantic data sets, that are a very good fit for retrieval.
9:07So graphs allow all sorts of interesting approaches to retrieving data. There's actually a very mathematical approach that one can take to retrieving data from a graph and traversing the relationships that are available. And there are all sorts of wonderful algorithms that have been developed that allow you to interrogate the data and also do so in a very intuitive way. So I absolutely love knowledge graphs. I think they're an amazing data structure for condensing information. The challenge with building knowledge graphs has always been defining the ontology or the types of things that you care about and how those relationships between the things can take effect.
9:58So what sort of relationships are allowed in a knowledge graph? Defining ontologies has always been a very challenging endeavor. And then also building the knowledge graph has previously had to be done manually or in some sort of not very flexible way. But with LLMs, we have an amazing new opportunity to interpret data at scale and extract entities from unstructured text or even structured data and understand their relationships in ways we've not been able to do before. And again, at scale, because previously it was hand-building knowledge graphs. So let's maybe dive into that. What does it look like to, let's talk about ontology, right?
10:45Right. So defining your categories of things. When you say with an LLM, you can just sort of automate that away. Is the LLM defining your categories? How are you thinking about that piece of this? Yeah. So let me speak to Zep's open source graffiti library. What's unique about Zep's graffiti library is that it has a temporal dimension, but let's not dive into that yet. What I will speak to, however, is how graffiti is able to build an ontology on the fly for data, which is very, very useful when you're working in a domain that isn't necessarily as contained or bound as we've seen in the past. So if you are building agents that interact with humans, humans say the darndest things.
11:42Yes. So do LLMs for that matter. You get all sorts of bizarre things coming out of them. Well, yes. Yeah. And so if you are, for example, trying to understand human preferences as an agent that maybe controls a home, building a well-defined ontology can be extremely challenging. And that's one of the things that has encumbered developers in the past. It's very challenging to fit your data into an ontology. And so graffiti builds this on the fly by doing named entity recognition across the data that it is seeing. So it is determining what the things are and then intelligently understanding those relationships.
12:28Now, that could become very, very problematic at scale if you're not being clever about deduplicating the things. So these two things are alike. They might be differently named, but conceptually they're the same things. So we should use the same type. Secondly, these relationships are very similar. And so we should use the same labels for the relationships rather than generating an entirely new label. Because otherwise your graph becomes very difficult to query into if there are many different types of entities that are very much related. So let's maybe talk about type, because I think one of the things I saw on Zep's website is you talked about pulling out strongly typed data from raw chat information.
13:22how do you think about that type extraction? And are those types themselves then morphable? Because I might have a conversation where I name three things about this entity, and then sometime later, the agent is having another conversation and discovers, oh, there's a whole slew of additional data that's attached to this type of entity. Yeah. So I spoke a little bit about Graffiti's organic development of an ontology. Graffiti also allows developers to define types or custom types that are well described, i.e. developer can specify what the type is about. So this is a person and a person does X, Y, Z, or this is a customer and a customer is defined as follows.
14:12And a customer has a company name, a first name, a last name, and several other fields. And that takes things to the next level in terms of being able to have a well understood ontology. And you can do so very simply using Pydantic. Graffiti is written in Python, and you can do so very simply using Pydantic models, which is something that developers have become very accustomed to when working with LLM APIs, such as OpenAI's structured output. And what that allows developers to do is build an ontology that is better structured for their particular use case. But Graffiti will still extract entities that don't match existing types.
15:01And so that allows you to get the benefits of this organic development and well-maintained ontology, but also have structured types that make sense for your particular business. now if you have for example a user type that you've defined and graffiti is interacting with an end user of some sort and it discovers oh there's a whole bunch of other contextual information that keeps showing up around users that it thinks is probably useful does it create a new type overlaying users does it extend your defined user like how do you think about this sort of like the fluidity of what you're interacting with people when you're dealing with the LLM extracted things, like it can come up with all sorts of stuff.
15:50Yep. So it would likely create associated entities with the user entity that describe an additional relationship. So for example, in the company type that I described earlier, if we hadn't put first name and last name of a particular contact at a company, it might create a contact entity type and that contact entity type has a name that makes a ton of sense so you have a set of types that you've defined as a developer that you are pretty guaranteed these are going to have these fields in this way and then there's a set of types that graffiti is interpreting creating on the fly relating to other types if for example it's a graffiti owned type and it discovers here's some new fields will it modify the underlying type or it will do the same trick of like okay this is a contact and now we have a contact phone tree because we know there's a set of phones or I don't know what the additional entities might be.
16:49Yeah, it would probably, a thing would be a phone number. So a phone number would be a thing, would create an entity for the phone number and relate it to the user or the company. So it's pretty clever that way. You mentioned changing data though, which is where graffiti really shines. And I hinted at that a little bit earlier with graffiti is a temporal knowledge graph. And it's a little bit of a mouthful where graffiti is specifically designed to deal with temporality and dynamic data. And this is very unlike other RAG frameworks. And in fact, I like to think of a major shift coming in terms of how we view building LLM-based applications.
17:47Well, it's actually here already. And that is we've shifted from these Q &A chatbots, so question and answers over a document corpus, powered by semantic databases providing RAG and RAG frameworks. RAG frameworks work really well with static data, but the way I describe agents and the way I strongly believe agents exist in this new world is in a sea of dynamic data. And I think we're in a post-RAG era now. Look, RAG didn't last very long, three years, but we're in a post-RAG era now and graffiti is a post-RAG framework. It deals with dynamic data and it deals with or supports a constant stream of unstructured or structured text, structured objects like JSON.
18:39And it's able to integrate this data into the graph in a way where it is understanding whether the new data or knowledge that it has created from the data is in conflict with existing knowledge in the knowledge graph. So for example, just a very stylized example, I purchased a pair of Adidas shoes six months ago from an e-commerce agent and the shoes fell apart. Sorry, Adidas, I shouldn't have actually used the brand. And I send them back to the return system, using the return system. That's not part of this particular agent that I spoke to. It's a different agent that manages returns. and I sent a nasty gram back.
19:30You know, I'm very upset that the shoes fell apart. And so my previous brand preference, because I'd had a conversation with the e-commerce agent and said, hey, I love Adidas shoes. So my previous brand preference was noted as being Adidas. But now I've sent these shoes back with a nasty gram saying, I'll never buy your shoes again. I'll never buy Adidas again. My brand preference has changed. And so when we integrate that knowledge, stream of JSON from the returns agent into our memory store, we've got to update that brand preference. And that's what graffiti does. It understands that the brand preference has changed and that it needs to invalidate that previous relationship where Daniel loves Adidas shoes to Daniel loves Adidas shoes, fact is no longer valid.
20:26So this is interesting. And there's a bunch of different pieces we can explore on this. I guess, first off, just like at a very vanilla layer, under the covers, does this essentially look like timestamps on any piece of data of when it was invalidated or when it became known and when it became invalidated? Or like, how is this implemented? So graffiti, and I'm going to get a little bit technical here, has a bitemporal model. And it has both episodic memory and semantic memory. And we're actually, in Zep's implementation of graffiti, we've also implemented procedural memory. So episodic memory is an event or a chat message or similar, and you send it to graffiti, and it has a created date, the event that you created that episode.
21:23And you can think about it as, you know, the episode might be a conversation that we're having today. It might be a single utterance in the transcription of this particular conversation. And it has a timestamp, and that becomes the created timestamp of any sort of relationships that we add to the knowledge graph. But I also might have mentioned that I purchased a pair of Adidas shoes six months ago. And so there's two entities there, Daniel, Adidas shoes purchased, and that relationship has a created date from today, but it has a valid at date from six months ago. And then when we now send the shoes back and we say we're really angry, we create a new episode.
22:14That's a JSON data event coming from the returns agent. And we update the knowledge graph by integrating this new information. In graffiti parlance, there's another date that we can add to the edge or relationship. The fact that Daniel no longer loves Adidas Shue, so we can invalidate that, which is an invalidat date. And this is how we capture the time dimension of changing state. And it allows graffiti to reason with the new event data that it receives. So I love this distinction between semantic memory and episodic memory. When you put those timestamps in place, are those also a timestamp and a link to the episode that resulted in this change?
23:06Okay. Yeah. Okay. Interesting. So in the graph, you have an episodic node, then you have entity nodes that are related to that episodic node. And so what you end up with, it's really interesting. You have multiple episodes that link to the same entities. And you can see longitudinally over time, state changes. and very importantly graffiti allows your agent then to reason with state changes. Oh Daniel used to love Adidas shoes but he's back again and he wants to purchase some shoes. I can still see that he is a roadrunner, pronates and has wide feet but the fact that he loves Adidas shoes is now invalid so I won't recommend those to him and he just mentioned that he thinks he wants to try out the Puma shoes.
23:57So we can add another entity to the graph that there's potentially a preference for Puma shoes. Yeah, yeah, yeah. Really interesting. Okay. So next question related to this is how do you conceptualize like multiplayer or things like that? So like in this case, let's go back and just say you haven't sent the nasty gram yet. So we know So Daniel said, Adidas is great. I love Adidas. Maybe Kevin said Adidas is terrible. Is that tied to each individual? Like, how do you consolidate knowledge? And is there a way in which, like, if you have specific user graphs, can they be consolidated or shared in some way across a set of users or an organization?
24:38Like, how would you deal with those layers? Yeah. So in graffiti, you would, in the data that you provide, ensure that graffiti understood that there was a different speaker or the data was related to a different user. You could do that in the JSON if you're passing a JSON event in, or you can do it in unstructured text in a transcription format. Maybe this is a good segue to speak to what ZEP is versus graffiti and how things work in ZEP. So graffiti is a framework for implementing temporal knowledge graphs. It can be used as a memory layer, but it's a very generalized framework. ZEP has first-class support for projects, users, sessions, or chat history threads, role-based access control, data governance, and privacy functionality that's all layered on top of graffiti, plus SDKs for Python, TypeScript, Go.
25:45So it's an end-to-end memory service that is framework agnostic, and we have the ability to implement ZEP within LandGraph agents, within Autogen, or without any agent framework. Here in ZEP, you can have user-based knowledge graphs. And by default, if you have single user agent interactions, it would go into a user knowledge graph. And it's well contained and ensures that user data is managed appropriately from a privacy and security perspective. But ZEP also has the concept of group graphs that can be used for multiplayer scenarios. So for example, you can stream Slack messages into a group graph and query the group graph independent of users, specific users.
26:43So get results back from multiple users. And so there's a lot of flexibility there. So my answer is yes. Let's maybe talk a little bit about then what this looks like from a developer standpoint, because I think one of the things that was interesting to me was this idea about like, yeah, automatically inferring, changing things like, Like, what am I sending to Graffiti or to Zep? And what do I get back? How do I query? Like, what does that actually look like for a software developer? Yeah. So in Graffiti, and I'll speak to Graffiti because I love the fact that it's open source. It's easy for developers to implement as well.
27:21So let's talk about Graffiti quickly. Any unstructured text, JSON, or even transcriptions you can send to Graffiti. It's a very simple API. and you can create graphs on the fly. You can have multiple graphs in graffiti. So you can approximate what we do with Zep in terms of putting a firewall between users, namespacing your user graphs. And what it looks like on the query end is, let me speak to a design decision that we took. Developers don't have to learn Cypher or some other graph query language to query graffiti. There are two reasons why we did this. One, we think that we can provide developers with a very powerful search framework without requiring immersion in a language like Cypher, which is a common language for knowledge graphs originally developed by Neo4j.
28:28And the other design decision that we made is we're not going to have an agent or LLM in the query path because that's super slow. So GraphRag can take tens of seconds to get a response back. And GraphRag is that static RAG framework that is built around a knowledge graph and was developed by Microsoft Research. And it is slow because there are all sorts of things that it's doing in the query path with LLMs. And that doesn't really scale in production. It means that you can't use something like Graphrag for voice agents. And those are becoming more and more common. We don't want to be sitting at a keyboard and typing.
29:13We want to talk to a device, for example. And so Graffiti offers a number of different ways to retrieve data. and we do so in a way that's super scalable. So the entry point into a graph is typically through an index, not through a graph search query where you have to then traverse the graph to find things that you're looking for. And these indexes are semantic and full text indexes. And so what we've done in graffiti is every time you add new data to the graph, I mentioned you have entities and you have edges between those entities. We place a fact onto the entity. Daniel loves Adidas shoes. And that gets indexed both semantically and from a full-text perspective with BM25 indexing.
30:10We also generate summaries for those graphs, which is for the entities in the graph. And those summaries basically are summaries of all the relationships that a particular entity has. So Daniel loves Adidas shoes. Daniel pronates. Daniel has wide feet. Daniel is a roadrunner. And we index the summary. And so that allows you then to search edges and nodes semantically and with full text or both. And your entry into the graph is in near constant time as a consequence. and so you're querying subgraphs so you're querying subgraphs rather than the entire graph got it okay so let me make sure i understand this so first data entry you throw over your stream of messages or documents whatever this doesn't have to be fast so this is when you use an llm to do extraction or do whatever other things that you're doing there it's done asynchronously yes done asynchronously you can do a lot of essentially pre-compute of all these different things and some of you pre-compute a set of summaries, you put things into an index, and the index also then points to the nodes.
31:21Then query time, you go straight to this fast index, it loads up for you a summary of a set of things that you can throw into your agent's context right away, and a sub graph if you want to do more searching that's there for you. Yep. Well, it's not just the summaries, it's also the fact on the edge, which is pretty well defined. And then you can also do graph search based on that. So things like retrieve the nearest nodes to the node that was a hit. So using something called breadth-first search, which allows you then to get a more comprehensive view of the relationship that you have retrieved from the graph.
32:04So you can see adjacent entities and their relationships. And that's done explicitly? Like I will get back and I can then do this? Or is that done as a part of that internal first query? So graffiti has a number of different recipes that have been developed that are kind of like best practice pipelines. But you can also create your own pipelines. And the types of things that you can do are pull together BM25 and VectorSearch and use something like Reciprocal Rank Fusion to join the two search results and then use a LLM re-ranker to re-rank the results. Or you can use, we have graph-based re-rankers as well.
32:50So you can, for example, re-rank results from distance from a centroid node. So if you want to get Daniel's specific facts back, you can say, re-rank all the facts by semantic. Well, looking at the semantic search results, re-rank by distance from the Daniel node. And so these are all baked in as simple recipes in graffiti. And they're super powerful. In Zep's implementation of graffiti, where we run our own embedding and re-rank her services, our own GPUs. You can search a graph in under 200 milliseconds. And the most expensive pipelines have a P95 of 300 milliseconds. This is super cool. I'm doing some pieces of this type of work right now.
Read the full transcript
33:38And I'm like listening to what you're saying. I'm like, holy smokes, I got to get my team on graffiti. Like this is cool. Yeah. So we actually benchmarked, We published a paper last month that describes how ZEP works and a deep dive into graffiti. It's available on Archive, and I think we can probably link to it in the notes of this podcast. And we benchmarked ZEP and ZEP's implementation of graffiti against the state of the art, the prior state of the art in memory, which is MemGPT. we actually found that the memgpt evaluation suite was too trivial. And so we used that to compare ZEP to graffiti, sorry, ZEP to memgpt, and ZEP was a far stronger contender there.
34:31But we also selected a far more comprehensive and larger evaluation benchmark called long mem eval. And as the name implies, it's for long memories and evaluating performance of retrieval across long memories, recall across long memories. And by long, I actually refer to 100 plus thousand token memory, so filling up a context window. And Zapp outperformed the baseline of putting the entire conversation and business data into the context for both GPT-40 and GPT-40 mini, outperformed GPT-40 by 18%, a mini by a larger margin. across a battery of 20 or so evaluations. And some evaluations where temporal understanding was required, it outperformed recall over the entire context window by almost 100%.
35:36This makes total sense because you can essentially pre-compute the relevance and you're only feeding in what is actually going to be really useful and relevant to the LLM. So it doesn't get distracted with all the other noise and all these different... No, it makes perfect sense. Yeah, so the needle in the haystack problem has not yet been solved beyond... Look, I don't like betting against models and the future of LLMs or any other architecture, whether it's diffusion or some other unknown architecture. However, today, the recall problem is not solved. And even putting aside improved recall, what we found is that with stronger models, they have performed better.
36:20So having your agent running with GPT-40 versus GPT-40 mini improved its understanding of the context that they're provided. So there's the promise of being able to offer far denser sets of information to the LLM and for it to then be able to make more cogent decisions based on that data. Absolutely. And you can load up only the relevant parts of the context for it. Exactly. I agree about not betting against models. But I mean, if you think about the way that we work as humans, we get distracted when you put too much in front of us as well. I don't expect the models to be that much better at it.
36:58Exactly. In fact, there's mental health conditions where if you have a storm of memories come back and other things that flood your consciousness, we also have issues. But the other dimensions that are very important to recall are latency and cost. And so as you described, if you're only pulling out the most important and relevant information that the agent needs at that moment, you can reduce cost and latency dramatically. So one of the other dimensions that we benchmarked in the paper were latency, reducing latency by 90 % and reducing token cost by 98%. This makes a ton of sense. So all this makes sense.
37:46I love it. I actually am really excited learning about this. What are the big problems you still see in agentic memory? What's not solved yet? And what are you kind of working on looking forward? Yeah. So we're very focused on production and not research. And so when you're building systems for research, you don't have to worry about real world consequences, putting services in the wild. And so what the benchmarks demonstrated to us is that there's still work that we need to do on improving graffiti and ZEP on a number of different dimensions. So there's ongoing work there. We also recognize that in large enterprises in particular, pushing data into memory is fraught.
38:39And so in production, you have all sorts of other things you need to wrap around a service, like a memory service, that touch privacy, that touch data governance, that touch security, and touch many different parts of the business that you need to get buy-in from. And so the real world consequences of building comprehensive memory for enterprise agents aren't just in the technology itself, but in being able to demonstrate that the technology works along many, many different dimensions. Not just that it creates the right memories, but that you ensure that data gets removed when it's supposed to get removed.
39:26Yeah. That it's secure, that it is safe, etc. So in a lot of ways, what I'm hearing is your focus right now is navigating the gap between cool prototypes and production. Exactly. Yeah. And so we have a lot of earlier stage pre-IPO companies as customers. Many of them have got compliance requirements as well. So we're SOC 2, Type 2 certified. We're working on HIPAA certification as well. We've seen a lot of interest from the healthcare domain. There is so much opportunity there for automation. But in the enterprise, we're still seeing folks moving from that prototype phase to stack selection. So building out reference stacks that can be rolled out globally across these large enterprises.
40:22And there, it's pretty early for large enterprises and agentic adoption. Yeah, this highlights one of the reasons it takes a long time for all these amazing things that we're seeing with LLMs to trickle out into impacting the whole world. Now, with those customers that you have right now, you probably have kind of an inside view on what's happening in the agent world. Can you share anything about what are the domains in which we're really seeing cutting-edge agentic applications? Yeah. So something that I'm finding most exciting is a rise of ambient agents. We've spoken a lot about human in the loop.
41:05Daniel wants to buy some shoes, etc. etc. Ambient agents are super cool, because what they're doing is they're just monitoring their environment and taking actions based on these changes in their environments. So you can think about an agent that is monitoring telemetry from a car and being able to understand that telemetry and provide proactive support to the driver around what they should do when something goes wrong. You can think about in a household environment, home automation, learning intelligently from occupation sensors, learning intelligently from your personal calendars as to when you're home and when you're not, coupled with present sensors for pets, being able to turn lighting on and off, adjusting the AC in the future, making food on time, responding to events like presence when there isn't supposed to be presence, et cetera.
42:28And so ambient agents offer these amazing opportunities for human agency because they're taking over a bunch of work that we needed to previously and can extend our consciousness as well, our understanding of what's happening in our environment, but also our, one of the biggest concerns when it comes to what the future holds from an AGI perspective. So a lot of the time in the popular media. And when people talk about AGI, they're actually talking about ambient agents and agents taking actions, not necessarily without human oversight. And so that's a lot of fun and very interesting. And we've been talking to actually some larger companies, thinking about ambient agents.
43:24Then we've also seen a lot of development. I mentioned in healthcare, around really automating things that were very, very expensive, processes that were very expensive to run, things like insurance claims, understanding coding of insurance claims. I don't know if humans understand that part, right? Like much less. In a prior life, I investigated medical billing and it's an incredible environment, It's such an adversarial relationship between healthcare providers and insurance companies that they actually have teams in competition with each other. But that's a different discussion. We're seeing a lot of use across B2B and B2C of agents.
44:15And everything from mental health applications for consumers, e-commerce applications, you name it. In the B2B world, there's some really exciting companies that we've been working with that are building, and this again is one of the frightening things about AI, building analyst tools to be able to comb through very large amounts of data for Fortune 500s and the national security industry. So basically analyst work suites that operate autonomously to develop reports and to take actions based on vast amounts of both open source and proprietary data. And so that's really interesting as well. Yeah.
45:03Wow. All right. So we're coming close to the end of our time here today. Is there anything we haven't talked about that you think would be important to touch on before we wrap up? I want to circle back on is the reframing of what agents are, getting a little more precise about that, and in particular, talking about memory again. We have a lot of folks come to us thinking in terms of RAG and static document corpuses, I think that's a solved problem already. And there is so much thinking that needs to be put into a post-RAG world where we're dealing with more dynamic data. And there are both technological challenges there, as well as human challenges, some of which are mentioned in terms of compliance, but others in terms of technology, there are areas that we're still pretty uncertain about in terms of how we offer really cutting-edge capabilities in production environments in a post-RAG world.
46:21And there, you know, we as humans, each of us perceives the world very differently. So our semantic understanding of the world is very different. That's why having different people's opinions is so helpful in decision-making because we see the world with different perspectives. So how are we going to help agents understand the world if we ourselves have different perspectives on what reality is? And that gets into some of the philosophy of reality. And that's an area that I'm very excited to see people explore. And I think it is incredibly necessary. And it ties into AI safety. It ties into the impact on labor, etc.
47:13Yeah, there's something really interesting there. And it reminds me of work I was pushing for at a previous job that didn't end up getting there. But like, when you have these, you're deriving these, I was going to call them facts, but they're not these knowledge entities. When we start talking about trying to perceive reality, we need to move from a single user to like a multi-user view. And you almost want like a Bayesian approach to like, we have these different supporting factors that lead us to an 80 % belief this is likely. And here's the things that might validate or invalidate our priors on this, like kind of having essentially probabilities and support structures associated with our facts.
47:52Yeah, I actually like framing it as a Bayesian problem. If one looks at the graffiti architecture, it is conceptually doing something similar to that when it does reconciliation, where it is looking at priors to understand how to form new memories or facts. But a formal Bayesian layer over something like graffiti would be very interesting. So adding some sort of Bayesian reasoning. Probability is a very challenging thing for humans to conceive. Absolutely. And as a consequence, LLMs are incapable of perceiving probability, but there is a architectural aspect of how LLMs work that can help us understand why an LLM produced what it produced.
48:56And that is looking at things like logits, the probability of the tokens that were produced. And so there's opportunity there. It would be interesting to explore. Yeah. There's kind of two sides there, right? There's the, what's the probability on knowledge based on multiple views of a thing coming in, multiple interpretations. And then there's even just like, what's the probability on the evidence that we've extracted because the LLM itself is probabilistic. Exactly. Yeah. Honestly, it's interesting. We see that with customers. This looks wrong. And you sit down with them and you say, but this is what it says.
49:39This is the input data and this is what is built. Oh yeah. Yeah. No, I see it now. And so then how do we get it into the structure that you'd like? And that's why we've done things like built custom entity types, et cetera, because we perceive the world so differently. There's a piece of this too that like one of the core ways I think about LLM applications, because they're probabilistic, because they have all these things is like, if you need reliability, or even if you don't, it's useful to think of them in terms of like, there's an inference step, and then there's a validation step. And then you go to another inference or something like that.
50:15And that validation might be a human in the loop. It might be some sort of formal checker. It might be a second LLM or some other thing. But what does that look like in a probabilistic world? And how do you kind of insert these validation steps as you build? So graffiti does implement reflection on things like entities it's extracted, conflicts it's identified, et cetera. And we've done this because we started building graffiti before reasoning models. We actually don't think at this stage that reasoning models are necessary for graffiti. It's costly, slow, and again, we look at how we scale things in production and cost.
50:57I think folks in your audience who have worked with knowledge graphs and LLMs and have worked maybe with GraphRag will know that it's very costly to produce graphs of any size. And so that is something that we're working on because we want to be able to commercialize this at scale, not just for large enterprises, but also startups. And so there's a lot of work that goes on there. In terms of the inference that you're doing from you have a stream coming in, you're inferring out these knowledge graphs of different sorts. Can that be done on a range of models? Do you need the top cutting edge models to get good results?
51:32Like what does that look like? Yeah, so we use Frontier models for our work. We have not fine-tuned any models. We might do so if we have some very domain-specific use cases. Frontier models have been able to perform well enough for our use case. We do run our own inference infrastructure for some use cases, not for graph building because of just the sheer scale of it. And so we do use Microsoft and other inference services for our use cases. You can use smaller models with graffiti depending on your domain and the complexity of your data. So you could run, for example, GP4 mini, or you could use Anthropic Haiku, or Llama 3.170B, or 3.270B, etc.
52:36If you have a pretty constrained domain, you may even get away with using much smaller models. And I find that a very exciting, the prospect of using smaller models is very exciting. I actually think that we're going to start seeing more complex product architectures from inference providers where what we perceive of as a model that we're using actually isn't just a single model. It is a set of different models combined together at different layers of what we would traditionally view as a model and solving different problems. Absolutely. All right. Awesome. Well, this has been super fun. Thank you, Daniel, for joining me today.
53:21Thank you, Kevin. It was a lot of fun. And we'll call that a wrap.
From the publisher
Contextual memory in AI is a major challenge because current models struggle to retain and recall relevant information over time. While humans can build long-term semantic relationships, AI systems often rely on fixed context windows, leading to loss of important past interactions. Zep is a startup that’s developing a memory layer for AI agents using
The post Knowledge Graphs as Agentic Memory with Daniel Chalef appeared first on Software Engineering Daily.
