In short
The episode argues that standard Retrieval Augmented Generation (RAG) is limited by irrelevant/duplicative lookups, repeated work per query, and inconsistent answers. It introduces Pinecone’s Nexus “knowledge engine,” which treats curated context as a precomputed, versioned artifact (analogous to a materialized view) rather than assembling context on the fly.
Guests and backgrounds
Jörg Shad, VP of Engineering at Pinecone; long career in database systems and distributed query optimization (Hadoop-era work), SAP HANA, Mesosphere/Mesos and early Kubernetes, CTO roles including RanguDB (graph/early graph retrieval, built a vector store), and NextData (data mesh/data products for AI/agents). Kevin Ball (KBall), VP Engineering at Mento; independent engineering coach; co-founded/CTO of two companies; founded San Diego JavaScript Meetup; organizes “AI in Action” via Latent Space.
Key claims
Nexus curates context once into versioned artifacts with schema, metadata, permissions, and lineage; improves reproducibility and auditability; routes agents to the right information using metadata/semantic layers; supports multi-modal context (vectors + structured fields + optional knowledge-graph-like relationships); uses a NoQL-style interface with structured queries and enforced response constraints.
Notable examples
finance/department/company context layering with permissioning; “yearly revenue” consistency; freshness metadata (e.g., context updated two weeks ago); extracting entities like dates/person names; pointer/lineage back to original documents; constraints like “revenue must be non-negative”; using confidence/probabilities in context edges and deciding whether to expose low-confidence facts to agents.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOMeet Jörg Shad
1:40 to 3:20
Discover Jörg Shad's background and journey in database systems and AI.
“Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.”
Exploring Nexus and Its Solutions
3:20 to 8:00
Understand the concept of Nexus as a knowledge engine and its role in data retrieval.
“So I recently checked all my code is out of the code base by now, but those were like interesting times.”
The Value of Pre-Computed Context
8:00 to 11:00
Learn how pre-computed context improves retrieval efficiency and reproducibility.
“And then for a particular task, I can actually go in and assemble those different contexts together into a meta context, if you want, or a set of contexts.”
Dynamic Context Metadata
11:00 to 14:06
Explore how dynamic context metadata enhances the usability of agentic systems.
“lookup pattern than was enabled by your previous database.”
Understanding Precompiled Context in Data Retrieval
14:06 to 18:52
Explore the concept of precompiled context and its role in data retrieval systems.
“So you have a description of the data set with some set of metadata.”
Understanding Precompiled Context in Data Retrieval
19:44 to 20:17
Explore the concept of precompiled context and its role in data retrieval systems.
“Your GitHub Actions bill is now a function of how much AI code you generate.”
Defining Context Datasets and Curation Processes
20:17 to 24:14
Learn about the specifications and processes for defining and curating context datasets.
“And when you're designing one of these context datasets, how is it defined?”
Quality Assessment in Curated Data Artifacts
24:14 to 28:00
Discuss methods for assessing and improving the quality of curated data artifacts.
“So you anticipated where my question was going to go, which is like around quality, right?”
Exploring Constraints in Data Sets
28:00 to 30:26
Discussion on the importance of identifying constraints in data sets and how they evolve over time.
“In the RDF world, I don't think we want to get into a discussion around which representation is best, but there's a long debate, shackle, et cetera.”
Bayesian Approaches to Confidence in Data
30:26 to 31:51
The potential of Bayesian methods to express confidence levels in knowledge graphs and data relationships.
“I'd love to dream a little bit more on this because this is one I've been thinking about as well.”
Show all 20 chapters
Designing Data for Agent Interaction
31:51 to 34:08
How to design data systems that empower agents to make informed decisions based on confidence levels.
“Yeah, well, and it gets to this interesting question that we sort of alluded to earlier of like, how do you design this for effective agent use, right?”
Metadata Integration in Queries
34:08 to 36:53
The significance of combining data with metadata for effective querying and response generation.
“because they iterate so quickly and they can form like a new plan on a much, I don't like the term cheaper, but just kind of like faster level, right?”
Curation and Context Management
36:53 to 38:19
The role of context curation in enhancing interactions and reducing computational costs in data systems.
“I think we may just, just to, just to step out.”
Abstraction in Software Engineering
38:19 to 41:27
How abstraction principles in software engineering can improve data handling and agent performance.
“I'm usually like in my mind, I'm still seeing that as a loop, but I know what you mean with linear.”
Common Patterns in Data Shapes
41:27 to 42:01
Identifying the similarities in data structure across different datasets and their implications.
“that's interesting coming back to what you're doing with nexus is you're essentially doing that same thing for data, right?”
Understanding Context Datasets
42:01 to 45:29
Learn about the similarities and differences in context datasets and the identification of core entities.
“I think the shape, especially if it comes to the entities we're talking about, this is different between different datasets, but I don't think the recipe changes.”
Evolving Knowledge Graphs
45:30 to 47:56
Discover how knowledge graphs can evolve and adapt over time even with static datasets.
“I think what is interesting is that even though imagine the data set being static, right?”
Context Artifact Schema Changes
47:57 to 50:49
Explore the impact of changing schemas in context artifacts and how it affects user interfaces.
“because it leads me down a path of like, can you even change the schema of that context artifact?”
Integration and Governance of Data Layers
50:50 to 54:18
Examine the benefits and challenges of integrating different data layers and governance issues.
“So freshness metadata for the underlying vector data is as we know how to get them.”
Future Trends in Data Systems
54:19 to 55:31
Anticipate future developments in data systems and the evolution of use cases for emerging technologies.
“And I think that's, again, when we said, well, let's talk again in a year.”
Transcript
Automatic transcript. May contain errors.0:00Retrieval has become one of the central problems in building useful AI systems. The standard approach to grounding a model in one's own data has been Retrieval Augmented Generation, or RAG, where an agent searches a vector database for relevant information at query time. That pattern works, but it has limitations, such as retrieving information that's not truly relevant, repeating the same lookup work on every query, and producing inconsistent answers to the same question. Pinecone is a vector database that's widely used to power semantic search and RAG at scale. The team recently developed Nexus, which is a knowledge engine that reframes context as a first-class, pre-computed asset rather than something reassembled on the fly.
0:47The approach borrows the database concept of a materialized view and curates context once into a versioned artifact that carries its own schema, metadata, permissions, and lineage. Jorg Shad is the VP of Engineering at Pinecone. In this episode, he joins Kevin Ball for an in-depth conversation about the frontier of retrieval technology. They discuss pre-compiled context, how context artifacts are curated and versioned much like code, how metadata and semantic layers help agents choose the right information, and much more. Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders.
1:30He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action Discussion Group through Latent Space. Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.
1:58Jörg, welcome to the show. Thank you so much for having me. Yeah, I'm excited to get to talk with you about this. Let's start with a little bit of your background. So can you give us the TLDR of your history and how you ended up at Pinecone and Nexus? Yeah, sure. I would say I was really lucky I could follow my passion, which is probably database systems where I started out almost 20 years ago with grad school, worked on distributed query optimization back in those Hadoop days, like really, really long time ago, worked there on name nodes and basically how can you use data duplication and then route to different queries.
2:35So that was really interesting times. And just looking back, it's super cool to see how it developed since MapReduce days. I kind of like followed along with that passion, was over at SAP working on HANA in the early days. And then at some point figured out that large enterprises here, maybe they're fun for a while, but not for the rest of my life. And joined a startup backstreet called Mesosphere. So Apache Mesos was like an open source project, somewhere in between open source version of Google's Borg system, their internal cluster scheduler, and kind of like pre-Kubernetes. So I built a lot of large scale systems across like Twitter, Netflix, Airbnb, where you can openly talk about.
3:19Also worked in the early days on Kubernetes. So I recently checked all my code is out of the code base by now, but those were like interesting times. Went back into the database space, was a CTO at a RanguDB, which is kind of like a graph database. Interestingly, I also worked on early graph rec, graph retrieval, when that kind of started up. We even built our own vector store in there, where we can maybe come back to a bit later when we're talking about vectors. And then I've been over at NextData, also working on how can we connect large-scale enterprises with data mesh, with data products to AI and agents.
3:59And I feel now this is actually all coming together in this one role, right? all those passions from data systems over infrastructure management, over actually connecting, creating end-user value from agentic systems by combining it with data. Okay, that's awesome. So today we want to talk about Nexus and kind of the patterns behind Nexus, right? How we are designing data, data retrieval data in different ways for agentic systems. But let's maybe start with just kind of the big overview of what Nexus is solving. What is Nexus? I saw it's described as a knowledge engine, but what does that actually mean?
4:37Yeah, a very good question. So let's maybe start with a history to understand the problem which we are solving, right? If we are following along, how do you... I mean, for me, the value of Gen.AI is actually coming from by combining it with your own custom data sets, which is either specific to the task, which are specific to your enterprise. You might not want to share them for privacy reasons, or they might just be very specific to that task. So I think this is where I've seen really big value being created. I mean, those early discussions like, should I fine-tune my LLM? I think most value we have seen is actually giving it context and making it specific to my task.
5:21And I think this is where, in the early days, Rack came up. We mentioned that earlier in my introduction. I've been working on GraphRack in the early days, kind of explored that a little bit. And I think Rack is great if you're exposing information. It's basically that retrieval pattern, right? The LLM knows it can look up for a piece of information. We're going to find like all the similar bits and pieces. This is when vector databases came up. We started having either dense retrieval for kind of like semantic search. We had sparse retrievals for kind of more lexical search out there. We added, often combined with like full text search in addition to that.
6:04So we ended up basically building an agentic retrieval system where the agent could retrieve similar amounts of information. I think just what we've seen then is that this is, it's helpful, but there is kind of that layer above. And I think this is where most people nowadays are using context for. So context is kind of a more curated version, right? So what we have seen often when the agent or the LLM, Gen AI, would start retrieving that information, it would have to do that over and over again. And it would then do operations on it. It would curate it kind of like ETL. It would just do that ETL kind of like on the fly and over and over again.
6:47So I think this is one of the steps where we can actually come up with meaningful pre-computed context. It's kind of in databases materialized views, right? I have a pre-materialized view of that information, which is pre-curated. And I think that's by itself enabling a number of different use cases. First of all, because we treat context as its own entity. we can put permissions on it, like similar is in a database system. You'll see I'll jump back to database analogies multiple times because that's simply my background. So apologies up front for that. But basically I can treat them as like a first-class citizen of the overall system.
7:30I can give it permissions. I can share it. And so if I'm, for example, seeing that in a larger enterprise context, I can have my personal context, which is just really information context specific to me or to what I've been doing over the last hour on this particular task. There might be department level context. So, for example, for the entire finance team, for the entire engineering team, which is shared amongst them. And then there can be context, which is actually shared for the entire company, because this is company wide knowledge. And then for a particular task, I can actually go in and assemble those different contexts together into a meta context, if you want, or a set of contexts.
8:16And this is then the set of relevant contexts for my particular task. So I think by just starting to treat it as such, I'm getting the benefits of system engineering we have been doing over the last decades from permissioning to aggregation on that. Second aspect of SADS, also with materialized views, I simply get more reproducible results, right? If we go back to what we said about RAC, if I'm doing that at each individual query, I'm doing it over and over again. And LLMs or agents, therefore, they're just probabilistic systems. They are choosing something on the fly. And so if I actually care about reproducible results, if I'm asking my data body question, what has been the revenue last year?
9:05I really want a consistent answer over time. I don't want it to be like fluctuating depending on which conditions are being used or what value is being taken for yearly revenue, for example. So I think that kind of reproducibility is another important piece, which makes it more reliable. And with that, I think I'm also getting in addition, I'm kind of getting, this might be more just an enterprise concern by itself, but I'm also getting this audit trails, I'm getting lineage trails. I can actually see this context because it's a curated asset has been generated with the data set from two days ago, 2 p.m.
9:51UTC, for example. Right. So I can basically go back there and I can also justify why is it in there. And I remember spending a lot of nights kind of like trying to rebuild this lineage tracing because we had to debug production issues and basically justify why was this here in the feature, right? If you're going back to like feature stores, why is this information ending up in features? And I think by just having the option for this lineage trail and being able to go back and connect it back to the source data, this is also, I think, from an observability governance perspective, is a big asset of using a context over repeated probabilistic retrieval steps.
10:35Yeah. So there's a lot of different things to dig in here. So I think some of the things you're highlighting are around these trade-offs between determinism and non-determinism and what's traceable versus what's not. But I want to kind of start, I actually really like using the database metaphor. So if we look at the concept of a materialized view, usually you're building that out because you have some sort of feature or functionality that you have planned where you need a different lookup pattern than was enabled by your previous database. One of the nice things about agents in general is they're so flexible, they maybe can handle many, many different lookup patterns.
11:10So how do you know what are the right conceptual materialized views to build out for agent context? I think one thing to add, I think it's different lookup patterns, but can also relate back to different permissions, right? If I'm building a materialized view over my secretive finance data, which I don't want to go out to any LLM, but I'm building a materialized view where I'm making sure I'm not including sensitive data, but only, I don't know, monthly aggregates, for example. And this is the interface where I can have my agents access that. This is, I think, another aspect to what you mentioned with the access.
11:48But I think it's not just the data access or the data shape. for that, but it's also just from a permission perspective. So this is where we spend a lot of time, for example, with Nexus as well, is trying to find that routing to identify from the metadata what is the relevant information to use, right? And I think that if you look back at this agentic development, for example, MCP endpoints, whatever MCP endpoints tools, what resource is so popular, it made it really easy. You just give it a description string and then it actually, the LLM could decide, right? I think LLMs in general, they are great planners.
12:30So you give them a tool and that tool is being described that I can do this. And then if you're not giving them too many tools, they're pretty good at discovering the right tool to use for the right job. You might give them some help with progressive tool discovery, for example, or like very like narrow set to not overload them with too many tools. But otherwise, they have pretty good planning tools also for like more complex pipeline steps. So I think this is the nice thing you can do with context is you basically, description sounds now just like a very long string, but you can actually generate like a also more structured description.
13:08What is in there? What is the freshness of it? What is potentially even lineage information? And because you're often just generating context, right? It's not that with MCP tools where someone goes in and is actually hacking a tool together and it's a static asset. Think context, even though we compare them with materialized views, even the description can be dynamic over time. So even that description can change with different versions of that materialized view to stick with that analogy. And therefore, I can really make sure that the description is always up to date. Imagine, for example, when you're trying to get the latest sales numbers from a context, and you know that a context last has been updated two weeks ago, then probably this is a sign that this is not very much up to date.
13:59And this meta information in the context metadata can actually, well, it gives the LLM a really good indication it might not want to use that. Yeah, that makes sense. Maybe let's actually break down, like when we talk about pre-compiled context, as contrasted to, you know, a lot of folks are probably familiar with some sort of vector database or Elasticsearch or other type of RAG lookup. Like, what is it that's happening here? So you have a description of the data set with some set of metadata. And then like, what is the shape of the data? How is it loaded? How is it even generated? Like, how are we doing the materialization?
14:33Like, what does this look like? There are different ways. Let's maybe just talk in general what I want. And then I can talk a little bit how we do that for Nexus in particular. But I think there are several things coming together, and I think this is the nice part. So when we earlier talked about how retrieval evolved over time, we talked about that there are different modes of retrieval, right? There's dense vectors, sparse vectors, there's full-text search. And I think this is very similar to what we are seeing in a context. So for us, a context consists of, for example, a vector index. So we can look up and do similarity search in that.
15:13It also has schema. So it can also have more structured information. I think one example we are, for example, often seeing as a pattern, especially if there are some sample queries around it. And we know this is a meaningful information. We would extract the dates. We would extract like person names. So basically one of the steps in context curation for us is we try to identify meaningful entities across it. Interestingly, just going back in like my history, this is a lot when you're trying to build like knowledge graphs. Right now we are not using a graph database here, but I can imagine like a knowledge graph, for example, also be a very useful tool.
15:57I think there's something interesting there too, because with similarity search and the ability to do embeddings on different fields, you can build kind of a fuzzy knowledge graph. Like it's not a strict graph, but you can have those relationships there. Exactly. I'll come to that in a second. We actually, we are building kind of like, when we come to Nexus, we are actually building a kind of cheap version of knowledge graphs. Exactly, exactly. Very similar to what you just described. But I think the key point is knowledge context is not like necessarily a single representation. It can be a combination of vector.
16:31It can be a combination of structured fields. It can be then also, as we just said, it could bring in a knowledge graph, for example. So I think it can contain multiple modalities of data and potentially even the same information in different modalities depending on what different query patterns. I think that the important piece next to that is that you have metadata associated with that. So I think we already talked about kind of like freshness metadata, lineage information. I think what we are also often seeing is just a reference to a semantic layer in there. So having like very well-defined terms so that I actually know if I'm talking yearly revenue, what is the definition of that field?
17:17What do we mean by yearly? Are we following calendar years or are we having like an offset fiscal year for some startups that have in the Silicon Valley? I think it's mostly startups going with that. But yeah, is actually calendar year equals fiscal year is such kind of meta information, which might be stored as part of the context or then as a reference to another semantic layer out there. But I think this combination of multi-modality representation of the data, of curated data, together with metadata, together with either an embedded or external, I think for scales, usually seeing externals are at least becoming more popular, is kind of that key part of bringing all of that together.
18:04Think about your mobile app source code. Once it hits the App Store, it's out in the wild. And without the right protection, decompiling is easy for malicious actors looking to steal your IP or tamper with your software. That's where GuardSquare comes in. GuardSquare provides the highest level of mobile app security for Android and iOS applications and SDKs. Their advanced tools integrate seamlessly into your CICD pipeline. We're talking polymorphic, multi-layered code hardening techniques and automated runtime application self-protection. paired with mobile application security testing and real-time threat monitoring to deliver the highest level of mobile app security without compromise.
18:47Don't leave your hard work exposed. Secure your mobile applications today. Go to guardsquare.com to learn more. If you're running Postgres in production, you've probably felt the moment analytical queries start fighting your transactional workload. Most teams end up adding a second database and all the pipeline complexity that comes with it. Tiger Data, creators of TimescaleDB, takes a different approach. We extend Postgres with hybrid, row, and columnar storage, so one table handles both writes and analytical scans. Native compression cuts storage costs up to 95%. Continuous aggregates keep dashboards live without bash jobs.
19:22And it scales to petabytes without you re-architecting. Companies like Cloudflare, Octave Energy, Schneider, Axpo, and Flowco run production workloads on Tiger Data today. No stale data. No second system to operate. Just Postgres. Managed for you, ready for the workload you're building toward. Try it free at tigerdata.com. This episode of Software Engineering Daily is brought to you by Warp Build. AI is writing more code than ever, which means GitHub Actions is running more than ever. Your GitHub Actions bill is now a function of how much AI code you generate. And every engineer knows the feeling.
19:53You push a commit, and then you wait. Warp Build makes GitHub Actions twice as fast at half the cost, with a one-line change to your workflow. Linux, macOS, and Windows Runners in Warp Builds Cloud or your own, enterprise ready, SOC 2 type 2 attested, and trusted by teams like Sky from Comcast, Bitcoin, and Braintrust AI. Get started with$50 in free credits at warpbuild.com. And when you're designing one of these context datasets, how is it defined? Do you have a spec? Do you have a pipeline? What goes into that materialization process? I think there are two different modes. Either I can create a general purpose, a context.
20:36So what we do for that, we analyze the data set, we're trying to extract and identify the key entities. And then, for example, this would be where we kind of like create like a more structured view on that with a fixed schema, we can also check. So that's kind of like a general where we try to identify general terms. The second aspect is if we actually know in which areas questions might come in. So if we either get emails or at least questions up front, which might be answered, this allows the curation step to be a bit more focused, right? So if I know I have to answer these questions, I can ignore things which are a bit outside.
21:18And I just, I basically have a direction. And then it's basically, I always imagine there's someone sitting there and curating documents, kind of like in the old days where you underline the important concepts or mark them in a different color. And then you actually bring them up and store them in another format. So it's kind of that step nowadays by an agent, obviously, where we go through the set of documents. In step one, we identify the format we want to create. So both the unstructured plus structured aspects of it. And then we basically compile, we actually version this kind of this step of how we want to curate it, right?
22:04So the curate step in the early days, it was an actual Python program. Imagine that just being generated. But having a version of that, it's actually another very powerful tool because the curation artifact by itself allows you to do that iteratively if the document, if like your databases evolves over time, right? I think so far we mostly talked about static data sets, but in reality, most of the data sets will be dynamic on one scale or another of time. And so I think by just having this two-step process of identifying how do we want to curate it and then actually apply the curation step, that allows us to get from a raw data set to a curation artifact to then the actual curated context.
22:55Yeah. Well, and it's interesting because we were sort of talking about this in terms of the different ways that agentic pieces play in, right? You just highlighted wanting to version the program that is doing the curation piece. We're sort of moving to a world in which like you can write code is a primitive that is available and code becomes like data where we version it. And it's just like kind of going through this. I mean, it was to a degree before, right? If you had like even ELT pipelines or the ETL pipelines, or just a Spark shop, right? Even that was code, you would actually version. And that was another kind of curation step.
23:35So I think it has existed before. It just was someone either writing the code or dbt generating that code or some kind of function generating it. And now it actually moves to more probabilistic tools, more agents generating it. And I think this is just why I think that versioning step becomes even more important. Because it might vary. And also having access to the old versions, it's fairly helpful for the agent to generate those tools. Ideally, even generate, if it makes sense, kind of like eval sets to be used later on for training. So you anticipated where my question was going to go, which is like around quality, right?
24:19How do you judge and iteratively improve quality of these curated artifacts? And I presume the curation process, as we've described, some of it's deterministic, right? You're writing code that's doing a set of things. Some of it's non-deterministic. It's got an LLM extracting or doing the highlighting that you mentioned. So how do you get quality and iterate on quality? First of all, even the code generation, right? The curation generation is, I think in most days right now, it's also being generated. So it's not someone writing code anymore in most cases. Ideally, I have a spec. I have a spec from where the code is generated.
24:53I think that's kind of the ideal standard. And this is also how we kind of see it. You have a well-defined specification. And then from there, you can generate the code actually doing it. In terms of quality, I think, as mentioned, there is, I think, the two use cases we are mostly seeing right now. One is kind of like this general thing. I want you to be able to ask, like, any question out there. Often it's a bit tough. We try to train the system wherever possible. We try to find some ground truth in there. And then in many cases where users come in with their datasets, they actually have a few questions they want to answer.
25:33And just by having a few short examples, we can actually generate a larger dataset for quality maintenance and iteration. And I think once you have that, it's basically an iteration step. Keep iterating, keep iterating, improving it, and move from there. Make sure that even with an updated data set, this is still true, for example. Now, when we talk about that check for quality, right, like we have these examples, is that at the level of the generated context? Or is it that at the level of the agent consuming it and using it to do the right thing? We see it as part of the system, so kind of stored alongside the context.
26:15Because you might have very different agents, right? You might have different consuming agents on the outside consuming it from an interface. If you see that again in just a larger context, you have your company-wide context. There might be a finance agent accessing it. There might be a marketing agent accessing it. So I think this should be part of the general inside system and stored alongside. Whether it needs to be part of the context, that's a good question. But I think it needs to be at least stored alongside the context. That makes sense. You highlighted the question of data sets changing.
26:56I think there's some interesting things to navigate in terms of dealing with stuff like conflicts between different sources or drifts between this particular data sources out of date or things like that. So how do you think about finding truth with imperfect data and keeping it up to date? Very, very good question. I think to start out, I think this is something you can encode in that curation process, right? So for example, when we talked about this first use case where we're trying to identify the entities. When we identify the entities, again, coming back into knowledge graph, in an ideal world, like how I would imagine that, not exactly what we're, for example, doing right now, but I would imagine you can also generate constraints, right?
Read the full transcript
27:45In knowledge graphs, especially RDF, you're even separating that kind of schema layer from the fact layer. In the first step, by generating the knowledge graph schema, by identifying the interesting entities, you can also generate the constraints around that. In the RDF world, I don't think we want to get into a discussion around which representation is best, but there's a long debate, shackle, et cetera. So you have a lot of options there to express constraints, but I would imagine this is probably that way. Right now, I said, I think you started the easy part, you generate a schema, which is already a set of constraints, right?
28:24You can put on then constraints on that. Hey, you always know that, I wanted to say, I mean, revenue needs to be like greater or equal zero. I have to think about it now from a business context, but just to make up an example, like you can, for example, say certain numbers need to be like strictly positive because otherwise they don't make sense. And you know that because you have identified that entity, like number of attendees in a meeting, for example, cannot be negative to come up with a good example there. And others might be then related. And I think the more time you spend analyzing that data set and also the bigger it is, right?
29:04It might be something you actually discover over time when you get more and more data in. You discover like new constraints. Again, there's a lot of work on knowledge graphs trying to identify these constraints and separate that schema layer and constraint layer from the actual fact layer in Knowledge Graph. So I would imagine this is probably a way where you can get that next step of always adding more boundary conditions. And then if we're a conflicting data set, you discover them, you can actually call that out. And I think that's valuable information by itself. If you can flag, hey, there's a conflict in the data set, probably at the beginning, that will require human interjection.
29:51But if you're learning that like, hey, this one data set, it's lagging behind and that's why it's not reliable. You can actually start, for example, ignoring the last two days because you know that second data set you're joining with, it's not reliable for the last two days. And then you only generate it like two days after if you have, for example, daily data ingestion. So I think this is where this curation step can then learn over time from kind of constraint violations. Right now we're kind of like dreaming. So I don't, I'm not aware of like any system doing that as of right now, at least.
30:26I'd love to dream a little bit more on this because this is one I've been thinking about as well. And in particular, like in my day job, I do a lot of work with things that are like essentially information about people and relationships where there isn't necessarily a ground truth that's going to be correct. And so one of the things I've been pondering is like, Can you create some sort of like Bayesian approach or confidence interval where you're trying to like, including that in the metadata for each particular piece of context and then like looking for ways to up your confidence or maybe even asking the end user?
30:57I mean, if we stay in the knowledge graph field, right, you can just have edges with confidence, right? So you can basically have an attribute on the edge. To take your example, right, I imagine you would build kind of like a network of people and then they have some relationships, know each other to start with our following to take like a Twitter slash X example. I think there you probably would have a ground truth, but you can just add an attribute on each edge saying this is the confidence I have there. And again, this is something which might evolve over time. If you're getting new information, Bayesian, as you mentioned, you can update your beliefs about that edge being true or not true.
31:41So I think that's basically if you view the context assets knowledge graph, you can embed the probabilities as part of that. I think that's probably the clean approach I would imagine at least like in quick as of now thoughts. Yeah, well, and it gets to this interesting question that we sort of alluded to earlier of like, how do you design this for effective agent use, right? Like what is going to make the agent most able to take advantage of this? And it might be that you don't want to give it those shades of gray. You just want to like have a cutoff. If it's above this confidence, it's in. And if it's below this, it's out.
32:20Yeah. And I think this is where agents are a bit different than humans, right? If we look at how is data consumed, so also just from a general database systems perspective, we as humans or even like dashboards, we're going to send over a SQL query or any query. We expect like one result and then we're going to fly with this result and just go forward. I think agents are both on the one hand empowered to use more because they can iterate, right? So an agent, I might give that result back. I'm only 5 % confident that this is actually true. And depending on which scenario I'm in, this might still be valuable information for the agent.
33:08And I think this is then where I think we can view that from two perspectives. I think, first of all, if we come from a governance perspective, and I think this is something we are seeing a lot with like Nexus as well. I actually, I want like a constraint, like anything where I'm less than 50 % certain, don't even give out to the agent. From a generalist perspective, if I trust the agent, I just want to give it as much information as possible. even this like, hey, I'm only 5 % certain in that can be something useful for the agent to just decide, like, what is the next step? What is my next planning step?
33:44Hopefully, it's not going to use the information, but it can use that meta information to actually take a different decision for its next step going forward, rather than just getting like, oh, no information available around that, right? So I think this comes down a little bit to the different use kits, whether you want to expose that information. But I think with agents compared to human users, because they iterate so quickly and they can form like a new plan on a much, I don't like the term cheaper, but just kind of like faster level, right? Because as a human, I start thinking about it and then I'm in like a second latency range.
34:21But from like an agent perspective, if I'm getting the outcome, oh, it's only 5 % certain in that, I can then turn around my question and not ask for specific data, but I could ask like, what facts are you more certain than threshold, for example? So I can actually change my approach for ring about data. Well, that brings us to a question about like, how do we expose this to agents? And I'm kind of curious, I saw that with Nexus, you have like your own little dynamic no QL language or something along those domains. But yeah, I was thinking about like, to what extent is this a pre-query that's pre-loading some things in context if we know it versus this is just a tool exposed to the agent?
35:01Or how do you think about the right ways to give agents access to this context? So I think from an agent perspective, I need a combined answer of data plus metadata. So for example, that answer we just talked about earlier, like I believe the answer is five, but I'm only 10 % certain about it. First of all, if I can get both those facts back, This is something valuable. NoQL, for example, and also we were trying to evolve and also get like more standardization around that. This is, for example, the other thing we have seen, which is super valuable for agents. If I can expose schema, at least schema on the output I want, I want that result that should be US dollars.
35:53So if I'm asking, like, what is the price of X? And I, as an agent, I can specify in a deterministic format and really want to make sure that I'm getting it in U.S. dollars as opposed to euro, for example. And I can specify that as part of my structured query, right? So no QL is kind of that pair of structured query where I can also interject metadata. It's not just a SQL query. I can actually put more in there. And secondly, I'm getting the same on the response side. I'm not just getting like the structure is the table back as with a SQL table, but I'm actually getting data combined with that metadata plus certain constraints are automatically enforced.
36:40So we can automatically check that all of those are numbers because I said like that price field, it should be a positive number lower than 10 ,000, for example. Got it. I think we may just, just to, just to step out. I think in that, in that agent interaction, how do you think about it? Because in most cases, it's not going to be like a one shot interaction. I'm not like getting one response back. I am typically trying to give back and interact in this pairs of data and metadata, because that allows both sides to actually give more relevant answers and understand the context to reuse that term once more around the data which has been given, right?
37:25It's not like one shot interaction, one result going back, but it's actually the combination which makes it powerful. Of course, also on the downside of those loops, they can get expensive, right, if we just talk token usage. And I think maybe this is coming back to where we talked about earlier, like this curation. I think we want to keep that loop of planning as, we know we're like super deep on that right the less layers we can have in there we have to do on the fly the better so i think if there is pre-curated context you have to do that only once you have to do this loop of actually discovering you do that once up front that might be expensive but then in the next interaction you can actually benefit from it and we already know what to give you back and we don't have we have to do like one iteration pipe less because there's still going to be plenty of set going on in the system yeah i've heard it described as like llms are really good at things that they can sort of describe linearly and the more layers of abstraction you have to handle all at once the more they are both it's expensive it takes a lot of their own like internal expensive key value memory to to track that context going through but it's also less reliable and so when you can like flatten those steps so So for each LLM interaction, it's only having to track essentially a smaller linear interaction, but it may sub-call out to things that are doing those depth.
38:56I find your linear abstraction. I find that interesting. I'm usually like in my mind, I'm still seeing that as a loop, but I know what you mean with linear. I'm just trying to avoid like a sub-loop at like one step in that loop. Exactly. So if I can avoid that level of nesting of loops, the better. Yeah. Yeah, well, and you want to offload that to your context system, right? So it's doing that sub loop. It also has to deal with exactly one layer of loop. Exactly. I mean, sorry, it's just coming to my mind right now. If you look at the software engineering, right? It's kind of like it's an abstraction, right?
39:30In software engineering, I'm also trying to pull out like common functionality into like one function. Because first of all, for us, it makes it easier to maintain code. I can reason about like one sub part much better because I can also treat it independently. So I can treat this like one sub loop, which is focused on curation or preparing that like independent of that larger loop retrieving it. So it's kind of like composability and software engineering as well. Oh, absolutely. Well, and I think one of the like superpowers for dealing with LLMs and programming is thinking in terms of domain-specific languages because LLMs are very linguistic, right?
40:11They think in language. And so if you can create a language that's the shape of your problem, LLMs are very, very good at utilizing that. And then to your point, right? You decompose that. Okay, now I have a new set of abstractions I need to build. I need to build those different functions that create my DSL. Yeah, I mean, I'm always a bit careful because you can go overboard with that as well and then make it too complex. But I think to a degree, the less we as humans have to, I mean, this is the power of abstraction, right? We don't have to keep everything in mind, but we can only talk about an abstraction here.
40:45And I think the same is true for the LLM. I think you described it very nicely. Like if you don't have to go in the sub loop, which is also different contexts, right? So basically the LLM has to change its role potentially and still wants to store kind of its previous context, but it's doing actually a different role. the more we can avoid that the easier same as for us humans if i'm looking at a code and there's an abstraction for a class with certain rules i only have to kind of rationalize and think about that abstraction and then at another time i can actually go into the implementation of that if i want to but from like that high level code i can actually stay at that interface level so i think i think that's interesting coming back to what you're doing with nexus is you're essentially doing that same thing for data, right?
41:32You're creating a domain-specific data object, a thing that is not only abstracting out the code layer that generates it, but abstracting out kind of the form of data that's going to be useful for a particular problem domain. Yes. So I guess the thing I'd be wondering about there, and this is more of an observational thing if you've seen a bunch of, how much variation is there in what makes for that shape to be useful? Obviously the details of what's exposed, the permissions, those things are going to vary, but like how similar are the shapes across these different types of context datasets? I think the shape, especially if it comes to the entities we're talking about, this is different between different datasets, but I don't think the recipe changes.
42:18If you had to represent it as like a class hierarchy of what's playing together, I think that underlying structure is going to be very similar. So the scheme of it is going to be similar. the actual instantiation is going to be different. So what entities do I have? But usually I'm having some core entities. And I think we said earlier, if we promise to come back to like that narrow knowledge graph representation, we actually do have and we do detect relationships between entities as well. So I think this pattern of identifying a small, fairly small core subset of entities, which are central to that, right?
43:00It's going back to like document analysis, like different scores of identifying what's relevant to a document. It's still like very, very much here. I think we're using different algorithms, different ways to detect that. But I think overall, we're still trying to identify what is core to this document, what is core to this data set. And then from that, we can identify that small, potentially related subset of entities We parse it into that. And on the side, we're having the vector search for finding similar instances of these entities. We're having the kind of more structured representation, and then we're having the metadata next to it.
43:41So I think that it doesn't vary so much. It varies just from the different instance creations. I'm curious. So if I use the very particular domain I'm in, one additional thing we often find useful is having essentially quotes that support whatever the extracted entity is. I'm curious if that's a pattern that you see. Yeah. I mean, what we do is, I think we talked about in the very beginning, right, the lineage. So we always keep a trace back to where it's coming from in the original document. So kind of like a pointer back to the original document. because yes, you're 100 % right from like multiple angles.
44:18I think, first of all, in many cases, we actually would recommend the system to retrieve the ground truth in addition, right? So if it has a chance, if it's not too much content, it retrieves the ground truth as well because it might even give you more information. And I think this was already true in retrieval, right? I think a typical retrieval pattern is you find the similar vector, similar entity in the vector space, but then you still kind of retrieve the original data to back that up and feed in. And then I think this is still the case also with context. I think with context, because we already have a structured representation, I have to do that less.
44:56But we definitely, we give back a pointer back to the original data, to the original place in your document. And again, coming back to then also the version document, right? If my data set is evolving over time, we need to make sure we are referencing the right version of that. Yeah, that makes sense. I'd love to dig a little bit into this knowledge graph concept, and in particular, the fuzzy vector-based or semantic knowledge graph. So how are you thinking about those relationships in a world that maybe is not as perfectly structured as your classic knowledge graph? So I think the interesting aspect of this, I think that's something we didn't talk too much into deep about yet.
45:36I think what is interesting is that even though imagine the data set being static, right? I'm creating like one context artifact. I might create like one context graph, one knowledge graph instance associated with that. Over time, when I'm traversing that, I can actually keep modifying my knowledge graph. When I'm initially creating my context, I have a certain budget of compute. Imagine every night your GPU cluster lies idle, and you're actually going to show that on your previously curated context, even though the initial data hasn't changed. You can try to find, I like your example earlier, like imagine in your knowledge graph, you have certain probability values where you think this is true, that edge should be in the knowledge graph or not.
46:28And then overnight or whenever you've got free compute time or idle compute time, you actually keep curating your knowledge graph and either collapsing it, expanding it, using graph machine learning to identify new connections, try to group things into certain clusters, and actually keep curating it, potentially also a Bayesian approach with additional questions I've seen come in. So I actually get more data, not by the initial data changing, but by actually just seeing the questions coming in. So I know what I might want to answer. And that can help me curate that over time. And I know right now we talk a lot about knowledge graphs.
47:11I don't think that necessarily has to be a knowledge graph. I can do the same curation over time by spending more effort on it. If it's in different representations, I can do the same. if it's in a more relational format, potentially with additional edges in between those, with like semi-edges, right? I can always put a graph format into relational schema as well by encoding that as joins. Might not be efficient to retrieve, but I can do it. So I think even there it's true. So even if I don't have like a traditional knowledge graph, I can still maintain that over time and kind of keep curating my context artifact and making it better, especially if I've seen some of the questions being asked against it.
47:55So that's interesting, actually, because it leads me down a path of like, can you even change the schema of that context artifact? Because you see, for example, hey, we're reliably getting questions that end up having to query across two different sources of context or things like that. Maybe we should provide a combined view or what have you. Denormalization, right? And I remember that, well, it's like 20 years ago, teaching the different normal forms in relational database systems. Most certainly, I think this is so nice that you have the abstraction of a context, right? A context is telling you, hey, this is actually the schema I have in there.
48:34So from like a consumer perspective, I'm still, I'm not changing the interface, even though I'm changing the schema inside. I just have like version two, version three, version four of my context. but it can have a different internal representation. That's awesome. Maybe kind of like a container. I mean, I've worked on data products before. You could even see that it's like a small data product combining the data and metadata into one entity. And then as a data product or in the microservice world, like the Docker container, it gives you still like an interface. You can change the internal representation and evolve that over time as well.
49:14Yeah, very good question. Well, and it's interesting too, because like you have some sets of data and especially metadata where you're wanting to do repeatable deterministic types of things on it. And so you don't want that to change too much. But a lot of this is like prepared for an agent to consume as a programmer, you can treat it as a black box and change that, that internal representation. We talked about NoQL before, right? So kind of like our query interface there. I think this is one of the nice things there because in NoQL, I can basically also give you a schema I expect back, right?
49:48So I'm keeping that interface static. But again, because the interface is given by NoQL and we know how to map it, we can change the internal representation and might be more or less efficient to retrieve. But I have that flexibility through that abstraction. So all of this, we're kind of talking about a layer that is developing now that is on top of these graph databases like Pinecone originally was, or we talked about being on top of a relational database or something like that. Do you see this as being a separated abstract layer that there's going to be a set of different options out there in the world?
50:26Or is this something that's going to be tightly integrated and vertically stacked with these kind of underlying data stores? So I think for us, it's actually we benefit a lot from owning the entire stack. If I look at our implementation, at least, we benefit a lot from owning that end-to-end because we can just expose metadata. So freshness metadata for the underlying vector data is as we know how to get them. We don't have to duplicate that potentially getting out of sync. Maybe just to give one example. So I think for us, we actually benefit from implementing it really end-to-end. And I think this is giving us better performance.
51:11I think the other aspect where it's going to come in is governance. We talked about different access, permission levels, different service accounts, having access to different things. So in theory, yes, you can definitely develop an abstraction layer over that. I think that will take quite some time to identify what do we all need to expose there. So what I would imagine is probably at the beginning, different people will develop different solutions. And then certain standards are going to evolve. Like how we're thinking about it, I think for us, the layer, which probably is going to standardize first, is going to be a little bit on this query front, of course, the NoQL.
51:59And then from there, let's see where it takes us, right? We can also probably standardize some of the knowledge definition, the knowledge spec. But I said, like right now, also just from like a governance perspective, keeping track of lineage, it is a very beneficial of for us controlling the entire stack. I would at least give it like another two years of iteration. And then maybe we have identified all the patterns and we can drive that out into a general spec. Yeah, that makes sense, right? We're very much in this phase of the tech isn't actually good enough yet. And so there's a lot of benefits to squeeze out every piece from that vertical integration.
52:37and we'll get to a place where, oh, okay, now we've over-served that and it benefits us to split it apart and optimize different pieces. So I think it's going to start at the query front. With query, I mean actually query and response, right? So that kind of format. And then it's probably going to go down to that stack. It's just going to make, I think, some of the benefits you get from that, except like lineages, I think it's something super helpful to be able to trace it back to the initial days. That's going to get a bit tough if we go into a general abstraction to just get it through like these different systems.
53:14But who knows, lads? I'm looking forward to talking to you in a year again and see how we have evolved then. Yeah, absolutely. Well, we're coming close to the end of our time and we've covered a lot. But is there anything we haven't talked about that you think would be important to discuss before we wrap? No, I think we actually covered it very nicely, right? I think we looked a little bit at the history, what actually led us to develop something like Nexus. I think that problem like agents actually retrieve, how can we build a system for agent retrieval? How can we make that efficient? So I think that's something very nice where we started.
53:51I think we branched up into many different areas. So I think we really covered most of that. Also, we already covered what I had here in my last notes, kind of the outlook. I think the outlook where this is going to grow is we're going to see more and more use of integrations with different semantic layers. I said, like right now, we do that internally. We're going to see like a lot of iterations. And I think the other side which we're going to see grow is just more people using it. And I think that's, again, when we said, well, let's talk again in a year. I'm really curious what use cases we'll have discovered in a year.
54:28So already, when did we launch Nexus? That's like two months ago. We already have seen so many different interesting use cases, which actually also helped us shape some of the internal implementation. So I think that has been a super interesting learning experience for us. And I'm personally really curious how that's going to continue going forward. So right now, I can just say super exciting times. It really is. I mean, I think it's fascinating how we're all kind of trying to rediscover how do we package data for agents as a primary consumer. And we've talked about a bunch. I think, honestly, the coexistence with the metadata and having that be key both on the definition and retrieval side, or like query and retrieval and how all of that tracks, like that's a huge step forward.
55:14And we talked about semantic layers. I'm hearing all sorts of people were talking about, oh, we need much better semantics for agents because people just put that in their heads, but agents need it right there. Like, so yeah, discovering these patterns of what needs to be co-located now and what needs to be described. It's a fun time. Yeah.
From the publisher
Retrieval has become one of the central problems in building useful AI systems. The standard approach to grounding a model in one’s own data has been retrieval augmented generation, or RAG, where an agent searches a vector database for relevant information at query time. That pattern works, but it has limitations, such as retrieving information that’s not truly relevant, repeating the same lookup work on every query, and producing inconsistent answers to the same question.
Pinecone is a vector database that’s widely used to power semantic search and RAG at scale. The team recently developed Nexus, which is a knowledge engine that reframes context as a first-class, precomputed asset rather than something reassembled on the fly. The approach borrows the database concept of a materialized view, and curates context once into a versioned artifact that carries its own schema, metadata, permissions, and lineage.
Jörg Schad is the VP of Engineering at Pinecone. In this episode, he joins Kevin Ball for an in-depth conversation about the frontier of retrieval technology. They discuss precompiled context, how context artifacts are curated and versioned much like code, how metadata and semantic layers help agents choose the right information, and much more.
Sponsorship inquiries:
sponsor@softwareengineeringdaily.com
The post Moving Beyond RAG with Precomputed Context appeared first on Software Engineering Daily.
