In short
Summary of TWIML AI Podcast Episode #664: Are Vector DBs the Future Data Platform for AI? with Ed Anuff
Episode Details
- Podcast Title: The TWIML AI Podcast
- Episode Title: Are Vector DBs the Future Data Platform for AI? with Ed Anuff
- Episode Number: #664
- Host: Sam Charrington
- Guest: Ed Anuff, Chief Product Officer at DataStax
- Date: [Insert Date of Episode]
- Show Notes: [twimlai.com/go/664](https://twimlai.com/go/664)
Episode Overview In this episode, Sam Charrington interviews Ed Anuff about the evolving landscape of vector databases (DBs) and their implications for AI and machine learning, particularly in the context of retrieval-augmented generation (RAG). They discuss the architecture and performance of modern vector databases, the role of embedding models in data retrieval, and the challenges in ensuring relevance and performance as datasets scale.
Key Topics Discussed
- Introduction to Ed Anuff
- Background in startups and tech, with experience in notable companies like Wired, Google, and DataStax.
- DataStax is noted for its Cassandra database, which is popular among major companies like Uber and Netflix.
- Vector Databases and RAG
- Ed explains how DataStax is pioneering the use of vector databases for handling large-scale, unstructured datasets.
- Vector search capabilities have been integrated into Cassandra to enhance retrieval for RAG, AI assistants, and other applications.
- Vector Database Technologies
- HNSW (Hierarchical Navigable Small World):
- Initially used by many vector databases for efficient querying, but not optimized for disk I/O.
- DiskANN (Disk Approximate Nearest Neighbor):
- Newer approach adopted by DataStax for better performance with large datasets.
- Addresses the limitations of HNSW in distributed settings, particularly for disk-based operations.
- Challenges of Relevancy and Performance
- As datasets grow, the challenge of maintaining relevancy increases. Larger datasets lead to potential drop-offs in accuracy and performance.
- Importance of embedding models in generating context for queries and retrieval processes.
- Embedding Models
- The role of embedding models in transforming queries into vectors for database retrieval.
- Trade-offs between smaller, cost-effective models and larger, more performance-oriented models.
- Data Preparation and Context Construction
- Emphasis on the importance of data preparation and context building in RAG applications.
- Discusses RAG as being more complex than just using an LLM; it requires careful consideration of how data is ingested, chunked, and presented.
- Future of Vector Databases
- Predictions that vector databases will evolve into both features within traditional databases and standalone platforms.
- Discussion on the need for abstraction in data retrieval processes to make it user-friendly for non-experts.
- Integration of GPUs
- Exploring whether vector databases should be GPU-native and how they can leverage GPUs efficiently without incurring high costs.
- Optimizing for RAG
- Importance of optimizing systems for retrieval-augmented generation and the competition among vector databases to meet these demands.
- Predictions that RAG will be a dominant use case for vector databases in the near future.
Key Takeaways
- Vector databases are evolving: They are increasingly integrated into traditional database systems and are becoming essential for handling AI-related tasks.
- Performance vs. Cost: There is a critical balance to strike between using large, powerful models and maintaining performance within cost constraints.
- Relevancy is critical: As datasets increase in size, maintaining the relevancy of the output becomes a significant challenge which needs to be addressed through better design and tuning.
- Future Integration: The future of vector databases will likely see better integration with existing systems, making the technology more accessible to a broader range of users.
Conclusion This episode provides deep insights into the implications of vector databases for AI and RAG, highlighting the technical challenges and innovations in this rapidly developing field. The conversation emphasizes the need for continued optimization and the potential for vector databases to transform the data landscape for AI applications.
For more details, visit [twimlai.com](https://twimlai.com/go/664).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:09All right, everyone. Welcome to another episode of the TwiML AI podcast. I am your host Sam Charrington. Today, I'm joined by Ed Enough. Ed is Chief Product Officer at DataStax. Before we get into today's conversation, be sure to hit that subscribe button wherever you're listening to today's show. Ed, welcome to the podcast. Thank you. It's great to be here. I'm looking forward to our conversation. We've got a bunch on the agenda. We'll be talking RAG and vector databases and assistance. But before we do that, I'd love to have you share a little bit about your background. In fact, we've got RPI in common.
0:44Yeah, we do. Yes. Which I think is probably a very chilly place this time of year. So it's been a while since I've been back. Have you been there recently? I wouldn't say recently, probably five years ago was the last time. Yeah, same for me. Great school for those who haven't been there, though. Small, great tech school, but in upstate New York. And one of the reasons why I chose to move out to the West Coast when I graduated was the winters there. absolutely so tell us a little bit about how you got from there to to here yeah so it came out to the west coast really wanted to get into startups and everything that was going on and of course this was the early days of things like internet well it was in pre-internet multimedia and all that but shortly thereafter internet happened and was over at wired in the early days doing the search engine and did a whole bunch of stuff, started a company in the enterprise Java space called Epicentric that had a great run, went on to do some other cool stuff, social media, advertising, blogging, was at Six Apart for a while, the company that made Movable Type and TypePad and ended up part of Apigee, the API management company.
1:58We had a great run there too, did an IPO, got acquired by Google. and after a few years at Google, decided to come over to DataStax, which is the company that makes Cassandra, the Cassandra database, and have been doing that for the last few years. So a bunch of cool, fun stuff, primarily making stuff for people that are building websites, building applications, building content. That tends to be the type of stuff I like to do. I totally forgot about your epicentric connection. I was a very early employee at Plumtree. Yes, that was an exciting time. Awesome. So tell us a little bit about DataStax has been kind of active in helping organizations kind of take on this challenge of using LLMs and RAG.
2:43Tell us about DataStax's kind of angle in that. Yeah, so as I mentioned, DataStax is the company behind Cassandra. And Cassandra was really the original cloud-native database. So an awful lot of companies, whether they're using Uber, whether they're using Netflix, Apple, these are all companies that use the Cassandra database. And when you do something like FedEx package tracking, that's all on topic standards on data stacks as well. And so we knew pretty early on that as people were looking to first with ML and then as AI and Gen AI became a big thing, we knew that that was going to be pretty important.
3:19That people would want to use the data that they had in these systems that power all these interactions. They'd want to add AI to it. And so we looked at how to add the vector capability, the vector search capability to the database. That's something that we did. We did it both within our AstroDB, that's our cloud service, but we've also, everything we do is also in open source. So Cassandra 5.0, it's part of the Apache Foundation, has this vector query capability. And as you know, when you want to get a database to work well with an LLM, and we'll get into it, I think, probably in depth in a little bit when we talk about things like RAG, but the starting point is to allow you to go and have your LLM within some architectures, and we'll talk about RAG, I think, in depth.
4:04But you want to be able to go and retrieve information from the database on a vector-based query. And so what we've done is, and we actually did a couple iterations of this. At first, we implemented, much like most of the other vector databases that you see, we took HNSW, which is the hierarchical, navigable, small worlds approach of adding vector query capabilities. We brought that to Cassandra. We've since switched to something that's called disk ANN, which is a different approach that is more oriented towards optimizing disk I.O. We've talked a little bit about that, too. But the goal here is to go and give you Cassandra, which is one of the most scalable databases for people that don't know, built on technology from Facebook, from Google, from Amazon, and gives you a scale-out data capability.
4:55But how do we bring this vector capability to these large data sets? That's really our angle on all of this. Can you drill into the HNSW versus disk ANN and what those mean and what are their implications? So let me talk a little bit about sort of the origins and why everybody has been using these things. One of the amazing things about, as companies have gone and approached a lot of the build out on AI infrastructure has been that they've been using a lot of stuff that's been around for a while. And in fact, if you talk to folks, they'll be like, oh, well, you know, approximate nearest neighbor and so on.
5:31Like this has been around for a while. And in fact, there are better approaches that are coming down the pike on this. But the HNSW implementation that most of the databases, most of the vector databases out there ended up using, what came out of Lucene. And Lucene was one of the original, sorry, I really asked the search engine. But at this point, it's more of the search library, search infrastructure, index creation that a lot of folks use. And I'll definitely get this wrong. So actually, I won't name names. But other than just say most of the big names that you would have heard about and that we all talk about when we talk about vector databases, they all started with the code that was in the Lucene HNSW implementation.
6:11And in some cases, they may have ported it over to their language of choice because not all these things are written in the same code base. but they started with that. And that was a good starting point. Like if you were the fill in the blank, the original vector databases that were named in the OpenAI blog post that kicked off the whole vector database race, or the people who came out very quickly afterwards, it's a good starting point. The problem was that HNSW was not particularly optimized for disk IO. So what we found was that as we started to deal with the performance of this, and particularly with Cassandra that's a distributed database, we went and looked at sort of the structure and essentially the levels and edges of the graph that gets constructed.
6:58Because basically you end up breaking these things down. This is where the hierarchical part comes from. And then you end up having to do a partitioning, particularly where, you know, in a clustering, basically, this really comes home to roost when you're using a distributed database. Because when you look at the way something like Cassandra works, it actually goes and takes a query, vector non-vector, this is just the way Cassandra works, it takes that query, farms it out to a whole set of nodes. Like people who are running Cassandra have thousands of nodes. And so if you don't have something that first of all allows us to map that hierarchy of where the semantic position is of something in that cluster, and then further doesn't go and reduce the number of disk IO operations, you end up hitting a performance wall.
7:46And one of the dirty secrets of vector databases is that they all perform really good on small data sets because they effectively end up defaulting to what's in memory. And in fact, in the early days of when these vector databases were coming out, a lot of developers you'd read on Hacker News and elsewhere would just joke that like, for these demos people are doing, you can just do everything in memory on your laptop and you'd get better results. So that's where a lot of the stuff that we end up looking at for us, particularly for Cassandra's role in this stuff, we want to be the one that's doing really large data sets.
8:20For example, one of our demos, one of our rag demos that we use has the entire Wikipedia, Wikidata data set, and that's something like 500 million documents. And you can actually start to see the issues start to crop up in as few as 100 ,000 or so documents. But definitely when you get into the million plus documents, you start to see that these trade-offs actually have very tangible results in terms of relevancy, not just in terms of the performance, but actually starts to cause fairly significant drop-offs in relevancy, and you just start to get to junk back. And you see particularly where sort of all of the vector databases right now are sort of in an arms race with it, because really the business payoff on the other sort of important comment about vector databases to understand.
9:02And databases, I won't spend a lot of time talking about the database business today, but all the databases, whether you are us, whether you're Mongo, whether you're Pinecone, whether you're whatever, these are consumption-based businesses. And it means that we all love and we all do the hackathons and all that, but generally the businesses are built on having very large data set. If we can't vectorize that very large data set and make it possible to do like production rag at scale, it's going to be really hard for anybody to build a business around this. And so if you talk to anybody who's running a vector database, a lot of the focus is about how do we make it feasible and cost effective.
9:37There's a big cost dimension to this as well to do these really large ASINs. And that's why all of these choices, like we're really happy about disk ANN right now. We have a lot of stuff that shows it outperforms in real world, a lot of other options, but we're already going and looking at some of the stuff that some of the different approaches that will move beyond that. Some of them specifically to distribute a database space. disk ann is a an algorithm and approach for doing approximate nearest neighbor on disk presumably it's an infrastructure aware approach for approximate nearest neighbor exactly that's just principal difference there are a whole bunch of other pieces to it that again i would do too crappy a job of in terms of going into detail on but the principal biggest difference that we've seen has been we've been running this stuff as a service since early this year and so we've we've had the privilege to see a lot of the sort of real world, what happens when you do rag in a real world setting.
10:34And so when you start throwing a lot of IO operations, disk accesses, the problem, you start to see a lot of things that you wouldn't see otherwise. And you wouldn't see, for example, in a classic chat GPT scenario, because they're not, they do now because now chat GPT is actually does a lot of RAG as of a month or two ago, but you wouldn't have seen previously in sort of the conversational scenarios that we're really just throwing things at the model. And in as much as the database is involved, it's being used for conversation history, which is a part of RAG, but it's not where sort of the bread and butter of RAG really comes home.
11:12You made an interesting comment about the implication on not just performance, but relevancy as you scale vector databases. I'd love to hear you elaborate on that a little bit more. For a bit more context, I've shared here on the podcast and elsewhere, one of my observations is that it's really easy to get from kind of zero to a POC with RAG and with Dialog Agents based on LLMs, but getting from that POC to a system that you'd be willing to put in front of customers is like a lot harder. And one of the big elements of that gap is relevancy and all the details that go into like embeddings and constructing the context for the LLMs.
11:57But at the heart, like it's a relevancy challenge. And as a corollary to that, the folks that I've run across that have deep experience in search and getting search results tuned up seem to kind of get this and know how to fix this problem. I'm wondering, A, does that resonate with you? And then B, like talk a little bit more about relevance and kind of the way you see that playing out in vector databases. Yeah. So there's a couple of different pieces to it. And in fact, you're going to start to see a lot more of this just in general as people talk about. You have, you know, your precision, your recall, your accuracy.
12:35You've got ways of measuring those. You've got things like F1 that combine precision and recall because they're trade-offs. And so as we look at those, these fluctuate over time, and they fluctuate as a function of the data set that they're operating under. So generally, we go and we look at stuff, and most of your interactions, so it's probably worth sort of, there's two ways of looking at this. You can look at the vector data. Remember, most of the action around vector databases prior to this year was in using the vector databases as search engines. If you go back in time, take a look at what, for example, like the names we all throw around today, when you're looking at essentially what were these companies doing?
13:17Folks like Chroma, folks like Weviate, all that. Like a year ago, they were primarily looking at being search engines. And the idea was that, and as you just pointed out from search concept, that a lot of this stuff was rooted in search. And the idea was that keyword search, which is what a lot of folks are most familiar with, has a bunch of limitations. It actually works well enough in most situations, but generally people want it to move from a keyword to a semantic approach. And so then that's where doing a vector search comes into play. And by the way, you can actually do really good recall with no LLM really in the loop.
13:56It's more of an LM rather than LLM. Nowadays, we just call it an embedding model that is just for the purpose of going and reducing your text input into a vector that I can now go and do a similar search. So now we get into this stuff and we're putting in much larger data sets and we want to go and try to figure out how do we go and how do we measure, for example, how accurate are these results? Basically, are we getting a lot of false positives when I go and do a search? Am I getting things back? And between accuracy and precision, we narrow down sort of we get a good sense of sort of we're able to calibrate the sort of the false positives.
14:34And then we also get into sort of the recall pieces of it, which translate down into like things that should have come back. So it's actually sort of more of the false negatives. And again, this is very overly reductive, but we start to get into that. Again, you can measure these things and you can measure them. And then again, I also mentioned F1, which combines basically you have a trade-off from these and from recall to be able to see. So people have created other measurements around these. And again, this is, if you go and look sort of over the next six months, just for context, This is where a lot of action is going to be.
15:05You're going to go to, if you go to any one of these, it's not there yet, but I can predict within a month or two. There you go. There's my prediction for the new year, which is when you go to a vector database homepage, what you're going to see is for any of these open source or commercial, you're going to see all these great graphs that show number of documents and what precision recall F1 score and all of that are. Because that's where, like when you get past, you know, you mentioned sort of the demos, It's like past sort of going and loading a bunch of this stuff in. That's where the action becomes.
15:35Now, that said, apart from just the raw database capability of doing these results, there's a big piece around whether actually a big data prep piece that goes directly into it. And that's one of the things because you'll hear people say, oh, but real world this and on there. Right. What they're saying is like, it is garbage in, garbage out. And so that's where we get into things like chunking and so on, where I go and I take this piece of content and I need to break it into a set of things, not just from the fact that you want to get the appropriate piece from a storage standpoint, but more importantly, when I chunk these things, we take these documents, depending on how I chunk them, I could lose.
16:17And chunking is exactly as the name implies. Taking this, one might imagine, a large PDF and breaking it into bite-sized pieces. I can lose a lot of context that the smartest fellow alum in the world is not going to be able to put back if I happen to break things apart, you know, if I did it naively or whether I did it with a better understanding of the structure of the document. So chunking itself, actually, in terms of putting the data in is where you see a lot of work. And if you, for example, you know, we look at some of the frameworks people use for the stuff, Langchain and LOM Index, they put a lot of effort into going at the ingestion stage.
16:51And again, that's part of the reason why Lama Index is named what it is. And actually, as an aside, sort of you get to the differences between those two projects and the philosophies behind them. Langchain, as the name would imply, is about chaining LLM invocations. Lama Index does that. But in order to do that, Langchain does do a bunch of ingestion stuff. Lama Index, as the name implies, was really about getting that content and building those indexes. Of course, they also do orchestration. But you see when you go and look at the folks doing those projects and, you know, whether you're talking about Harrison or Jerry, like you can sort of tell they both had sort of a difference problem they were trying to solve.
17:27And again, when we sit down, we have to think when to build like a rag app, we got to think about the intent. They both become big problems. And depending on when you want to go and improve the accuracy of your results from the vector database, some people when correctly will go and say, oh, focus on getting the data in there, because that's going to have a big impact on your scoring. And other folks come in and say, well, that's true, but a big other piece of it is kind of how we break apart and construct the context that we're using to get to generate the vectors that we want to look up on and then what post filtering we do.
18:01And the answer is both of those are very important. Do you think since you offer it up predictions, do you think we'll get to a point where this all happens automatically kind of in the infrastructure and the user just needs to kind of bring their documents and get great results? Or will there always be a degree of fine tuning and tweaking that needs to happen in order to get desirable results out of a RAG type system? It's a good question. And as we tease that apart, we're actually going to get into a bunch of the subjects that I know you wanted to talk about. So what's our holy grail? Our holy grail is that the process of getting my data and using it within a Gen AI conversational experience.
18:49And as we've seen, conversational doesn't mean chat anymore because we've got multimodal and we've got drawings and pictures and actually dynamically generating a graphical user interface is something that we can do now. But I think probably you and all of your listeners probably already know that, which is that Gen AI doesn't just mean chat anymore. It just happens to be the simplest way to prototype. The first really important piece is, and this is not self-evident to a lot of folks. In fact, actually a lot of folks in the AI research domain don't fully get this, which is that real world AI, particularly like business use cases, needs you to be able to bring your own data.
19:24And that data, and again, this is really important, is not static corpus of data. What people are trying to do, and this is why people are using RAG, you're talking about data that is live data that is changing, that is oftentimes maybe confidential or proprietary, things like electronic medical records or people's financial statements and stuff. That is never going to get fine-tuned into a model. It is never going into the model, which means it's always going to go into infrastructure that's around the model. So once we do that, we're talking about RAG or some variant of RAG. If you looked earlier this year, there was this charmingly naive debate on RAG versus fine tuning.
20:06And seriously, you had really smart people getting into this debate and saying things like, oh, fine tuning, we won't need RAG as a fine tuning. And I was like, you're never going to train a model on people's electronic medical records or bank statements. If you want models leaking personal information, that's how you get models leaking personal information. right but that was it was the intersection of the fact that you had a lot of people were just very focused on building these really cool models and hadn't sort of zoomed out and said how are we going to use this stuff in more applied ways and so this year has been about the collision course of research and applied in the ai domain in a really exciting way but you see it playing out in like a day by day so we do know that but then the question becomes okay what do i have to do to get my data into a way that i can effectively retrieve it and that's where you've got the frameworks, and I mentioned a few of them, and by the way, there's a whole bunch more, but the kings at this point are Langchain and Lama Index, who have just been moving really, really fast on this.
21:02And you get a lot of people complaining about the code bases there, like, man, like, they're just, I'm like, yeah, but these guys are following the stack. You got to admire that, because that's what that's all about. But as a consequence, there is a lot of trial and error. So I see a lot of rag projects. And so going back to what you're saying, like, is it just going to be, somebody's just gonna make this dead simple. Well, yeah, ultimately there aren't a lot of people working on it. Right now, when you go and do a rag project, you spend a whole bunch of time. And it's not dissimilar, I think.
21:36So in that way, there's a lot of ways that an AI project this year, a Gen AI project is similar to the ML projects that you and I have been looking at and involved in in previous years, right? And then there's a lot of ways that they're completely different. And the ways that they're very similar is that for all the talk you're talking about, about like the cool stuff, like the day to day is really a lot of data engineering, which is a euphemism for just like data cleansing and a whole bunch of. And this was again, this is why people are using Python and such. Just a whole bunch of data munging just to get that data from the state that people have it in into a way that's going to yield the best results.
22:15So is that going to go away? Yeah, absolutely. Absolutely. And people are going to put a lot of wrappers around what currently requires you to do a lot of spaghetti architecture. And a lot of folks will, you know, and you'll see a whole bunch of companies that focus just on that because you've got a whole bunch of data in the systems, in the databases that people are already using. They're using streaming architectures, whether using things like Kafka or whether using their databases, relational databases, whether using non-relational databases. Like all that data needs to go into this stuff. And then you've got more importantly, like that's all your structured data.
22:51You've got a lot where AI starts to really blow things away is with unstructured data. And the unstructured data in most cases, you know, again, I spoke to this earlier, the hello world that for most RAG apps or RAG frameworks and most vector databases is chat with PDF. And the piece of that is, again, how do we take apart that PDF, which is at this point is now the canonical piece of unstructured data that everybody tests with? Not the most interesting one, but the most pervasive one. And then we go and say, OK, how do we take it apart? And then you start to get really into the world of mundane data at that point.
23:29It's like, is this an insurance claim? Is it a legal contract? Is it a research report? Is it a whatever? And every one of those has a set of heuristics involved that may be actually defined procedurally. Because, for example, in the case of a research report, or excuse me, a legal document, you don't actually need like AI to take it apart. Like every contract has exactly the same structure. So you can do a bunch of string munging and extract it and turn chunk it, which people do. There are entire companies and startups that do that. Or you can do things more cleverly, which is to go and you can have an ingestion loop where you're feeding this stuff into the LLM, having the LLM guide the ingestion flow.
24:08So you have the LLM supervising the dismantling of this unstructured content into chunks that then later at rag time, you're then going to go and be able to grab out. Right. So again, this is and I'm sure because you've got a lot of folks, a lot of your your listeners and stuff are probably in the middle of rag hell right now. And so they're probably like hopefully nodding like, yeah, no, I just did that this week. All that stuff has to go away. But part of the problem becomes it's like everybody's learning as we go along. And so like you don't want to prematurely optimize and automate for what the use case was last week.
24:44that was just the Hello World app when what we try to do next week is something more complicated. Because every time somebody brings in a new data set or a new type of data, it's like, okay, can we reuse the approach we use last time? Can we abstract on it? Can we build a new framework? And hence, there's a new BRAG framework every week, right? That's a fair and interesting point. Like you could over-optimize on PDF retrieval and get to anchor down into that and totally miss multimodal, for example, is what you're saying. Like there's a danger to premature optimization, which is kind of a truism in software, right?
25:22No, I mean, your example exactly was right. So again, most of what people are doing right now is smart knowledge bases, which by the way, not a bad thing. Like a lot of use cases, a lot of practical applications, a lot of happy users and a lot of businesses are gonna save a lot of money and make a lot of money just having AI knowledge bases. And for the knowledge bases, a lot of that stuff, the content that they're sourcing from are support contracts, things like that, or whatever documents in PDF, right? But like multimodal, the stuff you see from multimodal that people are building now, because we now have multimodal, obviously we have multimodal in GPT-4, but you actually have multimodal now, even in the open models.
26:04That's like the next phase of magic, right? To be a little hand wavy, but like that is also rag based or can be, but it has a completely different ingestion flow. It has actually two pieces of ingestion flow, because a lot of this stuff that we're talking about here is data prep time. But in multimodal, you have a real-time ingestion flow. You see all the examples where people, I'm drawing a blank right now, the name of the draw program that everybody's using to do the, you draw the thing and then you circle it. A sketch thing? Yeah, the sketch thing. Right. You know what I'm talking about. Yes.
26:34So the magic on all that, on that particular use case, is that they built a plugin for, it's basically an ingestion plugin, right it's like selecting the piece of the drawing area that it's then they've gone and wired in and sends it up to gpt4 right so that is a nice hello world example the minute i start applying that in to different use cases now i'm going to be like okay how do i capture is it something from my screen is it something from this app oh i'm taking something from you know so there's going to be a whole bunch of software engineering that goes into into that we've talked a lot about rag I guess, you know, a question for you that I've been grappling with a little bit is, you know, as a vector database is kind of this vector capability.
Read the full transcript
27:15Is it a feature or is it kind of a new platform? You know, I think when folks think about vector databases for better or for worse, they think about kind of some of these upstart companies. but there's PG vector for Postgres and data stacks now has a vector capability. And, you know, all of the traditional database vendors will have a vector capability. The question still is open for me. So obviously I'm at a vector database company. So I think about this multiple times a day. So my best thinking right now draws from a couple of things. So simple to the TLDR is it's going to be both, right? And the longer answer is you have a confluence of new things that create an opening for a new type of database.
28:00And so we go back 15 years ago, whenever it was, and you had people building with new languages, predominantly JavaScript. You had people using new types of APIs, predominantly REST APIs. So they're building new languages on the client. They're building new types of ways of moving the data. They were building a new runtime in the back end, people using things like Node.js, but other dynamic languages as well. And so what you had was you ended up having this data, JavaScript object notation, aka JSON. And so you had JSON on the client, you had JSON on the wire, you had JSON on the server, and then it made sense that you had JSON in the database.
28:36And that end to end created an opportunity. Now, Mongo was not the only JSON database, but there was a need. And so at least one pure JSON database then emerged and is now. And so Mongo is doing just fine, right? Like, like, but at the same time, you had JSON as a data type, like Postgres has wonderful JSON support now, and let's do most of the other databases. And there's other things. So the answer, if you were to go back 15 years ago, it was like, oh, is JSON, is it a feature or is it a new type of database? The answer was both. Were there 10 new JSON databases? Well, there were 10 new JSON databases.
29:11There's only one now that we remember. At the time, there were. Exactly, but there's one we remember. So of this current batch, one or two of them is going to go and be the Mongo of this age. I like this answer. Yes, that makes a lot of sense. But at the same time, it is also a feature. Like a bunch of other people are going to add these things, and it's going to be important. And everyone's going to bring to it kind of their special sauce. Your special sauce is kind of horizontal scaling. Someone else's special sauce might be this underlying document orientation. Someone else's special sauce would be something else.
29:43I think the real important piece, and this is the, it's like my refrain these days, I'm more concerned with, and for us and where I look at the people who are doing this right, it's like follow the stack, follow what people are building, because that's the important piece. Like going back to the Mongo example, the important part wasn't JSON as a data type. important part that Mongo did was it was JSON as queries, like going able to things that the other databases, like you go into Postgres and yeah, they added it as a JSON type and so on. And they've improved it over time. But the thing that Mongo did better than anybody else wasn't just that they could store and retrieve JSON.
30:20It was like they treated it as a first class citizens when you did a query, right? So it's the same thing. So again, from a product strategy standpoint, I go and look at it and say, okay, what is the equivalent of that now? Like, it's not just a question of, So everybody right now has gone and said, okay, I've added vector as an indexed column. So that's great. Very good starting point. But remember, all of the vector databases were out there, all of them, whether the ones that everyone associates as the pure plays or whether it's people who have added the capability, this all happened before RAG became a thing, right?
30:53And so now the question becomes like, which of these are like, follow, that's why again, I go back to say, follow the stack, follow the application, right? Like, are you adding features that are designed to make RAG better, right? And what's involved in that? And we've spent a bunch of time talking about this stuff, and we can go into even more detail. But that's the other piece, going back to, like, again, if I had my prediction of, like, if you go to a vector database website, basically, it's already true. Like, again, that website's going to have two columns to it. One is going to be, here's our recall stats, because that's your new stat.
31:27Databases used to talk about, like, oh, I can handle this may request a second. Now they're going to be like, this may request a second with this level of precision and recall. But the other hand, on the other side of that web page, it's going to be all about RAG, right? Because that's your canonical use case. And how do they make it uniquely easier? That's where all the innovation is going to be. And that's where you're going to see the difference. The databases that treated us a feature, if you as a developer, if I sit down to write something, I'm going to go to the ones that I can tell that whether it's an open source or commercial, it doesn't matter.
31:57I'm going to look at it and be like, are these folks focused on trying to make this easier for me to have a RAG application? So, yeah, we'll see. So long and short of it is, yeah, I mean, it is both a feature, but you'll have one or two folks that knock it out of the park and build a business on it. And that's always the case when we see this stuff, right? And is there an infrastructure element or clearly there's an infrastructure element to this? I guess more specifically, I was having the same conversation with someone and they mentioned, they seem to suggest that there were vector databases that were kind of GPU native or GPU enabled and take advantage of the GPU.
32:35And there are others that the implication was that they were not pure plays or something and they weren't able to take advantage of the GPU. Is that something that you're seeing? So when we look at the process of retrieving data from a vector database, your goal is not to have to do a whole bunch of GPU-dependent vector comparisons, first of all. So that comes down to the efficiency of how you built your index and so on. But anybody who's going and hitting the GPU in an unbounded way from a vector retrieval standpoint... At query time? At query time is not going to be... You could make the argument like, oh, that's a good thing to do.
33:13but it's not, you know, in terms of being more GPU native. And certainly the GPU vendors are very excited about that. So at query time, you are going to hit the embedding model. Your goal is to hit that once, not on a per row basis. When you take apart that query. Meaning you embed the query, turn the query into a vector. Yeah, it gets a little more complicated than that, but yes. So at their basic level, I'm going to take my input and I want to run it through an embedding model, right? and it's going to generate some germatic. My vector comparisons, my vector reversal, you can involve the GPU in that.
33:48Like I said, you're going to put yourself in a cost prohibitive situation. And one of the key pieces of the other key metric is cost. A lot of these things that work really well on your laptop, you price yourself out in terms of going into production from because it just costs you more money than, you know. You mentioned that your customers have thousands of nodes. If all of those have to have GPUs, that's another class of of infrastructure cost so we go back into that so yes you do hit the embedding model and as you do that by the way that becomes a big big selection problem and it directly goes into overall like generally you want your your embedding model is is generally partnered or or derived from or optimized from the main model that you're going to be using at generation time it doesn't have to be but because what happens is again you know and i know you've talked a lot about RAG and LLM chaining.
34:38But the reason, again, why we have things called Lank chain is because typically I go and I ask something from the agent and what the embedding model does is it breaks down my request. And one of the things that breaks it down into is a set of vectors of things that I want to know more about. And it farms that out into a set of queries, right? So I get maybe five, maybe I get 15 vectors and I do the lookups of that, all the things that my generation model should know about when it produces the answer. And so when I do that, that first model does not have to pair with the second model. Oftentimes, again, when you're doing things like OpenAI, you're going to use the same embedding model that is vector compatible.
35:15But you don't have to do that because you just retreat everything and then feed it textually into that. By the way, depending on what you're doing, like that's your simplest sort of rag model has two LLM invocations. But the reason why people call it, again, Langchain, where the name comes from, is I get multiple LLM iterations with branching. And you get into all sorts of things like chain of thought and so on for doing very complex answers or instruction following where you can get into some really cool stuff. But one of the questions that you see, like the idea, do I actually need it in the database retrieval loop?
35:48No, I don't want a GPU there. But the bigger question becomes, can database invoke the embedding LLM directly or do I have to do it in my application tier? That's more of a convenience thing, but that's an important. And so you do see at this point, most of the vector databases offer that as a feature, as a capability. You know, we do as well. Meaning they'll take the text as opposed to the vector and do the embedding for you. Yeah. Again, this will be a big deal next year is a lot of the databases are going to offer you a natural language query capability because it turns out that these models actually do a very good job of text to SQL, text to query language types of generation.
36:31So we are going to see that as well, which is going to be really interesting when that happens because it's going to blur the lines between, for example, NoSQL databases and SQL databases. A lot of effort goes into creating your queries and like I said, a lot of effort. It's actually a lot of foundation models are putting a lot of effort into it. They already do a very good job because they all, of course, use like, whether Google's using its crawl data set or everybody else that's using the common crawl, there's a lot of SQL priors on the web. So these models are already very good. I mean, you can actually use Llama 2 and get, for that matter, I mean, you go to Code Llama, but just even just, you know, regular Llama 2 will give you very good SQL, which is an interesting thing to see.
37:11When we were talking about embedding text, you mentioned it's more complicated than that. What was underneath that comment? Let's talk about it from the ingestion standpoint. Let's talk about the query piece of it. From the ingestion standpoint, we generate the embedding vector, the add inserts or upserts, right? And those are when we either create a new record or we update something. The embeddings that we create, again, a lot of stuff happens in the application tier with chunking, which is we're figuring out relevant piece. But we also get into which embedding model to use because the dimensionality of the vector is going to have a lot of issues from both the cost and performance standpoint.
37:50You see, for cost purposes, a lot of people end up wanting to go and use a smaller model that does, for example, maybe a 300 dimension, right? Because we know that the open AI, you know, open AI is, of course, the gold standard if I want to do it. If I've got unlimited money, I don't want to do it right. But it's going to give me a 1500 dimension vector, right? Each one of those dimensions is a floating point, right? So it's a big thing. And so ideally, maybe I'm going to use one of the small 300 dimension models off a hugging face. Problem is whatever I use at ingestion time is going to be the first or one of the first models that I invoke at query time.
38:25And so now I've got a trade-off because my first model that I hit is generally is taking your raw input where you're like, you know, where should I go for lunch? Whatever with all the additional injected context. but these smaller models, yes, they will give you that list of vectors to look up, but they may not necessarily be that smart in doing it. What happens is in that first pass, the goal is to generate, to build the context, right? Because everybody, users think in terms of prompts, but the LLM takes a context, which has your prompt, but all of the additional information that you choose to supplement it with, right?
39:01And so all of that then gets fed into the second model, which then again, typically restates your original question in depending on what you're trying to do. If you're trying to get to a zero hallucination type result, you may have a system prompt. So it's got the system prompt, which says you're going to get a question from the user and you've got the prompt underneath it, which is here's Sam's question. But the system prompt says you're going to answer this question for Sam, but you're only going to use the additional information I supply, right? So the context has system prompt, user prompt, and then a set of the rag retrievals.
39:37But the set of the rag retrievals are only as good as what that initial earlier model, what we call the embedding model, was able to retrieve. And if that embedding model is just not very smart, then particularly in that situation where I'm limiting my response to the stuff that was retrieved from the vector database, it might not be very good. And then further, again, we get into the chaining situations. Like you look at the stuff people are doing, and this is the stuff that people love, that particularly like the Langchain folks love showing off in their demos because it's really cool. Or for that matter, if you get into like the auto GPT stuff, which ends up being that with sort of an outer loop around it, then you're actually going in and you may be doing like several LLM, essentially input, prompt, generation the reduction, summarization, look up list of vectors, then feeding that in again, then feeding that again, then prompting back.
40:33Part of what it's doing is it's powering a conversation, not just a magic, give me the best answer. It comes back and says, hey, Sam, do you mean this, this, or this? And it gets to the point, again, through one of these chain of thought branching structures. It's like, let me go back to the user and ask him now, and then it goes and does it further. But the quality of each one of these steps, And by the way, we call these embedding models, but one of the things that's important at RAG, and I want you people to use a vector database, but the models can generate a lot of other types of lookups.
41:03Like you may also be like, give me a set of keywords for a conventional search. Like vector is not the be all end all. The important part of all of this is, this is all about iteratively building a smart context. And vector lookup is one of your best tools for building your context, but other forms of lookup as well. And again, this is where the intersection of the vector database and what the system prompts, and then the logic around it in the case of your lang chaining is where it goes into. Because part of it might be, you know, one might imagine a trip planning thing. And it's like, and by the way, give me a list of zip codes to look up.
41:41And that's perfectly valid. Like it's not a vector lookup. It's just a zip code lookup. And that's perfectly fine. Or give me a numeric range of prices based on the iteration, based on things that the model, and again, the model may be fine-tuned on concepts. Like we may have some concept of affordability that the model is able to opine on that says if the user meant they wanted something low cost, that they meant in this context between$5 and$25. In which case, again, that might be a lookup, a separate lookup from a product catalog or a restaurant pricing lookup. Like options on this are endless.
42:15Like you're going to have a lot of stuff that people are, and you already see this, like a lot of these apps that people are building are very sort of domain specific that are based around building domain specific context. And it would be great to imagine that somehow I can just throw more AI computing power. I mean, you can, but the model can't solve this on its own. This is a conversation between the model and the data, some of which is happening in the background while you're sitting there waiting for the response, right? So anyway, I don't know if I gave you too much information. Again, maybe what you need is, maybe my answers should be mediated by an LLM first to get them a little bit more concise.
42:54No, this is great stuff. And I appreciate additional context, I guess, and what you're seeing from a vector database and RAG perspective. It's fun because what it translates into is it is the intersection of a whole bunch of, I mean, again, there's a whole bunch of, you look at the things we talked about, there's a whole bunch of very mundane like data prep and data cleansing, data integration, a lot of stuff that, by the way, is not hugely a lot of fun, but it's still, it's a data engineering project. We have a whole bunch of stuff that's very model specific and one can lose themselves in the model domain because it's fascinating and so on.
43:33But then you just got a whole bunch of software engineering and architecture around this stuff. And the part that makes it hard is that when you start getting like each one of those, like, sure, there's half a million data engineers and data scientists in the world. You know, given that there's about 25 million developers in the world. So then we go and say the number of people understand software architecture pretty well and whatever. And you've got so who can architect around these things. And that's probably about 15 million developers. And then you have like number of actual like AI model data scientists, which is probably 100 ,000 to 200 ,000.
44:07Right. And then you've the problem is, is like the intersection of those. And you start to get down into like a very small number of them, the majority of whom are at hackathons in San Francisco. Right. And I think this is kind of the heart of my probably my first question, which is there's another element of this Venn diagram, which is like, you know, we talk about RAG like it's this new thing. But that retrieval, like this information retrieval, we've been studying this for decades. And there's a smaller number of people that are really experts at IR and have been working on search. And are we going to need that expertise for these systems to kind of fully meet their potential?
44:45Or are we going to be able to abstract that away? Or will the LLMs be able to do that for us so that, you know, I don't need to tweak my embeddings, my chunking and my context and my hierarchy? and all that stuff. I come back to that because it's a question that I've been thinking a lot about recently and asking a lot of people about trying to come up with some kind of forward-looking thoughts on. I mean, we've seen this plenty of times with everything new, like the first iterations of it. So we look at what happened over the last year, right? Like the first thing was you had this stuff and it was unoptimized.
45:20So when we look at Gen.AI, it was seriously unoptimized, but it worked. And you won't consider maybe it just barely worked, but it was unoptimized, which meant it was expensive as hell, right? And then we get to, say, mid-year, and quantization comes onto the scene, and it makes it possible to now actually go and run your models on consumer GPUs. And then shortly thereafter, you saw optimization coming in that allowed it to actually run without needing a GPU. You can actually run it. It's not very fast, but you can run this stuff on Intel. And then, by the way, we all love Python, but Python is like, or as a magnitude, slower.
45:55or this is not meant to be a performance language, right? So first thing you see is the optimization phase. The second phase is the important part, and I love the fact that you tie it back to like information retrieval and search, right? So the thing though is that is an extremely ubiquitous and mainstream use case. Like every company is going to need to take, every company generates information and knowledge and AI enabling it can't be like a, let me go and hire some AI research grad students. It has to be an off the shelf proposition. and it has to happen where that information sits. So yes, so that's going to happen too, by the way, right?
46:29So all of these things make sense. But I think we're right now, like at the point where it's, a lot of it is roll your own and it will be for probably another year. I mean, again, if you're like worried about this stuff, putting developers out of their jobs, it's like, no, you just need to start working with this stuff and learning it. There's going to be an awful lot of, an awful lot of coding that's going to be happening for a long time to come around. But as it gets easier, people are going to then do hard, like you just pointed out, Like people are just getting to the point where they can do text-based RAG in a fairly formulaic way.
47:00And now we got multimodal, right? Well, Ed, great conversation. Thanks so much for joining us and sharing a bit about kind of your take on RAG and VectorDBs. Awesome. This was a lot of fun. Thank you. Thanks. We'll look forward to next time. All right, everyone. That's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimla.ai.com. Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening and catch you next time.
From the publisher
Today we’re joined by Ed Anuff, chief product officer at DataStax. In our conversation, we discuss Ed’s insights on RAG, vector databases, embedding models, and more. We dig into the underpinnings of modern vector databases (like HNSW and DiskANN) that allow them to efficiently handle massive and unstructured data sets, and discuss how they help users serve up relevant results for RAG, AI assistants, and other use cases. We also discuss embedding models and their role in vector comparisons and database retrieval as well as the potential for GPU usage to enhance vector database performance.
The complete show notes for this episode can be found at twimlai.com/go/664.




