Vector databases (beyond the hype)

1 Aug 2023 · 51 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Practical AI - Episode on Vector Databases

Episode Overview Title: Vector databases (beyond the hype)

Guest

Prashanth Rao, Senior AI and Data Engineer at the Royal Bank of Canada Hosts: Daniel Whitenack and Chris Benson

In this episode, the hosts engage with Prashanth Rao to demystify vector databases, exploring their real-world applications, trade-offs, and how they fit into the broader landscape of data management and AI.

---

Key Concepts Discussed

What is a Vector Database?

  • Definition: A vector database is a specialized database for managing, storing, and retrieving high-dimensional vector data efficiently. Vectors represent compressed data, such as text or images, capturing semantic meaning.
  • Functionality: It retrieves similar vectors based on the semantics of a user’s query. This contrasts with traditional databases focused primarily on exact matches.

Importance of Semantics

  • Semantics vs. Keywords: Traditional search methods rely on keywords, while vector databases utilize semantic understanding, allowing for more accurate and meaningful results.

Historical Context of Databases

  1. SQL Databases: Originated in the '70s; structured data management using relational models.
  2. NoSQL Movement: Emerged in the mid-2000s to address the limitations of relational databases regarding flexibility and scalability.
  3. Vector Databases as Evolution: Vector databases represent an extension of NoSQL solutions, designed specifically for handling unstructured data and semantic searches.

---

Advantages and Trade-offs of Vector Databases

Key Trade-offs Identified by Prashanth Rao

  1. Purpose-built vs. Existing Databases:
  2. Purpose-built databases designed for vectors offer better performance and scalability than traditional systems like PostgreSQL or MongoDB, which have added vector functionality.
  1. Embedding Pipeline Options:
  2. Choosing between using existing models for embeddings versus integrated solutions provided by some database vendors.
  1. Indexing Speed vs. Querying Speed:
  2. Some databases prioritize fast indexing, while others focus on rapid querying. The balance depends on specific use cases and data flow requirements.
  1. Recall vs. Latency:
  2. Trade-off between the accuracy of results (recall) and the speed of retrieval (latency).
  1. In-memory vs. On-disk Indexing:
  2. In-memory solutions offer faster access but are limited by available memory. On-disk indexes provide scalability at the cost of increased latency.
  1. Sparse vs. Dense Vectors:
  2. The type of vectors used can impact performance and applicability based on the dataset and query characteristics.
  1. Hybrid Search:
  2. Combining full-text search capabilities with vector searches to enhance retrieval accuracy.
  1. Filtering Options:
  2. Deciding between pre-filtering data or post-filtering results based on the database’s capabilities.
  1. Self-hosted vs. Managed Services:
  2. Considerations regarding whether to use cloud-hosted solutions versus self-hosted options based on privacy, control, and scalability concerns.

---

Future Directions and Innovations

  • Prashanth expresses excitement regarding the impact of vector databases combined with large language models (LLMs) for applications like Retrieval-Augmented Generation (RAG), which enhances traditional search by integrating generation capabilities that provide more contextual responses.
  • He also highlights the potential synergy between graph databases and vector databases, aiming to leverage their strengths in knowledge retrieval and semantic querying.

---

Conclusion Prashanth Rao provides a deep dive into the nuances of vector databases, clarifying their significance in modern data systems and AI applications. The discussion showcases their transformative potential in semantic search and the growing landscape of database technologies.

Further Reading

  • Prashanth's blog series on vector databases:
  • [Part 1: What makes each one different?](https://thedataquarry.com/posts/vector-db-1/)
  • [Part 2: Understanding their internals](https://thedataquarry.com/posts/vector-db-2/)
  • [Part 3: Not all indexes are created equal](https://thedataquarry.com/posts/vector-db-3/)

---

This episode emphasizes the practical applications of vector databases and encourages practitioners to explore and understand the ongoing developments in this exciting field.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:06Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related technologies are changing the world, this is the show for you. Thank you to our partners at Fastly for shipping all of our pods super fast to wherever you listen. Check them out at Fastly.com. And to our friends at Fly, deploy your app servers and database close to your users. No ops required. Learn more at fly.io.

0:42Welcome to another episode of Practical AI. This is Daniel Whitenack. I am the founder of Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who is a tech strategist at Lockheed Martin. How are you doing, Chris? Doing well today, Daniel. How are you? I am doing so good because a lot of my dreams are coming true in terms of topics to talk about. I've been wanting to talk about vector databases on the show for quite some time. I know that we've mentioned them, but we haven't had a full episode on them. And I was scrolling through LinkedIn and saw a set of amazing posts and very practical posts about vector databases that I quickly shared and also sent a message to Prashanth Rao, who is a senior AI and data engineer at the Royal Bank of Canada.

1:36Welcome, Prashanth. Hi, thanks a lot for having me. Yeah, yeah. Well, now you have a three-part series on vector databases, a three-part blog series. What makes each one different? Understanding their internals and not all indices are created equal. I hope we can get into a bunch of that, but maybe to start out, could you just let us know what a vector database is? And in particular, why are people talking about them now? For sure. So I think the way I want to answer this question is I'd like to break it down into parts and answer each bit sequentially. Great. So to answer what a vector database is, let's start with what data is in the first place, right?

2:19The definition in my head is data is an organized collection of structured or semi-structured information, and it's stored digitally in a computer. Now, when you have data, you need somewhere to put it. So that brings us to the question, what is a database? So a database is a system that's built for easy access, management, and updating, and also querying the data at hand. We also need to talk about what vectors are. Vectors, you could call them a sort of compressed data representation that contains semantic information about any underlying entity. It could be text, images, audio, anything like that.

2:54So now we put all of these things together. What is a vector database? A vector database is a purpose-built database that efficiently manages, stores, and updates vectors at scale. I think the scalability is a very key factor there. And it also retrieves the most similar vectors to a given query in a way that considers the semantics of the query. So I think all of these terms holistically come together to form what we understand as a vector database. And when you say semantics, what do you mean in terms of semantics and how that maps onto a vector? So I'm sure everyone, most listeners are familiar with the concept of language models.

3:34NNNs are everywhere. So the thing with semantics is typically a query that you have. Like if you write a query to like a search bar on Google or something, you're thinking in terms of keywords. You're just thinking in terms of, okay, I want this particular thing, this item, whatever you're thinking about. and you type the word in there. Where semantics comes in is, did you type in something along the lines of what the data itself has? So can the query actually translate into something that the database actually understands and produces the result that is most meaningful to the query that you put in?

4:10So it's not just about the words or the features of that word, it's also about the meaning of that word and how that comes together in the underlying internal of the data. Cool. Well, so before we keep going, just because, you know, you have developers and data scientists, and they've worked with kind of all the other database types that most of us have worked with for decades. And we have multiple times over the years had to kind of like understand the new thing that's out and what the, you know, what the value is. No SQL. There you go. And so like, I'm going to jump on that. You know, we started with the SQL query language with these that are for relational databases.

4:47and then we went to NoSQL, which there are variants of and things called object databases. I understood your definition of vector, but I didn't understand how it related to the utility or lack thereof in some of those other approaches. Could you kind of lay the groundwork or the landscape of what that is? For sure. So yeah, I mean, I'm a total database junkie. I love thinking about the various kinds of databases out there. So actually, before we go into that, a quick summary in terms of where I'm coming from, right? So I started off as a data scientist. So I'm fully in your world, Daniel. And it's been a few years down that road for me.

5:25And I think for me, I've hit that point where I've been lost in the world of models and hyperparameter tuning and absolute data scientists will relate with that. But the more I began thinking about it, there are people who have entire PhDs in database theory, right? And their implementation. But then the more I've worked with data, I realized that you don't need a PhD to understand enough to build a working application built on top of a database. That's when you began thinking about what exactly are these different flavors that you have out there. I mean, of course, we've all come across SQL databases at some point in our careers if we worked in tech.

5:59So to answer that question, I think the general history of how these things panned out is quite interesting. I believe the origins of SQL databases come from way back in the 70s, I think, when this field called relational algebra was formalized. it's a kind of like you could say a formalization of the mathematics around what it means to join data query data store data in a database in a way that is queryable i think sql databases are so mature so tried and tested and the reason they've withstood the test of time is because they view the underlying storage or the underlying data as structured and in many cases you have structured data that is in the form of transactions and what a transaction basically means is some event happens in the real world and you log that information and you essentially build up a sequential chain of data which is basically a table and that's kind of what the relational data model came from and where relational models get interesting is you have tables that are related to other tables and that kind of maps into real world complexities where not all data is independent like some of the data depends on other things like a person's metadata could depend on what company they work at and things like that.

7:10So that's how relational data kind of became the norm. People were gathering data from digital systems and then putting them together. And SQL became the sort of standardized query language that you could use to query data. Fast forward to mid 2000s and the NoSQL movement starts to pick up. And where that comes from is there's a point beyond which relational data modeling can become a bit inflexible. Like it becomes a bit rigid because in the real world, you have data that comes in from various sources. Now, some of that data can come in very rapidly. With the advent of big data and streaming and all these rapid ways of gathering data that we have today, it became very obvious that the schema-based approach, a schema is basically what kind of data types exist in your table, right?

7:55So the way relational models were built was you needed to define a schema, and the schema kind of was the ground truth. The data has this data of this type only and that's what you expect in there all the time i think the no sql movement's sort of built on top of the limitations of the relational approach of being pre-decided by a schema because to be truly flexible in terms of the massive amounts of data coming in from various systems you need to have a schema-less approach at times and the schema-less approach basically means you store documents you dump data in semi-structured json blobs and things like that in a scalable way.

8:31And I think horizontally scalable became very, very important in that period. The earlier databases that were relational, I think they were more vertically scalable in the sense that you could just add more and more compute and you essentially scaled up your data that way. But now with NoSQL, the idea of distributing the data as documents across multiple machines and having those machines communicate with one another, that became a new paradigm. But I think the challenge with NoSQL is because of the underlying nature that the data need not necessarily be dependent on itself, like in the sense of relational tables, they didn't adhere to the SQL language standard.

9:08And they kind of diverged. MongoDB was among the first, and there were many others that came after it using JSON-based query languages. So there was a big bifurcation, I guess, in, you could say, the database community when on one hand you have SQL enthusiasts who swear by the declarative nature of SQL, and then you have the other community, NoSQL, who uses JSON, essentially, to query the database. They claim it was developer-friendly, and JSON is a developer-friendly interface, language agnostic, and so on. So in some ways, it does have its benefits, but then depending on your use case and depending on what you're trying to do, there are people who will argue on both parades that SQL should be the only thing you should use or no SQL should be the only thing you should use, and so on.

9:47So does that clarify aspects of both those camps before we move into the modern ones? It does, and then if you could distinguish as you go kind of how Vector is different from those others that would be helpful for me, I know, and maybe some other folks in the audience. Yeah. And I think maybe one thing that I loved about your blog posts is I see some of the players from the world that we just talked about represented within that landscape. And then also some that I'm not familiar with, or at least that I've seen only recently. And so you've got these different axes like Postgres, which is a SQL based query language to a relational database has some part to play in this vector database ecosystem.

10:34But then others which seem to have their own query language. So maybe you could also start to break down for us. So we want to store vectors in databases now to do these sort of semantic queries. Does that need to be soared in one or the other of these types of databases that you've talked about developed over time? Or how has that happened? And what are the sort of major categories of players in the vector database space? Absolutely. So I think, yeah, before we get into the specifics of databases, I think to answer Chris's point, we definitely do need to talk about the evolution, right? I see that vector databases are a natural evolution of the NoSQL class of databases.

11:14If you imagine a Venn diagram, you have like a circle that represents SQL and the other circle represents NoSQL. You have an intersection, that intersection point, I believe they're called NewSQL now. I'm not sure if you've come across that term. It's quite interesting, but NewSQL, they technically use SQL-like languages, but they also claim horizontal scalability and a bunch of other things related to asset compliance and all the other things. So it marries the benefits of both SQL and NoSQL paradigms. I was thinking initially, where do I place vector databases? does it go in that intersection or does it sit purely in the no sql camp then i imagine this as you extend that circle that has no sql it becomes like a blob like a fuzzy amorphous blob no sql is huge and i think in my head vector databases are like an extension to no sql and why they came about uh to understand what vectors are how they're stored in database i think it's important to understand what search is and what essentially you're doing when you query a no sql database So where it comes from is in the early days, I guess people were just submitting an exact query using a JSON sort of query language like how MongoDB has.

12:22And that query has to have all the terms or parameters in there that tell you what you want to fetch from the database, right? In a SQL world, it will be done with a declarative query in SQL, whereas in NoSQL, you typically do it in JSON. Over time, I think the idea of full text search became very important because I think everyone wants to be able to retrieve information from massive blobs of data fitting around. And how do you query that, right? If it's in a NoSQL sort of format, if you can't write a SQL query to retrieve it, how do you get that information? So the idea of a full text index came about.

12:55And what essentially that is, is it uses a concept of inverted indexes, inverted file indexes, sorry, where you consider the term frequencies of terms that appear in a certain document and obviously the relative frequency of how often those terms exist in a document versus the entire data set. So you combine all those things together, similar to how DFIDF is in data science. There's an algorithm called BN25, which is the most popular inverted file index algorithm. It's the most commonly used one for full text search. So the early days of search involved how do you scale that up because you have massive amounts of data, how do you build that index very, very efficiently?

13:32And then the querying interface sits on top of that so you essentially submit a query saying okay i want so and so term in the keyword that you put in and the inverted file index the bm25 algorithm it considers the words frequency and it considers subword features and a bunch of other things to intelligently retrieve relevant documents that contain that term while also throwing out you know useless words stop words and things like that so it was more of like a bag of words sort of which we consider an NLP analogy, it's kind of like a bag of words way of approaching text. Now, fast forward a few years, I think ever since the transformer revolution happened, people began observing the obvious power of transformers in encoding semantics, right?

14:15A transformer is way better at isolating meaningful terms in a document, especially when you're doing things like classification, retrieval, and so on. So how could you merge those benefits of a transformer with what you have in a database so i think vector databases the term got coined i think much later after transformers came about it was mostly called search engines before that a more generic term i think a catch-all term for anything that in word search but nowadays i believe search engine refers to a more like you consider semantics as a key component so essentially vectors are the only thing that can do that so to really describe what a vector is essentially you have a language model typically a transformer based language model that you use to embed the representation of a sentence into tokens and the representation is stored as a vector the vector that you have essentially for a particular sentence typically those are done using sentence transformers which is the most common kind of model you use that essentially embeds the entire semantics of that sentence in the vector and then the way this scales up is you consider the context of each and every token in that vector in a way that when you submit a query, the semantics of the query are mapped to the vector in your database, and you can find a similarity between what you entered as a query and what exists in the data.

15:36So a vector is a very powerful way of, you could say, compressing the representation of meaning in a sentence or a document in a way that scales up numerically, and you can rapidly query that in a digital.

15:56This is a Changelog Newsbreak. We've talked about prompt injection quite a bit since ChatGPT ushered in the LLM era. In brief, that's where you handcraft a prompt that tricks a chatbot into not following its own rules. Well, new research has uncovered some new LLM attacks on the block which aren't exactly that. Quote, large language models like ChatGPT, BARD, or CLOD undergo extensive fine-tuning to not produce harmful content in their responses to users' questions. Although several studies have demonstrated so-called jailbreaks, which are special queries that can still induce unintended responses, these require a substantial amount of manual effort to design and can often easily be patched by LLM providers.

16:45This work studies the safety of such models in a more systematic fashion. We demonstrate that it is in fact possible to automatically construct adversarial attacks on LLMs, specifically chosen sequences of characters that, when appended to a user query, will cause the system to obey user commands even if it produces harmful content. End quote. The biggest difference here is that they're achieving the jailbreak in an entirely automated fashion, and they make a case for the possibility that such behavior may never be fully patchable by LLM providers. Game over, man. It's game over. What are we going to do now?

17:24What are we going to do? You just heard one of our five top stories from Monday's Changelog News. Subscribe to the podcast to get all of the week's top stories and pop your email address in at changelog.com slash news to also receive our free companion email with even more developer news. worth your attention. Once again, that's changelog.com slash news.

17:52So Prashant, you kind of alluded to this, and I think that explanation was amazing of how this vector-based semantic search really exploded around the time that transformers and large language models did. I think even in this past, let's say, year, there's been this huge explosion of interest in vector databases. Could you maybe describe a little bit, so we know that you can search a vector database to find similar statements, let's say, or similar chunks of text where the similarity is based on semantics. How are people using them with regards to their AI workflows? And how does that kind of correspond to what's sort of popular right now in terms of what people are exploring with AI?

18:40So yeah, I think I need to highlight the fact that I'm both fascinated and frustrated by the current state of marketing in vector databases, both at the same time. I'm genuinely interested in the use cases, don't get me wrong. When combined with LLNs, large language models, like ChatGPT, you could say any sort of language model layered on top of a vector database can be used to build some very, very interesting applications. One of those interesting applications is querying your data via natural language. I think this has always been a dream of data scientists and people who work with data, right?

19:12Rather than writing my query by hand or constructing the query painstakingly from the ground up, can I just talk in natural language and have the database kind of respond to that query in natural language as well? The application we built using an NLM at the core, and essentially that would be powering the whole translation of human instruction to machine instruction and back to human. I could go into the details of specific applications, but one thing I do want to really throw back at you is, I know this is a practical AI podcast, right? So I guess what I was hoping to get into is I have an idea for a fourth blog post and in the series, basically.

19:50Part of it is the trade-offs, right? What really interests me about the various vector databases out there and why I began writing about these at length is when it comes to understanding what tool to use in the real world. When you have a business problem, when you have a particular case you're trying to address, obviously there's tons of information out there. You could go out and read a bunch of blogs and papers and come up with your trade-offs. But I think it makes sense to actually walk through some of these trade-offs. And my understanding is that as you go through these trade-offs, you actually begin formulating the value of these things much more clearly.

20:25And in my head, I think it makes sense to talk about the use cases once we go through some of these key trade-offs because in many ways using a tool depends on what goes into it and what you thought about the different options. You can dive right in because yeah I had follow-ups which were essentially what I think you're about to cover anyway so I'll just leave the mic with you man. For sure yeah so basically it makes a lot of sense to write about this and obviously read it at your own time but this is a great place for me to begin talking about it and eventually I'll put these down in words as well.

20:55So I've broken these down into I think roughly eight categories and the trade-offs. I'm specifically speaking about what do you need to think about when you're thinking about a database. And this will answer exactly what you talked about earlier, Daniel. So the first thing I think Daniel mentioned is the idea of deciding between existing databases that have been around, document format and things like that, versus newly designed databases specifically for vectors. So I'm going to call it purpose-built vendors versus incumbent or existing vendors. right so i think it's very important to understand in many cases you might just be looking to add semantic search capability or just retrieving information using semantics on top of an existing application right and that existing application could very well be built on a well-known tried and tested solution like elastic search postgres and so on there's many solutions out there and obviously in those cases it makes sense to just say hey but why can't i just leverage the vector index or the vector storage of that database itself like for example you mentioned postgres one real big concern with this is if you look at some of the material online on the performance of these the methods of pg vector pg vector is basically the vector plugin add-on to postgres and there's been enough documentation about this but it essentially the way it's been slapped on to postgres is as like an add-on it's like built by a third party called superbase and they add a vector functionality to the existing engine that Postgres has.

22:25So by its very nature, because it's not tightly integrated with the underlying internals of the database itself, like the storage layer, the indexing and all of that, you're going to miss out on a lot of optimization. Not you, but the technology is basically not optimized from the ground up to speed of indexing, performance during querying and so on. And this has been well documented. So that is a very big concern. depending on your use case and how much accuracy and what quality of results you want are you better off using an existing database that you already have in your stack or actually bringing on a new tried and tested purpose-built database for that very reason right and from my experience i've been tinkering around with quite a few options out there with purpose-built vendors in my opinion they are always a better solution in terms of scalability efficiency and also accessing the latest technology, like the latest algorithms out there, what indexing algorithms are out there, how do they get the best bang for buck in terms of your speed of indexing, the quality of query results, the latency of those results, and so on.

23:26So I feel like in the long term, if you actually are serious about building a vector search or large-scale information retrieval system that considers semantics, it makes far more sense to think about a purpose-built solution. Many, many database solutions are out there. I've listed some of those on my blog. And I think those are going to win out over the incumbent vendors who have kind of built, you know, vector offerings, as you can call them. What we're talking about is exactly what I had hoped we would talk about in this episode because your blog posts were so practical. In terms of how you think about the infrastructure that you work with day to day, would you recommend, because sometimes you don't know how much you need to optimize at the beginning and you can over optimize, right?

24:11So would it be a valid maybe stepping stone to say if I'm already working with Postgres, I could try out the vector capability of that. And if it works for my use case and I don't have, you know, three million documents that I'm searching over, maybe it's maybe it's fine. I just have a you know, I'm doing my personal blog or something and then kind of optimize as you hit a wall. Or is there danger in kind of trying to make that, put a square peg in a round hole sort of thing and get yourself in trouble? You hit the nail on the head. I was going to exactly say put a square peg in a round hole because I face those issues myself.

24:51I wouldn't name exact database vendors, but I work with SQL and NoSQL databases, which obviously have vector solutions. I think the challenge and the issue with saying that, okay, I already have something that works is you've got to remember that every single database that has existed for i think more than 10 years the databases come with baggage and they have their own tech debt that is associated with the underlying programming language they're built on there's years of decision making and architectural decisions under the hood that they've taken to implement solutions the way they have so they can't just throw all of that away and then build a vector solution that is optimized from day one right it's going to take a fair amount of time before these incumbent vendors are able to optimize their offerings to a point that perform as well as purpose-built vendors because these purpose-built vendors have spent thousands of manors, I guess, per offering in just tuning and building for a very specific goal.

25:44So what I've noticed in my experiments is that a lot of features that you take for granted in a purpose-built offering are not even available in the existing solution. Like PG Vector is a very, very young solution right now. Elasticsearch is vector offering. I've worked with that as well. it is also i mean considering elastichurch has been around for so long they only released their first vector like ann algorithm i think last year like 2022 so in terms of a database's uh capabilities that's very very young so i would say like there's a lot of things that you could potentially be missing or lacking and i'll cover some of those in my other trade-offs that i missed as we go forward yeah yeah let's go on to those i'm curious what number two is for sure the number two is uh i came across this in my first blog and reading some of the comments on them and one of them brought up this fact that the trade-off between using a database that allows you to build your own embedding pipeline or versus using a built-in hosted sort of embedding pipeline and by that i mean how do you generate these embeddings or these vectors right many people are familiar with sentence transformers it's available on hugging face and a bunch of other open source platforms so essentially it's quite easy or you could say it's trivial to put your data into these pipelines and generate sentence embeddings that you can just use to ingest into a database alongside your actual data.

27:03So you have your document data that has all the fields and attributes that you have in there alongside the vectors that encode the useful information in that that you want to query on. So that's a relatively trivial thing to do. But there are certain database vendors who offer convenience features on top of that where they embed the API of these models inside their own offering. So if you're just getting started and you don't know much about how vectors work or how, you know, NLMs work or any of these things, that might be something to consider. You might be better off using something like VV8, which has pipelines built in where you can just tell it, okay, connect to Hugging Face so-and-so model, and it will build the embeddings for you, as opposed to you writing your own custom transformer pipeline that actually takes in the vectors, generates the vectors, and so on.

27:48Now, if you have experience with transformer models, you might be far better off in doing all of the embedding work upstream, paralyzing, you know, and optimizing that portion, generating those at scale, and optimizing from a cost perspective, getting those done with the least resources and, you know, most quality that you can, and then just sending the vectors over to your database. So this is an important thing to consider depending on the level of experience that you, your developers have in your team, to actually bring the vectors in. Gotcha. That makes perfect sense. What are some of the other trade-offs?

28:22So then the other thing is the two key stages, right? You could say you could break down when you use a vector database as a developer. The first stage is the input, which is essentially building the index, right? I go into the indexing methods in a bit more detail. That's not really a trade-off. It's more about knowing what indexing even does under the hood. But what indexing means is you have data that you need to encode into a vector, right? Now, it's not as simple as just dumping a vector, which is like an array of numbers onto your database. You have to be able to search through those vectors.

28:56So the goal of indexing is to design efficient data structures and store the vectors using those efficient index data structures in a way that they can be queried efficiently and at scale. So that is an upstream process. You do that once upfront, you bring all your data in, it's indexed, and now you have a bunch of vectors in there that are searchable. the downstream portion of that is querying right it's basically like inference in nlp the query stage involves you taking the user input transforming that into a vector just like you did your raw data and the vector embedding that you use there is an embedding model that you use to basically transform your data so that they are compatible so that's a downstream step you're clearly separating the indexing step from the query step right so the trade-off here is is your database optimizing for indexing speed or query speed or is it mature enough that it has optimized for both and if you look through all the offerings out there many of the existing vendors have focused more on one end of the pipeline and not so much on the other some of them are faster at indexing and not so much at querying but some of them are way better at querying and much much slower during indexing so generating that index actually can be a very expensive step because it's not only about using a sentence embedding model or a transformer it's also about the database being able to translate those vectors into an index and it can actually query so depending on the size of your data this could take hours or even days like it's not unheard of here of indexing the uh periods of the order of days um and of course depending on your the amount of money you're throwing at it you could go use gpus to speed up the vectorization and and use multiple you know parallel instances of the database to scale that portion up.

Read the full transcript

30:37But that's exactly it. The trade-off here is how important is indexing speed? If your data is coming in in a stream at a very rapid rate, it's important to consider indexing speed as an important criteria. But then if you're not so interested in dumping large amounts of data very quickly, but more interested in serving results to a very large number of users asynchronously, then query speed becomes very, very important. I know we don't want to necessarily call out certain players in this space, but I think a lot of people are already familiar with a lot of the names here. So maybe if you could just highlight from your perspective, what are maybe some of the ones that are maybe more, like you were saying, mature in how they're thinking about both of those phases, whereas maybe certain ones that are optimizing more on one side or the other, which, like you said, depending on your use case is going to be a good thing.

31:28or it might be a bad thing. So it's really about use case. It's not so much about the goodness or how amazing a certain offering is, but more about use case. Yeah, absolutely. So as you say, I'm not going to call out specific. I mean, to be fair, everyone, every vendor makes trade-offs. They themselves are obviously juggling a lot of their own trade-offs when they build these things. But I obviously haven't used every single one out there. But the ones I have worked with, the most mature ones, I think Milvus is an open source purpose-built database. It's been around the longest, among the longest, I think, in the vector database market.

32:03It's extremely scalable. I mean, I like to think, I've written in my blog, I call it Milvus throws the kitchen sink and the refrigerator at the vector problem. So it can really handle billions of data points. I mean, it's designed for that. And obviously it has had time. It's been around for about four or five years. I wouldn't say that that would be my go-to first choice. That's my own personal preference, to be honest. It's more about, I guess, usability, how accessible their Python client is and so on. Then other vendors like VDN and Quadrant, I think they're also very, very optimized for this reason.

32:34So you could say that these are very, very powerful solutions. They scale really well, they ingest data really quickly, and they also supply query results very quickly and relatively accurately as well. To be fair, I think the existing database vendors like Elasticsearch, Postgres, they're not there yet in terms of the speed. And that's partly because they're general purpose databases. They're not specialized vector databases. So it makes sense that they have to deal with other priorities and they cannot optimize for all of these things with the laser focus that purpose-based vendors have. Thank you so much, Prashant, for helping us start to pick apart some of these trade-offs.

33:10And I'm starting to structure things in my mind in a useful way, which is really great because I've also been exploring a lot of these. And I agree with you. There's a lot of also new entrants into the field that show a lot of promise, even the ones that aren't quite as mature yet. What are some of the other, you mentioned eight. I think, I don't know if we've been through three or four yet. I wasn't keeping track. I might have to speed things up a bit. Just kind of list them off at least. Yeah. Okay. So maybe I'll quickly go through at a high level, then we can go into the finer details of which ones you think are the most interesting.

33:44So, okay. Let me summarize the first three. It's basically purpose built versus existing solution. That's number one. Number two was external embedding pipeline versus built-in hosting pipeline. Number three is indexing speed versus querying speed. So I think the others are going more into the actual indexes and generation of those indexes in more detail. So I'll go through them. Number four is recall versus latency. That's more related to how accurate are the results versus how fast am I retrieving those results. Number five is in-memory index versus on-disk index. I think this is a very big one for the future so we definitely want to go into i think some of the details of that number six is sparse versus dense vectors the kind of vectors themselves that are underlying the index number seven is the importance of hybrid search where it's full text search combined with vector search they both have their own trade-offs and i think the last one is the importance of filtering so pre-filtering versus post-filtering to decide the quality of your Yeah, I am very curious about this in memory or on disk one.

34:54Well, I'm interested in all of them. But I know one of the things that has come up in several of the applications that I've worked on has been, okay, do we self host one of these things? Do we use the managed service because they're going to be able to, you know, scale up and optimize things. There's also the choice of, oh, well, I could just load one in memory, you know, on the fly, ephemerally, right? I could have an embedded case where I load a bunch of vectors in and then there's some persistent file that I can pass around, right? And then there's, I think, more of what you are getting at, which is like, is this index represented on disk or in memory?

35:33Could you maybe help us parse through some of those things and go into a little bit more detail of what you mean there? So yeah, now that you mentioned self-hosted versus cloud, I think that's a number nine that I will add eventually. That's a very good point that you brought that up. Perfect, yeah. Maybe we can find a number 10 to round it out before the end of the episode. I'm sure there's way more, yeah. I could go on all day, but yeah. So yeah, going back to your in-memory, I think it's a very important one. So I think this is one of the things that is defining the, you would call it the race towards vector supremacy.

36:04I don't think the term is very accurate, but anyway, I think the challenge with most of the vector indexes out there. I think the most popular one by far is called HNSW, Hierarchical Navigable Small World Graphs. And I go into the details of the algorithm in part three of my blog. So I'd be happy to, I mean, discuss more with anyone else outside of this if required. But HNSW index is known for its relatively good trade-off between recall and latency. It's fast and it's relatively accurate, but it is also memory hungry. And where this becomes an issue is as data sets get larger and larger and larger, this is called the trillion scale vector problem now.

36:41A lot of vendors are talking about it. It's not too far away to imagine that you're going to have to, at one day, at one point, index a trillion vectors. And that is by no means a mean feat, right? It's a very challenging problem. So the data set in that situation would be way too large to fit in memory. Now, HNSW already does a lot of optimizations under the hood. The algorithm is designed to store a sparse graph in memory. and essentially you search through the sparse graph and then through the layers of that graph, you narrow down on the nearest neighbor to the query that you input. But as we go and get larger and larger in terms of data, even that sparse graph does not fit in memory.

37:18So databases have come up with different solutions as to how to deal with this out-of-memory issue. One example would be Quadrant. They use this thing called MemMap. It's like a sort of static RAM option where you don't actually store the vectors in memory, but you persist it to the page cache and it's still better than directly loading storing it on solid state drive which is one level below right so in terms of latency hit it's not as bad so you don't lose that much performance and you'll notice that a lot of vendors fight really hard to avoid persisting any vectors to or the index to disk because the moment you go onto a solid state drive there is a massive performance hit in terms of retrieval because the speed at which you're able to retrieve things from memory is, as you know, much, much, much faster than what you could do on disk.

38:03That's a general trend, I think, across the board right now. Most vendors are largely working with storing the HNSW index in memory and then adding some sort of caching layer to avoid having to repeat the queries and waste time in that sense. But this is an entirely new index called Vamana. So I've written about that in that log. It's optimized for solid state disk retrievals. And the algorithm they use is called disk ann not every database vendor has implemented this it's still in the early days but if i look at where the future is going there are many options that vendors could go down the road of right they could choose to implement hnsw on disk but record suggests that that's not a great idea because its performance would drastically reduce it would not perform that quickly as it does disk ann seems to be the agreed upon standard across many vendors but the challenge of this KNN is the original research paper that implemented it, the Microsoft team that implemented it, their implementation does not directly translate into the database internals.

38:59Depending on the language that the database uses, many of these are written in Go or this KNN implementation was written in C++. So it's not a direct transplant of the algorithm from the source to the database. It required a lot of rewrites and a custom approach towards optimizing for that speed. But that being said, I have to point out one particular vendor that I really stands out from everyone else on this strand they're called lance dd they're i believe the youngest database out there did they just come about i think end of 2022 early 2023 and they are the only solution as far as i know who only support on disk indexes they don't do an in-memory index at all and i was initially very surprised as to how they even do this how can they go about this but as i dug into it and i've spoken to some of the team as well they're really really open about their you know research that they're doing and all the models that they're building but essentially they innovate on multiple fronts but the biggest innovation is the underlying storage layer storage format they built this format called lance which is essentially optimized for on-disk reads of data and the database itself is built on top of this open source format lance so the whole thing is open source it's built in rust so the performance there is already you know close to bare metal it's really really fast they have already built an experimental disk in an implementation.

40:16So when it comes to this on-disk versus in-memory trade-offs, Quadrant is going about their own path in terms of how they achieve on-disk larger-than-memory data. VVH is going around its own path. LansDB is innovating on a different front. I feel like these are the three vendors who I've interacted with more and used. And I think the future is heading towards one where on-disk becomes a requirement and a standard way of implementing an index. But the engineering each added is instead on green. Let me ask you a slightly different question. It's not completely unrelated. The things that you've been, you know, kind of addressing there kind of are leading me to the next step on that.

40:52So when you're thinking about kind of environments that you want to put in, like, I know if you look at the other database types before Vector, you would have some that are, you know, scaled massively in the cloud, you'd have others as we've moved more and more intelligence and data out onto edge devices and they're either embedded or they're designed to serve in a very constrained environment. What are the options for vector databases in that? I'm assuming that there's obviously the cloud capability because that's kind of always the baseline. Do you also have, you know, as we're moving into an increasingly autonomous world out there and more and more things are being pushed out outside of the data centers in the clouds or at least the central parts of the clouds, are there options for either embedded or micro-serving, if you will, on the vector side?

41:44That's an amazing point, yeah. And I covered this in my blog post number one in terms of the architectures of these databases. And you're absolutely right. I think there is a lot of room for embedded databases to become the norm. I know DuckDB is making waves in the SQL market on this front. I think a lot of vendors are annulating what DuckDB has done in SQL. as you know DuckDB is an embedded database unlike Postgres which is a client server architecture data is right so what happened in the SQL world is now translating into you could say the vector world two databases that are following this embedded approach LanceDB as I mentioned and ChromaDB these are the most ChromaDB is quite well known people have been talking about it for a while but between the two of these I do think that LanceDB has it stands out more in the underlying technology because the Chroma from what I understand right now is it's still building out its underlying layer.

42:32It was kind of wrapped around an existing underlying internal database itself. It did have its own purpose-built offering to begin with, but they're kind of building that out as they speak. So I think between these two vendors, it'll be interesting to see how, you know, each of them rolls out their own features, a kind of target-specific part of the market. Going back to the point of cloud versus on-prem, that's another big thing I think that's going to come up. Honestly, like Pinecone and services like that that are completely on cloud, there could be real potential bottlenecks for companies to be okay with just you know sending out their data to some cloud even if pinecone says they would deploy on your infrastructure at the end of the day it is still a purely cloud-based solution there's a lot of infrastructure related hurdles around that self-hosted is i think as you see going to become more and more common and certain options like vv8 quadrant they offer self-hosted options in their licensing as well so the question for me that remains unanswered is which model in vectors database vector search will dominate in the longer term embedded versus client server right we are so used to the model of client server that's been working for more than a decade right now uh pretty much every database we've used is based on the client server architecture where the server sits remotely and i don't have to have the server running anywhere near where applications running but i think embedded databases especially with llns in the picture it makes a lot of sense in terms of data privacy and things like that and the scalability of these have i guess not truly been tested wd is just three or four years old advanced db is less than a year old crova as well so it'll be interesting to see how embedded databases compete on on that front like how well adopted they are because i think industry generally tends to favor things that are tried and tested at scale this sort of thing to catch on it would have to offer real real business value and the way these these databases monetize their offering i think that's going to be interesting to see and i guess we've already started moving this direction a little bit, but as we draw closer to an end here, I'm curious, you have explored probably more than, well, many people, certainly myself, in terms of how all of these offerings compare, what the trade-offs are related to vector databases.

44:42I'm curious, as you look towards the future, what are you excited to try that you haven't yet tried? And then maybe what excites you about this space? I know you mentioned, certainly there's things that are hyped or maybe different marketing that plays into this, but what are you actually practically excited about as a practitioner in the future of this vector database space? I think the low hanging fruit is the immediately obvious one. So I'll start with that. I think in the past, when it came to search, we imagine the Google search bar and idea that to build something like that was inconceivable a few years ago.

45:20Having a scalable, reliable search engine that you could build in-house on your own proprietary data was really, really difficult to do at scale. But today, I think with the combination of vector databases and LLMs, with GPT-4 now and all the other models out there, I really think that it's kind of become available to the masses, right? Like the average company who does not have massive compute is still able to build very, very valuable search solutions, information retrieval solutions on top of their existing data there are additional offerings like haystack and you know like the search engines that build on top of vector databases but i think the foundation layer are actually being enabled by vector databases which is why i'm so interested in in those use cases so those applications are very interesting at first the other thing is retrieval augmented generation this is a term that came about i think it was introduced by meta in one of the recent papers.

46:12Essentially, the idea behind retrieval augmented generation is typically information retrieval involved, you send a query and you receive a response that retrieves information relevant to your query, right? Where the generation comes in is now LLNs add an additional layer on top of that. You could send a query in natural language and you retrieve the most similar documents to that query, right? But rather than just retrieving the document itself, you could have the language model go through the document, look at your query, and then retrieve only the part of the document that is relevant to that query, and then generate a response that could potentially answer a question that you had.

46:52Like, what is the birthday of so-and-so person who runs this company, right? So these kinds of things were really, really, like, almost impossible to do before. But now I think it's really actually achievable with the kinds of tools and technologies that are available today. I think retrieval augmented generation is really skyrocketing right now as a term. I think everyone's talking about it. But what I want to add to that is I want to throw this out here to any of the listeners and maybe potentially I'm going to talk about this to other people in industry as well. Can we add another layer to retrieval augmented generation?

47:24And what I'm really interested in is how the two worlds of graph databases and vector databases come together. And I've posted about this a couple of times, but what's really interesting right now is most graph databases, like Neo4j, for example, they use declarative query language interfaces like Cypher. Cypher is, you could say, the SQL equivalent for graphs. The good thing about Knowledge Graph is they encode factual information and in a very human interpretable way. So the things that form nodes and edges in Knowledge Graph, they are something that we as humans put in there and encoded our knowledge of the real world into the data.

48:00Where vector databases sit complementary to this is, in many cases i might have connected data where let's say a person you know knows other person like a social network situation person follows another person person lives in a city and so on these are all meaningful connected entities in the real world but you add some layers of data on top of this uh you know data about a city you know data about a person you know where they worked you know what company information that has like there's so much additional unstructured data that attach onto the node in a grass that is actually hard to query using conventional grass algorithms or you know graph languages so i think vector databases are uniquely placed to add new value in that space in terms of i call it factual knowledge retrieval now the problem with knowledge retrieval is sometimes the queries that you have need to be exact the ability to submit a fuzzy query that does not exactly match your terms in the grass is something that you didn't have before.

48:58It's very difficult to actually generalize your query in a way that retrieves useful information. So I'm very interested to see how the power of natural language querying interfaces enabled by LNMs can be built on top of vector databases that store all the information related to an entity and then encode that entity into a knowledge graph. And then you tie all these things together in a way that you can actually retrieve information and explore and discover aspects about your data that you couldn't otherwise in a way that actually ties all these tools and technologies together. So I call it like an enhanced retrieval augmented generation sort of model.

49:36And this would obviously require tools like Langchain or Lama Index. I mean, these additional frameworks that allow you to compose these different tools together and pass data and instructions back and forth between the human and the different underlying databases themselves. So I'm super excited about those technologies. Yeah, I think it's great to hear that perspective. Also, you know, usually the answer is not like only this technology and nothing else is the solution, but a strategic combination of things often is where things end up. And I think those are really interesting topics to explore.

50:12And I look forward to your mini part blog post as you explore those. I'm definitely going to be following your writing now. And yeah, thank you from the community and from us for your work on this topic and sharing that work with the community. It's super practical and we're very privileged and happy to have you on the show. Thank you so much, Prashant. No, I appreciate it. Thank you so much.

50:44Thank you for listening to Practical AI. your next step is to subscribe now if you haven't already and if you're a long-time listener of the show help us reach more people by sharing practical ai with your friends and colleagues thanks once again to fastly and fly for partnering with us to bring you all change talk podcasts check out what they're up to at fastly.com and fly.io and to our beat freaking residents breakmaster cylinder for continuously cranking out the best beats in the biz that's all for now we'll talk to you again next time Thank you.

From the publisher

There’s so much talk (and hype) these days about vector databases. We thought it would be timely and practical to have someone on the show that has been hands on with the various options and actually tried to build applications leveraging vector search. Prashanth Rao is a real practitioner that has spent and huge amount of time exploring the expanding set of vector database offerings. After introducing vector database and giving us a mental model of how they fit in with other datastores, Prashanth digs into the trade offs as related to indices, hosting options, embedding vs. query optimization, and more.

Join the discussion

Changelog++ members save 3 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
  • Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs. 
  • Typesense – Lightning fast, globally distributed Search-as-a-Service that runs in memory. You literally can’t get any faster! 
  • Changelog News – A podcast+newsletter combo that’s brief, entertaining & always on-point. Subscribe today. 

Featuring:

Show Notes:

Vector databases blog posts from Prashanth:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Vector databases (beyond the hype)Practical AI · 51 min
Listen in VO