In short
The TWIML AI Podcast Episode #669: Building and Deploying Real-World RAG Applications with Ram Sriharsha
Episode Overview In this episode, host Sam Charrington converses with Ram Sriharsha, VP of Engineering at Pinecone, focusing on vector databases and retrieval augmented generation (RAG). Key topics include the integration of vector databases with large language models (LLMs), the complexities of deploying RAG applications, and Pinecone's new serverless offering.
Guest Background
- Ram Sriharsha:
- PhD in theoretical physics.
- Worked at Goldman Sachs (finance) and Yahoo (big data systems and machine learning).
- Experience at Databricks and Flink, focusing on machine learning research.
- Joined Pinecone to address overlapping problems in data retrieval and management.
Key Concepts
Vector Databases
- Definition: Databases designed to manage and retrieve data in the form of vectors (arrays of numbers), enabling semantic search rather than simple keyword matches.
- Relevance: Essential for effective information retrieval, particularly when working with large datasets and LLMs.
Retrieval Augmented Generation (RAG)
- Concept: Combines LLMs with vector databases to enhance the accuracy and relevance of generated outputs. LLMs serve as the intelligence layer while vector databases provide the knowledge layer.
- Importance: RAG systems improve knowledge-intensive tasks by ensuring access to relevant, accurate data.
Discussion Points
Evolution of Vector Databases
- Rise in interest due to the expanding capabilities of LLMs (e.g., ChatGPT) which now serve as orchestration layers for AI applications.
- Traditional information retrieval methods are evolving to incorporate dense retrieval methods using vectors, which enable better semantic search.
Trade-offs in RAG Implementation
- Challenges:
- Infrastructure complexity when scaling from demos to real-world applications (e.g., managing billions of vectors).
- Keeping indexes fresh as data evolves (adding, deleting, editing documents).
- Key Factors:
- The choice of embedding models and chunking strategies is crucial to the efficiency and effectiveness of RAG deployments.
Pinecone's Serverless Offering
- Introduction: A new vector database model allowing for on-demand data loading and flexible scaling without needing to pre-allocate resources. Currently in public preview.
- Benefits:
- Cost-effective query processing by separating storage from compute.
- Enhanced scalability for various use cases without the overhead of traditional database architectures.
Challenges in RAG Applications
- Infrastructure and Scalability: Managing performance and costs when transitioning from demo to production.
- Index Freshness: Dynamic updates to data require sophisticated solutions to keep retrieval systems accurate.
- Cost Management: Balancing the use of high-quality models with the associated costs of inference and data management.
- Quality Assurance: Addressing issues like LLM hallucinations and ensuring the accuracy of generated content.
Future of Vector Databases and RAG
- Anticipated advancements in:
- Streamlined workflows for embedding creation and document chunking.
- Improved integration between vector databases and LLMs to enhance overall user experience in knowledge-intensive applications.
Conclusion The conversation underlines the growing significance of vector databases in AI applications, particularly in conjunction with RAG. As these technologies evolve, frameworks like Pinecone's serverless offering aim to simplify the complexities of deploying practical RAG systems while optimizing cost and performance.
Additional Resources
- [Complete Show Notes for Episode #669](https://twimlai.com/go/669)
- Subscribe, rate, and review the TWIML AI Podcast on your favorite platform for more insights into machine learning and artificial intelligence.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:28All right, everyone. to be here. Thanks for coming on the show. I'm looking forward to chatting with you. We'll be talking about all things vector databases and retrieval augmented generation. Before we jump into that though, I'd love to have you share a little bit about your background and how you came to work in AI. Great question. So in another life, I did a PhD in theoretical physics. From there, I moved to different fields. So I kind of spent some time in Goldman working in finance. From there, I moved to the West Coast to join Yahoo. This was around 2010, and I stayed there until about 2014.
1:03That was my first exposure really to big data systems, large-scale data processing, and machine learning. At Yahoo, I worked in different areas around machine learning, eventually ended up in Yahoo Research, where I was focusing on scalable machine learning. From there, I left to, over time, join Databricks, where I spent some time working on Spark, but also starting new initiatives like Genomics and so on. From there, I went to Flunk to head the machine learning research group there. At some point of time in the course of doing all these things, I wanted to start my own company. And I was thinking about doing that when I reached out to Ido, who is the CEO of Pinecone.
1:40And we started talking and we realized that we are kind of trying to solve very similar overlapping problems. And that's kind of how I ended up teaming forces with Ido again and joining Pinecone. It strikes me that a lot of our audience probably has no idea how huge Yahoo was in terms of developing big data infrastructure and some of the early kind of commercial uses of machine learning in support of search and ads and so many things. Absolutely. I think some of the best work in cloud systems, online machine learning and so on, had kind of come out of Yahoo. People who worked at Yahoo who eventually went to a place like Google and so on.
2:22So yes, it's been certainly formative for the whole industry, but it's really the place where I learned a lot about basics that were needed to kind of work on big data systems, machine learning and so on. And we'll be talking, of course, about vector databases and why they have suddenly become so interesting. But you mentioned connecting with the CEO of Pinecone. Does this predate RAG and LLMs and all of that? Has Pinecone been working on vector databases for a while now? Yes. So Pinecone has been working on vector databases for several years now. I myself joined Pinecone about two years and six months or so.
2:59Now, while RAG and LLMs are kind of becoming very popular now, Now, the research and development on RAG and LLMs have been happening for a while. Language models have been getting bigger and bigger for a while. So it was only a matter of course that they would end up here. So we were aware of the LLMs. We were aware of the trends in language models. We were thinking already about how vector databases could be used in these sort of flows and so on. But the industry as a whole has started taking off really early in the last year around these topics. So elaborate on from your perspective, you know, when you think about how vector databases have suddenly come into the limelight and people who wouldn't have thought about deploying a vector database previously are now trying to figure out what this thing is and how do I use it?
3:49Like, tell us about the evolution of that from your perspective and the attention being placed on vector databases. What's the role of the vector databases and what folks are trying to do now? That's a great question. I think before talking about vector databases, it might help to set the context a little bit. So chat GPT and models like that are what we call large language models. These large language models are really just sequence to sequence models. So they take a sequence of text and produce a sequence of text, right? And now you can use this for a variety of tasks. You can use this for summarization, you can use this for search, you can use this for literally, you know, new use cases are coming up on a daily basis.
4:28What has made them very powerful is that these models are really large. They, in some sense, embed the world's knowledge in their parameters. So they are trained on a lot of data. They embed both the structure of language itself, as well as the ability to reason in their parameters. And then they can be used. What makes them really powerful is that now they can be used across tasks that they weren't specifically trained for. And this is what really made them very powerful. Now, the best way to think about large language models these days is that they are the intelligence layer or the orchestration layer for emerging generative AI applications.
5:07What I mean by that is you can use them to actually orchestrate intelligent AI flows. You can also use them to reason about things. So they act as the intelligence and the orchestration layer. But something is missing. If you just have an LLM, if you just have a large language model, you're missing what we call the knowledge layer. Okay. So what you're missing is the actual knowledge required to perform a lot of knowledge in intensive tasks. And that's really where vector databases come in. Now, large language models, to some extent, embed the world knowledge to some extent in their parameters.
5:39So it's not like they don't have any knowledge. They do have world knowledge, they do have the ability to reason based on that knowledge and so on. But what they lack is the ability to access accurate, relevant knowledge. And that's what vector databases provide. Now, to discuss how vector databases provide this, you kind of need to take a little bit of a step back and talk about the other piece of this entire workflow, which is retrieval. So if you talk about retrieving relevant, accurate knowledge, now this is something that people have been doing before. This used to be something we used to do even before there were large language models, right?
6:13If you think about search engines like Google, if you think about any sort of a search engine over text documents, for example, that's what it does. What it does is you take those documents, you put them in this sort of search engine, then you can ask questions about it. You can search over it. It does what's called retrieval. And the field of information retrieval is very rich, very well understood in the classic sense of retrieving text documents based on keyword matches. And this is something that's very, very well understood. But over the last maybe half a decade or so, there has been a new type of retrieval that's already been happening.
6:49And vector databases have already been used for that, which is what's called dense retrieval or retrieval using vectors. So instead of doing just keyword matches over whether your corpus contains keywords that match your query and how relevant are those keywords to the query itself, this used to be the traditional focus of information retrieval. What's happened recently over the last five, seven years is that people have realized that if you take those documents, if you embed them using these neural networks, what we call embedding models, and you put them into something like a vector database, now you can search over them much better.
7:24Meaning your searches give you much more relevant context, much better retrieval can be achieved doing this. And what do vector databases really do here? They act as the evolution of search engines, right? So what they do is, instead of searching over text by breaking up text into chunks of keywords and then creating posting lists and then somehow searching over them, this is what search engines used to do. In vector databases, retrieval works differently. What happens is that you take this same text, you chunk it up somehow. Each chunk is embedded with a neural network to produce an embedding.
7:56An embedding can be just thought of as an array of floats, an array of floating point numbers. And you take these embeddings and put them in a database, put them in a vector database. And vector databases are uniquely positioned to answer the following question, which is given a query, you encode the query as well. And now we have a query vector. And you can ask the vector database, give me the most relevant candidate document vectors corresponding to this query vector. And this particular interaction retrieves documents for you that are most relevant to the query and does what's called semantic search, semantic retrieve.
8:30So modern information retrieval already does this, which is modern information retrieval uses vector databases as a core of retrieving accurate, relevant knowledge this way. It's been my observation that as LLMs and interest in ChatGPT and the like have rapidly expanded the field of AI, that the connection between RAG as an approach to getting a chatbot to work and search and information retrieval is not really fully appreciated by folks. And the folks that I've seen have the most success in getting beyond a rag demo to something that's actually useful have a lot of experience, previous experience working on search and relevance types of problems.
9:18And all that knowledge comes into play in making a chatbot that a user is actually going to want to use because it's relevant and timely and provides the right information. Do you see some of the same things? Absolutely. In fact, all of this work that has gone into retrieval, relevance, ranking, and so on, is what is being used along with large language models to give them knowledge. So when we talked about it, we said large language models really don't have knowledge. The knowledge comes from retrieval. That's why you're finding that the people who are able to use retrieval augmented generation, RAG, workflows the best, are people who are really invested in doing the best retrieval.
9:55To put that another way, you can think of a chatbot that's based on RAG as executing a search on your documents based on a query, retrieving or pulling out the results that are most relevant to that query. And then the role of the LLM is just to summarize that and make it present it to the user. And so if you don't do a good job on that retrieval, you've got like a garbage in, garbage out type of problem and the chatbot's not going to be able to produce useful results. Absolutely. It's even more interesting than that, which is, remember, I told you that LLMs have been trained on large amounts of world knowledge.
10:31You could take the same world knowledge, put it into such a vector database. And if you do retrieval really well, you're making the LLMs better even on the knowledge that they were already trained on. Forget about the knowledge they had no access to, like your own documents or some corporate legal documentation or things that they would never have access to, right? So even on knowledge that they already have access to, retrieval actually provides better results. To maybe put that into context, what I'm hearing you say there is when you're using an LLM and retrieval, you're talking about two algorithms.
11:03One algorithm is a kind of sequence to sequence next token prediction algorithm. And the other is an embedding based information retrieval algorithm. And you're essentially saying that embedding based information retrieval is better than sequence to sequence at knowledge production? I think the best way to say it is that the two used together, the sequence-to-sequence models like LLMs used together with information retrieval through a vector database is strictly better than using LLMs by themselves or even fine-tuning LLMs through something, right? So purely just retraining LLMs on your data doesn't provide you as much of a value as using an LLM in the context with a retrieval engine like vector databases.
11:44So this is what's called RAG. RAG really is the combination of the two, and that turns out to be the best way to provide knowledge to these sort of knowledge intensive tasks. So I alluded to this earlier, you can go online and search out RAG demo and independent of your language or framework or database of choice, like there are tons of code walkthroughs and blog posts and things like that, that can help you create a demo of this. Like it's very easy to demonstrate, but that's just the kind of the beginning and delivering something that you would want to put in front of your users is a lot more complex.
12:21Can you talk about some of those complexities and some of the challenges that you see folks running into? Yeah, I think there's a multitude of complexities. Some of them are just on the infrastructure and kind of the scalability side, which is it's one thing to create a demo with a few hundred documents or a small number of documents and kind of do retrieval over that. Doing retrieval over billions of vectors is an entirely different problem, right? And that's kind of why some people like Pinecone spend most of their time building and perfecting vector databases because doing this at scale is very, very hard.
12:52The other thing is that outside of simple demos, your corpus evolves. So people are adding documents, deleting documents, editing them, redacting them, and so on. And being able to take the latest information and make it available for this sort of a RAG workflow is, again, a very hard problem. This is also where the database part of vector databases comes in. Because in this sort of flow, it's acting really like a database. In essence, you're synchronizing your enterprise data stores with this embedding store. Exactly. And you're expecting that every time you - And sync is historically really challenging.
13:27Very challenging, but particularly challenging for vector databases. I think traditional databases have gotten really, really good at doing this and kind of keeping their indexes fresh and so on. Vector database indexes are very, very challenging to keep fresh. In fact, very few people have cracked that problem. Why is that? I think the biggest problem is that it's easy to keep an index fresh if your index can be built incrementally. Okay. For vector databases, very few algorithms exist that actually build indexes incrementally. Because in some sense, when you put a document and encode it and put it into this sort of a database, you can't just incrementally add connections to the rest of the documents in some small region of space.
14:03Because if you do that, when a query comes, the query has no idea that there is a new document in this region of space it has to grow, right? So either the query has to scan everything for it to know that, or the query is going to miss the document out. And neither of those are good. If the query has to scan the entire corpus, that's actually computationally intractable. Whereas if a query is only looking at the places that it already has somehow indexed or pre-computed, it's not going to know that a new document was added here. So that process of keeping these sort of vector indexes fresh is fundamentally a hard problem.
14:34And this happens to be a hard problem for every known search algorithm out there. It's also where we spent the last good half of a year, actually close to a year, thinking about and solving and figuring out. Okay. And are these problems, like is this research frontier problems? Like you're implementing cutting edge algorithms, search algorithms, or are they engineering problems? This is equal parts, both research and engineering. So the way kind of typically it works, at least in Pinecone, but in general, I think also is that pretty much at the cutting edge, researching things and testing it out on lots of data sets, figuring out what works, what doesn't work.
15:10And then when we feel confident enough that we have something that works for the majority of the use cases and provides a lot of value, we start engineering and we start building. But at the same time, we are continuing to research. Something like keeping indexes fresh and so on will continue to be a research problem. There is a lot to do over the next few years to build world-class vector databases this way, but we know enough to be able to provide a lot of value today. So it's really striking that balance between research and engineering. So recapping kind of where we are with regards to challenges, there's these fundamental infrastructure challenges associated with scale.
15:42There are challenges associated with keeping your indexes fresh at scale. There's a couple of other challenges, by the way. There is challenges around just cost, right? A lot of generative AI workflows today, or if somebody is thinking about building a generative AI application, they have to worry about cost because a lot of these applications today are actually very expensive. solely because of hitting an expensive inference endpoint, or are there other reasons? It's part of it, right? For example, OpenAI endpoints, actually OpenAI has done a lot to reduce the cost per token, but even then it's still very expensive.
16:15And then you have these open source models, and I would call them today at least weaker models, which are cheaper, but they're weaker, right? So for knowledge-intensive tasks, people obviously want to use the best models out there. Here again, you can make weaker models more powerful by adding more context to them and by actually leveraging vector database and so on. So we spend a lot of time trying to figure out how do we make these cost prohibitive applications today 10 times cheaper so that you can unlock new use cases in a way that you couldn't before. And the last part I missed out in terms of the question you asked was, this is infrastructure and cost, but there is an even bigger question, which is quality.
16:53Now, historically, LLMs have struggled with hallucination. They've also struggled with things like attribution, right? So faithfulness. They're actually faithful to the documents that they have knowledge of, they're faithful to the context that you've provided to them, and so on. Measuring this and making sure that you are building RAG applications that are actually correct and giving you answers that you can be confident about, I think that's one of the big challenges as well. And even here, vector databases and RAG and so on can help significantly. So these are all I would consider equally important challenges.
17:21So, Ron, those are the high-level challenges that folks are running into. A couple of more in the weeds things that I hear about a lot are the choice of embedding model and the chunking strategy. And I guess there are a lot of these little things that someone has to, these decisions that people have to make as they build and deploy RAG applications. But what is your sense of how important these are? There seems to be some controversy out there as to whether these are key considerations or secondary considerations. That's a great question. First of all, there are key considerations. I think what's happening right now is that people are kind of reaching out for things that are easy at hand, right?
18:03Which is always the case when initially attendable applications that are trying to get off the ground. For example, embedding models that if you're using OpenAI for your LLM, kind of makes sense to use their embedding models and maybe the latest embedding models. That's what we see often, which is people are just using the things that are easy at hand, which makes absolute sense. However, the choice of the embedding model is very important. It's actually important in multiple ways. One, smaller, cheaper embedding models could actually do the same job, in which case you're kind of overpaying a lot.
18:32So it's important in that sense. In another sense, embedding models that are fine-tuned for specific things that you're trying to do can actually help a lot. So again, the choice of the embedding model is pretty important. I would say more important than that is actually chunking strategies. So how you go from text to vectors is probably one of the biggest parts of making the RAG workflow or improving the quality of the RAG workflow. So I would say there are two big parts to improving the quality of the RAG workflow, which is how do you go from documents to vectors, which is where checking strategies come in.
19:03And there's many ways to think about it, but fundamentally, yes, checking strategy is one of the most important things you could get right in this workflow. The other one is once you get relevant passages or relevant pieces of documents and so on, what do you feed to the language model itself? Do you just put it all together and feed it? Do you do something more on top of it? Do you re-rank? How do you re-rank? Basically, the second stage after retrieval becomes very important as well. So with all that in mind, from a challenges perspective, talk to us about what's happening in the vector database to accommodate these new styles of applications and the accordant challenges that have emerged that folks are trying to overcome.
19:45Absolutely. So first of all, we just released what we call Pinecone serverless, which basically unlocks an entire new set of use cases. So up until now, vector databases, unlike traditional databases, were very rigid. So if you go to the pod-based architecture in Pinecone today, you have to pre-specify the number of pods that you use and you have to pre-understand how much data you want to put in there. Every time you go beyond that, you have to kind of reshard and re, in some sense, reshuffle the data around. All of this is very expensive. And if you're pre-allocating it, you have to make it.
20:16So you've got to make these kind of fundamental infrastructure clustering type decisions before you even know like how far you're going with this app and what you need to, you know, what you're going to run into. Exactly. And if your use case changes, you have a serious problem, right? For example, if you want to take the same data and try to use it for an on-demand use case, now you're paying 1 ,100x more because you're only querying on demand. You're not querying all the time. So you're having these things sitting there and taking up memory and space and cost you don't need. So depending on your use cases, you would have a lot of inflexibility in the past.
20:53You would have had a lot of inflexibility in the past in how do you best leverage vector databases. Now, traditional databases have solved this problem over the course of the last 30 years. Vector databases are just catching up. And Pinecone took a very big step in this direction with Pinecone Serverless. The other thing that we are doing here is, remember I told you that you can actually improve LLMs on their own knowledge by putting that knowledge into a vector database. This again unlocks 10x cost reductions for generative AI applications and so on. But this again only becomes possible if the cost per query is very cheap.
21:25If you have to pay a lot to run a single query that uses a vector database. I'm only putting my most important data in there. Exactly. You're not going to put a billion vectors in there. So what we have done over the past year has worked on how do we rethink vector databases so that you can achieve 10 to 100x cost reductions across a variety of use cases. Secondly, you can make it very flexible. People can throw in their data and think about use cases later because we know that use cases are going to evolve. We know that people are going to find new use cases for this data and so on. You don't want them to think about all of that upfront.
21:57It's very inhibiting for generative AI workflows. That's what we've done. And to do this, we kind of had to reimagine the vector database itself. Traditional approaches don't work. And so what does that reimagining entail? Like how did a traditional monolithic architecture, well, I'm making some assumptions about the architecture, but assuming it was a traditional monolithic architecture, like how does that need to shift to become now serverless? That's a great question. So all vector databases prior to us releasing PintCode serverless, they operate on this, what's called the search engine architecture, right?
22:28So which is your data is split up into a bunch of shards. These shards contain a subset of your data indexed, always available, which means it's on these machines that between a combination of RAM and SSD, they're kind of always up and running. Now, this makes sense when you're running like tens of thousands of queries per second, which need to touch your entire corpus. If you're doing that, then yes, that architecture kind of makes sense. But if you're running queries on demand, or if your queries are not touching the entire corpus, which for a web-scale corpus, a single query is not going to touch your entire corpus, and you don't want it touching your whole corpus.
22:58For scenarios like that, now this becomes an extremely expensive work, and this is what customers of Pinecone were also finding because Pinecone was also using the same architecture, quad-based architecture. So the first thing you have to do to make these sort of workflows, the emerging generative AI workflows, as well as many other workflows that customers at Pinecone do, if you want to make them 10x cost-effective, you kind of have to decouple storage and compute. You cannot have all of the storage for the index sitting close to compute all the time. Now, this is something that traditional databases have started doing really well over the last several years.
23:31So if you're familiar with Snowflake, if you're familiar with systems like this, they're great and they've taken decades of database learnings and figured out how to separate storage and compute. So you can drive down costs for using those databases, right? What we needed to do was to rethink how to do that for vector search. Because remember, I told you that vector search has this fundamental problem that every time you're trying to update an index, you kind of have to have, in some sense, a global view of the index, at least in existing algorithms you have. Similarly, all existing algorithms that you're familiar with, whether it's HNSW, whether it's FIIs, whether it's practically anything that anyone does today, they hold the entire index between memory and local SSD.
Read the full transcript
24:09In fact, most algorithms up until two years back weren't even using disk. They were all memory-based. And memory is very expensive on the cloud even now. But even disk-based systems are actually very expensive because there is no way to page anything from a cheaper storage. You have to maintain everything on local SSDs. And on the cloud, you don't have the fungibility of SSDs being decoupled from compute or something. When you get a machine on Amazon or GCP, you're getting a lot of SSD, a lot of cores, lots of RAM. That's the whole box that you get. There is no fungibility even with the SSDs. So the best way to drive cost savings is to really decouple storage and have it in some cheap storage like blob storage.
24:47But the moment you put things in blob storage, you have to get very efficient at incrementally indexing that thing, incrementally pulling out only the parts of the index that queries care about. And this requires fundamental re-innovation in how vector search even works. That's kind of what we did. I would say that the core of the innovation is around re-imagining vector search in such a way that you can actually decouple storage and compute and page parts of the index on demand so that queries don't have to look at, basically the cost of the query is no longer the cost of the whole data under management.
25:17It's only proportional to the cost of the parts of the data that the query even needs to look at. And so presumably you're still doing this on the cloud, but you're just using different primitives than you were before. Maybe before you were using large machines with lots of SSD. Now you're using what? So it's not just the machines that are different. It's the algorithms themselves that are different. Before, what we used to do was we would take all of those, the entire index and preload them or even keep them fresh and keep them loaded onto these big machines. Today, what we can do is we can still have these big machines, but now they are only loading the parts of the index on demand.
25:54They're caching the parts that have been retrieved and found frequently used so that the frequently used parts of your data are actually cached and so you can have very low latencies and so on. But then you're not paying the cost for all the data that is not being touched. That's the core innovation. Got it. Yeah, when you mentioned blob storage, that made me think that data was being pushed to S3 now as opposed to online storage. Exactly. It is. But the trick is in figuring out what parts of the data do I need to keep close to my compute when the queries need to be answered. And so is this things like predictive query optimization and that kind of thing?
26:32So I think that what we do is basically when you create this, when you index or re-index, when you keep your indexes fresh and so on, you want a way of partitioning them, right? In traditional databases, you're familiar with range partitioning. you're familiar with different ways in which they can slice up ranges of columns so that you can say that if I'm getting a query for, let's say, timestamp greater than 100, I know that the range I need to look at is everything greater than 100. I don't have to even look at data that has timestamp less than 100, right? And you do this by creating these, breaking up data into chunks.
27:04Each chunk has some statistics around it. You then use that to figure out what do I even need to look at? And that general philosophy is called partitioning, right? It's the partitioning data. You want to have the same idea, except now you're partitioning these vectors. So these vectors are now in some space, and you want to kind of geometrically break down that space so that once you've broken down the space into chunks of regions of space, when a query comes, you can say that, actually, you know what? I'm only interested in these regions. I don't care about data that's in all the other regions.
27:35Go fetch me only the parts of the index that belong in those regions, and then I'm going to see what are the closest candidates that the query needs. So it is basically applying the same traditional partitioning ideas from databases, except for this new way of partitioning, which is geometry based, right? So traditional databases don't understand geometry, but vector databases need to understand geometry. Interesting. And so from the user perspective, I've got an existing Pinecone database. I've got documents in it. I've gone through the partitioning strategy that you talked about earlier, and I hear about this announcement.
28:14I get excited and I want to take advantage of lower costs, all the things that you mentioned. Does anything change for me? Do I flip a switch and all of a sudden I'm on the new serverless or do I have to reload my data? And even that, it's challenging at scale, but maybe worse is do I have to touch my code? Do I have to rewrite my application? What are the things that have to change from my perspective? That's a great question. So today we are in what's called public preview. So during the public preview phase, you do have to re-ingest your data if you want to use serverless. Between public preview and general availability, we're going to make it so that it's a flip of it.
28:53So basically sometime between now and when we are generally available, you shouldn't have to do anything to start taking advantage of serverless. But in this initial phase, you will have to re-ingest your data. There are no code changes involved for you, so it should be seamless from a code perspective to start using serverless. We also work very closely with some of our biggest customers who have billions of vectors, and for them, re-ingesting is not even an option. So for some of these customers who are trying out serverless, our field engineering team and engineers work together to figure out an option for them to do a seamless migration.
29:26So there are things that we are doing to help some of the biggest use cases, but in public preview, We expect people to re-ingest their data today if they want to use the virus. From the perspective of a developer or someone who's architecting one of these systems, it sounds like this changes a lot of the fundamental economics. And maybe I'm considering different use cases than I was previously, or I have fewer decisions that I have to make up front. Are there API changes that go along with this? I guess I'm trying to get at it. Are there new capabilities beyond the things that are transparent to me that as a developer, I might get excited about?
30:04So first of all, you're 100 % right. This unlocks some new use cases that people wouldn't be thinking about before, which is if you have data and you want to run on-demand queries and so on, you can now just start ingesting it because ingestion is cheap, storage is cheap, and you're only paying for what you query. This really unlocks an entirely new set of use cases. Alongside, we try not to break any APIs. We're trying to keep APIs as compatible as possible. But in serverless, you do get some extra information. For example, when you are using a serverless index, you're going to get information about, in some sense, how much did this query cost?
30:39So we have a proxy for this cost that we call read units. So every query returns to you what the cost of that query was. And now that's very useful for you to figure out, in some sense, to budget and to understand how much is a use case pattern going to cost you. And that's very important to know. The same serverless architecture allows us to, in some sense, put control over amount of your data that you need to actually look at for the quality of search that you need. What I'm trying to say is that it turns out most query sets and most experiments we've run and most benchmarks we've done tell us that about 95 % of your queries can be answered by looking at a small amount of data.
31:17The trick is, of course, figuring out what is that small amount of data to look at. And that's what serverless does really well. The 5 % of the use cases... Which is also a traditional database optimization type of a challenge, right? Exactly. Except it's a harder challenge for vector databases because you have to understand geometry really well to do this. It's geometry and it's global as opposed to much more local. Exactly. But there are always these 5 % of the queries that need to scan more. In our old pod-based architecture or in any other vector database today, every single query would be paying for the least common denominator.
31:51It would be as expensive for you to run a query that should have looked at a small amount of data as it is to run a query that should be looking at enough data to be able to answer a question. With serverless, you don't have to do that. So one of the things we are trying to do over the next, between now and general availability, is figuring out what is the best API that we can give that puts this control in the hands of users. I think this is going to be very powerful because now it allows you to be in control of just how much are you paying for a certain, what you consider good quality. Maybe taking a step back to where we started and kind of returning to RAG and the challenges that folks experience trying to get to, you know, a really production quality experience for user facing applications.
32:35Tie what you've done with serverless to that core challenge and talk a little bit about what you see coming down the pike that will make it easy because, you know, it's got to get easier. It's still very difficult. Absolutely. There's so much attention being paid to it. It will get easier. How do you see it getting easier? And how will the Vector database contribute to that? I think a lot of work that we've done over the last year has been to make the Vector database really simple to use, really economical at scale, and something that you can just throw your embeddings in and search over. And we have put a lot of effort into that.
33:13Pymecon serverless is like a huge step in that direction. And I expect that we continue to double down on that and spend a lot of time refining that workflow. Now that said, so I'm very confident that we'll do that really well and we do that really well. But there is the other part of it, which is creating embeddings themselves as being complicated. We talked about chunking strategies themselves as being very much of an art than a science right now. We talked about very little being done in the area of re-ranking and information retrieval itself to connect all of these together really well. I expect that we'll find ourselves spending a lot of time making that workflow seamless and easy for people.
33:48So in some sense, I expect vector databases to really get to start kind of gluing together the areas around vector databases that today are very disparate and bespoke, whether it is what embedding models you're using, how you're really chunking versus what you're retrieving. How does that kind of get into the context of an LLM and so on? I think this is where a lot of engineering and research is going to happen over the next year, whether it's SpineCone or anywhere else. Awesome. Oran, thanks so much for taking some time out of your busy day to share with us kind of the latest and greatest with regards to vector databases and RAG topics that, as I mentioned earlier, are getting a lot of conversation, a lot of airplay here and elsewhere.
34:28It's a great conversation. Thank you so much. And thanks for having me. Thank you. All right, everyone. That's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimla.ai.com. Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening and catch you next time.
From the publisher
Today we’re joined by Ram Sriharsha, VP of engineering at Pinecone. In our conversation, we dive into the topic of vector databases and retrieval augmented generation (RAG). We explore the trade-offs between relying solely on LLMs for retrieval tasks versus combining retrieval in vector databases and LLMs, the advantages and complexities of RAG with vector databases, the key considerations for building and deploying real-world RAG-based applications, and an in-depth look at Pinecone's new serverless offering. Currently in public preview, Pinecone Serverless is a vector database that enables on-demand data loading, flexible scaling, and cost-effective query processing. Ram discusses how the serverless paradigm impacts the vector database’s core architecture, key features, and other considerations. Lastly, Ram shares his perspective on the future of vector databases in helping enterprises deliver RAG systems.
The complete show notes for this episode can be found at twimlai.com/go/669.




