Vectoring in on Pinecone

10 Jul 2024 · 44 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Practical AI - Vectoring in on Pinecone

Episode Overview

  • Podcast Title: Practical AI
  • Episode Title: Vectoring in on Pinecone
  • Hosts: Daniel Whitenack (CEO of Prediction Guard) and Chris Benson (Principal AI Research Engineer at Lockheed Martin)
  • Guest: Roie Schwaber-Cohen (Developer Advocate at Pinecone)
  • Main Topic: Advantages of vector databases in machine learning, particularly discussing Pinecone’s offerings.

Key Discussions

Introduction to Vector Databases

  • What is a Vector Database?
  • A specialized database optimized to store and manage high-dimensional vector data.
  • Distinct from traditional databases as it focuses on the geometric relationships between data points.
  • Importance in AI:
  • Vector databases provide a crucial layer between structured data and the capabilities of large language models (LLMs).
  • They allow for efficient retrieval and management of data represented as vectors, enabling more effective machine learning pipelines.

Pinecone’s Role in Vector Databases

  • Founding Background:
  • Pinecone was founded by Ido Liberty, who has significant experience in AI and cloud services.
  • Pinecone’s early recognition of the importance of vector representations positioned it as a leader in the space.
  • Core Features of Pinecone:
  • Ability to handle vector data at scale while ensuring speed and resiliency.
  • Offers features such as metadata filtering and namespaces for efficient searches in multi-tenant environments.

Comparing Vector Databases to Traditional Databases

  • Differences:
  • Vector databases are designed specifically for high-dimensional vectors, while traditional databases (like relational and NoSQL) focus on structured data and textual indexing.
  • Vector databases leverage embeddings for semantic searches, allowing for retrieval based on conceptual relevance rather than just keyword matching.
  • Advantages of Vector Search:
  • Can achieve better search results based on semantic similarity (e.g., "king" is closer to "queen" than "dog").
  • Supports complex queries in natural language, allowing for a more intuitive user interaction.

Practical Use Cases for RAG (Retrieval-Augmented Generation)

  • Definition of RAG:
  • Combines LLMs with external data sources, improving the accuracy and reliability of generated text by grounding it in real data.
  • Implementation Challenges:
  • Organizations must evaluate their specific needs and data structures to implement RAG effectively.
  • Continuous monitoring and evaluation of the RAG system's outputs are crucial for maintaining quality.

Pinecone's New Developments

  • Serverless Architecture:
  • Introduced to decouple storage from compute, allowing for scalable and cost-effective solutions.
  • Users can store significantly more vectors at a lower cost while maintaining performance.
  • Knowledge Assistant Feature:
  • Simplifies the process of integrating various AI functionalities without requiring extensive setup or infrastructure.
  • Users can easily manage documents and queries through a streamlined interface.

Future of AI and Vector Databases

  • Emerging Trends:
  • A shift towards integrating traditional AI methods with LLMs, moving Beyond just relying on LLMs.
  • Recognition of the need for a variety of data management solutions (vector databases, graph databases, etc.) to solve different problems.
  • Vision for the Future:
  • The potential for LLMs to serve as interfaces that coordinate various AI tools, enhancing the overall capability of AI systems.

Conclusion

  • The episode highlights the transformative role vector databases like Pinecone play in the AI landscape, particularly in enhancing the capabilities of machine learning models through efficient data management and retrieval strategies. As organizations increasingly adopt RAG strategies, the integration of LLMs with vector databases is seen as a significant step forward in AI applications.

Key Takeaways

  • Vector databases are essential for managing high-dimensional data in AI.
  • Pinecone’s innovations, including a serverless architecture and knowledge assistant, enhance accessibility and usability.
  • Continuous monitoring and evaluation are crucial for successful AI implementations.
  • A multi-faceted approach combining various AI methodologies will lead to more robust and effective AI applications.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related tech is changing the world, this is the show for you. Thank you to our partners at Fly.io, the home of changelog.com. Fly transforms containers into micro VMs that run on their hardware in 30 plus regions on six continents. So you can launch your app near your users. Learn more at Fly.io.

0:42Welcome to another episode of Practical AI. This is Daniel Whitenack. I am the CEO and founder at Prediction Guard, where we're enabling AI accuracy at scale. And I'm joined, as always, by my co-host, Chris Benson, who is a principal AI research engineer at Lockheed Martin. How are you doing, Chris? Doing great today, Daniel. How's it going? I know we're recording leading into a holiday weekend here. We are. And so many exciting things. Last week, I got the chance to briefly attend the AI Engineer World's Fair, which is sort of prompted in certain ways by our friends over at the Latent Space podcast.

1:22And that was awesome to see. And of course, a big topic there was all things having to do with vector databases, RAG, all sorts of retrieval, search sorts of topics. And to dig into a little bit of that with us today, we have Rowie Schwaber-Cohen, who is a developer advocate at Pinecone. Welcome. Hi, guys. Thanks for having me today. I'm really excited to be on the show. Yeah. Well, I mean, we were talking a little bit before the show. So Pinecone is, from my perspective, one of the OGs out there in terms of coming to the vector search, semantic search embeddings type of stuff. Not that that concept wasn't there before Pinecone, but certainly when I started hearing about vector search and retrieval and these sorts of things, Pinecone was already a name that people were saying.

2:17So could you give us a little bit of background on Pinecone and kind of how it came about and what it is position-wise in terms of the AI stack? So Pinecone was started about four years ago, give or take. And our founder, Ido Liberty, was one of the people who were instrumental in founding SageMaker over at Amazon and had a lot of experience in his work at Yahoo. And I think that one of the fundamental kind of insights that he had was that the future of pulling insights out of data was going to be found not exclusively, but predominantly in our capability to construct vectors out of that data.

3:02And that representation that was produced by neural networks was very, very useful and was going to be useful moving forward. I think he had that insight way before tools like ChadGPT became popular. And so that really gave Pinecone a great edge at being kind of the first mover in this space. And we've seen the repercussions of that ever since. With the rise of LLMs, I think people very quickly came to recognize the limitations that LLMs may have. And it was clear that there needed to be a layer that sort of bridged the gap between the semantic world and the structured world in a way that would allow LLMs to rely on structured data, but also leverage their capabilities as they are.

3:52And that is one of the places where vector databases play a very strong role. You know, vector databases are distinct from vector indices in the sense that they are databases and not indices. So an index basically is limited by the memory capacity that the machine that it's running on allows it to have. Whereas vector databases behave in the way that traditional databases behave and in the way that they scale. Of course, there's a completely different set of challenges, algorithmic challenges that come with the territory of dealing with vectors and high dimensional vectors that don't exist in the world of just simple textual indexing and columnar data.

4:37And that's where the secret sauce of Pinecone lies, right? It's ability to handle vector data at scale, but maintain the speed and maintainability and resiliency of a database. As you were kind of comparing vector databases to indices and then kind of bringing that compared to that, one of the things that I run across still a lot are people, you know, vector databases are really, you know, incredibly helpful now. But there's still a lot of people out there who don't really understand how they fit in. You know, they don't really get it versus the NoSQL versus relational databases. Or fine tuning.

5:17Yeah. And so, and they hear you say it does vectors and stuff like that. Could you take a moment since we have you as an expert in this thing and kind of like lay out the groundwork a little bit before we dive deeper into the conversation about what's different about a vector database that is storing vectors versus, storing the same vectors in something else? Like why go that way for somebody who just isn't quite, hasn't really ramped up on that yet? So the basic premise is you want to use the right tool for the job, right? And the basic difference, right, between a relational database, a graph database and a vector database or a document database for that matter, right, is the type of content that they are optimized to index.

6:02Meaning a relational database is meant to index a specific column and create an index that would be easily traversable. And in scale, it would be able to traverse it across different machines and do it effectively. Graph database does the same thing, only its world is nodes and edges. And it's supposed to be able to build an optimized representation of the graph such that it could do traversals on that graph efficiently. In vector databases, vector databases are meant to deal with vectors, which are essentially long, high-dimensional set of numbers, meaning you can think of an array with a lot of real numbers inside of that array.

6:45And you can think of this collection of vectors as being points in a high-dimensional space. And the vector database is building effective representations to find similarities or geometric similarities between those vectors in high-dimensional space. And that means that basically it would be very effective at given a vector, finding a vector that is very close, quote unquote, to that vector in a very large space. Right. So to do that, like you need to use a very specific set of algorithms that index the data in the first place and then query that data to retrieve that similar set of vectors to the query vector at a small amount of time.

7:30and also being able to update or make modification to that high dimensional vector space in a way that is not cost prohibitive, right? Or time prohibitive, right? And that's like the crux of the difference between a vector database and other types of databases. Just to draw that out a little bit more. So from your perspective, like what would be if you were to kind of explain to someone, hey, here I've got one piece of text and I'm wanting to match to some close piece of text in this vector space. what might be advantageous about using this vector-based search approach and these embeddings in terms of what they mean and what they represent versus doing like a, you know, TF-IDF has been around for a long time.

8:20I can search based on keywords. I can do a full text search. There's lots of ways to search text. You know, that concept isn't new, but this vector searches seems to be powerful in a certain way. From your perspective, how would you describe that? Yeah, I think that the linchpin here is the word embedding, right? The vector search capability itself is a pretty straightforward mathematical operation that in and of itself doesn't necessarily have value, right? It basically, it's like other mathematical operations. It's a tool, right? The question is like, where does the value come from? And I would argue that the value comes from the embeddings.

9:01And we'll talk about what exactly they are. We just point a flag and say embeddings are represented as vectors, which is why the vector database is so critical in this scenario. But why are embeddings helpful in the first place, right? So embeddings come from a different set, right? Like a very wide set of neural networks that have been trained on textual data. And they create within them representations of different terms, different surface forms, sentences, paragraphs, et cetera, that map onto a certain location in vector space. The cool thing about embeddings is that it just so happens, and we can talk about why, it just so happens that terms that have semantic similarity have a closeness in vector space.

9:50And that means that if I search for the word queen and I have the word king embedded as well in my vector database, and I also have the word dog, right? Because the word king is more semantically similar to the word queen, I will get that as a result and not the word dog, right? And that allows me to basically leverage the quote-unquote understanding of the world that machine learning models and specifically neural networks have, large language models have of the world in a way that I can't quite leverage from other modalities like DFIDF and other BM25, etc. that look at a more lexical kind of perspective on the world.

10:32And so when we talk about practical use cases, like RAG comes up very, very frequently. And the reason for that is because we are in semantic space. A user interacts with the system in semantic space. So that means that they ask the system a question in natural language. We can take that natural language and basically, again, quote unquote, understand the user's intent and map it again into our high dimensional vector space. and find content that we've embedded that has some similarity to that intent, right? And so we're not looking for a exact lexical match, but we're actually able to take a step back, right?

11:13And look at the more ambiguous intention and meaning of the query itself and match to it things that are semantically similar, if that makes sense. Would it be fair to say that you're essentially, because the output of those structures being embeddings and those are vectors and therefore you're essentially storing it and operating on it in a closer representation to how they naturally would be and so you're not doing a bunch of translation just to fit it into a storage medium and and to operate on it therefore it's going to be quite a bit faster since you're is that fair is that a fair way of thinking about it perhaps we're we're in a way we're compressing the representation into something very small in a sense, right?

11:58So you can think of an image, for example, right? An image would, that could be like a megabyte big, right? We can get a representation that in terms of like its actual size in terms of the vector is order of magnitudes smaller, right? And we can use that representation instead of using the entire image, right? To do our search. Now, it just so happens that again, when we're doing embeddings for images, right, we get like that same quality where we're not looking at, you know, we're not looking at an exact match or like pixel matching pixel to pixel with images that we have. We can actually look at the semantic layer, meaning what is actually in that picture.

12:41So if it's a picture of a cat, we would get, as a result, other pictures of cat we've embedded and saved in the database, right? And that will come out of the representation itself of the embeddings that were a result of like, say, like a clip model that we use to embed our image. So I don't know if it necessarily means that like it simplifies things in a lot of ways, it actually adds a lot of more oomph to the representation, right? So you can actually match on things that you wouldn't necessarily expect, right? And that's kind of like what the beauty of semantic search in that sense, is that users can write something and then get back results that don't even contain anything remotely similar in terms of the surface form to their query.

13:28But semantically, it would be relevant.

13:54it.

14:01hand without sacrificing control. Deployment is easy. Pipelines are live API endpoints. Eliminate the need for constant code redeployment and debugging by deploying complex AI pipelines as API endpoints. Team collaboration is easy too. Plum's declarative node-based editor enables you to build quickly while empowering non-technical roles to iterate on what you've without breaking it. You can build advanced AI features, get structured output every time, transform data and leverage validated JSON schema to create reliable, high-quality structured output. So Plum is built for builders. Early stage product teams are using Plum to go from idea to validation in record time.

14:45To get started, go to useplum.com. That's Plum with a B, as in plumber, to request access today. That's U-S-E-P-L-U-M-B dot com. Again, useplum.com.

15:15Well, Rowie, I really appreciate also the statement about adding oomph to your representations. I think that would be some type of good t-shirt that could be derived out of that. I see for listeners who are just listening on audio, Rowe's wearing a shirt that says, love thy nearest neighbor, which is definitely applicable to today's conversation. Well, this is great. So we've kind of got a baseline in a sense from your perspective, what a vector database is, why it's useful in terms of what it represents in these embeddings and allows you to search through. You mentioned RAG. We've talked a lot about RAG on the show over time.

15:59But maybe for listeners, this is the first episode that they've listened to. What would be the kind of 30 seconds or some type of quick sort of remember RAG is X from Rory? Right. So I love quoting Andrei Karpathy with his observation on LLMs and hallucinations. So people, when usually people talk about RAG, they say, oh, RAGs sometimes hallucinate and that's really bad. Right. And Andrei Karpathy says, actually, no, they're always hallucinate. Right. They do nothing but hallucinate. Right. And that's really true. Right. Because LLMs don't have any kind of tethering to real knowledge in a way that we can trust.

16:40Right. We don't have a way to say, hey, I can prove to you that what the LLM said is correct or incorrect based on the LLM itself. We need to go out and look and search. And RAG to me is that opportunity where we can take the user's intent, we can tie it using, for example, a semantic similarity search to structured data that we can point to and say, this is the data that is actually trusted. and then feed that back to the LLM to produce a more reliable and truthful answer. Now, that's not to say that RAG is going to solve all of your problems, but it's definitely going to give you at least a handle on what's real and what's not, what's trusted and what's not, and where the data is coming from, where those responses are coming from.

17:26And it shifts the role of the LLM from being your source of truth to basically being a thin natural language wrapper that takes the response and makes it palatable and easy to consume to a human being. Great, yeah. I think a lot of people have done a sort of, maybe they've done even their own demo with sort of a naive rag, maybe pulling in a chunk from a document that they've loaded into some vector database, they inject it into a prompt and they get some useful output. But one of the things that I think we haven't really talked about a lot on this show, we've talked about advanced RAG methods to one degree or another.

18:07But I know Pinecone, along with other vector database providers, offer more than a simple just search. That's the only function you can do. There's a lot more to it that can make things useful, in particular, like having, you know, you mentioned, Pinecone mentions kind of namespaces that can be used, metadata filters, sort of hybridized ways of doing these searches. Could you kind of help our listeners understand a little bit? So they may understand, here's my user statement. I can search that against the database and get maybe a matched document. But for an actual application, like an application in my company that I'm building on top of this, what are some of these other key pieces of functionality that may be needed for an enterprise application or for a production application that go beyond just the sort of naive search functionality in a vector database?

19:07Yeah, for sure. So we can take this one by one. So metadata is definitely like one of those capabilities that vector databases have that are above and beyond what a vector index would provide you. Right. And basically what they are is, again, the ability to perform a filtering operation after your vector search is completed. And so you could basically limit the result set to things that are applicable in the application context, right? So you can imagine different controls and selection boxes, etc. that come from the application that are more set in stone, so to speak. They're not just like natural language.

19:45They're categorical data, for example. And you can use those to limit the result set, right? So that you hit only what you want. That is something that is very common to see in a lot of different production scenarios. And could you give maybe an example of that? Like in a particular use case that kind of you've run across, like what might be those categories or what, just to give people something concrete in their mind? Yeah. For example, like you can imagine a case where, I'm not going to name the customer, but like You can imagine the case where you want to perform a RAG operation, but you want to do it on a corpus of documents, but not on the entire corpus, but rather on a particular project within that corpus.

20:27So imagine that you have multiple projects that your product is handling, like finance and, you know, HR and whatever, engineering, right? And you want to perform that search and then limit it only to a particular project. And in that case, right, you would use the categorical data that is associated with the vectors that you've embedded and saved in Pinecone to only get the data for that particular project, right? That is like a kind of super simple example. But it can go beyond that, right, and move into like the logic of your application. So like you can imagine a case where, you know, you're looking at a movie, a movie data set, right?

21:06Like, and you want to search through different plot lines of movies, but you want to limit the results only to a particular genre, right? That's another case, right? Like we can just leverage metadata. You can think of wanting to limit the results to a time span, right? A start and end date, right? Things of that sort that kind of have to do more with the nature of when and how and what category the vector belongs into and not specifically the contents of the vector. So that's one thing. Namespaces are another feature that we've seen as being incredibly important for multi-tenant kind of situation.

21:44And multi-tenant rag has become kind of like a very strong use case for us. And that's where you see a customer and that customer has customers of their own and not one or two, but many, many, many. And in that case, you definitely don't want to have all of the documents that all of the subcustomers have to be co-located in one index. And in that case, you basically break them apart. So they're still in one index. So management of the index overall is maintained under one roof. But the actual content and the vectors themselves are separated out physically from one another in namespaces. They're sort of sub-indexes to that super index.

22:25And that's another feature that we've seen as being super important to our enterprise customers. As you're looking at these enterprise customers and with maybe most enterprises, you know, getting into RAG at this point at some level and trying to find use cases for their business to do that. I know, you know, my company and lots of other companies are doing this. What are some of the ways that they should be thinking about these different use cases when we're talking about RAG and semantic search and multimodal things that Pinecone does? What are good entry pathways for them to be thinking about how to do this?

23:02Because, you know, they may have come up with their own, their kind of own internal platform. It might have some open source. It might have some products already in play. but maybe they don't have a vector database in play yet. And so, you know, how do they think about where they're at when you guys are talking to them and you're saying, let me, you know, because we've been talking in the show so far about kind of the value of the vector database and these use cases, but not necessarily kind of an easy pathway. So how do you onboard enterprise people to take advantage of the goodness on this? Yeah, that's an excellent question.

23:36And in fact, it's like a quite a big of a challenge because it ends up being a straightforward pipelining challenge that has existed from the beginning of the big data era. How do I leverage all the insight that is locked in my data in a beneficial way? And the sad part about this story is that it always depends on the specific use case, and it's hard to give a silver bullet. A sort of light at the end of the tunnel is that we've recently published a tool called the rag planner and its purpose is to basically help you figure out what do you need to do to get from where you are to an actual rag application and follow through all of the different steps that are required in between right and sort of like understand like from an understanding of like where your data is stored how frequently it updates like what the scale of your data is etc etc to the point where it could give you some recommendation as to like what are like the steps that you have to do, like in terms of, do you build a batch pipeline?

24:36Do you build a streaming pipeline? What tools should you be using to do those things? What kind of data cleaning are you going to need to do? What embedding models are you going to want to use to do this, right? Like, how are you going to evaluate the results of your RAC pipeline? So all of these questions are pretty complex. So what I would say is, as a general rule of thumb, first of all, like you have to evaluate whether or not RAG is for you, right? So for example, there are a lot of situations where RAG may be the wrong choice, right? Because the data that you have, right? And the actual capability of answering the end user's questions based on that data does not match up, right?

25:16And that's how you get to see cases where chatbots sort of spit out results that may seem ridiculous, but nobody catches it. and companies get into a lot of hot water because of it, right? There are a lot of scenarios where it's much easier to start that journey and to sort of develop the muscle memory that's required in order to set these things up. In a lot of these use cases, you see like a lot more internal processes, definitely in bigger companies, right? Where like there's a very big team that just needs access to its internal knowledge base in an efficient way, but it's not a system that is going to be mission critical, right, in any way.

25:55So like if a person gets a wrong answer, it's not going to be the end of the world, nobody's going to get sued, right. And so what I would say is, there's definitely a learning curve here for big organizations, for sure. It's usually recommended to develop, again, that internal knowledge of what the expectation versus the realities on the ground is going to be, to have like a really good idea of how you assess risk in those situations. And most importantly, how to evaluate the results that are produced by those systems, right? Because a lot of people are like, okay, you build the RAG system, great, and now produces answers, I'm done, right?

26:32Like, we're, everybody's happy. That's farthest from the truth that you could possibly be, right? Like, these systems need to be continuously monitored, and feedback needs to be continuously collected to the point where you can understand, right, like how changes in your data and the way that you're interacting with it, changes in large language models that you're implying are actually affecting the end results, right, are going to be, and how your users are actually interacting with the system overall, right, how all of these things kind of coexist and happen together. And are they working in the way that you want them to?

27:04And of course, you want to do that, you know, in a quantitative and not qualitative way, right? So like, there's a lot of instrumentation that has to go into it. I'm curious, as a little follow up to that, and obviously leaving specific customers out of it, are you tending to see more internal use cases of RAG deployment to internal, groups of employees and stuff, maybe from a risk reduction? Are you seeing more of an external, I'm going to get this right out to my customers and try to beat my competition to it? Where do you think the balance is as of today? I think that there's a widespread and I think that it's a journey.

Read the full transcript

27:36I think that the more tech native companies that we see that are more, I would say, forward looking or technologically adapt to do these things quickly are more ready to not only take risks, but take educated risks in this space with the evaluation that comes with it, right? So like, these are not just like, let's set and forget, but they actually know what they're doing. In those cases, you see them going out to production with very big deployments. That is our bread and butter, I would say at the moment, right? With companies that are more traditional, that have been like, they're not necessarily getting tech native, you see a more cautious sort of progression, which is only to be expected, right?

28:20Like I think that's kind of like natural to see. Well, Rui, I have something that I saw on your website, which was new to my knowledge, which I think is also really interesting. One of the things that I've really liked in experimenting with vector database rag type of systems as an AI developer is having the ability to run something without a lot of compute infrastructure, maybe in an embedded way or an on-disk index, something that I can spin up quickly, something that I don't have to deploy a Kubernetes cluster or something to, or set up a bunch of kind of client server architecture to set up and test out maybe a prototype that I'm doing.

29:04And I see Pinecone is talking about Pinecone serverless now, which is really, really intriguing to me, just based on my experience in working with people, these sort of serverless sort of implementations of this vector search, I think, can be really powerful. So could you tell us a little bit about that and how that kind of evolved and what it is, what's the current state and how Pinecone thinks about the serverless side of this? So serverless came about after we realized that tying compute and storage together is going to limit the growth factor that our bigger customers are expecting to see.

29:45And it basically makes growth kind of prohibitive in a space, right? And so we had to find a way to break apart these two considerations while maintaining the performance characteristics that our customers are expecting and are used to having from our previous architecture. So essentially, like serverless has been a pretty big undertaking on our side to ensure that, you know, the quality of the database is maintained. But at the same time, we can reduce cost dramatically for customers to just give you like an idea with like for the same cost of storing about, I don't know, around 500 ,000 vectors before, you can now store 10 million, right?

30:32And that's a humongous difference, right? Like it's an order of magnitude difference. I think that like to accomplish that, right? Like there was like a lot of very clever engineering that had to happen because again, now having compute and storage separated apart means that storage can become very cheap. But on the other hand, it requires you to handle the storage strategy and retrieval and a lot cleverer way. We have a lot of content on the website that kind of delves deeper into how exactly technically that was achieved. And we won't be able to cover that given the time that we have. But like the basic premise is that you can now grow your vectors index to theoretically infinity, but practically to tens of billions and hundreds of billions of vectors without the cost of the expense becoming prohibitive, which is the main drive for us with our bigger customers.

31:25and also with smaller customers. You can start experimenting. We have an incredibly generous free tier that allows you to start, like you said, if I'm just a developer on my own testing things and trying to understand how vector database works in my world, it's very unlikely that I'll be able to tap the entire free tier plan even several months in with many, many vectors stored. And it will work the same way that our pro serverless tiers work in terms of its performance. So it's not like a reduced capacity or performance in any way. So you get to feel exactly what it would feel like. And the effort that's required to stand it up is minimal to negligible, right?

32:07You just set up an account and the SDK is super, super easy to use. Yeah, and I understand it or sort of representing things, right? Like in terms of the massive, so there's a massive engineering effort, I'm sure, as you mentioned, to achieve this because it's not a trivial thing, But in terms of the user perspective, like if people use Pinecone before and they're using Pinecone now, you already mentioned the performance. Is the interaction similar? It's just this sort of scaling and sort of from the user perspective, scaling and pricing. And maybe also you could touch on, so Pinecone is people might be searching for different options out there and some of them would require you to have your own infrastructure or some of them are hosted solutions.

32:55is Heinkone, at least in its kind of most typical form, would be hosted by you. And yeah, could you just talk a little bit about the user experience pre-post serverless? And then also kind of the infrastructure side, like what do people need to know and what are the options around that? In terms of what happened pre and post. So before serverless, there was like a lot of possible configuration choice that you could do, right? There was, in fact, a lot of confusion with our users. What exactly is the best configuration for me? Should I use this performance kind of configuration? Should I use the throughput optimized configuration?

33:36What exactly am I supposed to use? And the pricing mechanism was a little bit convoluted. And I think that serverless, the attempt there was to simplify as much as possible and to make it really, really dead simple for people to start and use, but also grow with uh with us right so again like i said like the bottom line is you know the external view into what pinecone offers may have looked pretty similar right so like if you're just the user you may say like hey like i got like a cheaper pinecone bill this month and you know like i can store a lot more uh always a good thing right always a good thing right but not super amazing right like but the end result is uh the question is like what happens when you know you can actually store a lot more vectors right what what does that unlock for you and and again i think that like the end of the day right like the way that we see pinecone and this this may help us kind of talk about like what's next for pinecone is a place where your knowledge lives right and and allows you to build knowledgeable ai applications right and having more knowledge is always net positive right in that context right so the assumption is that like you know as ai applications grow they accumulate more and more and more knowledge and they become that more powerful with any additional knowledge that you can stuff into them and so there's like actual value beyond the fact that you can store more right like and and it's cool right your application actually becomes more powerful because it can handle more types of of use cases it has a better ability to be more accurate and respond truthfully to a user when they are interacting with it.

35:20And so I think that like in general, like there's like this blatant kind of value that is only going to be apparent once people really experience what it means to have, you know, a million documents that are stored in Pinecone versus 10 million documents that are stored in Pinecone. And that effect is going to be very powerful. I think that's the majority of the benefit that I see. Maybe that gets to the next thing which I was going to ask about, which I also see the announcement around Pinecone assistance. And I'd love to hear more about that. Of course, sometimes maybe that can be loaded language also for people in the AI space.

36:02But in terms of this assistance functionality for Pinecone, what are you trying to enable and where do you see it headed? So that has to do with the question that Chris had before, which is like, what is the journey, right, for customers, right? And I think that like, as a general purpose that we had around Assistant was to reduce the friction between me having a bunch of documents that I want to interact with, with an LLM or an AI in some form and capacity, to the point where that actually works, right? There are a bunch of ways of going about it. I think Pinecone wants to bring, on top of our very robust vector database, a very smooth experience that lets users really do very little and get all the value out of Pinecone without having to think too much about it.

36:53So for that purpose, we don't only have the ability to take your documents and then embed them and, you know, do the end to end process of creating that completion endpoint for you. Right. We're also the ones providing the actual inference layers as well. Right. And so it's not going to be again, if you ask like me, like this question is like, how do you build a rag pipeline? right like a year ago even right i'd have to tell you hey you have to go to like some embedding provider you have to find someone who would do like your uh you know pdf extraction or you know take the data and chunk it and do all this stuff right no more right like the reality here is you can take a set of documents throw them at this knowledge assistant and the rest is kind of quote-unquote magic right it just happens for you behind the scenes while maintaining the quality that you want to get and at the scale that Pinecone can deliver, right?

37:50Which is, again, another differentiator. So like I said before, Pinecone is built to withstand hundreds of billions of documents, right, of vectors that you would store with us and still be able to produce responses in a reasonable amount of time. And that's true for Knowledge Assistant because Assistant sits on top of the vector database. So it sounds like that may be a really good way, especially for small organizations. You know, we talked about enterprise and they have a certain infrastructure and teams to go with that. But there's so many more small organizations out there that have very little in terms of trained people necessarily to do that.

38:30And they don't have all the infrastructure in place. And they're looking, you know, with assistance and serverless, they're looking for simple ways to onboard and get utility out of it. would you say that the combination of serverless and assistance and then maybe whatever they might have in aws or whatever platform that they're using is kind of just made to gel easily for them so they can get to something working pretty quick yeah i mean at the end of the day like if you think about it right like the process shouldn't be as complicated as it is right it's just that there are many parts to it and nobody picked up the gauntlet of saying like hey we'll just do it all you know what i mean uh because all of it is quite complicated to do right right and so um yeah like i think that like initially we'll see smaller organizations kind of you know picking that up because they don't have the resources but as time moves along you know you're gonna have to ask yourself even as a bigger organization do i want to own this pipeline right is it something that i need to own right and what value am i getting from actually owning all this right and so um yeah like It would be interesting to see.

39:36So this is a very, very new product still in public beta. And it will be interesting to see how the market kind of reacts to it and sort of experiments with it. But my bet is that as time progresses and knowledge assistants themselves become more capable doing things maybe beyond RAG or beyond simple RAG, quote unquote, that more and more sophisticated organizations might want to actually give it a try. And that really brings us maybe to a good way that we like to end episodes, which is asking our guests to sort of look into the future a little bit and not necessarily predict it because that's always hard, but to look into the future and kind of what are you excited about?

40:19it could be related to vector databases specifically or pine cones specifically, but maybe it's more generally in terms of how the AI industry is developing, the sorts of things that you're seeing customers do that are encouraging, whatever that is, what sort of keeps you excited about where things are headed going into the rest of this year? I'm excited about the fact that we're seeing sort of like a resurgence of what you would call traditional AI kind of come back into the fold in the form of, for example, Graphrag. I think that the notion here is that for the longest time, and I think it's been since GPT-3.5, you basically saw this, I think, over-indexing on LLMs, right?

41:07For good reasons, right? They're super exciting. They're very powerful, right? And they can do really, really cool things, right? But with that said, it's as if every other technology that has ever existed before just like dropped off the face of the earth. And nobody has ever like talked about like, okay, wait, so what can we do with those things and LLMs, right? Like, and where do LLMs fit in the bigger picture? I think that vector databases kind of like put LLMs in their place a little bit in the sense that, you know what I mean? Like you're not thinking of the LLM as being the end all, be all, like this is the only tool that we need.

41:42I'm very excited to think of LLMs as these operators or agents that can tap into the capabilities that exist in other systems. And I think that what we're going to see more and more and more is that people are going to figure out like in what subset of the ecosystem does each tool belong? So what set of problems that each tool solve? For example, like a vector database solves like the problem of bridging the gap between the semantic world and the structured world. A graph database can solve problems like formal reasoning over well-structured data. Relational databases can solve a whole set of different problems that they used to be solving, like aggregation, et cetera, et cetera.

42:23And then you can imagine that LLMs and agents can sit as sort of like an orchestrating mechanism and a natural language interface mechanism on top of all those things together. And that's what I'm excited to see. It's kind of when the community as a whole is going to wake up from its LLM fever dream and sort of realize that there's other things out there and realize that it has so many more powers that it could yield to make really exciting applications. That's awesome. Well, thanks for painting that picture for us, Rui, and for taking time to dig into so many amazing insights about vector databases and embeddings and knowledge management in general.

43:07So yeah, appreciate what you all are doing at Pinecone and hope to have you on the show again to update us on all those things. Thank you so much. Thanks for having me.

43:23All right, that is Practical AI for this week. Subscribe now. If you haven't already, head to practicalai.fm for all the ways. And join our free Slack team where you can hang out with Daniel, Chris, and the entire ChangeLog community. Sign up today at practicalai.fm slash community. thanks again to our partners at fly.io to our beat freaking residents breakmaster cylinder and to you for listening we appreciate you spending time with us that's all for now we'll talk to you again next time

From the publisher

Daniel & Chris explore the advantages of vector databases with Roie Schwaber-Cohen of Pinecone. Roie starts with a very lucid explanation of why you need a vector database in your machine learning pipeline, and then goes on to discuss Pinecone’s vector database, designed to facilitate efficient storage, retrieval, and management of vector data.

Join the discussion

Changelog++ members save 3 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Plumb – Low-code AI pipeline builder that helps you build complex AI pipelines fast. Easily create AI pipelines using their node-based editor. Iterate and deploy faster and more reliably than coding by hand, without sacrificing control. 

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Vectoring in on PineconePractical AI · 44 min
Listen in VO