In short
Software Engineering Daily - Episode Summary: Turbopuffer with Simon Hørup Eskildsen
Episode Overview Podcast Title: Software Engineering Daily Episode Title: Turbopuffer with Simon Hørup Eskildsen Description: This episode focuses on TurboPuffer, a vector database designed for speed, cost-efficiency, and scalability, crucial for AI applications. Simon Hørup Eskildsen discusses his professional background, the origins of TurboPuffer, its technical design, and the economic considerations of vector storage.
Key Themes and Concepts
Vector Search in AI Applications
- Foundational Technology: Vector search is essential for applications such as semantic code search and contextual retrieval in large language models (LLMs).
- Challenges: The high cost of vector databases grows with data storage needs, making it challenging for businesses to adopt these technologies.
Introduction to TurboPuffer
- Founders: Created by Simon Hørup Eskildsen and Justin Lee in 2023.
- Adoption: Notable companies like Cursor and Notion are users of TurboPuffer.
- Focus: Emphasizes speed, cost, and scalability.
Simon Hørup Eskildsen’s Background
- Career Path: Began at Shopify, contributing to infrastructure and database scalability.
- Transition: After Shopify, he worked on various infrastructure challenges before identifying the need for a new database architecture.
Technical Insights
Database Design and Architecture
- Inspiration from Work: Simon’s experiences at Shopify informed his understanding of database bottlenecks and scalability challenges.
- New Architecture Needs: The need for a new type of database architecture became apparent, particularly for AI-related tasks.
Storage Mechanics
- NVMe SSDs: Utilizes NVMe SSDs, which provide a cost-effective storage solution compared to DRAM (Dynamic Random Access Memory).
- Cost Efficiency: Highlights the drastic cost difference between cloud storage options (S3/GCS) and traditional memory storage.
Performance Metrics
- Cold vs. Warm Queries: Discusses the difference in performance when data is cached (warm) versus when it has to be fetched from storage (cold).
- Round Trips: TurboPuffer aims for a maximum of three round trips for sub-second cold latency, balancing performance and cost.
Vector Indexing
- SPFresh Indexes: TurboPuffer employs SPFresh vector indexes, which are optimized for scalability and recall.
- Clustering vs. Graph Methods: Explores the trade-offs between cluster-based and graph-based indexing approaches.
Use Cases
Cursor and Notion
- Cursor: Reduced operational costs by 95% after migrating to TurboPuffer, supporting semantic search and embedding models for codebases.
- Notion: Also transitioned to TurboPuffer for its AI features, enabling cost-effective indexing across large datasets.
Economic Considerations
- Pricing Strategy: TurboPuffer does not offer a free plan; they emphasize a commitment to customer support and experience.
- Market Positioning: Suitable for companies with the ambition and scale to index large amounts of data efficiently.
Future Directions
- Feature Expansion: TurboPuffer is focused on adding more database features and enhancing full-text search capabilities.
- Team Growth: Currently a small team of under 20, with plans to expand to support customer needs and product development.
Conclusion The episode highlights the innovative approach TurboPuffer is taking in the database landscape, providing a cost-effective and scalable solution for businesses looking to leverage vector search in AI applications. Simon’s insights into the challenges and design principles of TurboPuffer offer a deep understanding of modern database requirements.
Additional Resources
- [TurboPuffer Official Website](https://turbopuffer.com)
- Follow Simon Hørup Eskildsen on [X (Twitter)](https://twitter.com/Syrupsin) for updates and insights.
---
This summary captures the essence of the podcast episode, highlighting the key discussions around TurboPuffer's technology, economic implications, and the vision for the future.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Vector Search has become a foundational technology for AI applications, enabling everything from semantic code search to contextual retrieval for large language models. However, a major challenge with vector databases has been the cost as data storage scales. TurboPuffer is a vector database that focuses on speed, cost, and scalability. It was created by Simon Harub-Eskelson and Justin Lee in 2023 and has seen adoption from high-profile companies such as Cursor and Notion. Simon joins the podcast with Gregor Vann to discuss the origin of TurboPuffer, its unique technical design, the economics of vector storage, and more.
0:42Gregor Vand is a CTO and founder, currently working at the intersection of communication, security, and AI, and is based in Singapore. His latest venture, Wintik.ai, reimagines what email can be in the AI era. For more on Gregor, find him at van.hk or on LinkedIn.
1:13Hello, welcome to Software Engineering Daily. My guest today is Simon Eskelson. Thank you for having me, Gregor. Yeah, great to have you here, Simon. We're here today to talk all about Turbo Puffer, which is a company some of our audience may have heard of equally, maybe quite a few haven't yet, and that's going to be why we're here today and talking all about it. However, you probably are using this product behind the scenes without you realizing. So we're going to get into that as well. But I think to begin with, Simon, you've got a really interesting kind of backstory, and we always dive into those with our guests.
1:48So just maybe talk us through very briefly how you kind of got to Turbo Puffer. And I'll just call out, you did spend quite a chunk of time at Shopify. And I think that'd be interesting to understand how that's helped lead into what Turbo Puffer is today as well. Yeah, that's right. I started my career at Shopify and moved to Canada as a result where Shopify was built. I moved from Denmark where I grew up and worked on infrastructure at Shopify for almost a decade. When I joined in 2013, it was in the hundreds of requests per second. And when I left, we regularly saw peaks in north of 1 million requests per second.
2:23And as part of that, I worked on more or less every single aspect of the infrastructure. Generally, the bottlenecks for that kind of scale tend to appear in the data layer. Those are the most persistent ones. So I spent, yeah, almost the entire time working on that layer between the Rails app and the databases and sometimes inside the databases themselves, but mostly on top of them. So, yeah, almost 10 years working on every single aspect of database scalability at Shopify. And then after that, I spent about two years working in small increments at my friends' companies on whatever infrastructure problems they had.
2:56It turns out in 2022, 2023, it's mostly tuning Postgres auto vacuum was the persistent infrastructure problem. And it was there through that I discovered how a new type of database was ready to be built. And a bunch of things had changed that allowed a new database architecture to be built. And that it felt like a lot, there were going to be lots of companies that wanted to connect a lot of data to AI. And that a new search engine on this new storage architecture was ready to be built. And so I started hacking on the first version of Turbo Puffer back in 2023, in the summer of 2023. Nice. And just sticking on that for a second, I think, again, audience might find it helpful to kind of understand like you left Shopify, but it was almost, I guess, two years before, as you call it, like you started on Turbo Puffer.
3:40What was the choice around not just like launching straight into something else? I guess it's around this idea that you're maybe waiting for kind of a problem to solve, so to speak. But how would you speak to doing what you did, which was just go and help other people with their problems for a little while, for example? I went to Shopify right after high school. So I'd spent my entire career inside of one company. And so it felt I couldn't develop high conviction on joining one particular company because I'd only been in one company basically my entire life. I did some little bit of work for another startup doing high school.
4:13But other than that, it's been a decade with one company. So I had a bunch of friends who had a bunch of infrastructure challenges that seemed interesting. So I wanted to bop around. and on small contracts solve a very actual problem at SaaS companies that had something and just over deliver on some infrastructure promise. And that was a really good way for me to see a bunch of companies. That was really the primary concern. And I took summers off for two years. That was nice. And it gave me a bunch of space to get in the best shape of my life and some other things. I didn't leave with the intention of founding a company.
4:42I didn't leave with the intention of really anything other than it was ready for a change. And this felt like the natural thing. I spent a lot of time at that time also working on my napkin math, which is a bunch of blog posts I've written about doing first principle, sort of, okay, the memory bandwidth is this much, so you should be able to do the task in this long. I spent a bunch of time writing that and building out all the numbers for it. And then at some point, this problem just started staring me in the eye and I couldn't stop thinking about it. Yeah, awesome. So let's get into Turbo Puffer.
5:10So let's start with the origins, if you like. You did sort of allude to it. You were helping some friends and you've written a few blog posts about Turbo Puffer. generally one of the first ones is kind of the origin story i believe you were helping friends at readwise do you want to just talk us through like where did this come from yeah so readwise run by some of my friends it's a company that their first product was essentially every highlight that you highlighted with kindle send you an email every day or every week with some of them and i love the product and always been close with the founders there they were launching at the time a reader product so read it later you take an article it's nicely formatted it's on your phone you can listen to it and for that they just needed their postgres database up to snuff and so i was helping them with that again tuning out of vacuum again it's the biggest problem we have i felt a little like gaslit by the orange site here having spent like a decade thinking that mysql was a worse database than postgres turns out that they're just two different databases and one of the biggest challenges with postgres is tuning out of vacuum which i spent some time doing for them anyway so for readwise one of the other things that we were working on this was like the fall of 2022, where I feel like everyone almost remembers what they were doing as ChatGPT dropped.
6:18And when ChatGPT dropped, it was while I was in the middle of doing this database tuning, and we decided to just build a bunch of AI functionality. And it's like, oh, like, maybe we can do this to help people comprehend what's going on in this article. And at the time, the context windows were pretty small. And so we started using vectors very early on to draw the right parts of the article in. I also played around with creating a little recommendation engine. So it would find one article and recommend another article that was similar. And so it's clear that there was sort of this foundational layer brewing.
6:46And as we had something working, I ran the back of the envelope math on what it would cost to take all of the articles that even at the time there was pre-launch was in Readwise Reader and put it into a vector index on one of the databases around then that could do it and did the math. And their Postgres database at the time was costing them a couple thousand dollars, 3k a month. And the back of the envelope math told me that on a reputable vector database that we thought could do the workload, it would cost 30 grand a month. Even if the company could afford that, it was clearly detracted from launching this feature, right?
7:24We were just not going to invest in something that was going to cost an order of magnitude more to store a subset of the data for just powering recommendations and some context into an LLM. And we just sort of put it in the bucket of, well, the token costs for the LLMs are coming down. So presumably someone's going to figure out how to make this cheaper as well. And I just sort of saw out the engagement there and got their database up to snuff and they launched and it was a very successful launch for them. And I couldn't stop thinking about how are we going to make this vector storage cheaper?
7:54And I was like, my memories of running the search clusters at Shopify were coming back to me. It was one of the worst clusters to get woken up by. and it very quickly became very obvious to me that all of the incumbents were storing everything in memory. And if you have a kilobyte of text, as you might have in an article, often there are multiple kilobytes of texts. In order to turn that into a vector embedding, you're going to put it into a bunch of chunks, right, that are semantically meaningful. So you maybe take a kilobyte of text and there's four paragraphs, so four chunks. You turn every one of them into a vector.
8:26Those vectors might be six kilobytes each. So now you have 20 to 30 kilobytes of vector data from one kilobyte of text. So at 20 to 30x amplification. For a normal full-text search index, the amplification to build all the structures to do full-text search is around two, maybe three on a bad day, depending on the data distribution, but somewhere between 1.8 and three. For vector indexes, we're now talking about 30x, right? Just to store the vectors, let alone build an index. And so that was why it was so prohibitive, why the costs were so high, let alone that all of these databases had to store everything in DRAM, which on a cloud provider costs somewhere between$2 to$5 per gigabyte.
9:09And it was clear to me that at Readwise, the economics didn't line up, right? It would have been like, I don't know how many users they had at the time, right? And there was a timing thing here, I guess, hardware-wise in terms of NVMe SSDs. So non-volatile memory express. Could you maybe just speak to that? as well, that sort of backend piece, if you like. Yeah, exactly. So I think that to build a good database, you need two things. You need a new workload and you need a new storage architecture. And the new workload here seems to be that, well, there's a new way to power these recommendations to pull context into another element.
9:39The second thing you need is new storage architecture. The new storage architecture that started to become clear to me around the time in early 2023, as I think about this problem, was that, well, NVMe SSDs are very fast. I think they got faster than anyone really expected, but they were actually not generally available in the clouds until the late 2010s, like around 2017, 2018 on AWS. So databases hadn't been built around them. And the phenomenal thing about NVMe SSDs is that per gigabyte, they're about 100 times cheaper than DRAM. But in terms of how much bandwidth and gigabytes per second you can get from them, they're only about an order of magnitude, maybe even only 5x slower than DRAM.
10:19So you've got this amazing set of principles, but most databases haven't built around getting all of that bandwidth through. You have to bypass the Linux page cast to get the maximum. You have to build your storage engine in such a way that it tries to get a lot of data for every round trip. So that was the first thing that had to be true to build a database like TurboCover. The second thing that had to be true was that OPEX storage like S3 or GCS, depending on the cloud you're on, need. It's a very nice primitive that they're consistent. You put an object and when you read it, it's the same object.
10:48and S3 only became consistent at reInvent in 2020, which is remarkably late, but a very nice property to have when you're building a database. You can build around it by creating new files and so on, but perimeters are needed. The other primitive that TurboPuffer was waiting for was, which was actually only released, we're recording in July of 2025, only at reInvent last year, so seven months ago, did S3 release the final API that is required for a TurboPuffer-like database, which is Compare and Swap. Compare and swap allows you to put a file on S3, read it, do a modification to it, and then write it back only if it hasn't been changed during that time.
11:25With that simple primitive, you can now build a database like TurboPuffer that doesn't have any dependencies other than OPEC storage, other than S3 or GCS. This API was available in GCS on GCP first, which is why TurboPuffer started in GCP because it was available there and we developed strong conviction that this was going to be a ubiquitous feature. because actually every other OPEX storage implementation other than S3 had it. And finally, they released it last year. And a month later, we went into AWS. And until then, we'd run workloads across cloud. So these were the three things that we needed to build TurboPuffer that became available around that time.
11:58And then suddenly, well, we had the two prerequisites to building a great database. We had a new workload and we had a new storage architecture. And a new storage architecture meant that there was differentiation from tacking this on to existing databases. Yeah. Yeah. And I mean, you obviously touched on sort of the prerequisites, I guess, but in terms of actually technical challenges in making sort of S3-like storage performant, like, were you aware of kind of those ahead of time, if you like, or is this still something where just through solving the problem, you've come across those challenges?
12:32And what would you say those were? I started working on ZerboHoffer in May of 2023. And the main challenge was that, And let's talk about vector indexes for a second, because it would be relevant of like, how do you put this on S3? So there's two broad categories of how you build a vector index. If you have a bunch of data, let's say, for example, that you're Spotify and you want to recommend music to people, then you take every song and you pump it through some kind of model and it puts it in vector space. And songs that are adjacent in vector space, as you plot it into this massive coordinate system, right?
13:06Imagining it's in two dimensions. songs that are adjacent in vector space would also be songs that might be similar and so for naturally clusters form there's a rock cluster and a pop cluster and a rap cluster whatever right as the semantic relationships sort of unfold in this space models have gotten very good in 2022 of taking a piece of data and plotting it into the coordinate system but not so much as actually storing the coordinate system and querying it so when you have a vector index which is essentially just a coordinate system in many dimensions. But if you're visualizing this in your head right now, you should visualize it in two.
13:39When I go on an e-commerce site and I search for red dress and they have a burgundy skirt, well, those two points would be very close in vector space. And this is a very nice problem to have solved in such a simple way. But searching for a query across all of the things that you have available, maybe the hundreds of millions of songs that are in Spotify, maybe the billion things that are on an e-commerce catalog, the only way to get the exact result of the vectors that are close in that coordinate system is to look at every single vector and compare with the query vector of like which ones are closest this is the only way that you're guaranteed that you're going to find the top 10 closest vectors but this is slow it's incredibly slow right a million vectors can easily be gigabytes big and searching at gigabytes in memory if you max out the node takes maybe a couple hundred milliseconds so that's maybe not bad in itself but now every node that's completely maxed out is doing five requests per second and that's going to cost you thousands of dollars per month right depending on the node sizes and so on so instead what we do is we have approximate algorithms we say okay we're okay with maybe a 95 or 99 accuracy over what we retrieve so in the vast majority of cases we're getting the exact results but in some outlier cases we're not there's two ways to build these approximate nearest neighbor indexes or ANN for short.
15:00There's others, but these are the two that people have productionized. The first one are the graph-based indexes. So you can imagine that as you're adding things into this coordinate system, you can come up with some heuristics to connect vectors that are adjacent in vector space or in the coordinate system, and you can connect them in the graph. And then you can walk the graph by dumping yourself in the middle of the graph and then continuously searching a graph to things that are closest and then find some good results. This works really well. And it was by far the most common method of doing vector search in 2023.
15:31Almost every production use case was using this. And it was what a lot of the vector indexes at the time were doing. The second way of doing it is a cluster-based approach. This is actually the classical method. This was the first thing I think people came up with and productionized, where what you do is you take all of the vectors that you have and you create clusters, right? So if you use the music example, there might be a cluster that's rock and the pop and the rap clusters, and you create these natural clusters. And then in every cluster, you take the average of all of the things that belong to that cluster, and that's the centroid.
16:06And then when you do a search, you say, okay, I'm searching for songs that are close to this song, and you say, okay, well, it's closer to the pop cluster, so we're going to search only the pop cluster. And that way, you're only searching 33 % of everything. And in reality, this way you can cut the space in such a way that you have to search only a small percentage of the entire data set. It's a very nice way to visualize. And there's some semantic meaning even in what these clusters mean. So these are the two competing ways. And at the time in 2023, I think it seemed like graphs were going to win.
16:36But the problem with graphs is that they're very fast in memory, right? In memory, you can navigate a graph where you dumped into the center. And then every time you have to read from memory, it's 100 nanoseconds. and you might have to go, okay, 100 nanoseconds to land in the center, 100 nanoseconds to go to the next part in the graph, 100 nanoseconds, 100 nanoseconds. And in aggregate, you know, that's still going to be very fast because memory is fairly fast and will CPU cache really well for these random reads. But when you're talking about disk and you're talking about S3, this doesn't work, right?
17:04If you're reading a couple of kilobytes off of S3, you're going to have a P90 in the hundreds of milliseconds. So you land in the middle of the graph, 100 milliseconds. You go to the next node, 100 milliseconds or 200 milliseconds. Go to the next one, next one, next one. And every single time it's 200 milliseconds because you can't predict which jumps you're going to make. You can make a bunch of rustics, but you can't really do much. So you can try then to shrink the diameter of the graph, right? So you can try to make it so that there's less jumps and make the graph like denser, for lack of better words, right?
17:30That will work on audio. And this can work, but it's very difficult on large graphs to get less than maybe 9 to 10 or so round trips on millions or hundreds of millions of vectors. It also becomes very expensive to write because every time you update one of these nodes, you have to sort of update a lot of things around it. And this is also problematic on disk. Again, it's not a big deal in memory, but it's a big deal on disk to do a lot of random reads, and especially on S3, editing a lot of random files. So graphs just don't work that well for it. They are amazing for in-memory, and it's almost impossible to beat the performance in memory.
17:59But when things are on disk or they're on S3, you have round-trip latencies into hundreds of milliseconds for S3 and into hundreds of microseconds for disk, neither of which is going to work well for having lots of these dependent jumps. But of course, if we go back to the clustering method, that seems very crude and primitive. Well, you could just download all those centroids, right? Here's the wrap centroid. Here's the pop centroid and whatever other genres. And you download the centroids.bin. And it's like a nicely packed binary back. You search through all of those and you find the closest few clusters.
18:30And then you do another round trip to S3 and you download just those clusters. You've gone to S3 twice. You fetched a lot of data, but S3 is fine with that, right? You can download a lot of data in 100 to 200 milliseconds. And this works really well for disk two, right? Because you go to disk and you get the chunk of centroids in one big page, and then you start getting the individual clusters again in two round trips. And so we built the storage engine all around minimizing the number of sound trips. We use a cluster-based index because we want to minimize the number of round trips, which allows to give this really, really good performance, but on these storage mediums that are so much cheaper than memory.
19:08Again, memory being in the about$2 to$5 per gigabyte, and S3 is$0.02 per gigabyte. and disks are about eight to 10 cents per gigabyte. Yeah, and we're going to get onto kind of how we even measure any of this. I've got kind of one of the more sidebar questions. So we'll come back to that in a second because just a lot of what you've been describing there, I feel like I can visualize it. And maybe for some of our listeners, it's still a little bit difficult to visualize. So we'll come back to that in just a second. Capital One's tech team isn't just talking about multi-agentic AI. They already deployed one.
19:39It's called Chat Concierge and a Simplify in car shopping. using self-reflection and layered reasoning with live API checks. It doesn't just help buyers find a car they love. It helps schedule a test drive, get pre-approved for financing, and estimate trade-in value. Advanced, intuitive, and deployed, that's how they stack. That's technology at Capital One. You did talk about this is all about round trip, and you have also written about this idea that I believe turbo offers should be a maximum three round trips for what you call sub-second cold latency. So again, I mean, could we also just talk a bit about like cold and warm here in terms of how does that also come into it?
20:18Because I think you did a very good job of sort of explaining how to think about these jumps between data, but maybe then how are we looking at, I think a lot of our audience are familiar with the concept of effectively cold and warm data. So how does that come into it as well? Yeah. So the canonical source of truth for all data in TurboPuffer is Obix storage. When you do a write and we return a success back from TurboPuffer, it's been committed to the most durable systems on earth, which is S3, GCS, and friends. When you do a query, we will send it to a node. And if that node has it in memory cache, then we'll use that.
20:52It's the fastest. And if it doesn't, then we'll go to disk. And if it doesn't, then we'll go to Obix storage. If we haven't seen a query for, say, a week, then it's not going to be in cache anymore. And we have to go directly to Obixstroid. So in the first round trip, we'll get a bunch of metadata files, like what's the schema, what's the most recent index, and a bunch of other metadata about the namespace. Then we'll get the centroids and a bunch of other things that might be relevant to any filtering that you're doing. And then we'll get the clusters and exactly what we need from the clusters to satisfy the query.
21:24And those are the fundamental three round trips. There are situations where we're going to do more round trips. If you're fetching a bunch of information and we can't make the decision that fetching all of that in that other in the third round trip is going to be the best performance, the query planner has to make a decision about whether to do another round trip or to try to fetch more data. So that's the scenario for a cold query. Again, these round trips, depending on the size of them and the mood of S3 on that day, take about 200 milliseconds or so each. So we get cold query performance just short of a second in 600 milliseconds, 800 milliseconds.
21:57But of course, S3 also has caches. and it depends, but that's generally what we see. In fact, when you do a query to S3, the variance is so high that once we cross the 98th percentile of the latency that we generally see for something of that size, we will send off a second query to try to minimize the variance in these query latencies. That's a cold query in the high hundreds of milliseconds. A warm query just can go to memory or can go to disk. And at that point, there's no reason why this can't be as fast as any traditional storage architecture. It's only that cold query that happens once in a blue moon.
22:32Some of our customers will do pre-flight queries, right? When you open the Q &A dialogue in Notion, it will send a query to hint to TurboPuffer that we should start warming that index. And we will do that in the order that reduces the latency as fast as possible to try to minimize this impact for the user. So generally, this is not an issue that we see, and especially with the cost that we can get by only having that single copy. This is also where the name comes from. right? It's like the puffer fish is fully deflated when the data is only on object storage, and we can inflate it all the way into DRAM or even CPU caches.
23:05As you query the namespace more, TurboPuffer gets faster. Got it. Okay, that makes sense. So yeah, sidebar question here. This is just actually jumping fully into the product itself from like a visual standpoint, like, because we're going to come back to a whole bunch of other things about performance, etc. But if you're a developer how can you imagine visualizing the data through turbo puffer i've used other vector store products so i can kind of think of how they try to represent this data to me some of it kind of worked and some of it i thought didn't make a lot of sense to me so like how has turbo puffer kind of approached this one do you mean visualizing it to the user correct yeah through like a gui or so forth yeah yeah turbo puffer's console is still fairly simple and operational.
Read the full transcript
23:49I would love for there to be more of a playground. What we see our customers do is generally export the namespace locally and then try to visualize it. I don't know how many of our customers do visualize the data and how many of them just start writing evals against it. I would love to help with tooling, but we've been very focused on the database itself. Okay. Yeah, no, I think it's really helpful to understand. It probably also speaks to where, obviously, where Togo Puffer has been focusing, maybe perhaps versus some other products that sort of are trying to like hit all the points at once and maybe not doing a bit of a general approach shall we say let's go to performance and i believe that the term in vector stores is recall so that effectively trade-off between latency and accuracy i believe turbo puffer kind of does measurement on i believe at least samples one percent of all queries internally and you have presented back that data anonymized through blog posts as well but talk to us about that?
24:45Like, how do you measure? Why are you measuring? And where do we go from there? Recall is extremely important, because if you're building a pipeline that searches, you don't want to think about in your evals, whether your search engine is inaccurate, you just don't want to think about that, that should be the job of the vendor, or this database that you've chosen, you should not have to manually tune it, and you should not have to manually run recall against it. Turbo Puffer, of course, as we reiterate on our ANN algorithm, we run it against, of course, various benchmarks that we have internally, right?
25:16Academic benchmarks and so on to make sure that it performs. But nothing beats real world performance, right? In a real world, someone is going to insert the same song a million times into a cluster with lots of other songs. They're going to put that burgundy skirt in there like 20 million times, right? And not realize they have the bug and they still are going to expect that the accuracy is going to go up. So overall, accuracy is very important to us, and we consider it our job. I want to just briefly explain what recall is, and then I'll talk a bit about how we work on it at TurboCoffer. So recall, if we go back to my original example, right, of the only way to get the absolute true result of the top 10 closest vector embeddings to another query and vector embedding is to look at the entire dataset, right, O of N.
26:01With the A and N, what you do is, if you issue a query with the approximates near and neighbor index, then you compare the top, let's say top 10 for the approximate results and the exact results. If you have a recall of 90%, it means that nine of the 10 results were correct in the ANN result. Of course, if it's one or 100%, then everything is overlapping. We find that our customers are very happy somewhere between 90 and 100 % erring on around 95 % recall, right? So it averaged across everything. That's what recall means. And generally anything above 90 % is really good. Recall is a tricky metric because you could search for banana in a cluster or like in a data set that only knows about songs.
26:46And it's going to give you a result because in some way, you know, banana is close to something, right? And so the vector distance might be very long, but it also in the clustered index, the longer you are away, the further the clusters are away from the query vector, the worse the results get. So it's a flawed metric, but it's like the best flawed metric that we've got and that we're all measuring against. So we're aiming for 90 to 100 % recall, erring on the side of 95. And when we, very early on in TurboPuffer's history, we decided to really take this on an hour problem. And as you mentioned, a percentage of production queries are sampled in production.
27:19So it means that for a random percentage of like somewhere around 1%, but it sort of scales how many queries a user does. We will send a query over to a different worker fleet and it will evaluate that like against the exact result, the approximate result and compare it and then report the number back to our instrumentation in Datadog. And so we have a dashboard, right? Where we look at every single customer and look at their recall and we will get the big red dots if someone's recall is below 90%. And that's how we reevaluate it because I think production is the only thing that tells the true story.
27:53You can't just evaluate against academic benchmarks. I think being on call for 10 years almost at Shopify taught me that nothing matters other than production. And it's going to be the same for accuracy, is that the only way that we will trust our recall is if we know that on all the production data sets, it's above 90%. Got it. I believe there is something that comes into this also, and native filtering. Perhaps you could talk to us a little bit about that. Filtering is the most important thing with recall, because this is where it gets tricky. And this is where just tapping on a vector index to an existing database is not quite enough to ensure that a recall is high.
28:30So let's take some simple examples and work through them, right? And explain what this pre-filter, post-filter, in-place filtering, whatever means, right? Let's say that you have a query that is, we'll keep using an e-commerce example here, and we're using it on a very large e-commerce data set. And we want to filter out any products that are not public. Let's say that 99 % of all the products are public, right? Like 1 % is sort of like held back as people are iterating on them or haven't released them yet or they've sold out. If only 1 % of the data doesn't match the filter, it's probably fine to just over filter by around 1 % and then filter it out after.
29:08That's a post filter. And with that, if you evaluate the recall, again, the only accurate way is to evaluate it against everything. You will get very, very high recall with a post filter. But you could imagine in a case where you're matching against an example that would only match, let's say, 10 % of the data set, it could be everything that, I don't know, everything that is the color blue, and that's only 10 % of the data set. Well, in that case, if that's 10 % of the data set, and I've overfetched 100 items, well, I'm not going to get that 100 % recall, right? The math just doesn't really work out for the precision that you need.
29:43So in that case, well, it's only 10 % of the data set. So maybe what we can do is we can just find everything and read another index, like a traditional B2E index or a bitmap index that's blue, and then evaluate all the vectors. And that works okay for something that's like 10%. It will be a little bit slow because that's a lot of vectors to look at, but it might be okay. But the trickiest queries are the ones in between, right? They're the ones that filter out maybe 50 % of the data sets. Imagine an example like you're in Singapore. Everything that ships to Singapore, maybe that's 50 % of the catalog, right?
30:13So it's like, okay, do you post-filter? you have to overfetch by a lot. You have to get a lot of stuff to make sure that you have the right recall because it might in a clustered index, it might completely eliminate some clusters. Like maybe, you know, food items are never shipping to Singapore, but you're searching for a banana and it's just not matching that much of it. But it might go into the next like food color clothing that might start shipping, right? But those clusters that you've cut off from. And when you start to think about it in the cluster sense, the query planner really has to be aware of like how much of different clusters match the filter and how much does that cut off and then you have to have some heuristics around how many vectors you have to look at in the order from the query vector to look at enough vectors that you can guarantee a high recall i realize this is hard to parse but really what i'm trying to impress here is that the query planner that plans how much data the query needs to look at to get high recall needs to be very aware of both the vector index and also the filtered index to get high recall and most databases that have just slap the vector index onto an existing database will only pre or post filter.
31:14Often the user has to choose. So you have to know about the selectivity of their data set, but it's not a trivial thing to choose. Yeah. And that leads quite nicely into what I wanted to touch on next, which is vector indexes. I believe it's SPFresh vector indexes that's used at TurboPuffer. This is probably a concept at a very high level that most of our audience are at least familiar with to some degree. If you've worked with databases, you probably understand the concept of an index, i.e. some way of saying to the database, hey, these are the kinds of queries we're going to be making on a very regular basis.
31:48So we need an index across, say, these three columns on this table, because that's the kind of lookup that we're going to be doing. That's a relational database index, for example. How does that look, obviously, in the vector context? And again, the choice here, SPFresh, I'd love to hear about that. So in the graph-based indexes, we talked about how there's graph-based and cluster-based indexes. In the graph-based indexes, they're really nice because you can add things to it and it fits neatly into the graph and you can just keep adding to them. And you don't have to worry that much about recall because the recall on a graph-based index is usually phenomenal.
32:23That's why they've been so popular because you just add to them. But given that we chose a cluster-based index for the read reasons that I mentioned before, when you do a clustered index, it's suddenly, you know, someone could have started adding products in a category or a cluster that didn't really exist when you created the initial clusters. And generally, a clustered algorithm sort of has to look at all the data and then decide what the clusters are. But when you get into millions or hundreds of millions or even billions of vectors, that can take easily days. Like on a state-of-the-art algorithm, it can take days on a very large machine to figure out what the clusters are.
32:56So you start using GPUs to try to do it faster, but it becomes very, very difficult to do. So basically these clusters can't be incrementally maintained, right? In a B tree, in a traditional database index, it's really nice because you just sort of like insert them into the tree and it nicely balances. And there's lots of properties that make sure you can just incrementally add to it like a graph. So in a clustered index to maintain it incrementally, you need a bunch of heuristics to maintain the clusters. You can think about it as like maybe you know if we continue with the e-commerce example you have a shoe cluster and this customer is just like they just keep adding shoes and at some point the shoe cluster has maybe a thousand items in it and it's like okay well that's a lot of data to search every time we search for something relevant to shoes so we have to split the cluster so at some point you know every time you write it's like well this this is a shoe this is a shoe and then it reaches some terminal size of the cluster and we split the cluster so you split the cluster and then maybe there's a sneaker cluster and there is a leather shoe cluster and that's what it's decided to make the two clusters and you can imagine incrementally maintaining the clusters like this where they reach some size we split them once in a while you remove enough items that we have to merge clusters and every time we do this we have to recompute the centroids of the clusters it is much more complicated than that but that is the general idea behind spfresh that with enough of these heuristics you can do this this is probably the most complicated part of the entire terrible buffer code base to make this work at very, very large scale and make it work with recall and make it work with filters.
34:24That is how SPFresh works. And this scales very, very well. Building agentic AI apps isn't just about choosing the best LLM. Agents need short-term memory, long-term recall, and lightning-fast retrieval. Without it, you're left with clunky prototypes that never scale. You know Redis, the world's fastest caching solution? It turns out fast data is the key to good context. and good context is essential for fast, accurate memory. It's what makes AI agents actually work with your data. Redis for AI. The right infrastructure, the right tools, the only way to scale. Learn more at redis.io slash genai.
35:03Building an app often feels like a balancing act. You want to ship features fast. Chat, activity feeds, moderation, video. But building them from scratch is slow and complex. That's where Stream comes in. Stream provides developer-friendly APIs that let you add real-time communication without reinventing the wheel. Why Stream? First, developer experience. Stream has clean, open-source SDKs and great docs, and you can get a proof-of-concept running in hours. Second, speed to build. Experiment, prototype, or launch features quickly. No credit card required to start. Stream also scales. Over 2 ,000 global apps, including Strava, Patreon, Nextdoor, Robinhood, and Peloton, rely on Stream to power in-app communication for more than a billion users.
35:55Whether you're a startup or an enterprise, Stream handles the hard parts so you can focus on what makes your app unique. Get started today at getstream.io slash podcast. and moving kind of to scaling generally so my experience with just this space i know that the people approaching it are doing incredibly different ways is in the email space and for us namespacing was kind of a like a hard requirement so let's take that example for a second you know we've been talking about e-commerce which i'm also quite familiar with but let's hard life into email for a second so email you know you've got however many thousands of users and you absolutely never want someone to be searching and someone else's email turns up in that search that would be sort of catastrophic so namespacing i.e the idea that every user's data is in a very siloed sort of environment was incredibly important to us i'm aware that namespacing is a part of turbo buffer and scaling that has i believe it's like 14 million namespaces or something so talk to us about that because i mean oh sorry 40 million 40 million namespaces and i believe that's probably Probably also something that, again, maybe your Shopify experience in terms of that was a huge scaling challenge, I imagine, led into kind of how you've been able to approach this as well.
37:13Yeah, look, the only way you can shard anything or you can scale anything is to shard it. And so we decided at TurboPuffer to make namespacing a core sharding primitive that we exposed to the user. Because if you can give us your sharding key in a simple way, well, then we can scale really, really far with you. In TurboPuffer, a namespace matched directly to one shard. and a namespace is just a prefix on obik storage so you can imagine that if you have gregor dash email well that's prefix number one right and it's literally just like you know that on s3 gregor dash email slash right and then all your files are in there then you have simon dash email slash and so on so on so on and that way we are only constrained in our scalability on how many namespaces S3 can have?
37:59Well, S3 can have a lot of namespaces. We have yet to see any limit and there's no documented limit on how big this can be. And TurboPuffer, yeah, has more than 100 million of these namespaces. And this works great because it also means that we can encrypt every individual namespace differently because you might decide that you want to encrypt the Gregor email namespace or prefix on S3 with a key that you have access to. And I want to do that with mine as well to add even another layer of defense that's basically as good of a protection as if you had that data in your own bucket. There's no logical difference.
38:31You can rotate the key at any time or revoke it. So that namespacing is a core tenant of Turbo Puffer. And it's a core tenant of, you know, like I'm thinking back to my Shopify experience, right? And most of the problems we solved by abusing the wicked shard on the shop and that shops did naturally have ways that they to talk to other shops other than in some edge cases so that's exactly what we did at turbo puffer and build that in as a core primitive in the future probably a namespace will map to multiple shards and things like that but it will be abstracted against this like single layer of completely horizontally scalable search nice yeah as i say there was as you got a core primitive that i was looking for surprisingly let's just say other people in the space take slightly odd approaches to it or virtually no approach at all so i think it's something that you know the audience should look out for when thinking about this kind of thing so we're going to take a bit of a turn into who's actually using this because there's some huge names using turbo puffer notion and cursor are probably i imagine two of the biggest names if you like but also two of the names i'm i think most of our audience will be very familiar with and can probably start to imagine sort of how from everything we've just been talking about how the technology actually works under the hood like how that's being translated through to what does that even mean for them as a user each day?
39:46So maybe could we talk about both of those cases? I think super fascinating to hear about those. Yeah, we can start with Cursor. So I got to know the Cursor team in 2023. And when they saw the first announcement of TurboPuffer, knowing the team so well now, it's very clear that they just had a conversation around the dinner table at some point thinking, well, why hasn't anyone built it like this, right? They're big fans of S3, just like I am. So it slotted right into their mental model. And we went back and forth with a bunch of bullet point lists on email and spend some time with the team. And we were just completely aligned on what needed to be built here.
40:19There was a couple of missing features. So we built that out for them and then they migrated. For them, the storage architecture of TurboPuffer just made sense, right? When you open a code base in Cursor, it indexes the code base, right? And one of the parts of the indexing is to embed the code base. And this powers a lot of the agents and a lot of the functionality inside of Cursor then has the ability to do a semantic search on the code base. I use this all the time to be like, where does it do this? And where does it do that? And I can just do that in plain language. So Cursor, yeah, will embed with their own embedding models that they've trained, will embed the entire code base and then use that as a tool call into everything else.
40:55And for them, the storage architecture of having everything in memory, which was the previous solution that they were on, just didn't make a lot of sense. You don't need every code base ever opening Cursor in memory at all times. That's incredibly expensive. and the per-user economics of that just weren't sustainable for them. So they needed a way to bring the per-user cost to something that made sense to them. And so the TurboPuffer model of inflate the puffer fish when you're querying the code base just made a lot of sense, right? Only some percentage of the code bases are going to be at active at once, but they're still valuable to keep around.
41:26So it's just such a clear fit to the storage architecture. So they've been great. And their first bill was reduced by 95 % to move on to TurboPuffer. And they've been amazing partners that have inspired a lot of how we built TurboPuffer since then. Yeah, Notion is another customer of ours. And they also, it's a very similar story where they were using a vector database that also had things in memory. And the per-user economics, again, just didn't make a lot of sense to them. Again, only a percentage of these workspaces are active at once. So that really appealed to them about TurboPuffer as they wanted to connect a lot of this data into their LLMs.
42:01So it's a very similar story to the Cursor One. But I don't want to give the impression that TurboPuffer is only good for these use cases. I mean, it's still fundamentally the cheapest way to store data in the cloud is to put it in object storage at$0.02 per gigabyte and then cache it on a compute node only when it's actually being used. Because generally, in a traditional storage architecture, you have to replicate the data to three nodes, right? So in TurboPuffer's case, you have it in$0.02 a gigabyte of object storage, and then you store it on disk and memory on a blended cost of maybe around$0.10 per gigabyte.
42:32but on a traditional storage architecture with memory and disk and all of that you easily run into dollars per gigabyte so this allows us to have this pricing that just makes a lot of sense and it allows us to have a truly serverless environment where you can just like push in the vectors and this was great for both notion and cursor because they have a lot of rights right every time you edit files every time you edit in notion it has to issue a right to update the vectors in turbo buffer yeah that's interesting i have noticed just how fast cursor is indexing you know i love using cursor workspaces so you know like dumping say like three different code bases into a workspace and then having it kind of figure things out across all three and yeah the indexing has always impressed me so now i know why i mean obviously credit to cursor as well but it's great to understand the technology behind the scenes and yeah i was at a notion event not so long ago and obviously they're doing a big push on their ai offering and a lot of this is look we're going to index across all these data types you know like google drive and i believe gmail is coming soon and slack and well probably slack i mean there was that announcement about the api restrictions but i'm sure maybe they're figuring that one out but at the end of the day it's a lot of data that needs to be indexed and then effectively ai rag search over that so that's also super interesting that turbo popper is kind of behind the scenes on that one yeah i think every company has a lot of data that they want to make it into context.
43:53And if you have a lot of data, then the economics of adopting another data store rather than perhaps using a vector index in your relational database could start to make sense when you get into the tens of millions of documents that need to be indexed, right? Same reason why we've always moved search workloads out after some point in time for full-text search as well. After a certain threshold, the downsides of having another data store start to make sense. So when you want to connect a lot of data to AI, we think the storage architecture makes a lot of sense. You touched on it, just talking through the cursor and notion examples pricing.
44:23I think it's good to call out Turbo Puffer doesn't offer any kind of free plan. And I think you've been quite sort of vocal about why you don't offer one. I think it'd be great to just to kind of hear the why's behind that. I fully support that just to be clear. But I think it's also great for developers to sort of understand as much for their own products as it is for the fact that they might want to come now try Turbo Puffer and then might be disappointed that there isn't a free plan, for example. Yeah, I mean, first of all, if you're not having a good time in the first 30 days and you've tried the product earnestly, then you have the right to cancel.
44:51But we think that the companies that benefit the most from Turbo Puffer, this is not a price point that's scary. And for small companies, it might make sense to use a small index that's tacked onto the existing relational database. But if you have ambition to index tens of millions of vectors, it can make sense to start on Turbo Puffer. But in a lot of cases, it's really about wanting to provide people a really good experience. And it's easier to provide people a really good experience when you have a commercial only offering and we can put behind the necessary staff to support people when they have questions and give them a really good experience so that you get the feeling that we're part of your team absolutely and you've announced i believe quite recently you're now ga so general availability how does turbo puffer sort of as an org now look i haven't done research and i always find it interesting are you two people are you 200 i mean where is turbo puffer yeah i mean we're less than 20 people and it's a very engineering heavy team.
45:43We've hired a bunch of database engineers, and that's the majority of the team, people working on the database. And now we're starting to hire other types of roles to expand the team and especially support our customers and so on. Yeah, that's the size of the company. We really want to have a dense P99 engineering environment. And so we try to hold the standards high on a team that we put in front of our customers and are developing this product. right this is a very nice number team-wise you know i think it's been a little bit over reported on sort of this idea of the you know the one person unicorn and all this nonsense so i think it's good to sort of call out that some of the the best companies coming out at the moment are maybe 10 but 20 also sounds good and a bit more and like from what you can share where's turbo puffer kind of going you know over the next say six months to a year like what's kind of top of mind for where the product needs to go one of the biggest frontiers we're pushing right now is just more database features.
46:39So for example, we just launched like conditional writes. So it's the ability to say, Hey, I only want to update this document if it's actually newer than the document that's in the database. So it's database features. It's a lot on the full-text search side. So Turbo Puffer doesn't just do vector search. It also does full-text search. And we see more and more of our customers doing that. There's a lot of expectations about what you can do in full-text search. And so we have to build that on the storage engine that we built from the ground up, right, to work. So those are the biggest things that we have.
47:04We have a very solid generic LSM storage engine and underpinning all of this that we've matured and now it's really about features and it's about it's about even more performance than we deliver now awesome well we're coming up for time one final question i had was just actually completely nothing to do with databases specifically it's actually around i mean you've talked about the name but also the visual language and just i guess the front-facing site very fun it's sort of pixel art i guess if you want to kind of give it some kind of name where did kind of the idea come from and like who kind of leads that on the team yeah i think it's like neo retro pixel art or something i don't know what the aesthetic is called yeah really this came from when i was back home in denmark and we had some some friends visiting and i was talking to to my friend who's a designer about this idea of of turbo puffer and i think he said later that he had no idea what i was talking about but i sounded really excited about it and he was so encouraging and in my mind lateral to the exactly how to lay out the bytes on disk and build a storage engine and index was sort of this just this aesthetic that i wanted and so we were riffing back and forth on what that would look like and kind of creating some sample sites and it was just like just like very bare bones right like i've you know at shopify i spent 10 years evaluating databases so i was on the buy side the majority of my career and every time i went to a website it's just like what does it cost what are the trade-offs who are the customers and what they're using for is these are the only questions i really care about and most database websites are just so filled with so many different other things right and i'm just like no what are the trade-offs like any database has trade-offs like is this the right set of trade-offs for me what's the architecture what are the guarantees what's the consistency model and so this show don't tell just was very important to me and then we wanted to breathe a little bit of fun into it i think turbo puffer in itself is kind of a whimsical name and so this aesthetic just sort of appeared over many many iterations that you'll see on internet archive and we're just really fond of it and will continue to iterate on it because it makes us happy when we go to the website yeah as you've called out if you go on turbo puffer and you'll see a whole bunch of like sliders etc so you get sort of into the meat of it very quickly but i think as you call out having this kind of fun aesthetic around it definitely sets the tone for who you are as a company as well reminds me a little bit in different form of you're probably familiar with tiger beetle and you know financial database and so they have a very very specific type of database but they also bring fun visuals to it which just and obviously the name so it's very memorable to that point i remember seeing something about turbo puffer and hacker news a little while ago and then this episode was suggested i knew the name straight away could visualize the website before it even come back to it so i think it's very smart so yeah thanks so much for coming on where can people find you or like where's the best place to go to kind of just get acquainted with turbo puffer yeah turbo puffer.com is a great answer point we post on all the social medias you'll find the link there.
50:00I'm Syrupsin on X where, you know, these days I mostly tweet when I go out for runs for some reason, but that's where you'll find me and find the company. Awesome. Well, thank you so much, Simon. Very deep technical episode today, which I think a lot of the audience will love and just really look forward to following along. Turbo Puffer, I think you guys are doing very interesting things. So thanks for coming on. Thank you so much for inviting me on.
50:29Thank you.
From the publisher
Vector search has become a foundational technology for AI applications, enabling everything from semantic code search to contextual retrieval for large language models. However, a major challenge with vector databases has been the cost as data storage scales. Turbopuffer is a vector database that focuses on speed, cost and scalability. It was created by Simon
The post Turbopuffer with Simon Hørup Eskildsen appeared first on Software Engineering Daily.
