In short
Eye On A.I. Podcast Episode #219 Summary
Episode Overview In this episode of Eye on A.I., host Craig S. Smith interviews Avthar Sewrathan, Lead Technical Product Marketing Manager at Timescale. The discussion focuses on the innovative ways that PostgreSQL (Postgres) is transforming AI development, particularly through Timescale's various tools and open-source projects.
Key Themes
- The evolution of Timescale from its IoT origins to AI applications.
- The role of Postgres as the go-to database for AI, including tools for vector management and search capabilities.
- The importance of open-source development in fostering innovation and collaboration in AI.
Sponsor The episode is sponsored by Netsuite by Oracle, a cloud financial system aimed at streamlining various business operations.
Guest Introduction
- Avthar Sewrathan introduced himself and explained Timescale's genesis in the field of IoT and data management. Timescale started as an IoT platform that evolved into a Postgres database company, particularly for time-series data.
Detailed Discussion Points
Timescale and PostgreSQL
- PostgreSQL's Unique Features:
- Reliability & Robustness: Over 30 years of operational history with many issues addressed.
- Extensibility: Allows developers to create extensions that enhance its functionality (e.g., TimescaleDB for time-series data, PGVector for vector data).
Open-Source Philosophy
- Timescale emphasizes an open-source approach:
- Core projects such as TimescaleDB are available under permissive licenses.
- Developed tools like PGVector and PGAI Vectorizer are open-source to encourage widespread adoption and innovation.
AI Applications
- Avthar discussed how Timescale's tools support various AI-driven use cases:
- Vectorization and Retrieval-Augmented Generation (RAG).
- Handling structured and unstructured data.
- Key Tools:
- PGVector: For managing vector data.
- PGVector Scale: Enhances scalability for high-performance workloads.
- PGAI Vectorizer: Automates embedding management and synchronization with source data.
Industry Impact
- Applications in various sectors, including finance, crypto, and IoT.
- Postgres serves as a unified solution to avoid the complexity of managing multiple databases.
Future Directions
- Avthar hints at the potential of integrating AI more closely with databases, possibly creating natural language interfaces for querying databases directly.
- The move towards simplifying data handling for developers to focus on building applications rather than managing infrastructure.
Key Takeaways
- Postgres is not just a database; it is evolving to handle modern AI workloads.
- Open-source tools are crucial for fostering innovation and avoiding vendor lock-in.
- By simplifying the complexities of data management, Timescale aims to empower developers in building robust AI applications.
Closing Remarks Avthar emphasized the importance of open-source in AI development and Timescale’s commitment to making powerful database tools accessible for developers of all skill levels.
---
Feel free to subscribe to Eye On A.I. for more insights into the evolving world of artificial intelligence!
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00I think what has really helped Postgres succeed is the reliability and robustness. I think having been around for 30 years, all the problems about the system have been known, and most of them have been fixed. And I think another defining characteristic of the Postgres ecosystem and Postgres database is its extensibility. So it has this notion of extensions where, you know, even if you're not part of the core Postgres team or a certain functionality that's not part of core database functionality, independent developers and companies can build these things called extensions, which essentially hook into the main Postgres planner and Postgres database.
0:37And that allows you to do all kinds of things beyond just kind of transactional workloads. What does the future hold for business? Ask nine experts and get 10 answers. Bull market, bear market, rates rising or falling, inflation going up or down. Can somebody please invent a crystal ball? Until then, over 40 ,000 enterprises have future-proofed their business with NetSuite by Oracle, the number one cloud ERP, bringing accounting, financial management, inventory, HR into one fluid platform. With one unified business management suite, there's one source of truth, giving you the visibility and control you need to make quick decisions.
1:24With real-time insights and forecasting, you're peering into the future with actionable data. If I were a larger organization, this is the product I'd use. Whether your company is earning millions or even hundreds of millions, NetSuite helps you respond to immediate challenges and seize your biggest opportunities. Speaking of opportunities, download the CFO's Guide to AI and machine learning at netsuite.com slash ionai. That's netsuite, N-E-T-S-U-I-E dot com slash ionai, E-Y-O-N-A-I all run together to get the CFO's guide to AI and machine learning. The guide is free to you at netsuite.com slash ionai.
2:20Great. So, Avatar, it's great to have you. Can you start by introducing yourself to listeners and tell us a little bit about Timescale, and then we'll get into some of Timescale's AI offerings? Fantastic. Thanks so much, Craig, for having me on. My name is Aftar. I'm the head of AI at Timescale. Timescale is really a Postgres database company. So the company started off in the time series space. It actually started before we got into databases. It was an IoT platform. And then in building that IoT platform, we noticed how difficult it is to build applications that require like time series and real-time analytics.
3:03And we couldn't find a suitable database. And so the team at the time decided, you know, why don't we just build our own? And that's where Timescale and TimescaleDB came from. TimescaleDB really extended Postgres, which is a well-known, well-loved relational database platform for time series and analytics without having developers use a separate database. And so that's really where the company came from and where we got our start back in 2015 or so. But since then, the company has really evolved outside time series to really extending Postgres for all these different use cases beyond relational data.
3:44And most recently in the team that I work on, we focus on AI and vector data and helping developers build these AI systems that every company seems to be working on these days, things like Rack, things like Semantic Search, and helping them build those just using Postgres as their vector database and as their database to power those applications. So that's a bit about the company. To summarize, we're really a Postgres database company. We have a lot of open source projects as well. And we also have a managed cloud, which is our primary commercial offering. I think open source is really at the heart of the company.
4:22And that's actually a thread through a lot of the work that we're doing in AI as well. But yeah, that's an overview of timescale and how we got into the AI space. Yeah. And just for listeners, Postgres was an open source project that came out of Berkeley, I think is that right? They had a project called Ingress, and then Postgres is the follow-on to Ingress. Can you just give listeners who are not database people a little thumbnail on Postgres SQL? Yeah, Postgres, it's a database that's been around for more than 30 years at this point. And it was seen as, I think, kind of the defining characteristics is it's open source.
5:16So it was seen as a great alternative to kind of more proprietary systems of those days. And I think what has really helped Postgres succeed is the reliability and robustness. I think having been around for 30 years, all the problems about the system have been known and most of them have been fixed. And I think another defining characteristic of the Postgres ecosystem and Postgres database is extensibility. So it has this notion of extensions where even if you're not part of the core Postgres team or certain functionality that's not part of core database functionality, independent developers and companies can build these things called extensions, which essentially hook into the main Postgres planner and Postgres database.
6:06and that allows you to do all kinds of things beyond just kind of transactional workloads that we're used to for SQL databases. There's extensions for geospatial data called PostGIS. There's extensions, as I mentioned, the ones that we've built at Timescale for time series analytics. There's PG vector for vector data handling. There's even extensions for document data handling. There's extensions for kind of everything under the sun. And so that's kind of one of the defining characteristics because that combination of this really solid open source core plus the different extensions in the ecosystem, that makes it attractive for a lot of developers out there and makes it a really compelling choice.
6:49I think in the past two years in the Stack Overflow survey, Postgres was the most loved and most used database among professional developers. And so that's a testament to kind of how popular it is and just what a great piece of technology it is. Yeah. But the extensions that Timescale has built are proprietary. Is that right? Or are they open source? Yeah. So it's actually not proprietary. So let me just take you through the list. So the extension that we started off with around time series, that's called TimescaleDB. That's actually, it has, there's two versions of it. One is open source under the Apache 2 license.
7:32And so that's completely open source. And there's another version that is under what's called the timescale license, which has some basically the complete feature set. And that is actually free to use. And it's not technically open source, but it's free to use for everyone, except if you're going to offer it as a service. So that's just some protections. You know, we see certain cloud providers. There's been a lot of debate in the open source community about, you know, what is the correct way to monetize. And so for that particular extension, we've chosen to have this really freely, it's quite a permissive license for the everyday developer.
8:12But unless you're kind of like a big cloud provider, that's really only when you're going to run into these kinds of constraints. And then for the AI extensions that we built, we built PG Vector Scale, which helps speed up and extend Postgres for these kinds of high-performance vector workloads, think like Boolean scale search. And then PGAI, which really is bringing AI workloads to the database. Those two are actually open source under the Postgres license. And so that's completely free to use. People can use it on any database of their choice. and it's an extremely permissive license. And so embracing open source is really at the core of our company.
8:53And we think that, you know, that's a defining characteristic of the Postgres ecosystem as well as like what's made us successful thus far. Yeah, although I have to ask, how then do you guys make money if everything is open source or free? Yeah, that's a good question. We get that all the time. The way that we do it, I think, is the primary way, as I mentioned, before we have a managed product. It's called Timescale Cloud. It's basically a database as a service or a database platform as a service. And so development teams who would rather not take care of things like managing a database, dealing with backups, dealing with replicas, and essentially all the hassle of what it means to do database maintenance and build an application in the modern era.
9:41We provide that product where we take care of a lot of that and taking that off developers' plates. So that's the primary way. And we think this balance of having open source products where any developer, no matter the budget, no matter your constraints, you can actually benefit from the software that we put out with also having a really opinionated, really compelling cloud and managed experience. That's worked well for us so far. We have thousands of customers across both kind of really innovative startups as well as folks in industry and like Fortune 500. And so that combination has really worked well for us and something we're going to be sticking with going forward.
10:22Yeah. And the cloud offering, does it allow people to build an application? And then if they want to migrate it to their own cloud or their own on-prem or whatever it is, they can do it. I mean, does it serve as kind of a sandbox as well? Yeah. So I think most of the time what we see is that developers are looking for a production, something to solve their use case where their application is moving into production. where then you need to take into account of things like high availability, you need to do read replicas, you need to have your backup and you need to have a backup strategy in place.
11:10You need to deal with like automatic failover and stuff like that. That's really where the cloud product shines. And I think just to be transparent, it's all just Postgres. And so what we pride ourselves on is to say, hey, if you're worried about vendor lock-in, you're not going to face that because it's just Postgres. You can take it anywhere. You can self-host it. You can move it to another cloud provider. But what we find in practice is that the kind of ease of use functionality and the functionality for production applications that I've talked about is compelling enough combined with, you know, we obviously have support and a really kind of really strong solutions architecture front that many customers are, And some customers might say, hey, we're just going to use this temporarily and they end up staying because it's just easier for them and it's less things for them to worry about.
12:04Yeah. Before we talk about the AI products and particularly the newest offering, can you take us back to the IoT time series issue that you were solving in the beginning? So you've got an IoT advice and it's producing a stream of data. And I'm imagining the issue is how do you get that into a structured database? Yeah, so I actually wasn't at the company at the time, but I'll do my best to kind of recount the story as I know it. So basically what was happening is that the team at the time was building, I think it was called IOBeam. It's basically an IoT platform for managing a whole bunch of sensors and devices.
12:54And so this is actually a use case we see a lot today where, let's say you have a smart factory. This is very common in manufacturing. Or if you have like a fleet of EVs for like logistics or in the automotive space. And you want to know essentially the latest about what's going on across your entire fleet of sensors. You want to encounter, you know, the latest status of it. you want to see if there's updates, if you want to see if there's problems. This is kind of like the platform that the team was building at the time. And in building that, they needed a way to scalably store and query time series data.
13:33And there's a couple of defining characteristics of this time series and analytics data. I think the first one is that it's extremely frequent in the sense that these sensors send readings every second, sometimes multiple times a second. And so this data piles up a lot. And so you're dealing with a large volume of data when you're ingesting. And then also when you're querying, you often are querying a lot of data, but you have these very specific query patterns where you either want to know, okay, for a bunch of devices in my fleet, just give me everything in the past five minutes, just as like a current kind of like status updates, these kind of like real-time updates, or for a few, like let's say you're looking at a specific location or you're looking at a specific factory ID, for the segment of devices, I want to go deep and I want to explore, you know, the monthly performance because I'm doing kind of historical analysis for a report or something like that.
14:33And finding a database that could keep up with those requirements of like you really need perform an ingest, you really need fast time-based queries. There's also these aggregations that happen where you want to know, okay, what is my monthly average utilization? What is kind of the seven-day average or the seven-day maximum of some particular metric? It turns out that that was really difficult to service on the existing time series databases out there. And so what the team did was ended up extending Postgres, which they had existing expertise with, existing familiarity with, for time series. And this is a story we hear a lot from a lot of current timescale customers where they start off with a certain specialized database, they realize like either the performance characteristics are not good, or it's just a pain to handle and maintain.
15:27and they look for the product that's familiar to them that their team might already have existing expertise on, which is Postgres. And funnily enough, we're seeing kind of a similar story play out in the AI space where developers are, you know, starting off with these different vector databases, but as they move into production, they're looking to consolidate, they're looking to go to something that's familiar with them. And that's where Postgres once again has a role to play. Yeah. And on time series, I immediately think of the financial markets. Can you, does timescale handle that kind of volume and speed?
16:17Do you have clients in the financial stock trading business? business. Yes. So in time series, finance and most recently crypto, which is just really a different version of traditional finance, is a huge industry for us. As I mentioned, in other in the IoT space, there's like manufacturing, there's also automotive, electric vehicles, logistics. But for finance, I think like that's, you know, we power some of the biggest crypto exchanges out there. I'm just trying to remember what I can actually disclose publicly, but there's a lot of, I think, you know, time series is so synonymous with finance.
17:02And so that's really where we see a lot of usage and especially for being part of data infrastructure that can support applications being built. That is a huge industry for us. Oh, okay. So then AI comes along, generative AI, and you have extended your offerings to include that. my understanding is uh i'm not going to be able to remember all the all the names of the products but there's postgres vectorize which in effect uh vectorizes uh data in the database uh so that people using rag for example don't have to have a separate vector database it can all be done within that database and then uh uh you you i'll let you go through them but the the one that kind of fascinated me i i think it's pgai uh where you actually have uh from what i understand an llm built into the database.
18:22Is that right? So that you can, when you query the database, it can produce a natural language response directly. But anyway, why don't you back up and start talking about when AI became a focus and then give us a chronology of the products that came out of that. And then the most recent one, I think it was just announced. And I'd have to look at my notes, but I'll let you explain it. Fantastic. Yeah, I think you hit on some of the big ideas. Maybe I can give a bit more kind of structure to what you introduced. So I think to set the stage, what we saw is that when, I think it all started with ChatGPT coming out in late 2022.
19:19The world was really taken by storm with what was possible with LLMs. And I don't need to make this argument. I think it's kind of, if you're listening to this podcast, you kind of understand the potential of AI and, you know, all the different cool things they can do. What people immediately wanted to do back then was build what I would call these chat with your data applications. It's basically, you know, having a chatty bt style interface, but with your own data. Now we know this as RAG or retrieval augmented generation. Now the question is, if you're trying to build these RAG or even semantic search, or most recently AI agents is kind of the main thing that folks are experimenting with.
20:00You need a place to store these knowledge bases that you want to augment the LLM's kind of training data against. Some implementations involve using search APIs to augment the LLM. In other cases, where you're dealing with private documents, if you want to really ground the LLM in certain documents for it to respond with knowledge of, using a knowledge base is the main approach there. That brought this question of, okay, I know what I want to do. I need to have a knowledge base for my RAG application, how can I easily build this? And so this is where we face a similar, you know, once again, I think it's kind of like a rhyming of history here where developers were looking at what choices of data infrastructure they have to build these applications.
20:50And they ended up with either to use specialized vector databases, which are really built for vector search, but they come with the additional headache of having to essentially add them into your existing infrastructure, which is fine in the beginning if you're just like playing around. But when you're building a production application, you have to deal with all the things that I talked about earlier of like maintenance and management and backups and high availability and stuff like that. And that was really difficult. And then on the other hand, you had databases like Postgres who were adding on vector support via extensions or via kind of first class support in databases like MongoDB and MySQL most recently.
21:35And what actually played out is that it turns out that Postgres was very popular. There's an extension that's open source called PG Vector that really took off in popularity. And we initially thought that, hey, if we just support PG Vector, this is great. This should solve all our problems. It turns out that that wasn't the case. PG Vector is great by itself, and it provides a really robust foundation for vector handling and vector search in Postgres. But what we started to hear from the community and as different developers started exploring what's possible with these AI applications, a common theme of scale being an issue came up again and again.
22:13And basically the problem was that, hey, as you scale past a certain number of vectors in your knowledge base, you get really slow queries. And to put it kind of in easily understandable terms, what this basically means is that as you add more documents or as you add more resources that you want the AI to talk about, you end up with this performance bottleneck, which is not great because ideally you don't want to have a limit of like, hey, I can only have, you know, 1000 documents or something like that. If you're building a, let's say, an application that services like a multinational company, you're going to hit that limit pretty quickly.
22:47And so what we decided to do was to try and solve this scale problem. And that's where we built this extension called PG Vector Scale, which takes some of the latest research in approximate nearest neighbor search, which is kind of the academic literature around vector search and vector databases and implement it in Postgres. And so what we did is we released this PG vector scale extension earlier this year in June. And what it does is it brings some really advanced indexing to Postgres and also some advanced compression methods. It's called statistical binary quantization. And that's been a huge, like really well received by the community.
23:27And basically what that does is work hand in hand with PG vector in order to give Postgres these really high performance, high scale capabilities and removing that kind of performance ceiling that was holding developers back from like building these applications. Yeah. And just, I misspoke earlier, PG Vector is a third party extension. It's not a time scales extension. Is that right? Correct. Yeah. It's an open source extension. It's not something that we've built at our company. Right. But PG Vector Scale is your extension that works together with PG vector. Yeah. And PG vector, does that vectorize the database?
24:14Because you have a tool for the vectorization, right? Yes. Yeah. So this is actually gets into the second set of problems, which is, you know, once you've solved vector search, which is about the actual like, hey, find me similar documents or find me kind of the nearest neighbors in a graph, you get into all the other issues around building an AI application, which gets to actually really tough engineering challenges. And so that's where the rest of the PGAI project comes in, which is a PGAI extension and the newest product that we just released last week, which is called PGAI Vectorizer. So the problem that these two tools solve is the problem of managing embeddings.
25:00So all these vector databases are simply a component in a stack or in an architecture where they expect vectors to be written in and then they can do search, then they output vectors which are then used in your application however you want. Now, what we saw is that developers were having a really tough time, you know, once you go beyond a POC to not just create embeddings, but to keep them in sync with underlying data. Now imagine an easy way to see why this is important is to imagine a, again, you have a company chatbot that is going to power and it's going to help your employees get the latest information about products that the company offers.
25:40And let's say you make a big update to the product where a sudden restriction has been removed or the product feature functions in a new way. You're going to want to change your documents in the knowledge base, and you're going to want your vectors to be updated as a result. But what happens with most vector databases today is that the vector database only handles kind of the vector embedding storage, and that it's on the development team and on the engineering team to actually go and update the documents and separately update the embeddings in the vector database. So there's all this synchronization, all this kind of engineering that needs to happen in order to keep your vector embeddings in sync.
26:27And that's really what PGAI Vectorizer does. It both creates embeddings, but it also ensures that as underlying data changes, whether it's additions, deletes, or even changes to the actual source data, your vectors will automatically be updated and it will reflect the latest state of the source data so that it prevents stale data in your applications. And this is not just important to keep your applications up to date, but it actually enables a whole bunch of easier experimentation and testing. And if you think about like in an AI application, the way to make it better is by testing and trying out different, whether it's new models, whether it's new kind of chunking formats, etc.
27:14because that's how you find kind of what works specifically for your application. And so what PGAI Vectorizer does is it makes it easy to create copies of your data with different embedding models, with different formatting, with different chunking specifications, with just a single SQL query. So what used to take engineering teams to say, hey, we want to test this model. Let me take a week or two to go and like set this up. Let me make sure we run all our documents in pipeline. Let me make sure that we, you know, all the updates that we made in order to get like a fair test. Let me reflect that.
27:48Now it's just a single SQL query and the system will take care of creating these documents and creating these vector embeddings for testing and for experimentation. So the whole suite, essentially, if I had to put these, all these things together, is making Postgres a better system and a better database for AI applications. and you have the performance and search component, which is PG vector and PG vector scale. You have the embedding management and embedding creation, which is PGAI vectorizer. And then as you mentioned, the PGAI extension has some of these functionality where you can actually call LLMs from the database.
28:28And that's useful. It might not be useful in every single application, but if you want to do things like in database transformation of data, where let's say you have, as your data is coming in, you want to do moderation or you want to do summarization and you don't want to do it in your application code and then write it into the database, you can actually do that within the database. And we interact with third-party models like OpenAI, like Coher, like Anthropic, and also open source models via OLAMA in order to make that happen. So this is really a cohesive set of tools for developers that are building AI applications so that they can use the Postgres that they know and love.
29:06Maybe they deploy it to their companies without having to spend time learning a new system and also spend time like maintaining and figuring out how to make these multiple databases work and play nice with each other. Yeah, I have a couple of questions here. One, I just want to say that I tend, because I talk to so many people about similar things, I forget that a lot of the listeners are not experts. So anyone listening, vectorization or a vector database is converting data into vectors, which are strings of numbers that are more easily searched. Is that a fair enough way to say that? Yeah, I think that hits on definitely the big ideas.
30:02I think one way that I found it easy to understand is what a vector is, it's essentially a compressed representation of some sort of data. So you might have text data, you might have image data, you might have a combination of text and images. and what it is is that it takes, let's say, a sentence or it takes a document and it compresses it into exactly what you said, a list of numbers, an area of numbers. And essentially what that is, it's a representation in, I guess, it's called high dimensional space. But when you have a lot of these vectors, let's say you have a thousand documents and you vectorize all of them, you can begin to see similarities and you can begin to see connections between them that it may be difficult to find by just looking at the actual source or the vectors itself.
Read the full transcript
30:52So this is the basis for semantic search, which is search by meaning, where you can do searching documents by the underlying meaning rather than actual keyword search. And then also for images, this is actually a very popular use case where you're doing object detection or you're doing image similarity where you don't actually have to like necessarily tag these images with like annotations and descriptions. The actual model captures the semantic essence of the image itself. So, you know, to bring it back, vectorization is just a process of turning whatever your source document or the text or image or something else into a vector that can then be used for this search in high dimensional space, which is super useful in a variety of applications.
31:40As I mentioned, it's kind of the core foundation for RAG and, you know, semantic search, recommendation systems, search engines, image search, that whole bunch of applications and like object detection and stuff like that. And so, yeah, it's super useful and it's kind of what's helping LLMs interact with data that is not part of its initial training corpus. Yeah. And then this being able to call an LLM from within the database. Can you talk about the plumbing there? Is that just a simple API call with a chat interface for users or how does that work? Yeah. And so this one is really, this is one of the more experimental functionalities we have out there.
32:32And it's basically saying, if you want to bring LLMs closer to your data, and if you have workflows where the source of what you're doing is in the database, and the end result is also going to end up in the database, this makes those kinds of workflows really easy. So I gave you an example earlier of one of our customers who actually does a moderation workflow in the database where, you know, they manage like a blogging site or a forum. And every comment that gets written in order to decide whether it actually shows up on the site or not, it gets moderated via an LLM. And all they do is that, you know, all the comments get inserted into a Postgres table.
33:15they have a trigger that when a new row gets inserted, perform this moderation. And then instead of having to do this in the application layer, they're storing it in the database anyway. Let's just run the moderation and call it from the database. And then we define our prompt. It makes an open AI call. And then you get a yes or no, whether this comment is flagged or not. And then that changes which table this then goes into. And it really makes it a lot easier for that particular developer to implement this versus having to store it in the database and then do the moderation on the application layer and then write it back to the database.
33:50This can all be kind of simplified. Another use case is like, let's say you're trying to do summarization and you have like a thousand rows or like hundreds of thousands of rows that you're trying to summarize. This kind of batch transformations where let's say you have a bunch of text or you have like a customer database or something like that and you're trying to like classify them or categorize them, do any sort of like LLM reasoning. This is really useful because then we can just have basically create a new column to say, hey, this is like a batch transform of data that's in column A and, you know, categorize it in one of these three categories.
34:26And then you'll get, you know, what the LLM thinks is a category that these one fits into. So once again, like that is really good for using LLMs within the database for workflows where the source of your data is in the database, as well as the end result is also going to get stored in the database to go forward. Yeah, that's fascinating. And it sounds to me like the next step would be to create a chatbot directly in the database rather than having to build a chatbot that then is using a database as the knowledge base. Is that possible that, you know, that then users, customers, whoever could be talking directly to the database?
35:25Yeah, I think this is where, you know, I think a lot of, there's a lot of exciting innovation going on in the space. And I think that's definitely one of the possibilities. I think one potential future that we could end up with is where you have these natural language interfaces over databases where, you know, to answer questions instead of necessarily writing like long SQL queries, you delegate that to the LLM to do this. There's an area of research called Text to SQL, which is about this. I think what we're looking at is rather than being very rigid in saying like, hey, we want everything to happen in the database, what we're trying to do with the PGAI project at timescale is build a set of tools for developers to use as a building AI application, where if they want to substitute, let's say, using a bunch of other tools and stringing them together and just have that complexity in the database, they can make the choice to do that.
36:28And so there's some applications where, as you described, maybe you don't need to really build an application-level chatbot. Maybe all your RAG and all your LLM calls are handled in the database. That might be good, especially if someone's getting started. Maybe they have a very scoped use case. But what we're trying to do is give developers the option where it makes sense for them. Like, for example, the PGAI Vectorizer is a good example of this, where a lot of the times what we found is that developers, when they're building these AI applications. And in order to create vector embeddings of data, you'd have to have like an ETL pipeline.
37:04You'd have to have a data sync service. You'd have to have queuing systems because you deal with API rate limits. You also have to have monitoring tools to catch stale embeddings, alerting systems, and also a bunch of custom code. And so we said, okay, that might work if you're at the scale of, at a very large company, let's say you're on Apple or Facebook or something like that, where you can have these hundreds of person data teams whose job is just to keep track of it. But if you're a startup or if you're a team within a Fortune 500 where you're resource constrained and you want to build an AI application that is kind of tolerant, not tolerant, but you want to in an AI application where you don't have stale embeddings and you can actually keep up with the demands of production, where it said, okay, PGI vectorizer is a great substitute for all that complexity.
38:01And so really giving developers this option to say, hey, where you want to substitute complexity for simplicity and Postgres, these tools really allow you to do that. Another example, as I mentioned earlier, is PG vector scale. Instead of managing a separate vector database, You can just use Postgres and use your existing database rather than having to manage a separate database and then deal with the problems of data synchronization, managing multiple systems, having multiple source of truth, you know, just eliminating that set of problems to begin with. Yeah. And I should have asked this earlier, but your databases handle both structured and unstructured data.
38:44We're not talking only about structured data. 100%. Yeah, 100%. I think, you know, bringing it back to something you talked about earlier in the podcast, the sheer variety of data types that Postgres supports is one of the reasons why it's so popular and why a lot of developers love it. And so you can deal with text, you can deal with JSON bees, you can deal with other kinds of unstructured data alongside your structured data in whether it's, you know, alongside your structured data in like tabular format. And so having that ability to store these two kinds of data in the same database is actually really interesting for AI systems because it gives you one place where, you know, if you hook up your database to an agent, for example, it makes it a lot easier because you have one environment for that agent to interact in.
39:35Yeah. And why then is not everybody using this? Why are people still stitching together, you know, Pinecone and, you know, and Llama 3 and, you know, putting it all. Why not just use, you know, Timescales Cloud with all of these tools within it? Yeah, I think we're actually seeing more and more developers, you know, as a meme going around of just use Postgres. And as you said, replacing the complexity of databases like Pinecone, like other specialized vector databases, with the simplicity of Postgres. I think what we're seeing, you know, some of the reactions that we got to the PGAI vectorizer announcement yesterday, which was like, hey, I can see all databases moving in this direction where, you know, the burden of embedding sync is not on the engineering team, but it's actually handled by the database.
40:42We think this is some work that is hopefully going to inspire, you know, I think for us, you know, it'll be great if developers use Postgres, it'll be great if they use our paid Timescale Cloud product. But for us, you know, we really just want developers to build better applications and to have a less stressful time and time where they can actually, you know, focus on self-expressing rather than like wrangling their different problems that come up with and like wrecking their brains with like, oh my gosh, I need to deal with this kind of really tricky data infrastructure problem. I think that what's interesting is that, you know, there's a growing movement among a certain set of developers.
41:27You can encapsulate it as like Postgres for everything, where folks are just like, hey, whatever you need, Postgres is probably going to solve your problems 90 % of the time. And in that other 10 % of use cases, maybe then you need something specialized. For example, you mentioned some of these other vector databases, maybe in the case of PGI vectorizer, maybe you do need to write your own thing. But the point that we're making is that like, that should be for the last 10 % of users or the last 5 % of users, rather than the bulk of folks who frankly have better things to do with their time, or they have more important priorities, but they're stuck, you know, dealing with all this complexity that's actually not part of their core application.
42:11Yeah. And you mentioned earlier startups that how small is too small to use timescales products? I mean, is there a certain scale that it makes sense to use it? And below that, there are simpler ways. Yeah. So I think this is actually a really good question. I think one thing I'll say is that, you know, we have literal like one person companies who use the Timescale card product. And we have everything from like a five person startup using our AI extensions and the PGAI suite of tools to, you know, thousands of people at a multinational company whose name obviously can't disclose because they, you know, have all kinds of like privacy.
43:07Sure. like all these procedures to go through if you want to name them. But what I wanted to just paint a picture of is that like, you know, while the name of the company is timescale and while the core time series product does have scale, I think the core theme is building for scale while also making it really easy to get started. And so, for example, like if you take the PGAI Vectorizer product, it's built for scale in the sense that you can process, I think it was like 20 ,000 embeddings per second. And so if you want to, you know, if you have like millions of embeddings that you want to process, we can handle that.
43:47But it also makes it such that, you know, when you're starting off, you can just process your embeddings and not have to worry about them being in sync so that you can focus on the 20 other things in your checklist as you're trying to get the product built. And so I think for us, that's kind of spectrum of use cases where I think, especially in AI, a lot of developers are worried about like, hey, what happens when my system hits a certain limit? You know, am I actually choosing the right kind of application for scale? And I think this kind of like future proofing is a big decision point. And, you know, one thing I'll emphasize is that like, you know, there's a lot of capabilities to make it easy to get started.
44:27And then when you hit that stage where scale and performance does become a concern, you already are on Postgres and you already are on the various extensions that we've built so that you can seamlessly deal with those problems without having to switch systems or without having to conduct a fire drill because you're worried about the impact on your users in production. Yeah. And your target market are developers, but I'm not a developer. I'm very interested in all of this stuff. So I'm particularly interested in no code layers of abstraction on top of a lot of this stuff. Do you see that happening with you guys so that people who are not database experts could build some of these or use some of these tools to build their own databases and applications on top of the databases?
45:32Yeah, so I think for us, I will say that Timescale is squarely focused on serving software developers. and I don't think that's going to change in the foreseeable future. I think one thing that's a trend that, you know, you mentioned things like no-code tools and stuff like that, a trend that we're seeing with AI is just making software development more accessible to folks. You know, there's so many tools out there. If you look at Cursor, if you look at Replit agents that have been out there, there's Devon, the AI software engineer. I think the barrier to being a developer and to being quote-unquote technical, Like, I really don't like that distinction personally, but to live in the world of code is a lot lower than it's lower than it ever has been, essentially.
46:19And I think that, you know, from my perspective, I think I'd love to see all these different no-code tools out there. And I hope that, you know, Postgres and PGAI and Timescale can actually power those developers who are using that. And I think that what we're seeing, for example, where we have a early access functionality in a database admin tool and data querying tool called Popsicle, that's part of the timescale portfolio of products. and what it has is this ability to, let's say you have a data in a database and you don't know how to write a SQL query, you can actually ask a question to say, hey, I want to know how can I get my sales for this month?
47:03And it will actually tell you the SQL query to run. And so just like, you know, that makes it a lot easier versus having to figure out like, oh my gosh, how do I write this? I don't actually understand the schema, et cetera. And so that really lowers the barriers to folks that want to use these more technical tools. But I do agree with you. And I think that for certain things, as much as I am a developer myself, for certain things, using these agents, using these no-code tools, it's becoming easier than ever to build something. And so that's really exciting because it means that you get more useful products in less time.
47:37And I think fundamentally, this is all a means to an end. And so it's really exciting to see how AI is kind of impacting this and making it a bit easier. Is there anything I haven't asked about that you'd like listeners to know? I think the final kind of message that I would bring up is just that power of open source and in everything that we do at Timescale and with AI. I think that what we're trying to do with all the different Postgres extensions and the AI products that we built is have open source at the foundation. And I think for a technology that's as important as AI that has, you know, there's all these debates going on about, you know, AGI and what to do when AGI goes rampant or whatever.
48:26But just speaking more practically, AI is going to transform every business and every product out there sooner or later. and having technologies that are open source, having technologies that people can build on, communities where communities can actually contribute, that kind of future is a lot more exciting to me and to a lot of developers than to rely on kind of proprietary products. And so that's where, you know, we think that open source has a huge role to play. That's kind of why we've made all our extensions open source and all our tools open source or have an open source version. And I think that, you know, that's kind of a bigger decision point today where developers do want more control.
49:09They want that assurance. They don't want to be vendor locked in or rug pulled. And so I think that that's, you know, one thing I'm very proud of. But also, I think, you know, more practically speaking, I think it's kind of the right way to do things in terms of ensuring that we're fostering innovation and doing so in an open and accessible way. What does the future hold for business? Ask nine experts and get 10 answers. Bull market, bear market, rates rising or falling, inflation going up or down. Can somebody please invent a crystal ball? Until then, over 40 ,000 enterprises have future-proofed their business with NetSuite by Oracle, the number one cloud ERP, bringing accounting, financial management, inventory HR into one fluid platform.
50:03With one unified business management suite, there's one source of truth, giving you the visibility and control you need to make quick decisions. With real-time insights and forecasting, you're peering into the future with actionable data. If I were a larger organization, this is the product I'd use. Whether your company is earning millions or even hundreds of millions, NetSuite helps you respond to immediate challenges and seize your biggest opportunities. Speaking of opportunities, download the CFO's Guide to AI and Machine Learning at netsuite.com slash ionai. That's netsuite, N-E-T-S-U-I-E dot com slash eye on AI, E-Y-O-N-A-I all run together to get the CFO's guide to AI and machine learning.
51:02The guide is free to you at netsuite.com slash eye on AI.
From the publisher
This episode is sponsored by Netsuite by Oracle, the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more.
NetSuite is offering a one-of-a-kind flexible financing program. Head to https://netsuite.com/EYEONAI to know more.
In this episode of the Eye on AI podcast, Avthar Sewrathan, Lead Technical Product Marketing Manager at Timescale joins Craig Smith to explore how Postgres is transforming AI development with cutting-edge tools and open-source innovation.
With its robust, extensible framework, Postgres has become the go-to database for AI applications, from semantic search to retrieval-augmented generation (RAG). Avthar takes us through Timescale's journey from its IoT origins to disrupting the way developers handle vector search, embedding management, and high-performance AI workloads—all within Postgres.
We dive into Timescale's tools like PGVector, PGVector Scale, and PGAI Vectorizer, uncovering how they aid developers to build AI-powered systems without the complexity of managing multiple databases. Avthar explains how Postgres seamlessly handles structured and unstructured data, making it the perfect foundation for next-gen AI applications.
Learn how Postgres supports AI-driven use cases across industries like IoT, finance, and crypto, and why its open-source ecosystem is key to fostering collaboration and innovation.
Tune in to discover how Postgres is redefining AI databases, why Timescale’s tools are a game-changer for developers, and what the future holds for AI innovation in the database space.
Don’t forget to like, subscribe, and hit the notification bell for more AI insights!
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Introduction to Avthar and Timescale
(02:35) The origins of Timescale and TimescaleDB
(05:06) What makes Postgres unique and reliable
(07:17) Open-source philosophy at Timescale
(12:04) Timescale's early focus on IoT and time series data
(16:17) Applications in finance, crypto, and IoT
(19:03) Postgres in AI: From RAG to semantic search
(22:00) Overcoming scalability challenges with PGVector Scale
(24:33) PGAI Vectorizer: Managing embeddings seamlessly
(28:09) The PGAI suite: Tools for AI developers
(30:33) Vectorization explained: Foundations of AI search
(32:24) LLM integration within Postgres
(35:26) Natural language interfaces and database workflows
(38:11) Structured and unstructured data in Postgres
(41:17) Postgres for everything: Simplifying complexity
(44:52) Timescale’s accessibility for startups and enterprises
(47:46) The power of open source in AI




