#153 Ed Anuff: Unpacking AI's Role in Data Management

12 Nov 2023 · 59 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Eye On A.I. - Episode #153 Ed Anuff: Unpacking AI's Role in Data Management

Episode Overview In episode #153 of *Eye on AI*, host Craig S. Smith interviews Ed Anuff, Chief Product Officer at DataStax. The episode focuses on the integration of AI with data management systems, particularly vector databases, and the implications of these technologies on business practices.

Key Sponsors

  • Celonis: A leader in process mining, offering AI solutions for improved business processes through enhanced data intelligence.

Key Themes and Discussions

  1. Introduction to Ed Anuff and DataStax
  2. Ed Anuff shares his background, including previous experience at Google.
  3. Discusses DataStax's role in the database industry, especially its popular NoSQL database, Cassandra.
  1. Importance of Vector Databases
  2. Definition: Vector databases store data in a way that facilitates efficient retrieval of high-dimensional data, which is crucial for AI applications.
  3. Impact on AI: The emergence of vector databases is driven by the need to address issues such as the "hallucination" problem in large language models (LLMs) like ChatGPT.
  1. Data Management Challenges
  2. Data Accumulation: Organizations struggle with vast amounts of data and the need for effective management and retrieval strategies.
  3. Data Expiration: DataStax offers features that allow automatic data expiration, helping manage data lifecycle efficiently.
  1. AI and Business Transformations
  2. AI Integration: Businesses are increasingly adopting AI for better customer experiences and operational efficiency.
  3. Future Growth: Anuff discusses the anticipated surge in demand for AI-driven applications, particularly in retail and other sectors.
  1. AstraDB: DataStax's Cloud Product
  2. AstraDB is highlighted as a key player in the cloud database market, offering robust vector database capabilities.
  1. Cost Considerations with AI and Databases
  2. Discusses the importance of understanding costs associated with AI implementation, particularly in vector retrieval which can be computationally expensive.
  3. Companies must evaluate the balance between experimentation and cost-effectiveness in deploying AI solutions.
  1. Importance of Benchmarks
  2. Anuff emphasizes the need for benchmarks in evaluating database performance, especially concerning relevancy and retrieval accuracy.

Key Takeaways

  • Adoption of Vector Databases: There’s a significant shift towards vector databases as organizations recognize their role in enhancing AI capabilities.
  • Data Management Solutions: Automation in data management, such as automatic data expiration, is becoming critical for organizations to handle large datasets effectively.
  • AI-Driven Business Strategies: Organizations are urged to think strategically about AI integration, focusing on practical applications that deliver measurable outcomes.
  • Conversations Around Costs: As businesses experiment with AI, discussions about cost will become increasingly vital, influencing decisions on production deployments.

Pivotal Moments

  • Anuff's description of how ChatGPT utilizes databases for managing conversation history illustrates the practical application of AI in enhancing user interactions.
  • His insights into the expected growth trajectory of AI in business underscore the transformative impact of these technologies on operational models.

Conclusion The episode underscores the intersection of AI and data management, highlighting how vector databases are poised to revolutionize the way organizations understand and utilize their data. With the rapid evolution of these technologies, businesses that effectively integrate AI into their operations will likely lead the way in their respective industries.

For further insights and a transcript of the episode, visit [Eye on AI](https://eyeonai.com).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I can also feed it the facts at the time rather than having it deal with its foggy memory So if I can give it the set of facts, then I can eliminate hallucinations as well. So those combination of things, it's basically if I can have the LLM, either I can tell it your preference is the information that I'm supplying to you. And if you genuinely don't know, say you don't know. That combination ends up with very low hallucination output. Hi, this episode is sponsored by Salonis, the global leader in process mining. AI has landed and enterprises are adapting, giving customers slick experiences and the technology to deliver.

0:46The road feels long, but you're closer than you think. You see, your business processes run through many systems, creating data at every step. Solonis reconstructs this data to generate process intelligence, a common business language. With process intelligence, AI knows how your business flows across every department, every system, and every process. With AI solutions powered by Solonis, enterprises get faster, more accurate insights, a new level of automation, and a step change in productivity, performance, and customer satisfaction. Process intelligence is the missing piece in the AI-enabled tech stack.

1:32Search Celonis, C-E-L-O-N-I-S, to find out more. Hi, I'm Craig Smith, and this is Eye on AI. This week, I speak to Ed Anuf, Chief Product Officer at DataStax, about vector databases and their role in AI applications. We explore how retrieving trusted data into large language models addresses the hallucination problem. Ed explains the growth and demand for vector databases and key factors like production costs and benchmarking relevancy. I hope you find the conversation as informative as I did. I have been over at DataStax for the last close to four years now, a couple of months short of that. Prior to that, I was at Google for about four years and joined Google when it acquired Apigee, which was the API management company that I was part of for a number of years.

2:47I've worked in a variety of companies, enterprise software, as well as things like blogging software and, you know, a few consumer things as well. So I've been in tech for a long time now. Originally an RPI graduate, Rensselaer Polytechnic Institute, way back in the day. Yeah, yeah. And what made the move to data stacks? And I'm curious, I mean, vector databases in particular are all the rage right now because of the hallucination problem with LLMs. But what was data stacks doing before the GPTs came along? Well, so we make the Cassandra database. It's a very popular open source database that is designed for scale out data.

3:42And Cassandra is the database that's used by Apple, by Netflix, by Federal Express, by any of a large number of companies that need to deal with scale-out data that they're distributing to users on a global basis. So part of the attraction of coming to DataStax was the fact that it was dealing with very large data sets. And the goal was to make that data available for AI and machine learning use cases. And that was something that we were already well underway on. a number of our customers were well underway on. Obviously, some of the users that I mentioned are very well known for their use of predictive AI.

4:36For example, Uber as well also uses Cassandra. It's one of their primary databases, and they've published a lot of great research papers and done a lot of talks about how a lot of the predictive AI that powers the Uber application is built on top of Cassandra. So our goal was always to bring that to the customers, as part of the product. And then in the middle of that, generative AI hit and obviously has really sent everything into overdrive. Yeah. Cassandra, what kind of a database is that? So it's a NoSQL database. It is similar to databases like Mongo and Couchbase or DynamoDB or some of these others, which means that it, You know, although it does let you use a query language like SQL, standard query language, which is what most relational databases use.

5:33But it's more designed for situations where you're dealing with large amounts of data and the data schema change can change or can be very flexible at runtime. So you find a lot of the, these types of databases became very popular for the large internet services. Google, Facebook, Amazon were really the pioneers of this type of database. And Cassandra itself spun out of a project at Facebook and the technology was open sourced by Facebook. So that's where you tend to see this type of technology is where people are dealing with large amounts of data that they're using in their interaction with their end users.

6:18And the, I mean, I've got a few questions.

6:25So DataStax builds itself as a real-time data company, a real-time database service company. or I guess you build on-prem as well. But is that exclusively for now that we're dealing with Gen.AI for vector databases? How different are vector databases from Cassandra? And why this sudden shift that I'm seeing to vector databases from other forms of storage? So a couple of different questions there. So first, we do talk about ourselves as a real-time database. We do so in the context of AI as well. And the reason is because Cassandra is designed for very high throughputs of data. You can write to the database very quickly, but more importantly, that data is immediately available for reading as well.

7:25And it turns out that becomes really important in Gen AI, and I'll talk about that in a minute. But along the way, what we saw, and basically starting last year and going into this year, is that a new application pattern emerged, and it was inspired by ChatGPT. And I think we've all experienced ChatGPT, So it's useful to sort of start with that because it helps explain why vector databases are important. When you show up at ChatGPT, part of what makes ChatGPT interesting is not just the question and then you get a response. It's actually the chat part of it. So you ask a question, it gives you a response, you ask it for some elaboration, or perhaps you're having it help write you some code.

8:17You'll go and say, oh, that isn't exactly right. Could you make it do this? And so it has a sense of history. And if you know about how LLMs, large language models work, they're completely stateless. The model itself, when you ask the model a question, it doesn't change. That model is frozen. It can later be tuned or retrained, but the model itself, the data comes in, it actually can be represented from a software development standpoint. It's a very simple black box. I have an input, which is often called the context or the prompt, which is a form of context. And then it has a response. So where do things like the history come from?

9:00Where do, you know, those actually come from a database. And so ChatGPT has a database that sits next to GPT-4, and it knows the history, it captures the response, and it's able to create a personalized conversation with you. And that's sort of where the magic happened. And then OpenAI, the company behind ChatGPT, did a blog post in December of last year. Because at the time, they weren't fully in the ChatGPT business. They actually had launched us an example. What they were really trying to do was get people to use their models. They wanted you to use GPT-3. They want you to use GPT-4. And so they said, if you like this thing that we've built and you wanted to build your own, what you need is a database that sits well with the large language model.

9:50And oh, yeah, part of what helps that database sit with the large language model is if it is able to store and query vectors. is the vector is the numerical representation of essentially the concept. What it does is it takes that whole string of text and it reduces it down to a very, well, I say reduce, but in some cases actually it can be much larger, a very long multidimensional number that represents what that word or phrase or sentence or paragraph reduces it down into a context and and or concept concept semantic concept and that's how llms think they don't really literally think but but they they represent these these you know what we consider ideas or concepts they represent it in a large multi-dimensional space and it's the distance between things in that space that says how similar or different they are.

10:54So this is an important capability that lets a database work well with LLM. Now, there are some other uses for vectors. Some databases have been supporting vectors for other reasons, basically for doing better search results for a while. And some of those were among the first crop of databases that people started using with LLMs. um all the other database vendors very quickly followed suit by by this time we're almost a year since you know a couple of months short of a year since since opening i wrote that blog post that kicked off this whole thing and at this point most major databases have some some vector facility because it becomes really important there's lots of people who are building these types of applications in enterprises and startups and so on and and they need databases and they need a database that can work with these models and that's why you're seeing everybody talking about vector databases sorry that was a little bit of a long response no that's uh that's wonderful that really clarifies it quite a bit so for data stacks uh at this point when did you introduce vector databases?

12:10And is your business, you know, 6040 with traditional relational databases and now vector databases? Where is the future of the database business going from your point of view? So really good question. So we introduced the vector capability, you know, since we're built on top of open source, we were able to get the code out there very quickly early in the year. We had it live in our cloud service in April of this year. And we're at a point right now where 40 % of new signups, people coming to our cloud service, are doing so for vector usage. Actually, let me amend that, actually. it's actually more than half are using it for vector usage.

13:09And about 40 % actually have never used Cassandra before. So there are completely new users who came to the service because they were looking for vector databases. So this is pretty important for us. Now, I suspect that some of the other database companies are seeing similar stuff. I'd love to say, oh, we're uniquely suited. I do think we've got some great important capabilities that others don't have. But I also do, I think that the bigger point is that there's a major catalyst this year with people trying to build these types of applications, and it's causing them to go and rethink what databases they're looking at and evaluate different databases and so on.

14:01So long and short of it is that's why, you know, if you're following the database industry, the entire conversation is around vector databases because new application types don't come out that often. The last time we had a new application type of this level was probably mobile. and that catalyzed a whole bunch of stuff. And then prior to that was probably e-commerce and websites and the web itself. And you start to get that hyperbole as to the, but I've been around for a while. I'm sure you've been following this stuff for a while. We all have seen this thing where there's a new application type and companies rush to respond to it and it really changes everything.

14:53And, you know, within the database space, again, probably the web. And then prior to that, probably client server being like the biggest catalyst. So very important to us, long and short of it. Yeah. Can we back up just a little bit? Because I want to talk about some of the products. But I have questions about data that I've never satisfied myself with answering. You know, I was close for a while to a company called LabelBox that is a platform for labeling data. And it fascinated me. This is, you know, during the supervised learning when that was the dominant technology in AI. And in reading, and understand my background is not as a technologist.

15:54I was a foreign correspondent covering politics and, you know, conflicts and stuff like that. So the amount of data that's being produced is increasing, it seems, exponentially. I don't have the numbers in front of me. Uh, and, uh, so you, you, when you talk about real time data, it's, it's, uh, it's, you know, writing and reading data on, uh, in real time, but presumably some of that data then gets stored for, uh, for long-term, uh, usage or retrieval that it's not being retrieved, uh, frequently, but you don't want to lose it. And where does all this data go eventually? Because it's not being erased.

16:49And at the time that I was looking at this, there's kind of a hierarchy where it gets moved from one medium to another until it's finally on literally tape drives stored in a mine shaft somewhere. But do you understand that? Where does this data go? And is it being accumulated to the point that someday there's going to be more data than any other artifact in the world? It probably already is. long and short of it is um it's somewhat somewhat depends on the use case and it really ends up um there there's two pieces to it there's there's sort of cost and policy um and um depending on the systems and depending on the so first of all you're correcting your basic statement which is particularly when we talk about real time um you're capturing data you are You're perhaps doing analysis over ranges of it.

18:00You're keeping it around. And you are accumulating it. And a lot of companies do accumulate it. And the question then becomes, how long do I keep it for? In some cases, depending on what you're doing, the cost of getting rid of it might be less than the cost of keeping it. The cost of storing data has dropped. over the years. And so what ended up happening was a lot of companies, particularly in the internet age, just kept on accumulating that data. Because again, querying and deleting a range of data is actually from a computation standpoint, from the process of doing that, be just as expensive as writing that data.

18:48So it could be much more actually. And then the cost of storing the data can be very low. So then you see a lot of companies just go and say, okay, we're just going to keep this data. just because it's not like a, it's not as much as a business decision. It's more just like, okay, the cost of doing something, we'll get around to it. And maybe they do it on certain basis. Then you have sort of the policy questions. And the policy questions come from either business or regulatory. And so then you have the questions of, in some cases, some policies require you to retain the data for a period of time.

19:32Other policies go and say, you shouldn't retain it. Many people on like corporate email systems and so on, things are auto-deleted after 30 or 60 days, right? For as a company policy, because if that's your standard policy, then you don't have the data, you know, then the data isn't there and so on. And it may not be desirable to keep it for various reasons. So that becomes a whole decision point in and of itself. And then there's security reasons as to why you don't want to keep data around for a long period of time as well. And so, for example, companies that are very concerned with privacy actually make it and advertise it.

20:11They advertise that they don't retain any data, personal data whatsoever. And so there's a whole set of things around that. And so the long and short of it ends up being that you probably could look at 100 different companies with their data retention strategies and find 100 different answers. Because a lot of it will have to do with the quantity that's being accumulated and the cost associated with it. But the companies that, and the ones that we talk about, for example, like financial statements and things like that, that only get, you know, updated and aggregate, you know, on a monthly basis.

20:56That's why you can get, for example, a bank statement. They increasingly made it becomes harder to do because they shift it from to what's essentially called cold storage or less accessible forms of data. But they still keep it around for 20 years in some cases if you need to get something. Whereas, you know, the internet companies that are capturing and literally storing a row in the database every time you load a page view because they're doing it for personalization and ad tracking purposes, those, they will collapse and compress and create summaries and delete the rest. And the interesting thing is that those summaries end up getting created through MLPRO.

21:47That actually ends up being an important usage of machine learning is to go and condense those 10 ,000 clicks into something that they can then just store and they can delete the records of the previous 10 ,000 clicks. So it is a whole thing in and of itself. I don't think there's sort of a one-size-fits-all strategy on that. It ends up, as I said, ends up being with how much data you're dealing with and how quickly it's changing. Yeah. And so with data stacks, is there a drop-off then that a customer can set that it's reading and writing, but after a certain period of time, it goes into some sort of a storage bin that's a little more difficult to access?

22:39How does that work? Well, so that you definitely can do that. There are a variety of mechanisms too, but it actually gives me an opportunity to talk about one of the cool features, which is we do actually have the ability for at the individual role level in the database, and it's pretty unique to Cassandra that you can make the data expire. And as a consequence, you have lots of users of Cassandra that take advantage of that. And they just they're like, OK, I only care about this data for a month or whatever. And then after that, I don't want to have to go through the problem of crawling through all my data and checking each row and saying, oh, is it older than this amount of time?

23:23Let me delete it or whatever. Rather than that, the database just automatically does it. It just drops it off of the database. And so then you have this data set that always contains your last 30 days or whatever the period of time that you want it to be. So that's one that is that. And why that becomes important is that, again, where you get into these situations is where you have these very large sets of data. And it is one of the areas where we play and where Cassandra is uniquely suited is for the extremely large data sets. Yeah. My familiarity with vector databases comes from addressing the hallucination problem with LLMs.

24:08And a lot of companies are building a vector database. Even after they tune an LLM, they still have this hallucination problem. So they build a vector database with trusted data and the LLM becomes simply a language interface. Is that how DataStax customers are using your vector database? Yeah. So that is a description of a process called retrieval augmented generation, or we often hear it called RAG, right? So if you're at a conference, an AI conference or something, everybody will be like, oh, RAG, RAG, RAG. and you're like, what is this? And so it is this process of, and as the name implies, retrieval augmented generation, goes and says that the generation is, that's the name would apply, being augmented by retrieval from data sources that are fed into the LLM at the point of inference.

25:07And in fact, what you can do is tell the LLM that it should only consider the knowledge that is supplied to it at the point of inference. In some cases, it's just to supplement it. But if you really want to eliminate hallucinations, you just go and say, look at, use your reasoning powers, but don't use your knowledge. And the reason is because one of the aspects of hallucinations is there's a lot of causes of hallucination, but one of them is that that LLMs have a foggy memory, if you will. Right. Just like humans. They although perhaps a different mechanism, although not that different a mechanism, but that would be an entirely different conversation.

25:50The, you know, so rather than it going and trying to, you know, remember exactly, you're like, here's the set of relevant content, choose among them and choose among these things and give me an output. And you do that for two reasons. As you said, it's trusted. But there's also another important piece of that, which might be that it's also, you know, sensitive. So, for example, you could have an LLM that is providing you with medical information, medical, you know, medical recommendation. And you don't want the language model fine-tuned on a set of electronic medical records. The LLM is not good with private information.

26:45Generally speaking, anything that goes into the LLM, it's going to leak out. Yeah. Either inadvertently or advertently. There's no way that you put access control on the knowledge that an LLM has. You could do something convoluted, which I see some people say is, oh, I'm going to have a second LLM spill the beans and filter it out. But now you're getting into this Rube Goldberg architectures. It's much better to just supply the LLM with that sensitive information as it needs it. Remember, an LLM has no memory in and of itself. So I can say, here's Ed asking a health care question and and here's his electronic medical record.

27:30And LLM looks at that, goes through and says, you know, well, you know, Ed, you know, probably, you know, need to exercise more because, you know, as I, you know, you know, your weight has gone up in, you know, over the last six months. Right. But it doesn't remember that fact and then later be in a conversation with Craig and say, hey, Craig, you know, you want to exercise more because you don't want to end up like Ed, right? Like, I don't want, I don't want Vell L.M. to know anything about Ed when he's talking to Craig. And I don't want it to know anything about Craig when he's talking to Ed, right?

28:04So that's an other piece of it. Now, the other aspect of it is, again, I can also feed it the facts at the time rather than having it deal with its foggy memory. So if I can give it the set of facts, then I can eliminate the hallucinations as well. So those combination of things, it's basically, if I can, you know, if I can have the LLM, either I can tell it, your preference is the information that I'm supplying to you. And if you genuinely don't know, say you don't know, that combination ends up with a very low, low hallucination output. A lot of these things are more art than science, which is frustrating to many as they're building these systems.

28:52And you can make it more exact. Again, we've done this. We do this for we have an AI co-pilot that helps you use our products. And what we've done is we supply content out of our documentation into it and we tell it you exclusively use that. and then it gives a result that doesn't have any hallucinations. Now, we also do a little bit of fine-tuning as well, but if you really want to eliminate the hallucinations, it's a good mechanism. Yeah, and right now, I know you have a product called AstraDB. What is your primary product and what are the use cases that people are turning to it for? Sure. So we have a cloud product called AstraDB, and we also have a self-managed software that people can run themselves called DataStax Enterprise.

29:53And they're both built on top of the Cassandra open source database. You know, our fastest growing product is the cloud product, as one can imagine. We do have many, many enterprises, actually. The majority of the Fortune 500 is using DataStax Enterprise, running it in their own data center. Both of those groups are very, very interested in the vector database capabilities. I talk to, so it isn't just the ones in cloud. I talk to customers every day that are doing things with their sensitive data that they need to keep in their own data center. But yeah, most people who are new to our product and our company do come to the cloud service.

30:47They come to AstroDB, they do a sign up and they're just right within a couple of minutes there that they are able to create a database and connect it to their applications. And the type of applications that people build are all over the place in terms of applications that are powering mobile apps, applications that are powering websites, applications that are being used for things that are controlling, Internet of Things devices. And as I started out by saying, at this point, for brand new people coming to the service, the number one catalyst is generative AI. Maybe that'll even out or slow down over time.

31:38but but i do think that just i think it just has to do with the number of new applications people are building in general that that a lot of that focus is ai related right now yeah uh the um does does the suite of products include the the vectorization of of uh data because uh you know in order to put it in a vector database they need to be vectors uh so that's that's a really good question. So what we do is we actually do have the capability within the products for being able to use open source models. The selection of the model that you use for the vectorization is actually a really important decision.

32:24Some people want to use OpenAI. Some people want to use Google's models. Some people want to go and use open source models. the different models all have different costs associated with it. And as a consequence, you see that this selection process, people will use a smaller model for vectorization because it's much cheaper. So we enable all of that. As I said, we do provide an open source model where that makes sense. but majority of people are using, you know, they're, they're, they're using something like open AI or they're using one of the Google models. Those tend to be the preferences. Yeah.

33:12But on your platform, you have those available or does someone have to? Well, there's, yeah. I mean, we, we, we let you integrate your account. You can go and put in your, your open AI account. Yeah. And I've been writing recently about the cost of inference and rate limits on large language models, which are constraints for enterprises. Does using a vector database lower the cost of inference and does it reduce the number of tokens that you're putting into or pulling out of an LLM because you've got the data vectorized already someplace else? Yes. So it depends on what it is that you're trying to do.

Read the full transcript

34:11If you're doing a simple search use case, the vector database can be a very good substitute for being able to go and take a natural language query, meaning a sentence that somebody's going to ask, and give you the result, but it'll be a search result, meaning that if you're asking a question about, for example, what is the statue somewhere in the center of my town, it's going to give me a standard answer. But what it's going to do is it's going to understand And the question is going to turn it into a vector. And then it's going to give me a result that will give me the top result. Or I don't remember how many results I want, but generally it will give me the top result.

35:03It tends to be the way people will build these things out that matches that question. So that gives me a little bit of the AI experience because I'm able to ask this human question. But then I'm getting standard response. If I want that response to be customized, the vector database can't write for you. What the vector database can do is it can read for you. But writing that response, then I need the language model. And so this ends up being one of these things that people look at for knowledge bases, for example, where they're like, you know, I don't really need a GPT personalized response to every query.

35:45I just want to give that result. And somebody goes and says to me, you know, ask, comes in and asks a question, you know, how, how do I, you know, whatever, you know, you know, change channels on my TV set. I just want to get that stock answer that, that says, you know, that gives me the page from the reference manual on the remote control. Right. But I want to handle the person asking that question one of a hundred different ways, including in different languages. the vectorization, the vector query process handles that very well. And so, yes, as a low cost option, you see that sort of thing quite a bit.

36:30But it really depends on, like I said, how much are they trying to do of having sort of a full conversational AI experience? If you need that, then you're going to invoke the model at least once through every interaction. Yeah. So I would imagine that this business is growing leaps and bounds. You were saying 40 % of all new signups are for use with an LLM. What kind of growth are we talking about? Well, so keep in mind, as we talk about these things, the usage growth is pretty dramatic. Yeah. But the interesting thing is that a lot of what people are doing right now is experimentation. So this is one of these interesting things as we talk about vectors and vector database and business growth.

37:30And obviously, we're running a commercial business. So I have this type of conversation all the time with investors and such, who are looking across the database industry. And they go and say, how soon are we going to actually see one of these public database companies announcing their results and see huge growth? And I said, well, we're probably another six months away from that because what we're seeing right now is a lot of experiments but it we're we're all what are all every database company is what's called a consumption business if people aren't actually using it live on their website and their business processes or whatever they're not going to be consuming more they're not going to be consuming more you know database software they're not going to be consuming cloud services they're not going to be consuming open ai it's like like Like, you know, that will all come from this stuff going into production.

38:23So what I would suggest for people as you're trying to look and understand this stuff is, you know, sort of look at the world around you. As you go into this holiday season, you're going to see, you know, we're about a month away from Black Friday, which is the Friday after Thanksgiving, largest retail day of the year. Right. And I genuinely don't know the answer to this. We're working with a number of those retailers and many of them are working at breakneck pace to try to get their stuff live. But as I said, something that sort of everybody who's listening to this can do themselves is sort of pay attention to that.

38:58Like when you go to, you know, when you start doing your shopping, you know, is there an agent on the, do you see a conversational agent on the website saying, hey, what can I help you with today that lets you do a chat conversation and is suggesting products to you? So by this time next year, very highly likely you're going to see those on the majority of websites, just the same way that when we saw the mobile transition, right? Remember that, you know, you saw, oh, it took about 18 months, but it was a steady progression. Everybody was like, hey, try out our app. And you started to do it. And now we're at a point right now where if you're like me, the majority of stuff that you buy, you know, from the store down the street, you know, you're going and looking at it like Home Depot, which, by the way, is a data stacks customer runs everything in their e-commerce site, uses our Cassandra database.

39:51But if I'm going to go and get something from the Home Depot, I've got, you know, some DIY, I've got to fix something. And I'm trying to figure out which, first of all, which store is closest to me and which one has that thing in inventory and what aisle and row that it's in, you know, all of that I do through the mobile app. Right. And so so we'll we're going to a lot of us think and with, you know, with a lot of evidence that we're seeing a similar transition to AI. And so what that means is over the next few months, you'll see a set of companies that do this and the rest will the majority of them, in fact, because we have what's called the crossing the chasm.

40:36You've got the early companies that do everything first and then the rest. But you'll see a few of them race out over the next couple of months with these types of AI services. And the expectation is, again, over the next year, you'll be every website you log into, every mobile app, when you open it up, you'll see various forms of this sort of personalized AI experience. Some of it will be classic chat. Some of it will be a little bit more. That's one of the things I often get people asking me, like, does it all have to be via chat? The answer is no. It actually can be done in a variety of different user experience modalities.

41:17But you start to see these things and you'll be like, oh, that's AI powered. And I think that's going to be the, you know, we're going to go through this process where it's pretty novel right now, but then it'll just become a standard thing. Keep in mind, the majority of the AI hype that we have experienced this year has been experiments. You definitely, you know, I suspect if you're like me, majority of the interactions you do, majority of things you buy online and so on right now, you're not seeing Gen AI in the loop on them. And that will change. That'll change pretty dramatically. But as it does, that's when you're going to see the business growth.

42:00We're already seeing the usage growth, as I told you right now, in terms of signups. But what are people doing once they sign up? They're building their demos. And the demos don't use a lot. So they neither consume a lot, nor are they paying for anything yet because they haven't launched these services. Yeah, well, actually, that's why I've been interested in rate limits in particular, because, and for listeners that don't know what rate limits are, the LLM companies limit the number of tokens you can use per hour, let's say, and that limits how you can scale. And so there's been a lot of experimentation, not a lot of enterprise.

42:45I shouldn't say not a lot, but it's mostly experimentation. And the question is, will this rate limit, which is directly related to the availability of GPUs, which we all know are, you know, all sold out. how do you increase, how many of these experiments will succeed in going into production? Do you have any sense of that? I know it's off topic from what we're talking about. No, it's actually, it's very on topic. And I get that question quite a bit. Because again, these questions, a lot of folks are trying to figure out what is the size of the vector database market. So I have variations of this conversation all the time with analysts and both industry analysts like Gardner and Forrester and financial analysts who are trying to make recommendations of which database stocks to buy and all that.

43:48And so what we know is this. We know that the cost for a similar amount of data and a similar amount of retrieval, and it doesn't matter which database you're using because we benchmark them all. There is a significant amount of compute used in doing a vector retrieval. It is almost 10 times the amount of compute that a regular database query would do. And so that's an increase in cost. Some of that compute is GPU-based compute. But there's a lot of work being done to shift that to minimize the amount of GPU compute that's necessary. And in fact, our systems can do all of the, in fact, they do all of the vector retrieval without GPU compute.

44:35So that's not true of all the vector databases, but majority of them are moving in that direction. We actually get a lot more requests in the opposite direction, which is if we want this to be even faster, can we use GPU compute, even though it's scarce. So the computer is expensive. And so doing Gen AI has an expense. And what we then go and say is, okay, as we look at that, everybody's doing these experiments, but probably some subset, probably anywhere from one in three to one in 10 actually have the business use case that justifies it. And when they do, it's not going to be a 10X cost. It has to be a two to five X cost.

45:22So where we look at what we're doing, and again, when I say 10X cost, we benchmarked everybody and we averaged it across, across whatever, every, you name it, like you throw names at me all day long. And we did, we, we, we worked it into our benchmarks. And because we needed to know this, where we are a database data at scale company. So, so all of our customers are very focused on cost. So what I think you're going to see is that, you know, every vector database company is going to be talking about and commenting and innovating around cost and performance. For example, we created a piece of technology called J Vector that is built on top of Java because that's the language we use that dramatically cuts the cost.

46:14And I know that you'll see similar approaches or announcements from other companies because they have to. So we have to get costs down. But even with that, you also have to have the use cases that actually drive a performance business outcome. Right. So going back to that's often why I use some of the retail examples, because they will they are the retail industry tends to be and the variations of it won't literally be people selling product might be people selling travel. but they have the outcome measurements where they're able to go and say, oh, we delivered this AI recommendation service and we converted, we got this much more business.

46:55The companies that don't have that are going to struggle because AI is a nice to have, no matter how cool and innovative it seems, it has a cost involved. And we're seeing this already. People come and say to me, oh, I use chat GPT and it was slower or the answer wasn't as good. Did they change something? And my response is, I don't really know what goes on behind the scenes there. But I say, look, what they've said has been that they haven't changed their model at all. So what I think that they're doing is they're doing things to control costs. And some of the things you do to control costs, what people perceive as quality and relevance from the model is not always a property of the model.

47:37It can be a property of the vector database you're using. It can be a property of the type of compute you're using. A similar thing, you go and see that Microsoft recently with some of their very popular copilot services, it's costing them more to provide it than what they charge users for. And so they'll also be, so all of these companies that, But, you know, the big tech companies that have been rushing AI out, they can do that and they can operate at a loss. Most enterprises won't do that. They might operate on an individual request, they'll lose money. But what they'll have to be able to go and say is our overall sales increased by X percent by using this model.

48:27And that increase of revenue that we got has to be greater than the amount we spent, right? Very simple lemonade stand economics, but that's the way enterprises need to operate or they're going to get beaten up on Wall Street, right? So we're going to see that. And this is what we call the production filter. And we talk about this a lot, which is everybody's looking at Gen.AI. and this is the year of experiments. And it should be. People should be getting out there, learning what this stuff is good for. That's the only way you can figure it out. But the production filter will be a much smaller set of those will be the ones that you see live on the websites or live in whatever format makes sense for the business.

49:17A lot of this stuff is stuff that won't actually be presented directly to the user. It might be something that your customer support agent that you're talking to they're using a gen ai based system that's giving them stuff to tell you right uh uh or or or how to process a claim or something of that sort so so yeah that that the cost will be one of the biggest gating factors yeah probably second to cost will be hallucinations but but it'll be cost by a long shot yeah and and to be fair too we're we're only a year in um the the gpt3 gpt4 uh era and uh and a lot of the you know everyone's working on uh reducing costs or increasing token yeah exactly they have to look i'm old enough to to remember you know the web 1.0 days when when you know startups were racking and stacking uh sun microsystem servers and Cisco routers and all of that.

50:26And, you know, as much as people sort of complain about the, you know, irrational exuberance of the tech industry and all of this stuff, the reality is that we've all seen these situations where, and it's baked into the institutional memory of most of these companies that you can only get, you know, so far ahead on your skis in terms of this stuff. And the big tech companies can afford to operate at a loss as they do these things. But even when they do it, they're doing it in very careful, measured ways to make sure that it's something that they have a path to sustainability and profitability around.

51:10And all of the enterprises, I interact with very few enterprises that are going and saying, look, we don't care how much it costs, we just want to get this thing out there. Every one of them is going and saying, okay, what is this going to look like at production? And frankly, actually, this is one of the ways that we're able to actually gauge as we're talking to one of these companies, our way of gauging how close are they to put it into production ends up being, is the cost conversation happening? If they're not talking about cost at all, then what we know is they're still just basically experimenting, which is okay.

51:47But obviously, as a business, you're always trying to figure out how far along are these folks, you know, in terms of where they are as customers. And so one of our checkoffs, checkboxes as we're going through this is like, okay, are they starting to ask the hard questions about what does production costs look like? And we want to get them there because we know that they're not going to go live until they go through that process. So obviously I think, you know, we as a company have very good answers for it, for that stage of it. But we know that they have to be asking those questions, right? And so, yeah.

52:21Yeah. Well, you know, this conversation, I've been kind of following my own interests, which is not necessarily the most sophisticated train of thought. What have I not talked about that you'd like people to know about data stacks? Well, I think, you know, I've dropped a few pieces into it. We do feel, you know, DataStax has been the company that has served most of the companies in, you know, in the Fortune 500 on their digital transformation journeys, meaning when they went to mobile, and they had the scale that they had to deal with, you know, DataStax and the Cassandra database was the database that was your production, reliable production at scale and at a cost basis.

53:16We've been the company that powers that. And the brands have name dropped in the course of the conversation. Whether you're using somebody like Priceline, whether you're using somebody like every time you scan a package or for that matter, every time you track your package on Federal Express, every time they scan the package on the 20 different or 100 different points on its journey. All of that is going through DataStax's databases. Every time that you use Netflix, that's using the Cassandra database, the open source Cassandra database. Apple uses Cassandra as well. So Cassandra and DataStax have played a really important role in making this data available at scale.

53:58And we've done all of the work to make it not just the best for those types of database use cases, but for vector database use cases. And so I think you've asked all the right questions, which are, what happens when you put this stuff into production? What does it end up costing? What are the challenges around dealing with these large amounts of data and preserving the relevancy around it? And as people are thinking about this stuff, I always recommend make sure you try everything. But the one thing that isn't happening enough in this conversation, in this AI conversation is the experiments are great, but asking those questions about what happens when I go into production and when you look at the vector database, as you pointed out at the start of this, every database company seems to be talking about vector databases.

54:48Make sure that you are thinking about what is the cost, the reliability, the accuracy, the performance of this stuff into production. Otherwise, you're going to have a really cool experiment that, you know, looks really good. And, you know, maybe you've presented it to your board of directors and everybody's like, ooh, ah, but it'll just be, you know, it'll be just the AI equivalent of the concept cars that, you know, that Detroit used to wheel out, right? It won't be actually something that you're actually able to get out there and have impact on your business on. Yeah. Just on that last point, are there benchmarks that people should be looking at when they're evaluating databases?

55:33There are standard benchmarks for relevancy and recall. And you should also be asking, you should be looking at, as you look at the databases, we're publishing these now and I'm seeing other people are too, but it's not widespread yet. So I would definitely ask a database vendor, whether it's an open source database you're looking at or whether it's one from a vendor, I would go and ask to see the benchmarks. And just not just from a performance standpoint, some of these benchmarks are also related to relevancy. There are standard tests that are done to evaluate the relevancy. Not all vector databases are created equal in terms of the relevancy of the results that they get.

56:15And that is something that you will end up discovering as you build this stuff out yourself. most common thing that we see is somebody loads in a bunch of data. And by the way, relevancy also changes with the number of records in the database. So any of the vector databases, when you're doing that little test where you load in a couple of hundred or even a thousand, you know, data entries, then you go and see, you're like, wow, this is really good. And then you go and load in a hundred thousand or more. And then you start to see this drop off. And at that point, you're in a real bind. And so we, you know, we think that's going to be important.

57:01Certainly, as we spend the next 18 months with people putting these projects into production, you're going to hear a lot more about this. And it already is a topic that you see on blog posts and articles and people write, like you see these, how to improve the relevancy of your vector results. You're going to see a lot of that. It's going to be the next hot button issue. Hi, this episode is sponsored by Salonis, the global leader in process mining. AI has landed and enterprises are adapting, giving customers slick experiences and the technology to deliver. The road feels long, but you're closer than you think.

57:41You see, your business processes run through many systems, creating data at every step. Solonis reconstructs this data to generate process intelligence, a common business language. With process intelligence, AI knows how your business flows across every department, every system, and every process. With AI solutions powered by Solonis, enterprises get faster, more accurate insights, a new level of automation, and a step change in productivity, performance, and customer satisfaction. Process intelligence is the missing piece in the AI-enabled tech stack. Search Celonis, C-E-L-O-N-I-S, to find out more.

58:31That's it for this episode. I want to thank Ed for his time. If you want to read a transcript of today's conversation, you can find one on our website, IonAI. That's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may not be near, but A-I is already changing your world. So best pay attention.

From the publisher

This episode is sponsored by Celonis ,the global leader in process mining. AI has landed and enterprises are adapting. To give customers slick experiences and teams the technology to deliver. The road is long, but you're closer than you think. Your business processes run through systems. Creating data at every step. Celonis reconstructs this data to generate Process Intelligence. A common business language. So AI knows how your business flows. Across every department, every system and every process. With AI solutions powered by Celonis enterprises get faster, more accurate insights. A new level of automation potential. And a step change in productivity, performance and customer satisfaction Process Intelligence is the missing piece in the AI Enabled tech stack.

Go to https://celonis.com/eyeonai to find out more.


In episode #153 of Eye on AI, Craig Smith sits down with Ed Anuff, Chief Product Office at DataStax. 

We take a deep dive into the world of vector databases and their integration with AI. Ed sheds light on the innovative ways AI is being incorporated into database technologies, with a special focus on the advancements and applications of Cassandra in this realm.

Ed elaborates on the challenges and opportunities in melding AI with database management systems. He talks about the evolving landscape of data storage and retrieval in the age of AI, and how these advancements are reshaping businesses and data strategies.

We also explore the broader implications of AI in database technology, including scalability, efficiency, and the future of AI-driven data solutions. Ed shares his insights on how companies like DataStax are at the forefront of this technological convergence, driving innovation and transformation in the industry.

If you find this episode insightful, please support us by leaving a 5-star rating on Spotify and a review on Apple Podcasts. 

 

Craig Smith Twitter: https://twitter.com/craigss


Eye on A.I. Twitter: https://twitter.com/EyeOn_AI


(00:00) Preview, Celonis and Introduction 
(02:19) Ed Anuff's Background and Journey to DataStax
(03:33) The Role of Cassandra in DataStax
(05:06) Understanding Cassandra: A NoSQL Database
(07:50) Impact of ChatGPT and Vector Databases in AI
(11:58) DataStax's Introduction of Vector Databases
(17:26) Addressing Data Accumulation and Usage
(22:18) Managing Data Expiration in DataStax
(29:24) Introducing AstraDB: DataStax's Cloud Product
(36:42) Business Growth and AI Integrations in DataStax
(42:18) Rate Limits, AI Experiments, and Enterprise Integration
(49:47) The Future of AI in Databases 
(55:20) Evaluating Databases: Performance and Relevancy

More from Eye On A.I.

All 266 episodes
#153 Ed Anuff: Unpacking AI's Role in Data ManagementEye On A.I. · 59 min
Listen in VO