90 | How To Use AI to Chat With Your Data with Joanna Stoffregen

21 May 2024 · 53 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Leveraging AI - Episode 90: How To Use AI to Chat With Your Data with Joanna Stoffregen

Episode Overview In this episode, host Isar Meitis invites Joanna Stoffregen, an expert in Large Language Models (LLM) and Retrieval Augmented Generation (RAG), to discuss how businesses can leverage AI to interact with their data effectively. The discussion emphasizes the cost implications of not utilizing AI tools and outlines a step-by-step approach for integrating LLM technology into business operations.

Key Themes and Concepts

Introduction to the Problem

  • Many businesses struggle to access and utilize their data effectively.
  • A common challenge faced by data users is the inability to find necessary information stored across multiple platforms (e.g., CRM, emails, ERPs).
  • This inefficiency can result in significant financial losses due to time wasted searching for information, missed deadlines, and poor decision-making.

The Role of RAG in Data Interaction

  • Retrieval Augmented Generation (RAG): A technique that allows LLMs to access specific data that they were not originally trained on (e.g., company-specific data).
  • This capability solves the "billion-dollar problem" of ineffective data retrieval and enhances decision-making processes.

RAG Implementation Steps

  1. Identify Use Cases: Determine specific scenarios where RAG can assist in data retrieval.
  2. Tool Selection: Choose appropriate tools available on the market or opt for custom solutions.
  3. Prototyping: Start small, testing out capabilities before scaling.
  4. Deployment and Maintenance: Implement chosen solutions while ensuring ongoing support and updates.

Cost Implications of Not Using AI

  • According to research by the Data Institute Corporation, employees spend an average of 2.5 hours daily searching for information, leading to significant annual costs for businesses.
  • For a company of 1,000 employees, this inefficiency can result in losses exceeding $2.5 million annually.

Examples and Use Cases

  • Joanna presented an example of a marketing manager named Julia who loses valuable time searching for a market research report across various platforms.
  • RAG solutions can enable users to ask specific questions and receive accurate answers without sifting through multiple data sources.

Technical Aspects of RAG

  • RAG systems leverage vector databases to semantically search and retrieve relevant data based on user queries.
  • The three levels of implementation:
  • Basic Level: Utilize existing LLMs (e.g., GPT-3, GPT-4) to test use cases with minimal setup.
  • Intermediate Level: Employ no-code/low-code tools that allow for data chat functionalities without heavy technical requirements.
  • Advanced Level: Custom development of RAG solutions tailored to specific business needs.

Ethical Considerations and Challenges

  • Data Privacy: Ensuring that proprietary or sensitive information is protected during AI deployments.
  • Accuracy and Hallucinations: Addressing the potential for AI-generated inaccuracies or misleading information.
  • Implementing safeguards and monitoring costs associated with using AI models is critical for successful deployment.

Key Takeaways

  • Leveraging AI through RAG can drastically improve data accessibility and efficiency in businesses.
  • There are multiple entry points for businesses of all sizes to start integrating AI into their workflows.
  • Careful consideration of costs, ethical implications, and tool selection is essential for successful AI implementation.

Resources Mentioned

  • [Glean](https://www.glean.com): A platform designed to help employees search through company data effectively.
  • [Verba](https://www.verba.com): A tool for securely managing and querying proprietary data.
  • [OpenAI GPT](https://www.openai.com/gpt): A language model that can be used for basic data interaction.

Conclusion This episode emphasizes the importance of utilizing AI to streamline data access and decision-making within businesses. Joanna Stoffregen's insights provide a practical roadmap for companies seeking to implement AI solutions and highlights the need for ethical considerations in AI deployment.

For additional insights and resources, listeners are encouraged to connect with Joanna on LinkedIn for ongoing discussions and access to tools and spreadsheets discussed in the episode.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hello, everyone, and welcome to a live episode of Leveraging AI, the podcast that shares practical ethical ways to leverage AI, to improve efficiency, grow your business, and advance your career. This is Isar Maitis, your host. And maybe the question that I get asked the most from people who are a little more advanced and started playing with AI is, how can I chat with my data? Like I have all these data sources and I have data in my CRM and I have data in my emails and I have data in my ERP and I have data in my Google Analytics and et cetera, et cetera, et cetera. And I want to be able to chat with my data and ask holistic questions.

0:36And I don't know how to do that. And most people don't know how to do that. But the reality is beyond the fact that I get asked that a lot, it's a huge need and not being able to access your data in an effective way actually costs businesses a huge amount of money in overhead, mistakes, missing clients, losing clients and so on and so forth, just because they cannot get the right information accurately and in a timely manner, a problem that having a capability to chat with your data across everything solves. So how do you do that? So I'm really excited to tell you that our guest today, Joanna, is an expert on exactly that.

1:21Her agency helps businesses develop these kinds of solutions. But what we're going to talk about today, beyond the custom developed solution, there are other ways that you can develop on your own with tools that exist in the market. You can start very small, test things out, and then slowly upgrade to more and more advanced capabilities. This is obviously maybe the holy grail for businesses when it comes to using AI and machine learning is the ability to really know what's going on across everything that you do in your business, from finance to customer service to marketing to projections to emails to etc, etc.

2:02And hence, this is obviously a very important topic. So I'm really excited and honored to welcome Joanna to the show. Joanna, welcome to Leveraging AI. Oh, hello. Thanks for having me. Excited for today. Yeah, listen, you and I have been talking about this for a while, and the content that you're sharing on LinkedIn on this topic is actually pure gold. And I think most people hear the technical aspects of this. Oh, it's called drag. I'm like, okay, I don't know what that is. Okay. And once people start diving into the technical stuff, they get really scared. and so they stop even though it could actually be very simple.

2:44So I know you prepare like a full kind of like step-by-step process to take us to first of all why and then how and then other considerations like safety and security and costs. So I'm sure everybody who's going to join us is going to learn a lot. So let's get started. I will give you the stage. Thank you. So yeah, let's get started. I will now demystify RAC for businesses and hope you guys will get value out of it. Let me share my screen when you see it. Yes. And by the way, those of you who are listening to this as a podcast afterwards and you cannot see the screen, a few comments about that.

3:24First of all, join us on the next live. We do this every Thursday, almost every Thursday at noon Eastern time. So you can join us and then you can see the screen and even ask questions and participate. So that's number one. Number two is we're going to describe everything that's on the screen so you can understand if you're driving and so on. And number three, we have a YouTube channel under the Multiply brand. It's just instead of a Y, it's spelled with an A-I in the end. So if you want to watch this afterwards, once you're done walking your dog or doing the dishes or yoga, whatever it is that you're doing, listening to this podcast, you can later on go and watch it again on YouTube.

3:59But let's dive in. It's Davey and I will also try to make it more understandable in case someone listens to it in podcast afterwards. Okay, so we start basically with the billion dollar problem that RAC solves. And I want to introduce you basically to Julia. So Julia is a product marketing manager that works for SoftTech.com. It's a tech company based in Silicon Valley, has 300 employees. She's working remotely from Berlin. And in her day-to-day, she actually talks with a lot of teams to do her job, like design team, tech team, sales, marketing. And of course, uses also a lot of apps for communication, task management, like Gmail, Slack for team communication, Jira for product tickets.

4:54She has tasks in Asana, Figma. So basically all these familiar apps that we all know if we're working in tech or not. So today, basically, Julia, she has she was working on a promotional campaign to launch their new product feature. And she arguably needs basically a market research report that, you know, her colleagues prepared a month ago or so. So she remembers that actually someone sent it on Gmail. So she looks there, but she cannot find it. Then she thinks, okay, maybe in marketing Slack channel. So she searches there, doesn't find it either. Then she looks again and maybe in Google Drive, she said maybe it's in the marketing folder.

5:43Does not have success either there. Common scenario, it happens to me all the time, even if I'm not even working in such a big corporate. Right. So then she remembers that actually her colleague Max might actually know for sure because he's also involved in this project. So she decides to Slack him. So she sends him a message. She waits now 20 minutes. There's no answer. Then actually she remembers that Sam is a current in San Francisco. It's 3 a.m. there. And actually she will take a lot of time and she will get her answer. So frustrated, she basically decided to move to another task, but actually she lost a lot of valuable time.

6:24She's now behind on her project and she might even lose the deadline for the submission of her project. So very common scenario that I guess a lot of us working in tech and in big companies or like a smaller company's experience, right? And we have also the data to support that. not our research, but from the Data Institute Corporation, they found out that 2.5 hours per day, on average, is spent basically on searching information. So imagine like one out of five days of work, we basically spend or the average worker to search through different apps to find their information that they're looking.

7:07On average, we use like around 11 apps and it takes around eight searches before basically we find the right information. And that obviously has a lot of costs to it because time is money, right? They estimate that a company that employs 1 ,000 employees loses around 2.5 million per year on actually this problem that we are not able to find, to search and retrieve the right information. That has obviously impact in not being able to find the right information for your project, like Julia, or you might miss an important update because you did not get the information on time. and fail to use basically important updates of your project.

7:58So all of that basically also making wrong decisions because you don't find the data. It costs a lot of money for companies. And this is basically the billion dollar problem that Rack solves. And also companies like Glean help to solve. Glean is a$2.2 billion company. so hence they're working on a very good problem to solve and just basically to explain you a bit on high level what Glean does and not just Glean there's few companies that do the same they help employees like Julia essentially search through their company data and retrieve this relevant information so for instance we can see here there's this employee that he's looking for the status of a project and normally what he would do he will search slack he will search google docs to find this status he might search the hub repository to see the commits of the code but now with the solution of that also cleaning of like technologies rack solve is that he just asks a query and then the system provides him basically all the information in natural language in one place to see.

9:16So imagine you don't wait. It's directly, you can ask anything you want the system because it knows basically what your company knows. So this is also their promise. Know what your company knows. Yeah. I want to touch on one short thing from everything you said. Every single company is sitting on a gold mine of data that is just not accessible today. And even small company people, I know a lot of people who are listening don't run billion dollar companies. They run a$20 million company, a$50 million company, a$100 million company. And they're like, I don't have any unique data or any proprietary data.

9:56And that's not true because every email of every employee, every phone call that you have recorded for customer relationships, every proposal you've ever written, every data in your CRM or your ERP, all of that can be helpful in making future decisions and understanding what's currently happening in the business and so on. And it's right now requires a lot of time and effort to do the analysis across these platforms, just because it's built into silos. So yes, we do this every now and then. There's like a project manager who will spend once a month and once a quarter writing a review by collecting all that data.

10:35It's time consuming, it's not 100 % accurate, and it's not accessible in real time when you need to make decisions. Like it's once a quarter when the project manager did that review. And it, again, may or may not dive to information that other people need. And having a solution like this literally enables, think about having a Slack channel with your entire data, where you can literally ask any questions about anything and you don't care where it's stored and what is the source. And this is exactly what this thing does. And then every employee in the company, if you're the customer service person, And you can find out the answer to answer the call you're on right now.

11:11If you are the VP of marketing, you can understand stuff about your recent strategy that you deployed. And if you're the CEO, you can understand trends across everything that's happening in the company. What are things that are happening? What's not moving fast? Like literally any questions you want to ask, if you connect the right data sources, you can ask and get answers, accurate answers in seconds. So it's a game changer compared to everything that we know. Exactly. And also another very common use case for salespeople before actually they're, before they're having their call with the lead, they need to access a lot of information in their CRM, in Slack, in their product to really understand what is their company and the lead about.

11:56and it spent a lot of time and now basically systems like rack systems you could just write hey give me all this information about elite x and it just pops pops it up for you so super helpful yeah and also you don't need to be a hundred million company or whatever even i know companies with 10 20 people that have exactly the same problems super helpful for any company size so exactly now so what what is RAC so yeah so RAC stands for retrieval augmented generation and it's essentially a way of for large language models like chat GPT to be aware of data that they have been not trained on so for instance chat GPT is trained on all of this public data but it's not trained on your slack conversations on your basically company data on your CRM data on your Jira tickets, or basically this data that are proprietary for you or for your company.

12:57So it's essentially a way how we can basically make a model be aware of this data. And of course, we have this problem of nations and that they're not, models are not up to date in their training data. So it's also solves that as well. And to better understand that, I want actually to showcase an example and use a case study that we were working on. So basically, we developed for car mechanics of Mazda cars, basically mechanics that repair Mazda cars, a way how they can basically get repair instructions from the system. So their problem was that it will take very long for the mechanics to search through manual Mazda instructions to find out how to repair the particular problem of a car.

13:50And that had huge business impact on basically unhappy customers because it would take very long time until they would receive their repaired car. And often they would also, the mechanics ordered the wrong repair parts because they found the wrong information that they needed to repair. Yeah, so you can see here. So let's say there's this mechanic and ask how to fix a broken airbag of Mazda G4 car. You can see here, if you ask this question to chat GPT, it will give you basically just generic data, generic information. And then it will also prompt you basically, hey, ask a specialist. The system that we developed is a rack system.

14:35So it has all the Mazda manuals. So it's able to tell you exactly how to repair, let's say, a broken airbag for the particular Mazda car. So now, obviously, the mechanics, instead of searching manually and doing all these mistakes, they have a system that directly basically tells them how to repair easily. So this is basically what RAC is and how basically it is able to become aware of your data and an example to showcase the difference between a normal chat with GPT that might have general knowledge about car repairmen, but not specified knowledge about the particular problem that we are looking forward to solve here.

15:21Yeah, so I'll touch on two things that you said just as a quick clarification to everyone. One, I like to say about all these large language models that they follow the common say that they're the jack of all trades but the master of none. as you start deep diving into specific topics and really getting needing like expert advice they usually don't know the data if you're trying to get deeper and deeper on a specific professional topic they just don't have it and even if they have it they may not have it for your business your company or thing so this is one thing that it's coming to solve is you give it the data that you needed to know and then ask it to retrieve just from that data and as joanna said, the other thing that it solves is hallucinations.

16:05The level of accuracy, and it would still hallucinate probably to an extent with RAG, but it's going to be a significantly lower percentage and a much higher level of accuracy when it comes to providing you the right information if you are using RAG versus just using a open, like a market available large language models. so how is the magic done? Take us through the process How is the magic? Yes, so on a very high level non-technical level there is the user that asks a question in our case the mechanic asks how to fix a broken airbag of Mazda car so now the system has basically all this Mazda manuals in a vector database stored, right?

16:51And then the system will actually semantically search for in the database for the relevant chunks or the relevant basically documents for this question asked. So you can see here, maybe search three, four documents. There have like the information needed in specific. And then this is the retrieval part. Then the augmented part is where basically the user query is augmented with the retrieved information from the first step. And then the last step, which is the generation step, is where basically the system generates the answer, right? So remember this example from before where you ask a question and the answer might be in very different sources.

17:38So this is the different sources. And then as what just a user is actually the generated answer from these different sources as a answer to your question. So this is a high level. Yeah. No, it's a great explanation. I think what we're missing here for the common person is, I assume people are not technical, the combination of the words vector database give them diarrhea. And so how do you get the data from your CRM, from your emails, from your Slack channel, et cetera, into a vector database that then the system can retrieve from? Yeah, sure. So basically, first, like a vector database is storing this, how I would say, it's basically storing the semantic information that is relevant basically to the query.

18:36So just take it from the beginning, how you would go about it, you will have all this data, you would need basically to use to load them into the vector database and you would use load data loaders like llama index or lang chain and you will always need to basically pre-process that so you would need to have basically unified format usually in markdown then basically you would need to split that data into chunks, that these chunks, it will be easy basically to create embeddings. And then these embeddings, we basically store them into this vector database so the system can retrieve from. And the vector database, basically these embeddings are basically the semantic meaning of basically the...

19:29So what is the semantic meaning? let's say a king and a queen it will be on the same kind of the same level we'll have the same semantic meaning as opposed to sound that it will be basically not relevant to king a queen so this is basically the semantic yeah sorry that the vector of databases i want to make this a little more specific is the common person can he use tools so people who don't know lang chain and don't know how to do this from a technical perspective, are there easy ways, like something that a common business person can do or specific tools that they can use in order to get to that outcome without maybe as a first step, right?

20:15Let's do a test case, see that it's actually working, see that it's giving us some things, and then come and hire a company like yours to actually do the more complex, sophisticated solution. Yeah, sure. Exactly. There might be levels to that. So first, as I mentioned, we would need to collect and prepare the data that we have available. It could be internal data. It could be your PDFs, invoices, all of that. But it could be also external data. It might be basically websites, blogs, or basically YouTube videos, transcripts of it, and so on. And you could basically get started very easily with, first of all, if your use case it might be you want to have your company knowledge database.

21:01So basically use Rack for your company knowledge. Use out-of-the-box tools. Don't try to basically start this from scratch, develop this from scratch. Use tools like basically Glean or similar, but depends on your use case. For instance, if you have a couple of documents with your company policies, let's say, and you want essentially to train a chatbot, to have it in the chatbot for the users on your website to ask questions about your policies, about how is your refund or about products, you might just prove, validate the use case with a simple GPT. So we start basically very simply with simple GPT from OpenAI.

21:48so take the documents and just upload them into a gpt and validate the use case there is it basically does it answer correctly my questions is it actually useful and then you can also let's say you want to get buy-in from stakeholders that's the perfect way actually to showcase them hey look it works it answers questions that our customers answer us and we have now a person answering them. We put the first level tool. So before you go to level tool, those of you who don't know what GPTs are, so OpenAI under their platform within ChatGPT, by the way, until a few days ago, it was only in the paid version as of this Monday, it's available to everyone, which is a magical tool that they now made available on their free platform.

22:38It allows you to develop these mini custom automations that again, sounds very technical, but you do them literally by giving it instruction in regular words. I want you to allow the user to ask you questions about the information in the attached documents. Questions they may ask might be one, two, three, four. I would like you to start the conversation by telling the user, what do they want to know about HR questions? And that's what's going to start. It's going to say, hey, what do you want to know about HR questions? And then the person will type. And if you've uploaded all the HR documentation into that GPT, which again is a simple drag and drop.

23:19There is no, all the stuff that Joanna talked about before happens in the backend magically. Like you don't have to set up the data, structure the data, convert the data. Like all of that, literally all you have to do is drag and drop the HR documents, your employee handbook, your guidelines, these kinds of things, drag and drop them into the GPT and that's it. and people will be able to chat with that data on a very high level of accuracy. Again, right now, it costs exactly zero dollars to do. And to figure out how to create the GPT, after a little bit of training will take you minutes, including the training will take you an hour, and you have something that you can chat with the data.

23:58The biggest problem maybe of GPTs is the amount of data you can upload to them. So the HR example is a great example because they don't have huge huge amounts of data, but you're limited with the amount of data you can upload into these GPTs. But as I mentioned, it's one of the most amazing capabilities we had in businesses ever. Definitely now that it's completely free, not that I think that$20 a month should make a difference, but now you don't even have that excuse. So it's definitely worth checking out and testing it. And as Joanna said, it's a great way to do a test case and see what results you're getting before you go to level two.

24:33So now let's talk about what is level two. Yeah. And before we go to level two, if you have still issues with data privacy, you don't want to share your HR data or whatever, you can actually generate synthetic data based on to model your real data and create a GPT to still basically validate the use case. So that's also creating synthetic data might be viable as well. Use, sorry, level number two. So in this basic level, you could use a drag and drop tool like Stack AI, for instance, that it's no code, low code, that you can create this RAC system or RAC chatbot by drag and drop. They do also pretty much everything for you as just more advanced version than a GPT, more capabilities there.

25:21And of course, if data privacy is a concern, you could actually cost locally, like a tool like Verba. Maybe we can show it. I don't know if I... Do I share now something? Yeah. I see a dog on your screen. Oh, here we go. Ah, cool. They changed also the user interface. Yeah, great. So basically, yeah, this is the no code. Sorry, not no code. basically a way how if you have proprietary data let's say if you're a lawyer let's say and you want to chat with your cases but the cases are proprietary data private data you can use a tool like verba it's yeah so you upload again your documents here and you chat with your documents yeah now they just change the user interface crazy okay anyway um so yeah that's level two And of course, level three is where you build that from scratch.

26:22The solution that I basically showed you with the car mechanics, that it's basically built from scratch. And there you use basically frameworks like LankChain or Lama Index. Yeah, you would need an embedding model to do the embeddings, yeah, to store the embeddings in a basically vector database. can use different ones, Quadrant, Pinecone, Beaviate, one of the most famous ones. And we use also Streamlit for the prototyping. So you can actually create basically these prototypes as I showed you on the use case, the case study basically before. I want to summarize very quickly the three levels. So level one is use an existing model.

27:05You can use a GPT in OpenAI. You can even just use Claude and upload documents, but that will be like a one-time thing, Like every chat will have to upload the documents again. By the way, if you want to chat with a lot of data as a one-time thing, Gemini 1.5 Pro from Google, as of this week for developers, so if you sign up, has a 2 million token context window, and I'll experiment with that in a minute. The platform that's available, if you don't sign up for the new program, is still 1 million tokens. So 1 million tokens is 700 ,000 words, which is a huge amount of data. It's way bigger than any of the other tools that exist today.

27:42So the second best one you're going to get is Claude, and that has 200 ,000 tokens. And the cool thing about Gemini 1.5 Pro is that it's free. So if you go to their AI studio by Google, you can upload documents and chat with those documents. So if it's a one-time thing, if you're working on a proposal and you want to look at what's in the RFP or what's in the proposal so far, or if you're working on analyzing existing data and it's not an ongoing thing that you'll need to do again and again, using a model is actually not a bad idea. Like I said, Gemini 1.5 Pro has two benefits. One, you can load a huge amount of data and two, they're getting a very high level of accuracy in getting correct responses, which is north of 97%.

28:27And so this is obviously very helpful. So that's level number one. Level number two is using tools that were built to do exactly that. There are no-code tools that allow you to upload and connect various levels of information and then chat with that information. And then there's, again, some of them work in the cloud, meaning you actually give them the information and it gets loaded over there. And some of them, like Verba, works locally, which means you don't have to give your privacy and you don't have to think about who am I actually giving access to my data in order to be able to chat with it.

29:01And 11.3 is a custom-built solution that is built specifically for your needs based on your exact data, based on your use cases and so on. The differences is obviously the amount of data you can upload and how much is it going to be customized to your needs. But the flip side is how much time and money it's going to cost you to develop it. So that's the thing. And it's not a bad idea just to go through all three steps. Start with step one, make a quick test, see if it works. Go to step two. that works well but you can solve everything with these tools then go and have somebody like joanna build a custom solution for you to solve that problem yeah no i totally agree like a hundred percent if you have a tool or out of the box solution uh to do your job don't even bother to build stuff from scratch yeah there's a question there's a question is verba the open source version of data stacks i don't know if you know that answer.

30:01Question from Verba. Sorry. Is Verba an open source version or like an open source option for data stacks? I don't know. What do you mean from data stacks? Yeah, it's probably a different tool. Yeah, it's from Vaviate actually. It's from the vector database Vaviate. So they basically created this Verba tool. It's open source. You can basically try it but i'm not sorry for not being able to answer yeah it's all good okay so let's continue yeah so basically you're still able to see my screen right okay okay what to be aware of when you basically the rack solution go into production besides basically the technical issues and not being able to retrieve stuff which basically but there's like a very technical solutions to those.

Read the full transcript

30:59But the main two things is actually the inference cost when you scale these applications to your production. So for instance, to generate one paragraph with GPT-4 costs the same money as if you would generate a whole book with mixtral model. And why I'm saying that, because maybe I can open that, Because some use cases of RAC, at least from my experience, you need to use models like GPT-4. Now GPT-4, oh, it's actually half the cost, so it's a bit better. But I need to click that and I can't click that. I will add something to what you said as you're opening the document. For those of you who don't understand how this works, all these models behind the scenes are actually using APIs to different large language models once you have the data in the database.

31:50in order to generate the actual responses. And each and every one of these tools, whether open source or closed source, has a different cost associated with using the API. And when Joanna is saying inference, inference is when these models are generating results, like the actual generation of content in tokens is actually called inference. So that's basically the generation of anything in these models. and you pay per token, which again is about 0.7 words. It doesn't matter why it's this way. Just take that as this is what it is. So the price range varies dramatically between these models. In some cases for a million tokens, so 700 ,000 words, you will pay 20 cents like a Mixtral 7B, which is an open source model from a French company.

32:40Lama is another open source model for Meta, similar pricing. If you go to a Claude 3 Opus, which is right now the most advanced model from Claude, from Anthropic, that's$70 instead of$0.20 for the same amount of generation. Quality, you're probably going to get a higher quality at least now. So you got to pick and choose in which cases you want to use which models. And I assume that's where you're about to go. Yeah, exactly. And I would have to say that the price will be commoditized at some point. The cost of intelligence will go to zero, probably. We see that basically with GPT-4-0, that it halved the cost with the same quality.

33:27By the way, that's after GPT-4 Turbo halved it from GPT-4 original. Yeah, it's crazy. Basically, as I said, I did not update my presentations on that. But yeah, this is how it did. So basically, just to give you an example. So for instance, use case like the Julia example that we were talking about, about having 300 users to use the system 10 times per month. it will cost with GPT-4 Turbo around almost$15 per user. So imagine a business that sells this. You need also to be viable. So yeah, it costs a lot, but you can also directly see how much more or less it's with smaller models. So one technique, basically.

34:15I want to go back to this just for a second. The numbers for 300 users were over$4 ,000 a month with GPT-4,$300 a month-ish with GPT-3.5, and$159 with Mixed Trial 7B. So it's 20 times more to do it with GPT-4 Turbo, but sometimes you have to. There's use cases where these better models just give you better results. But I agree with Joanna that this price is going to go down and down all the time. Yeah. For instance, this mechanic rack use case, we needed to use GPT-4 to actually. Yeah. And obviously, their quality matters, right? And of course, when you have little users, it does not even matter.

35:05Yeah. But in case you scale with users, one solution is just to have a router that depends on the complexity of the query it basically roots it to a cheaper bottle so this is one of the many yeah um then we look into security very important when you have especially a front-facing Iraq app we had this use case not use cases sorry examples that there was this chatbot that sold a car for one dollar and it was actually legally binding and also another example of air canada a chatbot that gave basically a bad advice on on on refund regarding a price yeah exactly there was this refund that the user wanted and it gave it to them and it was legally binding so who is liable there basically what happens here it's we need to imply some guard drills to eliminate basically any data leakage or prompt injections.

36:11And yeah, for the guide drills part, there's tools like Nemo Guard Drills or Guard Drills.ai, a lot of them, basically. So it's a way how before the, it could have both an input and output query, but it's basically a step in the middle where before it gives the answer to the user, it actually filters it for any harmful or basically inappropriate content and can also protect personal information. Very important step, as I said, for front-facing, customers-facing solutions. And of course, also for the token, to monitor basically the token usage. We also need a solution there. Very important to understand where costs are coming from.

36:57And there are tools like Landsmith or Phoenix from Arise that you can directly see how many tokens are queries from your users and you can even track cost per user and per query. So very important when you think to basically put the app finally into production, but also before, especially the token usage monitor. Yeah. I think that's it. No, it's great. Before I dive into, there's a bunch of questions from the audience. before we dive into those questions. I want to summarize that aspect of it. So we said there's three different levels that you can test this out. One is very simple. The other is a little more advanced, but still does not require any third-party custom development.

37:45And then the third requires custom development. In the custom development universe, you have to pay attention. And actually on the security side, probably on the other two as well. Like you got to understand what data you will be sharing with who and that company that you're sharing with, how are they keeping your data secure if they're keeping your data secure? And different companies will have different comfort levels with different solutions, right? So if you're Joe Schmo and you're selling shoes in the market, maybe you don't have any information that is problematic to give to Chachi PT or to Claude or one of those.

38:24but if you're a doctor, a lawyer, a financial advisor, all these highly regulated companies, you just cannot. It's not even an option, which means your company that you're working for, and if you own it, then it's on you to figure out how to run this in a secure way that does not expose your data, which means putting the right measurements in place. So that's problem number one. Problem number two is that that Johanna mentioned is how do you protect yourself from just mistakes that these models make, either because they just made a mistake or because people know how to manipulate these models in order to give them information that they shouldn't.

39:05As John has said, people can use it against you because if the model is going to commit to a price, it committed on your behalf. And so you're legally bind to whatever the model says, if it's in a chat with a client. And so you got to take these things into consideration with the solution that you're providing and make sure that the solution that you're putting in place is aligned with the needs of your business. And this could be anything, right? It could be, I don't care. That's a fair enough solution. It's fine. Like I said, there's cases where if it's an internal tool that people are going to use and you give them the ideas and they're like, okay, fine.

39:41So it's going to help you 90 % of the time. And the other 10 % people still have to do the old manual process. Okay. So there's different justifications and different environments will require different kind of level of security, both in means of data security, as well as means of protecting you from getting the wrong answers from a model. I want to jump into a few questions. I have one, and then there's people from the audience who have a few. So I'll start with the first. The first question is, can you please re-mention how AI systems are priced? so do you want to take that one yeah sure I can actually show you here do you still see my yeah so here we have basically all the costs for an LLM app and specifically for a rack app so we have here the embedding costs which are basically the embedding models which it's very minimal as you can see to upload 10 million documents, or let's say who uploads take a million, but let's say 10 ,000 documents, it just costs you 65 cents and 10 cents with a smaller model, embedding model.

40:57Very minimal. We don't even count them. Yeah, no point in mentioning it. Yeah, exactly. But then here, the inference cost, it's basically each of these companies like OpenAI, they charge per token. So for instance, tokens is basically you have a query and depends on how long is this query, it charges you basically for the token of the query. So for instance, here you could see that the cost per 1 million input tokens, it's$10 compared to GPT-352 or BOE is just 50 cents and 27 for Mixtro. Then they have different price for the output tokens. So basically for the tokens that they will generate to generate the answer to you.

41:49So it's 30 for GPT-4 Turbo and so on for the other APIs models. And yeah, we also have basically the costs from the document context and the prompt that the system of our app has. So, for instance, back to the example with the mechanic car instructions, the prompt template will be basically the instructions that are in the backend for them to tell, hey, now you are an expert in car repair. Here it's the question of the user, please check your documents and find the appropriate basically answer. So this is the prompt template. So this is also basically the costs. Yeah, again, to explain this, in all these solutions, even if you're using a basic GPT, you're basically adding an additional prompt to what the user is writing in order to explain the system exactly what to retrieve and how to retrieve and what data it's looking for.

42:51And all of these count in your token count. The question from the audience, the follow-up, is one token equals one byte? And no, the answer is no. It's just the way these systems work. And just take it into that a token is about 0.7 words. And that's what you're going to pay for. Most of the pricing models that you're going to look at are going to give you the price for a million tokens. So basically, if you're looking at a million tokens, it's going to cost you X. A million tokens is going to be 700 ,000 words-ish. Just depends on how long the words are. Yeah, we can actually also see it here.

43:29I just make a quick example of what tokens are. So we have this query, how to fix a broken air back of a car. You can see here, it's basically the tokens, how it splits. So each of those are tokens. Yeah, how it splits the words into tokens. So there's multiple tools online. What Joanna is showing now is called tokenizer. You can literally just paste your text in there and it's going to tell you how many tokens, but it's also going to show you how it's broken up. Totally unnecessary. If you just want to know roughly, just assume that every 0.7 words is one token and you're going to pay for the tokens you use.

44:02And there's different price for tokens coming in than there is for tokens coming out on some of these models. So what I mean by coming in is your input. So your prompt, the document you're loading and all of that stuff is tokens coming in. And then tokens coming out is inference or what the model is generating. and in most cases they're not equal and in all cases it's not equal the inference to generation actually costs you more money and sometimes a lot more money it's still very small amounts as I mentioned before the most expensive model right now is Claude 3 Opus and it's$70 for every million tokens of output so$700 ,000 words.

44:51I don't know how many books that is, but it's probably two to three books of 300 pages each or something like that's going to cost you 70 bucks. So if compare that to any other way of generation of that amount of content before, it's still free, right? But if you add that times 300 employees times 10 times a day, these costs start adding up. And so optimizing for cost is important. So that was one topic that we've covered. I have another question still on this. What's the cost of actually hosting this, right? So I need that vector database to reside somewhere. What is roughly the cost of hosting that?

45:32I have actually to host 50 million tokens in a vector database, it costs, I think, just 70 something. So also very minimal. Yeah, Very negligible. Yeah. Yeah. Okay. Yeah. Yeah. Awesome. The next question is what size of companies you feel are best suited for this solution? So I think, and I will let you answer, but I think we mentioned three different solutions. So I think we can reply on the different solutions and the different sizes of companies. Yes. any basically company even if you're individual you can and you have some documents that basically you want to chat with your documents or gain more information from the documents or you might have i don't know you might have five youtube videos about a specific topic that you want to create a new content about generating original content out of it you can basically use gpt to do that create a gpt and store the data there um but sorry back to your question it's so which side of the company yeah yeah so no go ahead yeah so but if you for instance if you use a company like green if you we go now to their website they're basically they don't even have the pricing online so it shows us that it's for enterprise so this might be for a company that are 50 100 300 000 employees yes i'm sure there's and the fact they raise whatever 200 million dollars also shows you yes yes exactly yeah so there are for sure other tools like Glean that are more specialized essentially for smaller companies I'm I don't I'm not really actually I think I haven't even spreadsheet with those companies yeah company there there are many chatbot tools that are basically the same thing chatbase or Dante are basically chatbots yes you connect your data they're really cheap and you can connect whatever data you want to them and you will be able to talk to that data so there's cheaper solution than Glean to give you an entry level to the level two.

47:58Okay, I'm not using ChatGPT or Gemini or Claude. I'm using an actual tool that does it. So Dante, Chatbase, there's a bunch of those that do the same thing and you can use them to upload your company information, to connect to URLs, to upload videos, connect to YouTube videos, like all these things, multiple sources, and you can do that. I will say something beyond that, and it's half a question as well. Both Google and Microsoft are clearly working in that direction, right? Where you'll be able to use their chatbots that is going to be integrated to everything within their universe and beyond.

48:35Even today on the Microsoft environment, on Microsoft Copilot, you can connect Microsoft co-pilot to external data sources. And it's not everything, but it's the big ones. So you can connect it to Slack. You can connect it to Salesforce CRM and stuff like that. So do you think this whole concept of RAG will become basically a given sometime in the next 12 to 18 months, at least to some level? Yeah. I actually had this discussion yesterday with someone about the fact that yes, Google also has. I think even the ability, even now you pay a bit more around 20 something dollars and you can search through your full drive and everything.

49:20Yes. However, now you cannot connect it with your Slack or basically with your other Salesforce or whatever, this kind of apps. That might be not sufficient if you actually want to have access to all your basically apps that you're using. You might use a tool like Glean that has hundreds of connectors. But definitely also with OpenAI, where they're going with their assistant API and also now they increase also context to their GPTs, you will be able actually to upload more. So definitely there, yes, I would say there is some competition and we will use that. Yeah, but there is room for other companies probably as well, at least in the long, short term.

50:09Yeah. Awesome. So quick summary. First of all, this was fantastic. Like we touched on a lot of things. I think we give a lot of people and both the comments in Zoom as well as the comments on LinkedIn are all very positive and people really appreciate all the information that you shared. the quick summary, you can use multiple levels of tools to communicate and chat with your data. Not doing it is by definition costing you more money than searching the old way where you're going to miss timelines, you're going to get the wrong information, you're going to potentially miss clients, lose clients, not win proposals, et cetera, et cetera.

50:55There's really very few excuses why not to do it. And you can start very small with tools that are, that requires zero technical knowledge other than literally connect your data sources and you can start chatting with them. Joanna, if people want to follow you, learn from you, work with you, what are the best ways to do that? Yeah, sure. So you just add me on Joanna Stofferigen. I basically can also see it here and ask me whatever you want. I can share more tools, more in my spreadsheet cost calculations and everything. So feel free to just shoot me a message on LinkedIn. Yeah, I think we'll do with a spreadsheet because a lot of people ask for it.

51:37Most of the people on LinkedIn said, yeah, I want it, I want it, I want it, I want it. So what we're going to do is I think I will ask you to create a shared Google Sheets with it and we'll just connect the link in the show notes once the podcast goes live. and then anybody who listens to the podcast can have it. I can't thank you enough. I want to thank also the people who join us. I'm going to just go, not with everybody because it's going to be a while, but the people who asked questions and participated. So Elsa Paul and Katie Cope and Kyle King and James Lindsay, my man, and Kursad Haratas, I hope I'm not butchering names here, and also China and Daniele and Igor and that have joined us on and asked questions on the Zoom as well.

52:20So thank you, everyone, for participating. Thank you so much, Joanna, for sharing your expertise with you. This was awesome. We'll do it again sometime in the future. Definitely. Thank you for having me. And thank you, everyone, who joined. And yeah, talk to you soon. Glad to meet. We'll take it from there. Ciao. Bye, everyone.

From the publisher

Joanna is an expert in leveraging Large Language Models (LLM) combined with RAG (Retrieval Augmented Generation) turning data into knowledge and action. Discover what it really takes—and the hidden but high costs of NOT implementing such solutions.

This webinar is tailored for business leaders who are considering or are in the midst of deploying LLM technology. We will cover every step of the project lifecycle, from initial planning and prototyping to testing, deployment, and ongoing maintenance. Joanna will share detailed examples and insights on tool selection, highlighting the pros, cons, and lessons learned from her extensive experience.

Expect to dive deep into the real costs and implications associated with NOT using AI to understand the insights hiding in your data. You'll gain invaluable perspectives to guide your decisions.

Joanna's recent viral posts on LinkedIn have sparked a widespread discussion on reducing LLM costs and optimizing project strategies. In this session, she will expand on these themes, providing clarity and actionable advice.

About Leveraging AI

If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!

More from Leveraging AI

All 330 episodes
90 | How To Use AI to Chat With Your Data with Joanna StoffregenLeveraging AI · 53 min
Listen in VO