30 | Unleash the Power of Your Business Data: An Actionable Guide To Training Models Leveraging Company Data. With Mayo Oshin, Founder & CEO Siennai Analytics, a private chatbot development company

19 Sep 2023 · 1 h 3 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Leveraging AI Episode 30

Podcast Title: Leveraging AI Episode Title: 30 | Unleash the Power of Your Business Data: An Actionable Guide To Training Models Leveraging Company Data Host: Isar Meitis Guest: Mayo Oshin, Founder & CEO of Siennai Analytics

Episode Overview

In this episode of "Leveraging AI," host Isar Meitis engages in a conversation with Mayo Oshin about how businesses can harness their existing data to drive transformative growth through AI. The episode outlines a comprehensive guide that covers the entire process of implementing AI, from defining business goals to deploying AI models.

Key Themes and Topics Discussed

Importance of Business Goals

  • Starting Point: Establish a clear business goal to provide direction for AI initiatives.
  • Clarity of Purpose: Understanding the intended outcome such as improving customer retention or operational efficiency is essential for team buy-in.

Data Utilization

  • Types of Data: Identifying structured vs unstructured, proprietary vs non-proprietary data.
  • Structured Data: Organized, easily searchable data (e.g., databases).
  • Unstructured Data: Non-organized data (e.g., PDFs, videos).
  • Proprietary Data: Sensitive internal data.
  • Non-Proprietary Data: Publicly available data.

Data Preparation and Cleaning

  • The process involves standardizing formats and removing corrupted or irrelevant data.
  • Examples: Converting scanned PDFs into usable text or ensuring consistent naming conventions in databases.

Model Training and Implementation

  • Model Selection: Deciding between custom-built vs out-of-the-box solutions.
  • Training the AI Model: Involves using cleaned data to train the model, ensuring the model learns effectively.

Deployment Strategies

  • Development of a user interface to interact with the AI model.
  • Ingestion Process: Transforming documents into a format suitable for AI processing.

Measuring Success

  • Establishing success criteria based on initial goals.
  • Evaluating the quality of AI responses against a benchmark to measure improvements in efficiency and customer satisfaction.

Cost and Timeline

  • Budget Considerations: Estimated costs range from $5,000 to over $20,000 based on complexity.
  • Timeline: Average project duration of 4-6 weeks from data preparation to deployment.

Ethical Considerations

  • Discussions around proprietary and sensitive data handling, including the importance of data security and compliance with regulations.

Practical Use Cases Presented

  • Training Company: Using video training modules transformed into a ChatGPT-like interface for improved user experience and retention.
  • Legal Firm: Automating document drafting processes to save time and resources.

Key Takeaways

  • AI Implementation Roadmap:
  • Define business goals.
  • Assess available data.
  • Clean and prepare data.
  • Train the model.
  • Develop the user interface.
  • Deploy and evaluate.
  • AI can lead to significant efficiency improvements across various business functions.

Notable Quotes

  • "All you're doing is replacing the manual tasks that a human being would do with the machine."
  • "If a machine can do it with high accuracy and consistently, then you don't need as many employees."

Conclusion The episode emphasizes the potential of AI to revolutionize business processes by effectively leveraging existing data. It provides actionable insights for business leaders looking to integrate AI into their operations responsibly and effectively.

---

Additional Resources

  • Mayo Oshin's LinkedIn: [Connect with Mayo Oshin](https://www.linkedin.com/in/moshin1/)
  • Mayo Oshin's Website: [Siennai Analytics](https://siennaianalytics.com)
  • Ultimate AI Course for Business People: [AI Course](https://multiplai.ai/ai-course/)
  • YouTube Full Episodes: [Multiplai AI YouTube Channel](https://www.youtube.com/@Multiplai_AI/)

---

Call to Action If you found value in this episode, consider leaving a five-star review on your favorite podcast platform, sharing your insights or what resonated with you most!

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hello and welcome to Leveraging AI. This is Isar Maitis, your host. In today's show, we're going to talk about questions that a lot of people are asking. themselves. How can I use my existing data like customer service, marketing data, CRM, sales calls that were recorded, financial data, project analysis, and so on, and train my own models so I can learn faster and get better results than my competition? We're going to do this with the help of Mayo Oshin, who is a global expert that has been doing this as a service to multiple businesses in different sizes for a while. At the end of the episode, as always, I'm going to share some exciting news and there were a lot of big news from most of the big players this past week.

0:40But for now, let's dive into the conversation on how you can train AI models on the data that you currently own to gain business benefits. In the next few years, AI technology will change our world dramatically. Whether you are a business executive trying to catapult your business forward or just somebody who refuses to be left behind and want to advance your career, this is the show for you. I'm your host, Isar Maitis, a serial entrepreneur and an AI enthusiast. You'll hear invaluable practical tips from innovative business leaders, AI practitioners, and some of the brightest AI minds in our world today on how you can leverage AI in ethical ways to advance your career and grow your business.

1:34Hello, and welcome to Leveraging AI, the podcast that shares practical, ethical ways to leverage AI to improve efficiency, grow your business, and advance your career. This is Isar Maitis, your host, and I've got a special episode for you today. You know, one of the big promises of AI is getting us better results in business and higher efficiencies and so on. But the biggest promise is allowing businesses to leverage their existing data, stuff that they've collected through the years on any kind of data in order to build their own models to get higher efficiencies, grow their business, and so on.

2:14And the problem with that is, and that's a question I get asked a lot. A lot of people ask me, how do I do that? How do I go about leveraging the data that I have today of CRM, customer service, marketing data, sales data, et cetera, in order to get better results. And while I don't necessarily not an expert on this topic, our guest today, Mayo Oshin, is an expert on that. So Mayo has been involved in this world of AI and developing these capabilities around it for a while. He's been one of the first contributors to LangChain, which is the platform that everybody's using two connect different AI systems to one another.

2:57He has helped companies large and small to do these processes and implement AI, as well as leveraging their models, anything from small businesses to large giants like PwC. He's currently the founder and the CEO of CNI Analytics, where this is what he does. He helps companies. He consults and train and implement companies, He's mostly B2B organization on how to utilize their existing data goldmine of any kind, whether it's documents or videos or databases or knowledge bases, et cetera, in order to improve customer satisfaction and reduce costs, accelerate growth, and so on. So he's the perfect person to have this conversation with.

3:40And like I said, since this is a hot topic that I get asked about a lot, I'm very excited to have him as a guest of the show. So Mayo, welcome to Leveraging AI. Thanks, glad to be on. Mayo, let's start with people who don't really know what to do with this. Okay, they hear about this concept that they can use their data, but they're not sure exactly how. Can you give us some examples, some use cases from clients that you've been working with or that have implemented this on how people are using their existing data and building AI models on top of that to gain business benefits? yeah so recently we just finished work with a client who that there's a pretty sizable training consulting company and they have a ton of videos and the issue from the user's perspective is that when you go through these training videos whether it's internal or external there's a ton of content to go through and so a lot of times people give up they don't even complete the training program or they just don't get the full value of what they paid for.

4:47So the task was to improve retention, to improve satisfaction, and potentially basically add more value to what the users paid for. How do we do that? We took the videos and each module typically has videos, has PDFs and other exercises and all these other things, and transform that into a user interface like ChatGPT, where the user can ask questions and have references where they can click and it will send them to the point in the video where they can continue to watch. So imagine instead of going through all the videos and watching one by one, you can ask a question and then it provides an answer based on the content and also source reference to where that video came from.

5:34Another use case would be a legal firm, one that draft tons of agreements, so maybe load agreements and all kinds of incorporation agreements. And that's something they would have to do manually all the time. And then they want to effectively get to the place where the task was to automate that process so that when another inquiry comes from the client, then the AI generates a draft of the agreement and then a senior lawyer can step in and effectively review that. And so, yeah, those are two use cases, but there's many others with PDF documents, databases, customer analytics, support documents.

6:16There's a lot that can be done. So as we talk, I can reveal more use cases as well. So I love the two examples you gave because I think we can generalize on both of them. One of them is more on video kind of content for any purpose, whether it's for training, for marketing, and so on. It becomes a lot more impactful if you can easily find content within those videos. So not just categorically, these videos are about this topic, but in minute 13 and 52 seconds, we cover this very specific topic. So now you can watch 30 seconds or three minutes instead of an hour and get exactly the answers that you need.

6:55This is good, obviously, for marketing, for customer service, for training, for any place where you have video content that now can be cataloged not to the content of the video, but to the content of every second of the video, and then can be chopped off to a much more useful, to many more use cases than just the content along for content. So that's one example. The other example, if we generalize it, you have an existing set of documents that you need to continue generating similar documents that are not exactly the same. Instead of starting from scratch or even starting with a template, you can start with something that is already adapted to the outcome that you need, the specific client, the specific case, the specific scenario, the specific market, like whatever the case may be.

7:43the AI knows how to take that quote unquote template, adapt it to whatever the scenario is, and now give you a very mature starting point, saving you a lot of time. So I think both of these use cases are relevant to a very large number of businesses. And hence, I love the fact that you pick these two examples. The next question that I think most people have is, what do I do? How do I get started? Like I'm a CEO of a company, or I'm in a leadership position or I'm in charge of implementing AI at my business. And we have a lot of data and a lot of companies say, we don't have any data. That's usually not true.

8:20Because even if you have a CRM that's been running for three years, I guess you have a lot of data. You have marketing automation tools, you have a lot of data. So anything that you have in the company, any Google docs that you have with interactions with different clients, it's data that you can use in order to train these models to get efficiencies. But the question I think most people ask is, okay, let's assume I understand I have all these data. What are the steps? What do I need to do to go from having that data to having a tool that can either be used internally or be customer facing? Yeah.

8:57So I think speaking to leadership, the first question is what's your goal here? And the reason for that is if that's not clear and it's just jumping on a trend or just trying to just be reactive to competitors doing things, then you might not get full buy-in from everyone in your team. They might get excited in the beginning and give up maybe when they start to see results and not as they expected it to be. So I think it's about clarifying that goal. And to clarify that goal thinking first and foremost about what is the business outcome? Is it that you want to improve retention of customers? Do you want to strengthen your brand?

9:41Where can it potentially show up in your income statement, for example, so that you can point to something tangible as the project goes on that the stakeholders would buy into? Because if it's just a case of, oh, this is going to help us, that's too big. If it's a case of here's how this could potentially cut support costs by 20 % and improve our gross margin, then that becomes a lot more interesting to other stakeholders as well. So I think that the clarity of that goal is very important. Ultimately, all you're doing is replacing the manual tasks that a human being would do with the machine. And so that's all that's being done here to make it more efficient and more scalable.

10:25And so once that clarity is there, then the next step is, okay, what data do we currently have? If you're in any type of business, even if you're a brick and mortar business, you have been collecting valuable data. It could just be email addresses. It could be the demographics of your customer. it could just be a process that you have going on internally that's a form of that they can assist with that and the best way to think about data is like a think of it like a tape like a table and one side you have let's say structured and unstructured and then you also have your proprietary and non-proprietary so proprietary data is obvious this is things in-house you know that you don't share with the public only your staff or stakeholders know this information can be data about your company or customers or research that you've done in your industry.

11:20The non-proprietary data is typically stuff open to the public or you will share with customers. This can include stuff on your website. Maybe you have other public stuff or stuff you share in conferences or speeches. Anything public will be in the non-proprietary quadrant. And then you also have the unstructured data. Unstructured data is just data that it's not in a table for. So anything that's PDFs, like I gave the legal drafting example, docs, anything that can be put in that kind of unstructured form. And then your structured data is more like table of stuff. This is your CSV, your databases, and so on.

12:02So you can see in this four quadrants, you can have combinations, proprietary database, for example. You can have non-proprietary PDF. And so if you're able to then figure out, okay, where's the valuable data here and see where they fit in this quadrant, that's another case. And the use case of that, which I didn't mention in the first portion of this, was an example where I built a chatbot to analyze Tesla's annual reports over the past three years. And so this was a case of going through SEC filings and parsing through and extracting tables because it contains the financial statements. And so this was an example of a non-proprietary because it's public, but also structured data, which I then turn into a chatbot that allow investors to make a decision on the stock.

12:53So that would be the next phase is about uncovering the data you have. Once that's done, it's about cleaning the data because a lot of times companies don't think about or they're starting to think about data, but they didn't start off thinking about data. So it's a whole digital transformation process to basically start to curate this data in a way that's usable. And so there's a whole process of getting the buy-in from employees to make sure they fit in certain things. and then your IT team to make sure they centralize the data in a particular place. And then there's the cleaning of the data, making sure it's a consistent format that can be used.

13:37So that'll be the first thing. Can you explain what cleaning means? I think it could mean different things for different people. And obviously, from your perspective, as somebody who implements this, it means something very specific. Yeah. For example, you can have maybe one PDF. PDF so a company can have maybe 1 ,000 PDFs and half of them are scanned the other half are just text PDFs and now the scanned PDFs are harder to deal with because when you do OCR and all that kind of stuff it can end up with special symbols or whatever or a lot of times companies do save documents that are corrupted those are then difficult to deal with then there's also the structure of the data, how has it been saved?

14:22Maybe they've got a database, but then it's all over the place. It's not really structured in a consistent way. Maybe the names or the way different things are named are different or they mean different things. So all these things basically make it difficult to just extract the data out of the gate without some sort of a way of making sure everything's clean, restructured, and then ready for training with the AI. Interesting. Okay. Question about this. Is this a process? Because this sounds like a very technical process that you need to know what you're doing. Is that something that typically a company like yours would do?

15:00Or is that something the organization themselves needs to do ahead of time? Or is it a mix of both? You mean a mix of whether it's - I mean, is it something you would come in, look at the data, say, okay, guys, you need to do this and that, work on this for a week, call me when it's done. That's to me, the mix of the two. it's up to some companies have the resources in-house to to do this kind of data restructuring and cleaning there is a point though where the person who's going to be in charge of training the ai model will probably want to be involved just to make sure that it's in the format that makes that would work best for training the model but ultimately most companies don't have that in-house capability to do the restructure and the cleaning.

15:53Okay. So what's the next step? So now I have a clean set of data. What's the next step? Okay. We're moving now to actually training the model. Okay. Yeah. So, yeah. So once the data is being cleaned, then the next step here is, okay, we want to decide on how we want to present this to the end user. Obviously, this would have been discussed initially, but the next process is actually building this out. So usually, typically, there's a UI. The process of that begins. But then at the same time, there's what's known as ingestion. Typically, with my company, my two teams work simultaneously. The UI team is working on the front end, and then the back end ML team is working on this ingestion.

16:40So what's ingestion? Ingestion is just taking all the documents, transforming it into a format that the computer can understand because the computer doesn't understand text. So we need to transform it into numbers so that we can perform different calculations on it down the line, which I'll explain. Then we store these numbers in a special database called a vector store. And this place will basically house your data but this data would also have associated numbers for those data. These are typically in chunks because it makes it easier to perform what's known as retrieval, which I'll speak on in a second.

17:22And so what happens is when the user asks a question, like for example, let's say we did ingestion on, give me a book you recently read. I'm reading AI for marketing right now. okay and so let's say in the book there's a question about there's a part of the book where you wanted to ask a question like okay what is a good ai marketing strategy for a small business and so what happens is if you ask that question to google or a typical search engine it what it does is it performs what's known as a keyword search it's going to look for keywords, similar keywords, and it's going to return back results to have those similar keywords.

18:09But what the ingestion process does is it effectively allows you to search by meaning, by semantic search. So it's going to look at the context of the question you're asking, and then it's going to look at what we, quote unquote, turned into numbers or that initial and gesture phase, and then retrieve the relevant chunks of your book that are associated or semantically similar to the question. And so what happens is the model then uses those semantically similar source documents to then produce a final result. So this will be the equivalent of opening up chat gpt and then you copy and paste sections of your book underneath the question you ask so it's basically the programmatic way of doing the same thing but no one is going to have time to do that or rather you're not going to be able to copy and paste your entire document into chat gpt so this is effectively how it goes from your document or your data to a UI like ChatGPT, the user asks a question, they get a response, and then they get a reference to the page, in your case, where the answer to that question came from.

19:38Okay. So I want to do a quick summary and then ask a bunch of follow-up questions. So you said, first of all, start with a goal. And I agree with that 100%. at the end of the day, you're trying to solve a business problem or to advance a business. So start with a goal. What is this supposed to help us do? Either reducing a pain or increasing a potential of something that has a potential we couldn't do before. Then you said, define the quadrants of the data. Is it structured or unstructured? Is it proprietary or non-proprietary? And this is one of my follow-up questions in a minute. Then you said, okay, go collect the data, clean the data, and we talked a little bit about what that is, train the model, define the user interface, meaning how is that going to be engaged with, and then do the ingestion process in order to get the data into the actual structure so you can then ask it those questions.

20:26So this is more or less the process. A few questions. Question number one is the whole aspect of proprietary information, regulated information, personal like PI information. What do you do if you have a lot of that? Do you separate it out? Do you run on a open source locally installed? What are the solutions that companies have if a lot of the data that they have is not stuff that they want to share with the world? Yeah, I think you basically touched on the key ones. I think this is where you have a choice as a company to choose between out-of-the-box, powerful, accurate, closed model like OpenAI versus an open source model that's not being widely tested and will require its own infrastructure and deployment and its own support.

21:29And so the company needs to think about two things. What resources do we already have in-house? If we don't have those resources, what is our budget to get someone who can do this? And so I'll start with the closed one because that's what everyone's familiar with. So if you're going with OpenAI or Anthropic, any of these closed solutions, you're going to get something that works out of the box. it's going to perform really well. It's going to be a very good proof of concept if you're just trying to get buy-in from stakeholders. And you're probably going to be relatively happy with the results.

22:09It's just an API that your developers can connect to and you get decent results. Now, the OpenAI recently has released reports, privacy policy documentation, just reassuring enterprises, look, we're not going to touch your data. and so on and so forth. Now, whether you believe them or not, that's a different question. Same question if you believe Microsoft or not. If your company is comfortable with Microsoft, then I'd say, why not? It's all within the same family anyway. In the middle, what we tend to do, so for example, we had a client who wanted to work with OpenAI. So we came up with a hybrid solution.

22:53We stripped away PI, So any personal information, emails, anything that could suggest any entity before we went through that process of working, basically showing the model, the information. And that was a hybrid approach that the client was happy with. But you do have clients that they don't even want the context or the surrounding information, even if it's not PII, to be shown to the model. And that's where we moved to the other extreme where there's an open source solution. You find one of these up and coming models. You have one called MPT by Mosaic, Llama 2. There's another one called Falcon.

23:37And you essentially deploy one of these models. Can cost a company anywhere from$600 to maybe$1 ,500 a month to host these models. So they'll be running 24-7 typically or per time of usage. And then you basically have to collect a training set from your company, which someone needs to come in, collect all the data, train the model, deploy the model. And then you need someone to do what's known as machine learning operations. Now you need maintenance of the model. Now, if you don't have the in-house capabilities of this, you either need to find a consultancy to do this for you, or you need to hire a machine learning ML ops personnel who can help you run through this process.

24:29And not all businesses have the budget for this. So that's the challenge. Question about that. I think you made it very clear. You go with one of the closed models. It comes ready out of the box. It's an API. It's ready to go. You don't have to host anything. You run on their servers, just talk to the API. It's easy to deploy. It's easy to maintain. But you're taking some level of risk of where that data is going versus I'm not willing to take any risk. I'm going to run the model locally. That means I need some kind of tech infrastructure in order to do this, which needs to be set up and maintained, which costs money.

25:02Two questions. One is, I know that all the big hosting platforms today, whether it's AWS, Azure, or Google Cloud, all now provide such infrastructure basically built into their tools. meaning you can now get these containers for that kind of data with the models already running on top of them and you basically can pick and choose at least in some of them what type of infrastructure you want with what kind of model you want on top of that that gives you that flexibility to do these things without having or maybe with less ml ops knowledge in-house is that a true statement or is this time will tell if they're really taking us there or not?

25:51Yeah, you still need someone. You still need someone to do the training. You still need someone to collect the data, to process the data, to make sure it's ready. If you don't get the training set properly done, then the results are going to not be great. Again, because these models are not as strong as the closed models, so you need to make sure that everything is done really well. You still need that machine learning personnel, in my opinion, anyway, or a consultancy that can help you set up the whole thing. Second question, that is more of a balanced question. So now I have a set of data that I give.

26:37Let's take ChatGPT as an example, because that and probably Claude 2 are the two most advanced models out there today. And I give it access to all this data and I will allow customers or maybe internally, and we can talk about this in a minute as well, ask questions and get results based on this data. Does the company have control on what percentage of the answer comes from Chachapiti as an engine versus what percentage of the answer comes from just look at the data that I gave you and all the answers have to come from there? Is that something as a business that I can control? Yeah, that's, in fact, I'd say that's probably 90 % of what we focus on is what's known as reducing hallucination.

27:26Hallucination is when the model effectively veers off the context, which is related to the company's data, and then just generates a response that's either inaccurate or it's not completely true. and a lot of the clients we work with are in industries where you cannot afford for inaccuracies right so there's a lot of time spent in terms of what's known as enhancing the retrieval augmentation and also to ensure that the model can see the context and if it's not sure so there's like a god clause that's typically used as you say to the model if you don't know what the answer is, just say that you don't know.

28:11So it's part of the prompt engineering to make sure that it doesn't just veer off and make things up. But there's nothing in the setup process of the model that increases its dependability on the new data versus what it knows as a model out of the box? That's mostly done through prompt engineering. So the more you, in the prompt, you reinforce that it should only use the context, the more it's going to do that. And also, we found many other people in the space have discovered that GPT 3.5 does a particularly bad job of following what's known as the system prompt. This is the prompt instruction.

29:01GPT-4 is very good at following the system prompt. So if you literally tell GPT-4, if the answer is not in the context, just say you don't know what the answer is. And it's going to listen to you nine times out of 10. Follow-up question to this. You said it's a lot about the prompt engineering, but if I'm going to make this an open to the public tool, let's say this is to help customer service, right? So I want to be able to have people chat with my customer service database. It's like a really smart FAQ, right? Because it sees everything that happened before. It can find similar events. It can find what was the answer and can give a good answer on people on what to do without me having to apply another person to answer that call or that ticket or whatever.

29:42This is a very obvious use case, but the end user doesn't know he needs to define those things. So So does that mean that when you implement such a solution, there is a prompt that the user doesn't see that's always there when you actually send it to the model, when you actually send it to the large language model, there is a backend process that happens that sends kind of like a pre-prompt thing that then you just take the end user prompt and attach to that? That's the way it works? Correct. Yeah. So what people don't realize is ChatGPT, for example, there was a leak of some people trying to hack around and find the prompts of all these popular things.

30:24And so they found the prompts for Notion. They found the prompts for Bing and ChatGPT. Basically, it's like a constitution. It's like a whole four or five paragraphs about what it should do and what it shouldn't do and what happens if someone asks a particular question. because if you do try and ask it certain types of questions, it just tells you, look, I'm not trained to do this. And that comes from the prompt that's hidden from the user. Got it. So basically any prompt that you send, you can set up an additional set of guides and rules that get sent to the actual model, regardless of what the user wrote.

Read the full transcript

31:04Yeah, it's called a system prompt. System prompt, perfect. Another follow-up question. So there's a big difference, I think, in my eyes between creating a tool for internal usage and opening the tool to clients. Definitely means of risk. If it's an internal tool, it could do stupid things. You still have somebody from the company looking at it and saying, this doesn't make any sense. Maybe I should go and do some additional research versus the same answer goes to the end user. Are there any best practices on what are the steps of deployment? So, okay, now I have done all this process. I have done all the stuff we talked about.

31:45I've created them all. I've created the user interface. I'm ready to go. What are the steps you recommend to customers, to clients to do in order to make sure that they're a enjoying the best benefits possible, but also reducing the risk of shooting themselves in the foot by deploying this to the world before it's fully tested. Yeah. So there's a, we spend a pretty good portion in authentication, right? Because if you just have the live URL, then anyone can go there and basically use your application. So the first question for it is typically, who do you want to have access to this information?

32:24And then we have a preset email. So it's usually provided by the company of the individuals. And you can see when they've logged in, when they've logged out, the questions they've asked, which is, again, is something very important for ongoing evaluation and testing, because you can see the questions the user has asked and the responses that came back from the model. So authentication is a big part of this. Another thing too is moderation. So if you have users asking inappropriate questions, they get flagged, right? And eventually you can blacklist them, right? So you can actually block them from asking questions.

33:04Another measure is known as rate limited. You don't want a situation where someone just keeps spamming questions, either themselves or using some sort of a bot, which will cost you a lot of money and cause your system to crash, right? So again, there's rate limit mechanisms in place. So those are the three typical things that are done so that from a security point of view, at least the authentication is in place and you don't have randomers who can just use the application. Interesting. Okay. So my next question is, how do you measure the success of these things? Because I go back to what you started with.

33:48I'm like, okay, I defined a goal. That goal hopefully has like a success criteria, right? I want to improve my customer service efficiency by 20%. That means that I can either spend 20 % less money answering the same number of customer service calls, or I can do 20 % more with the same amount of resources that I have. Doesn't matter. How do you go back and measure what's the action? And again, I go back to this problem that I just mentioned. There's five other moving things that are happening at the same time, right? I've hired additional more people. We've upgraded our ticketing system. Is there a good way to measure the actual benefit and impact on the business of deploying such tools that are quote unquote embedded into the actual solution so I can track what it's actually doing from a business performance perspective?

34:43So this is like also part of the prelim process would involve the client providing what we'll call an evaluation data set. So what this is essentially is what is a typical question being asked and what is a satisfactory answer for the end user. And so what the goal is here effectively is if we can replicate what a human being would respond like with 80 % or more accuracy, then you can deduce from that the amount of support you can cut down so we start with the data set provided by the client's evaluation data set and a huge part of the process is how do we keep optimizing this model to get closer and closer to what the client considers to be a good result so in the case with the training company they already knew the material inside out.

35:45They already knew what would constitute the typical questions that the users would ask and the answers they were expecting. And they knew the cost of having a coach or some other support person manually provide those answers. And so when we hit those goals, for them, it was like, oh, wow, this is essentially automated what a human being would do. So there has to be a benchmark in the evaluation process or in the prelim process from which the client or the company can look at and say, if we can replicate these answers, then we can automate the costs and effectively raise our business. Interesting.

36:31So basically you set the benchmark up front to the level that it needs to perform at, knowing that if you hit that level, that's the level of savings or growth opportunity that you're generating on the other end. Yeah, because again, what is AI and what is all the fear and excitement about it is automation of human manual tasks. So whether you have a sales team doing something over and over again, whether you've got junior legal petitioners drafting agreements over and over again, or in the courses or any other case. The point is, if a machine can do it with high accuracy and consistently, then you don't need those, you don't need as many employees or you don't need to pay as much.

37:23And you can do the maths on the machine because I think opening eye charges, depending on the model, but you're looking at anywhere from approximately$0.04 per, I think it's a thousand tokens, which is 750 words. And there's other mechanisms to save costs. We talk about cash in and stuff like that. So yeah, it must, because otherwise that means that the client didn't set their goals. Their goals weren't clear. Part of the setting the goals must be what does success look like? Awesome. I love that. Since you started talking about money and more the business aspect of this, roughly how long does a process like this take?

38:05And obviously it depends on the amount of data you have, but let's say to a small to mid-sized company who has, let's take a very simple use case that we talked about before, which is customer service. How long should a process like this take from starting to work on this to we have a model that we've tested internally that is mature enough that is providing the business result and we can actually start deploying it. So how long does it take? And roughly, how much does it cost to set up and how much does it cost to operate moving forward? Yeah. So if a client already has an idea of the data source they want to work with.

38:44So let's say the customer services, they have tons of website pages and maybe they have some other PDFs for internal docs. Yeah. And the ticketing system that they've used with all the open cases and closed cases and what was the process, like all of those things. So if they have that and then they want to use open AI, just keep it simple. Yeah. 46 weeks. Oh, okay. So it's not a huge project. no it's still yeah i'm speaking for my i don't know what other no i'm saying this what you gave is a great answer i was wondering if this is a month two months six months a year and a half so a month to a month and a half is a very reasonable time frame yeah because the main job for me is get get it because a lot of people just want to experiment with this stuff they don't really know what they want so my job is to provide clarity to ensure the data is processed properly.

39:42That's a lot of what I spend time on. Once all that preparatory stuff is done, we know what we need to do and it's a lot more straightforward. A lot of the issues come from clients not really knowing why. They're just doing it out of FOMO or they just want to keep up with competitors. They don't really know why they're doing it. Yeah. So. I have another question related to that. There are more and more no code, do it yourself kind of tools that do this process, right? Where I can go in and upload data to it or give it access to something and then get a chatbot that works out of the box. What are the biggest differences, things that I need to consider, good and bad, whether I want to do and use one of those tools or whether I want to hire a company like yours to do a more robust process?

40:33Or maybe what are the use cases or cut points where, okay, if you want to do this, you should go with this. If you want to go below the line, you should do that. Yeah, I think it's the same as any other industry or before AI, right? There's always going to be out-of-the-box solutions that you can use. But at the end of the day, if you want that personalized, high-accuracy solution, you have to go custom. There's no way around it. And I I know I've got friends who are running some of the most popular no-code solutions. It's impressive what they've done at scale, but on an individual level, you're just never going to have that.

41:16It's like food as well, right? You go to a restaurant where a chef makes food particularly for you versus you go to some mainstream chain, right? So what do you want as a company? And for me, for people who are new to the space, I think it's important they have a first experience that's mind-blowing, right? If they go and use some out-of-the-box solution or something that's just mediocre, then it proves their doubts. Because a lot of people are coming into this saying, okay, look, it's all hype, it's there. And so for me, I'd rather their first impression be something custom that kind of blows their mind.

41:58And then they get excited and they tell the stakeholders like, look, we need to bring this in. But if it's just some generic solution, which one is going to be cheaper for you. So I'll talk about the pros. It's going to be cheaper for you. They will, you can get up and running immediately. I'll say that's another benefit because it's a SaaS solution. Outside of that, the cons are, as I've said, whereas with the custom solution is definitely going to be a lot more pricey and it's going to take time because you need to be involved in terms of get everything up and running so basically what i hear you say that the out-of-the-box solutions are good if you want to maybe experiment knowing that the level of you know what i'll rephrase that i think if you're willing to accept a lower level of accuracy in the out, then it's a good enough solution, whether you just want it for experimentation or you want it for an actual solution.

42:58If you're in an environment or if your use case, if it works 70 % of the time, it's good enough, then perfect. Then you got a quick and dirty solution that you can do in-house without going through bigger project. But in cases where you want higher level of accuracy, which is probably most of the things in business where clients are involved, you probably want a more custom solution. Would that be a good way to classify the two solutions? Yeah, I would not even say it's up to 70%. I'm not saying this because I'm pro my stuff. It's closer to 50%. And also your options are limited. They focus more on simple stuff like PDFs, maybe what else?

43:42A website, stuff like that. It's more straightforward. Once you start wanting certain requirements. Maybe you've got some other types of your database or you've got some other types of quirks. You don't want certain names to show up, certain emails to show up. Maybe you want it to be in a particular design, for example. Maybe you want to, and a big one for a lot of clients is they want to be able to see the chat, the conversations of users and have control, basically. If you want more control, then I think that's the way to go. It's yours, it's your IP, you can do what you want. Last business question of this, we said it's about four to six weeks.

44:25How much are we talking about budget-wise that a company needs to be able to prep in order to say, okay, let's take this first step and see where it goes? Yeah, so on the lower end, you're looking at five to 10 ,000. That's just a basic plan. You can just imagine it as a lot of the stuff. It's similar to ChatGPT, but it's trained on typically PDFs and stuff that's from a data person point of view, it's not too crazy. But then you do get your own application effectively, authentication, all that kind of stuff. 10 to 20 is where you're starting to look at, you're more interested in a combination of two things.

45:10You're working with things like databases, APIs, things that are just a bit more complex that require more security measures. You're wanting to place a lot more emphasis on accuracy. So the basic one is just more, okay, building the solution for you. It's done, it's good. But then if you want to go from good to better, then you're looking at, okay, we need some sort of way to improve this. And then 20 up is when, yeah, now you're talking full on, basically full scale, working on ongoing evaluation, making sure that all the security boxes are ticked, making sure there's just a ton of testing. I'll say 20 plus is just usually enterprises, a lot of testing, a lot of security checks, a lot of trying to get as perfect as possible, I guess is the way to say.

46:04So obviously each of them require different resources from a developer point of view, which is why you have those different tiers. So I would caution anyone listening to not, I've noticed that some people try to shop the market and find lower and lower prices and all these kind of go offshore and do all these things. But then you risk having something that's not secure. You don't know what they've put in there. So be careful with that. I think it's very simple, right? You got to think about what is the level of risk that this can put you in, right? If this is an internal tool that's going to be an assisting something, okay, you can take some bigger risks and then maybe invest a little less upfront.

46:54If this is going to be basically allowing your customers to query your company in order to impact large business processes, whether it's sales, whether it's marketing, whether it's customer service. The future of your business depends on this thing working perfectly. And hence, you just got to define how much budget you're willing to put into this. And at the end of the day, I think if this thing works, if you applied it to the right problem, the ROI should be very easy to prove. And so it's invested$100 ,000 in this, but it's going to save you$1.5 million in manpower in year one, or it's going to make you$10 more million on top line in year one.

47:40It's a no brainer. It's a very easy investment to make. And so I think it goes back, and I love that we went full circle, it goes back to picking the right goal in where investing in building the right solution will yield a high enough ROI to make this a very easy business decision. Yeah, that's exactly it. Because again, I will say, as I said before, who are you replacing? Okay, let's say in the case of a private equity firm, you managed to replace two financial analysts. How much would that have cost you? We were talking about at least$200 ,000. So$200 ,000 versus the price range we're talking about, I think, again, like you said, it's investment.

48:28Awesome. Mayor, this was great. I think we covered a lot of stuff and we managed to do this without going too technical. So I think it's still a great conversation for any business person who's considering this and now can just go and understand the process and the steps and the pros and cons. And we talked about really a lot of stuff. I really appreciate you taking the time and sharing your information. if people want to follow you, learn from you, work with you, what are the best ways for them to connect with you? Yeah. So my website is enyaianalytics.com. So that's S-I-E-N-A-I analytics.com.

49:06That's the consulting website. I'm on Twitter as well. So at M-A-Y-O-W-A-O-S-H-I-N. So I'm pretty active on Twitter. I share the latest things I come across. I also recently launched a newsletter for leaders interested in building AI chatbots, which I think is our scene as well. Those are the three main ways. I'm also on LinkedIn as well, if you want to say hi. So yeah, I'll be happy to even provide just general strategy. If that's something that you're interested in for your business, I also offer consulting for that as well. So just AI strategy and how to set that up as well. It's a very exciting time.

49:42I know it can be overwhelming to know where to get started. So that's what I try and do. Simplify this stuff and take the noise away. Awesome. Thank you so much. This was really valuable. I appreciate your time and thank you for joining us. Awesome. Thank you. Great conversation with Mayo. And it definitely helps demystify some of the concepts and needs when it comes to training models and what you need to prepare or be ready to do if you want to go down that path. I definitely do not suggest that as step one for any business, you should start with low hanging fruits and gaining efficiencies by learning how to prompt better and using existing models that are out there across more or less every aspect of the business.

50:27but this is definitely a logical next step. And definitely if you're a bigger company and you have a lot of solid data across different aspects of the business, that is well-organized. And now to some exciting news from this week. As in previous weeks and specifically last week, there were a lot of really important big news. So I'll try to go through them quickly, but I will also add in the show notes, additional news that did not make the cut. So they're still big news, but they're not important enough for me to share through the recording of the podcast, but there will be links if you want to find out more.

50:58And we'll start with government regulations. Two senators, Bi-Parkisson, Blumenthal, and Holloway, are planning to introduce comprehensive AI regulatory framework. To do that, they've met with the leaders of Microsoft and NVIDIA and Elon Musk and Nadella and Altman and the biggest names in the AI world in order to come up with what the framework is going to include. The framework includes provisions for AI licensing and auditing and a new AI oversight office and definitions of company liability and privacy and civil rights protections and data transparency and safety standards and so on. The goal is obviously to reduce the risk that AI represents to the society while maintaining the innovation and the capabilities that we gain from it on both personal and business sides.

51:51It's definitely an important move forward and hopefully turning this into laws in the near future. There's obviously pros and cons in the fact that the leaders of the industry are the ones that are helping dictate the law because then it might become a self-fulfilling prophecy of what they want. I really hope that the lawmakers at the Hill take this a little deeper and maybe consult with additional people that are not from the industry in order to balance these people's views, but either way, I see this as a very positive move. Another regulatory action is the FTC is looking into potential antitrust problems in the AI space.

52:33They fear that incumbents may try to leverage their power against new generative AI companies, hence reducing competitiveness. This is obviously a serious risk because some of these big companies are gigantic and they control both software, hardware, engineers, and so on. So their ability to limit access to newcomers is significant. And the FTC I'm quoting, appears ready to intervene aggressively if they feel that anti-competitive actions being performed in or around generative AI space. This is already problematic if we are combining it with the previous piece of news, because just the fact that some of these large companies that has already violated multiple laws are now helping to write the new laws that will prevent from new companies doing the stuff that they did in order to train their models is already a problem.

53:29I see that as a positive move and I really hope that we'll be able to see a full ecosystem flourishing, including smaller startups, et cetera, that can compete with the bigger players, at least in niche specific markets or specific subjects. A few huge companies have released information about new AI capabilities in this past week. One of them is Salesforce. Salesforce Dreamforce conference happened this past week. And on September 12th, they introduced Einstein Copilot. Back in March, Salesforce introduced Einstein GPT, which was a smaller version of that, which allows users to use GPT capabilities to write and create content.

54:11but the idea behind Einstein Copilot is the ability to basically perform and query anything within Salesforce in natural language. As an example, a salesperson in a company can use this to research new accounts or newer customers, or a customer representative can look back and see what were results of similar cases of previous customers, or a product manager can create a storefront for a new product that he or she wants to launch. Really for almost any action within the Salesforce world, instead of the users having to ask questions from other people in the company on how to do that, or have to know all the set of clicks and menus that they have to go through in order to find the right buttons, send the right queries, et cetera, they'll be able to just ask for what they're looking for and generate whatever they need in a natural language while conversing with the Salesforce co-pilot.

55:10They also added what they call the Einstein trust layer with two goals. One is to reduce AI hallucinations and false responses, and also for data security. In addition, they've announced Einstein for Developers, which is a coding tool built exclusively for the Salesforce-specific coding languages, meaning developing new capabilities within Salesforce, will become significantly more efficient. My take on this is it's not surprising. Again, it continues a trend that we've seen from Salesforce and that we'll see from every other large software provider. Otherwise, they will lose market share to the competition that will offer these kinds of capabilities.

55:49And to prove my point, Ernest & Young, also known as EY, just announced that they have invested$1.4 billion in developing their own large language model and AI platform. They call it EY.AI EYQ, and it includes data management and use cases and framework analysis specifically for AI adoption. The goal of this is to position Ernest & Young as a global leader consulting company for large organizations who want to implement AI, which is basically any organization out there. This again, doesn't come as a surprise. If you look at their industry, other companies already announced this before. So peers like KPMG and Accenture and PwC and Deloitte all have made similar announcement in the past few months, all varying between one and$3 billion investment each.

56:40So they fall within the ballpark. And if we're already around the topic of large consulting companies, BCG, Boston Consulting Group just announced an alliance to deliver enterprise AI to clients with no other than Anthropic, the creator of Cloud2 large language model. So the goal is to give BCG's clients direct access to Anthropic's Cloud2 AI assistant, and BCG will be the advisory company that will help those clients implement these models in the most effective way. It's obviously a big win for Anthropik to be brought on board with BCG, one of the largest consulting companies in the world. But again, it falls in line with what we're seeing in the industry.

57:23I think we'll see more and more of these partnerships from the companies who generate the large, more successful models together with big companies who want to provide these capabilities and do not necessarily want to develop them from scratch, like we've seen in the previous example with Ernest & Young. The last two pieces of news relate to how effective is AI in different business components. The first of those pieces is research performed by professors at the Wharton School and the University of Pennsylvania. And they were trying to check who is better at generating innovative ideas, whether it's their MBA students or was it ChatGPT?

58:04and Chachupiti won by a landslide. And if you look at the statistics and read the article, they found the following. Chachupiti can generate ideas much faster and cheaper than students. On average, Chachupiti's ideas were a higher quality per survey. Chachupiti generated 800 ideas per hour compared to five by the humans. Now, this is obviously not a quantitative test by a qualitative test, But at the end of the day, when they picked the top 10 % ideas that were generated across human or Chachi PT, 87.5 % of the 10 % best ideas came from Chachi PT. And in addition, the cost per Chachi PT idea was 63 cents and the cost per human idea was $25 considering an average salary that they considered for these people.

59:01So ChatGPT won on speed, it won on quality, and it won on price. That's not very good news to anybody who considers themselves an innovation consultant or just somebody who is an innovator. The flip side of that, it will allow us as a society to innovate much faster, which overall is a great thing. And the last piece of news that has to do with productivity that AI brings is a research that was done by the Nielsen Norman Group. They did three different studies in multiple companies with the goal of evaluating the improvement in efficiency across different aspects of the business. Study number one was done around the topic of customer service.

59:45Study number two was done around writing routine business documents like sales proposals, et cetera. And the third study was around coding small software projects. The findings were not surprising knowing that it's going to be more efficient with AI versus without AI, but the numbers are pretty astonishing. On average, the productivity across these projects was 66 % higher than using AI versus not using AI. Also with a very big spread, customer service only gained 13.8 % improvement in answering inquiries per hour, which is what they were checking. Writing business documents, on the other hand, saw a 59 % increase in efficiency in creating documents per hour and writing code.

1:00:33So 126 % improvement when what they were measuring is how many tasks can a programmer complete per week. That more than doubled while using AI compared to not using AI. What does that tell us? Well, if you can improve the efficiency of things across your business by 50, 60 % or 100%, but let's take the average of 66%, one of two things can happen. Either you can find new clients that will allow you to use the extra capacity in order to grow the business, which would be amazing. The reality is that's not always the case. It's definitely not the case right now in a problematic economy. And it's definitely not continuously the case across all sectors, which means for the companies who do not have the flexibility to grow fast enough to make more use of the extra bandwidth that is generated, they will cut costs by letting people go.

1:01:36And anybody who says otherwise just does not understand how effective this technology is, meaning we will see significant job losses in the immediate future and definitely moving forward. By the way, I know I said that I've already covered the last piece of news, but it relates directly to this. Forrester's 2023 Generative AI job impact analysis anticipates that 2.4 million US jobs will be replaced by Generative AI by the year of 2030. I think this is totally underestimating it. I think the impact is going to be much bigger, and I think it's going to happen much sooner. But who am I to question Forrester's analysis?

1:02:15but they also anticipate that 11 million jobs will be influenced and will require retraining to work alongside AI. I 100 % agree with that. And again, I think that will happen way faster and at a much higher scale. The good news is if you're listening to this podcast, it means you're one of the people that actually cares, that understand this, that the world we know today in business is changing and changing fast. You understand you need to acquire new skills in order to learn how to work with AI in order to become more efficient, which means you have a much higher likelihood of staying ahead of the curve, keeping your job, or maybe even starting your own company or a better job within your company, because you will have those skills.

1:02:59That's it for this week. Have an amazing week. Explore AI, share what you find with the world, share it with me on LinkedIn, and have an amazing week. you

From the publisher

Could AI Really Transform Your Business Overnight?

In this episode, we have expert Mayo Oshin to reveal how you can leverage your existing data to unlock transformative growth.

We uncover the step-by-step process to implement AI in your business and start seeing results fast. From goal setting to data preparation to model training, you'll learn the full playbook.

Topics we discussed:
✅ Why starting with a clear business goal is crucial
✅ Taking stock of your data across structured, unstructured, proprietary and public
✅ Data cleaning and preparation best practices
✅ Model selection: custom vs out-of-the-box solutions
✅ Typical project timelines and budgets
✅ Evaluating success and tracking ROI
✅ Deployment strategies and measuring impact

Mayo Oshin is an AI consultant and educator. He helps companies implement AI to transform their businesses.  

About Leveraging AI

If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!

More from Leveraging AI

All 330 episodes
30 | Unleash the Power of Your Business Data: An Actionable Guide To Training Models Leveraging Company Data. With Mayo Oshin, Founder & CEO Siennai Analytics, a private chatbot development companyLeveraging AI · 1 h 3 min
Listen in VO