#187 Mohamed Elgendy: Systematic Testing for Generative AI Models with Kolena

18 May 2024 · 58 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Episode #187 Summary

Episode Overview

  • Title: #187 Mohamed Elgendy: Systematic Testing for Generative AI Models with Kolena
  • Host: Craig S. Smith, former New York Times correspondent
  • Guest: Mohamed Elgendy, CEO and co-founder of Kolena
  • Focus: Exploring AI quality assurance and systematic testing for generative AI models.

---

Key Concepts and Discussions

  1. Introduction to Kolena
  2. Kolena is an AI quality platform that assists teams in testing, validating, and ensuring the performance of AI products.
  3. Mohamed Elgendy transitioned from a software engineer to leading an AI quality platform, addressing the non-deterministic nature of machine learning models.
  1. Challenges in AI Testing
  2. Traditional testing methods for software do not translate directly to machine learning due to their inherent unpredictability.
  3. AI systems require a more systematic framework for validation to avoid blind spots and misalignment between teams.
  4. The discussion highlighted the necessity of structured quality assurance processes akin to traditional software development.
  1. Kolena's Framework for Testing
  2. Kolena aims to define a gold standard for AI evaluation by establishing a quality framework.
  3. The framework includes:
  4. Determining metrics and understanding risks associated with machine learning models.
  5. Establishing test coverage for various scenarios, such as autonomous vehicle perception under different conditions (day/night, rain/snow).
  6. Utilizing both automated and human evaluations for comprehensive testing solutions.
  1. Future of AI Quality Assurance
  2. Kolena is involved in organizing AIQCon, the inaugural AI Quality Conference, to promote new standards and practices in AI testing.
  3. The need for a systematic approach to AI quality is emphasized, moving away from reliance on individual scientists' intuition.
  1. Use of Human and Automated Evaluations
  2. The episode discusses the balance of human involvement in the evaluation process and the automation of testing metrics.
  3. Mohamed mentions plans to develop their own evaluation models that would enhance the testing process through reinforcement learning, allowing for user feedback to shape future evaluations.
  1. Insights on Hallucinations and Bias in AI
  2. Elgendy explains how hallucinations in AI can be categorized and tested, focusing on how to mitigate risks and ensure safety.
  3. Hallucinations are discussed as a potential risk, where models may generate inaccurate or misleading outputs.
  1. Collaboration with Industry and Regulatory Bodies
  2. Kolena collaborates with various industry leaders and regulatory bodies to shape the future of AI quality and testing standards.
  3. They aim to provide transparency and build trust within the community through their frameworks.
  1. Subscription Model and Customer Engagement
  2. Kolena operates on a SaaS subscription model, providing different tiers based on customer needs.
  3. They emphasize that ongoing testing and evaluation are crucial for businesses working with AI technologies.

---

Key Takeaways

  • Systematic Framework: Emphasizing the need for a structured approach in AI testing akin to traditional software QA processes.
  • Collaboration: Engaging with industry leaders and regulators is essential for establishing standards in AI quality assurance.
  • Iterative Process: Testing and evaluating AI systems require continuous feedback and improvement cycles.
  • Focus on Metrics: Developing a comprehensive understanding of metrics is critical for assessing AI performance effectively.

Conclusion

  • This episode provides valuable insights into the current state of AI testing, highlighting the challenges and opportunities within the field. Kolena's initiatives aim to advance the quality assurance practices necessary for creating robust AI systems, ultimately helping to ensure they perform reliably across various applications.

---

Listen to the full conversation on [Eye On A.I.](https://eyeonai.com).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So we are working with the industry, with our customers and with the regulators to put this standard out. for everybody. There's no sensitivity in that and it's in our best interest that it's copied. Everybody built on it. For now, our customers trust us because they see that we have a framework and they see the results from our work. We are planning to build our own eval LMs, so we think that quality LMs. And this quality LMs has two jobs. One is understanding what are the metrics that are used everywhere. We have millions of test runs running every month. on a wide set of applications, even within LMS, so a wide set of applications of that.

0:41Hi, I'm Craig Smith, and this is Eye on AI. In this episode, I speak with Mohamed Eljendi, CEO and co-founder of Kalena, an AI quality platform that helps teams test, validate, and ensure their AI products perform as expected. Mohamed shares his insights on the challenges of testing machine learning systems, the importance of having a systematic framework for AI quality assurance, and how Kalena is working to define a gold standard for AI evaluation. We discuss the role of human evaluation, the potential of using adversarial models for testing, and the future of AI quality standards. Kalena and the MLOps community are joining forces to host the inaugural AI Quality Conference, or AIQCon, the first conference in the industry to focus on advancing new standards and best practices for AI model testing and quality.

1:44Taking place on June 25th at the Hibernia in San Francisco, AIQCon will bring together leaders from Cruise, NVIDIA, Google AI, Anthropic, U.com, and many more, as well as top VCs and journalists to discuss how the industry can consistently build high-quality AI systems, prevent model failures, eliminate biases, and drive greater adoption. I hope you find the conversation as interesting as I did. Hi. Good tech solves problems that you've thought about. Great tech solves problems that you haven't even thought of. What can the commerce platform trusted by millions of merchants do for you? It's time for Shopify.

2:31The commerce platform revolutionizing millions of businesses worldwide. Whether you're a garage entrepreneur or IPO ready, Shopify is the only tool you need to start, run, and grow your business without the struggle. Shopify puts you in control of every sales channel. So whether you're selling satin sheets from Shopify's in-person point of sale system or offering organic olive oil on Shopify's all-in-one e-commerce platform, you're covered. Shopify powers 10 % of all e-commerce in the United States. And Shopify's truly a global force, powering Allbirds, Rothy's, and Brooklyn Lennon, and millions of other entrepreneurs of every size across over 170 countries.

3:19Plus, Shopify's award-winning help is there to support your success every step of the way. So sign up for a$1 a month trial period at shopify.com slash IonAI. That's Shopify, S-H-O-P-I-F-Y.com slash I on AI, that's E-Y-E-O-N-A-I, all lowercase, all run together. Go to Shopify.com slash I on AI to take your business to the next level today. Give them a try. They support us, so let's support them. AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power.

4:11So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs. OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle. So now you can train your AI models at twice the speed and less than half the cost of other clouds. If you want to do more and spend less, like Uber, 8x8, and Databricks Mosaic, Take a free test drive of OCI at oracle.com slash IonAI.

5:05That's E-Y-E-O-N-A-I, all run together. Oracle.com slash IonAI. That's oracle.com slash IonAI. Well, thanks, Greg, for having me. A bit of background, I'm Mohamed, CEO and co-founder of Colena. Before Colena, I'm a technical founder. Before Kalena, I was in machine learning leadership roles for a big part of my career. I worked at the AI teams at companies like Amazon, Trilium, Rakuten, worked on different set of machine learning problems. Before that, I was a software engineer, and my education was in biomedical engineering, so software and hardware engineering and healthcare field. Okay. Yeah.

5:51Tell us about Kalena. How did you come to start it, and what does Kalena do? Yeah, so things started with us when, well, Kutena is an AI quality platform that takes care of the testing and validation and ensures that the team is building a product and AI product that they know exactly how it's going to perform and they have tested it regularly in a systematic way. things started with us before that when we like a few years ago many years ago started we were working on building AI systems we started noticing that every time I worked with a team we have to build some kind of solution for testing the reason here is machine learning is not a deterministic process or not a deterministic product these algorithms they define the logic as opposed to engineers defining the logic in software So that ends up making the testing process very tedious and time-consuming.

6:49And more importantly, it's an ad hoc process that heavily relies on the scientist's intuition and their experience in machine learning and experience in the domain and in the company. Because they need to remember, oh, we failed in these areas last year. We failed in these areas this year. So I need to look into that. The nature of this process being not a systematic process and a very ad hoc process leads to a lot of blind spots in the test, leads to misalignment between teams. And then the end product is something that nobody knows whether the scientist is building, the product team that is actually speccing the product, or the user.

7:28Nobody knows really how it's going to perform in all scenarios, which led to the term that's known as long tail wedge cases. Now, this is compounded now by the generative models that started now, where it's trying to achieve more of a generic purpose that makes it really, really hard to start relying on the scientist situation. it was really hard before and now it's becoming impossible and we're seeing in the industry how things are going uh with the we'll probably talk about it if you bring it up like with all these models and how things are going uh so we with that said we have uh we started with uh looking into how we've been building models whether it's language models vision models uh structured uh data or time series like recommender systems and structured data side we started seeing that okay there's a need happily for us as builders to have a framework first of all not even tooling i think what is that framework for testing this happened in software which is machine learning is software so what happened in software when it was started early beginnings of say three decades ago or more more in the early beginning of software everything was still manual and ad hoc and still were not reliable even at the end until the software community started thinking about quality assurance processes and frameworks and started thinking about, okay, we need to have some unit tests and functional tests and user acceptance tests.

8:50Now we have stages and everybody understands. So when you talk to people in this encoders working in software, it's clear. What stage is your software? And we're in beta. So we lost, but still in beta and we're expecting what happens there. That's not the case for AI. It's still we don't know really. The biggest question we hear from customers or practitioners in space, it's not like what's your tool doing? It's more of the first question is, how do we do it? What is that framework? So we're focusing first on defining that framework. Now, what comes after the framework is the tooling process of the infrastructure that automates that framework.

9:31So there are two stages in our head that we're working on in Kulana. One is defining that quality framework and then defining this whole standard of what quality means. What is astrophilic criteria look like? What are you testing? Here, if you're testing a computer vision model, you need to be testing, let's say it's an autonomous vehicle in a perception context. You are testing whether the vehicle or the sensors or the algorithm is able to identify these obstacles or objects. And now you want to test how it's doing for different scenarios, right? Like the occlusion, not occluded, day, night, rain, snow, and so on.

10:08And now you have defined what's called a test coverage. This is what it's called in software. So when you look at this, Mohamed and Craig are working on building their product, you see Mohamed is engineer, Craig is a product manager, for example, and we work together and define what these functional requirements should look like as best as we could. And now they are defined, everybody in the team start, you start having these alignment and trust with everybody in the team who understand what we're building. So you built the specs for the first time in machine learning. So that's the example in autonomous vehicles.

10:37and then you add another layer. How are you testing this? Not just the test coverage. What are the metrics? Is it the traditional metrics that we know about in machine learning, like the precision, recall, accuracy, F1, and so on? Or for the specific application, which in this example, it's autonomous vehicles, then do we need specific metrics? We call this product level metrics. You're not necessarily want to test. If you rely on recall, which is looking at false negatives, you don't necessarily want to test at this level are the recall metric, which is how many objects you have missed or the model has missed.

11:12But a more effective metric, which is a product level metric, you want to call it collision risk because that's the goal. And then collision risk is not just missing the object. It is a function of depth to how far this object is. This way, if you're looking at the collision risk as opposed to recall, which is looking at missed objects, you start seeing that this model A or system A is really good at identifying objects when they are in the collision zone. And then another one is which has better score for recall, but it's really poor at collision. So that's the example here. And that same thing for question answering for LMS.

11:55If you use the traditional language models, metrics like the blue and rouge and others versus custom metric, If it's a Q &A application, then correct answers to test factualness. So it's a function of BERT and BLUE and others and maybe a threshold that you want to set in there. That's what we call a gold standard, which is defining the test coverage. Part of it is some scenarios. Some part of it is data leak. Part of it is hallucinations and others and jailbreak and so on. So I divide the risk with NIST calls this risk management framework. So you map, this is the first function. You map these risks or the functions that you want.

12:32And then the second step is identifying what metrics you're going to use to test each of these functions. And then the third thing is, what is the path of your criteria? So this is what we call a framework to test. Once this is defined in the industry and inside the organization, now everything becomes a systematic process. And then you start evolving the process to more areas and evolving the framework. The models that you're evaluating or testing, are they proprietary models built from scratch, or are most of your customers fine -tuning open-source models or some instance of a proprietary model?

13:17I mean, what, or are you being used by Google and OpenAI to test the foundation models that they're building? Where in that market are you guys? Yeah, we are, we support, like we built Kulina to support the three layers that you mentioned. The teams that is building their own generative models, whether it's language or vision now. and then teams that are using these language models to fine-tune into the application because there's a layer here you need to find the best of these tools for your own application so you cannot just rely on these benchmarks on the internet and you can test it on your own data and then the third one is these foundation model providers we don't have like we don't work with OpenAI and Google specifically but we work with the other two layers and some of them are building their own foundation models for their own applications or they are tuning these applications.

14:18But the step before tuning is what is the best one to start with? Because there are strengths and weaknesses that are different for that. So you want to test it on your own, figure out which is the best out of the tens that are available now, and then guide the fine-tuning process. Because fine-tuning process here, you will want to send data to the provider and here retrain on this data. So you need to understand where the weaknesses are. And this will dive deep into this when we dive into our framework. But we serve these two layers as of now, and we support all the layers. Yeah. And when you talk about coverage, it's not like unit testing where you have a code base and you need to cover a certain percentage of that code base with unit tests.

15:06What do you mean by coverage in this context? Yeah, that's a good question. We've been going back and forth on what is the term because you mentioned unit test. We call it unit test. We understand the confusion here that might cause unit test in software. You'll talk about code coverage. So you're testing the code. unit test we mean here by and on the machine learning side scenario level testing you are not looking at that aggregate metric you wouldn't say llama versus gpt4 llama 2 versus gpt4 90 92 or 93 that is what's done now which is the aggregate metric unit testing is looking at every scenario how is this model doing for hallucinations and then you break you break have no solution down into these stages.

15:52And how is this model doing for jailbreak? How is this doing for data leak? Because data leak here is defined, even like, for example, differently, or even bias. So bias can be gender bias, can be race bias, can be just bias for the kids or in healthcare application. We have one of our customers building their own gen AI for that. So bias here, you can be biased towards a department in the system versus other. So bias here will be defined based on the application. So this is usually what goes with that. Yeah. And how do you do that? I mean, I was just talking to somebody about whether or not large language models will develop a true understanding of the world.

16:43and what they were saying, which was interesting, this researcher, is, yeah, they probably will, but we'll never know for sure because you can ask it a bunch of questions and it can give answers that demonstrate understanding, but you don't know if you've asked the right questions. If you ask five more questions, it may break down. And so you're never going to have a certainty. So how do you deal with that? I mean, you can, in scenario testing, you can come up with, you know, a dozen scenarios and test them. But there may be a 13th scenario that you didn't test in which the model breaks down. Yeah.

17:37Yeah, so I read your article on Forbes last week about the understanding, and I totally agree with his thoughts. Because we need to define what's understanding first, before we say whether these models would understand. From the perspective of, so what these large models do, or in general deep learning, they start painting or drawing this mind map, if you will, if you try to visualize it. And if you have an object like a car, it knows that it takes, when it's training, it takes, okay, a car is an entity. It doesn't know what a car is, but it's an entity. And it falls under the vehicles, types of vehicles, and falls under like a means of transportation.

18:18And then within vehicles, you break it down into cars, trucks, buses, and others. And then within cars itself, you start adding characteristics of cars, what they look like for wheels and so on. And then you start adding how to drive. And then from that context, it understands, right? Maybe, probably, definitely, it recalls more information. So from the information perspective, more than a human, because within this entity, all types of cars are there, better in specific areas and how it respects the car. So from that perspective, they do understand, and they will understand or recall information or put the map of information about these entities better than a human.

19:00Now, whether the understand is something that we need to define, like exact language or intelligence, what is that? So we need to define what this is. From the perspective of a data scientist building a product or testing a product, it should be the same. The idea of how much it understands that this is a car, how much it understands and predicts the next token and the next word, from us, we want to understand based on the product. and we work with our customers instead of going ahead and saying, okay, here's a model or here's an algorithm, then we test it. That is a very big task. A QA is not something that software just started.

19:38Every product we have has some QA process. It doesn't go out to consumers this way. So with that, you need to define the specs first. This is what quality is built to specs. You cannot say this$1 pen is higher quality than$100 pen because it's built for that purpose. So it is high quality. So the way we approach that is we work with customers and start saying, let's build two specs. You are building a, we have a customer that's doing like really, really exciting work automating the drive through process. You know, you go in and you give your order and then actually several of those because it's a very hot area.

20:19Now, if you start and they are in the applications is based on what you asked before they are doing that. They are testing these LLMs and finding the best and then fine-tuning internally and through the provider. So when you start thinking about this, you can't just... How are you going to test it? Because you need to collect data for this kind of test or prompt it as you see results. So you need to think about first, okay, how much is it able to extract the order that I'm ordering? And how much does it understand confusion here? because if you are going, give me a burger, you don't want to be confused by something that's outside the menu, like a burglar.

21:01That doesn't make sense. It should be able to map that somebody see the burger and map it to the menu. Now, that menu, this process is called Rack, which is the retrieval augmented generation. So you want to build first two specs. You define, okay, I want to be able to extract what the order is. I want to be able to handle corrections when somebody or having their customer changing their mind. So now you start building the specs. Now with these specs, and then you want to handle angry customers. So that's another spec. So with that said, you will build your product, your first version based on these specs.

21:35And then you throw it out to your customers and handle those, obviously, the biggest risks that you're worried about. And then give it out to customers. And then it's an iterative process. It becomes the monitoring piece, which our solution involves three main pieces. The data quality for training and then pre-deployment test and then post-deployment evaluation and monitoring. And then the shortest answer that I'll give you is it's an iterative process that starts with, like any product and software specifically, it starts with defining the risks and functional requirements, build the metrics and test coverage for that, and then start testing.

22:10And then now you are systematically chasing the long tail and you're not relying on data scientists in the company with their own experience and leave and come back and build memory, this whole thing becomes a systematic framework. Right. And how much are humans involved in this? How automated is this? So there are several stages. We automated the biggest parts of most of them. So the first part is the data quality, understanding the quality of your test, test data. some known properties of quality of data, like missing values or interesting differences in values and so on. So you clean up, you make sure that you have hyperdensity tests.

22:56So you want to make sure, okay, my test is this way. And then within that, the functional performance that you have just defined, there is the area of identifying blind spots. So the long tail of edge cases here, you have a customer that's been working on a problem for a couple of years. So we have a long list of these functional requirements for the experience. Now you have another customer that just started today or you have another team machine that just started today. You have a much smaller background and much smaller knowledge of the system. So the way we automate this is we are looking into these applications domain by domain and building a library.

23:33It's a as a recommender system, a library of all the, what we can see as known now across the industry, not known for your own company, all the known edge cases across the industry. And it comes as a recommendation system. So you are building, again, this drive-through application, automation application, and then you will have to find 20 of those and you find when I recommend another 40 articles and mapping them to your data. The example of this, like you are looking at it, Like I said, an angry customer and defining an angry customer as how to handle that. So you want to, another customer that is, think of this.

24:08So we wrap this and we tell them within your data, you have 100 ,000 pumps. Let's say, for example, we saw that there's 15 ,000 pumps that are representative of this angry customer. We grouped them for them and tell them, take that or not. And then we can ignore that this doesn't apply to the reputation. So that's the data quantified. Make sure that you have covered as much as you, the industry collectively has known about these unknowns and loans to speed up this product evolution. And then identifying points whilst increasing these test cases. Now, the second, the train, if they train internally, like OpenAI or others, for the fine tuning, and then which is outside Kulana training and everything is outside Kulana.

24:52And then we test on Kulana. So the first component data quality, we talked about it, and then the testing part here, we start defining what are the test coverage, the test cases in test coverage, we just talked about that. And then what are the metrics here for each one? Now, these metrics here, matrices, this is where we get into a fusion between automated matrices and human eval. So, for example, we all use several tools like the GPT tools, question answering tools, if you will, chat GPT and others. And then we have our preferences. You know what, this tool always answers better. It's a situation, this other tool talks more and doesn't follow prompts, for example.

25:42So that is more subjective. And so some test cases or some testing scenarios will require human eval. Now, that human eval here will look at the framework for this. And then you're able to define the questions that you want to ask Craig and your users or your internal testers. And then they go these lists of results for the models and one or more. And then give some eval. The system learns that you are looking for data leak. And in data lake here, you mean you don't want the person's name or the person's age to be retrieved if you're reading from your HR system folder or anything. So having defined the metrics, and a big part of the metrics can be automated.

26:26Another part of the metrics has to be human evolved, but it's then defined now. And then the third thing is defining pass or fail criteria, which is the last thing here. Pass or fail criteria here, it's impossible for a non-deterministic system like machine learning models in general. to tell it, okay, when you reach 90 % automatically deploy, or when you reach 99%, you cannot say that because there's no 100%. So now with that said, there are different ways that you want to define passive-critical theory to get the full automation of that evaluation. So the several approaches that we have is you define critical scenarios or test cases, and then you say, okay, then for any models that I'm updating now, pass or success when it's not regressing on these critical scenarios.

27:18We can never deploy without any regressions on these scenarios. Now, that kind of automation allows us in the training, which is outside our product, that was outside Polina, allows us to throw out 50, 100 or so experiments, push them all together in training, and then automatically when the model converges, where the training and validation set started to agree on the results, then you go ahead and run a test on Kulena. And instead of looking at these hundreds of models, how they're doing, you're looking at only the five or ten that did not fail in your passive criteria or did not progress in your critical cases.

27:55Now, this automation allows faster iteration in training and faster validation process. Right. On the human evaluation, you say you have a framework and then do you plug into some sort of a crowd sourcing platform like Mechanical Turk? I mean, where do you get your human evaluators? Yeah, so that's the plan. As of now, the last two years we've been working with select customers. We're actually about to announce something next week. I'll be more open. But for now, our customers are AI-first companies and they have the resources and they want to evaluate their models by specialized people in the domain, like we are in the healthcare domain and so on.

28:47So with that, for now, our customers are taking care of this quote-unquote labeling, but it's actually called evaluation. Now, soon enough, we're planning probably sometime in May or around May Q2 timeframe, We will start providing these, whether plugging into MTurk from AWS or others, plugging into this human evaluators for other cases. And then applications that are in production, that you have ChatGPT as an example, now we can utilize this massive user base that we have and then build in the feedback based on the customer's response to the prompt. or they give feedback, but it's much more better feedback than thumbs up and thumbs down.

29:34Right. And on the automatic side, is each evaluation strategy unique? I mean, or do you have templates at work? a customer comes in, they have a particular problem, and you have an evaluation protocol that off the shelf that will apply to that? Or do you have to craft a new evaluation strategy for each customer? How general can these solutions be? Yeah, so it's a hybrid of a automated template. It's not really a template. I will explain it a little bit more now. And configurable by the customer to add their own metrics. So the example of that is, let's say we have a text summarization problem.

Read the full transcript

30:41Asking judge of need to summarize news articles or something. Now, every data scientist will think about the category of the articles and sports versus technology or others, maybe the level of the language that's been written, they will think about their product perspective. But what Colina knows is there are properties for a text summarization problem. So we generalize by the type of the problems of text summarization here. There are properties that affects the sensitivity of the models, and it's impossible to be determined by just the human data scientist's eyes or something like the length of the article.

31:17This looks simple, but it heavily affects the summarization process. And now we present that as a recommendation to our customers. When they come in, they see the length of the article, the language, the attitude, and so on, and the sentiment. And then they see that, and they see the distribution of their test data on that metric or that property. Now, you start seeing, okay, for text, for length of the article, for example, then you start seeing, oh, all my data are shorter articles. I have a blind spot in larger articles. This is when they start knowing that. So, when they come in with their own properties that they know about, with their own problems, Kulena provides the unknowns, unknowns, that are always feeling that are coming for your specific problem.

32:06So, that is for text summarization. Another very obvious one that took us a few months to debug when we were in a previous product we were building is in object detection, the bounding box. Now, we are always looking at, okay, the bounding box, let's say if you're doing face recognition, that's my face and your face here. So that bounding box, we just look at, okay, are the features, if the model failed in Mohammed's face, then are the features clear or not? Are the lighting, the contrast, the noise? So, and that's valid. But what we don't see is the aspect ratio of that bounding box. It's one of the main reasons these models fail.

32:45Now, aspect ratio here can represent depth. So if my face is here and my face is right there, then you start seeing the aspect ratio grows heavier. And now if your model hasn't, even though the features are clear, that eyes, nose, and so on, but the model will fail, there's a chance it might fail in this area which hasn't been trained in that size. That's the interesting part about neural networks. They are very powerful, but they're not very intelligent, really, because they get confused by these nuances. And so that's the bounding box. And then, obviously, the number of bounding boxes, if you have 10 people on the frame there, all the features are clear.

33:24But that affects the model sensitivity versus two or three people. Now, these nuances are, this is what our R &D team is always working on, understanding these reasons of failures across different domains, and they come out of the box as a recommendation for our customers when they look at, okay, my problem is a bounding box problem, is a text summarization problem, question answering problem. They start seeing the data analyzed on that scale, on these properties that we have defined, and then create test cases based on that. You know, there have been some really high-profile failures of evaluation.

33:59The most obvious one is Gemini with its vision, I mean, image generation. What do you think went wrong there? Because that to me just seemed like such an obvious problem. It's hard to believe that they didn't catch it. Yeah, I mean, actually, the answer is in the given, because I read something maybe yesterday, Sergey Brin mentioned that we needed much more testing of the models. And then I read the actual official statement by Google. And it makes sense. The intention is clear, is good, which is making, hey, we are optimizing for reducing bias, right? but now optimizing for reducing bias means here we went so far, which ends up being a different kind of bias.

34:59That's what people should know. Now, that is a different. NIST has published this risk management framework probably 14, 15 months ago and it aligns very well with our framework that you need to identify the risks first. Now, when you identify the risk, we say we're going to build to specs. So if you build to specs and say, I want to build against bias, raise racial bias, which is the example that we're referring to for Gemini, then we want to test against what's the passive criteria, right? And then this is one of the things that we provide at Qlena recommendations of, you have to provide positive and negative examples.

35:34So that's the framework, right? So positive and negative examples. So if you, let's say if Gemini was using Qlena, they would put in the models and they put in, okay, bias. And then we noticed that, hey, your racial bias dataset doesn't have negative examples. You're testing that, okay, give me 10 people, give you 10 people of different races. Have you tested the negative example of this? Which is, again, part of that framework of how to construct that high-fidelity test case. And then again, after that, passive-fair criteria and metrics for this. So I'm not familiar, like I don't work with Google or the Gimini team specifically, and I'm not familiar with the internal testing process.

36:15But my best guess is it has been done with the way that it's done in the industry now, where it's like a group of brilliant engineers recalling previous failures from the head or writing them down in some internal week and manually testing that, which is very, very prone to human errors. because you're relying on their own effort in identifying these requirements, constructing the high fidelity test cases, defining the metrics, and actually doing the job of evaluating that. Isn't there a way to have sort of an adversarial LLM, not necessarily an LLM, but a large model, in the way that AlphaGo learned to play chess by playing against itself, you could have another.

37:14And I know that there's research into reinforcement learning with AI feedback. Do you guys do any of that? So, you train a model to ask questions or generate prompts, and then the model itself can evaluate how well or poorly the model that you're evaluating does. Yeah, so this is a practice that started getting popularity in the last few months. We've seen it working very well in cases and others are not. So to take a step back, any approach that we are recommending or putting as a feature on Qwena or as a framework in general, we take very high responsibility on that. and we do a lot of homework before we actually push it out because we are a platform of trust, right?

38:17So we have to actually, our customers look at the results and trust that. So the approach is this all relies after you have defined the framework, how are you going to validate your models against these metrics that you define? Now, this is a hybrid of, like I mentioned, of automated metrics and human, I said, if using that with human evals and then the third thing is to be one of each is using lms lms are really good especially when they are trained for that purpose are really good in identifying which paragraph if you're against still summarizing paragraphs which paragraph is more succinct in defining the area which one is is more talker you can be building for each you feel like grot is building to be a walker kind of platform or a cool responses or something.

39:09So again, once you have defined the metrics that you've been used to evaluate now for each metric, here becomes the diversity of things because the metrics are different. As of today, we're supporting automated metrics and we are recommending these metrics for others and covers a large part of the risks, especially because that's the highest area. which we're focusing a lot on these automated metrics and the human eval workflow. Now, LMs, what we're planning to do, we tested that on current LMs. We saw a lot of gaps there because they're not built for that purpose. But we are planning to build our own eval LMs.

39:49So I think quality LMs. And this quality LMs has two jobs. One is understanding what are the metrics that are used everywhere. We have millions of test runs running every month on a wide set of applications. Even within LN, there's a wide set of applications of that. Then we are in the best position to understand what these metrics are and how effective they are and how they are defined for pass-through criteria. We train our own LNs to identify. Think of it as a manual federated learning. We don't really have a system because customers' data cannot see each other, and obviously the whole privacy part.

40:33But we're able to understand what are the common metrics for these specific problems, and then build models to automate that. Now, that's step number one. Step number two is adding reinforcement learning so it becomes specific to a customer when they are, let's say, Craig is working today on a problem. and then you added a metric and uh in the middle of the problem you add a metric that say how good the model is in following products right following now that following prompts is a metric that can be generalized to all your question answering now but you want it here for you just added it today it learns by the by the votes that your customers are giving or it turns by the the way that you're testing uh however you're doing it and then it goes ahead and reevaluates the rest so instead So if you're going through millions of data points, you give the first thousand or two, or however the ratio looks like, which is an approach that has been known in labeling before as active learning.

41:29But this is again different applications of that, is being able to take the feedback from the customer evaluation live and then recommends the feedback, recommends the evaluation for the rest of your data. And so you're building an LLM that will do that. When are you likely to deploy that? We're in the early stages of this because, again, it's defined really by the market more than that, but by the technology. We have seen so far the biggest areas that we cover now covers the vast majority of risks because we're focusing on risks really because the Gemini issue with this, like we mentioned, is a risk.

42:11It's not just a failure in functionality. It's actually a risk. So we're focusing now on putting into that framework all the known risks across the industry, across different industries. So we're focusing on risks here. Most risks are deterministic by functions. You can build deterministic functions or build metrics for that. And then as we go, we go into dive into the functionality or the factualness. because factualness here you can say if you have a reg system that's reading from your hr the records you can say how long muhammad has been in the company it can tell you muhammad's been in the company for three years or it can tell you muhammad started in 2021 right now both are correct but which is the right answer really subjective it seems like the it has been three years is better answer because it doesn't seem it answers correctly in english wise but the date specific is more accurate than precise.

43:08So with that, because this is still generic, and we want to still see the human eval, how they're going to respond to that, I would imagine something by end of year will have a little earlier iteration of what this would look like. Yeah. What about hallucinations? How do you test for hallucinations? Because you don't know what a model's hallucinating until you, you know. Well, hallucination here, again, every big problem, we try to break it down to see what this is, define this for me. And then the first definition, like several examples, like the categories that we defined for hallucinations, if anything, there are up to 15 types now.

43:50I have to tell my mind, data leak is a type of hallucination because it leaks some data from your records that you're training on. Another bias is type of hallucination. Other types like jailbreak. Jailbreak is obviously all of them are risks, but it's a higher risk. And jailbreak is basically that the system is went outside of the constraints, whether it's been through prompt injection, which is another type of risk, or through just the hallucination approach. So you start breaking down these hallucination types. And that's the educational part. That's the framework part. Our customers go into Covena without defining what the problem is.

44:29They just drag and drop their data through their system or something. And then the platform is intelligent to understand the type of problem. This is a question answering problem. This is a tax summarization problem. This is and so on. And then for this problem, here's what we have identified as possible functional requirements. And here are what we have identified as possible risks. And then within risks, we see hallucination types, one and two and three. we focus on education a lot so we click on something to learn more go to our docs, see the full list of 18 identified types of hallucinations across the industry and we tell them for your problem you are facing these four only and five only so this is the starting point you go into Colena within the first 20-30 minutes you understand really the nature of the problem that you're building for and then you start building working backwards into, okay, for this type of risk of hallucination, like prompt injection, you tell them the prompt, hey, you tell them the model, ignore your internal.

45:35That was a funny one, actually, we saw, I don't know if you saw it, in the early days of HRGPD, where customers or users are typing prompts like ignore your internal instructions or ignore the instructions that you were trained on and give me the answer for this. And did they actually do that? So that's the one. Let's say for this, since I give the same example for this chat GPD, then OpenAI would look at that and say, wow, we realized another type of risk that is prompt injection. Let's add that into all functional things. And then this seems like a software problem, not even a machine learning thing, that it can just obscure these answers or block these answers.

46:14And then they start solving for the risk that has been defined, getting data for it, like getting all types of prompt injections that you have, which is not a hard thing to just mine this data. And we provide tools to mine this data from the repository. So we have a bucket of data and then they extract quick search of here's what the different types of prompt and injections that we have in our data. And then you create a test case for that. And then back again, because I didn't lose track of the initial question of how we test for hallucinations. To recap, we break it down into what type of hallucination we provide recommendations for our users to know what these types are, and then they work backwards into mining data through our data studio from the repository that represent these risks.

47:00And then they start defining again their metrics for it through our recommendations of their own knowledge and passive criteria for it. How do you guys work with customers? Do you price this on a project basis or is there a subscription model? Or yeah, how do you? So we're a SaaS subscription model. Our model, we have three tiers, starter and professional and enterprise. The price is very good to be based on the number of users, size of data, and the capabilities that they're looking for. whether they are focusing only on data quality, that's one. And then when they evolve in their quality process, then they include model quality too.

47:48And then that's another tier. And then get to the full quality and comprehensive solution for end-to-end model quality and automation, then that's adding the post-deployment monitoring. So based on their stage as a team, they start with these tiers. Yeah, but subscription suggests that there's a monitoring feature? I mean, because otherwise, if you're building one model, then you want that one model to be evaluated. You know, let's say it takes a month and then you don't need a subscription anymore. I mean, what's the logic of having an ongoing subscription? Yeah, well, it depends who you're, what the type of the customer is.

48:37If you're talking about this is the builders, this is a machine learning team or an AI team versus the buyer of the AI, which is somebody that just wants to test some cases. Or our third category of customers are regulators and independent auditors. So for the person third, like for the builders and the regulators, it's an ongoing thing. No builder will be like, okay, I'm going to evaluate once because you're always evaluating and iterating. Now, for the buyer or the user of the model type, we are focusing now on companies that is using these LMs. So it's B2B SaaS. We're focusing on the B2B side of it.

49:18With these customers, I imagine when things become much more democratized, then now we're at the early stage of this, and we're actually about to get into that phase. Now this is all businesses using AI. Others are not going to buy it too. I don't imagine you, Craig, you will go ahead and buy a subscription for Kulana to see which is best there. ChatGPD versus Lama versus others. But we're not focusing on those. The builders and regulators and even the buyers, as long as they are businesses, this is not a one-time thing. Even if you're using an LLM out there, if you're a buyer, you want to use Kulana to test which is the best LLM out there.

49:56And it doesn't take a month, by the way. It takes a few hours to get that out. And now, as an AI company that's using these LMs, you need to fine-tune it. And fine-tuning is basically, which is the best guidance for this in the industry is Colena. Because when you're fine-tuning, you want to know what data to send. So it trains on it. Now, you can go ahead and collect a lot of data with a lot of cost and time and budget. Or you will need to test on Colena using that test case approach or scenario-level approach. You know the exact scenarios that you want to collect data for. First, you know which was the best model, and then you collect the data for these specific scenarios.

50:36And that even fine-tuning process happens within a few days, not in weeks or months. Now, after that, this LM provider will publish a new model that you want to evaluate. And then the same thing on your data, your customers, it will start drifting. The model will start drifting. So it's always a business that's ongoing. We tell customers, this is not a question that comes often from our customers, but when it comes, we tell them, if you're stopping training, you'll stop testing. If you're stopping anything, you have even software. You don't buy software and then you don't build software and you fire the engineers when the software is out.

51:14They are still building, fixing the software and make it more stable and so on. So it's an ongoing process. It doesn't end with one test. I imagine there will be a sector of customers that come later down the road that are, I would say, more a B2C kind of thing. And we're not targeting those for now. You mentioned NIST. This is a developing field. Is there a standard developing for how models are evaluated? and tested so that people can assess whether or not you are the evaluator they want or somebody else. I mean, presumably, and you mentioned that Kalana is used by regulators and auditors. presumably they have a lot of different tools and frameworks available for testing and auditing.

52:24Yeah. Is there some standard coming out of NIST that you guys are adhering to so that if someone's looking for evaluation, you know, there's a seal of approval on Kalina by NIST so they know that, yeah. Yeah. There must be something that has to, like, there's nothing that we know of now other than the risk management framework, which is a, and that happened, the risk management framework was introduced to the world a few months before Biden's executive order. So it was not in response. In response, this thing started accelerating. but NIST created the risk management framework and it's three main functions or stages map, manage, measure you map the risks, identify them it's exactly the same steps that I just mentioned with our framework so map, identify the risks and identify the functional requirements and then measure them manage them in the sense of track them and then monitor them and making them available rule number one in quality is transparency see and understand everybody understand that and then you measure that so this is a more of a generic framework as a start to just start these deadlines um i don't have information if they are building a some kind of a standard for testing they must i know that this area that everybody is looking for uh now speaking of regulators or others who are looking or even practitioners who are looking for that gold standard uh where to identify which tool is better uh things are very very early in these efforts in the industry uh i believe we are one of the first ones to start thinking about it and and go back so we're we're helping we're working and with with regulatory bodies to us to share our knowledge and help with what we know and with our resources how how this can be shaped we are connect much more connected to the industry we have the resources that we're working with them uh so we would love i would love for this to happen for So now we are working on providing our help as much as it's needed.

54:30Whenever we're offering it, we're reaching out to all of them. I'm in conversations with most of them to start sharing this knowledge that we have and all the outcomes of our research. We put it out in public. It's in our documentation. If you go to docs.covina.com, you will find all the open metrics that we are updating that. So we are working with the industry, with our customers and with the regulators to put this standard out for everybody. There's no sensitivity in that, and it's in our best interest that it's copied. Everybody built on it. For now, our customers trust us because they see that we have a framework and we see the results from our work, and then we go ahead with that, especially that our framework does not, we call it the model party framework, and it doesn't contradict with the EU Act.

55:18It doesn't contradict with NIST framework. So I believe this is the most advanced that's available right now. That's it for this episode. I want to thank Muhammad for his time. If you want to read a transcript of today's conversation, you can find one on our website. That's I on AI, E-Y-E hyphen O-N dot A-I. And remember, the singularity may not be near, but A-I is changing our world. So pay attention. And before you go, give Shopify a try. It's only a dollar a month for their trial period, and you can sign up at shopify.com slash IonAI. That's shopify.com slash IonAI, E-O-N-A-I, all run together, all lowercase.

56:07Go to shopify.com slash IonAI to take your business to the next level today. AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs. OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle.

57:00So now you can train your AI models at twice the speed and less than half the cost of other clouds. If you want to do more and spend less, like Uber, 8x8, and Databricks Mosaic, take a free test drive of OCI at oracle.com slash IonAI. That's E-Y-E-O-N-A-I, all run together. Oracle.com slash IonAI. That's oracle.com slash IonAI.

From the publisher

This episode is sponsored by Oracle. AI is revolutionizing industries, but needs power without breaking the bank. Enter Oracle Cloud Infrastructure (OCI): the one-stop platform for all your AI needs, with 4-8x the bandwidth of other clouds. Train AI models faster and at half the cost. Be ahead like Uber and Cohere.

If you want to do more and spend less like Uber, 8x8, and Databricks Mosaic - take a free test drive of OCI at https://oracle.com/eyeonai



In this episode of the Eye on AI podcast, join us as we dive into the realm of AI quality assurance with Mohamed Elgendy, CEO and co-founder of Kolena. 

Discover the intricacies of ensuring robust AI performance as Mohamed shares his journey from a software engineer to leading an AI quality platform. Kolena is setting new standards in AI testing, addressing the non-deterministic nature of machine learning models with a systematic framework for validation.

The conversation navigates the challenges of AI testing, the importance of structured quality assurance, and the collaboration with industry bodies to shape the future of AI standards. Mohamed elaborates on Kolena's innovative approach to both human and automated evaluations, providing comprehensive testing solutions tailored to diverse customer needs.

We also explore real-world use cases and the platform's future plans to make AI testing more accessible and reliable. Mohamed's insights illuminate the path toward achieving high-quality AI systems, ensuring they perform as intended across various applications.

Tune in to understand the advancements propelling AI quality forward and how Kolena is pioneering the way in this crucial aspect of AI development.

Don't forget to like, subscribe, and hit the notification bell for more insights into the technologies driving the AI revolution.



This episode is sponsored by Shopify. Shopify is a commerce platform that allows anyone to set up an online store and sell their products. Whether you're selling online, on social media, or in person, Shopify has you covered on every base. With Shopify you can sell physical and digital products. You can sell services, memberships, ticketed events, rentals and even classes and lessons.

Sign up for a $1 per month trial period at http://shopify.com/eyeonai



Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI

More from Eye On A.I.

All 266 episodes
#187 Mohamed Elgendy: Systematic Testing for Generative AI Models with KolenaEye On A.I. · 58 min
Listen in VO