In short
Practical AI Podcast Episode Notes: Metrics Driven Development
Podcast Title: Practical AI Episode Title: Metrics Driven Development Featuring: Shahul Es, Daniel Whitenack
Episode Overview In this episode, Shahul Es from Ragas discusses the importance of a "Metrics Driven Development" approach in evaluating and enhancing large language model (LLM) applications. The conversation explores how Ragas provides tools for developers to automate evaluation processes, understand performance metrics, and generate synthetic test data.
---
Key Concepts
Metrics Driven Development
- Definition: An approach that emphasizes the use of specific metrics to assess and improve the performance of AI applications, particularly LLMs.
- Goal: To provide developers with actionable insights to optimize their applications without requiring extensive machine learning knowledge.
Ragas Overview
- What is Ragas? An open-source library designed to evaluate LLM applications, providing tools and workflows for developers.
- Development Background: The founders, Shahul and Jaden, have extensive experience in machine learning and were motivated by the complexities of evaluating LLM applications.
---
Discussion Highlights
Difference Between Evaluating LLMs vs. LLM Applications
- LLM Evaluation: Focuses on assessing the models themselves, often through benchmarks and performance metrics.
- LLM Application Evaluation: Centers on how well the model performs within specific applications and real-world scenarios.
- Application builders need to understand how the model interacts with their specific use cases, rather than relying solely on generic benchmarks.
Importance of Metrics
- Metrics allow developers to quantify the performance of their applications, enabling:
- Faster debugging
- Better decision-making related to application adjustments
- Understanding the impact of changes made within the application workflow
Continuous vs. Discrete Evaluation
- Traditional software evaluation is typically discrete, with clear expected outputs for given inputs.
- LLMs operate in a continuous space, where multiple correct outputs exist, requiring a different evaluation mindset.
---
Evaluation Tools and Metrics in Ragas
- Types of Metrics:
- LLM-based metrics (high correlation with human judgment but can be non-deterministic)
- Non-LLM-based metrics (more reproducible but less correlational to human judgment)
- Ragas provides workflows to help developers select appropriate metrics based on their specific applications and domains.
Synthetic Data Generation
- The ability to create synthetic test datasets using both internal documents and production data to tailor evaluations to specific application needs.
- Developers can verify generated data quickly, saving time on manual annotation.
---
Future Directions
- Shahul expresses excitement about advancements in:
- Tool use cases for LLM applications that can enhance user experiences.
- Improved frameworks and structures in the AI ecosystem based on practical experiences over the past year and a half.
- Innovations in evaluation practices that can standardize how LLM applications are tested and assessed.
---
Conclusion The discussion underscores the necessity of a metrics-driven mindset for developers working with LLM applications. Ragas offers valuable tools to facilitate this approach, ultimately aiming to streamline the evaluation process and enhance application performance.
---
Additional Resources
- Ragas Website: [ragas.io](https://ragas.io/)
- Assembly AI: [assemblyai.com/practicalai](https://www.assemblyai.com/?utm_source=practicalai&utm_medium=podcast)
---
Call to Action Join the Practical AI community for further discussions and insights on AI developments and best practices. Subscribe to the podcast for more episodes on AI applications, metrics, and implementation strategies.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related tech is changing the world, this is the show for you. Thank you to our partners at Fly.io, the home of changelog.com. Fly transforms containers into micro VMs that run on their hardware in 30 plus regions on six continents. So you can launch your app near your users. Learn more at Fly.io.
0:42Welcome to another episode of Practical AI. This is Daniel Whitenack. I am the founder and CEO at Prediction Guard. And I'm really excited to talk to another founder today in the AI space. Today, we have with us Shahul from Ragas. He's one of the co-founders. Welcome, Shahul. How are you doing? Hey, I'm doing good. Hi, Daniel. Hey, folks. Yeah, yeah. Well, thanks for joining at a late hour in India. Appreciate that. But yeah, I would love to hear a little bit, maybe for those that aren't familiar with what you're doing, maybe share a little bit about what that is and also maybe how you came upon the types of problems that you're solving with your current work?
1:33Sure. That's a very good place to start. So I bought Raghas. Raghas is an open-source library for evaluating LLM applications. And what we are trying to do with Raghas as an open-source library is to provide the developers for any engineers who are building LLM applications, the tools and the workflows that are necessary to automate or partially automate the process of evaluation through using different techniques, different methods and techniques that we bring out to Ragas. And how we came up to the idea of Ragas is basically me and my co-founder, Jaden, we have been working in ML for the past six, seven years.
2:16And when LLMs came out, we were already working with natural-in-it models. Jaden mostly worked on the inference and infrastructure point of it. And myself, I was working as an applied researcher. So we were practically, we loved it. We were practically doing a lot of experiments with it. We were part of different open source initiatives, building LLMs and also using LLMs to build different frameworks at that point. Even Langchain and Lambda index was coming out as one of the earlier frameworks at that point. This is around early 2023. and we were also working with different applications and stuff and RAC was one of the most popular applications that went into production very easily with LLMs and this is one of the first things that LLMs actually opened up as a possibility for enterprises to build something on top of them which can save a lot of time and money for that.
3:12So we were also building RACs and we initially, after a couple of experiments with some clients we found out that, okay, we are able to build these LLM applications, this RAG application using any LLM. It has different moving components. And going forward, there will be a lot more moving components to any LLM applications. It will be something like a common system where there are, as you can see now, it's not only RAG, now it's like two use cases. There are different versions that constitute a common system kind of stuff. And we thought, okay, why don't we build something some evaluation metrics, few evaluation metrics that can be used to understand the quality of any rack application that any engineer is building.
3:59And we also found out that going through these answers or going through these intermediate results manually is not a scalable approach. I could rephrase it to say that it's a very boring thing to do and nobody is very keen to do it. This is something everybody will push into someone else's responsibility, but it is something very important too. So we need to have identifying methods or coming up with solutions that could help developers do the evaluation, but also save them a lot of time while doing it. So it should not be something like, you know, it is basically an evaluation of an LLM application, regardless of RAG or any agency workflow is basically a very tedious process that takes up a lot of time.
4:44If you go through manually, you want to make sure that this tedious process, this manual process that takes a lot of time is cut down into one by tenth of the time, and it should get you the same insights as you do as in Baguilly. So that's how we came with our initial MVP of progress, and we released an open source library in the middle of 2023, and since then, we have been continuously iterating, and we have been getting organic growth and usage from there, and that's what we are going to do. Yeah, that's awesome. I noticed that you very specifically refer to the evaluation of LLM applications, not necessarily the evaluation of LLMs.
5:29Could you explain what might be the difference between those mindsets for people that are maybe getting into this? Maybe they have looked at benchmarks, let's say, or a leaderboard for LLMs. And there's a certain level of evaluation or benchmarking there. But then you're talking about the evaluation of LLM applications. So could you help us understand some of the differences there, some of what needs to be thought of at the application level versus at the model level? So this goes to the ideology of AA as a consumer product as of now, because pre-LLMs, most of the enterprises or most of the startups who are building AA-powered applications used to have their own models.
6:23And they obviously also have a team of data scientists who is building and managing these models. So in that spectrum, the building of the LLM itself or the model itself and also the evaluation of the model itself was the responsibility of the researcher, the engineer, data scientist who were working on it. But now it's not that people are building LLM applications or AA applications is not really building their own models. They are consuming models from an external endpoint or even open source models that are being released and building applications on top of it. Now, this actually forms a wide spectrum where at the left end of the spectrum, there are people building the LLMS itself.
7:08They have a separate list of targets, loss functions, metrics, and benchmarks to evaluate them. And towards the right end of the spectrum, there are people who doesn't really care about how the LLMS build itself. They only care about, you know, can this thing which I'm consuming using an API do what I'm trying to do here using this as a technology. Now, when a researcher or an organization like who builds LLM, evaluates LLM, they don't really know what is the exact use case for which this A or LLM is going to be used against or the type of data this will be used against. know when the evaluation is done at the researcher or an llm builder part of the spectrum they can they are limited to the evaluation or you know testing of this llm's capabilities on a general purpose basis they are not really tailoring this evaluation or testing to your application because they don't really have a control over what you are trying to build with it they can only say that okay this is a general capability of the model and we know that even the general capabilities of the models and benchmarks that are highly leaked into the training data.
8:21And even that is questionable, but that's a different question. This is what a search over LLM builder does at his end. But when it comes to LLM application builder, he can't really, if someone builds a rag or a tooling agent on top of an LLM, he can't really say that, okay, the LLM builder has set X and X accuracy, so I should get X and X. That would be a wild assumption today. So what we are trying to do here is giving that user, a application builder who is at the right end of the spectrum, the power, the tools to evaluate this application without knowing so much about ML or getting into so much jargon or anything.
9:02You want to make it as easy as possible so that as intuitive as possible, because most of the people who are building applications are not from an ML background, they are from a software engineer background or application building application background where their their speciality is building and scaling this application not building and scaling the models so we are trying to be intuitive as possible we also want to make sure that we do not take a lot of their time while doing the evaluation itself we want to make sure that their time is valuable and we want to make sure that we you know we do the heavy lifting of this evaluation for them yeah that makes a lot of sense and and And also I'm realizing there's this kind of spectrum that I was thinking about while you were talking.
9:43On the one side, you have kind of data scientists or researchers who are building out benchmarks or metrics for models. On the other end of the spectrum, you have maybe software engineers who are used to writing unit tests or integration tests for their software. And then what we're really talking about is integration of LLMs into software applications. or into certain workflows. You talked a lot about kind of this distinction between LLM benchmarks and evaluating LLM applications. Could you talk a little bit about the differences? Maybe there are software engineers in the audience that maybe they're used to writing unit tests and integration tests for their software.
10:28Now as a consumer, like you said, they're integrating some LLM functionality or maybe a reasoning, you know, a chain of reasoning with LLMs into their software, from a practical standpoint, what are the new types of things that they might need to consider that are different maybe from the way that they've unit tested in the past or written tests in the past now that they're working with these LLM workflows? Sure, that's a very interesting question. So when it comes to the application builder slash software engineer who is building with AI, as you said, most of them are already familiar with unit tests and integration tests that they regularly write for their software.
11:10software. Now, the major difference is that I also have observed many of the software engineers who are relating to evaluation using this analogy. Okay, you know, this is testing. I am praying I will, you know, learn or understand evaluation using the, you know, my understanding of the testing. So the thing is that this is a fundamental thing about analogy itself. So now you are using analogy to understand one thing but it could be that these two things can have different physical and chemical different properties here for example now uh for when when it comes to valuation and versus traditional software testing traditional software testing is mostly a discrete space where you know you have an input you have an expected output and if you give this input you are supposed to get this expected output from the software there is no variation in that that satisfies the test.
12:08Now, you know, if the logic is basically to add one plus one, you should get, if the whole logic is basically addition, if you give input as one plus one, the output should be two. There is no other possibility that exists that would, you know, satisfy that test case is correct. But when it comes to natural language, there is more of a continuous space. If you have an LLM that does the same thing, let's say one plus one, or addition, LLM could even say two in actual language or two in decimal point numbers. So both are actually correct. So here there are continuous space where the output cannot be exactly matched against or asserted against, but you should have an understanding.
12:53But there is a, in this continuous space of outputs that are possible, there's obviously a subset that is actually, can be regardless, correct? Obviously, if the LLM gives an answer as three in natural language or something, that means wrong, that is out of the bay, right? So, out of work, and out of work, work, and so what software engineers should really understand is that when you deal with ML, when you deal with, when you integrate A into your application, you should try to think about it as in a continuous space rather than a very discrete space or in a black and white manner. You know, it's yes or no.
13:26It could be that it is like in the middle of here. There are like a lot of space in the gray area here. So, you know, that is what a software engineer should understand. And then there is also a non-deterministic part of the whole, you know, A thing. Basically, if you have a system, which is part of core and part of AI, the system is going to be somewhat non-deterministic whatever you do. So this is also something that traditional software is not like that. Traditional software is very deterministic. So you have an input, you will have an output, given that the intermediate states remain the same.
14:02So here again, you have this non-deterministic thing to take up for. So that's also one thing. These are the two major differences that exist between software testing, traditional software testing, and this A application or COM1 A application testing that makes it different from each other. Yeah, super interesting. And I know in the software world, there's all of these sort of frameworks for development around test-driven development or data-driven development or these different things. I notice in, Ragas, in your core concepts, in your documentation, you talk about metrics-driven development for these LLM applications.
14:41So for those that are maybe developing LLM applications out there, could you describe a little bit the mindset of metrics-driven development, what you mean by that? And then maybe we can get into a few more of the details of Ragas itself and how you enable that framework. Yeah, metrics-driven development is a concept that we take from the test-driven development itself. So the idea we want to educate developers and software developers who are not familiar with metrics, but they are familiar with testing to what we are trying to do here. So as I said, metrics versus test is like metrics is something that delivers a value or that helps you understand the performance of some application in a scale of let's say zero to one or something.
15:33Now, Whenever you have an application, if you want to iterate or change, let's say you are bumbling and pull an application that consists of different agent work flows plus Iraq work flow, etc. And now let's say you want to change one single prompt or one single function code or something, or even the retriever, how would you understand the effect of this change in your pipeline? It's a question that metrics the vendor element trace your answer. So if you have a metric, if you have a way to kind of objectively quantify or understand the performance of your system before and after this change, you can also understand and analyze the systems, you know, responses or the behavior of the system using these numbers.
16:19For example, if you are switching out the retriever and let's say you add a set of metrics that effectively quantifies the performance of your system, you could actually switch out the retriever or switch out any function calls or something and then run the NREA pipeline once again for the given test set and understand the change in those metrics in particular dimensions. And, okay, let's say once you observe this change, you could again easily dig down. OK, you could easily fetch these samples for which the change is reflected the maximum. And then you could easily analyze and understand the manually.
16:58That actually helps and reduces a lot of time when it comes to debugging and testing these applications. So that's the idea of Matrix-driven development.
17:22what's up friends i'm here with a new friend of ours over at assembly ai founder and ceo dylan fox assembly ai is where you can turn voice data into insights chapters transcripts summaries and so much more with their leading speech ai models so dylan give me a glimpse into what you're doing with speech AI models at Assembly AI. So at Assembly, we're building industry leading speech AI models for various tasks like speech to text, streaming speech to text, speech understanding to help developers easily convert voice data, whether it's live or pre-reported into super accurate text. And then to help developers extract a ton of information and metadata around voice data or even around the text that they just were able to convert from that audio data.
18:09So these These are things like picking out entities or PII that was spoken in voice files or summarizing voice and audio data down into custom summaries. It's things like being able to detect how many speakers spoke and who said what and what the names of different speakers were. So we bundle all those things into a super simple API with really great docs that developers can just sign up to for free to start, use the API, build into their apps, and then build these really cool AI apps and products and workflows and automations on top of voice data with. I dig it. Okay. Can you take me a little deeper into the opportunity for developers?
18:48Because it seems like there's a lot of voice data out there and there's a lot of trapped value in that voice data. There's so much voice data being created on the internet now. Podcasts, videos, phone calls, voice messages, audiobooks, virtual meetings. It's crazy. And you can now transform and understand all this voice and audio data in ways that were not even possible a year, 18 months ago. So what we're seeing with the help of these new AI models that we're creating at Assembly, developers and organizations are just racing to build all these new applications, workflows, automations that leverage the voice data they have either within their organization or within their product to build really cool new products and services workflows that are just like taking off at the market.
19:35So at assembly, we're building the industry leading models for all those different apps and workflows, whether it's speech to text or speaker diarization or speech understanding capabilities to summarize voice data or extract entities, voice data or mask PII from phone calls for various types of automations that might be built. And we're exposing that through a super simple, super scalable API that's just constantly being updated and constantly getting better. And so we're seeing a crazy amount of developers and companies just build really cool apps and services on top of our API every day. It's really only just getting started, especially with the model updates that we have planned over the second half of the year that are coming out.
20:15They're really excited to launch to the developers on our API. Okay, constantly updated speech AI models at your fingertips. Well, at your API fingertips, that is. A good next step is to go to their playground. You can test out their models for free right there in the browser. Or you can get started with a$50 credit at assemblyai.com slash practical AI. Again, that's assemblyai.com slash practical AI.
20:58there's a lot that you've already integrated into ragas which are really interesting and maybe some people have heard of or some people haven't like faithfulness and context recall noise sensitivity aspect critique summarization score and more that i'm not listing i'm wondering if you could maybe just share a couple of those that you think are maybe most utilized or maybe interesting from your standpoint, just to give a sense of people, like what types of metrics are we talking about here? I think it would be beneficial to answer that in an abstract level. For example, if I say metrics, the whole concept of V-devicing or giving that out approach, these metrics is that it's not extremely hard to come up with these matrices, but it is when a software developer or an application developer thinks about evaluation, because he's already familiar with, you know, this concept of putting up metrics versus deciding the right flows and everything, it could take him like two to three days to figure out, you know, the right matrices.
22:04So what Raghers does is, if you are an application developer, okay, you are developing this particular type of application, you can come to Ragas and you can go to Ragas Metrics. And there you will find enough workflows, enough parts of the documentation that can help you navigate according to use cases, according to your requirements, that can land you up in the right metrics that you should use. And the documentation will also provide an intuitive understanding of how this metrics is calculated underneath. So these are the two value props that we provide there. It's not that a developer can easily, a developer or a data scientist could easily make up these metrics as their own.
22:47But sure, if it is mostly, if you are a software developer application, it could be very hard for you to figure out how to think around all these stuff. So we are effectively taking that load of you and providing you the ways to easily navigate and understand the right metrics to evaluate your application, tell it to your application. Now we are also expanding the restore metrics to use case and identity workflows. And we are also going to reformulate the whole load on that particular so that people can easily find the metrics. What are the metrics? It can be LLM-based or non-LLM-based. These are all, you know, whenever a developer comes to our capacity, obviously, ask a lot of questions or even think about metrics.
23:33He will ask a lot of questions like, for example, what is my use case? What are the parts in my application that I want to evaluate? And then what should I go for? LLM-based metrics, which has high correlation with human judgment, but also has its own lack of its own issues like, you know, non-determinism and everything or non-LM-based metrics, which are traditional. It has less correlation with human agreement or human judgment, but it is, you know, more reproducible. So all these questions, all these doubts both pop up in mind when another person thinks about metrics. We are trying to abstract all that into Ragas metrics and provide the developer a way to think about metrics.
24:15When they adopt a metric from Ragas, they also have a related list of items or features that they can use because they adopted the metrics from Ragaz. For example, let's say you designed a metrics for language English. Now, let's say you are evaluating for Portuguese or Spanish, you'll have to convert it to other language or whatever. Whenever you're evaluating different language, you'll have to do that. Now, if you're using a metric from Ragaz, Ragaz could take care of that. Then there are also issues like if you are building an LN-based metric, which is the trend here and most of the Ragaz because that has high correlation with human judgment.
24:59Then, okay, let's say the LLM, you are evaluating with LLM-based metrics and you are finding that the LLM-based metrics is actually performing. It's actually not performing very much in your case. You should have a way to allay in these LLM. People will have different expectations for different metrics. For example, let's say you are trying to do something like truthfulness. And truthfulness is basically the amount of hallucination that is happening in the answer given a context. And if you look at it, different developers from different domain has different strictness towards hallucination. I could have a statement like, I have a blue car versus I have a car.
25:45And I could, for me, maybe in my domain, I could say that these are two equal statements. But for someone who is working in fintech or something, it might not be closed. So there is also a domain kind of bias when it comes to metrics or in rejectments, because developers from the different domain expects different level of strictness and everything. So bringing out this alignment with these metrics depending upon your domain is also one extra thing that we are trying to tackle here, which is called, basically we call it as a metric alignment that is trying to align larger language model judges to your specific judgments, using the feedbacks that you can give to elements.
26:25It is also an upcoming feature regardless. I want to just dig into a couple of those things. So one is like you talk about alignment, but also there's like the idea of, it's very important in this case for me to be thinking about data and examples from my domain in terms of how I'm evaluating these. So I'm just looking through, you know, just to make things for concrete for people, looking at your example around answer relevance. and there's data samples in there that have a question, an answer, and then a set of contexts, and then you can evaluate with Ragas using the metric answer relevancy.
27:04So obviously there's a data component there. And in that component, there's answers that are there and contexts that are there. So one question would be like, how much and what type of data will people need to configure such that they can appropriately evaluate their LLM applications? And are there metrics that require sort of data upfront or metrics that maybe are reference-free that wouldn't require data upfront? What's the perspective there in terms of kind of, I guess, the cold start? Like if people are starting an LLM application from scratch, they haven't run it in production, what's the data burden and the path towards both getting this data in place and getting the alignment in place?
27:50So regarding the data itself, the number of samples people generally used to evaluate is around 100 to 500 when it comes to off-lay evaluation. So it really depends on how you formulated your test data. So for example, you could have the same results or same kind of dragging when you have 100 samples data point, or in many use cases, even 100 may not be enough to include all the kind of use cases or all the kind of distributions that you see in production. So it really depends on the variety of items that you see in production, that you expect in production. Then let's say you have a very niche use case and variety is very skewed, your test data can be very small, yet still serve the purpose.
28:38But if your use case is very broad and you see a wide variety of users coming into your application or a wide variety of records, your test data has to be broad enough to include all these different distributions in the data sets itself. So that's about the number. It's not really about the concrete number, but again, basically making sure that you understand what are the different distributions of queries coming into your system, and also making sure these distributional queries are also represented in the test data set so that you know when you change a probe or when you change a probe, which type of queries are being affected.
29:17It could be that when you change a particular retriever or a particular tool or something, when you start out a particular tool or something, it could be that a specific set of queries are being affected. The overall pipeline almost never gets affected if you want to do a huge change. And for the small changes, it's mostly a subset of queries, a subset of distribution that gets affected. And it is very important that you identify these items. So that's about how to formulate a data set itself for evaluating any applications. And regarding the second question of metrics and reference-free and the reference metrics, we usually provide both reference-free and reference-with reference metrics.
Read the full transcript
30:00But there is a big shortcoming to reference-free metrics because reference-free metrics are basically, you know, we could estimate things there, but there is an error to estimation with reference-free metrics. For example, answer 11c, you could estimate with some way that if the given answer was correct or something, but if you don't really have the exact thing, somewhere in order to know it's really possible to say that, okay, the NLM application arrived at the right thing at the end of the day. So there what we provide, even with production data, it's very hard to occur in a test data set because, again, with production data, data is very, very messy.
30:39It's a common thing that people, if you ask any researcher or a guy who has put MLs into production, MLs into production, and the method of occurring test data set, But obviously they are going to look at production data. But the way in which test data are being accurate is there is a long way to go from production data to formulating it as a test data. Because production data is incredibly messy because it's an uncontrolled environment. It's not a human-entated controlled environment. People can come and say anything there. And they have like zero consequences. So the thing is that production data becomes incredibly messy depending upon the application.
31:19If you have, let's say, if your application is serving an internal set of users, like an internal company, employees or something, still, you have an amount of control. But if your application is strictly B2C and you are opening up the whole application to, you know, anyone using your application, there can be users who are basically trolls who just like to use your system and, you know, leave out, you know, quotes, feedback, and so forth. So production data being incredibly messy, there is a long way from going from production data to a good test data set. So there again, what we are trying to do is providing a way to synthetically create these test data sets.
31:57Now, these synthetic creation of test data sets will be grounded on things like production data. Then there are your internal documents that should be taken into account when creating a synthetic test data set. Because again, you are trying to create a test data that's really tailored to use case and not a generic one. So there are two points where we ground it. Basically, the set of internal documents or whatever you have that you ground your application and also the production data where the user engages with your application. And this is again one of the upcoming features. The test data generation is already there.
32:33but we are also trying to extend it to the way that we could ground these into the from production data. Basically, we call it seeding from production data. Basically, if a user has been already using Gagas tested generation, but he wants to take motivations from production to imitate more behavior in the test data set, that is also happening in production. Now, there is a lot of things that has to be done to first understand what is happening in production. There can be different distributions, again, as I said, different set of users. We have to understand different set of queries that's coming in, different set of interactions that's happening, and everything.
33:14And once then we understand that, we could also have a way to synthesize these type of data points using LL. Now, the developers will not be to annotate these data points, but to verify these data points. Once this data or test data is synthesized by Ravas, you could export it to a simple UI tool, and then you could simply, or even Excel sheet, most of the users could export this data to an Excel sheet and basically they go through it. And then once they go through it, they can easily cut out the bad data points that they think is LL messed up or something. Because again, we can't really guarantee 100 % efficiency while synthesizing these data points.
33:57So what we are trying to do is improve the efficiency. Let's say if you generate 100 data points, our goal is to make sure that all the 100 data points are equally valid and good. But it might not happen like that in every use case. So the developer can take 10 minutes of his or her time and go through this data set manually. It is a fairly quicker process than annotating these, annotating or creating these datasets, because creating these datasets would easily take a day or two or even a lot of money when you give it to other human annotators. So developer basically can go through this data points that's been synthesized and then cut out the points that he thinks that is not valuable.
34:36So again, that is again one, these are actually the real ways in which we save a lot of developer time in evaluation, because again, and formal data set is very, very, it's a very, you know, cucumber stone, it's a very time consuming, boring process that nobody wants to do. And it's mostly falls upon one developer and the team to do it and it's a messy process. So these are the type of innovations what we are trying to bring in the evaluation space as well to save a lot of time of the developer. And this ideology or this philosophy is why we are being getting this organic growth. Because it is really kind of hard to come up with these kinds of solutions.
35:13But when we come up with these kind of solutions, there are, you know, a lot of developers who are, you know, who wants, who badly wants this. And they are, you know, they are the people who motivate us to continually bring these kind of innovations to the evaluation space. Yeah, that's great. I love the innovations around synthetic data use in evaluation and utilizing kind of the LLMs and these models to help in the evaluation process, but in a way that's validated and still fitting with improvement over time. So yeah, as we kind of close out here, this has been a fascinating discussion, but maybe just to close out here, what are some of the things that you're excited about moving into the next six months or so to either explore or maybe it's things that you see happening in the AI ecosystem more broadly?
36:10What really excites you about the direction that people are going with their LLM applications? So with LLM applications, with early 2023 or mid-2023, RAG was a big thing. And if you remember the time when RAC became popular, there was, again, a lot of limitations. There were only 4K contacts or, you know, LLM was kind of saying anything and everything. I think with HND Code tool use cases, we are at the same level as of now. You know, we have been trying to bring more and more, you know, tool use cases. Tool use cases are actually incredibly useful when it comes to building a whole LLM application experience because combined with internal knowledge, internal knowledge is what you can infuse with RAC, and then taking actions is something that infuse with tool use cases, tool bindings.
37:03So if in a tool binding, I'm very excited to see the next class of models performing better on tool use cases. Again, now even these recent models have been bringing abstractions or using tool binding to facilitate this, but still it's a little bit shaky as of now. But the next class of models, I didn't expect it to be better. And then the whole thing, the drag plus tool use case, we don't see no more enterprises adopting and using another applications at different capacities and it's saving a lot of time and resources for both people who are at the back of these applications and also people who are interacting with these applications.
37:50So that's something I'm really excited when it comes to the very next six months of scheduling applications itself. And also, I think when it comes to the frameworks or libraries that are being built around applications, I'm seeing more and more better structures these days when it comes to these frameworks and libraries because people now have almost one and a half year of building with another application experience. Now that experience is yielding more and more understanding of what kind of, what is the best abstraction to be used to build these applications and everything. And then there's an overall agreement that is coming up with what is, how to format outputs, how to build these combo systems itself.
38:33Because earlier, the first year, early 2023 and everything, people were really, really confused on how to build these applications, how to use anything and everything. Now that clarity, I can see that more clarity is happening at that end too. So at the model building stage, again, the model building spectrum, we are again having more clarity on things like the data of the models, the type of data needed, more and more papers and more and more research happening at the data processing and pre-processing stage, what is the type of data needed to train the higher quality and best quality and all of itself because we see we think that uh we have mostly settled on the architecture itself from that point of element point of view now the main thing people are you know working on this data again when it comes to data people are again now uh mostly we will finish up the free data that's available on the internet very quickly and now it's again synthetic data that has a really good chance of improving these models itself so the idea of models output being used to feed models and improving the models itself on different use cases is itself a fascinating thing regarding the model building stage.
39:42When it comes to evaluation, there needs to be more and more innovations happening in that evaluation space itself, whether as an open research or as an open approach to bringing down the time involved in testing and assessing these applications. Because that's one big pain point on why the enterprises or, you know, big companies could, a big barrier for these big companies to adopt LLM applications, LLMs or AAs in their system because these people, they have high responsibilities. So when, with high responsibility, you have to do these testing and evaluations. And there is a very good way of doing it and a great upon way of doing it.
40:27It becomes a pain point. And that's something we are also trying to bring in. You know, we are trying to, you know, formulate all that research, all the innovations that are happening in the evaluation space to come together and build an open source standard for evaluating LLM applications itself. So there is an agreement between everyone on how to evaluate LLM applications. And that's the long-term vision of the company itself. Awesome. Yeah, well, thank you, Shahul, for taking time to join and again joining at a late hour where you're at. I'm really excited about what you're doing with Ragas.
40:59And this is a really interesting space that we'll be interested to follow. So thanks for taking time and hope you can have a good rest of your week. Sure. Thanks, Daniel. This was fun chatting with you and I hope your users learned something from this conversation. Yeah. Thanks. Bye-bye.
41:24All right. That is Practical AI for this week. subscribe now if you haven't already head to practicalai.fm for all the ways and join our free slack team where you can hang out with daniel chris and the entire changelog community sign up today at practicalai.fm slash community thanks again to our partners at fly.io to our beat freaking residents breakmaster cylinder and to you for listening we appreciate you spending time with us. That's all for now. We'll talk to you again next time.
From the publisher
How do you systematically measure, optimize, and improve the performance of LLM applications (like those powered by RAG or tool use)? Ragas is an open source effort that has been trying to answer this question comprehensively, and they are promoting a “Metrics Driven Development” approach. Shahul from Ragas joins us to discuss Ragas in this episode, and we dig into specific metrics, the difference between benchmarking models and evaluating LLM apps, generating synthetic test data and more.
Changelog++ members save 5 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
- Assembly AI – Turn voice data into summaries with AssemblyAI’s leading Speech AI models. Built by AI experts, their Speech AI models include accurate speech-to-text for voice data (such as calls, virtual meetings, and podcasts), speaker detection, sentiment analysis, chapter detection, PII redaction, and more.
Featuring:
Show Notes:
Something missing or broken? PRs welcome!




