Generative Benchmarking with Kelly Hong - #728

23 Apr 2025 · 54 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

```markdown

Episode Notes

Generative Benchmarking with Kelly Hong - TWIML AI Podcast #728

Podcast Overview Title: The TWIML AI Podcast Description: The TWIML AI Podcast focuses on the transformative impact of machine learning (ML) and artificial intelligence (AI) on business operations and daily life. Hosted by Sam Charrington, the podcast features insights from industry experts, covering technologies in ML, AI, deep learning, natural language processing, and data science.

Episode Summary In this episode, Kelly Hong, a researcher at Chroma, discusses "Generative Benchmarking," a new method for evaluating retrieval systems like Retrieval-Augmented Generation (RAG) applications using synthetic data. Kelly critiques traditional benchmarks such as MTEB for failing to represent real-world query patterns and outlines the limitations of embedding models that perform well on public benchmarks but poorly in production. The conversation covers:

  • Generative Benchmarking Process: A two-step method that includes document filtering to focus on relevant content and query generation that mimics actual user behavior.
  • Application Insights: Kelly shares insights from applying generative benchmarking to Weights & Biases' technical support bot, highlighting the value of domain-specific evaluation.
  • Alignment Challenges: The need for aligning large language model (LLM) judges with human preferences and the effects of chunking strategies on retrieval effectiveness.
  • Real-World Query Differences: The distinction between production queries and benchmark queries in terms of ambiguity and style.
  • Evaluation Approaches: A call for systematic evaluation methods beyond "vibe checks" to enhance the effectiveness of RAG applications.

Key Concepts and Discussions

Limitations of Traditional Benchmarks

  • Public Benchmarks (e.g., MTEB):
  • Often generic and do not reflect real-world use cases.
  • Datasets tend to be clean, while real-world data is typically messy.
  • Models can perform well on benchmarks due to memorization rather than genuine retrieval ability.

Generative Benchmarking Process

  • Two-Step Process:
  • Document Filtering:
  • Focuses on selecting only relevant documents for the evaluation.
  • Filters out irrelevant content that does not align with user queries.
  • Query Generation:
  • Creates test queries based on the filtered documents.
  • Ensures queries mimic realistic user behaviors and styles.

Importance of Domain-Specific Evaluation

  • Insights from the Weights & Biases technical support bot showcase how a tailored evaluation can lead to more accurate assessments of embedding model performance.
  • Highlighted the discrepancies between model performances on public benchmarks versus real-world data.

Alignment of LLM Judges

  • Human Preference Alignment:
  • Emphasizes the necessity of aligning LLM judges with human assessments to ensure meaningful evaluations.
  • Discussed how prompting strategies directly impact the effectiveness of LLM judges.

Future Directions and Applications

  • Iterative Improvement:
  • Call for continuous improvement of benchmarks as more queries and usage data are collected.
  • Encourage developers to adapt their evaluation sets based on emerging user queries and information gaps.

Key Takeaways

  • Systematic Evaluation is Vital: Developers should focus on systematic evaluation methods rather than relying on informal checks.
  • Generative Benchmarking Offers a Path Forward: The approach is designed to aid developers in creating realistic evaluations without needing extensive prior data.
  • Adaptability is Key: Engineers should prepare to iterate on their evaluation datasets to remain aligned with the evolving nature of user interactions.

Final Thoughts The episode wraps up with Kelly emphasizing the significance of generative benchmarking as a starting point for developers looking to evaluate their retrieval systems effectively. The discussion underscores the importance of moving beyond surface-level evaluations and adopting more rigorous, systematic methods for assessing AI performance.

Full transcript and show notes available at: [TWIML AI Podcast](https://twimlai.com/go/728). ```

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If you're familiar with the retrieval space, MTAP is a pretty popular benchmark for embedding models. If you see the release of any new ing button model, you'll probably see something like, oh, our model achieved like this score on MTAP. We outperformed this other model on MTAP. And people would just kind of take that score and assume, okay, this model is better just because it did better on this benchmark. But there are so many problems with these public benchmarks.

0:35All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Kelly Hong. Kelly is a researcher at Chroma. Kelly, welcome to the podcast. Thank you so much for having me. I'm really looking forward to digging into our conversation. We're going to be talking about a project you recently published called Generative Benchmarking. That one really resonated with me for a few reasons. It touched on themes like synthetic generation data for evals, LLM as a judge, data leakage and overfitting on public benchmarks. So a lot of interesting themes touched on there.

1:14Before we dive in, I'd love to have you share a little bit about your background. Yeah, of course. So yeah, as you mentioned, I'm currently doing research at Chroma. We're a company working on vector databases, but we also work on research around retrieval as well. just to give a little bit more background about myself I started my undergrad at UC Berkeley around three years ago studying CS about two years and I got pretty interested in this space and like search and retrieval so then I spent some time like building projects around it going deeper into the research and I would come across Kermal a lot because like I would use it in my projects but I also saw that they did a lot of good research around like chunking strategies like embedding adapters, just ways to make your retrieval system better.

1:59So I got pretty interested in Chroma specifically, and then I started talking to the team, got this opportunity, and now I'm here. So I'm on a gap year right now from school and working here basically full-time. And Chroma does come up, particularly in the context of vector databases and RAG. And RAG continues to be a topic of interest for folks listening to the podcast and the folks that I talked to in enterprises. So I think this will be an interesting conversation in that light as well. Talk a little bit about the origin of the project. Was it already framed up when you got there? Or did you take some time to kind of scope it out?

2:42Yeah, yeah. Happy to talk about it. It's something that the team has been thinking for a while. And I've personally been thinking about like on my own as well. It just so happened that like we were pretty aligned on the same problems. But when I was like building out my own projects, like one thing that would constantly come up is how difficult it is to like systematically evaluate AI systems. Like if you think about this analogy with like software systems, it's very easy to evaluate them. You have these like very simple code tests, you like run them and then you get like a deterministic output.

3:12But if you have like AI systems, all of your outputs are like probabilistic. You don't really know what it's going to output. So then like how do you exactly like evaluate them? I think initially when I started building out in the space, I would just evaluate based off of Vibes, which is really bad. You know, like an output would either feel good or it didn't look good. So then I might change the prompt a little bit, maybe change out my embedding model. But there was really good, like no good way to systematically evaluate it. So I kind of felt the personal frustration there. And that's something that Chroma aimed to target for a while as well.

3:44Like we wanted people to be able to systematically debug their AI systems. So that's kind of where generative benchmarking came about. One of the main motivations for that was to address this problem of people not being able to evaluate their AI systems in an easy way. If you think about evals as a whole, it's not a very friendly topic to a developer who's just getting started with RAG. They have to go pretty deep into learning about how evals work, what a benchmark is. And it's a lot of work that people don't really spend a lot of time on. but it's very important if you actually want to improve the performance of your AI system.

4:20So yeah, one of the core motivations was to just create a really easy way for people to evaluate their AI systems and understand the importance of it. And I think another motivation for that was to address the problems of a lot of current benchmarking methods. If you're familiar with the retrieval space, MTub is a pretty popular benchmark for embedding models. If you see the release of any new in button model, you'll probably see something like, oh, our model achieved like this score on MTAP. We outperformed this other model on MTAP. And people would just kind of take that score and assume, okay, this model is better just because it did better on this benchmark.

4:59But there are so many problems with these public benchmarks. I can go into these in like a bit more detail as well, but just to give a high level overview, they're very generic. So they don't reflect like the domain-specific use cases of what people see in production. The data sets are also very clean and polished, whereas in the real world, we have very messy data. And a lot of this data has probably been seen by embedding models before. So you don't really know if you're testing for like true retrieval or if they're just, you know, doing this based off of memorization. So just kind of like going off of, you know, all these problems of people not really knowing how to evaluate their systems.

5:38And also like the problems with like a lot of like public benchmarks kind of inspired this idea of generative benchmarking. And you mentioned that vibe checking is the predominant way that developers are kind of assessing the performance of their systems. There is a lot of, I think, kind of a broad call for folks to take a more rigorous approach to evaluation in the space. How do you see folks evolving from Vive coding to a more rigorous or scientific process to evaluating the performance of their AI apps or RAG apps? What are the steps that folks typically take? The first thing that people do is they probably look into, you know, benchmarks like MTEP and they kind of see scores of that.

6:28And then they make the decision on like which embedding model to use. but I think like with these methods of like generative benchmarking what we'll hopefully start to see is like people kind of like taking advantage of these tools that make evaluation like really easy like you don't really need to know like the details of like the entire evaluation pipeline like the generative benchmarking like notebook that we set out is just a very simple notebook you just upload your data you fill in like the context and then you run a few cells in a notebook it makes the process very very easy so I think like hopefully the trend we'll start to see as people start using these kinds of tools that make evals very approachable.

7:05And I hope that from that, people get more familiar with the idea of doing evals and also understand the importance of it as well. And I think from that, people can get more into the details of maybe aligning their LLM judge or adding more human-in-the-loop processes. But yeah, I think initially, we'll probably see people start to favor more of these approachable eval solutions. Yeah, yeah. So you don't see generative benchmarking as kind of an advanced stage that you evolve to. It's meant to be kind of an easy thing that you can adopt pretty early on. Yeah, for sure. It's kind of like helping people with that first step if they don't have a golden data set to start with.

7:49Because for a lot of retrieval systems, people probably don't like log their queries when they're getting started. They just have their like set of documents, right? For the retrieval system. so they don't really have like a set of like query document pairs to test their system on so we kind of help with like uh starting the process of like generating um that golden data set so people can start to evaluate but the eventual goal there is for people to like continue iterating on that data set you know for example if they have like incoming user queries hopefully they use that to align their golden data set even further maybe like improve their reachable system like overall based on that um but yeah i think this sort of more serves as like like a starting step rather than like the end state of evaluation.

8:32Yeah, what I've typically seen in my own projects and in speaking with folks is like you mentioned, you start with kind of vibe checks, you put some prompts in, see what comes out. And then after you've done that for a while, like you collect those and now you have a short list of, you know, evals that you can kind of repeat if you want to tweak it, change a different model or some parameter. And it sounds like generative benchmarking might be the next step. Or historically, I think of the next step is like doing synthetic generation. Synthetic generation is a part of the overall generative benchmarking process.

9:14But the way you incorporate synthetic data generation is specifically tailored as part of this process. So this might be a good time for you to kind of talk through the overall generative benchmarking process. At the high level, generative benchmarking is essentially generating a custom eval set based on your own data. And we do this for Retrieval specifically. So you can kind of think of it as like, okay, user inputs, a set of documents gives us some context on what their application is. And then from that, we generate queries to test your Retrieval system on. And this is broken down into two main steps.

9:51we first do document filtering, and then we do the actual query generation. So I can kind of go through these two steps in a bit more detail. Before you do that, why that two-step process? As opposed to, I've got a bunch of documents, give them to an LLM and say, hey, generate a bunch of queries, and now this becomes my generated test set because I've got relationships between those documents and the queries. Yeah, so we have this two-step process of generative benchmarking. And we do this because we realize the document filtering step, which is the first step in our process, is actually very important because usually a lot of the time there are cases where a few documents aren't very relevant to your use case.

10:33For example, we were working with data from Waste and Biases, their technical support bot. So it's basically like this chatbot where you can talk to Waste and Biases documentation. We noticed there were a few documents in there that were talking about like startup funding news or just content that wasn't relevant to any kind of developer using this technical chatbot. So we first want to filter out those documents because they're not reflective of what users will actually ask about. and kind of the main goal with generative benchmarking is to create a benchmark that's representative of your true use case.

11:05So that means, you know, we want to filter out content that users wouldn't ask about in the first place. So we find that document filtering is a pretty like critical step in this. Like we kind of compared this to like a naive query generation approach where we did no document quality filtering, no like query string anything and then compared it with that, which I can talk about it like in a bit more detail later on. But yeah, that's kind of like the core motivation for breaking this down into two main steps. Like we want to make sure that we're selecting for content that users will realistically ask about and also generate queries in a way that's realistic and reflects like queries that you'll see in production.

11:42And so is that filtering, do you see the filtering step as in and of itself kind of a core differentiation for this approach? Or is it, you know, one of several? I think that's definitely one of the core differentiators. I say that and the way that we are generating queries. Because I think like these two parts combined kind of like make up what a realistic user query is. Like first you have like the content, like you want to be asking about content that would actually be relevant for a user that's using your application. And also the style that you ask the query in. Like for example, a lot of users using a technical chatbot, like the one that Waste and Biasis has, they probably wouldn't be asking like queries in complete like sentences.

12:25is. It's more of like statements, which you typically see in a lot of like chatbot applications. Like people are very vague with their queries. It's not like super comprehensive. So we want to reflect that as well. So I think these two steps like go hand in hand. A big part of the two steps is injecting the context into both the filtering and the query generation. What is that context? Yeah, the context is basically like what your ROG application is for. So for example, in the case of Waste and Biasis, like the context would be something like, this is a technical support bot for Waste and Biasis.

13:00A lot of the users would be people who are, you know, maybe training models or like using these kinds of like platforms. So just giving some context to kind of telemodel like what to focus on. It's kind of like a system prompt. Yeah, yeah, kind of like that. I was thinking about this process and its applicability to my favorite use case, which is I have a bunch of podcast transcripts and I want to create a rag chatbot for those. And I'm evaluating this as like, oh, how would this help me? And one question that I had was, do you think it's applicable to that use case? Because I'm not sure that filtering is, filtering could be interesting at the chunk level.

13:50It's not necessarily interesting at the document level. There's some chunks, probably a lot of chunks that don't add a lot of information and aren't reflective of user queries. But the document corpus as a whole, those are all relevant documents. I guess just to clarify, like we're both doing like chunk filtering and document filtering. So we would filter for chunks as well, because I think in a lot of cases, like even if a document is relevant, like, for example, this case of like podcast transcripts, all documents are relevant. but they're probably chunks that wouldn't be good to generate a query from.

14:23Like, for example, say you have like the end of a podcast where you're saying like goodbye or something. Goodbye, goodbye. Yeah, you probably don't want to retrieve that. And so we would filter out chunks like that because like what kind of question would you ask for a chunk that just says like goodbye, goodbye? So, yes, I think like in this case, it would be more focused on like filtering at the chunk level. You mentioned filtering and that second step is query generation. What do you do different with the query generation part? Yeah, for query generation, it's pretty simple. We first give it like the same context that you would give to like document filtering.

15:01So this would just be like, okay, this is a technical support bot and then et cetera. And we also give it some example queries as well. And that's mainly for the purposes of kind of steering the LLM to generate like queries in a realistic style. So like I mentioned before, a lot of people who use these chop out applications, the queries are very big. They're not like formed as like a complete question with like a question mark and everything. So we want to be able to reflect that as well. So, yeah, we incorporated both like the context and some example queries. Is the key distinction around kind of the sample queries relative to what you see users doing when they're using LLMs to generate test queries?

15:45Or is it more distinct from what you see in kind of the big public benchmarks? Are you asking about like how the example queries differ from like public benchmarks? One of the points you referenced earlier is how with the public benchmarks, the queries aren't representative of actual user queries. Users are, you know, they take shortcuts. They're not always grammatically correct. The public, you know, queries are more structured in some ways. and I'm asking if that is the key thing that you're trying to, if kind of the real user nature of the query is the key thing that you're trying to model or are you trying in this step to correct for the way you see developers or engineers like manually creating sample queries to give to an LLM to generate more queries?

16:50We're, yeah, we're trying to model like what you would actually see in production, which is like pretty different from like the public data sets that I spoke about earlier. Because I think, yeah, like in these public data sets, like a lot of questions are very like well formed and polished. So they would contain a lot of like complete questions and they'd be highly relevant to the documents as well. but in like real production data which we got from like Waste and Bias is like these weren't queries that were you know like written by the Waste and Bias team there were queries that they actually took from like users who used their technical support bot so we know that it's reflective of what they will actually see in production so yeah we wanted to make the distinction between you know like what you actually see in production versus like what I don't know you was just seeing like a typical benchmark data set yeah and I get the distinction between the benchmark data set.

17:45I'm wondering if you mentioned that you pulled this information from the weights and biases, like a query log. And I'm wondering if you have a sense for if someone at weights and biases was like building out these queries manually, would they be different from what they would see in the test log? And the style of queries, I imagine it would be very similar. I think it's pretty easy once you see like a list of like actual queries from like users, you can like replicate them. But I think maybe one distinction there would be like the content that the queries focus on. Because when we're looking at the data, one interesting thing we noticed was that there were some queries that just cannot be answered by any documents in the corpus.

18:30And if you were just like a waste and biases, like engineer, like writing queries, then you wouldn't focus on queries that are not answerable by your document corpus. So I think that was like another interesting point that we saw. Like if you are able to work with like real production data, then you can also kind of use that to, you know, evaluate how well your document corpus reflects what your users are asking about. So I think like the style of queries will probably be very similar, but maybe the content, the content of queries that users ask about will be pretty different. Again, you filter your chunks, then you generate your queries.

19:05What's next? Yeah, after you generate your queries, you have your query document pairs. So this is how evaluation typically works for evaluating an embedding model. You have a query document pair, you embed your document corpus, you embed your queries, and then you see whether you retrieve that document in the top K results. That's how you get recall a K. So then we basically run that evaluation for our generated dataset. And then we compare it against the ground truth queries as well. that kind of serves us like our ground truth for kind of testing whether our generated queries actually reflect like what you would see in production.

19:40And those ground truth queries are the examples that were given or where did those come from? Oh yeah, so these were the actual queries that were logged from like their technical chatbot. And the waste and biases case, got it. Yes, for waste and biases. So it would be like the log queries and then we like manually labeled all the data and then use that to test for representativeness. But this was mainly for like our research report. But if you are maybe like a developer using this tool, you would just end up with that set of like query document pairs and then use that to evaluate your retrieval system.

20:12Because kind of like from our end, we kind of did the work of demonstrating that these are representative of what you'll see in production. So then users can just take that generated data set and use that for evals. And you mentioned that one of the things that having an eval system allows you to do is to swap out different embedding models, to play with chunking strategies, that kind of thing. A question that comes up frequently is how much does that stuff really matter? What's your take on that based on what you've seen? I think it matters a lot. I think we've seen some pretty distinct differences between embedding models.

20:53And I think one of the key things we saw when working with like production data was that, like we saw that Gina AI's model like performed slightly worse than like OpenAI's text embedding large model, which actually conflicted with like the MTEP scores that they released. So in spite of performing better on the public benchmarks, it performed worse on real world data. Yeah, so that kind of goes to show like, oh, like these performance on like public benchmarks don't necessarily translate to your specific use case. And yeah, we saw some pretty like clear differences between like the rankings of embedding models.

21:28We saw that Voyage's three large model performed best. So we'd recommend them to use something like that. And we didn't extensively test out chunking strategies or we have another report on that. And I think there are some other things you can do as well, like contextual rewriting. We noticed that a lot of chunks were missing context, which is a pretty common problem in a lot of these ROG applications. You can just chunk up documents, but some chunks might just not be very valuable if they don't have any context on where they came from. And trying to fix problems like those actually boosts retrieval performance a lot.

22:03Anthropic did a research report on that as well. So yeah, these strategies make a pretty big difference in retrieval performance. Is that typically done in your chunker, just asking an LLM to capture the context and prepend it to whatever the chunk is? I wouldn't say it's super common because it is pretty expensive to do. I mean, you're like passing it into an LLM every time. So it really depends on your use case. If you have like tens of thousands of chunks, probably doesn't make sense. You probably want like a more effective approach to maybe like selectively choose chunks to like rewrite. But yeah, I guess it really depends on the use case.

22:40Do you focus on a particular set of metrics in talking about this work? We mostly looked at recall at K just because there are other scores that are used in a lot of these retrieval research papers. People use NDCG, but one thing that has been brought up is that the ranking of chunks in the context of an LLM doesn't really matter. It doesn't affect the performance of an LLM if you have the relevant chunk in the beginning versus the end. So we don't look at NDCG. We just look at recall, which is like, was the relevant document retrieved or not? Presumably, if it's in there somewhere, the LLM will find the information it needs and create the right response.

23:25Yeah, yeah. So, yeah, I think like rather than looking at the rankings, it's more important to just see like whether it was retrieved or not. Are other metrics like groundedness and things like that, are those out of scope for this work? Yeah, we didn't look at groundedness. We purely just looked at like numerical metrics. I think like for groundedness, a lot of people use like LLM judges for that. But I think for that, you do need a good level of alignment, which I think could be interesting for maybe some future extensions of this. But yeah, we only looked at metrics for these. In the report, you talked about LLM as judge and alignedness.

24:08Can you elaborate on where that comes into play? Of course. So alignment is a very important part of this work. we specifically use this framework called Evalgen, which is basically like this, I guess, like framework for aligning an LLM judge to human preferences. Because if we just blindly use an LLM judge, it's not guaranteed to judge in the way that humans do. For example, when we were first using an LLM judge to filter documents, we compared that to like ground truth human labels. And I think it only got around like 46 % alignment, which is pretty bad. So we basically use this framework or like EvalGen to like iterate on our LLM criteria and get it up to like over 70%.

24:47So I think a mistake that a lot of people make is they just blindly use an LLM judge without any alignment, just assuming that it is aligned. But I think we're too far from that right now. And we do need some level of like human alignment to actually like confidently put the LLM judge to use. And to dig in on that, what you're doing there is you're asking the LLM to determine whether a chunk or document is relevant. And that's the filtering process. So basically, is it good to generate a query from or not based on this use case? Before we dig in on EvalGen, I'm curious how sensitive your results were to the prompting that you use for LLM as judge?

25:31I guess we noticed the prompting strategies through using EvalGen. I guess they both kind of go hand in hand because EvalGen is essentially evaluating how well does the prompting strategy allow your LLM judge to be aligned with your preferences. So yeah, we noticed it was pretty sensitive because we were just iterating on three criteria. And based on how we were wording those criteria or how we were maybe taking out one and adding in another, it would range quite a bit. So we sat on a report, it improved from 46 % to I think like somewhere like in the 70s. So that's a pretty big difference just based on like how you're prompting the LLM judge.

26:11So yes, we did notice that LLMs are very, very sensitive to prompting. Can you talk about EvalGen in a bit more detail, how that process works? Yeah, yeah, of course. So for EvalGen, you basically need a set of criteria. So in our example, we just use criteria like relevance. So like, is this relevant to our context? We use completeness, which is just how is the overall quality of this LLM judge. I think one other criteria as well. But yeah, basically, you would write out these criteria. And then you would also have ground truth labels for your documents. So we had 200 documents. And then we manually labeled them as either good or bad.

Read the full transcript

26:51So it's very simple. And then for the 200 documents, we would pass them into the LLM with their criteria and just ask it to evaluate based on that criteria. like does it meet this criteria or not and then we would have a threshold as well like if it passed like two out of the three criteria then we would judge that as um being like a good quality document so that's how we got the um lm judge labels and once we had those labels we would compare them with our ground truth and then that's how we get the alignment score so based on the alignment score we we would iterate on our criteria we would also have like criteria specific scores as well So that kind of helped us to determine like if one criteria's alignment score was very low, we would maybe like rewrite the criteria or just like take it out like all together.

27:36And we would kind of like use that iterative process to, you know, modify the prompt essentially and increase our score over time. How sensitive was all of this to the document size? Or you mentioned earlier that you're, at least for filtering, you're doing it at the document level and the chunk level. Were you also judging relevance? That's also, well, that's part of the filtering. So that's also a document and chunk level? Oh, we were doing the filtering like all on the chunk level. So - All on the chunk level. Got it. And so the labeling of relevance, that was all done at the chunk level as well.

28:16Yes, yes. So you would - Yeah, you would pass the chunk into the LLM. Yeah, I think maybe I should clarify that because I think a lot of people, they don't immediately know like what a chunk is. So I would typically just say like document. So yes, so it was just a chunk into the LLM judge. And the chunk size, I'm trying to get a sense for, you know, if someone is looking at applying this, like what the data labeling process looks like, are they applying a, you know, say you've got, you know, 500, you know, actual documents, you then chunk those, what chunk size are you looking at? Does it matter for these purposes?

29:00And how many of those chunks are they labeling in order to kind of kickstart this process? Yeah, yeah. So in our case, we had around like 13 ,000 documents. And then out of those, we would manually label 200, which is a pretty small percentage. Like it doesn't require that much human labeling. So I should clarify, 13 ,000 chunks from documents. And then we would manually label 200 chunks. So yeah, we would do that. And all the documents were pre-chunked for us. So we didn't really do any chunking on our end. But I say they were typically around maybe like 400 tokens, like somewhere around that range.

29:44But yeah, I think your chunking strategy really depends on your use case as well. Like if you have a very information-dense document, then you probably want to have smaller chunks. So you're able to like kind of like isolate like those like semantic differences. But I think for like waste and biases, like this chunking strategy made sense for them. But yeah, basically like out of like 13 ,000 chunks, we would manually label 200. And then after we aligned our LLM judge, we would extend it to the rest of our document collection. You then have your data set and then you can apply this eval kind of repeatedly against your documents.

30:25Are there other steps in the process that you need to do as an engineer trying to apply this? No, that's pretty much it. We try to make it pretty simple for people to follow. So it really just comes down to like filtering out your chunks and then generating queries from them. You mentioned in the report you referred to Ragus and another tool as kind of, you know, existing work, prior work that or existing tools and prior work that folks use. Do you see generative, this generative benchmarking? like do you see it ever being packaged as a tool or is it more a process like you offer it as a it's a library that you import in a notebook and you kind of run through it um do you think it can be more packaged in its delivery yeah yeah that's something we're definitely thinking about for the future right now it's just packaged as a notebook just to like have it along with the report but i think in the future it makes a lot of sense to have it as like a tool like along with their product as well.

31:38I think, yeah, like Roggets and the other tool is like Airbench by Gina. Yeah, they have a pretty similar like motivation as this, which was to like, kind of create like a custom eval set based on your data. But I think I wanted to like emphasize this like a bit more that our approaches are pretty different. I think for Roggets and Airbench, they focus a lot on like generating a diverse set of queries. They don't really do much. They don't put much focus on like how representative it is of like your actual like production data. So I think that was like the main like differentiator that we had.

32:15Like we know there were a lot of like previous methods on like synthetic data set generation, but I think our work is like pretty unique in the fact that it like focused a lot on like how representative the like generated queries are. And again, And kind of going back to this scenario of an engineer that's starting to work on a system like this or, you know, a rag based system or retrieval oriented AI system, but they don't have historical data to use. Is generative benchmarking still useful to them? Yeah, yeah. I think that's definitely like one of the key use cases of this. Like if you don't have like a set of like queries that you can test your retrieval system on, like we would generate that for you.

32:59Like all you need is just like a set of like documents that you're using for your retrieval system, which people already have. So like you have all the tools. So you just need to like run this and then we'll generate the eval for you. So I think, you know, like we were kind of hoping with this that it would be an easy way for people to get started with evals. just start to like get more familiar with it and like understand like why it's important to use. You mentioned that part of the inspiration for this was things that you ran into in your own personal projects. Have you tried using it with projects beyond the weights and biases dataset?

33:37We've tried it like on our own like Chroma documentation as well. Like one thing we're just like, you know, testing out for fun was like, how well would this perform on like Chroma's, like our like chroma documentation. So yeah, we ran like evals on that for like a few different models. And yeah, like we got some like results for that just to like kind of indicate like which embedding model is like best to use. I haven't tested it out like extensively on like other projects, but yeah, I think - Did it change anything about the way you used, the way you built a system for your own docs? Yeah, yeah.

34:13Yeah, for us, like previously, we just used like OpenAI's, I think, like small model. But after looking at this, it actually like performed like, it was like one of the worst performing models out of the ones that we tested. So yeah, I think like from now on, we'll probably go with something like Voyage's model, which works pretty well on like technical documentation like this. So yeah, that was like pretty interesting to see. I think like once you actually like test on your data, you actually have like more confidence in like determining like which embedding model is best to use for your use case.

34:43Any other thoughts on what's next for the project? Yeah, yeah. I think in terms of next steps, one thing that would always come up is how you would iterate on this generative benchmark once you have your initial set. Because I think, like I mentioned before, it's not necessarily the end goal. We want people to iterate on this further to get it better aligned with what they see in production. So I think if you are logging queries, one interesting thing you could do is query alignment. So maybe you could kind of use that to determine what kind of topics users mostly ask about, align your eval set closer to that.

35:19Also, it could help in determining if you have any information gaps in your document corpus. If users are asking about this one topic, but your document corpus isn't able to answer that, then people can maybe proactively act upon that and fix it before more users run into that problem. I think another thing that came up when we were working with production data was we noticed that retrieval performance drops a lot when you're working with these very domain-specific data sets. Because we've spoken to a few other people who are working on RAC systems as well, just maybe for internal company documents and things like that.

35:56And these are very, very specific compared to the public benchmarks that you see. If you just have a Wikipedia data set, it's going to focus on a variety of topics. But if you just have, I don't know, like an internal tool that has, that's just focused on like one very specific area of your company, then it's very hard to like differentiate between like different like chunks, right? So I think one area that could be interesting to explore is like, how can we like improve performance in these like very, very domain specific use cases? Because I think this isn't really revealed in a lot of like public data sets where retrieval is like a lot easier.

36:31But yeah, in these cases, like we see performance drops a lot because of this. It's hard to like disambiguate like between documents. So I think, yeah, that would be interesting to look into as well. Yeah, I was thinking a similar thought in trying to mentally apply this to the podcast transcription use case that I mentioned. And like, it might be more useful for me to change the context on a per transcript basis than to tell the LLM, you know, you're trying to judge the relevance of documents for a podcast. Yeah, yeah. I think, yeah, context definitely improves performance a lot. So I think, yeah, I would recommend like getting as specific as you can.

37:17I'm curious, maybe taking a step back from generative benchmarking, when you think about the engineer, the enthusiast who is, you know, building a system still in that, you know, vibe, vibe check regime. team, what are the things that they need to, you know, they're, they're, you know, not an information retrieval, you know, researcher or anything like that. What do they need to understand about information retrieval to build really useful systems? I think, I think one of the important things is like understanding how important retrieval is in the context of your entire like ai system i mean a lot of people just use it for like rag where you like retrieve like relevant documents and then you have an lm output a lot of people just tend to focus only on the lm output um just like is it performing like good or bad and then they might like change your prompting but maybe like the problem is just in like the retrieval itself if you can make that better maybe your lm output performance will increase a lot more too so i think um maybe like one important thing is to like look at the individual components of your like ai system.

38:35Don't just look at like the input and output and like try to like figure out your way from there. I think there is like a more like systematic approach to this. And, you know, like one of the best ways to just like start off is like looking at the outputs of like each component of your system. So if you have like RAC, look at what documents are like actually being retrieved and try to like figure out what's going on from there rather than, you know, just looking at the output and then like seeing what the wipes are. You kind of pointed to the garbage in garbage out scenario where the LLM can't do any better than what it's given.

39:05Yeah, exactly. So yeah, I think like, yeah, just focusing on like retrieval specifically could be like an important point because yeah, a lot of people just focus on like the final output, but there is a lot more that goes on. Any other thoughts along those lines? One thing that comes up a lot in like ROG systems is that if you have a lot of distractors in your LLM prompt, it leads to very degraded performance. What's an example of a distractor? If we take this example of a technical support bot, maybe you just have... Let me try to think of an example. Maybe in Chroma Z's case, you're asking, okay, how do I create a collection?

39:48And then it retrieves a bunch of chunks around creating collections, like querying a collection, like filtering a collection. It has so many specific points about collections in general. So then the LLM might get distracted and it might focus on how it might mistake maybe a function for filtering a collection when you were really asking about creating it. And we noticed this comes up pretty often. So that's why we think retrieval is pretty important because your LLM can get very distracted by which documents are retrieved. So, you know, being able to like debug that first is like a pretty important step in actually like improving the overall output.

40:27Have you come across useful ways for folks to think about the number of retrieved documents to give to their LLMs for generation? Seems related to the problem, the distraction problem you are just describing. Yes, that again depends a lot on the use case as well. I think like you want as like few distractors as possible. There's no like one number that I would recommend to use. I think like you really just have to like look at your data, test out different K values, like different like numbers of like documents to retrieve. And then, you know, see like based on that, like if you run like evals, like you can run evals with like different like K values.

41:11You can do like recall at one, recall at three, recall at 10, see how like the scores differ and maybe like go off of that because it does change a lot for people. Like, for example, when we were working with like the public benchmarks, we only just had to look at like the recall at one scores to get like a pretty good idea of how they're performing. But for like the waste and biases data, the retrieval performance like went down a lot. So we would look at like recall at K, I mean, recall at 10. So it definitely differs a lot like based on use case. So definitely like try out like different values and like see what works best.

41:40And is that because the real world queries are either less specific or I guess a point that you mentioned in the work itself is that in the benchmarks, the queries are like taken verbatim out of the benchmarks in many cases or the data sets in many cases. So it's easier to do to recall at one. Yeah, definitely. That's definitely one of the core reasons. A lot of real user queries are very ambiguous. so in um yeah and like a lot of like the polished data sets that you see like the query document pairs are highly relevant so it's very obvious that like a query matches documents whereas in the real world maybe a query is only um relevant to like the first sentence of a chunk so yeah that's why reachable is a lot harder because we have these ambiguous queries the query document pairs aren't as relevant to each other um so yeah we see that a lot um and we we tested this out with like a naive query generation method as well, where we wouldn't give any like example queries.

42:46We wouldn't give it any context. We would literally just feed in the chunk and tell the LLM to generate a query. And in those cases, oftentimes it would generate like pretty relevant queries, like more relevant queries than like real production queries are. And we noticed a performance like increased by a lot, which if you're just looking at the numbers, it looks good, but it's not really reflective of what you'll actually see in production. So yeah, like what we actually want to see is like - Meaning your retrieval performance on that data set does well, but your retrieval performance on your real world queries would not be as strong.

43:22Yes, exactly. Like you don't just want to get like high numbers, like you want something that's more realistic. Yeah, in the section of the report where you talk about the naive query generation, you talk about kind of near identical matches between the generated queries that are like rewordings of the ground truth queries. Yeah, so that was when we were working with like public datasets because we wanted to first demonstrate that our generated queries were like representative of these like widely accepted benchmarks. So yeah, one thing that we initially tried doing was just giving like the chunk to an LLM and telling it to, you know, generate a query.

44:08And when we did that, we noticed in a lot of cases, the LLM would generate like identical queries. And identical queries are like near identical queries, which basically meant that the queries were like reworded. And this basically shows that like the LLMs - Identical queries to what? Oh, identical queries to like the original public dataset. So for example, like the Wikipedia dataset, It comes with the query document pairs. So if you fed that document slash chunk into the LLM, it would generate the exact same query that would appear in the Adrenal Hugging Face dataset. The example that you give for that is from the Wikipedia dataset.

44:51And you show that the ground truth query is like where was the movie deliverance filmed? And the generated query was where was the film deliverance shot? But the target documents, like the very first thing it talks about is where the film was shot. And so it seemed like an obvious query to generate and didn't strike me as proof that there was like some kind of data leakage. Yeah, that's a good point. I think like we showed that example just to demonstrate what like a near identical query was. But we provide some like examples in our appendix where it shows like identical queries that were generated like word for word, where that isn't necessarily.

45:30Where it's less obvious from the document itself? Yes. Like the LLM could have asked about anything, but like if you see like constant patterns of like identically generated queries, then I think that's a pretty good sign that like the LLM has seen this data set before. And did you find that different LLMs had different degrees of, you know, presumably exposure to the various benchmarks? We didn't do much extensive work on that. We mainly use quad. So I can't speak much on that, but I wouldn't be surprised if it also did memorize a lot of these datasets. You have a chart in there where you're trying to demonstrate that the generated queries are representative of the ground truth.

46:25And you do that by showing that there's similar relative performance across the embedding models. Can you explain that methodology a little bit? Yeah, yeah. So that's one of the methods that we show representativeness is, does our generated evals that generate the same ranking of embedding models when we run the original task over them? So yeah, basically how this would work is we have our ground truth dataset and then we have our generated dataset. So for each of those, we would run the retrieval task and we would get scores like recall at k. And we would compare those scores. And you can see in the charts there, we kind of compare the scores for generated queries, the scores for ground truth queries, and then we do that for each embedding model.

47:10And we see that we have consistent rankings for embedding models, regardless of whether you're looking at the ground truth queries or generated queries, because ultimately this is what matters to a developer building a RAC application. It's just like, which embedding model should I use? So if there is no like distinction between... If they started switching positions and something performed really well on the generated set, but not on the real world set or vice versa, or ground truth versus generated or vice versa, then that would not be good for choosing an embedding model. Yeah, like we want to see similar performance on the generated and ground truth data set.

47:47So because like we saw like similar rankings, we saw like similar scores. Yeah, that kind of like supported our like argument that our queries were representative. Yeah, one of the things that I saw pop up a little bit in the commentary around this work was the suggestion that it was like auto evals, that you would hit the easy button and this generative benchmarking process would just generate hands-off, lights-off kind of benchmarking data for you from your data set. Is that how you think about it? No, and I think that's a very common misconception that people have. Like if you just hear a general benchmarking, you might just think that, okay, I just need to like press a button and then it generates like an eval for me.

48:31But that's definitely not the case. We do need some like human in the loop to actually make this process like more reliable in terms of like how aligned it is to like human judgment. And this kind of takes form in like a few ways. Like one is like the user provided like context and example queries that helps like steer the LLM in a way to generate realistic queries. Otherwise, if you just ask the LLM to generate queries, it's going to generate queries in a way that it's probably not reflective of what your users will actually ask. And we also have some human alignment in the whole LLM judge process as well, where you're filtering document chunks.

49:05Yeah, so I think if we didn't have that, as we saw initially, we only had 46 % alignment. That's not very reflective of how you would want to evaluate your system. So yeah, we definitely have human in the loop in this entire process. So it's not 100 % auto-generated. Like we do have some generation, but we still need a good level of like human involvement to, you know, make this evaluation process like truly reflective of like your retrieval system. It's interesting that you're going for multi-party alignment, and it makes me think, it makes me curious about whether there's specific research on like how to align an LLM not just to a particular party, but to two parties.

49:53In this case, the user that's issuing queries, you're aligning to that user through the real-world query data set. But then you're also aligning to the creator of the system and their input to the LLM as judge part. Yeah, that's interesting to think about. I don't think there has been, or at least from like what I've seen, there hasn't been a lot of work in like, like LLM alignment on like multiple like parties, as you mentioned. Yeah, I've not come across it either. I'm going to look for it. You know, just the question kind of opens up thinking about like, you know, game theory implications and all kinds of potentially interesting stuff that could come out of it.

50:36Yeah, yeah. I'd definitely be curious to like hear about it further. But I think that's definitely like an important area of work, because I think, you know, like some of the research that I see, like a lot of people just use like full automation, like without any human involvement. But I think like, you know, like the trend we're like starting to see is like, you know, we do have more like human in the loop and like all these like processes. So yeah, I think like definitely like multi-party alignment would be interesting to see for sure. Yeah, yeah. I guess thinking about it more, you could argue that or it may be the case that like a lot of the core LLM alignment work is fundamentally multi-party in the sense that like, if you think about like instruction following, you want to follow the developer's instructions, but you also want to be useful to the user.

51:31Like, so folks are thinking about this. I'm so curious about like, you know, research formulations of it as a multi-party problem. But yeah, I'm curious, like what kind of research is being done in this space? Because I feel like, you know, you can't just look at like one like component of this. You kind of have to maybe look at like the isolated parts of like each party. Like how do you, you know, kind of like approach that in a more like systematic way? I think that's a, yeah, definitely interesting to think about. And so is this the future steps that we mentioned, are those your near term focus areas at Chroma or are there other research projects that you're involved in there?

52:10Yeah. So currently we're working on another research project. I think like there are some interesting directions of this, which, you know, we've been talking to a few people about. But I think, yeah, like immediately after this, we're looking more into like agent memory. I won't go too deep into it yet. It sounds like you don't want to throw the beans. But I think memory is becoming a very hot topic nowadays. I mean, you just saw ChatGPT memory. Everyone's very excited about it. There's a lot of work around just memory in general. But we want to do, I guess, kind of like a more robust research into how agent memory is implemented.

52:55What are the best practices in that? and yeah, just like agent memory broadly. But yeah, hopefully you'll see that in the coming months. Well, Kelly, thanks so much for sharing a bit about the generative benchmarking project and what you've been working on there. Yeah, thank you so much for having me. It's great talking. Thank you.

53:27Thank you.

From the publisher

In this episode, Kelly Hong, a researcher at Chroma, joins us to discuss "Generative Benchmarking," a novel approach to evaluating retrieval systems, like RAG applications, using synthetic data. Kelly explains how traditional benchmarks like MTEB fail to represent real-world query patterns and how embedding models that perform well on public benchmarks often underperform in production. The conversation explores the two-step process of Generative Benchmarking: filtering documents to focus on relevant content and generating queries that mimic actual user behavior. Kelly shares insights from applying this approach to Weights & Biases' technical support bot, revealing how domain-specific evaluation provides more accurate assessments of embedding model performance. We also discuss the importance of aligning LLM judges with human preferences, the impact of chunking strategies on retrieval effectiveness, and how production queries differ from benchmark queries in ambiguity and style. Throughout the episode, Kelly emphasizes the need for systematic evaluation approaches that go beyond "vibe checks" to help developers build more effective RAG applications.

The complete show notes for this episode can be found at https://twimlai.com/go/728.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Generative Benchmarking with Kelly Hong - #728The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 54 min
Listen in VO