In short
TWIML AI Podcast Episode Summary
Episode Title: Why Your RAG System Is Broken, and How to Fix It with Jason Liu - #709 Host: Sam Charrington Guest: Jason Liu, Freelance AI Consultant and Creator of Instructor Library Release Date: [Link to the Episode](https://twimlai.com/go/709)
Overview In this episode of the TWIML AI Podcast, host Sam Charrington converses with Jason Liu about the intricacies of Retrieval-Augmented Generation (RAG). Jason shares insights on common challenges organizations face with their RAG systems, effective strategies for diagnosing and fixing these issues, and the importance of robust evaluation methods. The discussion touches on embedding strategies, data quality, and the evolving landscape of AI applications.
Key Topics Discussed
- Introduction to RAG
- RAG systems combine retrieval and generation components to enhance the capabilities of AI models.
- Jason's background in machine learning and experience with recommendation systems provides a strong foundation for his insights into RAG.
- Common Challenges with RAG Systems
- Organizations often struggle with the performance of their RAG systems, leading to loss of customers and revenue.
- Common signs of a malfunctioning RAG system include:
- Misalignment between user expectations and system output.
- Over-reliance on external embedding models without sufficient context or data specific to the organization.
- Diagnosing Problems in RAG Systems
- Jason emphasizes the importance of moving away from subjective adjectives in performance discussions and focusing on measurable metrics.
- Initial steps for diagnosis include:
- Establishing clear, data-driven metrics for evaluation.
- Understanding the specific questions the RAG system is meant to answer.
- Building Robust Test Datasets
- The significance of constructing reliable test datasets for evaluating RAG systems is discussed.
- Jason encourages engineers to experiment with generating synthetic questions to test the system’s retrieval capabilities.
- Evaluation and Metrics
- Emphasis on precision and recall as critical metrics for assessing retrieval effectiveness.
- Jason shares insights on maintaining low-cost, fast evaluations to facilitate rapid iteration and improvement.
- Fine-Tuning Strategies
- Fine-tuning re-rankers is cited as a valuable strategy, especially when organizations lack extensive data.
- Discussions on when to consider fine-tuning embedding models versus relying on off-the-shelf solutions.
- Chunking and Context Strategies
- The impact of chunking strategies on retrieval performance and the way context is managed in RAG systems.
- Jason’s experience shows that the structure and presentation of prompts play a crucial role in improving system output.
- Multimodal AI and Future Trends
- Jason expresses excitement about the potential of multimodal models for enhancing retrieval tasks.
- The integration of structured data alongside text in RAG systems is highlighted as an area with significant potential for future development.
- User Experience (UX) Considerations
- The importance of UX in gathering feedback and improving system performance is emphasized.
- Jason shares examples of how small changes in prompting can lead to significant improvements in user interactions and data collection.
- AI Consulting Course
- Jason reveals his initiative to teach others about leveraging AI, focusing on enhancing personal and business productivity through AI strategies.
Key Takeaways
- Organizations need to prioritize understanding user needs and aligning their RAG systems accordingly.
- Building and maintaining high-quality datasets is crucial for effective RAG performance.
- Evaluation metrics should focus on precision, recall, and user experience to drive continuous improvement.
- Fine-tuning and context management are essential strategies for optimizing RAG outputs.
- Emphasizing user experience can lead to better feedback loops and improved system efficacy.
Conclusion This podcast episode provides valuable insights into the challenges of RAG systems and offers practical advice for diagnosing and fixing these issues. Jason Liu's expertise and methods serve as a guide for organizations aiming to enhance their AI-driven processes and user interactions.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Like the big smell I usually like to call out very early in the beginning is just you know customers saying, man, I really wish these models were capable of more complex reasoning. And it's like, do you want the complex reasoning because you haven't reasoned about what the customer wants? Because the more we really think hard and we really reason ourselves like what the customer wants, the product becomes much more clear.
0:35All right, everyone, welcome to another episode of the Twimble AI podcast. I am your host, Sam Charrington. Today, I'm joined by Jason Liu. Jason is a freelance AI consultant and advisor and creator of the Instructor Library. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Jason, welcome to the pod. Hey, it's hard to be here, man. Really excited. I'm excited for this conversation as well. We've spoken a few times before, and I'm looking forward to picking your brain on all things retrieval, augmented generation and more. And I think a fun way to dig into this conversation is to talk through when you come in and people's rag is broken, how they go about fixing it.
1:24But before we dive into that, I'd love to have you share a little bit about your background. Perfect. So, you know, graduated from University of Waterloo when we were doing a lot of physics. And basically after my first physics term, I realized, oh, man, machine learning is definitely going to be next. And I think the Nobel Prize, I nailed it, right? Absolutely. So once I started doing more machine learning, a lot of my background was mostly in computer vision and recommendation systems. So image embeddings, text embeddings to do recommendation systems. and it just so happened that now where we do RAG, it's kind of the same thing all over again, right?
2:02It's all text embeddings into recommendation systems that now feed into language models versus people. And so it was a very nice transition into RAG when I started doing more consulting. And you did some of that at Stitch Fix, if I remember correctly? Yeah, most of my background was at doing like multimodal embedding that Stitch Fix from like 2017 onwards. So like taking an outfit, putting it into some outfit embedding space and trying to predict the next outfit that the person would want in their box yeah it was you know how you do how do you do replacement how do you do these like recommendation carousels where's a similar item all that fun stuff and so um what were your first steps into uh you know rag and helping folks with uh their gen ai challenges it's funny so you know i basically took a year after i took a year off after i left stitch tricks And when ChatGPT came back out and they were talking about RAG, everyone was really amazed by the power of text embeddings.
2:59In my mind, text embeddings was my intern project in 2016 because I didn't know how to set up Elasticsearch. So it's been a very exciting ride to come back and say, oh, I have like eight years experience doing this kind of stuff. Let's jump in and figure out how can we actually learn new embeddings and how do we actually improve and measure these kind of search systems? Especially when more of these people now are just plugging in OpenAI embeddings and just doing some kind of vector search. There's not much room for improvement when it comes to retrieval. And so a lot of the companies come to me and they just say, hey, it's not really working.
3:32We're losing customers. We're losing a bit of money. How do we make this better? And that's kind of how the conversation starts off. Got it, got it. And when you say there's not much room for improvement around retrieval, what do you mean by that? I think a lot of companies, if you think about how we used to do embedding at Stitch Fix, Netflix, Shopify, Spotify, a lot of it is using user and product interaction pairs to train embeddings that are optimized for maybe a click-through rate or maybe for some kind of relevancy metric. But to me, at least, it feels pretty crazy that we're going to use these external embedding models from OpenAI and just assume that, oh, yeah, of course, my question about some law is going to be embedded very similar to exactly the paragraph that answers the question about the esoteric legal statement.
4:20That's not really true. If you think about even just the sentence, I love coffee and I hate coffee, are they similar or dissimilar? right they could be similar on a dating app because they're both preferences about coffee but maybe they're dissimilar because they're negative preferences against each other but you know i should be able to choose which one that looks like and i think that's where a lot of people are getting tripped up it's really a big assumption to think that we know what is and is not similar in this in this embedding space interesting interesting i think the reason why i zeroed in on on you mentioning that is because when i talk to folks that are working on rag and their rag is broken, they are often trying to fix it via tuning the generation.
5:06And that invariably is not the right way to do it. So there's often a lot more headroom in making sure that the LLM has the right context than fine tuning the prompts. But But let's maybe, you know, when you're called into those situations, like how do you begin to diagnose the problem? Yeah, so the first thing I do is I basically just ban adjectives as a word that you can use during stand-up. I think a lot of companies, you know, when software engineers - As in good, bad? It looks better. It feels better. Okay, better. 10%, 20 %? What does that look like? Like, that's usually the first step is really getting away from this, like, vibe-based estimate of the generation and really thinking about retrieval.
5:54I like to think about this bias, as I learned in this MBA book. The first one is called Absence Supplyness, which just says, like, you can't really think about the thing you don't see. And you see the generation. You always see the text coming out of the language model. So you think that's the thing that I got to control because I don't see the generation, the context. the second thing is intervention bias, which is I want to change things to feel in control. And so if you want to feel in control of a RAG application, all you got to do is just twiddle with the... Change some text and you're propped.
6:26And hope that all the relevant data is in there. And usually that's not the case, right? I think that's where a lot of the issues stem from, right? You're looking too much at generation and not really thinking about recall or precision and whether or not a language model is confused or even finding the right information. And that opens up the whole conversation around evals and evaluation loops and pipelines and flywheels and the like. And at least from my perspective, yeah, nine months, 12 months ago, like folks were trying to figure out how to spell eval and now like it's coming up in a lot more conversations.
7:05Are you seeing a similar shift? Yeah, except the thing I've been noticing is we've almost also delegated the scoring back to language models, right? The LLM as judge idea? Exactly. I think it's very useful to get some kind of proxy for what is good and what is bad. But what ends up happening is instead of trying to solve the relevancy problem, we're just solving the problem of, again, prompting another language model, right? I basically said, don't fiddle with the generation of the language model. You feel in control because you can fumble with it, but you're not going to get any results. and they said, okay, well, let me work with a different prompt instead, rather than building a precision recall data set.
7:47I think there are many, many tasks that take milliseconds to compute that we could run tests across thousands of examples and figure out what the relevancy looks like, rather than looking at LLM as a judge. Literally, a couple of weeks ago, during stand-up, a team was like, hey, who spent$1 ,000 this weekend? And some junior engineers are like, oh, I was trying to run evals to see how good the new changes are. Should I not do that? And say, oh, wow, they're so expensive. I've just incentivized someone to not run more tests. That's really not how I want things to be. I want tests that are really fast, really cheap, that you should be running every 10, 20 minutes when you make a line of code change in your system.
8:29And so it's been pretty funny to sort of see that transition and really push to be just a lot faster, fail faster, and do very cheap, cheap evaluations. And part of that is just that building data sets is hard and time consuming. And hey, pre-trained models was supposed to get me out of that business, right? Yeah. I think every data scientist, every data engineer is like, yeah, you know, I'm kind of the janitor, right? But then what happens is you kind of get like, you get the Roomba and you spill something, the Roomba just smears it all over the floor. and you're like, oh, I should have done it myself.
9:04That's kind of how I think about these things. But also, in earnest, I think what's really happening is because there are so many more engineers coming into the space with less data literacy, it actually is just very hard to even describe what a good data set looks like. Eugene Yang and I were basically trying to figure out what does good data literacy look like, and we really struggle. We can only come up with 10 reasons why it was data illiterate, but it's totally hard to describe. What is the intuition and the vibe of when do you give up, when do you try new things? If the model said it's 98 % accurate, you probably did something wrong.
9:43All those kinds of things are pretty undocumented, I think, when it comes to making this mental shift. So when you're talking to folks and you're encouraging them to take this initial step of building a data set that will allow them to measure their retrievals? Like, do they always know what that process, you know, needs to look like to do that? And, you know, how do you, you know, guide folks that don't through the process of building that data set? I think they have some ideas, but what ends up happening is some ideas feel so intuitive, you almost need someone else's permission to believe your or trust your gut, right?
10:21Like the simplest thing to do, for example, is say, you know what, given a text chunk, can I generate a synthetic question with a language model and then save these two pairs? And then let me check whether or not the question I just generated finds the text chunk. A lot of times I think there are engineers on the team that have this idea, but they just be like, hey, Jason, does this make sense? Nobody really knows for certain if what they're doing is ridiculous. And I think a lot of it, at least with the really great engineers I work with, they almost just need permission to trust their gut and do these tiny experiments because in engineering, there is just like the edge cases.
10:57You have to enumerate everything right away and then you sort of build out your test. But here it's a lot of it. It's like, well, we really do just have to try it. We just have to try 10 different things and figure out what works and what doesn't. And giving the engineering team the permission to trust their gut has also been a pretty valuable lesson on my end. And are those the same 10 things in every case or is it 10 edge case specific things that are different for every person who's trying to build a system? Yeah, what I find is one thing that is often missing is they don't actually understand what the workflows ought to be and what kind of question type they ought to serve, right?
11:36I think everyone wants AGI, right? G stands for general, and they want to solve every... But, you know, the open AI definition of AGI has something to do with like economic value. Like, are we unlocking economic value for our customers? and the big smell I usually like to call out very early in the beginning is just customers saying, man, I really wish these models were capable of more complex reasoning and it's like, do you want the complex reasoning because you haven't reasoned about what the customer wants? Because the more we really think hard and we really reason ourselves what the customer wants, the product becomes much more clear for example day one, we have a bunch of user questions coming in.
12:21If we can do some kind of clustering and segmentation, we might find out that, oh, wow, 30 % of all the questions are looking for contracts and whether or not they're signed. 10 % of the questions were just who modified the document last. That's not even in the text chunk, but if we just append an additional token that says modified by JSON, we can now just serve 10 % of our question base. If we just parsed out the dates and like an is signed boolean variable again we could now serve like 30 of our questions and so i think that the real trick is just developing the habit of looking at the data but also trusting that your your job is to make these hypotheses and your job isn't to be right all the time right and you have to be wrong you have to do these experiments and fail fast that that seems like it needs to be first like really understanding what the questions you're trying to serve with your system, whether it's a chatbot or something else?
13:22Because you can't even really build the data set until you know what those questions need to look like. I mean, sometimes if you just have a bunch of PDFs, you could try to have the language model answer these questions. But to my surprise, there have been times where, you know, if you use like Paul Graham essays and you generate synthetic questions off of random text chunks, you get like 96 and 97 % recall. The problem is too easy so I have to make it harder. But there are other data sets where I do the same task and I get like 60 % recall. For example? Right. If you just take all GitHub issues.
13:55GitHub issues? Okay. Yeah. Like a common question that gets generated is like how to get started. Well, it turns out if you don't have a filter on repo like repository you can't answer the question how best I get started because now there's filters involved. It turns out if I just say best ways to get started in repo, I now have to sort of parse things out and do some filtering. And maybe if it's not the exact filter, I need to do some string matching and all that other stuff. Are you saying when people are asking you the best way to get started with RAG or when people want to be able to serve or answer the question for their users, the best way to get started?
14:34I'm not following your example. Answering the question, like imagine doing like GitHub issue search and I search best way to get started. Right. How could I have possibly found the trunk that came from? There's thousands of best ways to get started documentations. Right. And then you do the, oh, OK, actually, in order to do this problem, well, I have to do some kind of like repo matching mechanism. I probably need to do some kind of like filtering mechanism. And now you slowly add complexity into the system that you build. I think too many people just sort of throw the data into a bunch of PDFs and go, well, obviously I can just ask you what the systematic risks of this investment is.
15:13And that doesn't seem to be the case. So if we're building up to steps, then, you know, one step is like, know your question. The, you know, next step might be build out your test set. And a third step is to, you know, think really hard about like metadata and like sourcing data that the LLM can use to answer the question or really that a preprocessor can use to get the right information to the LLM. Actually, we're still in retrieval at this point. Yeah. So I like to think about it this way. I'm going to do some segmentation. If I was going to do marketing, it might segment against men's and women's and East Coast, if it was West Coast.
15:56Every problem you want to solve, you kind of want to segment in some way. So we're ultimately going to find these segments in the question space. And ultimately, there's two kinds of segments. There's going to be segments that don't do well because we have capabilities issues. So for example, if I ask who modified this document last, if I don't have that metadata, I can't answer that question. So I need to improve my capabilities, right? The row exists, but I need an extra column. The other world is like inventory issues where the column doesn't exist. If you think of maybe like a DoorDash and you find out that Greek restaurants near me is a terrible search query, the solution might be to buy iPads for Greek restaurants so they can come onto the platform.
16:45So there's been other times when we do this kind of debugging and we realize, oh, wow, we don't have the data to answer these kind of questions. We don't have the scheduling information to do this. We don't have the tables extracted to do this. And so there's usually two kinds of solutions you got to start integrating. Are there more capabilities by adding like more rows to your data set or just like more columns or add inventory and include the number of rows we have on our data set? And that's kind of like how I think about these things. Adding rows in all of the examples you've given is a much longer process than adding columns.
17:18Exactly, exactly. And I think sometimes it's like, oh man, we need to figure out contracts or we need to figure out, you know, if someone is responsible, the feature that people really care about is can we contact them in some way? And now you start building an actual application because you know what the customer wants, right? It's not just to ask a question, but to take some action, make some decision. You mentioned this issue of using off-the-shelf embedding models. Are you advising folks to build their own embedding systems or just get better data to the ones that are there? Or is there a decision matrix around?
18:02And in fact, there are a lot of similar questions that I run into that I don't, I've heard mixed, you know, mixed opinions as to like, if they're even important. And it's like the embedding scheme, the chunking strategy, like headings and other, you know, contextual information around chunks. Like, do you have, is it, are there one size fits all answers to these or like a decision matrix? Or like, how do you think about the space of like, you know, implementation details around embedding? It's a really good question. Primarily because why guess when we can test, right?
18:49I think this is a matter of really just, you know, investing more in a data set and just going, right, well, let me just run like 30 experiments, you know, over the weekend, I'll come back and I'll just know the answer. And I think that's kind of, even the nature of that question to me is a symptom of sort of the lack of having really good evaluations on your own data sets. But likewise, the answer to the question is an indication that there aren't clear patterns and it's very data set and use case specific. And, you know, if you said, for example, yeah, like, you know, we really only ever see chunking strategies giving a, you know, one, 2 % lift.
19:33It's not usually worth it. Like that would be really informative. But you didn't say that. So, you know, there must be cases where you change your chunking strategy and you, bam, get some great results. Exactly. A good example of that might be thinking about chunking and then completely processing like tables within PDFs differently than regular text chunks. It's like, OK, like paragraphs you chunk and then if you see a table, we have to save the entire table somewhere else as a separate index. Right. That would be a good example of like when chunking really matters. I would also say, you know, if you have even thousands of examples of questions and labels on whether or not a chunk is relevant, it's probably pretty fruitful to fine tune a re-ranker.
20:17But even then, I think oftentimes I surprise myself in whether or not certain interventions perform better. There are times when using hybrid search with embeddings and BM25, BM25 is like 3 % better, right? That often is the case if the person who is searching the data is aware of the file names and the text that's in the data. Like if I wrote my own essay, it'll be easier for me to find it if I use full text search because I know what I wrote. Right. There's been times when re-rankers don't improve the performance of the model. And then there are times when the re-rankers do. And again, it just becomes a superstition.
21:00but it is very easy to sort of absolve myself of the superstition by having these tests that run really, really fast, right? It's just, is the order better? These things are really great. Whereas if you go into factuality or like self-consistency and like context recall, who knows, right? Maybe the model just wants to choose itself. Now you're doing a whole set of other experiments to prove that the model is aligned with a metric you never made up. And that's when I think things get really good. Yeah. and expensive yeah even time-wise I feel like there's been times where we have summarization prompts, so for example when we want to retrieve images, do we use a clip embedding?
21:46Well that means we also have to use a clip embedding for the text. What I've seen go really well is actually using a visual language model to give a detailed description as like a paragraph of the image and then just text embed the paragraph. But what this means is now the describe this image prompt is another hyperparameter to experiment against. And I've seen situations where if you just say describe this image, recall is like 27%. But if you can teach this concept of recall to an engineer and you just make them hill climb for like a day and a half, we've been able to get to like an 87 % recall just by improving the prompt.
22:27write a prompt tell me what doesn't recover find the blueprint but also count the number of rooms now it's like 35 % also transcribe the street addresses and include that in the description 70 % also describe whether it's north facing and east facing also describe the positions of the cabins and all of a sudden you have a 96 % recall system for finding blueprints because you actually worked on the prompt It sounds like feature engineering. I can't say that because then they got confused, but it ends up being it, right? It's like, oh, this is a classical machine learning, but we're able to hill climb because our eval is very fast.
23:11It takes, you know, 50 milliseconds to try again and try again. And I think that's where a lot of things can be really optimized for. But yeah, it's definitely just feature engineering, but that's also another word. are you um you know one of the questions that that comes up a lot around the the whole idea of evals is like tooling like um do you have you know go-to answers for that i'm guessing it's going to be yeah build your data set and like you know some silly eval in a you know notebook but uh do you find that there's a point at which it becomes more complex and there's you know some you know open source or off-the-shelf tooling that makes a difference for folks i would say if you are working independently you are likely best off just like writing things through a json line file or like a sqlite file primarily because you're just building out these very fast devals, right?
24:11Like if you're just comparing like length of summary divided by length of input and figuring out if there's a compression rate I want to set a goal against super fast. What I do when I work with bigger companies is I use Braintrust primarily because like Braintrust was basically built because Anker has just been like sharing screenshots of like results from a Jupyter notebook every once in a while. You're like, oh, I need a tool that does better than this. And so oftentimes if you need to collaborate on sharing data sets, collaborate on sharing results and getting feedback and, you know, coercing your team to label data with you, I think that's when a tool really, really shines.
24:49Just for the collaboration aspect, not because the evaluations are better in some kind of way. Exactly. Because the evaluations, you have to build yourself. Right. If you're pulling evaluations off the shelf. Yeah. The actuality, the self-consistency, that stuff to me is crazy. It's like having someone grade their own assignment. It's like, oh man, I hope.
25:18And then the other stuff that's valuable is like, okay, how can I downsample my production traffic to also run these evaluations to make sure that things are running productively? Can I monitor these things over time? A really simple example is just, I have a company that we do a meeting summarization and we, we plot the average length of a transcript and we plot the average length of the summary divided by the average length of the transcript. And every once in a while, there's like a blip. And like, why did that blip happen? Well, it turns out, you know, we ran a marketing campaign and we got a whole new set of users and these new users are doing three hour long podcasts.
25:57And it's really bad because the summary is just like, they talked about AI. right you're like oh man like the the compression rate is too high now let's go do something great when you build a rule that says if the call is less than an hour we can use this prompt if we use greater than an hour can we use that prompt okay the ratios are like recovering a little bit and can we monitor that and i think that's how i think about building these systems have like the dumbest evals possible to tell you what to look at uh you mentioned compression rate previously before talking about this specific example is that a metric that you've applied broadly or is it just this transcription summarization thing yeah i mean i've just found a lot of the applications i tend to work on are ones where we're doing a lot of summarization and summarization is a very uh interesting task because like llms are good at summarization in the sense that indeed the output is shorter than the input but it's actually very hard to evaluate like what is a good summary and like when do we lose nuance and all that kind of stuff and so you know obviously we can have the entire like llm as a judge model of doing things but ideally we have much more like much simpler metrics right so i have metrics of just you know length of summary divided by length of transcript.
27:25I also have counts of named entities, right? For example, if the summary is all mentioning my name and versus the summary just going like, they thought it was, you know, it's like, it's very like ambiguous. And so can we, can we, can we preserve some kind of information density there? They're all proxies for some, you know, satisfaction or nuance that we can also use a language model against. But looking at like odd examples of just simple numbers still can tell you a lot of information. Like what we found was when we plotted summary length by transcript length, it would go up, and then after like 20 ,000 tokens, it actually got shorter again.
28:08I was like, okay. That took six minutes to plot out and write the data for, but now we know that there's some weird behavior when the transcript's really, really long. Great, let me change my prompt, rerun this, oh it's straight again perfect we're good was there a step before changing the prompt that was trying to understand like the intuition for why that might be happening or was that ancillary to actually getting the problem i don't remember like what we did in that example i i think we we just kind of saw that like oh wow not only is it getting dropping the variance is also increasing as we drop and so what we want a prompt that has lower variance in the like compression rate that seems like a very like healthy and quantifiable goal where we can just say hey jason i tried three different prompts and i was able to drop the standard deviation by like 40 and it now is like monetarily increasing as a function of context like that becomes so scientific and so quantifiable that we don't have to worry about some of these like bigger things with it.
Read the full transcript
29:14Obviously, we might still lose nuance, but setting a goal against that is very easy. My sense is that folks coming to this
29:29fresh and being told, hey, you should build a test data set, kind of wrapping their head around precision recall, whether you can keep those two straight without looking it up that's another issue but like you know that's like oh that's probably something that i need to be able to measure and test against is like um you know an obvious thing uh compression rate feels like less obvious or more nuanced in some way are there other kind of nuanced types of things i think either i guess i'm thinking there are two ways that you get this It's either one like, you know, banging your head against your problem.
30:12And, you know, this is probably the best way to come up with these things. But, you know, part of what we're trying to do is like accelerate learning and provide shortcuts. Like, what are the shortcuts that you've come across for different problem classes? Like, oh, these four metrics, like you probably wouldn't think about them. But, you know, when you did, like you discover that these come up all the time. Do you have that list? another one that is pretty reasonable in this like summarization task is just whether or not it adheres to a certain schema and a certain formatting but again i try my best to just write a regular expression that tries to capture this as quickly as possible right and you know i could have like a six point rating scale on whether or not it fits the the markdown format that I want, but then I lose all, like there's too much nuance, right?
31:02Really, I just want to have a bunch of pass-fail tests that are very binary where I can say, great, show me 10 examples where I failed, 10 examples where I succeeded. Let me just go like think really hard and figure out what is happening and how can I change that. Outside of that, I find a lot of it ends up being very, very specific, right? There's an example where I generate action items, but I want to evaluate whether or not the action item is correctly assigned to the person, right? That's just a very specific eval that you have to build. And it ends up being very challenging sometimes to correctly assign something like that.
31:39Are you finding that you're always running all of your eval suite whenever you're making any change? or do you find like running specific, you know, feature specific evals and then running your broader suite less frequently? Yeah, it depends on like what kind of eval, to be honest. If I can afford to, I would just rather run them all the time because you never know what kind of cross influence there is. Like a really simple example was we had both an executive summary and a list of action items. and the action item description was too long. And you're like, great, well, make the action items shorter.
32:28And then we got an equal length action item, but just fewer action items. Right, but that test is easy because you can just like parse out the asterisks and count them and that's one evap. I literally had an evap that was just count the number of action items. The second one was like, what is the average character count of the action items and what is the average summary count? So then you do, okay, well, just make the description of the action items shorter. And then all of a sudden the summary is also shorter. Right? So there is cross-contamination. And one of the things that's valuable is to go, okay, well, I don't know why, but controlling one and not the other is so difficult.
33:12I'm going to break this down into two tasks. I'm going to have a summary task and an action item task. And the reason I've done this, the reason I've added this extra complexity is because I have all these experiments to prove that I can't figure out how to combine them. Maybe if a new model comes out and it's better and more steerable, we can reevaluate this. But I'm going to separate these to two different tasks because I cannot get the evals to match what I want in terms of performance, right? So now you can have this idea of like, I'm going to segment to make this simpler, but there are conditions when I would recombine these tasks.
33:53And maybe when Haiku 3.5 comes out, I'll rerun my old evals, see if I can fix these things and justify some of these investments. But the idea really is you're making your resource allocation and how you spent your time and how you designed your system and its complexity based on the trade-offs you're making with evals. And again, these evals are just regular expressions or not anything fancy or you're not calling any LLM. Fine-tuning comes up in the context of RAG and GenAI more broadly. What role do you see for fine-tuning in the types of systems that we're typically building for RAG? I would say if you're going to start fine-tuning, the first thing to fine-tune is likely going to be something like a CoHero re-ranker.
34:43That's where you're going to have the least amount of data you required. It's going to be very easy to label this data. If you have a bunch of questions and text chunks, you can probably ask the smartest, most expensive model you have and just label thousands of examples. So just transfer learning that task is pretty affordable. Probably for$50, you can get a fine-tuned re-ranker that outperforms anything off the shelf. I don't know whether fine-tuning embedding models is worth it just because of like, it's just annoying to like own inference. But I think that's the second easiest thing to fine-tune is fine-tuning an embedding model to do search better.
35:23After that, the only thing I would really fine-tune is any kind of rewriting step. So, you know, can I, given a question, parse it out to, you know, query, start date, end date? Can I map it to metadata filters? That, I think people should be fine-tuning because it's a very specific task. You can fine-tune a LLAMA model, you can host it in a way that has fast inference, and it's usually going to be pretty effective. Whereas I find it's pretty challenging to really think about how do you fine-tune 4.0 to do answer generation. You would need to be pretty justified to explore that, especially because these models are going to get better in a way that we can't control.
36:06And they're always going to have better recall. They're going to have better robustness towards like low precision text chunks. It's hard to beat them because they actually have all the data. Whereas for something like re-rankers, you know, they don't have that data. Like we are the ones that are able to capture the value, have this data set, fine tune and outperform the public benchmarks. One of the model-related questions that comes up all the time, especially as the big models get better, is do I need to think about any of this in a large context-length regime? like do I need to re-rank do I need to you know embed like chunk can I just throw everything you know of course for some definitions of everything it's going to be too big you know bigger than whatever context when do you have but like assuming a large context and you know assuming context sufficient for you know a lot of your context um yeah you get where i'm going with this yeah like yeah i mean i think what's really going to happen is we're going to go in the same way that the iphone battery life has gone right like we've never had a better battery and then longer battery we've just had more powerful applications and so i think as context increases we're is going to have way more complex instructions with like different personalities or you know maybe not only is going to have the context length it's going to have my you know my history all this kind of stuff that said i think there's a great place for long context models especially when we have a few documents i would almost rather always shove everything in the context right but we're always going to run into latency trade-offs if we think about the recommendation systems or you know e-commerce systems, we know that even 100 milliseconds, 300 milliseconds of latency could be a 1 % revenue hit.
38:21I think that'll be the same thing for these language models. There's always going to be some frontier of context and latency and business outcome that we're going to have to make trade-offs against. I think that's what's really going to happen. Yeah, I was wondering if you had more nuance around the way you think about the generation side. I guess my observation is like the length of the context itself is insufficient as a determinant of success. Right. Then, you know, that's why, for example, we have re-ranking because, you know, within a given context length, the model can't really follow the plot all the way from the top to the bottom.
39:03Right. And so like just saying like this is the number of context length doesn't say enough about how well the model, how good a job the model does at attending to all the various things in the context. And so, you know, that gets to, you know, concepts like precision and recall and other things. Like, do you have a structured way that you think about that or like the way that you would approach evaluating a different context length? So when I use a longer context model, I'm usually working with a very few set of documents. So the question is, is the relevant information in text chunks split across many documents, or really is it going to be a very few documents?
39:46And a good example of this is we have an agent that's job is to take sales calls, reference your pricing pages, and give you a compelling personalized pricing on a certain service that you provide. So we have a one hour long transcript. We have a 16 page PDF that describes our pricing options for like different add-ons and whatnot. And the prompt goes as follows. It says, here's a transcript. Here is 16 pages of our pricing. First, list out all the variables that are required to determine whether or not you can personalize the price. And then it does that. And then for everything that we list out, extract out exactly what part of the transcript they mentioned this variable.
40:30So first it lists out the variables, then it lists out the variables hydrated by excerpts. Then, you know, reread the transcript and the page and list out the resulting, like, price number that you can give it, and then construct a follow-up email that offers a personalized price. so what we're really doing is we're trying to just push the language model to do a lot of very specific chain of thought where you're kind of extracting the data then organizing again in a smaller package and then as you generate the email you assume that we're kind of only attending over this like prepared notepad that the language one of them and that's mostly how i think about using long context models when it comes to very few data which is just to say i want you to attend over everything, reorganize the information in your chain of thought in your scratchpad, and then finally give me a final result.
41:23And that has usually worked pretty well. That's been the difference between something we could ship to something that is actually sending follow-up emails right now in production. That's really interesting. So it's kind of speaking to like long context doesn't necessarily say that you're going to be able to one-shot your answer, But if you can reduce your context systematically, you know, through a lens of the way you thought about breaking down your problem, then you could, you know, long context can be a convenience for you. Exactly. And it makes the problem really, it's still a single prompt now, right?
42:00It's just generating like scratch pad one, scratch pad two, scratch pad three. and it actually lets us work in a world where when the next long context model exists, we can just replace the model number rather than going, you know, we have this like six prompt agentic system and I hope - Okay, yeah, I was envisioning a six prompt agentic system. Oh yeah, this is a six prompt. Tell me what does the scratchpad mean in that context and how is the prompt incorporating that? So basically I said, okay, first list out the variables, then list out the variables and the transcripts but it's just doing it so just in the generation you're asking it to show its work that kind of deal yeah but like so it's like a very very long show your work right it's like it's you know maybe 3 000 tokens of planning of just going like well uh we can offer a per seat model if the per seats is greater than 30 uh this person mentioned 30 was the minimum seat number and they said they had 48 seats so now and just sort of like goes down this but as a single generation interesting and are there is there anything that you need to do or prompt magic to get it to kind of stick to the steps or does that generally work pretty good for you know sufficiently advanced models so for the advanced models because we have this long context we just have like four or five examples of this entire reasoning protocol right that's another reason that the long context value matters because now we have the transcript 16 pages of calls and four examples of reasoning about the variables needed to create pricing pages and and you know like if it's lower than this price offer this package to do this um that just becomes like way way more context that we can use because we have a longer context model but then ultimately you run the thing it's like 178 000 tokens of prompt use right and i think what's going to happen is as context models increase, we're going to have much more sophisticated few-shot examples.
44:04Maybe we have full examples of transcripts in the past and how we reason about them. We're just going to saturate everything as much as we can. So got all our basics lined up, closed loop evaluation, considering techniques like fine tuning. Are there other optimizations that beyond fine tuning that someone might think about once they've got the basics lined up? In the model, maybe less so, but I think a lot of people are sort of ignoring the UX and the product facing side of things, right? For example, if we focus on streaming, we can make the perceived latency decrease, right? If we just focus hard on building great copy and great UI to collect feedback, we might be able to start fine-tuning Reranker sooner rather than later.
45:01For example, if I generate an answer with a bunch of files, what if I gave the user the ability to delete one of the files and regenerate an answer? That becomes a negative sample in your Reranker because now we know that was irrelevant. There's a lot of UX features that we can do. One of the great examples I discovered with Zapier, working with them for a couple months, was we changed a copy of how did we do to did we answer your question today? And that in itself 5x the amount of feedback we were able to collect per day. And that basically, you know, within, you know, within one month, we got enough data that we could get together as a team, review all the examples and figure out what we want to do next month.
45:46Just because we have volume. And stuff like that, I think is really underlooked, especially when the UX can also be used to educate the user. If we discover question types that are low volume and low success, maybe we just say no to answering those kind of questions. If we have question types that are low volume but high success, maybe we preview that as an example question that you can ask and teach users that we can actually do this very well and we should be using this to answer those kind of questions. right a lot of that education in the ux i think is something that is often underlooked at smaller teams yeah along the lines of the ux one of the things that i've been uh talking a bit about recently is this idea like hey we've all started with like trying to replicate chat gpt for our business's data but that chat experience isn't necessarily the best experience for everything In fact, it might not be the best experience for a lot of things.
46:49And at least in an enterprise context, maybe it's different on a product context, but in an enterprise context, a lot can be gained by integrating what you're trying to accomplish with RAG into an existing workflow as opposed to creating some new standalone chatbot. Is that something that you see in the folks that you work with? Yeah, one of my most popular takes on RAG is that question answering is very low value and cost-centered centric. Whereas, one of the big things I see in the companies that I've been advising, like Vantager, for example, like Vantager.com, for example, they do report generation.
47:32So instead of saying, given a data room, can I ask a bunch of questions about how the founders met and what is the TAM of their business? Vantage just says, if you give me a data room, I will just pre-generate every report that you use to make a decision in your business. And now you can just use the workflow of reviewing reports. But now you can start processing 40 businesses a quarter. You can do 80 businesses a quarter. right and now the question is instead of capturing the percentage of the cost of labor we might be able to capture a percentage of the roi of the decision and i think that's where a lot of really great rag applications will will come about right can we capture the roi rather than the cost of doing this kind of work and that also lends itself to uh kind of progressively inserting rag into multiple places in a long-running workflow or business process yeah and it goes back to this idea that if you if you wish the agent had complex reasoning is because you have not thought hard about the problem yourself it's a spicy take but i oftentimes you know i think people admit to uh agreeing with it even just a little bit yeah multimodal's uh popular topic is we've talked a little bit about like extracting tables from reports and some of the ways that you are extracting metadata from images.
49:04Are there other ways that you see multimodal coming up? Yeah, I think, you know, if you follow like Joe from Vespa or Ben from AnswerAI, they're all very excited and me included, very excited on the models like Copali, where we use visual language models to do search effectively. And not only can you do search, you can then use visual language models to give in the images, answer the question. And because it's all local, well, not all local, but it's open weights, you can also inspect the attention mechanism. So when I ask a question on a PDF, I can tell where the model is looking to determine its relevancy.
49:45I think there's a lot of features there that can be very useful in the context of maybe, you know, if we have hundreds of PDFs with hundreds of pages, we can use something like Copali to really be great at, you know, reading diagrams and understanding structure without thinking about the OCR and the table extraction and all that kind of work. So that's something I'm very excited about exploring. we've talked a little bit about uh agents and um you know there's one dimension of agents that is like breaking up your prompt into a bunch of steps and using that as a kind of a reasoning mechanism um but there's you know i guess you could argue whether this is an agentic thing or not but like function calls and tools and stuff like that do you see those capabilities coming into play in the rag systems that you're building?
50:37Yeah. I mean, I think the real question is like, how many hops is my rag agent allowed to take? Right. Like, can I do retrieval and then determine that I still need to do more retrieval? Or do I only have a couple attempts to answer the question when I retrieve data? I think the general idea is that if we can segment the problem space or the query space into these different buckets, it probably benefits us to build specific indices to serve each set of questions, right? If I know 40 % of the questions I ask are going to be around scheduling, I might just develop a data structure optimized for querying schedules and then have a function call hit that API, right?
51:22And then I think function calling is effectively just building out routers that can combine these separate indices into a single API. and letting the language model determine what's going on. I think there's also a world where if we have many, many tools, we might want to do retrieval and search to figure out what tools are relevant. Imagine if we had 200 tools at our disposal. It's just another precision recall evaluation to figure out whether or not the question is finding the right tool. But I think for the most part, I've just been able to really just push precision recall to almost be the hammer in a world where I think LLMs are the hammer for everything.
52:01I've just gone back to basics. And your last example spoke to another interesting thing that I've seen is trying to get beyond building right systems just with a bunch of text, but also including structured data, which can be incorporated via tools. It sounds like you're seeing at least some of that. Are you seeing that grow? I guess in the report generation example that you mentioned, that would be a big part of it, right? Yeah, exactly. I think ultimately it will just be function calling plus the messages are right. And I think that can probably do a lot of these cases. And outside of that, one thing I like to say, I forget what theorem or paradigm this was, but it was the idea that all complex systems are derived from pre-existing complex systems.
52:58And if you think about chatbots and finite state machines, you know, LandGraph is covering that basis, right? It's kind of the LLM extension of a system that already works. You know, if you think about like OpenAI Swarm Library, it's very much like message passing and distributed systems and like the actor model of programming. So I think we're already slowly seeing these different forms of agentic programs being remapped to known, successful working paradigm for building out these kind of complex systems whether it's like line graph or swarm or anything like that i think we've sort of figured out what works and now our job is to scale that better and better i guess maybe changing topics slightly yeah uh beyond rag another thing that you're very excited about is like helping folks kind of tool up as AI consultants?
53:55Like, where did your interest in that come from? Well, I just struggled so long myself personally. You know what I mean? I feel like I didn't work for like a year. I came back to you and I was like, oh man, people are asking me for help. I don't really know how to turn this into a business. You know, even a year down the road, I really feel like through a lot of like AI augmentation, there should be more and more individuals who are able to scale up their own knowledge work with LLMs. So I was like, well, I just think there's going to be more businesses, like more solo business, like entrepreneurs making six or seven figures.
54:29And so, okay, if that's true, I should try doing it. But I just realized that, you know, I think the, like if you're a technical person and you enjoy technical work, it is very hard to do the sales and do the writing and figure out how to write proposals to like charge more. I think everyone tells you to charge more, but you don't know what that means. And there's no playbook for that, right? to just say things like, well, just look in the mirror and name a price and then double it. And if you don't laugh, keep doubling it. And at some point you can just ask that number. And that number worked for me.
55:04And so I basically took a bunch of courses, I write a bunch of books, and I'm trying to distill everything I know into a little package on Maven and kind of just like sort of save the regret and the embarrassment of undercharging for so long. and, you know, sort of pass it forward and help them have everyone else. I'll figure it out. Yeah. You know, I feel like the first job I did, I asked, they asked me how much I charged and I said, oh, between like 150 and 170 in an hour. And they just said, great, we'll do 170. Just I mean the paperwork. And I was like, oh, you answered in three seconds. I just, I called my girlfriend.
55:42I was like, hey, I just took food out of both our mouths. I'm really sorry. Like I'll do better. Maybe I'll double it. I don't know. I'm nervous. And yeah, the course I'm running this, this next month is my goal to not do that again. Yeah. I can say that everything that you mentioned is true for being an industry analyst and a podcaster, you know, either slash both, which I am, you know, I've been at it for quite a long time, but you know, there's definitely a learning curve and it changes all the time too. So it's awesome. Exactly. Well, we will link to that course. Are you still doing the RAG course as well?
56:24The RAG course, we're running it again in February 4th. It'll also be six weeks. I'm pretty excited. We already have some folks from OpenAI who's taking the course now. So hopefully I can help with their solutions as in years, improve other people's RAG systems. And so I'm very excited for the new cohort. There's a bunch of really amazing companies involved. Well, we will be sure to link to those in the show notes and maybe we can work out some kind of discount code for listeners or something. Yeah, let's do it. Awesome. Awesome. Jason, it has been great catching up and I feel like we probably could have continued on for another hour, but we should make sure to keep in touch.
57:04Thanks so much for jumping on and sharing a bit about your experiences with us. It's been super fun. Thanks so much. Awesome. Thank you. Thank you.
From the publisher
Today, we're joined by Jason Liu, freelance AI consultant, advisor, and creator of the Instructor library to discuss all things retrieval-augmented generation (RAG). We dig into the tactical and strategic challenges companies face with their RAG system, the different signs Jason looks for to identify looming problems, the issues he most commonly encounters, and the steps he takes to diagnose these issues. We also cover the significance of building out robust test datasets, data-driven experimentation, evaluation tools, and metrics for different use cases. We also touched on fine-tuning strategies for RAG systems, the effectiveness of different chunking strategies, the use of collaboration tools like Braintrust, and how future models will change the game. Lastly, we cover Jason’s interest in teaching others how to capitalize on their own AI experience via his AI consulting course.
The complete show notes for this episode can be found at https://twimlai.com/go/709.




