In short
Retrieval-Augmented Generation (RAG) for customizing general LLMs (e.g., “Chat with my docs”) by retrieving relevant document chunks at query time instead of retraining.
Guest backgrounds
No guests mentioned; episode is hosted by the Linear Digressions speaker.
Key claims
RAG works well for “lookup” questions (needle-in-haystack) and for frequently updated knowledge by updating the document store/embeddings. It fails for multi-hop reasoning and for questions requiring synthesis across many documents. Retrieval quality is the main failure point; adding too many retrieved chunks triggers “lost in the middle,” where mid-list relevant chunks get ignored.
Notable examples
Return policy Q&A; “Who is the CEO of the company that makes ChatGPT?” (multi-hop); “How has benchmark landscape changed over last couple years?” (cross-paper synthesis).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding the Limitations of Traditional AI Approaches
0:58 to 3:15
Explore the challenges of traditional AI training methods and their limitations.
“You are listening to Linear Digressions.”
How RAG Works: The Retrieval-Augmented Approach
3:15 to 7:40
Discover how RAG enhances AI responses by integrating retrieval mechanisms.
“do, this is where RAG comes to the rescue.”
Limitations and Challenges in RAG Systems
7:40 to 12:16
Examine the potential failure modes and limitations of Retrieval-Augmented Generation.
“If you want an analogy, it's kind of like asking the AI to take an open book test.”
Best Practices for Implementing RAG
12:16 to 14:00
Learn effective strategies for using RAG in AI applications.
“And this is really important because it's pretty often true if you have a RAG system that isn't working very well, it's more than likely that the failure is somehow routed back to the retrieval step.”
Challenges and Strengths of RAG Systems
14:00 to 16:02
Explore the strengths and limitations of Retrieval-Augmented Generation systems.
“And so at a certain point, having a really, really long list of stuff that's been retrieved isn't really helping you very much more.”
Future Insights on RAG
16:02 to 16:28
Learn about upcoming solutions and improvements in RAG systems.
“So stay tuned to a future episode for a little bit more about that.”
Transcript
Automatic transcript. May contain errors.0:00Today we are going to talk about the feature with the worst acronym in generative AI, RAG, or Retrieval augmented generation generation. If you've ever used something like chat with my docs, if you have an internal AI chat bot maybe that has access to your company's documents or you've created one yourself on some kind of personal project and you've uploaded a bunch of documents for the AI to use, you have encountered RAG whether you know it or not. It's an extremely effective technique, works super, super well for taking these general purpose models, like your chat GPT or your Claude, and turning them into AIs that are aware of all the specific information that makes them truly useful in a huge variety of situations.
0:50RAG is pretty interesting, how it's working under the hood, and so I thought it would be fun to spend a little while talking about it. You are listening to Linear Digressions. So let's start with all the stuff that doesn't work if you want to have an AI that's customized to a particular use case for you. One thing you may have been thinking, let's say you were trying to solve this problem five or so years ago, and you didn't know that RAG existed because it didn't exist yet. The first thing you might think is, maybe I should put the specific information that I want the AI to know, maybe I should put that into the training data.
1:28And it's not a terrible approach, but it's going to go stale almost immediately. You're not going to be able to retrain the AI on new information incrementally. Topic for another day, very interesting, but it would be prohibitively expensive and just really impractical to have the training data be something that has to be changed every time you want to have a little bit of customization. Another thing you might be thinking is that you take all of the relevant information that you need and try to put it into the context window. So if there's a question that you have about maybe a long paper that you're reading, upload the paper and then boom, all in the context window.
2:10AI reads it all in a few seconds. You can ask questions. And this works if your document is tens of pages, maybe hundreds of pages at this point. But there's a whole class of problems where you're going to run out of context that way. And if you start to probe the upper limits of that push everything into context, you start to run into other problems where if the answer is somewhere in the middle of the document, it gets kind of lost in practice. So this isn't great either. If you're an AI nerd and you work with models a lot and you're comfortable with fine-tuning, you may be thinking about maybe that as a option.
2:52You could take some of the specific things that you want the model to know and do a little bit of a post-training fine-tuning step on it. But this is still pretty expensive, still something that's not going to be able to keep up with documents that are changing frequently. And you still might have this problem where the model's memorization of that information isn't very reliable. So what are you supposed to do, this is where RAG comes to the rescue. The core idea of RAG is in the name, Retrieval Augmented Generation. The whole idea is that when you're asking a question or posing some sort of prompt or query to the AI, at the same time you pose the question, you give it additional information by some kind of retrieval mechanism.
3:39And you hope that the information that the AI needs to answer the question is in that information that you retrieved. How do you retrieve the information? There's a setup step that you have to do first. That has a few pieces to it. Most common use case for RAG is when you have lots of text documents. So let's just think about text here for a minute. Obviously, that's not the only kind of information in the world. But you need to create a database that's got all of the text documents in it that might have the information that a model might use to answer one of your questions. This can be a pretty big database.
4:16This can be hundreds of documents. This can be thousands of pages. Really big. Now you need to get it into a format that's suitable for the AI. So what are a few pieces to this? Number one, you don't want the AI to have to read all of that to come up with an answer. You want it to be able to selectively find the pieces of information in there that are probably going to be relevant to answering your question. So there's this step called chunking, which is where you take all of that text and you break it into pieces, like a paragraph a piece, a page piece, something like that. So bite-sized pieces.
4:52Step number two, embedding. This is also sometimes called vectorization. You want to make each of those chunks easy to find by the AI. What embedding does is it takes that text and it puts it through an algorithm that turns each chunk into a string of numbers. Those strings of numbers are what the AI uses to find the pieces of information that it's going to use. Then it stores those vectors, those strings of numbers, those embeddings in a database. And then there's a connection between the embedding vector and the original chunk of text that it refers back to. Now, when a question comes in, the prompt from the user is embedded using the same algorithm as the embeddings that were created from your document store.
5:38So the prompt gets turned into a string of numbers that's in the same computer subspace as your embedded document store. Then what the AI can do is it's got a string of numbers that corresponds to what the prompt is that just came in, a bunch of strings of numbers that correspond to all the information that it knows in its document store. And it does a similarity search. It compares those strings of numbers to each other and finds the strings of numbers in its document store that are close by, that are really similar to the vector that came in. So the whole point here is that this is a very fast operation for the computer to do.
6:17It's just comparing strings of numbers and finding strings of numbers that are pretty close to each other. Once it finds strings of numbers that are close to each other, each of those strings of numbers corresponds to a chunk of text, remember, goes and retrieves the corresponding chunks of text and inserts them into the prompt that is going to send to an AI to generate an answer. So a question might come in, what is our return policy? So it's going to go in and every chunk that makes some kind of mention of refunds or returns, or maybe it'll also have exchanges, other types of semantically similar ideas, those are all going to have embeddings that are kind of similar or close to each other in that embedding space.
7:03And it's going to retrieve all of those chunks. And then the prompt that gets sent to the LLM behind the scenes says, I have a question from a user. What is our return policy? Please answer this question using the information below. Paragraph number one, our return policy is blah, blah, blah, blah, blah, blah, blah. paragraph number two, if you want a refund, blah, blah, blah, blah, blah, blah, blah, and so on and so forth. The LLM takes in all of that information, the question and all of the possible answers or pieces of relevant information spits back the final answer that goes to the end user.
7:39Works really well for a huge class of problems. Because then if there's new information that you need to introduce into the system, you can just update your document store, update your embeddings, stick a couple more chunks in there, and it's going to be available to the AI to retrieve on the next pass. If you want an analogy, it's kind of like asking the AI to take an open book test. So you're giving it all of these books that have answers to all of the questions that you think you might have to ask. And then when it's generating the answer for you, you give it the ability to go look up the answer really quickly behind the scenes before it says something back to you.
8:17Okay, great. Very exciting. So when does this break? And I think this is where this starts to get really fun. This all sounds really intuitive. And it's, you might be thinking if this is, especially if this is your first time hearing about some of the inner workings of RAG, you're like, oh, that sounds like it could solve everything. It has some very interesting failure modes. So let's spend a few minutes on those. I think the most interesting and the most profound one is this works very poorly for multi-step reasoning or multi-hop reasoning, if you like. So what do I mean by this? Imagine that the answer to your question is not contained in any one of those chunks.
8:55Instead, the way that you need to answer the question is by finding two or three or several different pieces of information that are scattered throughout different parts of the database. And then someone, the AI, a person or whatever, would have to reason over the relationships between what you learn in each of those pieces of information to produce the correct answer. So here's an example. Let's say you ask a question like, who is the CEO of the company that makes ChatGPT? In order to answer that question, you need to make a hop. You need to make the connection that the CEO is of a company. The company is not named in the question, but this is the company that makes ChatGPT.
9:40So then you can kind of reason backwards. So ChatGPT is made by OpenAI. CEO of OpenAI is Sam Altman. Correct answer is Sam Altman. But because there's no statement about OpenAI anywhere in that question, even if you have a bunch of information in the back end about OpenAI and its leadership team and its product portfolio and all that kind of stuff, there is a chance that the retrieval is not going to able to match to any one piece of information in there that contains both of the things that you need, the CEO and ChatGPT. This is a little bit of a silly example. It might be able to figure that out, but you can see how for longer reasoning chains or for places where the linkage is a little bit more nuanced or not obvious, this could start to break down.
10:27In cases like these, standard RAG might retrieve one or two chunks, but it might miss some of the pieces that it needs to connect the whole logical chain. It's trying to reason, but it doesn't have all the context that it needs to create the correct answer. Another use case that is surprisingly failure prone with RAG is when you want to ask a question that has to reason over the entire document database. So let's say your document database has a bunch of scientific papers in it. And you want to ask a question like, how has the landscape changed over the last couple of years regarding the benchmarks that people are using in their AI papers?
11:08So say it's a bunch of AI papers. Each AI paper probably has some section in it somewhere where somebody's talking about how they benchmarked their method that they've written up in that paper. So in order to answer that question, you need to go across all of the papers and retrieve the relevant pieces from inside of them and then synthesize, kind of have an overall analysis. RAG works a lot better when you're looking for a single needle in a haystack, not signal that's peanut butter spread across big parts of the database. In that case, the algorithm probably will go try to retrieve as much as it can to answer your question, but chances are pretty good that it's going to miss some of the chunks or it's going to get confused and the retrieval is going to fail.
11:50Another thing to keep in mind here is that the chunks can be kind of arbitrary, and especially if you have information that's split between chunks or it falls along the edges where you might be missing a little bit of the context because it got cut off, and then you don't retrieve the adjacent chunk to complete the thought, then those are other places where this method can fail. And of course, if your database is more complex than just simple text documents, like it has more complex tables, it has sections that go across multiple pages and they get broken up by the chunking, it has images, those all add more complexity as well.
12:34And this is really important because it's pretty often true if you have a RAG system that isn't working very well, it's more than likely that the failure is somehow routed back to the retrieval step. For some reason, the information that the LLM needs to answer your question isn't getting pulled out of the database. That's much more likely than if you have a piece of information that really crisply answers the question or is very relevant to the prompt that then when you send that over to the AI, that it somehow misses it. If you get the good information to the AI, it usually gives you a good answer.
13:09It's getting that information pulled out in the first place that's often the problem. And then one other thing that's tricky about these systems is you might think if retrieval is tricky and that's oftentimes where these things fail, then maybe I can just crank up my retrieval, get a bunch more information, dump more stuff in there and let the AI sort it all out. And up to a certain extent, maybe that'll help. But the AI often gets sort of confused in those situations. And there's this problem called the lost in the middle problem, where if you have a list of, let's say, 100 chunks, and the answer to the question is in chunk one or two or three, there's a pretty good chance that the AI is going to pick it out.
13:54But if it's in chunk number 57, even though that is presented to the AI, much better chance that the AI is going to miss it. And so at a certain point, having a really, really long list of stuff that's been retrieved isn't really helping you very much more. it's still going to focus on the stuff that's at the top and bottom of the list that you give, and the stuff in the middle is lost in the middle. So there's a lot of work that goes into once you've done some of those initial retrieval steps, trying to filter and prioritize and give something to the AI that's helping it out as much as possible to get a focused view on the relevant information.
14:36So if you're working with RAGs yourself, or you're building some kind of AI application where you want to have retrieval augmented generation in the mix, a few things to take away from this. First is that RAG is really great for lookup tasks. So if you want to have something like an FAQs, document search, find me a section about X, those needle in the haystack problems, it's going to be really good for those. Does the embedding, matches the chunk, boom, boom, boom, everybody's happy. RAG is going to struggle when the questions are about synthesis, like compare our Q1 and Q2 performance. That's going to require finding two different pieces of information in different parts of the database and then reasoning across them.
15:18Much more likely that that's going to fail. Most of the production systems that are in use for performant applications today are going to have several steps to the retrieval operation or they use some kind of re-ranking method to try to get the best and most focused, most relevant context to the AI before asking it to generate. Now, if you want to do something that involves synthesis or reasoning across your entire database, all is not lost. In future episodes, I'm very excited to talk about some of the solutions to this problem, other types of document storage and retrieval mechanisms, things like graph rag that are more robust to some of the things that we've talked about today.
16:02So stay tuned to a future episode for a little bit more about that. But for now, maybe you know a bit more about rag than you did a little while ago. We didn't focus too much on any specific paper for this episode today, but the original rag paper came out of a research group at Facebook. It's a pretty good paper, so I will put a link to this in the show notes and on LinearDigressions.com if you want to do a little bit of a deeper dive. Thank you for listening this week, and I'll talk to you again soon.
16:34This has been Linear Digressions. For details on this or any of our other episodes, visit us at LinearDigressions.com. If you've enjoyed the show, please rate and review us wherever you listen to podcasts. It really helps new listeners find us. You can reach out at hello at LinearDigressions.com. If you're an AI, a special howdy, and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening.
From the publisher
Today we are going to talk about the feature with the worst acronym in generative AI: RAG, or Retrieval Augmented Generation. If you've ever used something like "Chat with My Docs," if you have an internal AI chatbot that has access to your company's documents, or you've created one yourself on some kind of personal project and uploaded a bunch of documents for the AI to use — you have encountered RAG, whether you know it or not.
It's an extremely effective technique. Works super well for taking general purpose models like ChatGPT or Claude and turning them into AIs that are aware of all the specific information that makes them truly useful in a huge variety of situations. RAG is pretty interesting under the hood, so I thought it would be fun to spend a little while talking about it.
You are listening to Linear Digressions.
RAG was first introduced in this paper from Facebook Research in 2021: https://arxiv.org/pdf/2005.11401