905: Why RAG Makes LLMs Less Safe (And How to Fix It), with Bloomberg’s Dr. Sebastian Gehrmann

15 Jul 2025 · 58 min · 22 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Retrieval-augmented generation (RAG) is widely believed to make LLMs safer by grounding outputs in trusted documents, but Bloomberg’s Dr. Sebastian Gehrmann argues it can instead make LLMs less safe by breaking built-in “helpful/honest/harmless” defenses and creating new, hard-to-enumerate attack surfaces—especially when long context windows and enterprise data are involved.

Guest backgrounds

Dr. Sebastian Gehrmann is head of Responsible AI at Bloomberg (Philadelphia call-in). Previously head of NLP at Bloomberg, leading language technology adoption for Bloomberg Terminal products. Before Bloomberg, senior researcher at Google working on LLMs including Bloom and PaLM. PhD in computer science from Harvard.

Key claims

RAG can circumvent safety mechanisms; “helpfulness” (task usefulness) is distinct from “harmlessness” (abuse resistance); longer context can increase likelihood of safety guardrails being forgotten; off-the-shelf guardrail classifiers may fail in domain-specific risks.

Notable examples

“Insider trading” style unsafe queries paired with innocuous Wikipedia documents can still yield unsafe responses; financial-services risk taxonomy with 12 categories (e.g., financial misconduct, unsolicited advice, personal data leakage, defamation/fake narratives); guardrails like LlamaGuard/ShieldGemma/Aegis failing in Bloomberg’s domain.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Overview of the Episode's Focus

0:31 to 2:10

Discussion on the episode's focus, including RAG and LLM safety.

“Welcome back to the Super Data Science Podcast.”

Guest Introduction and Background

2:10 to 2:44

Introduction of Dr. Sebastian Gehrmann's background and expertise.

“All right, you ready for this exceptional episode?”

The Shock of RAG's Impact on Safety

2:44 to 3:22

Exploring the surprising finding that RAG makes LLMs less safe than expected.

“We'll test the air conditioning systems for the first time this year.”

Understanding Retrieval Augmented Generation

3:22 to 6:10

An explanation of RAG and its application in various contexts.

“understanding is that it was supposed to be exactly the opposite.”

The Importance of Grounding in AI

6:10 to 7:22

Discussion on how RAG helps ground AI models in trusted information.

“So you might be a law firm with millions of contracts that your firm has saved from over the years.”

Risks Associated with RAG and Attack Surfaces

7:22 to 13:20

Exploring attack surfaces created by RAG and the importance of context in AI applications.

“I think just to put some more color on this, if we're talking about large language model, very much this is a technology where at one point in time, You train a model.”

User Responsibility in AI Applications

13:20 to 14:01

Emphasizing the need for users to understand risks when using LLMs.

“So does that end up being the case that then you end up having this kind of list of risks and mitigations for your specific use case?”

Understanding Risks in LLMs

14:01 to 15:05

Explore the risks associated with data and user queries in LLMs.

“They understand how data is being retrieved and how it's all being put together.”

Hallucinations vs. Safety in RAG

15:06 to 18:28

Discuss how hallucinations relate to safety in retrieval-augmented generation.

“I can look at, okay, what are users typically doing?”

Mitigating Risks in RAG Systems

18:29 to 21:00

Learn about strategies to mitigate risks when implementing RAG.

“Can a question answering system that helps financial analysts answer question and test hypotheses?”
Show all 22 chapters

Evaluating RAG Systems Effectively

21:01 to 23:27

Understand how to evaluate RAG systems by context and safety.

“I think that's a great opportunity to talk about a little bit about how should we evaluate systems.”

Context Length and Safety Challenges

23:28 to 28:00

Examine how context length affects the safety and performance of LLMs.

“Can we have a classifier on inputs and on outputs that identify violations of our rules that we set up ourselves?”

Context Length and Retrieval Challenges

28:00 to 29:31

Explore how context length affects the performance of language models in retrieval tasks.

“Being able to handle longer context also allows you to give much more contextualized answers.”

Challenges of Needle in a Haystack Tests

29:31 to 31:26

Discuss the implications of various context lengths and the effectiveness of testing methods for LLMs.

“And you explained all of that very clearly.”

Trade-offs in Context Window Expansion

31:26 to 33:39

Understand the trade-offs in expanding context windows for language models and their effects on latency and costs.

“Yeah, there are additional considerations to address your needle and haystack point.”

Optimizing RAG System Efficiency

34:24 to 39:26

Dive into how RAG systems work and the techniques to enhance their efficiency in retrieving documents.

“And over a long enough timescale, over many years, maybe many decades, compute costs over millions of tokens might be trivial.”

Model Size and Safety in RAG Contexts

39:26 to 42:00

Examine the relationship between model size, safety, and the implications for RAG applications.

“Speaking of snappy and fast, we've talked now about context window length.”

Understanding Refusal in Model Responses

42:00 to 43:50

Explore how refusal to answer questions impacts model safety and reliability.

“To what extent is refusal to answer a, you know, how does that factor into these kinds of assessments?”

Risks of Generative AI in Financial Services

43:50 to 46:05

Discuss limitations of LLMs in finance and the need for domain-specific safeguards.

“There's a second paper that you also recently published.”

Best Practices for Selecting LLMs

48:01 to 51:04

Get practical recommendations for selecting LLMs in regulated domains.

“If we're trying to select an LLM for a particular use case, what do you recommend we do?”

Book Recommendation: The Unaccountability Machine

51:04 to 52:08

A recommendation of a book discussing accountability in organizations.

“That's a great soundbite at the end there.”

Following Dr. Sebastian Gehrmann

52:08 to 53:36

Learn how to keep up with Dr. Gehrmann's work and insights.

“to know if something gets blocked, but you really need the answer.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:This is episode number 905 with Dr. Sebastian Gehrmann, head of responsible AI at Bloomberg.

0:12Jon Krohn:Welcome to the Super Data Science Podcast, the most listened to podcast in the data science industry. Each week, we bring you fun and inspiring people and ideas exploring the cutting edge of machine learning, AI, and related technologies that are transforming our world for the better. I'm your host, John Krohn. Thanks for joining me today. And now, let's make the complex simple.

0:46Jon Krohn:Welcome back to the Super Data Science Podcast. Today, our guest is Dr. Sebastian German, an exceptionally gifted individual at Thoroughly and Clearly Explaining Cutting-Edge AI Research. research. Sebastian is head of Responsible AI at Bloomberg, the New York-based financial, software, data, and media company that, with 20 ,000 employees, is huge. Previously, as head of NLP at Bloomberg, he directed the development and adoption of language technology to bring the best AI-enhanced products to the Bloomberg terminal. Prior to Bloomberg, he was a senior researcher at Google, where he worked on the development of LLMs, including the groundbreaking Bloom and Palm Models.

1:24Jon Krohn:He holds a PhD in computer science from Harvard University. Today's episode skews slightly towards our more technical listeners like data scientists, AI engineers, and software developers, but anyone who'd like to stay up to date on the latest AI research may want to give it a listen. In today's episode, Sebastian details the shocking discovery that retrieval augmented generation, RAG, actually makes LLMs less safe despite the popular perception of the opposite, why the difference between helpful and harmless AI matters more than you think, the hidden attack surfaces that emerge when you combine RAG with enterprise data, the problems that can happen when you push LLMs beyond their intended context window, and what you can do to ensure your LLMs are helpful, honest, and harmless for your particular use cases.

2:10Jon Krohn:All right, you ready for this exceptional episode? Let's go.

2:20Jon Krohn:Sebastian, welcome to the Super Data Science Podcast. It's great to have you on the show. Where are you calling in from today? Thank you so much for having me. I'm calling in from Philadelphia. Nice. It's going to be a scorcher here in New York coming up at the time of recording, heading up to 100 degrees Fahrenheit. I hope you're able to stay cool down there. Yeah, not too much different here. We all got the heat warning yesterday evening, and it's going to be a rough couple of days. Yeah, yeah, yeah. We'll test the air conditioning systems for the first time this year. For sure. All right. So you are the head of Responsible AI at Bloomberg, a hugely well-known financial software, data and media company.

2:57Jon Krohn:Long before you started at Bloomberg, you were researching at the intersection of natural language generation and responsible AI solutions that are trustworthy, transparent, reliable. And now you've kind of brought all those things together with your latest paper, which, of course, we'll link to in the show notes. It's called RAG LLMs Are Not Safer. It's a safety analysis of retrieval augmented generation for large language models. And this finds counterintuitively, and this is the main reason why I wanted to have you on the show, is because it blew my mind when I discovered that RAG actually makes LLMs less safe and their outputs less reliable because my understanding is that it was supposed to be exactly the opposite.

3:40Yeah, absolutely. Yeah, I should mention this was a paper that we had an intern work on last year who has been working with colleagues from our CTO's office and AI engineering organization. and this was part of our broader responsible AI research theme that we have going where we want to make sure that if our clients or our customers use our AI that if they use it for something that they shouldn't that we can identify this we can block this we can monitor this over time which is incredibly important especially for a such heavily regulated industry such as the one that we are operating in. There are so many rules that apply to our clients.

4:21They really want to make sure that people can't abuse purposefully or completely accidentally our technology. So as part of that research direction, we were interested in the safety of RAG because in the end, RAG is such a ubiquitous technology and it is absolutely necessary to ground responses and trusted sources of data. That is, there's no way around that if you want to answer questions that are grounded in this broad and really challenging area that we operate in where there's hundreds of billions of pieces of data coming into our systems every single day. The only way to do this is if you use RAG or some kind of other similar retrieval augmented technology or document grounded technology.

5:09So what our paper did was it coupled unsafe queries. So think of the worst thing that you might want to ask a large language model, right? How do I do insider trading? And we coupled that with RAG. So we retrieved completely innocuous documents from Wikipedia. And while most large language models that we looked at didn't respond to the original question, when coupled with these completely harmless documents from Wikipedia, The response was often then unsafe, which is why we gave the paper this title that was very, very strong, because everyone has been talking about how RAG makes things more safe.

5:54They are more grounded. They are grounded in actual factual information rather than using the large language model to kind of make up things. And what we found is that actually RAG can circumvent these built-in safety mechanisms that large language model providers put into these models.

6:10Jon Krohn:I got you. I got you. Now I understand. And I guess we should also, it occurs to me as we're talking about RAG and probably a lot of listeners out there are aware of what retrieval augmented generation is, but maybe we should spend just a couple of minutes explaining it as well. I don't know. You can probably do a better job with this, but at a high level, it's, it's, you already kind of gave an example there where you can be using documents from Wikipedia, the public internet, or it could be a common use case with RAG is to have lots of documents, internal documents. So you might be a law firm with millions of contracts that your firm has saved from over the years.

6:48Jon Krohn:And you could put all of those millions of contracts into a RAG type of system. And you can search, you can use natural language to query over all of those millions of documents. The most relevant ones will be pulled out from the millions. And then you can typically fit that small number of documents that's pulled out, maybe it's half a dozen or a dozen, into the working memory of a given large language model. And then you can have a natural language response come out as a result of that. What do you think about my RAG explanation there? Absolutely correct. I think just to put some more color on this, if we're talking about large language model, very much this is a technology where at one point in time, You train a model.

7:33And then once you're done training this model, it is stuck. You freeze the model and you say, okay, this is the model. The knowledge cutoff is January 2021. And then if you ask a question that requires knowledge from beyond 2021, the only way that you can get this into the model is by actually providing it in its context. So researchers a couple of years ago, when they developed this RAC technique, four large language models in particular said, well, the most natural thing to do is to couple a search system with the large language model. The large language model is very good at synthesizing, summarizing, extracting information, but this kind of freeze in time is really bad.

8:13To overcome it, we plug in a search tool. And that's why it's called retrieval augmented generation or before RAG was a paper, a lot of search engines already gave answer snippets. And so at that point, it was kind of search enhanced question answering. But they're all kind of this have become this overall term that we that we use for not just this vanilla setup where we have a retrieval step and a generation step. Often we use RAG as kind of the overarching term for anything where you ground a large language model response with some kind of data that you retrieve from somewhere. That could be structured data, unstructured data.

8:53It could be unstructured or structured data from multiple different sources. all of these are kind of under this whole umbrella term of retrieval augmented generation, which is why this has become such a big topic to talk about, because this is one of the predominant ways in which applications today are built.

9:10Jon Krohn:For sure. And I like how you're using this term grounding. I think that that makes a lot of sense. Another term that you've used in a recent podcast, you were emphasizing that reg is essential for grounding GenAI products in actual trusted information. However, you've described REG's architecture as creating unpredictable attack surfaces. So that's kind of the other term that I want to get into here, this idea of attack services. So on the one hand, we need REG for grounding GenAI products in more recent or an actual trusted information, or maybe in confidential information that an enterprise or organization has.

9:46Jon Krohn:But at the same time, this creates these unpredictable attack surfaces. So what the heck are attack surfaces? Yes. So for that, let's kind of explore what we mean by attack surface. Obviously, if we think about something like cybersecurity, attack surfaces are everywhere where you have unsecured endpoints or you have code that you can kind of inject into. You have databases that are just out in the open. Latch language models are very different because they, unless you give them the information, They don't have that kind of attack surface, but there are other attack surfaces or risks, specifically what we call them are typically content risks, where either inputs to the large language model can be asking for a very harmful output or the output itself could lead to something that is harmful or against established regulation, rules, laws.

10:39so this attack surface very much is grounded in the way that you want to use this language model or this application that you're building so take us for example we're a financial services company our content is very much focused on helping people conduct analyses in the financial space synthesize information summarize information ground all of this and attribute all of this to our just-in-time data that we have on our Bloomberg terminal. So for us, the attack service is very much linked to the domain that we operate in. So what we care about are things like, can someone conduct financial crimes with the help of light language models?

11:26Can it facilitate insider trading? Does it give unsolicited financial advice that people might trade on and lose a lot of money? Does it create an information imbalance or an asymmetry where it discloses trading strategies from one company to the other? But we are just one of many, many examples of people who use these light-changuage models. So when we talk about an unanticipated attack surface, the people who build light-changuage models, they can't enumerate every single way in which they are being applied. They're being applied to healthcare, they're being applied to law, They're being applied to education.

11:59They're being applied everywhere around the world now. And on top of that, the jurisdictional differences, the different geographic locations, they might have different laws even applying. So now we have this whole web of applications that are built on top of large language models. But we only have one provider who says, yeah, actually, our model is safe. there's no way where that would be possible to really anticipate every single use case around the world grounded in all the relevant regulation so all this coming to the main point that we're really trying to keep making over and over you need to really evaluate your AI application in the context you want to deploy it in because in the end only the organization developing the application that you integrate the large language model in understands their risks So that is really why we're doing this kind of research.

12:53We want to understand what are the risks that are specific to us and the way that we're building applications and the clients that use our applications. So when we talk about unanticipated risk, this is very much it. We don't understand the risk unless we measure it. The people who build models, especially if we use third-party models, they don't understand our risks. So we need to study this to really understand. And instead of having an unmeasurable, unanticipated risk surface, we want to have one where we very much understand what are the risks and how do we mitigate them.

13:23Jon Krohn:So does that end up being the case that then you end up having this kind of list of risks and mitigations for your specific use case? So it sounds like basically it's the user of a reg setup or the user of an LLM that has to be cautious in their particular circumstance, given that the LLM creator couldn't have anticipated, like you said, all the possible use cases. Is that right? Slightly. So what I mean really is the people who build the RAC system, they usually know what they're building it for. They understand the data that goes into the database, or they should at least. They understand how data is being retrieved and how it's all being put together.

14:06And then there are a couple of sources of risks. Number one, you have the data. If there is dangerous information in the data and write an instruction on how to shelter money in places that money shouldn't be sheltered, and you ask the large language model, hey, give me information, and it adds that information in the context, that could be one vector of getting harmful information out of the language model. But there is another vector in which the large language model itself could be unsafe. It could be unsolicited, give that advice. But then the user is another attack vector. What if the user types that in?

14:43How do I shelter money wherever? How do I commit tax fraud? But in this case, we can rely on the large language model to kind of block these kind of queries, which really puts a lot of trust in this developer of the large language model that often there is a different entity from the organization that builds the RAC system. Or as the builder of the RAC system, I can take a look at those three. I can look at, okay, what are users typically doing? What are attack vectors that users might be trying with our system? What data is in the databases? How do I make sure that I understand this data and safeguard us from exposing information to users that shouldn't have access to this information?

15:26And how safe is the large language model really for our use cases? Our paper is very much in that third realm here because we wanted to know the built-in defenses of the language model. How good are they actually? And how will they stand up to a RAC setup? And in our case, we found that for the general purpose, dangerous queries that we looked at, actually, RAC safety breaks down. Because large language models really are only secured for non-RAC setups. Which means that we need to build guardrails around our applications that go beyond what the large language model builders are actually providing to us.

16:08Jon Krohn:Nice. Okay, yeah, that all makes sense to me. So a question that occurs to me as you're talking about this. So, you know, when I originally heard about this paper and about LLMs being less safe in rag situations, the thing that popped into my head, I guess, is I was thinking about how it's my understanding is that hallucinations are much less likely in a rag circumstance than outside of a rag circumstance for an LLM. So when you have that, you know, to use that word grounding again that you've been using, when you have that grounding, it seems to lead to fewer hallucinations. So I guess in my mind, maybe I kind of and so I'd also I'd welcome your input on the hallucination point if that is actually true.

16:52Jon Krohn:But I guess I kind of ended up conflating those two things in my head and thinking, OK, if it's hallucinating less often, then it's surely safer. But yeah, and again, you can cut into my thoughts on the hallucination in a second. But now it's becoming clear to me. It's becoming clear for me to understand how maybe even if hallucinations are reduced, the issue here is that the kinds of safeguards that an LLM creator put in, say, meta put into some LLAMA models that they release, those safeguards that are built in, in a reg setup, often break down. Exactly what you're saying. That is absolutely the case where we say, look, there are typically the three H's.

17:39They were first developed by Anthropic. Many companies are now adopting them. The application you're building, it should be helpful, it should be honest, and it should be harmless. and hallucination very much goes goes into this honesty bucket which often is also then combined with the helpful because how can you be helpful if you're not honest so that if we if we just focus on being helpful and being being harmless everything that goes into hallucination and the advantage that rag brings to to helpfulness they're they're vast and they make a rag a necessity that's why we're not saying that rag is dangerous we're just saying it is not necessarily safer It is absolutely necessary, but the harmlessness angle is something that is completely separate.

18:27And the way that organizations think about it is often to kind of split them and look at them through two different angles, because the helpfulness angle, it's very much grounded in a specific application. Can a question answering system that helps financial analysts answer question and test hypotheses? The answer to that is, does it help them? Yes or no? but there's always this angle of a malicious or an unintended abuse of that system where the same system that helps people assess hypotheses could then also be used to, right? Who is the worst broker and who should I exploit? Which is maybe not something that you want to answer in turn.

19:09So now we have these two angles and often you can say, okay, the harmlessness angle is something that is usually consistent across an industry, across a sector, across a domain, that there might be some application-specific risks to. And then there's the helpfulness angle where you provide answers that hopefully help. And here, RAG really helps because it leads to you being able to build things like transparent attribution. So transparent attribution for us means every time I produce some kind of tidbit, some piece of information, I need to ground this in some kind of document or some kind of structured data.

19:48If I say the price of meta is so-and-so today, then I want to be able to look at that and say, where does that number actually come from? Is it hallucinated? Or do I actually know the query that produced that number? If I'm saying the following analyst said the following statement, then I want to be able to hover over this and say, oh yeah, this is where that statement came from it is not hallucinated. This to some degree can prevent the harmfulness or the harmlessness issues as well, but it is a somewhat separate topic. As you were saying, you talk about hallucination and that's where the big advantage for RAG comes in, but then there's the harmless or the harmfulness angle as well.

20:28Jon Krohn:Okay. Now, so maybe recap for us again. So there's probably lots of listeners out there who are sold on RAG then they're like, great, I want to reduce harmfulness. I want to increase the honesty of my LLM. And so I'm going to use a RAG system. What are the kinds of mitigations? I mean, you went into this a little bit, but kind of recap for us again, the kinds of mitigations that our listeners can take home away from this episode to be able to use RAG so that it is safe for their particular circumstances. Yeah. I think that's a great opportunity to talk about a little bit about how should we evaluate systems.

21:08Because in the end, what you see talked about a lot publicly and probably the most is benchmarks. And you see, oh, hey, this new model that large language model provider A or B or C produced, it achieves a better score on the large language model arena or on the following benchmarks. And it's better at reasoning. It's better at code. That doesn't necessarily mean that it's better for all the downstream applications that are being integrated into. And I think we often conflate this kind of view where it's actually really, really important to measure and to evaluate the system in the context it's deployed in.

21:44And that includes things like safety testing and specific guardrails. So we also released a second paper in addition to our RAC LLN paper where we developed our own content risk taxonomy for financial services, where we say here are 12 categories of risks that we really, really should address with applications in our area. And for each of these, we can measure because we can collect data against this and say, okay, here are 100 queries that try and do financial misconduct, 100 queries that try and make the LOM generate financial advice, 100 queries that try and get at personal information, 100 queries that try and defame someone or create fake information or fake narratives.

22:32In all of these categories, you can create data, you can measure it against this, and you can not only test the large language models themselves, but rather you can test the entire end-to-end system that is deployed in a specific socio-technical context. And here, typical guardrail systems, there are a lot of open source solutions. NVIDIA has their own, Lama has their own, Google has their own. They provide open source guardrails. But again, these open source guardrails, they're shielding against these general purpose risks. In our paper, we found that if you apply these LamaGuard or ShieldGemma or Aegis is what they're called, if you apply them to then to our specific risks, they're classifiers to say, is this input or is this output safe?

23:16Yes or no. And they also fail in our domain because similar to our Rack paper, it's just not a use case that people necessarily have thought about before. But it gives us then the idea of, okay, how do we build our own guardrails? Can we have a classifier on inputs and on outputs that identify violations of our rules that we set up ourselves? So now, instead of having a vanilla rack system where it's retrieval answer, we have guard rail retrieval answer guard rail. And in practice, to prevent hallucination and to add attribution, real systems that are deployed to wide ranges of audiences, they have many more components.

23:55And it's really this end-to-end application that should be evaluated. and where you need the subject matter expertise to also know, is it helpful? And is it harmless?

24:04Jon Krohn:Nice, okay. So yeah, so a combination of subject matter expertise, people digging into their particular circumstances. You mentioned earlier in the episode, being aware of your data, making sure that there aren't data in there that are going to be harmful, that could be used as grounding by the LLM. And then you mentioned just now this idea, this flow of guardrail retrieval answer, and then another set of guardrails. And so, yeah, so very practical advice there.

24:36Jon Krohn:Changing the topic a little bit, like still staying on reg and still staying on helpfulness, but something that we haven't talked about yet is context length. So I have a couple of points here, a couple of questions here. You observed in your paper that LLMs are often optimized for short prompts, but deployed in long text environments like reg. So going back to the example that I gave earlier, I talked about there being a million legal documents that the RIG system searches over. And then it pulls out a dozen documents. Those documents could each be 10 pages long. So then you're talking about 120 pages of context.

Read the full transcript

25:18Jon Krohn:And if the LLM was optimized for a question like, what's the capital of France? Then, yeah, I could imagine you run into issues. So do you want to fill us more in about this and the kinds of issues that arise? For example, it seems like there's tradeoffs between the benefits of longer inputs versus this becoming a new risk surface, a new tax surface. Yeah, absolutely. And I think you're really hitting the nail on the head here. RAG is incredibly powerful. There's a lot of investment from companies that build large language models into increasing the possible context length. At the same time, increasing the possible context length also requires developing methods and developing the models to actually be able to handle such a context length rather than just being able to technically being able to handle them by having an attention that goes long enough back.

26:14So here, there are a couple of considerations. In our paper, what we evaluated was how does context length really influence the safety angle? And we found that, especially for safety, there seems to be this effect where the more context you put in, the more likely the model is to forget the built-in safety guardrails or this alignment that people talk about, which absolutely shouldn't happen if you're considering how RAG is set up because again you're adding innocuous completely harmless information and just because there's a long context doesn't mean that the language model needs to behave any differently from if you just post in what's the capital of France.

26:55At the same time the context length question actually has also massive implications of how we build RAG systems. In practice retrieval systems have multiple components too. We've been kind of glossing over this point where you just say, oh yeah, there's a search system and there's a database. But often you have things like, okay, how do I parse the query? If I ask a question, to what timeframe should I limit the search results? Am I filtering to particular industries, sectors, companies, any kind of other metadata? There's usually the way that search engines are written, there's usually multiple steps as well, where you do a first pass retrieval, where you go from hundreds of millions of documents or even billions of documents down to just a handful.

27:39And then there's usually a re-ranking step that's much more computationally intensive, where you really pick out which snippets in the documents are you actually trying the answer in. So all of these components here have an influence on the context length. When you say you have 10 legal documents of 10 pages, how do you get to them? And what is the effect of pasting entire documents versus just a paragraph or two from each document? and typically what people find is the less context you need to provide the especially if the answer is in that context the easier it is for the large language model to find the right answer seems pretty obvious but that then becomes a retrieval problem how do I find within the you know if I have 100 million documents each 10 pages long how do I find the two paragraphs that actually answer the question so a hope really has been in large language model development to just increase the context length and to rely less on more and more accurate retrieval, but rather let the language model figure it out.

28:36Being able to handle longer context also allows you to give much more contextualized answers. If you have the entire 10-page document, even if the answer is found in just one paragraph, it can still give you the context from page one or tell you what is this document type? Where is it from? What was in the intro? What did the executive summary of this document say in contrast to the actual document. If we go into voice, like what was intonation? If we go into video, there's so many opportunities to take advantage of longer contexts. But again, we have to really consider how are the models being deployed?

29:11How are they, who are the users of this? What is the application? And every single one of these design decisions I just talked about can influence the helpfulness and the harmful or harmlessness in this case of the entire system. So this is a massive undertaking and requires a lot of research.

29:28Jon Krohn:That was a fascinating answer. I learned a lot. And you explained all of that very clearly. Thank you. Something that you talked about there was how longer context windows allow us to have more context in our answer, more subtlety, more nuance in the answer that comes back. And you also mentioned there how context windows are getting longer and longer, which actually ties in perfectly to the next question that I was going to ask you, which is that we're at a point now where it's getting reasonably common to see LLMs that have a million token context length. And we've seen some release that are multiples of that, many millions of tokens in context.

30:10Jon Krohn:And typically when these are released, it comes with these kind of needle in a haystack tests where you try to hide a small amount of information. Like I think a common one is like a pizza recipe. Like there'll be something like the world's best pizza is, and then these like, you know, random ingredients like anchovies, um, you know, something that, that is, that is unique. And then you'll have, you'll say, okay, you know, over our 10 million token context length, um, the, the model was able to successfully retrieve pizza information. Now I've got, I'm getting a little bit deep in the weeds here and on a little bit of a tangent, but some people have also said, you know, that isn't a great test because if you have millions of legal documents and then there's one pizza recipe, that's quite unusual.

30:57Jon Krohn:And so it's probably, you know, that's then maybe something that the LLM is going to take notice of. And so, yeah, so there's controversy about needle in a haystack test, but we don't necessarily need to get into that too much unless it interests you. The question that I'm getting to is, do you think we'll get to a point where context windows expand so much that it is effectively like an infinite context window? And then that means that we don't need reg at all. Yeah, there are additional considerations to address your needle and haystack point. I think this is a perfect example of the difference between developers of large language models and integrators of such language models in actual applications.

31:42As someone who might be considering, okay, which of these long context models do I use? I can look at Needle and Haste. I can say, oh yeah, this model got it 100 % of the time. This one did not. Clearly, I'm going to look at the one that gets 100 % first. And that is really the decision that these benchmarks, even if they're artificial benchmarks, can influence. But I think no one is making the point that just because a model can find some information in 10 million tokens, it is going to be able to help you with a research-heavy legal task. That is really on the integrators in the end to identify.

32:18And in this case, it might be a much, much, much harder task. If you have a system that we just released earlier this week at Bloomberg was a tool that helps research analysts search through over 400 million documents and news articles and analyst reports to answer questions and to help with the identification of and solution of hypotheses. It's very much hypothesis driven. You ask a question, it goes through all these 400 million documents and then synthesizes an answer. This is much, much harder to do than a simple find this piece of pizza information or pizza recipe information. And instead, you really need to, again, evaluate in the context of the deploy system.

33:01But to then go back at your question, there is obviously strong advantages of models that are technically capable of handling more and more complex situations. If I'm able to just paste in more documents or more of a context, I don't need to rely on as many tricks to really narrow down the context window. I can just rely on a large language model. There's a trade-off here though, where longer context usually comes at a cost of significantly increased latency in cost. So even though long-context models are available, they might not necessarily be the best for the task if you want a snappy, direct answer.

33:39Jon Krohn:This episode of Super Data Science is brought to you by the Dell AI Factory with NVIDIA, two trusted technology leaders united to deliver a comprehensive and secure AI solution customizable for any business. With a portfolio of products, solutions, and services tailored for AI workloads from desktop to data center to cloud, the Dell AI Factory with NVIDIA paves the way for AI to work seamlessly for you. Integrated Dell and NVIDIA capabilities accelerate your AI-powered use cases integrate your data and workflows, and enable you to design your own AI journey for repeatable, scalable outcomes. Visit www.dell.com slash superdatascience to learn more.

34:20Jon Krohn:That's dell.com slash superdatascience. Of course, that is such an obvious point to make. And over a long enough timescale, over many years, maybe many decades, compute costs over millions of tokens might be trivial. but at least for the foreseeable future, it isn't. And so, yes, that makes perfect sense. So something that I guess we could make a little bit more explicit for people who aren't familiar with RAG because we haven't talked about this is that when the way that RAG systems work is, like let's go back to that example of the million legal documents, what we would do in advance before running any RAG queries is we would map each of those million documents into a high dimensional space.

35:07Jon Krohn:called a vector space and the location in that high dimensional space, like you could, you can, you can only visualize in three dimensions. So, so in your head, you can kind of imagine on like an X, Y, Z, an X, Y, Z plane, um, you know, in three dimensions, you can kind of imagine, okay, you know, in the top right corner near the front of this space, we have, you know, commercial law documents. And then, um, nearby there, there are some, you know, other kinds of related legal documents. And as you move further and further away from a given point in space, you'll get more variety in the kind of document that you're looking at.

35:48Jon Krohn:So the closer that things are in this space, the more overlap and meaning there is between the documents. And so this allows us then in real time to take the user's natural language query, map it into that same high dimensional space, but you can imagine it's three dimensions. When I say high dimensional, I mean, hundreds or thousands of dimensions, which you can't visualize, but for which the linear algebra is basically identical for a computer relative to a two or three dimensional space that you can visualize. And so, yeah, so we take the natural language query that a user makes to the RAG system.

36:23Jon Krohn:we can convert that into the same high dimensional space, find its location, and then retrieve the documents like you talked about, kind of a cheap, fast first retrieval step, which could be something like I'm just describing where we take the closest documents to wherever the query gets mapped to in the high dimensional space. And then we can do more complex processing after that, but that kind of gives us our initial results. And so doing that is very, very fast. You know, we only need to convert one query, which might be short, into a coordinate in a high dimensional space. And then we could use a very fast mathematical operation, like a cosine similarity score, to find the closest documents in that space.

37:06Jon Krohn:And that's all computationally very inexpensive, very fast. It allows the rag system to work in real time, even over, like you said, billions of documents. and yeah in contrast if all of that text if our millions or billions of documents were in the context of an llm even if it all fits in you'd then have instead of this computationally simple calculation this fast calculation you would have to have tons and tons of really high-end gpus running to comb across all of the meaning in that huge context window. So yeah, correct me if I... No, absolutely. And what you described here, commonly known also as semantic search, because you can really search based on the meaning of a query.

37:58To add to your point, often, even for commercial systems, it is still the case that you rely on keyword retrieval, just because it's even cheaper, it's even faster. There are techniques from the early 90s or even late 80s that are still around just because they're so computationally efficient because you really want the retrieval to be as cheap as possible. And often you use semantic search for the more toned down. You do a first pass keyword retrieval. You do a second pass semantic search within the keyword retrieved steps. So there's a lot of engineering over the years that has been developed to just get this retrieval step as efficient as possible.

38:35and we're still a while away from large language models being able to do anything even remotely as efficient as this step. And as you said, this can lead to a massive use of GPU power for something that you can't solve otherwise. And again, grounding this in the individual end-user experience, it could be that in the future, if someone does a side-by-side comparison, I do this really expensive process where I just pipe everything into a large language model, I do a cheap process, And it could be that in terms of helpfulness, users actually prefer the one that's snappy and fast rather than the one that's maybe five points more accurate in the end.

39:15This is all something we need to evaluate in the end. And those are all design decisions that are all being evaluated all around the globe right now as people are building their own Gen.AI and Rack solutions.

39:26Jon Krohn:Speaking of snappy and fast, we've talked now about context window length. How about model size? That's something that we haven't talked about yet. So I know that you investigated in your paper differences between small models and larger ones in RAG contexts. What did you find? There are differences. And generally what we found is if a model is safer from the get-go, even without RAG, it tends to be more robust to adding RAG. A large factor of this is the model size or the model capability in general. I think at this point, model sizes are a little bit of a misnomer because we have so many models that rely on a mixture of experts and that have architectural advantages that even though on paper they have more parameters, they actually are using fewer of them when you actually run them live.

40:17So it's very hard nowadays to actually compare the parameters. And we often compare based on active parameters or there might be ways in which models are compressed, which again changes the representational power. But generally speaking, to answer your question is, yeah, we found that models from the get-go safer. They are also harder to break through RAG, although we found that basically every system was breakable, regardless of whether small or large. And I think we specifically call out LAMA for being relatively safer than many others, both LAMA 70B and 7B. but LAMA 70B I believe was a little bit better than 7B even.

41:02Although again, it can change from time to time because what we found really was the guardrails were broken because of this increased context length. It could be that once the next generation of this model comes out that this is being prevented, that there's an active component of the post-training, of the alignment step that looks at how can someone use this with a longer context and our exact setup could be one of those test cases where you can just continue training the model on and we'll just inherently protect against this particular angle of attack. Nicely said.

41:37Jon Krohn:I've learned a ton from you in this episode so far. We still have a little bit to go. So I'm excited for that. I'm curious, what's the effect of refusing to answer? in these. So it sounds like it's clear that bigger models are generally better. You know, they're more capable. They fare better in red contexts, generally speaking. To what extent is refusal to answer a, you know, how does that factor into these kinds of assessments? Like in your paper, you called out Gemma 7b in particular for showing low unsafe outputs, but largely because it refused to answer questions. Yeah, so there are a couple of different considerations here.

42:23If you just refuse to answer, it could be because you don't know, or it could be because you actively find the input to be unsafe. And if you can't distinguish between them, it's very hard to know whether, well, your model is just bad or whether it's unsafe or safe in this case. so model sizes and model capabilities again they're all so intricately linked where you want to build a system that is helpful so you always need to pair an analysis like ours with one that actually evaluates the how helpful the model is and if in the end you call jama here if in the end jama is also refusing to answer completely safe questions and it's completely safe and correct rack setup, it's not going to be helpful.

43:12So even though it is harmless, it still would not be able to be used. So that's, I think, just highlighting the need for having a multifaceted evaluation. You need to consider those. Similarly to how model sizes will also affect latency and cost of running a system. It could be that the fast, cheap, small model is completely up to the task. And in that case, why would I use this completely overblown model to do the same task just because it is performing better on things that are completely not relevant to your particular application?

43:45Jon Krohn:Nice, nice. So changing gears now a fair bit. There's a second paper that you also recently published. So your first author on a paper that was submitted to Archive in April called understanding and mitigating risks of generative AI in financial services. So mostly so far in this episode, we've been talking about generally how models fare under RAG. But in that paper, it's related to risk of gen AI and finance. You emphasize that most foundation models are not trained on finance-specific corpora bodies of knowledge. So what are the limitations this creates for LLMs in general, but particularly for RAG.

44:33And I'm assuming that this same kind

44:37Jon Krohn:of sentiment, you looked at it with finance specifically because Bloomberg is a financial services company largely, but do you think that the same kind of limitation would apply in other sectors as well? Yeah, absolutely. So yeah, I gave a little bit of a teaser of this paper earlier in an answer as well. And what the way that we wrote our paper very much should be seen as a case study in finance here or financial services, in particular capital markets and asset management, is the case study that we use to make the point that we really need to think about risk and risk taxonomies and risk management in our domain, in what we are trying to build.

45:20And as you say, we made the point, yeah, models are not necessarily trained on financial domains. We see that both in the helpfulness and the harmlessness angle. Often, you know, complex financial tasks are not being able to be sufficiently handled by large language models by themselves. But also in our paper, we make the point that even safeguards that are dedicated models or systems to provide these kind of first paths, like is this safe? Is this unsafe judgment? they're also not trained on financial services. And if you use them out of the box and say, look, I use LamaGuard, I use Shield Gemma, I use Aegis, I'm safe now, right?

45:57You're protected against a particular view of safety that is very much grounded in categories that are relevant to broad populations, to things like chatbots that help you do productivity day-to-day tasks. The typical applications that you would see in those AI productivity tools, no matter which one you use, They all have similar mechanisms, but those are not necessarily the same risks that we are under in financial services. Those are not the same obligations that companies, organizations in healthcare are under or law or any other highly domain specific knowledge intensive domain that has a lot of specific regulation, jurisdiction specific regulation, considerations about whether just refusing to answer or giving disclaimers is enough or whether questions should be blocked altogether.

46:46And there's just this difference of view that can be encapsulated in a single model that a provider can give that very much is focused on a different use case.

46:59Jon Krohn:Regular listeners will already be aware that I'm obsessed with Anthropik's Fable 5 model, and it has taken over my working life. I'm writing a technical book that includes LaTeX files, mathematical notation, Python code examples, and Fable 5 and Cloud Code handles requests I make across whole chapters. with accompanying Jupyter Notebooks end-to-end. Work that a few short months ago would have been dozens of separate requests with way more manual fiddling required. With Fable 5, it just works, essentially like magic, first time. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you.

47:35Jon Krohn:Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai slash superdata. That's claude.ai slash superdata. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Claude.ai slash superdata. Nice. Yeah, so I don't know. Do you have guidance for us? If we're trying to select an LLM for a particular use case, what do you recommend we do? I mean, practically, how can we move forward with all the information that you provided in this episode in selecting an LLM for a particular use case, for a particular domain, particularly if we want to be applying it in RAG situations?

48:27Yeah. So in our paper, we also have a list of best practices and recommendations that we have for, especially for knowledge intensive domains and regulation heavy domains. Not necessarily everything has to be followed if you're building something for a much broader general population. but especially for these kinds of domains all i can do is spray my mantra evaluate the system in the context that it's deployed in if you are building something for healthcare well you better evaluate in the context of healthcare if you are building in the context of financial services you better evaluate your subject matter experts in financial services and specifically on the safety angle our paper makes a couple suggestions here there are very good starting points there are taxonomies such as the NIST risk management framework for AI.

49:16There are other industry collaborations ongoing. There's ML Commons. Those all provide more general purpose taxonomies, but just taking them as a starting point and then from there adjusting them to your domain can often save a lot of time. And especially if you're a large organization with a compliance or risk department, it will help them also understand how one can classify and then categorize these kinds of risks. Another recommendation we make is to organize red teaming events or do any other kind of red teaming. Red teaming in this case is this practice that had to start in the Cold War where you have users trying to be malicious.

49:59So we get people in the same room and we say, hey, look, for the next couple of hours, try and break the system. Try and play evil. Here are some instructions on how to do this. And then afterwards, we can look, how often was this actually broken? How often did the system give financial advice? How often did it refuse? And from there, we can quantify the risk surface. Since we were talking earlier about this unknown risk surface, well, just measure it, and then you have it. So that's kind of the main takeaway that we have. We give pretty specific advice for how to go about this and how to set up risk management frameworks.

50:33And all this needs to go hand in hand also with, again, this evaluate in the context that's applied. Make sure you invest a lot in evaluation. Don't just take the word of the large negative model providers that their benchmark scores are going to translate into all the downstream applications. And if you follow that advice, you're going to have a system that is in the end much more trustworthy, reliable, robust, and you're going to have users that are going to keep using it rather than trying it twice, getting really bad answers both times and never touching it again.

51:05Jon Krohn:Perfect. That's a great soundbite at the end there. I'm sure we'll be making it into a YouTube short. So before I let you go, we are pretty much out of time here, but I always ask my guests for a book recommendation before they leave the podcast episode. And I usually give guests a warning, but we rushed into recording and I forgot to tell you. So hopefully you have something on hand in your mind. It doesn't need to be something AI related necessarily. All right. I'm going to give you the recommendation of a book that's currently sitting right here on my table, which I'm reading right now, which is The Unaccountability Machine.

51:39It just came out a couple months ago. It talks about how organizations are failing to build processes that act as accountability things. If you've ever talked to customer service and you couldn't escalate and the rep you've talked to couldn't solve your problem, you screwed. This book is talking about why from a cybernetic perspective, this is a bad design and how to set up organizations and processes that can help us, which is also applicable to AI because you want to know if something gets blocked, but you really need the answer. Where do I go? How do I escalate?

52:12Jon Krohn:That's a great recommendation there, Sebastian. Thank you. And then final question for you is, how should people follow you after this episode? I learned a ton from you. I love the way you explain information. How can people continue to get your thoughts after this episode? You can follow our publications on our blog called Tech at Bloomberg. You can follow me personally on X or Blue Sky at Sapgehr. So just the first letters of both of my names, just because it's a little bit long, or obviously on LinkedIn. The first syllables. The first syllables even, yes. Yeah, yeah, yeah. It'd be amazing if you got SG on either.

52:49Jon Krohn:Yeah, it's nice. you know, it's been a while since I've heard a Blue Sky one, because it seems like, yeah, it seems like most guests these days are focused on LinkedIn. But it's great to hear, you know, I actually, I'm really rooting for Blue Sky. Me too. And we'll see what comes out of it. A lot of academics have moved over. So I have to, at this point, still follow X and Blue Sky at the same time to get my deep technical news. But we'll see how it develops in the future. Nice. All right. Thank you so much, Sebastian. Yeah. And hopefully we can get you on the show again in the future when you have some more brilliant research insights for us.

53:28Thank you so much for having me.

53:34Jon Krohn:What a great guest Dr. Sebastian Guermanois. In today's episode, he covered how RAG can circumvent built-in safety mechanisms in LLMs. While RAG reduces hallucinations, improving honesty, it can compromise harmlessness. How organizations must evaluate AI systems in their specific deployment context because general purpose safety measures often fail for domain specific use cases. How effective RAG safety requires a guardrail retrieval answer guardrail architecture, not just vanilla retrieval and generation. And how financial services and other regulated industries need custom risk taxonomies and red teaming exercises to identify domain-specific vulnerabilities.

54:16Jon Krohn:As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Sebastian's social media profiles, as well as my own, at superdatascience.com slash 905. Thanks to everyone on the Super Data Science podcast team, our podcast manager, Sonja Breivich, media editor, Mario Pombo, our partnerships team, which is Nathan Daly and Natalie Jaiske, our researcher, Serge Massis, writer, Dr. Zahra Karche, And yes, our great founder, Kirill Aromenko. Thanks to all of them for producing another exceptional episode for us today.

54:49Jon Krohn:For enabling that super team to create this free podcast for you, we are deeply grateful to our sponsors. You can support this show by checking out our sponsors links, which are in the show notes. Otherwise, share the episode with someone who would like to have it. Review the episode on your favorite podcasting platform. Subscribe. oh and if you are ever interested in sponsoring an episode yourself you can find out how to do that at johncrone.com slash podcast but most importantly i just hope you'll keep on tuning in i'm so grateful to have you listening and hope i can continue to make episodes you love for years and years to come until next time keep on rocking it out there and i'm looking forward to enjoying another round of the super data science podcast with you very soon

55:38Thank you.

From the publisher

RAG LLMs are not safer: Sebastian Gehrmann speaks to Jon Krohn about his latest research into how retrieval-augmented generation (RAG) actually makes LLMs less safe, the three ‘H’s for gauging the effectivity and value of a RAG, and the custom guardrails and procedures we need to use to ensure our RAG is fit-for-purpose and secure. This is a great episode for anyone who wants to know how to work with RAG in the context of LLMs, as you’ll hear how to select the best model for purpose, useful approaches and taxonomies to keep your projects secure, and which models he finds safest when RAG is applied.

Additional materials: ⁠⁠⁠⁠⁠⁠www.superdatascience.com/905⁠⁠

This episode is brought to you⁠ by, ⁠⁠⁠Adverity, the conversational analytics platform⁠⁠⁠ and by the ⁠⁠⁠Dell AI Factory with NVIDIA⁠⁠⁠.

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

In this episode you will learn:

(03:28) Findings from the paper “RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models”

(09:35) What attack surfaces are in the context of AI

(38:51) Small versus large models with RAG

(46:27) How to select an LLM with safety in mind

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
905: Why RAG Makes LLMs Less Safe (And How to Fix It), with Bloomberg’s Dr. Sebastian GehrmannSuper Data Science: ML & AI Podcast with Jon Krohn · 58 min
Listen in VO