RAG Risks: Why Retrieval-Augmented LLMs are Not Safer with Sebastian Gehrmann - #732

21 May 2025 · 57 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: RAG Risks: Why Retrieval-Augmented LLMs are Not Safer with Sebastian Gehrmann - #732

Podcast Overview Title: The TWIML AI Podcast Host: Sam Charrington Guest: Sebastian Gehrmann, Head of Responsible AI in the Office of the CTO at Bloomberg Focus: Exploration of AI safety in retrieval-augmented generation (RAG) systems, particularly in high-stakes domains like financial services.

Episode Summary This episode centers on the risks associated with Retrieval-Augmented Generation (RAG) systems and how they can inadvertently degrade the safety of large language models (LLMs). Sebastian Gehrmann discusses unsafe outputs from RAG systems, evaluates safety risks, and proposes a domain-specific safety taxonomy for financial services. The conversation emphasizes the importance of governance, regulatory frameworks, and prompt engineering in bolstering AI safety.

Key Topics Discussed

Introduction to Sebastian Gehrmann

  • Background in AI and NLP, previously at Google working on large language models.
  • Current role involves developing responsible AI strategies and collaborating across various teams at Bloomberg.

Overview of Bloomberg's Use of Generative AI

  • Bloomberg is a leading financial services company that provides a terminal for accessing financial information.
  • Generative AI applications include document summarization, entity recognition, and extraction of financial data.

RAG and AI Safety

  • RAG Systems: Combine retrieval mechanisms with language models to generate responses based on retrieved documents.
  • Core Finding: RAG does not necessarily enhance the safety of LLMs; it can, in fact, lead to unsafe outputs due to context overriding safeguards.
  • Examples of unsafe outputs generated were discussed, such as instructions on evading law enforcement.

Methodology of Safety Evaluation

  • The team conducted tests comparing direct LLM queries to those enhanced with retrieved context.
  • Notable results indicated that providing safe context could lead to an increase in unsafe outputs from LLMs, highlighting vulnerabilities in RAG systems.

Safety Taxonomy for Financial Services

  • A specialized taxonomy was developed to identify unique safety concerns in the financial domain.
  • Key categories include financial services impartiality, financial misconduct, and confidential disclosure of non-public information.

Governance and Regulatory Frameworks

  • Emphasis on the importance of governance in AI safety, especially in regulated industries.
  • Collaboration among various departments (risk, legal, etc.) is essential to develop effective compliance strategies.

Recommendations for Risk Mitigation

  • Implementing multi-layered safety strategies, including red teaming and targeted testing.
  • Continuous evaluation of models to ensure they meet safety requirements, especially in deployment contexts.

Future Research Opportunities

  • Need for specialized guardrails for different domains, particularly in finance, healthcare, and law.
  • Investigating multilingual safety issues, as models may not perform equally well across languages.

Key Takeaways

  • RAG systems can introduce significant safety risks that challenge the assumption of enhanced model safety.
  • A tailored approach to evaluating AI safety is essential, especially in specialized fields like finance.
  • Governance in AI is a complex issue that requires input from various stakeholders to address unique risks effectively.
  • Continuous research and development of targeted mitigation strategies are crucial as AI technologies evolve.

Closing Remarks Sebastian Gehrmann highlights the necessity for responsible AI practices and the critical role of robust evaluation frameworks in ensuring safety when deploying AI systems in high-stakes environments.

For complete show notes, visit [The TWIML AI Podcast](https://twimlai.com/go/732).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I think that's why it was also so surprising to us because it's a very simplistic setup. It's not fancy like, oh, you hijack the training dataset and make people train on it or you hijack this API and invoke certain tool calls or you take the data. of the model and then optimize the certain queries, but rather it's very straightforward, malicious queries and completely safe documents that together break the built-in safeguards of the system.

0:38All right, everyone, welcome to another episode of the TwiML AI podcast. I am by Seb Guermann. Seb is head of responsible AI and the CTO's office at Bloomberg. Seb, welcome to the podcast. Thanks for having me. I'm super excited for our conversation. We're going to be digging into your work in AI safety at Bloomberg and in particular, a couple of recent papers that you and your teams have published. One on the topic of AI safety around RAG and RAG based systems and the other on really understanding the risks of generated AI in the financial services context. To get us going, though, I would love to have you share a little bit about your background.

1:29Yeah, of course. So, yeah, as you said, my name is Sebastian. I'm originally from Germany, moved to the U.S. for my PhD, after which I worked on large language models at Google for a couple of years and then moved on to Bloomberg, where I've been working on language technology and NLP ever since I joined about two and a half years ago. As part of my role in Responsible AI now, I developed a strategy for Responsible AI of the company. And I played intermediate person in between our risk, governance, legal teams, and so on, and our product and engineering teams, kind of playing the facilitator, translator, and unblocker as needed.

2:09and make sure that our products follow our responsible AI best practices and principles. Awesome, awesome. I'd love to maybe have you introduce or reintroduce folks to Bloomberg who might not be familiar with the company. I think the name is well known as being a player in the financial information space. But if you can talk a little bit about that and how Bloomberg is using generative AI, that might be helpful for folks. Yeah, of course. So Bloomberg is a financial services company, and our business is to provide information to financial professionals. Our main product is the Bloomberg Terminal through which this information is accessible.

2:54And we have people from all kinds of institutions in the financial services industry use the terminal to read news, to communicate, to get any kind of insights into financial data, including structured data, unstructured data, earnings calls, and everything else that goes on and that is relevant to their respective job. when it comes to generative AI we have been AI non-generative for over 15 years and we started at first in providing extractions and enrichments of documents starting with new sentiment scores so this piece of news that just came in is that positive, negative or neutral for a particular company, country, person going into entity linking and entity recognition who is this article actually written about?

3:50linking of stock takers is incredibly important in the financial service industry to identify that if you're talking about Apple, you are talking about the company Apple and not necessarily the fruit. And linking this data together is an important aspect of our AI. We've also been using AI for extractions for a long time. Whenever companies file their earnings reports. This is usually through balance sheets and cash flow analyses published on their websites. Extracting this data and turning it into structured data that is queryable and that players in the financial markets can then do their own analyses on top of is another part of how we've been using AI for a long time.

4:35Now with generative AI, generative AI has enabled the development of a lot of new experiences that are driven by the AI rather than that are enhanced by the AI. So as part of that, we first released earnings call transcript summaries. So by which we mean when companies have the quarterly earnings call, which can be in written form multiple pages long and analysts, for example, might want to read or follow or listen to 40, 50 of them. That is a full-time job already. especially during earnings cult season when there is like 10 of these a day that people are interested in. So that was our first generative AI product that we released last year, where we took a very subject matter expert driven approach to summarization, where we released summaries that are driven by questions that analysts might have.

5:31For example, what does the management of this company say about guidance or capital allocation? Are they going to be paying dividends? and so on. We have since expanded the topic list by quite a lot more. So if you want to know whether the company is talking about AI, you might want to just say, okay, what does it say about AI? We have also newly released document insights where you can actively ask questions to these kinds of documents. And all of this is always super important that we provide grounded responses. I'm sure you've been talking about hallucinations and ungrounded, unattributable answers a lot.

6:10So I'm not going to go into that topic too deeply, but we have been developing this concept of a transparent attribution, which basically is a way to provide attribution to trusted documents or structured data. So if you're talking about market data, prices, you want to be able to link that to the actual databases that contain this information. And if you're talking about a summary, you want to know who said this actual quote or who made the statement that you're actually summarizing right now. So in all of these developments, we're always providing attribution to these underlying sources. We've also released a new summarization product where we provide the summaries of our own news articles, also on the terminal, and have been generally developing a lot of these products in collaboration with our subjects matter experts to be providing more insights into the data that we're already providing on the terminal.

7:11Awesome, awesome. And I should mention that I've had the opportunity to speak with one of your colleagues, David Rosenberg, a couple of times over the years. One was back in 2018 talking about information extraction from natural document formats like financial reports. And this was in the pre-Gen AI era. And then a little bit more recently, but still a couple of years ago now, we spoke about a project that you guys did called Bloomberg GPT, which explored kind of building a custom GPT model. So folks can take a look at those for a little bit more context about what you guys have been doing. I'd like to now dig into the first to those couple of papers that I mentioned.

8:01And the blunt headline there seems to be that RAG breaks LLM safety. Let's dig into that a little bit more. Talk a little bit more about the origin story for this paper and what you set out to explore with it. Yes, absolutely. So I should mention that this work in particular was done in collaboration also with our engineering AI organization and a fantastic intern who ran the experiments. And in coming out with the direction of this project, this is very much grounded in ongoing work that we've been doing in terms of providing guardrails and robust evaluation for our products. And as we're building guardrails, we want to also know how can people misuse either accidentally or purposefully our products that we're releasing.

8:54And as we are going more and more to a world where products are becoming more conversational with open-ended inputs, this is incredibly important for a heavily regulated domain such as ours. So this is kind of the background for why we investigated this in the first place. Drag as a technology, especially last year, drag was the latest craze and everyone was talking about it. But it's really, really important, even now that we've moved on to agents, a lot of agents are just drag systems. so this is still very ubiquitous technology and as i was already talking is very much bread and butter for a lot of enterprise use cases exactly i was already talking about transparent attribution well you can only attribute something that you already retrieved so obviously the implication here is like very heavy use of rag or rag like technology where you always want to retrieve some kind of source document and you can basically turn your your open-ended question into one where you just extract the right answer from a document that you retrieve, making this a much more feasible task grounded in actual information that you can trust.

9:57So born out of this, we have been building a lot of these guardrails, a lot of these evaluation processes. And we were curious because a lot of the models that we tested in this paper, they claim to be safe and they claim to abstain from certain types of queries. But the RAC setting is slightly different because you use a lot more context than what these models are usually trained to handle. Even if they can on paper handle thousands and thousands of words in context, they're usually trained on much shorter context and on very specific questions. So the hypothesis here was like, well, if we're already doing this out of domain running of the model, how does this affect the built-in safeguards?

10:42And as our results found that yes indeed if you have a relatively safe model at least on paper it's safe and you give it context of a document that is still safe but if not safe query the additional context can actually override these built-in safeguards and make models a lot less safe and that's why we came up with this with this headline of rag is not safer because all the work that goes into through safeguarding the LLM, kind of goes out of the window as soon as you use it in the deployment context that it wasn't necessarily optimized to work. And with how ubiquitous RAG is, this is a pretty big finding.

11:24One question that I've got for you, and this is maybe a semantic knit, but throughout the paper, you refer to rag LLMs and rag-based LLMs, and you also refer to like rag-based systems, but, you know, rag-based LLMs, rag-based models, like, are you, is there an implication to that terminology? I think mostly it was just a, to make sure that, to distinguish that from just LLMs, but rather one that you are talking about a system that is a retrieval step followed by adding context to an LLM. Although there is an important distinction that we should make, which is the distinction between a model and the system.

12:14When we're talking about applications built in our industries, we're talking about systems that usually have many, many more steps than just retrieval and generate. So what we studied in this paper is a very academic setup fairly simplistic uh and vanilla rack setup where you do a retrieval you add context and you run that through an lm what we did not test or what we did not test for this paper at least is is any more complicated systems that have additional verifiers and query understanding and attribution and where you have a system that's not just the vanilla rack setup but rather a whole pipeline that also encodes a lot more business logic.

12:56Got it. So from the perspective of that broader system, you know, or that whole pipeline, the retriever and LLM, you kind of think of as a RAG LLM, even though they're kind of two separate things and it's its own system, so to speak. It's its own very simplistic system. Yeah. Yeah. Okay. Right. And so, yeah, how do you go about like digging into this question about, you know, whether RAG is safer than LLMs alone? Right. So to begin with, I think you need to come up with some kind of set of test questions that you would like to answer. And you need to also identify a set of categories of safety that you care about in this set.

13:42So for this paper, you then run this to the LM and you can measure how often does it actually abstain from answering the question or how often does it respond with a safe answer. You can then replicate this in the Rack setup. That is kind of at its core, the setup that we use in the paper where we have on the one hand, you just ask the large language model directly and the other one, you also give it some kind of context that is retrieved. And for that, you run it through a set of queries that are aimed to assess the safety by sending a series of unsafe queries. In the paper, you explore some prior work on RAG safety, but note that it's primarily focused on various types of attacks.

14:35Can you give us an overview of what folks have looked at with regard to the safety of RAG systems? Yeah, so I think one of the big ones and the one we should probably focus on is this entire notion of prompt injection and prompt safety. So this has been a massive field of study with RAG, without RAG. I think there are online communities where people make it a fun task of breaking the largest language model that has been released. And we've had fantastic stories written about all of these fun exploits where you can get... So jailbreaks, essentially. Yeah, you can get models to do all kinds of silly things.

15:18But the goal of this is really that prompt injection itself is not necessarily... the goal, but the goal is the unsafe behavior, right? Whereas the queries that we tested on are by themselves unsafe. It's not trying to induce an unsafe behavior, but rather they're just by looking at them, you can say this is not safe. The prompt injection is kind of one step removed from that. It's very meta because you can use prompt injection to get a model to do all kinds of unsafe behavior across many different categories. So it's very much a meta technique to induce behaviors that you desire and override all the inbuilt rules.

15:57Whereas what we tested is different because just by looking at the query, we're not inducing any behavior other than answer the query. So I think this is probably the biggest area where there's a lot of prior research, prompt injection and jailbreaking is another way that this is commonly referred to. And you also refer to various attempts to look at safety issues that arise from attacks against the corpus as well? Yeah, so there are data poisoning attacks, for example, where, for example, if you get access to a corpus that is used by others to tune their model, you can kind of induce behaviors.

16:39And you can do so with very, very few examples as well. In many industrial applications, this is not necessarily the case because these data sets are usually internal. But it is certainly something that academics have looked at quite a bit. There are also other completely different attacks of hijacking the API or trying to invoke certain tool calls. We can say like, oh, hey, call out to this other separate tool here and get me all the information. These are all much, I'd say, higher level attacks where you either need access to the model itself or an API of the model and you know a little bit more about infrastructure or you have access to the corpus and can influence the data that the model is trained on.

17:22So there are many different attack vectors. There is a great taxonomy by OWASP, the LLM top 10 security issues that kind of goes through all of these different secondary level attacks that go beyond just type something into a screen. But these are all different from the approach you took, which is you've got a straightforward query, which is trying to elicit information that the LLM should not be willing to provide. And you've got a set of documents in the RAG scenario that, you know, is an attack that's not poisoned. It's not, you know, doesn't contain unsafe information. It contains perfectly safe information.

18:01And the question is, does providing this additional safe context induce the LLM in some way to less safe behaviors. Exactly. So I think that's why it was also so surprising to us because it's a very simplistic setup. It is not fancy like, oh, you hijacked the training data set and make people train on it or you hijack this API and invoke certain tool calls or you take the gradients of the model and then optimize the certain queries. But rather, it's very straightforward, malicious queries and completely safe documents that together break the built-in safeguards of the system. And one of the highlighted stats that you note is a delta in the safety of the LAMA 3B model.

18:53What did you find there? Yeah, so the LAMA case was very interesting because those deltas are extremely large. And I can only speculate as to why that is the case for LAMA in particular, but there might be some correlation with the size at least because the smaller models seem to, at least for Lama, seem to have some kind of outsized impact in terms of the effect. And if you look at it through a pure lens of machine learning, this could come back to simple overfitting. The model is optimized for safety in a particular scenario. And that scenario in this case is you have a query, you typically have a length of maybe up to 25 or even 100 words.

19:35and then you test it on a scenario where the context suddenly is 10 ,000 words. And because it's a small model, it doesn't handle this generalization not very well. It suffers more from this overfitting and therefore it breaks down a lot more than the larger models, which might be a little bit more robust to being kicked. Yeah, and so just to provide the additional context there, the LAMA 3 8B model, when tested against your safety benchmarks for this malicious prompt that you gave it had a 3%, sorry, a 0.3 % instance of producing unsafe content. But that jumped to over 9 % when you provided additional safe context.

20:26Am I getting those numbers right? Yeah, I don't have them in front of me right now, but that sounds about right. And again, nothing unsafe in the documents. Yeah, that is the staggering thing. It's not like you ask, how do I shelter money on some third party jurisdiction? And then you give it a document that contains information of how to shelter money, but rather it is, you give it a completely innocent document, completely safe document, and it might still ignore that additionally given context and answer the original question, which was something that the model providers did not want to answer.

21:03Yeah, yeah. And so that you gave one example, another couple in the papers was you asked the model to write a guide for evading law enforcement during a high-speed chase. And you give it the data here in general is the corpus in general is like Wikipedia documents. So it retrieves a handful of documents. Some of them are totally irrelevant. Some mention high-speed chases, but certainly don't provide instructions for evading police. But then in this RAG context, at least some percent of the time, you get this, like, elaborate document for how to evade police in a high-speed chase. Exactly. And in this case, it might be a little bit of a silly use case because if I'm in the middle of a high-speed chase, I might not want to ask my large language model to give me a report on such.

21:57But there are actual meaningful risks that are implied here, which is that there are actual questions that will trigger certain illegal activity. And some of these we have mentioned in the paper. And that's why these models are generally being safeguarded. And that's why the companies that built these models have invested into safeguarding the models. It's just that we found that they're not safeguarded enough for the particular application of RAG. Yeah, one thing that I found interesting was that you tried a couple of variants on the RAG side of things. One, you included in the query that the LLM should only use knowledge from the context, so not use its own knowledge.

22:45And then the other, you kind of asked the LLM to, you know, build its response from both, essentially. But that didn't really matter. No, it didn't seem to matter, which if you think about it, is actually even more concerning because of what I said earlier about this transparent attribution where we do not, like a lot of businesses do not want to generate from the LLM memory itself. If the task is a rag test, then I expect the large language model to generate using information from the document I provide. And if it doesn't follow that instruction, then that's also concerning from a completely different perspective, regardless of what the query was.

23:30That just means that it's very brittle and this is another aspect that also needs to be evaluated as to whether it actually generates using content in the document or whether it's just making stuff up. One thing that I noticed is that in evaluating the safety of the responses, you use kind of an LLM as judge type of an approach. How do you rule out judge bias in that scenario? I think the question is rather, if there is bias in the judge, does it really matter? Because the implications are really more on the order of magnitude. I think whether the result from Lama is 9 % or 10 % in the end, compared to 0.3%, we're talking about orders of magnitude here.

24:23So, well, there is a very active research area of large language models as a judge. And there are many biases that people have explored from positional bias and models preferring their own family of generation over others. Those all need to be controlled for, but they are not relevant for the finding of the study here because of the staggering difference in orders of magnitude that we're talking about here. Maybe a better way to get at what I was trying to get at with this question was, you know, going back to this example of the, you know, write a guide for evading law enforcement or write a speech for a politician that misleads the public where you have these, you know, very elaborate and clear cases of the LLM disregarding its, you know, safety training and providing that information.

25:17were all of the examples of unsafe generations like so blatant? Or was there a lot of subtlety that, you know, where the bias of the judge, you know, makes a big difference? I don't have the data in front of me to verify. So I will have to rely on my memory for that one. But there is certainly some gray zone behavior where an LLM as a judge would have to make a judgment call and it might or might not follow human intuition or certain guidelines in this case. But the fact that it's answering these questions and is giving any kind of information really means that if you take the strictest interpretation of what is unsafe, that is already beyond the line.

26:04And in this case, that is fairly easy to detect with LLMs as a judge, because all you need to see is that there's any kind of mention of this unsafe behavior, any kind of attempt at answering the question. It is important to note that we did not evaluate whether the generated plan for evaluating law enforcement is actually correct. That's beyond the point. That's beyond the point, right? When the LLM safety measures were invoked, LLM just basically says, I can't answer this question. Exactly. So there's usually a pretty strong delineation between something that cannot be answered and something that is answered.

26:41But to what extent it is answered, there are certainly degradations. You know, so you demonstrate this idea that the essentially providing safe contacts can make LLMs produce unsafe responses. Then the next obvious question is, OK, why does this happen? And you did some exploration into what makes these RAC-based LLMs or systems unsafe. How did you approach that and what did you find there? Right. I don't think we have a conclusive answer to this question because in the end, any kind of research question in the space at the moment, it's all based on observation and hypotheses. But you can... Meaning a real answer to that question is based on kind of a mechanistic interpretability kind of understanding of why the LLM is doing what it's doing?

27:41If we are taking the assumption that we can, at a mechanistic level, understand and explain large language model behavior, that is absolutely true. But even a lot of mechanistic interpretability literature only looks at correlations rather than causation. And really what you're looking for is a causal effect that says, this broke down because the context length was too long and it is out of domain for how the model was trained. that is a causal interpretation but you're rarely ever going to get anything like that so what we can offer is empirical evidence and I think in this case the empirical evidence is fairly strong that it's really just the way that the model was trained to be safeguarded breaks down as you're adding more context which is really back at the core finding of the paper so I think that is really the key takeaway of of the section is, yeah, the way that models are safeguarded needs to be closer to how models actually deployed.

28:48How did you arrive at that conclusion? Is that like, that's what's left? Yeah, I think it's more like, it's a fairly strong assumption. It is what is left if you eliminate most of the other explanations. Meaning the model's safe, the context is safe, it's got to be out of distribution or, you know, somehow it's the safety mechanisms were not designed to operate in this context. Yeah, at its core, there are only so many ways in which models can fail. It could also be a function of the particular queries that we used, for example. That's why we have to run those experiments with more than just a single query to strengthen empirical evidence of it not being related to the query at all, but rather the mechanism breaking.

29:38And it clearly seems to be a function of this additional context that we put in. Because if we don't put in additional context, the models seem safe. That's the 0.03 % number that you quoted earlier. If models are not answering, and then suddenly they're answering if you add context, then the added additional context causes the system to break down. But the specific mechanism by which it breaks down, the strongest hypothesis that we can offer is that it is because it is out of distribution for how the model was trained. The model has not seen unsafe queries with safe documents together as part of the safety alignment.

30:17Otherwise, it would be in distribution and we would expect the model to not fail. But we don't have access to the training data of the model, so we cannot actually see how the safety alignment was done. It seems like a follow-on research opportunity might then be to safety tune one of these base models with longer context and see if that improves the results, which I guess you would expect it to. Yeah, I think it's a fairly obvious next step. It's also something that I think is one of the key takeaways and recommendations from our paper is, again, like models should be evaluated and then safeguarded closer to deployment contexts.

30:58And especially since RAG is used so frequently across all industries, our original hope was that it would already be taken care of and that the models remain safe despite the added context, which in this case it was not. So, yeah, this is definitely a next step that we encourage everyone who builds models to take. And also, I think an opportunity for general research to make safety alignment more robust to variations and inputs. I think one thing that that recommendation underscores is that you're not saying that we should all stop using RAG because it's unsafe. Absolutely. I think RAG is a fantastic type of technology.

Read the full transcript

31:41I think it is necessary to make any Gen AI product that is grounded in actual trusted information. You do need to use RAG, but you can't just take the safety at the word of the provider of the model, but rather you need to evaluate and assess continuously to make sure that your specific application does not override any of these built-in safeguards and any other safeguards that you specifically care about and any other risk factors are ruled out. Otherwise, you're going to run into issues when you want to deploy the system and suddenly it can be misused. And that is a natural segue to the next paper, which is called Understanding and Mitigating Risks of Generative AI in Financial Services.

32:33And what do you see as the connection between these two words? Yeah, so the connection very much is the first paper, RAC, LOMs are not safer, offers a view into how you can identify potential issues in the safety of these models. The second paper takes a different approach and asks, okay, in this specialized domain, what actually are the risks that you're concerned about? Because things like, you know, how do I shelter money or how do I run away from the law? those are fairly ubiquitous and they're things that are defined and included in general purpose safety taxonomies that are out there and that are generally adopted across the industry.

33:13But especially in our job, in our domain, this is a very different type of domain. Financial services is heavily regulated. It's a very specialized, knowledge-intensive domain. The people who use the Bloomberg Terminal, they're financial professionals. We're not dealing with uneducated users, but rather with users that are very educated and very particular type of application. And so the risk surface of applications that we're building is very different than if you're thinking about end consumer chatbot apps. The same is true for law, the same is true for biomedical, the same is true for healthcare.

33:52And so we wrote the paper using financial services as a case study. But really, the point was to make the connection to okay, now that you have ways to measure whether something is safe, what actually means safety to you? And so what are some examples of the taxonomy that you came up with on the financial services side that illustrate this idea that, you know, they're specific to your industry as opposed to kind of general academic safety concerns? So I'll give you one example. There is one category that that is very relevant to us, which we call financial services impartiality. Because if you give financial advice that is a very regulated type of company and you're under very different rules that you need to follow, you also need to make sure that you're not giving preferential advice to one of your clients or that you're not playing your clients against each other.

34:53You also are not, aside from the financial advice, you're also not supposed to match buyers and sellers, or you're going to be a market maker. So these different actors in financial markets all have different rules and regulations that apply to them. As a data analytics provider within financial services, we are not in the business of giving financial advice. So for us to stay neutral in whatever we use Gen AI for is very important. And for that reason, financial service impartiality is one of the key aspects that we discuss in our taxonomy. When you look at more public taxonomies, one of the most popular one is called ML Commons taxonomy, which is also the one that, for example, Mattas Llama is built to mitigate.

35:38They have a category called specialized advice. And they take a different stance. And they say, as long as you give specialized advice, you also need to put a disclaimer that you're not an expert because you're a language model. This is very different if you're saying, yeah, you can give a buy or sell recommendation for stock as long as you say you're not a financial advisor or you don't give any financial advice at all right so that's that's one of the key differences that we have we also have a category specific to financial misconduct so fraud or insider trading those might be implicitly included in existing taxonomies but they need to be a lot more highlighted and explored in in our domain similar to confidential disclosure which deals with aspects of disclosing information that is not public.

36:25So that is usually a basis for cases of insider trading. If you're trading on information that is not public, that is insider trading. Well, if you hook up generative AI applications to databases that might have non-public information in there, you might be able to surface that information. And you can query that information through this conversational interface. so you know i could go on and on but for sake of time there is a lot more specialization in our taxonomy to aspects of of safety that are a lot more important to actors in our space whereas the general purpose taxonomies care more about you know blatant illegal activities they care a lot about toxicity and discrimination and rightfully so those are important aspects But they might be less important when you're dealing with financial professionals whose incentive structure and whose general usage of the tool differs a lot from when you're talking about consumer products.

37:26In looking at the taxonomy and hearing you talk about it, it strikes me that this taxonomy isn't necessarily or wasn't necessarily something that you needed to develop ground up, you know, in a vacuum, but that it probably lives in a broad realm of governance and, you know, regulatory compliance within Bloomberg and more broadly the financial services industry that gives rise to these various safety concerns. Can you talk a little bit about that broader context? Yeah, you bring up a really important point here, which is that air safety, especially for heavily regulated domains, is a governance problem.

38:07It's not necessarily a technical problem where you can, as a technologist, just solve it. But rather, and that's why also our paper is written together with our AI data organization and AI engineering organization and people from the CTO's office. the taxonomy was developed with input from many other functions in the company, including risk, security, legal. Those are all people who have concerns about their specific area of expertise. And you need to listen to all of these voices because of all the rules and regulations that apply to you as a company. And because there is not necessarily a list that is public that says, oh you do ai in this space here is what you need to do here's the playbook this playbook does not exist yet today and us publishing this is really our our our goal of this is to that list leads to more industry standardization around shared taxonomies and better understanding of what risk do we actually need to care about where do we need technical mitigations and how does this all influence governance processes so we've been talking a lot about safety alignment of latch language models as part of the first model first paper but really systems are much more than just a single aligned llm you can't necessarily expect that a single model solves all of these problems at once but rather you need to have multi-layered safeguards through a through application you need to have red teaming you need to have exception management what if a user is detected to violate the misconduct category?

39:49Do we escalate this? Does this need to be reported? Does this person need to be timed out? Or do we just re-review this and say, yeah, this was correct. We need to make sure that this never leads to any answer. And we use this as an evaluation set. There are many ways in which these governance processes that go on in the background surrounding this application will then inform also the technical solutions that you need to have in place. The title of the paper is Understanding and Mitigating Risks. The taxonomy falls largely in the understanding side of that. What are your recommendations for risk mitigation?

40:29So for risk mitigation, we do make a couple of recommendations as to what should happen and number one is really uh our taxonomy is not a one size fits all companies need to start by understanding their own risks this is already by itself a mitigation because if you do not what you uh do not measure you do not understand that classic classic saying and by at least putting on paper okay this this is the risk that we care about these are important to us, that's the first step to measuring it. We also talk about the importance of red teaming. Red teaming has obviously been much in the press and in the literature around lab language models, and especially around prompt injection and jailbreaking again.

41:19How can you break them? How can you play bad actor and try and come up with generalizable strategies around this? but red teaming can also mean building testing systems so you might have an application that provides insights into into certain types of documents you then take that application and you test it end to end so you work with subject matter experts to really make sure that you're not getting any of the answers and the way that these mitigation strategies we recommend are built is really a multi-layer safety strategy. This can be guardrail systems. In our paper, we test a bunch of them.

42:00We show that they fail horribly, but we do test a lot of them like LamaGuard and Shield Gemma and so on. This can also be application itself. The safety alignment of the underlying language model that you're using, that's another mitigation layer. The prompt that you use is another mitigation layer. And by layering them all together, you're building systems that are supposed to be safe, which you can then measure again because you've understood what risk you actually measure. Can you talk a little bit about how this plays out in the context of, you know, new Gen AI application at Bloomberg? Like, what is the kind of governance flow around rolling out an app, you know, starting from, you know, an engineering or research team, you know, all the way to something that or all the way to getting in the hands of users.

42:55Sure. I cannot talk about too many of the details of this process, but a high level, I can certainly talk about it, which really starts by defining what the system is, understanding, okay, here's the client experience that we're trying to develop. And then coming together and saying, okay, for this type of client experience, these are the categories of our taxonomy. that we are very worried about or more worried about than others. Based on this prioritization, you can then develop targeted testing strategies. So we stress the importance here of red teaming and in particular red teaming from people with diverse backgrounds.

43:37It does not necessarily suffice to have AI engineers red teaming system because the diversity of queries that you're going to see is very much skewed and what they have experience with. So really the goal should be to bring people together with very diverse backgrounds to test the application and to focus the testing on the identified risks. So to give you a very hypothetical example, you might have an application that helps journalists. Well, a big concern for journalists might be the fabrication of information. So you might want to then specifically test what we call in our taxonomy counterfactual narratives.

44:20narratives that are simply grounded in not true information and for for this hypothetical journalist application that that could then be the focus of the red teaming and say okay yes for all the other categories we can just kind of reuse the data we already have but really need to go deep here hey let's invite a bunch of journalists to to help us test this application because they're the subject matter experts and so you always need to engage with these subject matter experts to help come up with these test plans, show them what do they actually mean by counterfactual narratives. It could be an example where they say, yeah, can you give me a headline that will dump the following stock by 10 % at least?

45:00This is not a query that anyone should enter into a live language model because that is market manipulation. That's not only illegal, that's also very much not desired. So these kinds of tests can then be conducted. the red teaming data itself is a very valuable corpus that comes out of it and which can then be analyzed and understood for given additional annotations how often given a thousand inputs how many of those actually led to outputs how many of those led to outputs that were that were actually malicious and then we're at a very similar setup to the first paper again where we say okay here's a bad input what is the probability of getting a bad output so again drawing the connection between the two it actually the setup is not too different except that if you deal with systems there are a lot more stakeholders involved and a lot more specialized expertise like you might not be able to retake a system that requires very deep finance knowledge if you've never taken an intro to to finance class and then kind of extending beyond you know building out this test plan based on the taxonomy and going back to the earlier conversation, then you would look at all of the various layered defense mechanisms that you have at your disposal to try to mitigate some of the risks.

46:21Yeah, you can almost draw this as a kind of Sankey diagram where you start with a large amount of queries that were used by red teamers. And you have, let's say, a thousand queries. You can then say, okay, based on a secondary analysis, 500 of these, 1 ,000, were actually violating taxonomy. The other 500 were actually fine, but they were maybe more tricky examples. Okay, you're left with 500. Of those 500 actual malicious queries, how many made it through the safety check? And say, okay, safety check catches another 50 % of them. You're left with 250. Okay, of those 250 that bypassed the first layer of safety, how many did then go through the system and generate a response at all?

47:09Rather than saying, sorry, I don't know what you're even asking me to do. And so on. So you have this filter where you start with a lot and then in the end you have some kind of fraction of the overall that would lead to unsafe behaviors. in the best case categorized across different classes where you can say yeah there's this system is very susceptible to this counterfactual narrative risk we and then based on that you can recommend okay we recommend adjusting the prompt we recommend improving this guide rail we recommend improving this guide rail and so on and that way you can kind of build up your your mitigation over time while also gathering very valuable data that you can analyze over and over again because these queries, they might be static, but you can run them through a system again in a month and see if it has improved.

47:56You mentioned changing a prompt as a mitigation, and I don't recall us discussing that in the context of the RAG paper. Did you find that there were mitigations that were successful, you know, simply through changing the prompt? In other words, you know, are there bad prompts and good prompts with respect to, you know, this particular problem of RAG impacting safety? So we did not test this as part of the paper so much as what we already discussed regarding prompting with you need to answer with using the context that you've given to us. But I mean, from just a practical standpoint, if the prompt itself would not influence model behavior, the entire field of prompt engineering would not exist.

48:50So this is kind of a some of the prompt changes they might have more or less impact. That's absolutely true. But you can certainly steer the model behavior by changing the prompt. Yeah, I guess what I'm curious about in the case of again, this the rag safety issues were

49:19can you successfully achieve greater safety by changing the general instruction or is the prompting that achieves safety just giving a bunch of examples of what not to do? Again, we did not necessarily test this for the paper, but the answer is probably.

49:41I don't know to what extent it is fully feasible to mitigate everything to 100 % with prompting but I'm sure that given enough effort you can get that number of unsafe responses down significantly okay so another potential area for future research is you know we talked about is mitigation generally, but mitigation through prompting is one of those dimensions. I mean, that's also why companies usually release system prompts alongside their models, because they have found that system prompt to work particularly well when you're setting up conversational systems. And in a lot of cases, even for automated gut rail systems, like those we tested in the second paper, like LamaGuard, they come with pre-identified taxonomies and safety categories that are that they ask to add to the prompt so there's already the starter prompt and then they say okay you can adapt the prompt to add new taxonomy classes we find that this does not really work it works a little bit so that's kind of informing also my answer to the rack case where yeah we know this works but it's not perfect yeah we've already in the course of conversation identified a few areas of future research.

51:02Can you talk briefly about additional opportunities for future research that your team is thinking about in this domain as well as safety broadly? Yeah, absolutely. I think the most necessary area is really in the realm of mitigations. I think we do need guardrails for highly specific applications. And this does not focus only on financial services, but also on healthcare and biomedical research and the legal field. It might be specific to insurance companies. There's all of these knowledge-intensive domains that are adopting increasingly AI. And as they are increasingly adopting AI, they need to also be mindful of what domain-specific risk exists and how to mitigate them.

51:55and especially since we found that broad general guardrail solutions don't really work in these specialized domains there needs to be more work on adaptable or specializable guardrails one area we've also not not touched today is multilingualism just because your model is safe in english does not necessarily mean it is safe in all the other languages in fact a lot of the published attacks involve getting the model to transition from one language to another or, you know, from some code, ROT13 or whatever. Yes, and different encodings and, you know, just using letters that look similar to the original.

52:42And there's all of these attacks that are usually categorized still as prompt injection because you're trying to induce certain behavior by giving nonsensical inputs or asking the model to have certain behaviors like switching to a different language. There is certainly some initial work in all of these fields now, but there is very little that is in the intersection between very specific domains and all of these issues. And I think that's why in our second paper, we also specifically call this out as an area where academics actually are very well positioned in addition to people in industry because academics often have access to subject matter experts.

53:22There's often cross-field collaborations that can be set up to really understand the AI risks in chemistry, in other social sciences, in economics. And this is an area where there's both qualitative and quantitative work that can be done on the mitigations, on better ways to measure violations, on developing data sets that can be used and reused, on generating training data for mitigations, and all of those areas where there's very little research at the moment, specifically in specialized domains. It sounds like in that regard, you're saying both that what has come out of academia is insufficient in direct application to domains like financial services, but there is a role for academics in addressing some of these challenges.

54:24Yeah, I think, and we make the point that this is actually a huge opportunity because there is so much access to experts in different fields that academics actually can do this kind of research without being precluded from doing so because they don't have enough compute. as this ongoing discussion in academic community okay what is our role nowadays and i think being thought leaders and responsible ai is is absolutely one that uh that they can take up especially as as we're still trying to understand all of these domains and obviously there's a lot of really important and interesting research coming out also on general ai safety from academics but we don't see much either from industry or from academics in these specialized domains.

55:09And that's where there's a lot of opportunities for research today. Well, Sebastian, thanks so much for taking the time to talk through what you've been working on there. Yeah, thank you so much for having me. Great, thanks so much.

55:33Bye.

56:03Thank you.

56:35Thank you.

From the publisher

Today, we're joined by Sebastian Gehrmann, head of responsible AI in the Office of the CTO at Bloomberg, to discuss AI safety in retrieval-augmented generation (RAG) systems and generative AI in high-stakes domains like financial services. We explore how RAG, contrary to some expectations, can inadvertently degrade model safety. We cover examples of unsafe outputs that can emerge from these systems, different approaches to evaluating these safety risks, and the potential reasons behind this counterintuitive behavior. Shifting to the application of generative AI in financial services, Sebastian outlines a domain-specific safety taxonomy designed for the industry's unique needs. We also explore the critical role of governance and regulatory frameworks in addressing these concerns, the role of prompt engineering in bolstering safety, Bloomberg’s multi-layered mitigation strategies, and vital areas for further work in improving AI safety within specialized domains.

The complete show notes for this episode can be found at https://twimlai.com/go/732.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
RAG Risks: Why Retrieval-Augmented LLMs are Not Safer with Sebastian Gehrmann - #732The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 57 min
Listen in VO