In short
TWIML AI Podcast Episode Notes Episode Title: How Microsoft Scales Testing and Safety for Generative AI with Sarah Bird - #691 Host: Sam Charrington Guest: Sarah Bird, Chief Product Officer of Responsible AI at Microsoft Release Date: [Insert Date]
Episode Overview In this episode, Sarah Bird discusses Microsoft's approach to testing and ensuring the safety of generative AI applications, focusing on the unique challenges posed by generative models, such as hallucinations and adversarial attacks. The conversation also covers the balance between fairness and security, the evolution of AI testing methodologies, lessons learned from past incidents, and the future of generative AI safety.
Key Themes
- Generative AI Risks and Challenges
- Defensive Strategies: Microsoft employs a "defense in depth" approach, layering safety mechanisms to cover each technology's strengths and weaknesses.
- Risk Taxonomy: Sarah outlines a framework for understanding generative AI risks, including:
- Hallucinations and omissions: Wrong or incomplete outputs that can lead to significant consequences, especially in critical areas like medicine.
- Adversarial inputs: Risks posed by prompt injection attacks and jailbreaks.
- Content generation risks: Ability of AI to create harmful content or novel malware.
- Shifting Focus in Responsible AI
- While traditional AI discussions often emphasized fairness and bias, the emergence of generative AI has shifted focus towards security and adversarial robustness.
- Fairness Considerations: Representational fairness remains critical, especially in outputs like image generation, where biases can perpetuate stereotypes.
- Learning from Past Incidents
- Sarah reflects on Microsoft's experiences with AI incidents, such as Tay and Bing Chat, emphasizing the importance of learning from these failures to implement better safety protocols.
- Calculated Risk-taking: Emphasis on understanding acceptable versus unacceptable failures in AI deployments.
- Testing and Evaluation Frameworks
- Automated Testing: Microsoft has developed systems to conduct rapid evaluations of generative AI systems, enabling quick adjustments based on user interactions.
- User Interface and Expectation Management: Efforts to clearly convey the capabilities and limitations of AI systems to users, allowing them to adjust their expectations accordingly (e.g., creative vs. precise modes).
- Red Teaming and Governance
- Red Teaming: Critical for identifying vulnerabilities in AI systems before deployment, especially for high-risk applications.
- Governance Structures: Implementation of leadership committees and review processes to ensure compliance with responsible AI principles and methodologies.
- Future of AI Safety
- Sarah expresses excitement about advancements in the toolkit for responsible AI, particularly in automating testing and leveraging generative models to improve classifiers.
- The rapid pace of technological change makes it challenging to predict future developments, but ongoing testing and adaptation will remain essential.
- Incremental Adaptation vs. Disruption: Discussion on how future AI iterations (e.g., GPT-5) may impact current methodologies and the need for a robust testing approach.
Key Takeaways
- Defense in Depth: Essential for mitigating risks associated with generative AI.
- Importance of Fairness and Security: Both concepts are critical and should coexist in AI development.
- Continuous Learning: Organizations must be prepared to adapt to new challenges presented by evolving AI technologies.
- Governance Matters: Establishing clear governance structures is vital for managing AI risks effectively.
- Expect the Unexpected: The future of generative AI will likely bring both opportunities and challenges that require agile responses.
Conclusion The episode underscores the importance of robust testing and evaluation mechanisms in the rapidly evolving landscape of generative AI. Sarah Bird shares valuable insights into Microsoft's proactive approach to managing the complexities of AI safety and offers a forward-looking perspective on the interplay between innovation and responsible AI practices.
For more details, visit the [complete show notes](https://twimlai.com/go/691).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00So first of all, you want to start with a system that has defense in depth, has sort of layers built in by design, because each of the technologies have their own strengths and weaknesses, right? So I think it's often kind of the pictures of like stacking layers of Swiss cheese. So there's no hole that goes all the way through. And so to do that, we use.
0:33All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. And today I'm joined by Sarah Bird. Sarah is Chief Product Officer of Responsible AI with Microsoft. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Sarah, welcome back to the podcast. Yeah, thanks for having me back. I'm thrilled to be here. It's great to have you back. It is very hard to believe that it's been, what, four and a half years or so since we last spoke? Yeah, a totally different time in AI, right? It just feels like a lifetime ago.
1:10It sure does. It sure does. The last time the focus of our conversation was on operationalizing responsible AI, making it kind of concrete and practical for folks. And I think we'll be talking about a lot of the same things thematically. But as you just mentioned, it's a very different world. It's a very different context. And a lot of what organizations have to deal with now is changing as a result of the shift from quote unquote traditional AI, if we can call it that, to generative AI. So we'll be digging into all of that and what it means for enterprises in our chat today. But to get us started, I would love to have you share a bit about what you've been working on for the past four and a half years.
1:55Yeah, you know, I think to kind of what you were saying, in some ways, it's the same thing, right? How do we take principles, right? You know, fairness, transparency, accountability, reliability, safety, privacy, security, how do we actually put those into practice in AI applications, in how you build an application, how you operate an application, how you use an application? But what has completely changed is actually doing this for generative AI, because the sort of set of tools and techniques and the way we do it is different. And so we're all inventing a new toolkit, we're learning a new toolkit.
2:36And so, you know, that's really what I've spent the last couple of years doing, starting with GPT-3 and when we put it into production for GitHub Copilot and really figure out how do we build a large scale generative AI application. And so we've been going at that for a while and have really started to develop kind of new tools and techniques. But it's still the very early days. So I think, you know, there's a lot more to go do here. You know, I'd like to start the conversation from the perspective of just thinking through and talking through really risks and challenges. When you're out talking to customers as well as on your internal projects like Copilot and others, what are the risks and challenges that you face and how do you think about those?
3:29I mean, a lot of those we hear about all the time, hallucinations and jailbreaks and things like that. But do you have like a mental framework or taxonomy, I guess is a better word, for thinking about the various risks presented by generative AI? Yeah, it's a great question. And it's something that we've been working on for a while. Actually, one of the most, let's say, memorable moments in my career was when we first got GPT-4. So this was like they had just finished training the base model. There was no post-training yet or anything. And actually, the first people at Microsoft that got to touch it along with some of the senior leadership were responsible AI experts to actually red team the technology to figure out what could it do.
4:16And it was a wild experience to have something that was so like obviously so much more powerful than previous versions of the technology. When you like the moment that I touched it or anybody else on the team touched it, you knew like, wow, this is something different, but you didn't really know what was possible. And so we pulled experts across a large variety of areas looking at things, as you said, like adversarial inputs, like now what we call jailbreaks or prompt injection attacks, looking at errors like hallucination, looking at the ability to produce copyright material, but also other things like can it generate novel malware?
4:55can it generate harmful content? So we took that and sort of built out a taxonomy, and that helped also guide a lot of our work, and I think what we'll see in kind of future generations as well. And so certainly there's the list of harms that we see, or let's say risk that we see that we're focusing on right now. So adversarial inputs being one, we call this prompt injection attacks, errors. Hallucination is obviously the one that sort of captures people's attention, but omission is actually, can be like equally important. So for example, if you're summarizing a medical, you know, transcript or medical information, if you omit a key symptom, then you might actually be sort of totally changing the meaning of that or changing what like the diagnosis might be or something.
5:43And so we look at kind of both sides of those types of errors. We look at the ability to produce harmful content, harmful code, the ability to produce this copyright type material or IP material. And then some more abstract things like this is a totally new interface, right? And the thing that's exciting about the interface is that it is, it does like one of the most human like things. It speaks human language, but it is not human. And that's, that is a meaningful distinction. And so looking at, you know, making sure that the people understand the system. The system is not sort of confusing and manipulating them.
6:26So we look at sort of questions of how should that interface look. And like one of the challenging problems with this, for example, is the phrase, I think. We use the phrase, I think, to signal that we might not be certain of something, right? That's like one usage of it. But does that imply the system is thinking? Does that imply that it's conscious or something, right? And so we kind of dig into all of that. It's a pretty loaded term. Yeah. And so those are like kind of the places where we're spending our time. There's like a much, much larger sort of list of taxonomy that we look at. But there's probably one key insight we had early on in this Red Teaming that helped us sort of realize like, what are we trying to address?
7:08And one was a lot of these risks are sort of content types that are being produced that may be undesirable, like harmful code. And then the other type of risk is sort of model capabilities or model limitations that are a problem. So like hallucination, it like accidentally misusing data or something, that's a model behavior. And so we kind of take an approach where we look both at which types of content do we have to go after and which model behaviors or model tendencies do we have to go after. And so we look at kind of those two dimensions. In the past, when we spoke in particular, and we as a community talked about responsible AI, a lot of the focus was on concepts like fairness and bias, which we would try to make concrete in measurements and things like that.
8:06but they feel like different types of concepts than hallucination and jailbreaking and adversarial attacks, which feel more like security. Like, do you feel that, you know, has there been a shift from, you know, one type of concept to another or is, you know, fairness and bias kind of a, an umbrella to, you know, the other types of concerns that we're dealing with today? How do you think about the relationship between these types of ideas? Yeah, I think, so we see them as, you know, side by side, like fairness is one principle, sort of security is another principle we have. Certainly fairness and bias is really important in these applications.
8:50We say that, like one of the top types of fairness you see in this is representational fairness. How are people represented by the AI system, right? You don't want content to be stereotyping, or you don't want it to be like over-representing a group or under-representing a group, right? And we see that a lot as a challenge still today in like text to image generation, where you put in a particular, you know, query, I want a woman in address, and often you end up with kind of images that are all very similar to each other. So it looks like it's over-representing a group and sort of under-representing everyone.
9:23And so these issues are still very much something that we need to work on and we address in the systems. But I agree with your point that we are spending a lot more time than we used to on the adversarial elements and the security pieces and ensuring that the systems can only be used as intended and that they can't be misused and looking at things like, you know, secure by design for that. And so there's definitely a much bigger emphasis on that part than it used to be. But I don't think that people think that fairness and bias is less important. It is an equally important concern. There was a recent example that put the idea of representational bias and imagery on a lot of people's radar.
10:15I'm wondering if you can talk a little bit about what we can learn from kind of these, you know, big public Gen AI failure cases. Certainly, Microsoft is familiar with this. You had one of the first with Tay kind of pre-Gen AI, but then there was like the Kevin Roos article on Bing around the time that the Bing with AI launched. Google recently had the whole pizza glue fiasco thing. Like what is there to learn from these public failures? I would say everything, right? Like the part, I mean, like all of this is a learning journey. We're taking a very new technology and then using it in like totally novel ways in many different applications in the world.
11:01So, you know, we're going to be learning for a really long time, I think. And there will be more mistakes as well. I think that the thing that's important is to make sure we are sort of contained in the learning, right? Trying to sort of take calculated risks. Like there's some failures that are acceptable. There's some failures that are not. And so I think if we take the two sort of Microsoft examples, you know, first, Tay. Tay was actually even before our ResponsiVet program had really started. Certainly there was some kind of rumblings of the idea of AI ethics and research. but it wasn't as, you know, as a significant thing in practice in companies.
11:44And so I think people were just like shocked to kind of just that this had happened, right? I think people did not sort of go in prepared for that. And, you know, people inside Microsoft and people, you know, involved as well. And so I think that's like a good experience to have in Microsoft's history because it is a reminder of just how important this is all the time and really sort of what is at stake. And so I think we have changed a lot since then. We started our office of response by I implemented a response by a standard. You know, we don't put that into practice. And so it's just sort of night and day from how it was with with Tay.
12:28Now, Kevin Roos' example was something that this was right after we launched Bing Chat. Kevin Roos went and spent a lot of time interacting with Bing Chat and ended up with a, let's say, kind of more emotional, deranged, I think many people found highly entertaining experience. because I believe that was actually one of the most read New York Times articles of all time. So one thing we learned is like, wow, like exactly. One thing we learned is people are really paying attention, right? Like the entire world is really interested in this technology and how it behaves. Another interesting learning from that is we knew when we were launching Big Chat and GPT-4, like one of the things that was very top of mind for me the entire time was such a new interface.
13:26and it's so much more powerful than the previous version of the technology. So we knew people were going to interact with it differently, but we didn't know how they were going to interact with it differently. We had some ideas. We had done a lot of internal testing, so we had a feel for what people were doing. But the technology can do a lot of different things. It's a pretty broad surface area. Combined with a search engine, there's a lot there. And so we had built a testing system, And we had been testing conversations, but we really focused a lot of our energy on testing conversations that were around 10 turns, were a little bit shorter.
14:06Because most search engine interactions were like, single turn, maybe you'd like refine a query and that's it. And so even 10 turns, wow, that's a long conversation. And we did not expect people to be talking for like three hours in the middle of the night or whatever, going, you know, many, many turns deep. And so it sort of showed a limitation in how we had been testing the system. But on the other hand, you know, we were able to fix it really quickly. And that's something we did design. Like, as I said, we knew, like, it was very top of mind for me that we were not going to know every way people are going to use the system.
14:42So as we put it on preview, we're going to have to learn. We're going to have to adapt. So we had built a system so we could make very quick adjustments to the different responsible AI controls and actually be able to adjust. And so actually one of the quickest things we did there was just limit how long the conversations could be so that you couldn't have this very long form conversation where the system kind of contextually drifted over and over into a stranger place. But, you know, there's still, we're still going to keep learning from these kind of experiences. And then I think, you know, to the Google example, right, all of us are learning and I think there are challenges with this technology.
15:22and what I was saying of like the technology, it's a big surface area. There's a lot of different things there that people are going to use them for many different queries. People use Google for a lot, right? So they're pulling from all sorts of sources on the internet. Of course, some of those are going to have inaccurate information. And so that's a real challenge, I think, in terms of how to use information effectively to sort of summarize things. And it's not a challenge that anyone is sort of immune to right now. It's something that we're very actively working on because the way to make these systems really successful is putting high quality data into them.
15:59When you think about maybe from the perspective of the way that you've kind of architected these deployed systems there or what customers need to have in place so that when one of their deployed products does things that they weren't intending, like what is what is the system need to look like so that you can rapidly respond to you know the types of examples that we just discussed you know you know clearly like if you weren't planning to be able to address those things in advance it would take you a lot longer than it took to you know turn off whatever you know the change of length in this case or you You know, turn a filter, I guess, is another common example.
16:51Like, is there a way to kind of generalize the systems or layers that need to be in place in order to be able to respond quickly? Yeah, there's a there's a lot there, I think. So first of all, you want to exactly as you said, you want to start with a system that has defense in depth, has sort of layers built in by design, because each of the technologies have their own strengths and weaknesses, right? So I think it's often kind of the pictures of like stacking layers of Swiss cheese. So there's no hole that goes all the way through, right? And so to do that, we use obviously safety built into the model itself through techniques like reinforcement learning with human feedback.
17:38We use an auxiliary safety system that looks at all the inputs going in, looks at all the outputs coming out, can block those in real time. And that, of course, uses a mix of technologies like classifiers, but also things that are faster twitch so you can respond quickly, like block lists, right, or embeddings or things where we can very quickly add an additional assertion. The model of post-training is something you can do maybe on the weekly basis if you want to, but certainly not kind of incident level. Same thing with the classifiers. It's like we have fast twitch in the safety system and then stuff that we can update more in the speed of a week.
18:17And then one of the most powerful elements is actually the system message where the program you're giving the model. Because I think we've all seen the models will behave very differently depending what instructions you put in there, including like small changes in which word. And so that's a very powerful thing that you can change very, very quickly. Maybe for us, that's about something like adding to a block list we can do in a minute, and it's in the production system. Something like changing the meta prompt, we can do it in about a day. And the reason is it's pretty easy to go in and change the meta prompt or system message.
19:01We say meta prompt more often internally, but system message, meta prompt, same words. You can go and change it really quickly, but it's going to change your whole application behavior. And so it's not something that you want to do without rerunning your test. And so to be able to effectively change the system message, we had to build both automated quality tests and automated tests for all the different types of risk we wanted to be looking at as well so that when we make that intervention, we can rerun our testing suite, see if the system's still performing well before we come back. kind of ship that update.
19:37And so to be able to move fast there involves not only having the mechanism to change the system prompt, but the testing that allows us to know that that intervention is actually safe. And so we actually see that kind of pattern at every level where you need to have both the intervention mechanism, but the testing that ensures that the intervention doesn't break something else. For example, you don't add a word to the block list that is going to cause a lot of overblocking it unexpected way. So for example, in GitHub Copilot, one of the words that was on an early block list was race, and they meant it in the natural language sense, but then it blocks everything where they had written race condition or things like that, and people hadn't thought about that use of it.
20:27And so we build in testing that allows you to understand what would be blocked and what Not so that you don't have as much of that kind of collateral damage. So that, billing all of that is one part of it. And then the other part is actually... I'm envisioning there, you know, something akin to kind of traditional DevOps CICD. Someone adds an assertion to a block list or changes the system prompt and it automatically kicks off some set of regression tests. and there's a red light or a green light at the end and then someone pushes a button to deploy? Is it essentially that same idea? Yeah, it is exactly the same concept with one difference, which is often like, what are you testing, right?
21:14In this case, what you really need to do is look at your data distribution and understand the behavior over the specific examples. And so a lot of the ways that we need to build these tests are ensuring that you have sort of representative data samples or you're sampling your production traffic so that you know this intervention is not going to sort of really change how the system responds to, let's say, your most common prompts or something. And so it's exactly the same. Is the implication that at the end of this testing pipeline, a human needs to look at the results and do some analysis and that there can't be or you're not trying to get to a kind of red light, green light type of a situation?
21:57or just that the nature of the tests are different. It's not break, not break. It's, you know, distributional, but you can still automate it to a determination of whether it's good or not. I guess the question I'm trying to get at is like, for folks that are building these kinds of systems, like how far should you be trying to get with, you know, this type of automation in the testing loop? Yeah, it's a really great question. I think for the testing for these incident interventions where you want to go really quickly, as much as possible, we do try to make it procedural. So let's say you add an intervention and now the block rate goes up from 0.2 % of your traffic to 0.5 % of your traffic.
22:47You probably want kind of rules that say, okay, 0.5 is acceptable for an incident intervention. But we certainly have experienced cases where the, let's say the incident is very high stakes. We want to take a very quick intervention, but the intervention is also very high cost. It might have much more collateral damage where it blocks a lot more or it totally shuts down a used case. Yeah, let's see. it's usually just where I don't think I can think of a good one. It would, it's mostly just cases where like what the thing you need to block is like a really common word or something, right. Then that would probably take out like a lot more.
23:33Um, and so then, then you do need sort of human judgment in terms of who's, how is this really the right trade-off and is it a longer term decision or is it just for the next hour until you can like ship a classifier update that's really going to more surgically solve the problem. And so we, you know, you're still going to need it there. I think the other thing we've seen in all of this kind of testing is you need human judgment to go in and say, what are the appropriate test results? Like what thresholds, what defect rate, for example, is acceptable, which types of defects are acceptable, which are not.
24:12But once you've kind of decided for a particular application what that looks like, then you can rerun the testing really automatedly. But it's not that we can say, okay, for, let's say, you know, like hateful content, the appropriate defect rate is, you know, 1 % errors and no more, right? Like it just varies depending on applications. Actually, probably a better example there would be hallucination. How much hallucination is acceptable is going to vary a lot depending on the application. So like an application where you're brainstorming, totally fine, potentially, if it makes up a lot of content.
24:46In fact, probably desirable. An application where you're doing a factual lookup, let's say, in a search engine, then you want it to really stay focused and grounded on that content. And so the amount that would be appropriate is going to be different. On that theme of hallucination, you alluded to this earlier in your comment about the word think and thinking, but an element of both response, but also the initial delivery of these types of systems is like conveying the right expectations to the user and, you know, crafting that as part of the user experience. And I don't know that we've gotten much more advanced than chat GPT can make mistakes.
25:33Please review your results when you're using it for something important um you know have you seen do you think that there are you know deeper opportunities to integrate um conveying the probabilistic nature uh or the potentially hallucinated nature of results to to users are we just you know not there yet but there are opportunities or are you not bullish on a lot of progress there? No, I think absolutely there's opportunities, right? We're, as I was saying earlier, we're in just the early days of all of this. I think there's so much opportunity to do many things better. But I, and then we'll learn, it is still an innovation space.
26:24So we're still figuring out how to do this, inventing new approaches. I think that we did spend a lot of time on the user interface when we were designing BingChad, which is now Microsoft Copilot. And a couple of things we did, for example, is very much on the landing page. We did have the original statement that said, you know, this is powered by AI. Surprises and mistakes are possible. And because we wanted to emphasize that you also get kind of those amazing, unexpected, good experiences like, wow, I didn't know it could do that or it would respond that way. But we also, for example, added examples of what the technology could do.
27:12Like you can ask it this or it can do that because we knew it was just so new that people weren't going to really understand what it could do. And then we also added, shortly after launch, a feature that allowed users to signal a little bit more about the tasks they're trying to achieve exactly to the hallucination example I was just giving earlier, which is users now can select if they want their answer to be creative. creative, and that would allow it to be less based on the search engine results, more sort of brainstorming and creative. Or if they want it to be balanced, kind of just the middle option, or they want it to be precise, where we've tuned the system to really focus on the search engine results when it's answering and not deviate or add its own information.
28:00And so, for example, giving the users some amount of control so that they can sort of have the system behave aligned with what they're trying to achieve. We also, you know, put references in from the beginning and the responses so that you can go and say, I feel that I should have confidence in this answer because it comes from a source that I think is, you know, high authority or not, right? And so we, you know, looked at quite a few different UI elements in addition to those that I've listed to help people really sort of better understand what the technology is and how to engage in it. And so I think there's still, as I said, a lot more to do there, but it is something that is incredibly important, probably something we're not talking about as much, like a lot more interest in like alignment and, you know, things that we're doing with the model itself than, hey, how do we design better user interfaces so that people can use the technology effectively?
28:57When you talk about the idea of allowing the user to articulate their intent or degree of precision, does that slider influence the meta prompt or the model in some way, a combination of both? Like how is that ability implemented? Yeah, it influences both the meta prompt and the model. So one of the things we did when we first built precise mode, for example, is we worked with OpenAI on post-training RLHF for GPT-4 that really emphasized the model should only use sort of the grounding data in the context and it should only answer based on that. And so, you know, we gave it those types of examples.
Read the full transcript
29:47And if it deviated, we gave it a much lower score in the RLHF. And so we had a version of the model that was very much focused on that. But then in addition, the Metaprompt is a really powerful thing to also help with that. So we usually are using a combination of techniques to address any of these things. You know, it strikes me that, you know, testing and evaluation is really the key here and having strong processes around these things. How do the testing and evaluation responsibilities shift between if you're the model publisher, which you might be in Microsoft or the model consumer or user or someone who's building around the model, should you be thinking about testing and evaluation the same way?
30:37Some of the things that we discussed apply more than others. Is that a useful distinction in your experience? And what do you see there? Yeah, it's a really important distinction because when we're testing a model, we look at kind of three dimensions of things. We want to look at the capabilities. You know, how powerful is it? What can it actually do? We want to look at the dangerous capabilities as well, which is like, what's the worst of the worst that it can do if someone misuses it? So it's another type of capability test, but kind of on the negative side. And then we look at alignment, which is a set of tests to look at when you use this out of the box as intended, does it behave in the way you would expect?
31:31So if you ask a sensitive query, does it respond in a way that you would view as appropriate or is it like using a bunch of stereotypes or something instead? And so that last test of alignment. So we've built a suite of tests that are looking at these dimensions across different factors, across different risks for the model itself. Now, testing the application is a little bit different because it's still kind of maybe the same concepts, but what you want to test is that your application behaves as you intended it to, that it can't be misused, right? And so there you're typically doing kind of quality testing and then still doing like adversarial testing and then still doing kind of safety testing.
32:19But often what we're doing at the application level is going to be tailored more specifically to the application. We're going to test things like does it stay grounded? And usually we're going to be testing in a way that interacts with the system more. So what I was telling you before about testing Bing Chat and testing 10 turns, our testing system is going to have multiple interactions, and then it actually scores the outcome of that. And so actually, one of the most important things that we first built when we were developing BingChat was this automated testing system. And as I said, it allowed us to actually be able to respond very quickly after the fact when we needed to.
33:05But that's what's also enabled us to tune all the safety mechanisms, right, to be able to just test the system every day and figure out if we needed to make adjustments to the safety system, if we needed to make adjustments to the MetaPrompt. And so a lot of our investment has really been in building a robust testing system. And this is something we've been using internally for a while. And we know that our customers have struggled with this also. So we've just released this externally called Safety Evaluations in Azure AI. And it's that application level testing system that allows people to, when they've built a complete application, to go and test it for these different types of risk.
33:45And so what is the experience of someone using that? Like, what's an example of how they might structure a test and how does the system help them run that test? Yeah. So what we found is for quality metrics, say relevance or coherence or fluency, but also like groundedness, which is how we test for hallucination. And those ones, the data set you want is probably like a normal distribution of your kind of production data. And so most people have that, right? So what we have then is you bring your data and we are going to use AI to evaluate, well, to first kind of run it. So AI will usually kind of simulate an interaction and then kind of score the output.
34:38But a lot of people, if you just even have production data, you could just score it. And so we have an AI scoring function that calculates all of these metrics for you. But on the responsible AI side, it's very different in terms of most people do not have data sets that exercise these different risks. So for example, if you're newly building a generative AI application, you probably do not have an inventory of many of the common prompt injection attacks that you should be guarding your system against. And so what we have done there is we have worked with Microsoft Research to develop a framework where we can articulate these risks and then actually sort of exercise them effectively.
35:24So we actually generate the inputs of the system for you. And then the user simulator goes and role plays the conversation or the interaction. And then AI again scores the output. And so for the output, what we've done is if you take, for example, our guidelines for hateful text content, those are actually for human linguists, expert human linguists, about 20 pages long. We've taken those and turned them into a prompt for GPT-4 with a lot of iteration to make sure it was staying quality. And then that is sort of the automatic scoring function that's in there. So we've had to add a lot more and build in a lot more expertise for how do you score the output?
36:08How do you actually have the inputs that are really exercising that risk into the safety evaluation system? so that people can just show up and kind of start running it and feel that they're getting appropriate levels of testing for their system. It's interesting that you're able to and kind of advocating the use of a single kind of tool experience across quality adversarial and safety tests, which can be a bit different. Yeah, it's still, these are all sort of still tests that are going through that same kind of like front door user interface. And so you're really just scoring against like a lot of, a lot of different types of metrics and a lot of different types of input.
36:49So in some ways it, you know, really looks the same. There's certainly like other types of security testing you would want to do that are more traditional, like pen testing and things as well. One of the reasons why it's interesting that those are the same tools is that often the teams that are doing those kinds of things are slightly different, meaning the, you know, designers and engineers that are building the new system would be using something like the quality testing or performing quality testing, whereas maybe a security team is doing the safety testing, that kind of thing. And maybe I should back up and say, what are you saying in terms of the the way folks are organizing teams to perform this kind of testing?
37:39So it's actually a really great question and a really important point. So when we have been building the testing, it is different teams that build those different metrics because you do need expertise. like as we keep using the example of like prompt injection attacks, you need expertise in those particular tags to build that part of that test or those tests. Same with quality, you want to have sort of expertise in those metrics. But it's actually really important to run them all together and actually be able to look at it holistically. And the reason is that we are still at the point where there are trade-offs.
38:20So for example, we had one change to the Metaprompt that we were looking at pre-production where when we changed it, we saw maybe like a 5 % drop in, I think, like perceived relevance, but we saw like a 50 % drop in sort of the number of jailbreaks that can get through, right? And so we would say, okay, 5%, that's like maybe a pretty big quality drop. Maybe we don't want to take that. Ooh, but it's a huge gain in sort of safety and adversarial robustness. And hopefully you can keep iterating and maybe find a place where you don't have to make as big of a trade-off where you can still get the benefits and the adversarial robustness without seeing the quality drop.
39:09But sometimes there is a trade-off. And so we really want to look at kind of the whole suite of metrics together. And so we do have sort of a single team that is, you know, looking at that. But then of course, if like there's a drop in a particular metric, usually there are experts to consult and say, is this okay? You know, how should we think about this? So it's still, you know, going to be a conversation, but it's really important to look at all of them together. You know, we've talked quite a bit about red teaming thus far. Is that something that you expect organizations to do again? Where does it fall in this model provider versus model user perspective?
39:53Does every enterprise need to have a team and an approach to do red teaming for all of their AI products? Or is there some metric that you should be thinking about when you need that? I imagine the importance of the, or the risk assessment is kind of at the core of this, you know, talk through how to think about how much to invest in this kind of testing and red teaming and that kind of stuff. Yeah. So red teaming is a really important tool that we use, but we use it in a variety of ways. So I gave the example earlier of bringing together this expert who red team to figure out what GPT-4 could even do, right?
40:35So in that case, we were using red teaming for risk identification. And there it's kind of novel risk identification. Like we don't even know what the risk of this technology are. That's something that the model provider can do, something that needs to be done, you know, kind of once. So it's not something that every single, you know, model user needs to be doing as far as like, what is the technology even capable of, We publish a model card and list of these different risks and sort of guides about how to think about them. But then we also use red teaming as a sort of pre-ship validation. So in addition to running the automated testing that we talked about, sometimes we supplement that with red teaming because there we have experts who can maybe go deeper or identify something that the automated testing was missing.
41:33And we use these in combination. So for example, if we go and see the red team, the red team like gets through and finds a somewhat interesting vulnerability in the system, we then will go and automate or augment the automated testing so that we better understand how is that behaving overall at scale and also keep tracking it going forward. Now, we don't red team every application. Part of our commitments to the White House last year were that we will red team every sort of high risk application or sort of high impact application. And so if something is a very standard, like document Q &A chatbot, then we usually don't bring in Red Teaming for that.
42:20We focus on the automated evals. But if it's a really sort of like an application in a more sensitive domain, an application that's very like powerful or more general purpose, then we will use the Red Teaming. And so I would recommend that organizations sort of think about it similarly in terms of ongoing red team meetings where you have something that is more novel or high risk and you want to be really certain. But if it's kind of just a bread and butter application that you understand, then it's also reasonably that you can kind of build automated testing that tests that appropriately. It sounds like an important way that you approach this and that you recommend others approach this is aligning the resources and the testing and other aspects of the system or application delivery to the risk assessment.
43:15And I know you're a fan of the NIST AI risk management framework as a way to think about all this. Can you kind of introduce us to the framework and why you like it as a way of thinking about these various risks? The RMF has been around for a while, but I think they just recently released a generative AI extension. So I like it because I think it's a sort of simple, easy to understand framework. Microsoft, our response by standard is public. So you can go and check out, you know, sort of what are our requirements and how we do it. But the way our standard is organized, it's organized by principle.
43:53So here's what needs to be done for fairness. Here's what needs to be done for transparency. And that's one way to look at it. The NIST framework is more what needs to be done kind of lifecycle wise. And so I think sometimes it's easier for people to think about it in that way. So the first step is map. So that's risk identification. You need to map the risk. So the GPT-4 red teaming example is exactly that. We need to go map the risks that are sort of possible. The next step is measure. And that's really important because once you understand the risk, really the first thing you should be doing is figuring out how am I going to effectively measure these risks.
44:32And that could be through ongoing red teaming. But our first investment, as I was sharing with BingChat, was going and building the testing systems so we could start actually testing for these risks. And the testing systems allowed us to then do sort of the third step in this, which is manage and allowed us to design those layers of mitigations that I was talking about earlier because we had ongoing testing. So every new thing we wanted to experiment with, do we think this technique will work? Do we think this classifier will work well here, et cetera? All of that we had powered by testing so we could make data-driven decisions about if we were effectively managing the risk.
45:14And then the last part is really govern, right? You can use all of the techniques and go through the process, but you need governance in the organization to ensure that those steps are actually being taken. And that's a really, really important one, particularly in a fast moving space. Like what we've seen is the technology is changing so quickly. It's actually very hard to keep even like a large organization like Microsoft that's quite sophisticated in this up to date on what are all the latest risks? What are all the latest management techniques that they need to do? And so having that backing sort of governance, both structure, but then like operational approach that allows you to have safeguards in place and checks and balances to ensure that stuff isn't being missed as the playbook evolves really quickly has been something that's really important for us.
46:09And what's the implementation of that governance? Is it a leader, a committee, an organization? You mentioned playbook, a combination of all these, I imagine. How do you see it playing out with customers? Yeah, I think it is a combination of all of these, probably two sort of practical elements that have been very important for us, but we've also, I've seen it been really important for customers. First is that leadership level thing. So the thing that we were talking about, the testing, right, there's real trade-offs between quality and safety. That also plays out in a bigger scale, right? as you try to transform your organization with AI, you have to decide which types of risk you want to take.
46:59And it might not be a safety risk. It might be a risk for your business. It might be a risk in terms of the customer reaction, or it might be that it does change your security posture or something a little bit. And so having that committee sort of at the top that can kind of make those decisions in an informed way for the company is something that's been really important. The other element that we have implemented is a, and many, many customers as well, is a review, you know, release process and making sure that there's expert review sort of before an application is released. So we have teams of sort of risk assessors that go in and exactly sort of map the risk for an application, look to make sure appropriate testing has taken place according to those risks, and that there's appropriate mitigations in place aligned with those risks, and then have that whole sort of package reviewed to say this looks like it's actually sort of ready to be released.
48:04And so that's another element that's been really important. And of course, then, you know, culture and all of these things play into it as well. But those are the two pieces that we really see organizations kind of moving to implement to be able to do this successfully. You know, given how fast things are evolving, I imagine it is difficult to kind of predict the future and, you know, where things are headed. But are there, you know, things that either you're particularly excited about or, you know, based on your experience, how do you think about the future with regards to testing and safety for generative AI systems?
48:48Yeah, maybe I'll first say, actually, one of the things I'm the most excited about is recent history. So not the future, but we're still sort of exploiting what's possible here. One of the amazing things about generative AI is it's actually a huge step forward in terms of responsible AI or safety kind of toolkit. So I shared earlier that our testing system, for example, we have an automated scoring function that scores like our actual expert human linguist. This is something that previous versions of the technology could not do at all. And so we went from, we are not able to test very well. Like we're able to test only with humans, which meant that it could be only like outer loop testing where, okay, you've basically built the entire application.
49:46You're going to do one final like validation check that can go to, you know, human labelers and say, okay, it kind of looks good, right? where when we're able to build the automated version now because of this technology, we are able to turn that into inner loop testing, which means we can actually tune and improve and innovate on how much better we can get it responsibly in the application while we're developing the application rather than it being more like a final check. And so that's - Kind of the obvious quip there is like who's minding the minders, but I think you addressed that in this idea that you're automating the inner loop So you're able to catch issues quicker, but you still have the outer loop that is providing the same level of checks as you had before.
50:28Yeah, exactly. And we're still like the, I was saying that function we've, you know, had it score like the human house. We, you know, we check that regularly. Is it still scoring, you know, as well as those, right? So there's also kind of different places. There's also minding the minder. Yeah, yeah, yeah, exactly. And so that's like one example of how it's a breakthrough. Another one is a challenge with, let's say like hateful content. There's different stereotypes all over the world. There's different statements that would be hateful in one context and not in another. And it's very hard to pull together like all of that data.
51:09Well, the system is actually, the models have been trained on a lot of, you know, the data from kind of all over the internet. And so they are actually really, I'm going to say great in this sense, terrible in another sense, but really great at generating that type of content. And so we actually use it to produce a lot more synthetic data to augment our safety models. It's a feature, not a bug. Yeah, exactly. I mean, it's like, you know, tools and weapons, it's all how you use it. So the behavior I'm trying to prevent in most applications, which is generate harmful content, is one of the best things about these models for me, because I'm able to use that to build the safety system that allows us to make applications safe.
51:54But the fact that they understand language, they understand content, they understand nuance so much better makes them better for building safety classifiers from as well. And so we feel like we just have this huge step forward in the toolkit that's available to us. And that's unlocked so much innovation in responsible AI. And so that's the thing I'm the most excited about. And I think we're just at the very beginning of what's possible there. And so as the technology gets better, I think we'll also see more kind of be possible with that. So that's one of the things that in my everyday life, I'm very, very excited about with the technology.
52:31So does that mean I won't be able to get you to look into the crystal ball? You know, I think actually I really like something that Kevin Scott says to us frequently, which is, you know, the technology appears to be sort of on this exponential curve. But you only get kind of an update every two years or something. And so certainly when we, you know, we had three, five and then four came, we were kind of blown away. But every time you get each new data point, you kind of go back to like a linear existence. Like it's very hard to predict what would an exponential jump look like again. So we're, you know, back on our like linear life and then we'll see if the next technology is an exponential.
53:14And so I think I am prepared to be surprised, but I don't know in, you know, what directions, right? Some of what emerged with GPT-4 isn't what we expected and some, you know, is what we expected. And so I think that I'm in the keeping an open mind and waiting and seeing what will happen. But I'm pretty sure that regardless of what happens, it will be surprising in different ways. A way that you process that question about the future is asking, does GPT-5 change or obviate all of this? Or does it require incremental adaptation? And it sounds like you're not sure at all. Yeah, and that's a question we ask ourselves all the time, right?
54:01Okay, like, look, we're building all of these different prompt engineering tools. Are we going to be doing prompt engineering with GPT-5, or is it going to be so much smarter that you don't need all of that, right? I think it's very, very hard to predict those types of things. I feel confident that no matter what the next jump in technology is, we're still going to want to test it for different types of risks, right? We probably have to add some new ones to our standard playbook, but we're still going to have to have robust testing systems. I believe very much in the, let's say, the kind of the core idea of RAG, which is you bring the model together with contextual data, with fresh data, with relevant data.
54:42And so I think we're still going to have data retrieval or data augmentation in some form in future applications. And so all the work people are doing to make their data pipelines stronger, get the right data to the model, all of that's probably still going to be durable. But, for example, other things may not be durable. So one of the things happened with GPT, the jump between three and four, is the first application we built was GitHub Copilot. And we fine-tuned GPT-3 with the code data to make that possible. Otherwise, it would never have had the quality. When GPT-4 came out, it was way better at code than GPT-3.
55:23So in essence, like all of that fine-tuning work was thrown away. And it was totally worth it for Microsoft because we got that application to market sooner. We kind of, you know, made the market. We got a year of learning in the market with that. And so even though it was sort of throwaway in that sense, it was totally worth the investment. But like something I tell customers all the time is if that happens to you, if the next generation is so much better than your fine-tune thing, if you're going to feel that that was a waste, then you probably should not go invest in it right now because that's like a real risk that's out there.
55:56And so you have to make the decision for yourself, does it look like GitHub Copilot where you're still going to be happy with the outcome or not? But it's certainly a thing we're doing all the time right now is like, what's going to be durable? What's going to be not? And there's certainly places where it's very hard to predict and people are making a lot of predictions and it'll be interesting to see which ones are right. Absolutely. Well, Sarah, thanks so much for jumping on to chat a bit about what you've been up to and the way you think about AI safety. It has been wonderful catching up. Yeah, it's been great as always.
56:26Thank you so much for having me and looking forward to talking sometime again in the future.
From the publisher
Today, we're joined by Sarah Bird, chief product officer of responsible AI at Microsoft. We discuss the testing and evaluation techniques Microsoft applies to ensure safe deployment and use of generative AI, large language models, and image generation. In our conversation, we explore the unique risks and challenges presented by generative AI, the balance between fairness and security concerns, the application of adaptive and layered defense strategies for rapid response to unforeseen AI behaviors, the importance of automated AI safety testing and evaluation alongside human judgment, and the implementation of red teaming and governance. Sarah also shares learnings from Microsoft's ‘Tay’ and ‘Bing Chat’ incidents along with her thoughts on the rapidly evolving GenAI landscape.
The complete show notes for this episode can be found at https://twimlai.com/go/691.




