In short
Podcast Episode Notes: Practical AI - AI Incidents, Audits, and the Limits of Benchmarks
Episode Overview
- Title: AI Incidents, Audits, and the Limits of Benchmarks
- Description: The episode discusses the transition of AI from research to practical deployment and the real-world implications of AI incidents. Guests Sean McGregor, co-founder of the AI Verification & Evaluation Research Institute and founder of the AI Incident Database, talks with hosts Chris Benson and Daniel Whitenack about AI safety, verification, evaluation, and the shortcomings of standard benchmarks.
Key Participants
- Sean McGregor - Co-founder and lead research engineer at the AI Verification & Evaluation Research Institute
- Chris Benson - Principal AI Research Engineer at Lockheed Martin
- Daniel Whitenack - CEO at Prediction Guard
Episode Structure
- Introduction
- Hosts introduce the topic and guest, Sean McGregor.
- Discuss the importance of AI safety and the relevance of documentation and understanding AI incidents.
- Sean's Background and Journey
- Sean’s academic background in machine learning with a focus on reinforcement learning.
- Experience with energy-efficient neural network processors.
- The evolution of his work leading to the AI Incident Database.
- Defining AI Safety and Incidents
- Discussion on what constitutes an AI incident:
- Incident: An event where harm has occurred due to AI actions.
- Different terms like accident, adverse event, vulnerability, and their nuances.
- Importance of understanding safety in AI, particularly in high-stakes contexts.
- The AI Incident Database
- More than 5,000 human-annotated reports of AI incidents.
- Role of journalistic reporting in sourcing incidents.
- Future need for mandatory reporting to better understand AI risks.
- Third-Party Auditing of AI
- Challenges in verifying safety in general-purpose AI systems.
- Importance of third-party audits to assess AI systems accurately.
- The necessity of statistical backing for claims made by AI systems.
- Benchmarks vs. Audits
- Discussion on the difference between benchmarks and audits:
- Benchmarks are often produced for research, not practical applications.
- Audits verify the real-world applicability and safety of AI systems.
- Risks in AI and Emerging Issues
- Identification of various risk categories such as misuse, unintended behavior, and emergent social effects.
- Discussion on how emerging AI technologies create unique challenges.
- Practical Examples and Lessons Learned
- Insights from the DEF CON event where AI systems were tested against hackers.
- The need for rigorous, systematic evidence in identifying vulnerabilities.
- Future Directions in AI Safety
- Sean's vision for the future includes improving methodologies for risk assessment and safety.
- Emphasis on the necessity for institutions to invest in developing safer AI systems.
Key Takeaways
- AI Incidents: Understanding and documenting AI incidents is crucial for improving safety and preventing future occurrences.
- Third-Party Verification: Independent audits provide a necessary layer of scrutiny to ensure AI systems function safely in the real world.
- Shifts in Reporting: Transitioning from voluntary to mandatory incident reporting is essential for accountability and risk management.
- Benchmark Limitations: Relying solely on benchmarks can be misleading; practical evaluations are necessary for real-world applications.
- Risk Categories: The AI field faces diverse risks that require ongoing attention, particularly as systems become more complex and integrated into daily life.
Conclusion
The conversation emphasized the evolving landscape of AI safety and the critical need for rigorous evaluations to ensure that AI technologies are deployed responsibly. With more incidents occurring, the focus must shift towards learning from these events to build a safer AI ecosystem for the future.
Additional Resources
- [AI Verification & Evaluation Research Institute](https://www.averi.org/)
- [AI Incident Database](https://incidentdatabase.ai/)
- [Upcoming Webinars](https://practicalai.fm/webinars)
Community Engagement
Listeners are encouraged to connect with the Practical AI podcast community via LinkedIn, X (formerly Twitter), and Blue Sky for ongoing discussions and updates on AI developments.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOMeet the Hosts and Guest
0:45 to 4:30
Introduction of hosts Daniel Whitenack and Chris Benson, and guest Sean McGregor.
“Welcome to another episode of the Practical AI podcast.”
Understanding AI Incidents
4:30 to 9:00
Sean McGregor discusses the significance of AI incidents and the need for careful documentation.
“And during this whole time, I had a project that just kind of got away from me because it filled a niche.”
Defining Safety in AI
9:00 to 14:09
A discussion on the definition and implications of safety concerning AI systems.
“And so I guess like AI incidents have been happening for some time in the sense that like you could have a computer vision model detecting maybe it's problems and things coming off of a manufacturing line.”
Understanding AI Safety Challenges
14:09 to 17:08
Explore the complexities of verifying AI safety in general-purpose systems.
“Do you want to explain a little bit about kind of how that side of things might fit in the landscape and why it's important?”
Real-World AI Incidents and Implications
17:09 to 19:08
Discuss real-world AI incidents to highlight the unpredictability of AI behavior.
“And we've all been on development teams and pushed to prod, pushed to the real world, pushed to real world circumstances and found like, oh, I didn't think about that.”
The Importance of Audits in AI
19:09 to 24:48
Learn about the significance of audits in ensuring AI system reliability and accountability.
“and they're trying to put this in the context of their own operations.”
Benchmarks vs. Audits in AI Evaluation
24:49 to 27:28
Understand the differences between AI benchmarks and audits and their implications for AI deployment.
“we're starting to see a valuation ecosystem pop up and that being a discrete task and something that's supported by organizations separately.”
Hacking Insights from DEF CON
27:29 to 28:00
Delve into the experiences of testing AI models at a major hacker conference.
“and listen to that a lot and always love hearing stories of hackers and the kind of world that is there and all of what people try to do at DEF CON and other places.”
Introduction to DEFCON and AI Village
28:00 to 28:48
Learn about the history and significance of DEFCON as a hacker convention and its AI village.
“with a declaration that the model is flawless, and obviously certain things happened after that.”
Generative Red Team Challenge Details
28:48 to 30:42
Explore the generative red team challenge and its implications for AI model security.
“talking about how risk architectures would change everything.”
Show all 14 chapters
Establishing Rigor in Hacker Submissions
30:42 to 32:59
Discover the importance of rigorous data in evaluating AI model vulnerabilities.
“And that is actually a surprisingly difficult thing.”
Surprising Exploitation Modes Revealed
32:59 to 35:38
Uncover unexpected exploitation methods that emerged during the AI challenge.
“And then one time out of 100, they'll get something.”
Future of AI Security Practices
35:38 to 38:38
Consider what the future holds for security practices in AI model deployment.
“the configuration of that handoff was looser than it necessarily needed to be.”
Reflections on AI Incident Management
38:38 to 41:19
Reflect on the ongoing challenges and aspirations for safety in AI systems.
“What are you, what are those things that as you, you know, end your days or you're driving back to home or whatever, what are those things that stick in your mind?”
Transcript
Automatic transcript. May contain errors.0:03Welcome to the Practical AI Podcast, where we break down the real world applications of artificial intelligence and how it's shaping the way we live.
0:30behind-the-scenes content, and AI insights. You can learn more at practicalai.fm. Now, on to the show.
0:48Welcome to another episode of the Practical AI podcast. This is Daniel Whitenack. I am CEO at Prediction Guard, and I am joined as always by my co-host, Chris Benson, who is a principal AI research engineer at Lockheed Martin. How are you doing, Chris? Hey, doing great today, Daniel. How's it going? It's going really good. Lots of fun things in the news to follow and lots of fun things to work on. By the way, for our listeners who might need a reminder about this, we are doing some webinars recently on a variety of topics, some of which are maybe even related to some of the things we're talking about today around security or safety.
1:29So if you want to find out about those, go to practicalai.fm slash webinars. That's where we have some of those things listed out. But I'm really excited today because I have an amazing set of things to talk about with Sean McGregor, who is co-founder and lead research engineer at the AI Verification and Evaluation Research Institute and also the founder of the AI Incident Database. How are you doing, Sean? Thanks for joining us. I'm doing well. Thanks for having me here. Yeah, yeah, of course. And this is interesting to think about AI incidents, verification, evaluation. How did you find your way into a day-to-day where you're thinking about and documenting and studying AI incidents, among other things.
2:19How did I get to work on what you might call impractical AI? Yeah, well, I guess where practical AI turns into problematic AI, let's say. Yeah, and really practical AI is the AI that has consequences and matters in the world. And those are the ones you actually care to look into where it goes wrong. So the kind of really quickie professional tour of things is I'm about a 2017 vintage PhD in machine learning. I focused on reinforcement learning as applied to wildfire suppression policy. So fire starts in a forest, what do you do about it? And how does that impact the development of the land and the values we get from it over the course of 100 years?
3:04really in that setting I had a very strong sense of the power of the technology that we were developing but also the brittleness of it and just how difficult it was to even know whether the system was doing what you wanted to do particularly in a reinforcement learning system and I just never wanted to wander through a forest and have a sense of you know this is a charred wasteland because of the the forester took my simulation and didn't really I realized It was research code and that he should actually apply his own human expertise about what is good fire and bad fire and that. So I went from that.
3:41Then I worked on energy efficient neural network processors, which that was a great effort. The organization called Sentient that shipped when I was there millions and has continued to ship many edge neural network processors. And that gave me a really strong impression of the kind of power and brillness once again inside these consumer electronic devices, hearing aids and the like. And so I left that. I started a company dedicated to the Tufts and Evaluation of Machine Learning Systems. And when that happened, or when I started that, the kind of LLM explosion happened. And, you know, that's the kind of Cambrian explosion of sorts of, you know, practical AI technologies that I'm sure we're all still feeling.
4:34And we sold the assets associated with that to an old safety organization called Underwriters Laboratories and spent some time there doing safety stuff before leaving and starting up the AVERY, the AI Verification Evaluation Research Institute. And during this whole time, I had a project that just kind of got away from me because it filled a niche. It's served a need. And that's the Einstein database that you really need to have systems that collect and produce usable data sets, motivating safety practice. And you see this in aviation, a plane crashes, you record what happens and use that to make sure effectively in the aviation industry, you have a form of regression test.
5:24You don't want that pass crash to happen again. You also have similar things in food safety, you have medical adverse event reporting. this is the fundamental primitive that you need to make sure you have in safety. As a bad thing happens, you make sure it doesn't happen again. And so we've, in that project, collected more than 5 ,000 human annotated reports of AI incidents. Those are collected across more than 1 ,000 discrete incident records at this point. And we've formed a lot of the training data of what is an incident, and we've been interacting with a lot of intergovernmental organizations, and GRC community governance, risk and compliance community companies and like just trying to motivate what is, how to motivate safety and like why very often that can actually be a business imperative because a lot of the incidents that we have in the database, you could actually look at the stock price before and after and there's an impact.
6:23Could you talk a little bit more, kind of expand on what the nature of safety means in the context of AI? and because coming in as a neophyte on this, I don't know if there's kind of a standard definition. It seems to me there isn't, at least not that I'm aware of. So how do you define it? You mentioned kind of defining an incident. Could you define what you think an incident is and how these different concepts relate together in a way that those of us who are not part of that kind of world, that specific world can kind of adopt the ideas and understand what you're talking about, kind of give us a background on it?
7:00So this is a three or a four hour podcast, right? It's like the question, like used to always we would have the question like is something AI or machine learning? And like eventually you just learn that those words are completely meaningless in terms of the way people apply them. So yeah, we just need to treat them as like a pointer and you know, like the memory might be allocated differently or the pointer can point to different memory or maybe the memory swap. So like there's a lot of safety terms out there. And like I've actually had a poster for this at one point because this comes up so often.
7:38You could have incident, accident, adverse event. You have vulnerability, exposure, harm event, controversy, issue. Like each of these have like slightly different nuance to them, but you can kind of bring it back to the concept of you don't want a bad thing to happen and that you don't want that bad thing to produce a harm. You don't want someone to say like, I've, I've been impacted or some organization been impacted. And so the, you know, the intergovernmental definition of incident that developed was effectively an event that a harm has taken place. That's an incident. Your mileage may vary in different contexts and like different communities.
8:21There, there are different terms that resonate, But the reason that we went with incident to begin with is basically it's sufficiently vague while still meaning something. Because you can't say accident. Some of these have intention. You can't say exposure or compromise or, you know, the ones that are in the computer security world because some of them don't have intention. And so incident covers them all. And, you know, I could gesture and give you some Venn diagrams. But this is a podcast. So we'll avoid that. I'm wondering, like in particular, when you narrow it down to AI incident, obviously sort of AI, depending on your definition of AI, kind of has existed for some time.
9:08And so I guess like AI incidents have been happening for some time in the sense that like you could have a computer vision model detecting maybe it's problems and things coming off of a manufacturing line. It misses one and that causes harm maybe downstream to someone who uses a product or something like that. How has this maybe, you know, you talked about the shift to, of course, the kind of expansion of what, you know, how AI is impacting our life. How, from your perspective, has that idea, I guess, been stressed in new ways in recent years? We have constant debates over what we would ingest into the AI and SYN database.
9:53and we because we have to operate and it's a you know database that we have to make these decisions day in day out have very difficult decisions and it's not just in our case it's not just what should be considered an incident it's also if we start ingesting these are we lassoing an infinity that's just going to like pull us in into some extreme direction and maybe we would like to index, catalog, and present to the world that infinity, but there are these kind of like little harms that are repeated maybe millions of times a day when you have these systems that indexing all of them isn't necessarily useful.
10:36So we do have a little bit of a bent towards, is this informing the production of a safer AI or safer world for AI? but a lot of things meet the fundamental criteria of involving AI and minor harms have taken place. There are also increasing number of high-scaled harms that take place. Like if you make a billion people slightly more depressed, non-zero number of people probably have died as a result of that. And that's where, you know, it's not just you're building a bridge that needs to bear weight. You're building a bridge that needs to bear the weight of all humanity flowing over it. I'm curious, as you're kind of selecting things to go into the database, you know, per the criteria that you just talked about, how are you sourcing that?
11:29And especially when you consider the fact that so many AI things today are proprietary and kept secret in organizations for obvious reasons. How is that sourced and how does the different sourcing affect the utility of the database and how you make those kind of selections? Most of what we have at the moment is journalistic reporting. The reason that we are living in that space at the moment is journalists actually put in a tremendous amount of work to validate the base facts involved in it. And that is where the degradation of the journalistic community, or at least the ability for them to make wages, has actually been quite harmful.
12:11We also do get direct reporting at times. We do get people that email us or submit direct to our forms. We get people that create blog posts and things like that and submit it to the Instant database. We do encourage people to submit to the incident database. The simple matter of fact, though, is the volume that we're dealing with is one where eventually we do need to switch from this kind of voluntary reporting to more mandatory form of reporting. In the EU Code of Practice, there is a requirement for severe incidents to be reported. It hasn't reached implementation as of yet. But there's a lot of back and forth on that front.
12:53But the insight you can derive from mandatory reporting is greater than that of this voluntary reporting. Like we can prove existence. We can prove it's happening. It's very difficult for us to assign a rate to incident events, though, particularly as it becomes non-newsworthy through time. So we do have some works that we're trying to make the most of that. Uh, this bears some resemblance to public health practices, um, where you don't have perfect insight into disease spread, but you do have these kind of indirect measures of, um, doing tests in the sewers and things to see if the, uh, the, uh, how much of the viral particles are there and whatnot.
13:33Well, Sean, I appreciate you taking us through this kind of idea of AI incidents. Obviously, we are practical AI and in a practical way, we would very much like to prevent incidents or at least understand kind of the security implications of using certain AI models. I know with the AVERY, the AI Verification and Evaluation Research Institute, in my looking at that, you're talking a lot about kind of third party auditing of AI or kind of frontier AI or foundational AI, however you want to frame that. Do you want to explain a little bit about kind of how that side of things might fit in the landscape and why it's important?
14:18Certainly. So a fundamental problem that we have, particularly at the frontier model level with things like OpenAI, Anthropic, Google's Gemini and so forth, is it's very difficult to even know how safe your systems are because they're general purpose systems. This basically broke the safety frame. All the safety processes that we have are built around starting presumption of there's a specific context it's operating in and you reason about its safety within that context. Well, if your context is just wildcard star everything, where does that leave you? Do you need to verify across all circumstances that it's going to be safe?
14:58Because the answer is going to be no. So how do you approach the task of verifying claims? How do you encapsulate something that a customer or person or organization would rely upon and say like, okay, this has gotten this level of assurance applied to it. So I know I can apply it within my college age student population, teaching them how to do the maths. And getting that level of signal at this point, you know, a real practical problem that you have is no one's going to believe you when you say, you know, this is good for college age math remedial level. Like you have to run a pilot program and see how it works before you actually can believe the representations being made by companies because they probably haven't evaluated in your exact circumstances.
15:57They might have something that's analogous to that, but they probably won't have it exactly in your circumstances. And, you know, this is a problem on the practical AI side of things, but it's also a problem on just the increasing power of these systems and the scale in which they operate. You actually do want some really strong top-level guarantees that the system is going to try and steer away from catastrophe and doesn't actually have an active propensity for doing the bad things. And the science of establishing that is actually still a work in progress. There's a lot of work, not just in establishing evaluation, but meta evaluation, basically deciding, is this evaluation saying the thing that it says it says or not?
16:47And Avery, as an organization, is concerned with doing third-party audits of basically saying, is this thing safe? Is this thing doing the thing it should be doing? and there's a premise in there that a third party is better positioned to do that than a first party. And we've all been on development teams and pushed to prod, pushed to the real world, pushed to real world circumstances and found like, oh, I didn't think about that. Like one of my favorite incidents in the AI incident database is this gentleman who got a traffic citation mailed to him. And he looked at the traffic citation and said, like, you know, this isn't my car.
17:38It not only isn't a car, it is a woman who's wearing a shirt that says knitter. And there's a purse strap going across it. So it like kind of played with the letters. So it looked like K-N-I-9-T-E-R or something like that. and I don't think the people making the traffic camera were thinking about a woman walking through the field of view of the traffic camera uh with like a shirt perturbed in such a way that uh it would catch it and register it as a license plate and then the traffic station went out like this the world is hard uh the real world is real hard I'm curious that I like that example it's a it's a very personal example.
18:23And I can certainly imagine that happening to me actually, bizarre as it is. But I'm curious as you, as you are looking at applying this process, you know, that you're describing to maybe a larger incident, is there an incident that comes to mind in the database that where, where you could kind of compare what your estimate, like if relative to not having gone through third-party process versus if it had, how you might envision that playing out in a different way. I'm kind of trying to get a sense, and you can interpret it any way you want or find whatever example, but trying to get a sense like when organizations are listening, you know, they have employees and their leadership are listening to the podcast, and they're trying to put this in the context of their own operations.
19:14Like, what do you think is a kind of an aha moment, potentially, for the listener. And maybe a Fortune 500, maybe not that big, you know, maybe a midsize company where this process will make a substantial difference to the outcomes involved. Probably a good way to work here is by analogy to other segments that have audit. And audit is something that no one truly enjoys. Like, you don't necessarily want to be audited. But you actually do. You do want to be audited if you're an organization that wants to receive investments from other organizations. Having audited financials is table stakes. You're not going to put your money into an org that hasn't been audited.
20:00Or actually, I should say, there are famous instances in which an org has not been audited and people subsequently regretted it because the reason they weren't audited is they were improving disbursements by Slack emojis. And, you know, that it didn't go very well in those circumstances for anyone. And so we do periodically get reminders that, you know, these processes are important and that audits solve a problem. And that's, you need to be able to trust the fundamental representations about the state of an organization. And it's the same thing for the model, because the model is similar in a sense that it's taking actions, it's doing things and has impacts.
20:42And so there's value to knowing when you should or should not trust information. And I'm looking through some of what the Avery Institute has been involved with, and in particular, kind of recent work of this meta-evaluation of benchmarks, this sort of bench risk, which is actually how we got connected through a colleague, Ashoria. And I'm wondering, this is like a meta evaluation. And I know a lot of people make decisions about models based on benchmarks. Could you maybe help us kind of walk through mentally, what is the difference between maybe like an audit and a benchmark and maybe like tying in some things that might make people think about whether benchmarks as they look at them on leaderboards and that sort of thing are really a, yeah, if they're really relevant to the real world behavior of models.
21:46So maybe overextending the financial metaphor on things, when you audit a financial organization, you say, okay, I see that this organization has a balance sheet indicating they have$10 million in the bank. You're like, great. Now as an auditor, I need to go and check the balance with the bank and see, is this money real? Is it actually there? And similarly, an audit is concerned with, I've seen the balance, I've seen the balance sheet, and now I need to check the evidence and make sure that I actually believe that representation in some form. And so where meta-evaluation or evaluating the evaluations or checking the benchmarks, evaluating the benchmarks, are useful is you're basically checking those receipts.
22:33And in the course of the Ventress project, we found that a lot of the receipts were just kind of like, lol, trust me, bro, like written on a piece of paper and that there were real substantive issues of like, okay, maybe there's like an IOU and there's some gold buried out in a field somewhere, but I have some doubts. And at the very least, I need to look into this a little bit more before I rely on it for real-world purposes. And the dichotomy that we identified is a lot of the benchmarks that are produced are produced for non-practical purposes. They're produced for knowledge generation purposes.
23:14They're produced for research purposes, where people are wanting to understand systems that are making sense of it, but it's not produced with the intent that someone's going to then and say, all right, I'm going to deploy this in my environment. And I know now that it's unbiased because it scored well on BBQ. And BBQ is a great benchmark. It was really foundational in the bias benchmarking community. But it's also not produced with the intention that every time a frontier model is released, you say, like, look, we've improved on bias. and it's 10 BBQ points better than the prior ones. And that's a good thing.
23:57You do want that, but it's not a practical AI purpose. It's not saying it's unbiased for my specific application because there's a distribution associated with BBQ. There's characteristics that aren't necessarily generalizable to the environment to which you're deploying it. The prompts can be in a very particular space for BBQ that you're just not operating in. And all these things are associated with, in this research work that we put out, they're associated with failure modes that we identified. We collected a list of failure modes, and then we looked at how well benchmarks had expressed mitigations against those failure modes.
24:38And the results were mixed, like some did better than others, but by and large, most benchmarks historically have been produced for research purposes, not for practical AI purposes. and this is a problem because this is what we're working with and informationally and this is why we're starting to see a valuation ecosystem pop up and that being a discrete task and something that's supported by organizations separately. I'm wondering as we're kind of going through this and discussing the process if as you're looking at some of the different risk categories that you guys, you know, outline on your website and stuff.
25:17And, you know, a couple of examples are misuse and unintended behavior and infosec and emergent social effects. Which of these do you think is kind of like maybe the most underestimated? You know, the thing that people aren't really expecting that has consequence that may be outsized relative to what those of us not in this field would expect. Do you have any sense of that in terms of like this kind of a little bit of the surprise that people might not realize is there? That's tough because I'm a little bit of the fish in water being told it's wet. Fair enough. I will say in like kind of rewinding my own experience on this is I feel like I'm a pretty astute observer and I've been more predictive of what problems are coming around the corner than I think most people I've encountered.
26:07But even with that, I still am regularly surprised. And I wish we could be a little bit more predictive about it. And I think we are going to get there and we're going to learn from the past and hopefully push that into the future. But this is a new thing. We don't have a great operating history, particularly for the general purpose AI systems. And I think that the split in this community that you should probably think a little bit about that question in terms of expectation is there's security and safety people and that there is a little bit of difference between those. Security is safety. Safety is security.
26:46But what you care about tends to be slightly different in what you're expecting. And security people have a much greater capacity for the presumption of bad actors doing bad things and that producing bad outcomes. And then the safety people have a much stronger presumption of just the world is its own adversary. We're all in our own kind of game against nature and bad things will happen regardless. and maybe in the bad actor side of things, they're a little bit more insistent upon it though, particularly when the scalability of doing bad things is actually substantially increasing. So we have to solve all of these.
27:23Well, Sean, one of the things I wanted to ask you about, well, in particular, I'm a big fan of the Darknet Diaries podcast and listen to that a lot and always love hearing stories of hackers and the kind of world that is there and all of what people try to do at DEF CON and other places. I was really intrigued by your study titled To Err as AI. I see in the description you kind of took some models to DEF CON, and I'm reading your description, had the pleasure of taunting a room of hackers with a declaration that the model is flawless, and obviously certain things happened after that. So I would love to hear, I mean, maybe how this came about or what the idea really was.
28:16And then would love to hear some of that story. Sure. So to set the stage a little bit, DEFCON is a hacker convention conference that's been run in Las Vegas for a good number of years. And it's, I don't know if listeners have seen the 1990s movie Hackers, which has like a very early, like Anjali Jolie and some others in it. And it's like these young men and women like rollerblading everywhere, you know, like that's just how the future would be. And I'm talking about how risk architectures would change everything. And there's this kind of hacker aesthetic that kind of noisily was portraying of, you know, joyfully compromising systems.
29:01And in this particular instance, having a mixture of people that are doing that with good aims and bad aims. But DEF CON is this community organized thing and you have villages. And one of them is the AI village that's been run for a few years now. And in that village, you have presentations, but there's also a section of it dedicated to some sort of a challenge that's conducted over the course of the conference. And people filter through and figure out how to break things and they're treated to cash prizes in this case. So this was something called the generative red team two. So it was a change from the prior year and form.
Read the full transcript
29:49And we basically asked people, there's documentation for the open language model produced by the Allen Institute for AI. So a 7 billion parameter, large language model, along with its guard model put in front of it. and we said, here's the documentation, here's the representations about what this model is supposed to do and your job is to show how that's wrong. That kind of unleashed a few days of chaos and I don't know what else was happening in the conference, but I was at the judging table for the entire time with reports flying in and needing to decide is this actually a violation of the documentation?
30:28It did the thing we wrote, which we did put some Easter eggs in there of like we expected people to be able to break it. But did they submit sufficient evidence of a violation of that documentation? And that is actually a surprisingly difficult thing. We had all these people that were very good at the compromising systems and saying like, yeah, I know how to prompt this LLM and get it to tell me how to burn down this convention center violation of what the model should do. And we kept on needing to say, oh, that's great, but you haven't told us anything. Anecdote does not equal data in this instance, and we need you to show that it's systematically pushing towards arson, that it's always going to burn things down, or that you can put some sort of more universal jailbreak on it that makes it so that it underperforms in that particular filter of things.
31:24And so what we basically had to bring to the security world, where the security world does not like this, and I don't blame them for it, is statistics. We had to say, it's not just about the fact that there's an exploit. It needs to be it's systematically vulnerable in some form because it's always rolling dice. That's just the nature of it. And that made for a wild few days, paid out a good cash purse and learned quite a lot about the collision of disciplines that we need to foster when it comes to security, safety, statistics, machine learning, the world's getting more complicated. Prior to kind of talking about some of the chaos that ensued at that point, as you talk about that collision of culture, if you will, you know, with kind of taking the hacker out of their thing of saying, hey, I have an exploit, you know, look at me, and requiring that level of rigor.
32:27First of all, Can you describe why you required that, why that level of rigor was important in this case, and how did that force different behaviors out of the hacker community that were applying themselves against the problem that you had presented? We had a few people that were unhappy initially, and then we explained it, and then they went back and came back. And then they started turning the crank and getting a lot of payouts. So people that really wanted to encamp and figured out the modality did very well. The problem and the reason why anecdote doesn't equal data here is if you say something is 99%, filters out 99 % of the bad thing, if they wanted to, they could roll up to one of those stations and just keep on issuing the same query or adding one period to it, dot, dot, dot, dot.
33:16And then one time out of 100, they'll get something. They'll be able to walk up and say, money, please. The problem is that's not really useful. If you're designing the system, You need some idea of what is the systemization here because you know it's 99%. To some extent, you're going to keep on working to get to 100, but you care about the cases where it's not 99%, but you made a mistake and it's actually 70 % or it's 1 % or 0%. And that's a statistical argument. That's a, here's a description, a higher level description, then here's an example. Like this isn't me prompting it. This is a description of a attack that is not accounted for in your documentation that you're vulnerable for.
34:01And you need to solve this as a strategy. Like you, you never solve talk like a pirate. So anyone can talk like a pirate to the system and get it to do bad things. And that's the thing you care about rather than I asked it and pirate ease something a hundred times in one time it gave me bad. Were there certain modalities like you described of failure that were a surprise kind of to the judges in terms of like the or maybe to like the like, I don't know how much interaction there was with the model building team, you know, at Allen Institute and that sort of thing. But, you know, any any kind of modes of failure that that weren't expected or just the ones that were continuously problematic, but even if they were known before.
34:47Somewhat unexpected. We characterize this in the research paper as well. And I do recommend people take a look at it because it's very practical in nature. The biggest one and the one that I think most people just instantaneously exploited and they didn't have full information on this one. We didn't give them full information on the integration of the guard model and the foundation model or the chat model underneath. and there's multiple ways of configuring that handoff. You can either do like a hard reject and say like this doesn't pass the guard model and we're going to give you nothing or you can reprompt the underlying model and say like, here's what you're being prompted but don't answer it.
35:29There's variations of that and then you can still get like a somewhat useful reply out of the system. It just should reject it. the configuration of that handoff was looser than it necessarily needed to be. And exploiting that handoff between the guard model and the underlying model was the kind of big vector that you could use that to print cash over the course of the competition. And that's perhaps the thing people need to think about, particularly as deploying these systems are very often a collection of multiple systems that present as one, is the interface between them is very often under-tested.
36:10It's very often difficult to know how those systems will interact. And particularly when the benchmarks are expressed at a lower level than the whole system, you don't know a lot about what will happen at that point. You've muddied the waters a little bit. I'm curious, as you went through this exercise with the attendees, beyond just the specifics of this problem, what was the value of going through the exercise and what you know in terms of the outcome for you and what you got out of it beyond the immediate guard model with the you know the model it's protecting um in that process like what was that like and i'm also curious if you go and do a similar exercise in the future which i'm assuming there will be some some variation of that what can you envision being a good next step to tackle in terms of trying to secure value and get to a next level outcome?
37:09You know, have you thought about what you would do as you came out of the previous event? Should you do it again? So I think one of the really strong takeaways, and there's actually been some work since then, was the need for tools and to kind of scaffold this in some form and to make it where you have a clear form of expression of what these things are called, which is flaw reports, because it's not an incident, like no one's harmed, you're kind of, you know, testing things in a laboratory. And so what you have in classical computer security is you have bug bounty systems, you have the ability to submit those, there's kind of clear adjudication processes that often lead to payouts to people.
37:52and the extensions that are necessary for machine learning-based systems, data-driven systems, distributional systems is, I think, clearer as a result of the activity that we ran at DEF CON. Like that's the kind of research value that's produced by that. L Institute probably got some operational use as well that they learned a lot because they're at that judge table dealing with the onslaught as well. That was the vendor table for things. And so I think if this is run again in the future, getting past those pain points of, you know, you're kind of rolling your own code and all that. and that's not necessary for this, scaffolding it a lot more will carry it a lot more distance and then move this towards something where there's one or more companies that are operating a flaw reporting business in the same way that there is security bug reporting.
38:57Well, Sean, as we kind of get close to the wrapping up point here on the podcast, I'd love for you to share anything that comes to mind in terms of like, as you look out kind of at the next phase of maybe your own work, but maybe, you know, the ecosystem a little bit more broad in terms of the area of research in which you are involved. What are you, what are those things that as you, you know, end your days or you're driving back to home or whatever, what are those things that stick in your mind? Like you're excited, like what if this happens? Or I'm really interested to see how this develops, or I'm encouraged by kind of this going into the future.
39:37What's kind of top of mind for you in terms of the community that you're involved with coming into this next year? There's an interesting phenomena that every time we pass a milestone in the incident database and we're like, oh, we've crossed a 500 or a thousand incidents. We're like, yay. And then we're like, but it's not a great thing. Should you celebrate? Yeah. Should we celebrate? And I feel like my dream utopia might be I need to go do something else because the safety problem has been solved and there's no more to do. The unfortunate fact though is this is only going to get more complicated and there's more to do.
40:21There's more to solve. There's not going to be a perfect AI system out there. So what I'm hoping for in the next year is to develop a greater sense of what institutions, what techniques, what methods we need to develop and scale and apply so that we can have a safer AI ecosystem. And this is important, not just from the perspective of wanting to prevent harm, but you can't deploy unsafe systems to clients that want safety, that care about outcomes. And so our ability to make a safer system is very heavily involved in our ability to ship product to solve problems in the real world. And so I hope to highlight that.
41:22I hope to have a greater accounting for the risks and make it so that we have better and better measures of those risks because you manage what you measure. And so let's measure these risks and then we can separate the actors that are investing in safety and stronger AI systems, safer AI systems from the ones that aren't doing that. And right now that's not where we are. Well, really appreciate that perspective and appreciate the work that you continue to be engaged in. Sean, thank you for that. And from the community, it's really, really important. So appreciate that. And thank you very much for taking time to chat with us about it.
42:04It's been great. It's been a lot of fun. Thank you for having me.
42:15All right. That's our show for this week. If you haven't checked out our website, head to practicalai.fm and be sure to connect with us on LinkedIn, X or Blue Sky. You'll see us posting insights related to the latest AI developments, and we would love for you to join the conversation. Thanks to our partner, Prediction Guard, for providing operational support for the show. Check them out at predictionguard.com. Also, thanks to Breakmaster Cylinder for the beats and to you for listening. That's all for now, but you'll hear from us again next week.
42:49Question from The
From the publisher
AI is moving fast from research to real-world deployment, and when things go wrong, the consequences are no longer hypothetical. In this episode, Sean McGregor, co-founder of the AI Verification & Evaluation Research Institute and also the founder of the AI Incident Database, joins Chris and Dan to discuss AI safety, verification, evaluation, and auditing. They explore why benchmarks often fall short, what red-teaming at DEF CON reveals about machine learning risks, and how organizations can better assess and manage AI systems in practice.
Featuring:
- Sean McGregor– LinkedIn
- Chris Benson – Website, LinkedIn, Bluesky, GitHub, X
- Daniel Whitenack – Website, GitHub, X
Links:
- AI Verification & Evaluation Research Institute
- AI Incident Database
- 38th convening of IAAI
- BenchRisk
- State of Global AI Incident Reporting
Upcoming Events:
- Register for upcoming webinars here!




