Assessing the Risks of Open AI Models with Sayash Kapoor - #675

11 Mar 2024 · 40 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The TWIML AI Podcast - Episode #675: Assessing the Risks of Open AI Models with Sayash Kapoor

Episode Overview In this episode of The TWIML AI Podcast, host Sam Charrington interviews Sayash Kapoor, a Ph.D. student at Princeton University, who discusses his paper titled "On the Societal Impact of Open Foundation Models." The conversation explores the risks and benefits of releasing open model weights, the need for a framework to assess AI risks, and specific examples such as biosecurity and non-consensual intimate imagery.

Key Participants

  • Host: Sam Charrington
  • Guest: Sayash Kapoor, Ph.D. student, Princeton University

Episode Highlights

Background of Sayash Kapoor

  • Initially focused on theoretical machine learning, later pivoted to ethical AI.
  • Worked at various institutions, including IIT Kanpur, Columbia University, and Meta.
  • Currently a researcher in the Center for Information Technology Policy at Princeton.

Motivation for the Paper

  • The paper addresses both the benefits and risks of openness in AI models.
  • Highlights the lack of common ground and clear frameworks in ongoing debates about AI risks.
  • Aims to create a structured approach to discuss and assess risks associated with open foundation models.

Definition of Open Foundation Models

  • Emphasizes that "open" refers specifically to model weights being freely available, rather than being open source, which includes code and documentation.
  • The authors aim to clarify what constitutes open foundation models to foster better dialogue in the AI community.

Framework for Assessing Risks

  • The framework aims to evaluate marginal risks compared to existing technologies and closed models.
  • Utilizes a six-step risk assessment process inspired by cybersecurity threat modeling, focusing on:
  • Threat identification
  • Existing level of risk
  • Existing defenses
  • Marginal risk of open foundation models
  • Ease of defenses
  • Uncertainty and assumptions

Application of the Framework

  • Discusses two case studies:
  • Biosecurity Risks: Concerns about open models enabling malicious activities (e.g., bioweapons).
  • Non-Consensual Intimate Imagery (NCII): Highlights the increase in NCII generation due to open models, such as Stable Diffusion.

Marginal Risk Analysis

  • The analysis indicates that the risks associated with open foundation models, particularly in the context of NCII, are significant.
  • The panel discusses how interventions at different levels of the risk pipeline can help mitigate these risks.

Open Versus Closed Models

  • Emphasizes that the debate should focus on the space for both open and closed models, rather than framing it as an exclusive competition between the two.
  • Advocates for continued openness to promote safety-critical research and foster innovation in AI.

Future Directions and Policy Implications

  • Suggests applying the risk assessment framework to other domains such as disinformation and cybersecurity.
  • Urges policymakers to consider the framework in discussions about regulating open foundation models, noting that there is minimal evidence for high marginal risks.

Conclusion The episode concludes with a discussion on the importance of creating a safe harbor for researchers investigating AI risks and the need for open dialogue in addressing the challenges posed by AI. Sayash Kapoor encourages continued exploration and rigorous analysis of the societal impacts of AI to ensure responsible development and deployment of technology.

Key Takeaways

  • Openness in AI can drive crucial safety research but requires a balanced understanding of risks.
  • The risk assessment framework provides a structured approach to evaluate the societal impacts of open foundation models.
  • Policymaking around AI should prioritize evidence-based assessments of risks rather than blanket restrictions.

For complete show notes, visit [TWIML AI Podcast Episode #675](https://twimlai.com/go/675).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:03All right, everyone. Welcome to another episode of the TwiML AI podcast. I am your host Sam Charrington. Today, I'm joined by Sayesh Kapoor. Sayesh is a PhD candidate in the Department of Computer Science, as well as a researcher in the Center for Information Technology Policy at Princeton University. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Sayesh, welcome to the podcast. Thank you so much for having me. I'm really excited to dig into our conversation. We'll be talking about a paper you recently published on the societal impact of open foundation models.

0:39Before we do, I'd love to have you share a little bit about your background and how you came to work in the field. So my background is in computer science, and I started off my research career working in theoretical machine learning. From there, I sort of quickly pivoted to investigating fairness and machine learning algorithms, and then more broadly to the societal impact of AI. Over the course of this research career, I've worked at IIT Kanpur for my undergrad, then did research stints at Columbia University and EPFL in Switzerland. I worked at Meta for two years in their community integrity team, building machine learning in order to curb some of the harms that we're talking about.

1:15And then I began my PhD at Princeton University. With regards to this paper, it's interesting timing, I think, at least in the context of the podcast. We recently published an interview on Ulmo, which, as you well know and mentioned in the paper is an open model that is really seeking to promote openness strongly. And the paper also reminds me that when GPT-2 came out, we held a big debate as a kind of a webcast on whether open AI should make this available, should it be closed, and should there be very limited access to it. We've since come a long way and models much more powerful than GP2 are readily available.

2:00And what this paper seeks to do is to really ask the question or a set of questions around that fundamental openness that is to a significant degree open to debate. I'd love to start by having you share a little bit about your motivations for the paper and how it came about. So I guess the main impetus for the paper was we were in all of these conversations where we were talking about the benefits of openness on the one hand, and then in other separate conversations when we were diving into the risks of openness. And in particular, there was a lot of attention to papers that claim these catastrophic risks from language models.

2:40So there were a couple of studies that came out of MIT last year that pointed to the biosecurity risks that having access to an open language model could allow malicious users to create bioweapons. So I think in the midst of all of these conversations, one thing that stood out was just the lack of a common ground when it came to even talking about these risks. So what is this sort of risk we're talking about? What is it being compared to? Is it compared to like the risk of, you know, someone finding out how to create a bioweapon on the internet? And if so, how do we go about conducting these risk assessments?

3:12So in this almost fractured debate, all of us, like me and Rishi and some of the other authors, felt that it was necessary to come up with a framework that first removes a lot of the misconceptions about openness and what risks it enables. And second, it also opens the floor for more constructive debate going forward on what the risks are and how we should band together and collectively address these risks. One of the things that struck me about the paper is the diversity of authors that are named in the paper. You know, several folks have been interviewed on the podcast before from a variety of organizations, academic and non-academic.

3:53I'd love to have you share a little bit about how that came about. Absolutely. So I think the origin story of the paper is that last September, we organized a workshop called the Responsible and Open Foundation Model Workshop. So this was a joint effort by Princeton and Stanford. And as a part of that, we really brought together like a wide range of perspectives. But all of these had one thing in common. They thought deeply about the impact of openness on society. So we looked at perspectives from the industry, from academia, from civil society organizations. And following that workshop, it became clear that there is a need to write something more cohesive about the impacts of openness.

4:32And it really became clear that a lot of the evidence that we were sort of talking about was not very well substantiated, or a lot of the claims did not really have the bar for empirical evidence that we would have liked. And so really, like in some sense, the paper is a follow up to that workshop, wherein a lot of these experts came together and tried to set some common ground for how we should discuss the merits and benefits and risks of openness. And one sort of other type of organization that's sort of referenced in the author list is policymakers. So I think people who are trying to influence policy or influence how policy around open foundation model plays out in the real world.

5:12And that's because apart from all of the debate that's happening within the AI community or within the research community, policymakers are also very actively considering the question of openness. So we have ongoing policy proposals in the US, the UK and the EU specifically looking at the impact of open foundation models. In some cases, it's to impose stricter barriers, such as in the White House executive order. In others, it is to provide carve outs, for example, in the EU. But I think having this framework to assess the societal impact of foundation models can really help these policy conversations as well.

5:47So many of the people who wrote the paper with us, we were very happy to have the policy expertise on board. One of the very first things you do in the paper is define what you mean by an open foundation model. And you define that in terms of the weights being freely available. I'm curious the degree to which you think that very definition is contentious or has nuances or is worthy of discussion in and of itself? Absolutely. I think the definition of what constitutes open foundation models or even open source foundation models, I think is one of the most contentious debates that's going on today.

6:23The reason we chose the specific definition and the words we did was first, we did not want to call these models open source foundation models because open source implies something completely different about an artifact. It means that the code and to some extent the data used to train the model is made available. You have the documentation to reproduce the data. There is no use restrictions on the model at all. So that's why we deliberately chose not to use the word open source, but rather just open foundation models. A second part of it was, you might've seen the author list going back to that for a second.

6:55The executive director of the open source initiative, Stefano Maffoli, is also one of the authors. He's actually leading this charge of the open source initiative, defining what open source really means for AI. So to some extent, the term open source AI is not even well defined right now because the open source initiative does not really have a definition of that. So both of those meant that, you know, open source is out of the window. Now, when it comes to the question of like the risks and benefits of models, we chose to focus on model weights because that's where a lot of the perpetrated risks of these models come from.

7:29So while for many of the benefits of openness, you might need access to the code and the data. And projects like Olmo and Pithya from Eleuther.ai are like, you know, standout examples of how far you can take openness with checkpoints and model weights and data. I think for a lot of the risks, all you need is access to the model weights alone. And I think this also comes through in last year's executive order from the White House, which focuses on foundation models with widely available model weights. So I think we really wanted to hone in on this specific question because once the model weights are released, to some extent, this decision is irreversible.

8:07And this fact about releasing the model weights openly has caused a lot of concern about releasing open foundation models. And so that's why we chose to stick with foundation model weights that are widely available and with the term open foundation models, but not open source foundation models. Many of the concerns that are referenced in the paper are still concerns in a world where the waits aren't released, but you do reference the idea that with the models behind an API of sorts, there can be monitoring, there can be services can shut down particular uses that are abusive. And to some degree, you kind of call out that once the waits are out, it's a bit of a Pandora's box being opened, whereas behind a service, there's still some optionality.

8:56Was that a core idea in the way you thought about this idea of openness? So I think, again, going back to the concerns that people have had with openness, a lot of these are tied to the fact that once foundation model weights are out there, basically anyone who can download these models off the internet can do with them whatever they please. So there are some mechanisms you might envision for reducing the harmful impacts. So for instance, if a model has been known to cause harm, you might get it removed off of Hugging Face and GitHub so that the model weights aren't hosted. But to some extent, a lot of these interventions can be easily circumvented, for instance, through model weights being uploaded to Torrent websites.

9:35One of the reasons we focused on the concept of model weights was precisely because the risks around models are at least theoretically amplified when you have no take backs, when you cannot take back the model weights or control who uses them or for what purposes. One of the things that is very clear when you read the paper is that you're really trying to create a framework for folks to communicate about these various risks. And you get the picture very quickly that you've been in conversations where people are talking about risks, but they're really talking about, you know, very different types of risks and they aren't able to reconcile how to bring them all together.

10:16How did you conceive of a way to reconcile that? So I guess one of the things that was most helpful was seeing how two papers that came out of MIT last year dealt with concerns around biosecurity risks. In a follow-up work, though, I think a couple of months after this paper was released, a group of researchers from Stanford looked at what this information was that the authors were arguing would lead to future pandemics being caused by language models. And it turns out that all of the information that was at stake here was easily available on Wikipedia. We can have a lot of harmful outputs from open language models, but if we can get the same exact information from web search on the internet, perhaps looking at curbing the release of open foundation models is not the right policy proposal.

11:06So this really informed us on the need for creating this common ground, because if we don't have this common ground, it'll essentially be researchers on the one hand pointing out all of the risks of AI and on the other hand pointing out how they're just equivalent to past releases or basically equivalent to information widely available on the Internet. The results of that second paper are fairly intuitive to anyone who's familiar with language models and the idea that to some degree they are pulling together information that's already in their training data sets. Why do you think there was so much contention around that idea?

11:42So I think one of the reasons was that for a long time, language models just did not work very well. So if in 2015, someone had come up and said, you know, language models can help someone cause biosecurity issues, I think that concern would not be taken seriously at all because the state of progress of language modeling was such that the state-of-the-art language models essentially could not offer any help. I think that has sort of changed in the last five years or so. And especially with instruction-tuned models, it has become easier for people to rely on language models for assistance. To be clear, I'm not saying that the paper's results are flawed or anything.

12:19I think it's important empirical evidence. It's important to know that we can use language models to extract information about biosecurity risks. And it's very important to know what the capabilities of these models are. But at the same time, our main point is that we shouldn't be caught up by focusing on just the part where we have a language model or access to a language model alone. So for example, there's a follow-up study by Stephanie Batalis in foreign policy where she points out all of the things someone might want to do in order to create a bioweapon. And essentially finding out information about the bioweapon is a very small part of this pipeline.

12:56A lot of this information is available on the internet. It's available in AP bio courses. And so if our aim is to look back and curb biosecurity risks, then perhaps looking at the language model is not the biggest choke point that we have. There are other things that we can do, for instance, imposing bans on DNA synthesis or more restrictions on genetic screening and so on. And so our risk assessment framework is essentially trying to open up this space of risks that have been called about because are called to have been coming out of language models and opening it up to see where the most useful choke points might be.

13:34One of the big contributions of the paper is defining the risk space in terms of marginal risk. Can you dig into that a little bit more? Absolutely. So marginal risk really means what is the risk of language models or foundation models more generally compared to previous technologies. So when focusing specifically on open foundation models, we think there are like two things against which we should calculate marginal risk. One is the existing state of the affairs with existing technologies like the internet. And the other is foundation models that are not released openly, because a lot of the policy efforts at regulating foundation models are specific to open foundation models.

14:14So the second sort of comparator is what if the foundation model was not released openly? What if it is a closed foundation model? And to come up with the framework for assessing marginal risk, we bank on the cybersecurity threat modeling framework. So threat modeling is a concept in cybersecurity, which basically allows cybersecurity researchers to come up with the entire pipeline of how a risk actually materializes in the real world. It has several steps like reconnaissance of figuring out what the systems are, what the loopholes are, and essentially coming up with an entire framework of how to do this in a repeatable way.

14:49So we take inspiration from this framework and we create this six-step risk assessment framework for identifying the marginal risks of open partition models. The first step is threat identification. So it's important to know what specific threat we are looking out for and who this threat is from. So this latter point is important because it's very different if a threat is from, let's say, a few individuals or a small organization versus from state-backed actors. They have very differing levels of resources. So it's important to understand where this threat is coming from. The next two steps are about the existing level of that risk, because in many cases, the risks we're talking about also exist in the real world, as well as the existing defenses that we have against those risks.

15:31So taking together these three points allow us to then think about the marginal risk of releasing foundation models openly. So this is compared to existing risks, for instance, from web search on the internet or closed foundation models, as well as how open language models or open foundation models might allow us to supersede the existing defenses that we have. And similarly, like once we've sort of come up with the marginal risk of releasing open foundation models, it's also important to look at how easily we can defend against this marginal risk. Because in some cases, while open foundation models might allow new risks to materialize, they might also be useful for defense.

16:11And then finally, our last step in the framework is very precisely stating what the uncertainty and the assumptions in this entire analysis are. Because in many cases, it's these uncertainties and assumptions that lead to the most prevalent points of contention between people arguing on both sides of the open versus closed debate. I'm wondering with that framework defined, if you could walk us through an example, whether it's the bioterrorism or another example of how you would apply the framework step by step. So in the paper, we carry out this analysis for two areas. The first is cybersecurity risks, and the other is the risk of non-consensual intimate imagery, such as technologies like deepfakes.

16:56So maybe I can talk through the latter example, which is non-consensual intimate imagery or NCII. So first, I mean, the first step of the framework is threat identification, where you identify who the threat is from. In many cases, we've seen that the threat of NCII is actually from individuals or small organizations. So this might be someone you know who's creating deepfakes. In some cases, it is deepfakes of celebrities, in others, it is people they know in real life. And this threat, I mean, in terms of the existing risk, this threat has been around for a while. So we've seen many types of digitally altered NCII images being created through tools like Photoshop, for example.

17:35And similarly, in terms of existing defenses, there are a few defenses that people have. The first is social media platforms like Facebook and Instagram and YouTube have channels to report NCII. If someone posts your image, you can report it and it will be reviewed. Similarly, there are some federal statutes against the sharing of non-consensual intimate imagery, both in the U.S. as well as outside, like in the U.K. and the EU. So this brings us to the question of marginal risk. In this framework or in this current scenario where we have Photoshop and people do share NCII of other people at some rate of prevalence, what is the marginal impact of open foundation models?

18:11So for this specific risk, we find that the marginal risk of open foundation models is actually pretty high. Several analyses have shown that since foundation models like stable diffusion have been released openly, the amount of NCII prevalent on online websites has gone up dramatically. Compared to tools like Photoshop, the use of this model might not require any expertise in digital technology at all. And compared to earlier enforcement strategies like going on social media and reporting your image, I think advocating for taking down AI-generated NCI is a little bit harder simply because of the legal status of AI-generated images right now.

18:50This is an interesting sort of policy rabbit hole. But back in the 90s, the Supreme Court ruled that virtual pornography is protected under the First Amendment. People have a First Amendment right to create and share these pictures, even if social media platforms later take them down. That's their prerogative under Section 230. Like just the creation of these deepfakes is as of yet still a contentious First Amendment issue. And it's unclear if the use of foundation models for generating more believable or more realistic imagery will sort of cause the court to change its opinion. That's a tremendous example of legal frameworks not keeping up with technology.

19:26Exactly. And in some sense, the Supreme Court verdict in the 90s case was actually, it had great foresight because at the time, the state of virtual pornography was such that no one would mistake a virtually generated image from a real image of a person. But now when we have foundation models like generative AI models, widely available and easily downloadable. I can run Stable to Fusion on my MacBook. It just becomes really hard. It's like a difference in kind, not just a difference in quantity. And so I'm sure we're likely to hear many cases on virtual NCII and generative AI very soon. You've mentioned factors including the ease of creation, the prevalence or the increase in occurrences.

20:10It sounds like many of these factors are very much themselves kind of up for debate or discussion. And it's ultimately a judgment call on behalf of the person who's making the argument. But what you're trying to do is ground it in, at least we're going to talk about this factor. Is that a fair assessment? That's absolutely right. And that's why the last step of the risk assessment framework is specifically about the assumptions built into the risk assessment framework throughout. Like whether you think these models will continue to improve at the same rate or whether we're seeing like somewhat of a plateau.

20:43There are also assumptions about the quality of the generated image and so on. And so I think that's the step where you sort of ground all of your subjective calls in this risk assessment framework and explicitly call them out. The final point of the marginal risk assessment is comparison to closed foundation models. And here, too, at least based on the empirical evidence we've seen so far, closed model providers like OpenAI have been quite successful with their guardrails. There was this incident involving Taylor Swift's NCII that was actually using Microsoft's tool, which is a closed model. But at the same time, what this allowed Microsoft to do is very quickly patch their model so that it wouldn't generate NCII of real people.

21:23And this is just not something that's available as a redressal mechanism to developers of open foundation models. In general, I think it's clear that for NCII, at least, the marginal risk of open foundation models is pretty high. And when it comes to the ease of defenses, I think this allows us to look at this whole pipeline of the creation of NCII and come up with a few spots where we can reduce the harm. So the pipeline for creating generative AI based NCII is you have a model. You might have platforms where those models are hosted. You might have downstream platforms like social media where people might upload these images.

22:01Interventions at the model level are quite hard. It's very hard to eradicate the spread of stable diffusion simply because you can download this file as like a 4 gigabit file. You can run it locally on your MacBook. You can also run it on your iPhone, actually. And so it's really hard to track the usage of these models. But when we come to the next two sort of layers of the pipeline, interventions become somewhat easier. So one example for the middle layer, platforms where models are shared, is this startup called Civit AI. Civit AI is a platform for people to share foundation models that can generate images of various sorts.

22:37One of the uses of Civit AI in the past has been to post bounties for models that generate NCII about specific people. So this example is basically like one where tangible harm is being caused because an online platform is failing to take down models that can cause harm in the real world. A few weeks after this very nice investigation from a news outlet called 404 Media revealed this fact, Civiti has imposed stricter guardrails on how these models can be distributed. And similarly, I think for social media platforms as well, I think social media platforms need to do a much better job of curbing down on, clamping down on NCII.

23:16And one way they can do this is by allowing users to take some of the power back in the hand. So there is this nonprofit organization called stopncii.org, where if you fear that one of your intimate images is about to be shared on the internet, or perhaps has already been shared, you can upload a hash of that image to stopncii .org. and when this image is uploaded to any social media platform that coordinates with Stop NCII, and I think as of today, this includes Facebook, Instagram, Reddit, and so on. Any of these platforms, if that image is later uploaded, that image will be automatically and proactively removed so that people don't get a chance to see it.

23:55So I think interventions like these are much more tenable compared to interventions at the model level, simply because of the high costs of non-proliferation of these models. So this is what the risk assessment framework, and especially the ease of defense step in the risk assessment framework tells us, where we should focus our defenses. But at the same time, some of the harms would likely still continue. So even though public platforms and AI model hosts might take down these models, you still have end-to-end encrypted chats and you have Telegram bots that create NCIA of people. So to some extent, this marginal risk of open foundation models for creating non-consensual intimate imagery is tangible and it has already caused real world harm.

24:36And then when we come to the last step, which is the uncertainty and assumptions, I guess some of the assumptions are around the legal state of affairs. So in particular, that end-to-end encrypted chats would still continue. I think there are a couple of initiatives, especially in the UK, which are trying to undo end-to-end encryption in the first place. I think that's a very dangerous and risky policy proposal, but nonetheless, it is one of the assumptions built into our analysis. Similarly, the assumption around these models already being useful for creating NCII means that even if future models can have guardrails, current models will already always exist and they can always be used to create NCII.

25:20So again, this assumption says that even if we can somehow curb the harms of NCII in future models, that doesn't really matter if current models are already good enough to have this huge margin at risk. You referenced in your explanation the overwhelming benefit of proliferation of open foundation models. Is that part of the analysis? Is that your opinion? Is that something that would end up in the assumptions phase of an analysis? So the risk analysis framework does not take into account the benefits. It's not meant to be like a cost-benefit analysis. I think the cost-benefit analysis is hard to do in a generalizable way, simply because different organizations might have very different incentives.

Read the full transcript

26:05They might view different benefits differently. So in order to scope the framework to something that was generalizable, but also specific enough to be useful, I think we focus specifically only on the risks of open foundation models and to get clarity on the state of the risks of releasing foundation models openly. The paper is useful in the context of these broader arguments in which people are arguing cost and benefits and risks versus benefits. You're focusing on the risks. It's up to the people that are making those arguments one way or another to assess the benefits relative to those risks, essentially.

26:41Yeah, we very much did not want to offer like a guide for when you should release a foundation model, because I think that decision depends very much on the specific stakeholders. That having been said, do you present by way of example, specific foundation models and assess whether they should or shouldn't be available or provide any further guidance into the application of the framework in the context of an end assessment? Or do you stay away from that? So I think to the extent that we focus on specific foundation models, it's always tied to a specific risk or benefit. So I think our main focus is clarifying what the risks are and what the benefits are.

27:20and not on the foundation models themselves. But the other risk that we analyze in depth is the risk of cybersecurity. Specifically, lots of people have claimed that cybersecurity risks would be hugely exacerbated by these large language models, simply because we now have a tool that allows us to automatically find vulnerabilities in code. Now, we look back on the field of cybersecurity for the last 20 years, and we see that automated vulnerability discovery has been much better than humans. So what someone might call superhuman for the last 20 years, there have been these tools called fuzzing tools, which look at specific pieces of code and try to find vulnerabilities.

28:01They can do it in a way that's much faster and much broader compared to most human analysis. And so in this way, they can also improve automated vulnerability disclosure. So why haven't we been faced with a constant scurry of hacks? but it's because defenders have access to the same tools. So defenders can also use fuzzing tools to preemptively block out a lot of the harms and fix these bugs. And we think the same will hold true for language models as well. In fact, this is not just a hypothesis. We've already seen language models like Google Spam being used for security and improving security across their own code, but as well as for open source libraries on the internet.

28:43So they're using their LLM combined with a previous fuzzing tool called OSSFuzz, which used to basically automatically scan leading open source libraries on the internet. And they found that using LLMs can drastically improve the number of bugs caught and the number of bugs later fixed by developers. So I think this offense-defense balance will continue to be somewhat tilted in favor of defense, even with these large language models like Palm, as long as we keep investing some amount of effort and energy into building tools for defense. Stepping back from the paper for a moment, I'd love to have you talk about the way you think about the broader cost-benefit analysis of open foundation models.

29:27One of the main reasons why openness is important is because it allows safety-critical research. So I think a lot of reasons for our understanding of where foundation models can be used, how they can go wrong in the real world, have been built upon a lot of work done to make these models openly available. We've seen jailbreaks that are applicable to open models, but transferable to closed ones. We've seen new methods for prompt-based attacks. We've also seen new methods for defenses being developed using foundation models being released openly. So to some extent, it seems like openness is extremely valuable to fix the very problem that many critics of openness say it causes.

30:09There was this very nice letter that was released by Mozilla, I think last year, which says that when it comes to AI safety, openness is not the poison, it's the antidote. And I think I very much believe that claim simply because the alternative is that AI and foundation models more generally are built just by a small handful of companies who are licensed to create these models. Such licenses have appeared in serious policy proposals. So Senators Blumenthal and Hawley have this bipartisan AI framework. And I think one of the provisions of that framework is that we need licenses for AI developers to deem a few people fit for creating these models.

30:49I think that's absolutely the wrong approach when it comes to AI safety. For one, it is centralizing this power of deciding what's safe and unsafe in the hands of a few companies. But more importantly, I think even if we ignore that part entirely, I don't think it leads to necessarily safer outcomes because it stops all of this crucial safety research. And there is going to be a time when the dam would break, right? There's going to be a time when compute would be cheap enough that these language models can be trained by just about anyone. So like the frontier models of today are the LAN party models of tomorrow.

31:23Like gamers at the LAN party would be able to train that model five years from now. So restrictions based on like how much compute a model needs are in my view wrong, because they do not leave us with the resilience of societal interventions that would be robust to the point where these models are released openly and widely. Do you, in your research, or broadly thinking about the space, look into the idea of regulatory capture by some of the large technology incumbents and that as a motivation for these anti-openness proposals? The way to sort of counter bad proposals is by focusing on the content of the proposals themselves.

32:04So a constant theme in my work has been that I try to avoid looking at the intentions of the actors who might be sort of arguing for various proposals and instead try to focus on the content of the proposals themselves. And I think in that way, it leads to a stronger defense because rather than just relying on hypotheses or speculations about why someone might be doing something, I think we can actually respond to the substance of it. So your argument then is openness is strong enough on its own merits and closeness is weak enough on its own merits. It doesn't really matter why it's being proposed.

32:38Absolutely. To some extent, the interventions can be debated on their own merits too. I would not go as far as to say that openness is strong enough on its own terms and closeness is not. Today, it's very clear that releasing language models and foundation models more generally openly seems like the right call. But I think we need to keep investing in assessments of marginal risk. And the other thing is like this debate is often framed as open versus closed. But to be clear, no one is really arguing that there should be no closed models. The only argument is about whether there should be open models or not.

33:12And I think the framing this question this way means that my stance is reframed to there is a space for open foundation models in this ecosystem of closed and open. So it's not open versus closed, it's closed and open versus only closed. How do you see this work being built upon? There are many different future parts. In fact, as a little backstory for how this work came about, the first draft of this paper, which now is around like 10 pages, I think the final length, the first draft was 60 pages long. We tried to do in-depth risk analysis of every single risk that we mentioned. And at some point, we realized that we're just not qualified enough to do this.

33:53We do not have the expertise needed in analyzing these specific risks to be able to confidently state what the risk assessment scenario looks like. And so one of the main things I would love to see is using this risk assessment framework in different domains, like biosecurity and disinformation, to see what the marginal risks are. And to some extent, clearly lay out this argument for why there is or isn't marginal risk from openness. On the policy side, I think one of the main ways this framework can be helpful is in guiding what type of research policymakers find. So we're already seeing AI safety institutes in the US and UK start to look at the question of openness.

34:34And I think the framework can be helpful in informing them what type of research is important and urgent. And I think the second area is to give policymakers who are extremely convinced that we need to lock down these models, a reason to think again. Because I think we really do see that there is very little evidence of marginal risk from open foundation models. And once we sort of hone in on this question of marginal risk, it becomes clear that at least for the time being, clamping down on model weight releases might not be the best intervention. To be clear, we do need interventions. For biosecurity, the executive order released last year had a provision for stricter screening of biosecurity materials.

35:15That, I think, is absolutely a step in the right direction. Similarly, for cybersecurity, we need resilience on the downstream surface. For NCII, we need these platforms to coordinate with each other so that they can take down NCII as soon as it's posted. So there is a lot of urgent action needed. It's just not at the level of the AI model alone. It's at the entire pipeline of how these risks materialize. You were recently involved in an open letter that was published on a related area. Can you talk a little bit about that? Absolutely. So this is somewhat related to our work on openness. In the open letter, we argue for a safe harbor for researchers who are investigating the risks of AI.

35:56So here's what I mean by a safe harbor. A lot of the legal terms and conditions that are attached to closed foundation models, as well as open foundation models, mandate that certain types of activity conducted using the model is prohibited. So this might come from a place of caution. OpenAI might not want cyber terrorists to use their models for creating malware. Similarly, it might not want people to create disinformation campaigns. It might not want all sorts of other materials to be generated. And this is all fine and good. But what these terms of conditions also mean is that when researchers are investigating the risks of these models, perhaps by looking at things like jailbreaking or prompt engineering, they might also run afoul of these terms.

36:39In imposing this one-size-fits-all terms of condition, there is very little room left for independent safety research and independent trustworthiness research when it comes to a lot of these AI models. And our open letter calls for a legal safe harbor for AI research that might violate the terms of condition, but is still done in the spirit of AI trustworthiness and safety. And in fact, there is a long history of protecting similar types of research. So in particular, in the security community, there is this well-established tradition of a safe harbor for security researchers who responsibly disclose these vulnerabilities to the companies in advance.

37:17We argue for something very similar for AI safety and trustworthiness research. If researchers are responsibly disclosing these vulnerabilities, doing good faith research, then they should be protected in terms of legal indemnification and also from the accounts being banned or removed completely. In the case of the security example, you mentioned tradition. Is it strictly tradition or is it legal precedence? I was under the impression that it's not infrequent for a security researcher to get in trouble with, for example, the DCMA as part of their work. Yeah, absolutely. I mean, I think it It is very much more than just a tradition at this point.

37:58So you're right that in the last few years, security researchers have had troubles with technology companies they're investigating. But very recently, I think it was the Department of Justice that rolled out this notice saying that security researchers doing good fit security research are exempt from the provisions of the Digital Millennium Copyright Act. And similarly, I think people have in other security areas too, fought for and won these safe harbors from company. So I think we're very much relying on this long tradition of security researchers having done the work of coming up with what responsible disclosure looks like and how we should go about articulating it.

38:38I'd be remiss if I did not mention this other safe harbor that was proposed, which is also very close to the one we proposed, which is for social media platforms. So many researchers in the last few years have tried to collect data from social media platforms. Let's say about the ad transparency or the transparency of how often users see certain types of posts. And in a lot of cases, social media platforms have gone after them. In response to such concerns, I think the Knight First Amendment Institute wrote a letter calling for a safe harbor for independent investigations of social media, which I think was very influential in informing our line of thinking here as well.

39:13And in particular, it outlined this framework for how to think about safe harbors. While the specifics vary between these two efforts, simply because the prior effort was aimed at social media and this one is aimed at AI, I think the overall spirit is very much the same, that we need researcher access and protections when people are doing research that is societally beneficial, even if it is not useful to the companies or even if it might impose some liability on the companies. Well, Sayesh, thanks so much for joining us to share a bit about your work. Thank you so much for having me. It was a pleasure.

From the publisher

Today we’re joined by Sayash Kapoor, a Ph.D. student in the Department of Computer Science at Princeton University. Sayash walks us through his paper: "On the Societal Impact of Open Foundation Models.” We dig into the controversy around AI safety, the risks and benefits of releasing open model weights, and how we can establish common ground for assessing the threats posed by AI. We discuss the application of the framework presented in the paper to specific risks, such as the biosecurity risk of open LLMs, as well as the growing problem of "Non Consensual Intimate Imagery" using open diffusion models.

The complete show notes for this episode can be found at twimlai.com/go/675.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Assessing the Risks of Open AI Models with Sayash Kapoor - #675The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 40 min
Listen in VO