In short
TWIML AI Podcast Episode #702 Summary: Stealing Part of a Production Language Model with Nicholas Carlini
Episode Overview In this episode of The TWIML AI Podcast, host Sam Charrington converses with Nicholas Carlini, a research scientist at Google DeepMind. They delve into Carlini's award-winning research on adversarial machine learning and model security, focusing on his paper titled “Stealing Part of a Production Language Model.” Their discussion covers current AI security, ethical concerns about model privacy, and remediation strategies proposed by key industry players.
---
Key Themes
- Introduction to Adversarial Machine Learning
- Adversarial Attacks: The conversation begins with an overview of adversarial machine learning, noting that the landscape has evolved significantly since the advent of large language models (LLMs) like ChatGPT and PaLM-2.
- Model Stealing: Carlini explains that model stealing has matured into a distinct field, with practical implications for AI security.
- Insights from Carlini's Research
- Stealing Language Model Layers: The primary focus of Carlini's research is on successfully extracting the last layer of production language models. This extraction can reveal critical information, such as the token embedding matrix.
- Technical Breakdown:
- Carlini describes a method for querying the model to retrieve probability distributions over token outputs, which are essential for reconstructing the last layer.
- He highlights how the model's internal workings—specifically, linear transformations—facilitate this extraction.
- Practical Implications and Industry Response
- Remediation Strategies: Following the publication of Carlini's work, OpenAI and Google implemented changes to their APIs to mitigate such attacks, including adjustments to output probability reporting.
- Impact on API Functionality: Carlini expresses satisfaction that the attacks prompted significant changes in API functionality, marking a notable advancement in the field of adversarial machine learning.
---
Key Takeaways
Adversarial Machine Learning Landscape
- The conversation illustrates how adversarial attacks have become more relevant in real-world applications as companies deploy LLMs in production environments.
- There remains an ongoing need for awareness and proactive measures against potential attacks on these systems.
Model Layer Extraction Techniques
- Carlini's work demonstrates that extracting the last layer of a model is feasible and can provide valuable insights into the model's structure.
- The research emphasizes the interplay between theoretical understanding and practical implementation of attacks.
Ethical Considerations in AI Security
- The discussion touches on the ethical implications of model stealing and the potential harm to individuals whose data might be implicated in training these models.
- Carlini urges the community to consider the broader consequences of releasing models that could inadvertently leak sensitive information.
Future Directions
- There is interest in extending the techniques to recover additional layers of models and to explore the implications of newer architectures such as mixture of experts.
- The ongoing research aims to refine methods of model extraction while ensuring ethical considerations remain at the forefront.
---
Conclusion In this episode, Nicholas Carlini provides valuable insights into the complexities of adversarial machine learning, particularly in the context of large language models. The discussion elucidates the challenges and advancements in model security, highlighting the need for ongoing research and ethical scrutiny in the rapidly evolving AI landscape.
For further details, listeners are encouraged to check the complete show notes at [TWIML AI Episode #702](https://twimlai.com/go/702).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00The way that GPT-4 works would be functionally identical if the entire field of adversarial machine learning dropped off the face of the earth. like nothing that we learned at ever so machine learning impacted the way that we designed the biggest models today like no one trained them to be absolutely robust open ai and google changed their apis they broke functionality to make their apis a little bit less functional because of this attack because they have viewed the attack as sufficiently advanced that they want to make sure that this can't be done
0:44all right everyone welcome to another episode of the twimble ai podcast i am your host sam charrington today i'm joined by nicholas carlini nicholas is a research scientist at google deep mind nicholas joined us early last year to discuss his groundbreaking paper extracting training data from large language models, and his most recent work in the area stealing part of a production language model earned him a best paper award at the 2024 ICML conference. We'll be digging into that paper as well as his other best paper winning work focused on large scale application of differential privacy in machine learning.
1:23Before we get going, be sure to take a moment to hit that subscribe button wherever you listen to today's show. Nicholas, Welcome back to the podcast. Yeah, thanks for having me here. I'm super excited to jump into our conversation. I will refer folks back to our last convo for your full background, but I did notice something in your bio that I hadn't noticed last time, and that is that you like to write interesting coding puzzles like tic-tac-toe games in a single printf call or doom clones in JavaScript or a fully functional CPU on top of Conway's Game of Life? What is that all about? I feel like programming should be fun.
2:07And most people get into computer science because they find joy in writing fun programs. And I wanted people to understand that people can both do serious research and also have a good time with things that have absolutely no practical relevance but at the same time are are just entertaining and like you know like it's it's a thing that you the people should should have fun with the things that they're working with and so yeah i wanted to to call it some of the things that i do that are interesting and fun but have they don't they shouldn't exist it's like there's not no no one asked for this to happen uh it's just it's entertainment that's awesome awesome uh so as i mentioned the last time we spoke we were talking about some of your early work in exploring this intersection of security and large language models large language models as a field has exploded since that time maybe a a good way to ease into the conversation about the specific paper is to have you talk a little bit about that explosion from your perspective.
3:21Like how has the landscape evolved since you started studying it? Yeah. So yeah, in the last year and a half, so the last time we spoke was just after chat GPT's launch. And so in some sense, very little has changed because in the field of adversarial machine learning, we're still worried about the exact same problems. We're still worried about attacks that we used to call adversarial examples and people now seem to call prompt injection. We're still worried about poisoning where someone manipulates the training data set. We're still worried about data stealing where someone tries to recover data model was trained on, we're worried about the same set of problems.
4:05But on the other hand, everything has changed because it's no longer a field that is focused on problems that may exist in some hypothetical scenario. It's now a question of attacks that someone actually would want to implement because people are using this in real production environments. And so a bunch of our recent work has been trying to think about this problem from this angle of why would someone who is an attacker actually want to achieve some goal and not why might a security researcher want to write a paper on some problem that they think might be interesting in the future. And so the problems have been the same, but the applicability of these problems to the real world has gone up by quite a lot.
4:55and i inferred from uh the paper that the this idea of model stealing has become its own kind of standalone field of research yeah so model stealing as a field has been around for for quite some time uh in in machine learning the very early paper on this early machine learning world i think was 2016 by one of the co-authors in this paper florian tramer and his collaborators where they showed how if someone deploys a simple machine learning model as some prediction service, and I interact with the model, that I can actually steal the model the person has trained by making API queries. And the way that they did this was they focused on some very, very simple classifiers, like a linear classifier that takes the input, does a dot product with a single linear sort of scaler, and then returns you the output.
5:54some very simple classifiers like this, and they showed that it worked in those settings. And there has since then been a very long field of work that tries to make this more and more practical. And so the way that this tends to go, that the field basically has two directions. There's one line of work focused on trying to steal a model that is roughly as good as you can by making queries. So what this looks like is you have your model and I make queries on like the same data distribution I care about. And then I do some kind of distillation, or I use the fact that your model is good to train my own copy that's like as similar as your model as I can make it.
6:39And here what all I care about is getting high accuracy, like on the same data distribution. There's another line of work that tries to do something more like model extraction, where i'm trying to actually get at the internals yeah like the exact weights of your model like i want a bit for bit copy of what you have on your disk and it turns out that this is this is actually impossible like i'm in in that there are many different models you could have that represent the same function you know you could scale one layer by two and the next layer by one half right and so when you when you perform your computation you'll initially have your computation going one way, you'll multiply by two, then divide by two immediately afterwards, and the output will be identical.
7:23Like I can't tell. And so I can't actually recover the same bits that you have on disk, but I can recover what's called a functionally equivalent model that like up to some invariance is as similar as possible that I could hope to achieve. And so in prior work, let's say In 2020, we and another one of the co-authors on the same paper, David Rolmich, wrote independently papers that showed that if you have a neural network with only ReLU activations that only has, let's say, under 50 neurons that works in floating point 64 precision, that I can feed in arbitrary inputs that i get full arbitrary outputs then it is possible to steal this model in a functionally equivalent way okay but no one uses models that have like 50 neurons that work in float 64 and all these things yeah and so after we both wrote this paper at the same time florian the person who originally did model stealing david the person who wrote the other paper concurrent and i got on a call together and we're like okay like what can we try and do how can we extend this work and we tried to come up with some ideas and one of the ideas we came up with was actually the idea that eventually became this paper from six months ago um but the so we had an idea but the idea had no practical utility yet and so we just like we didn't we said okay fine like we can't we can't find something to do let's sort of like let's wait um and then you ask sort of what at the beginning, like what's changed?
9:04Well, the language models became a thing. And it actually now is the case. Like if you think about it, it's now actually a thing in the real world where someone will deploy a model that you care about, that you can make queries to, but you don't actually know what the model is, right? This is what a language model API is. You know, all of OpenAI and Google and Anthropic and everyone like have these production language models that do great, but they're not going to tell you exactly what the weights on the models are. They're not going to give it to you. And so we went back and we started thinking about this model stealing problem again.
9:41And we came across this idea we had from probably 2020, 2021, and realized this is now an important problem that we had some ideas on how to make some improvements. So that's roughly how this paper happened. Now, was the idea problem focused or approach focused? Yeah, both. Right. So, so we had the approach. So, okay, so maybe let me tell you what we can do in this paper. In this paper, we show it is possible to recover one layer of a language model. So language models, like, you know, like any other machine learning model are a sequence of layers. and what we can do is we can steal the last layer of the model which turns the internal representation into a sequence of tokens or words or whatever you have and in many models the first layer and the last layer are the same they're shared and so this also gives you what's called the token embedding matrix which turns the input tokens into the hidden state now not all models are this way but many models are they do share these two layers and so what we learn is the last one and if the last one's the same as the first one we also learn the first and a couple of comments yeah one is the to contrast the paper that we talked about last time again this is february 23 this was focused on extracting training data as opposed to the the model itself Yes.
11:07And I guess I want you to speak a little bit to the like the relationship between those two and the way you think about them and the approaches and their utility. Yeah. So so they're they're both about stealing, but they're about stealing entirely different things. So one of them is. Although like presumably the way you're the way and the reason you're stealing the training that is that it is somehow embedded in the model. also extracting training that is kind of model stealing yeah okay sure yes okay so from this from this perspective yeah but so maybe let me say talk about the motivation so in security this is a question of like who is being harmed by an attack like when i'm doing an attack like who am i trying to get at when i'm doing model stealing i'm generally harming the person who trained the model because like i have spent you know whatever millions of dollars training my very fancy model and someone else can try and get a copy of it without having spent that much money.
12:08Versus training data extraction is more privacy concerning your... Exactly. I'm worried that some hospital is going to train a model on some patient records and they're going to try and do a good thing for the world and make it so that some small hospital in the middle of nowhere can feed in their medical scans and get a reasonable diagnosis when previously they had no specialist on hand. And that would be great if they could do that, But it turns out that given access to the model, I can learn about the people the model was trained on. And so there's concerns of patient privacy. And so I'm not harming the hospital directly because they've actually published their model.
12:46I'm harming the patients who participated in this scheme. And so this is the big difference on the motivation here. And then the techniques are also fairly different. So when you're doing data stealing, for the most part, the way that we do this is we try and mimic the kind of data the model may have been trained on. And we try and get it to emit that kind of data at test time. whereas for for training for sorry for the model stealing piece what we try and do recently is use sort of we treat these things as mathematical objects and we treat this like a cryptographic problem where like you have this thing that exists and we try and steal it and so like so this is what we tried to do so the paper that i mentioned that we published in 2020 we published our paper at Crypto, which is like the conference where people submit new block ciphers and new theoretical encryption algorithms and whatever, these kinds of things.
13:53And the reason we put it there is we tried to draw this analogy between model stealing and cryptanalysis of ciphers. And this is the way that we framed the problem. We framed the problem not as some machine learning thing, but as you have this mathematical object. It happens to be a language model or a vision model or something, but like what it's doing internally is some matrix multiplies with some non-linearities like this is a thing that is just a mathematical object you can study independent of how it was derived and cryptographers are really good at that kind of thing and so we tried to frame this as that kind of problem to get some of that community focused and working on this problem the second question i had you were talking about the extracting this um this final embedding layer which often correlates to the initial embedding layer you know there's a lot of excitement about embeddings embeddings uh have are like increasing importance does that mean that like extracting that piece in and of itself uh is outsized outsized has outsized value relative to other parts of the model that you might be able to extract yeah we don't know um this i think is a good question so the the question of what you would do with it so we have some things that some applications of what i could do with stealing this part of the model um you can speed up some forms of generating adversarial examples in some ways.
15:16You can speed up some other attacks that need the ability to make arbitrary queries to the model. But a big part of what we wanted to try and do was to say, previously, the literature has focused on, let's imagine a hypothetical threat model. And then let's see, what is the strongest attack that I can make possible there? and like you change the threat model so that your attack becomes really strong. And this is, you know, I said we can steal a full model as long as it had this long list of assumptions that were all false in practice. What we wanted to do is we wanted to say, maybe on occasion we should write papers that look at what is true in the real world and then try and see what is the best that we can achieve in the real world.
16:09And obviously, it's not going to be nearly as strong as we can achieve in some hypothetical future where like various things were true but like this is the world we're living in now and so let's let's look at that world and try and see what the strongest attacks we can do are here because we know it will have some real world impact and that was like the setup for this paper is maybe this model is so maybe this layer is very interesting maybe it's not but it's it's the best that we can actually do right now and my hope is that or maybe my hope isn't i don't know depending in which you're viewing this, you could extend this to a second layer, right?
16:42You know, once you've done one, maybe you can do two. Once you've done two, maybe you can do three. And this is how the attacks that we wrote previously worked. Previously, we would go first layer, second layer, third layer, fourth layer, and we'd sort of cover them one by one. This time we're going backwards. We've got the last layer. Maybe you can use that and get the second to last layer. Now, our techniques are not going to help you there, but someone else who's smart might be able to come up with a way of doing that. And this is the general question and the way we want to try and reframe the field a little bit so that we could get some people looking at some real world problems.
17:11The implication there being that you were very much targeting that last layer as opposed to you had an idea of an attack and, oh, look, it produced the last layer. So this is where it comes to the two. So we had this idea in mind of let's try and attack real things. And I think even in the last time we talked, I might have mentioned this a little bit, where this is the thing that we've been trying to think about for a while. and in general the way that i try and find research ideas is i try to first ask what do we want to achieve and then ask what are the things that we can do and then you just like compute the cross product you look at all the things that we're able to achieve we'll look at all the things that we actually want and then we try and see which ones match and so we had this idea in 2021 or so of a way to recover part of a model and then we had this sort of realization that we we want to be able to to steal these these production language models and so we already had both these ideas we just needed to connect them yeah and that's what happened sometime uh almost a year ago today and then we we spent some time writing the paper up and then i got the paper into ICML.
18:34I guess I'm really trying to get at like the, you know, I want to steal part of a language model, a production language model is different than I want to steal the embedding projection layer of a language model. And so was that just a low hanging fruit or was that what appeared as you started attacking or did you look at the model as a white box entity and say, i can probably get that thing great okay um maybe to be more clear stealing the last layer is the thing that we learned how to do in 2021 ah okay so like the the attack idea that we had in 2021 was exactly this of how to recover the last layer of a model that had certain properties okay and no interesting model at the time had those properties and it was not interesting to recover any part of this model if we could make one up and then when it turned out that like these language models became a thing and we tried to think about stealing we went like well what are all the things that we know how to do and was like oh here's this idea that we came up with in 2021 that was not very useful at the time and it has to be adapted because at the time we were looking at value networks and language models aren't value networks and so there are some tweaks we had to make but we had the idea for how to steal the last layer of a model with certain properties at that time and then it turned out that many of those properties happened to be true some not that we had to adapt some that some that were that we could just transfer directly over uh so how does the attack work yeah okay um so uh i will hopefully not scare anyone away um but you will need a little bit of linear linear algebra knowledge here uh so let me just try and say this in the simplest possible way.
20:26So a language model emits, when you make predictions from it, a probability distribution over a bunch of tokens. So it says, you know, if you have like, my name is, it will have very high probability on valid names and very low probability on like the word cat. Okay. So many models have, let's say, 100 ,000 possible tokens that they know. This is roughly like words or something like this. And so you can visualize the space of outputs from a model as a 100 ,000 dimensional space. Now, one thing you could ask is what happens if I query on many different random inputs and plot the outputs of the model in this high dimensional space.
21:18One thing you could imagine is maybe I'm drawing points randomly in this hundred thousand dimensional space. And that's what people usually think about. It turns out though, that's not the case. Actually, despite the fact that these these logit vectors live in 100 ,000 dimensional space, they actually all lie on a much smaller dimensional subspace, let's say like, maybe 8 ,000 dimensional subspace. The way you should visualize this is, you can have a bunch of points in three dimensions, and you can have them all put on a plane, a two dimensional plane. So you can have a bunch of points in some very high dimensional space, but they all live on some lower dimensional subspace of that.
21:57And that's the case for language models. And there's this very particular reason why, which is what language models do is they work in some small hidden space. The hidden space of the language model is the width of the network. And this is roughly 4 ,000, 8 ,000, something like this for various models. And the final thing that the model does is it projects from the hidden space up to the output embedding space with a single linear transformation. And if you remember your linear algebra, you cannot increase the dimensionality of a space by a linear transformation. And so what you do is you take the hidden space and you just linearly scale it up into an 8 ,000 dimensional subspace of this 100 ,000 dimensional vector space, which is the outputs that the model produces.
22:44OK. Now, this is true. And now I'm going to skip over a bunch of details, except to say first one thing. First, let's observe that this lets you learn the size or the width of a model fairly easily. This is sort of the first warm-up attack that we have. First thing that you can do is you can learn how wide the model is. And the way that you do this is you make many, many, many different queries to the model. And then you just plot them all. And you look, what is the dimensionality of the subspace that all these points lie in? and it will tell you, let's say, 8 ,000 or something, or 4 ,000 or something.
23:24And now I've learned a single number about the model which was secret. Now, this number in itself is also a little bit interesting, because the size gives you something which people don't publish, right? No one knows how big GPT-4 or GPT-3.5 or any of the cloud models are. But if you run this attack in this way, you can learn not the exact size, but you can learn the width. Only the height is missing. And the height and width are usually highly correlated. And so I can learn some significant information there. And I think this in and of itself is also interesting because, okay, so for example, we worked with OpenAI to run this attack on their models.
24:05We got their permission. Lots of excitement there of like, you know, asking Google lawyers to ask if I can steal OpenAI's models, at which point they say, tell me no. and then like you know okay um but uh after we got we all agreed on this um i ran the attack um and what this means is there are now two groups of people in the world who know how wide gp 3.5 is there are like current and former employees of open ai and me like the way like this is like an actual thing you would want to learn and like actually the best way to learn it is through this attack which is something i think is interesting because this is not the case for many other attacks in security right or in adversarial machine learning where there are like other other ways you can achieve the same goal but like here like the best way to achieve this is through this attack and so we we have this piece of knowledge you talked about um querying the model.
25:05What is the nature of these queries? Are they, you know, human readable, a collection of human readable queries, a collection of nonsense that has statistical properties? Like, what are we talking about here? So it actually doesn't matter. In the language of cryptanalysis, all we need is known plain texts. I don't need to choose the inputs that I make. I just need to know what they were. So all that I need is a bunch of different inputs. And so in practice, what I did, because if you had a large user, you could passively observe their queries and responses and perform your attack without actually executing new queries, or at least this initial part of it.
25:46With one caveat, which is that I need to know the entire probability vector for a given input. And so most language models don't reveal this to you. They only give you, let's say, the top five probabilities. so what i have to do is i do have to make many queries to the model with the same prefix with the same prompt and ask the model to reveal to me the probabilities of all the tokens one after the other so i can't just do it entirely passively but which prefixes i choose don't matter so you're asking the model uh you're presenting a fair degree of repetition so that you can reliably extract the distribution of yes output tokens correct so i'll take the same prefix um whatever the prefix is and i'll ask the model for the output probabilities and i will do this you know a hundred times or a thousand times so i can recover the probability distribution over outputs And there's like the majority of the technical work in the paper is exactly in how you do this efficiently.
26:57Because there are lots of strategies you could take depending on exactly what APIs are made available. And do you also, presumably, I guess, you know, you also need to. So that's one dimension where you're like crafting your sequences. But then another is like some degree of coverage so that you can broadly scope out the distribution. So that matters less. Oh, really? Okay. Yeah. So all we need is to be able to just, because models are highly nonlinear. And so as long as you make any arbitrary queries, you'll cover basically the whole space. Dig into that a little bit more. How does the nonlinearity translate to covering the whole space?
27:39I guess when I hear something is nonlinear and you're trying to extract information about it, my intuition would have been that, well, you need very broad coverage because learning something would be very local. Okay, yes. So this is a good distinction. We are recovering the last layer, which is just a linear projection of the hidden state. So what we are learning is just a linear thing. All we need is inputs that in the dimensionality of the final hidden state are sparsely distributed around. And because the rest of the model is very nonlinear, we can rely on that to give us points that are essentially randomly sampled from the embedding space.
28:26But we're only learning the final linear projection. And that's what makes it easy to do that piece. Got it. So the linearity of the final layer kind of says that you only need a you know small number of parameters like if you can learn all the points of a line with you know your m and your b and then the internal non-linearity says that you're going to get a broad sampling on this whatever that space is you don't need to craft that yourself exactly right yeah you're relying on the non-linearity of the first piece to give you the sampling and you're relying on the fact that there are not many parameters for the second piece.
29:07Now, okay, in a simple world where you're doing a line, yeah, this is like mx plus b, you know, you do this in pre-algebra kind of things. What we're trying to learn is a matrix times some x plus some bias b. And it turns out that the techniques you could do here are essentially the exact same things as the things that you learn in algebra. They're just over some big giant matrix and it turns out that there's an algorithm for it it turns out the singular value decomposition gives you this it's not immediately obvious why but that's like very in the weeds it turns out that this is the thing that you can compute and and you can recover it in this way so the attack is a matrix inversion essentially yes yeah what we're trying to do is we're trying to invert some matrix and yeah there are reasons why it makes sense to try and do this and um no It's one of these things where the attack is not hard because of the math.
30:09The attack is hard because of the scale. Yeah, right, exactly. You need to get all the things working in the right way. Most of the technical work, as I said, was recovering these probability vectors because no one wants to give you these directly. You have to do some analysis on how to actually recover these things. And yeah, and so that's what makes things a little bit tricky. And then there's a lot of analysis that happens in the paper where we try and formally characterize exactly what's going on here. And this is where it's nice to have theory people as collaborators. I am not a theory person.
30:44My knowledge does roughly stop at singular value decomposition. But all of them are able to tell me, here's this attack that works. I go to them and say, why does this work? and they can do the analysis and explain the query complexity and here's what you should in principle want to be able to do. And then we can go and take that and actually implement it. That's awesome. And so getting from the collection of query responses or the pair of queries and responses to the matrix that you need to invert, what's involved there? Or is that an obvious thing? No, yeah. So it's not obvious, but the answer is nothing.
31:25So once I have the, so I'm going to form a sequence, I'm going to form a matrix. Here's what my matrix looks like. For each row, it is the logic vector for each token, one after the other after the other. So I sort of look at the token for the first token, second token, third token, fourth token. I place the probabilities of the logits in each of these for a row and I form the matrix that has many rows. Each row corresponds to a different query and then this is the matrix that I invert. On this matrix, I compute the singular value decomposition, and it turns out that this gives you what the final projection matrix is.
32:09And yeah, I won't work through the math there, but the 10-second explanation that should convince you of this is this is a low-rank matrix, and the singular value decomposition gives you the low-rank approximation If it actually is low rank, it gives you the actual low rank values. That's all I'll say for people that this made sense for. And this didn't necessarily make sense to me a year ago. I had to like brush up on my map. Low rank is essentially the dimensionality of your embedding space is much shorter than the number of samples you have. Yes, it means that these vectors live, even though that they are in a 100 ,000 dimensional output space, they actually live on a small, low dimensional hyperplane in the outputs.
32:53Yeah, so there's some fun math here that I hadn't touched. So ostensibly, I have a math degree. I haven't used all that much of it in the last, let's say, 10, 15 years. And so it's fun to occasionally go in and get to actually use some of that stuff. and so a lot is known about uh at least for gpt2 like the tokenization scheme and things like that like do you rely on any of that knowledge uh no we don't but but they actually publish it too um the reason why is because you want to know how much you're going to pay on a per token basis and so most providers give you a way to tokenize and compute what the cost would be of a given query OpenAI actually just gives you the full tokenization directly.
33:46They have this library called TickToken that is open source, and they just give you the full tokenization of all their models. And so this is not usually where the secret sauce lives. But if it were to be, we have another paper that shows how to recover this new. If it were secret, it wouldn't be harder for you to do what you're doing. Yeah, it's not. We use a couple of facts about the tokenization, but not a lot. but people would in general want to reveal this anyway. I was going to ask about a mixture of experts, which is kind of an emerging architecture for these large-scale models. Do we know enough about those architectures to say whether this attack might or might not work?
34:27Yeah, so it should work identically because we make no assumptions on the internal representations of how those are computed for the model. As long as the final output projection is a linear transformation, everything else will behave the same. and you will of course need to make very different approaches once you try and start to scale this attack up to recover let's say the layer before because if that has some mixture of expertise kind of thing that might behave very differently but as far as ours goes we don't make any assumptions about what happens inside the model you wrote about this you have a paper someone can't go and like perform this attack because turns out there's a remediation that you shared with OpenAI and Google, like talk a little bit about the remediation approach.
Read the full transcript
35:08Yeah, sure. Yes. Yeah, so this I think is one of the big successes from this paper that I'm happy about. What do I mean by this? For a very long time in adversarial machine learning, we have been producing attack after attack after attack. And the way that you know if an attack is really actually powerful is if people who are not security people decide to change it the way that they do things because of your attack. So let's look at other areas. The way that we have, the way that your processor works, like Intel and AMD have put security features on your processor because of particular attacks that security researchers have discovered.
35:50When you're on the network, you're connecting over HTTPS because of attacks that security researchers have demonstrated are possible. All of the pieces of every computer that we use today are informed by attacks that have previously existed. And we have built them in because people who are not security researchers have decided this attack is sufficiently worrying that I will change the way I do my things to stop that attack. In adversarial machine learning, we didn't have this for a very long time. The way that GPT-4 works would be functionally identical if the entire field of adversarial machine learning dropped off the face of the earth.
36:27like nothing that we learned at ever so machine learning impacted the way that we designed the biggest models today like no one trained them to be absolutely robust what what what are you saying that is saying something about i'm saying that we were thinking about problems i guess i'm thinking about like uh the old um you know put a sticker on a stop sign and it turns into a picture of a graph and like people like theory you know autom av makers in theory like thought about these kinds of problems and did something to maybe but i i have seen no evidence that any car manufacturer actually has done this why suppose that i wanted to make a car to crash the adversary of someone who has the idea that they want to like cause your car to crash to kill you has many other ways to do this than like put a sticker on a stop sign they could put a trash bag over the stop sign that's probably a lot more effective than trying to construct a transferable physical world adversarial patch so it's kind of a critique on the the field more so yes yeah yeah yeah okay yeah exactly i'm not critiquing everyone else i'm saying that like we as a field have not produced a compelling enough attack that other people have been willing to change the way that they do things because of the attacks that we have.
37:46And so the reason I was excited that we sort of open AI and Google changed their APIs, they broke functionality, they made their APIs a little bit less functional because of this attack. And so it tells me that like, I have convinced someone else who is not a security person to do something that makes some of their customers less happy with how their API works because they have viewed the attack as sufficiently advanced that they want to make sure that this can't be done. And so I think like sort of at a high level, independent of what the actual methodology was to fix it, I think it's, I really liked seeing the fact that we could finally convince someone to make a breaking change to fix some security problems.
38:27So I think that was very nice. Okay. Now let me talk for a second about what they actually did, which is a relatively simple fix. um what what they do is they make it so that you can no longer for the same query actually read off the probability distribution over all possible output tokens was it previously all literally like literally all i thought it was five yeah so it was five okay but what we showed how it was possible to use another parameter called logic bias to bring different tokens into the top five at different times. So you would get, so you would query on some random prefix and you would get the top five.
39:09And then you would say, okay, what about those five? I'm going to bias those tokens to be plus 100, bring those into the top five. And then you get another five. Then you bring another tokens in the top five and you repeat this and you do this many, many, many times. So you just have to, you know, iterate it, you know, by N over five or whatever. Exactly. N over four because of some, okay. Windowing and overlaps and stuff like that. Exactly. Yeah, okay. with very fancy math you can get over five to work some of the theory people on the paper showed this it turns out to be numerically unstable and so we did on over four but yeah okay like there are lots of fun details here but we don't have to get into all of them but yeah so the way they fixed it is to make it so that if you send a query that has both asking for top p to like learn what the probabilities are and you ask for a logic bias then you can't do both at the same time.
39:59Like you can, the way that the logic bias works now is it sort of works after it tells you what the probabilities are. So you can always learn the top few probabilities. They increased it from top five to I think now top 20, but you don't get to bias these with the logic bias. So that's how the defense basically works. Meaning the top 20 is invariant to the logic bias parameter. caress and i think you worked also with google did they take the same remediation approach uh yeah so google actually just banished logic bias entirely uh it was like um let's just be extra safe about this and they may at some point put it back at some some in some um safer way um but uh yeah so there is an attack in our paper that works with logic bias only without probabilities.
40:57It does some binary search. So what you do is you say, well, what is the value of token 12? Well, I'll increase its bias to be plus 100. Now it's almost certainly the most likely output. And then so now I go to 50, and maybe it still is. Now I go to 25, and maybe it's not at this point. So now I go to 37.5, and you sort of do binary search to recover. And this is a much less efficient version of the attack, but in principle, it still works. And so the API that was at the time on Palm 2 just removed entirely the logic bias. And they may return it. I mean, this is an ultimate question of what is the cost of the attack and what is the value of the information that you learn?
41:38And the cost of the attack can just be, you can count it for the version that we implemented. It was like a couple hundred to a couple thousand dollars. If you were to do it with this legit biased binary search trick, it's now, I think, you know, five or 50x more expensive or something on this order. And so it ends up being$10 ,000 or maybe$100 ,000, depending on exactly what you want. And there's some value dollar amount at which it is cheaper to just train your own projection layer. And it's up to the company to decide like, what is that value for them? And which defenses do you want to put in place?
42:15the implication of that last statement you made is that the reason why someone would be doing this is because it's cheaper to steal a projection layer than to create their own but is that really the reason why somebody would do this sure okay there are other reasons too you might for example just be interested in learning the size of the model and training my own is not going to help you here um it might be better for some attack in which case training my own might not be helpful but But like, very rarely is it the case that what you want is the projection layer. You probably want to do something else.
42:47You want to make another attack more efficient. And so in this case, it may be better to spend your$10 ,000 training some other transfer prior model and spending more on queries. Like there's the thing that you care about, and this is sort of instrumental in that way, but it's not like the ultimate thing you care about. So the question is like, as an adversary, where would I want to spend my money? And you just need to make sure that any piece of it is not like the most vulnerable piece, right? Like, you know, in principle, anything is broken under enough of a dollar value, you know, like no one wins against the NSA.
43:25Like, so it's just like how much money does someone willing to spend in order to get the given thing? And as long as you've put it above the threshold of what it's worth, then it's not economically efficient for the adversary to try and steal it in this way. As a security researcher, do you have a sense for whether there exists a GPT-X zero-day market kind of thing? Yeah, this is a good question. I don't know, actually, right now. I think this is a thing that exists in classical security, right? Where people can find zero days in some software and they'll then go and sell them to the highest bidder.
44:09Or if they're a nice person, they'll go to the bug bounty program the company provides and the company will pay them for having found some security bugs. I don't know if this is the case right now. It may be and I just don't know about it. but I think what's probably more likely is that there's just not yet a huge amount of economic value in any of these particular attacks there's some but it's like it's not quite there yet and also many of the attacks are relatively easy to discover and so it's not the case that they would command very high of a value because you could probably have found it for less money by just doing the work yourself.
44:50All these attacks that we're finding, we publish the papers on them. And so, in principle, you could pay someone a large amount of money, but once the paper's on archive, you could just go and read the paper and then implement it. And so, for many of these things, it's not the case that I think there's a reason for this to exist yet. But I think this is mostly a function of the fact that the total amount of value being driven through language models is only a very, very, very small fraction of the amount of value being driven through like the entire internet. And so it just makes sense. The value of what you might pay for an exploit is corresponds to what you can actually get out of it.
45:27And you just can't get that much out of these language models now. But, you know, it's a possible future if it turns out that this language model thing really does scale to be like truly amazing. I think it's entirely reasonable that the value of an exploit would scale with the value of what you can sort of get out of it. And then it's another question whether, you know, those zero days would be like model stealing versus privacy targeting versus some other thing. Exactly, yeah. And that's a question of what is the thing the adversary actually wants to achieve at the end of the day. Right. And I think this is a very important question that's nice to try and figure out what it is that we actually want to do.
46:08Interesting. And so where do you see this particular research direction headed? We're going in a couple directions with this. One direction we're working on right now, we have some preliminary results, is how to achieve the same thing with fewer assumptions. So I mentioned that we were trying to do previously this with this logic bias and, you know, this learning the probabilities and these things. And there's a question like, what if you don't have that? Can you still try and implement the attack? In other words, can you do it post mitigation? Exactly. Yes. And this is the question you always ask.
46:43Like you've removed one thing for me to do. Like what else can I do? There are other questions of, can I extend this to two layers? Can I take a different approach that lets me to learn the first and second layer of the model instead of the last and second to last layer? You know, there's lots of these kinds of questions. And at the same time, we're also doing, you know, still some basic research on the stealing attacks that try and scale it up from, you know, these tiny 40 neuron models to let's see if we can do 400, 4 ,000 neurons, like just trying to approach the problem from all directions at the same time.
47:19And, you know, most of these things won't work, but eventually you'll find an idea of it has one thing that has some promise, and then you'll push the field forward a tiny bit more. And then you hope that other people can do the same kinds of things. Are there kind of qualitative merits one way or another for the kind of bottoms up to you know back down approaches to these things in terms of what you might expect to learn about the models no i just think that you try and approach the problem for as many different angles as you can and you try and see what value you can get from each and you know you don't want to spend all of your time thinking about just one approach even if it's been fruitful in the past because there might be some you know very easy thing you can do in a different way and that's we're trying to spend a little bit time maybe it turns out that the last layer is like where you stop and like going the second to last layer is just not worthwhile and you know but even still like then maybe it's easier going so it turns out that the reason we were doing this for the previous work is recovering the last layer in some cases if you do go forwards can be kind of challenging and so getting the last layer in itself is actually useful for the other attack that just went forwards because you go forward up to the last layer and you get the last layer for free that's a win even by itself and you know this is the kind of thing we were hoping for i mentioned at the very beginning that you won a best paper uh for another paper this is like a position paper on differentially private learning with large-scale public pre-training uh we might not be able to give this one the full treatment, but...
48:57No, that's fine. It's a much simpler paper. All right, let's go for it. What's this one about? Yeah, so this one has no technical results. It is just making an argument, just making exactly one argument. So the paper is trying to say something about privacy. So first I have to tell you what people do. So back to data stealing, this question of you train a model on some private data and you're worried about someone who can recover the data. There exists a defense that prevents against data stealing called differential privacy. It is not from the machine learning field that predates most of deep learning.
49:41And it comes from some other areas of statistical analysis of data. And what it does that guarantees that no individual example can influence the final trained model by very much. And so you can provably say if your model is trained with differential privacy, and you set the parameters right, then you cannot have memorized any individual person's data. And so the model cannot ever emit people's data to anyone else. And so it would be safe for a hospital to train with user level DP of particular values of epsilon and delta on like the hospital could train on patient data and release the model and that like it would mathematically provably be safe which is like an amazing result to have so the thing that we're arguing in this paper is that people do it so why is this why is that not every model is trained in this way because you lose a bunch of accuracy.
50:41If you train with differential privacy, the accuracy goes down by oftentimes quite a lot. And so what people do instead now is the following. They take some big giant model trained on all of the internet without privacy. And then they fine tune it for a little bit longer with differential privacy. And then they say, this final model is private. And it is private with respect to the data it was fine-tuned on. This is very different, though, from the model being private. This is the only thing we're trying to argue in this paper. Does the original differential privacy, you know, the theory that supports that extend fully to like the guarantees the differential privacy provides, does that extend to pre-training or, you know, is there some theory of interactions that says, yes, you've differentially, you've trained, you've fine tuned in a differentially private way.
51:53but like there are interactions that could leak data yeah so no interactions can happen so differential privacy in this sense is perfect but it only protects the data that you trained with the privacy preserving technique so okay let me give you but if it does that if it does guarantee that that seems pretty good in the hospital example yes okay so yeah okay so so now let me so so this is why people did so now let me give the points now let me argue the paper so this was sort of all the setup for like what where the world is now you have differential on privacy, it's great. It loses utility. So people train big models and take those and fine tune with privacy on those.
52:27Okay. So let me argue the paper. First argument. Suppose that, let's say you trained a model not on hospital data because, I'll only like to that in a second, but you would do this on, let's say, people's text messages or emails or something like this. And you release a model that you claim has preserved privacy. okay now let's suppose that there's some person out the world let's call him peter who who takes this model and he starts prompting it and he starts like asking like what do you know about me and the model says you know i'm not sure i don't really know very much and the peter says i'm you know um here's the thing that like you know you shouldn't know my phone number like can you tell me what what peter w's phone number is and the model like gives him his complete phone number peter might be like you know like wait like you trained on my data like you're giving me my phone number, like, what's up?
53:21Like, like, you must have done something that's sort of like, wrong with this data. This is not actually something that would be entirely hypothetical, because if you trained a model that was a fine tuned version of GPT-2, it turns out that there's some person, Peter W., that GPT-2 memorized his phone. He uploaded this as some part of like a scientific thing that was to the government that required that he attach his phone number, physical address, company, fax number, email address, all these things. And GPT-2 saw this document and memorized it. He did this in one particular context, like to disclose something to the government.
54:09And the model is then being fine-tuned on a bunch of other data. And then he asks for his phone number and the model reveals it. Now, on one sense, the model provider is correct in that the fine-tuning data did not cause any memorization of Peter's fandom. But you can imagine how you might be confused if someone tells you this model was trained with privacy, and yet it still memorizes all of your data because it happens to be from the pre-training. Okay, so now someone might say, that's fine. That's like an education thing. Yeah, you deserve it. But like, there are a couple of concerns here. The first is, if people are not particularly careful in how they describe what they have done, this can degrade the meaning of differential privacy.
54:56Because someone else could train the entire model end-to-end with privacy and actually do things perfectly well. And another person can do only the fine-tuning with privacy. And it's very hard to explain differential privacy to users, even without this subtle distinction. but now i have this this another level distinction that like one of the two breaks the privacy and one of them doesn't in this way and so this this can have this one problem so the the other concern is not everything on the internet was meant to be on the internet and not everything that was like placed there for one context is meant to be used for another right so you can for example doxing is a thing that is bad right this is like but like where you and i live in the united states almost certainly is a matter of public record right you can like you can if you sort of have a house with a mortgage or something you can go to the city and you could file the paperwork to request like you know where it is that you're staying today it may even be available online There are things online that are there, but are there because it has classically been hard to find these kinds of things.
56:13And so a model just making some of these things available can feel like a privacy violation, even if technically the information was already still available. On the other hand, you also have things that were never intended to be put on the Internet, but were put up there for a short period of time. a model was trained on them. They were then taken down. And now the only knowledge of this is contained in the model. So for example, GPT-2 was trained on some uploads of people's conversations that were on IRC that were made available for a short period of time online through 4chan, then taken down.
56:50So the best of my knowledge, the only way to access these transcripts now is by querying GPT-2. So someone like had some private data, they uploaded, they took it down, and now it's only there. And so if you had a fine-tuned model of GPT-2, you would be revealing still this kind of data that is no longer available anywhere else. And so it is very much the case that not all data that is public is necessarily intentionally public. And so you have this world where people might be saying that things are privacy preserving because they want to make sure that they can get their utility back. but you end up sacrificing some other kind of privacy because you're preserving the mathematical definition.
57:37But privacy is one of the fields that how the user feels about their privacy matters like almost as much about as whether or not you're actually mathematically giving us a particular kind of thing. Because in many cases, the harm is not a real, like in many cases, I don't lose monetary value for like some of these things that are private to me. And it's private to me. I would like it to be kept sensitive. And I can't necessarily put a dollar value on that. But so if a person feels like they have lost their privacy, in some sense, it is similar to what they actually have. this is the challenge of this field is you can't make this entirely technical problem and so we're not going to try and make the social arguments in our paper we're just going to try and say there are lots of people who know how to make those kinds of arguments and we need to be careful when we're using these words because we're talking about a mathematical definition but it actually relates to real people and we should acknowledge that people exist in the world and so to kind of restate that it sounds like what you're saying is that applying the marketing terminology differential privacy to this partially differentially private pre-training scenario is confusing and potentially degrades the perceived value of differential privacy itself yeah and we have make a couple other smaller arguments in the paper that that talk about so you asked about the hospital example um it turns out that it may be the case that you don't actually get better at doing hospital classification if you take one of these big pre-trained models because the benchmarks people use for private learning are the same as the benchmark people use for standard learning and we already know that models like are just zero shot good at all the benchmarks right and so like to what extent like this is a period of like two or three years so you're doing this bad thing that you probably don't even need to do anyway it's not helping you solve your problem yeah like there's a period of two or three years where the state-of-the-art private technique was to take whatever the latest advanced best language model was and repeat the paper with a better language model and like you get like a new state-of-the-art privacy and you have done no privacy improvement and so like you should be careful this is the other problem it's like we should you know have these like benchmarks that actually measure private learning, not measure accuracy on tasks that can just be zero-shot solved by big, giant foundation models.
1:00:10And now probably one of the most, I don't know if popular is the right word, but one of the applications of federated learning potentially involving differential privacy is some of the new Apple intelligence stuff that was released. Have you looked at what they are saying about what they're doing and how it applies to the considerations you're describing here? Yeah, I haven't yet. I think this is like, once it gets into the weeds of like, how does the implementations actually match user expectations? I think this is one of the areas where people who study humans should really start to take a look at these kinds of things.
1:00:54We wanted to write this paper. So this is the paper was mainly driven by Florian, one of the people who actually wrote the previous paper with me, and Gautam, who's another professor. And the reason we wanted to write this paper is we were sort of mostly attacking some of our own work in this paper. Like we sort of felt like we can make this argument because we had worked on some things that did this kind of thing. And so we could reasonably convince people like, look, we have done this too. we understand that this is a problem and so you should like we're not like attacking you for doing the bad thing and so that's why we wanted to write this paper initially but we do very much think that this is the kind of problem that should be studied by people who actually have the ability to like ask ask users what you feel about this thing like we're not i mean you're saying people might confuse these things but people could actually go out and figure out if people do confuse these things and that's not what you've done yeah no exactly yeah like you could You could ask people, you know, maybe it turns out that you ask people, do you mind if a model learned on something that you put online, even if it wasn't intentional?
1:02:02Do you care about this? And they say no. And then fine, great, like we get to do it. Like, but like, you have to be careful about how you set up that experiment. Like I, I am not someone who feels qualified to study people, because I know that the way that which you phrase questions matters quite a lot. And someone could almost certainly answer this question much better than we can. And we just And a person may care. And so we should be careful on how we use these words. Well, Nicholas, thanks so much for taking some time to catch up and get us up to date on what you're up to and how you're thinking about the field.
1:02:35Yeah, of course. It's been great to talk to you. It's been a pleasure. Thanks so much. Thank you.
1:02:51Thank you.
From the publisher
Today, we're joined by Nicholas Carlini, research scientist at Google DeepMind to discuss adversarial machine learning and model security, focusing on his 2024 ICML best paper winner, “Stealing part of a production language model.” We dig into this work, which demonstrated the ability to successfully steal the last layer of production language models including ChatGPT and PaLM-2. Nicholas shares the current landscape of AI security research in the age of LLMs, the implications of model stealing, ethical concerns surrounding model privacy, how the attack works, and the significance of the embedding layer in language models. We also discuss the remediation strategies implemented by OpenAI and Google, and the future directions in the field of AI security. Plus, we also cover his other ICML 2024 best paper, “Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining,” which questions the use and promotion of differential privacy in conjunction with pre-trained models.
The complete show notes for this episode can be found at https://twimlai.com/go/702.




