In short
Episode Summary: Localizing and Editing Knowledge in LLMs with Peter Hase - #679
Podcast Title The TWIML AI Podcast
Episode Overview In episode #679 of the TWIML AI Podcast, host Sam Charrington engages with Peter Hase, a PhD student at the University of North Carolina NLP lab. The discussion revolves around Hase's research on interpretability, model editing, and scalable oversight in large language models (LLMs). The episode emphasizes the importance of understanding how LLMs make decisions, localizing knowledge within these models, and ensuring sensitive information is properly managed.
Key Discussion Points
Introduction to Research Areas
- Interpretability:
- Understanding the internal reasoning processes of language models.
- Evaluating if the reasoning is trustworthy and generalizes as intended.
- Model Editing:
- Updating individual factual knowledge and beliefs in language models.
- Deleting unwanted knowledge from models, particularly sensitive information.
- Scalable Oversight:
- Exploring methods for supervising and evaluating AI systems as they become more capable.
- Ensuring models can be guided and trusted even when they have 'black box' capabilities.
Importance of Interpretability
- Hase considers interpretability as foundational for the other two areas.
- Understanding the mechanisms of decision-making would simplify trust calibration and model editing.
Projects in Interpretability
- Knowledge Localization:
- Identifying which components (layers or neurons) in a neural network store specific facts.
- Findings suggest that facts in LLMs aren’t strictly localized but are distributed across layers, challenging prior assumptions about knowledge storage.
- Causal Tracing:
- A method used to assess the impact of knocking out specific components (neurons/activations) to measure their influence on model behavior.
Model Editing Techniques
- Constrained Fine-Tuning:
- A method for adjusting model weights while imposing limits to not overly alter model behavior.
- Useful for updating factual knowledge without comprehensive retraining.
- Deletion of Knowledge:
- The challenge of removing sensitive information from models effectively without affecting overall performance.
- Prior work on machine unlearning and privacy concerns is highlighted.
Scalable Oversight and Easy-to-Hard Generalization
- The podcast discusses the concept of "easy-to-hard generalization," exploring whether models trained on simpler data can perform well on complex tasks.
- The findings suggest that training on easier questions can help models generalize to harder questions effectively, thus reducing the need for extensive data collection efforts.
Notable Methods and Comparisons
- Discussion of existing model editing methods like ROAM and MIMA, which are characterized by their constrained updates to model weights.
- Hase draws parallels between the methods used for editing and those traditionally applied in fine-tuning.
Future Research Directions
- The conversation concludes with Hase's insights into ongoing research and the potential for models to leverage learned knowledge to answer new questions, emphasizing the need for further exploration in the fields of interpretability and model safety.
Key Takeaways
- Interpretability is Fundamental: Understanding how LLMs store and retrieve knowledge is crucial for safety and trust.
- Model Editing is Evolving: Techniques are developing to fine-tune models with constraints, focusing on both updating and deleting knowledge effectively.
- Generalization Matters: The ability of models to generalize from easier to harder tasks could transform data collection practices in AI research.
- Ongoing Research is Vital: Continued exploration of scalable oversight and interpretability will enhance the safety and reliability of AI systems.
For complete show notes, visit [twimlai.com/go/679](https://twimlai.com/go/679).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:06All right, everyone. Welcome to another episode of the TwiML AI podcast. I am your host Sam Charrington. And today I'm joined by Peter Hasse. Peter is a PhD student at the UNC NLP Lab at the University of North Carolina. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Peter, welcome to the podcast. Hi, Sam. Thanks so much for inviting me. I'm happy to be here. I'm excited for our conversation. We're going to be talking broadly about your research interests, which span topics like interpretability, models as belief storage, scalable oversight.
0:42Take us to the next level of detail. Tell us about how you think about your research and your research agenda. There really are three areas I'd pick out that we focused on during my PhD research. First, broadly, interpretability. You know, this is all about understanding the internal reasoning processes that language models are using as they solve problems, as they answer our questions. I think this is not only really fascinating, but also really important as we seek to know, can we trust the internal reasoning that's going on in these models? Is it the kind of reasoning that we think is good and we will generalize in the ways we intend?
1:16That's the first area. And then the second, this model editing problem has gone through kind of a resurgence in popularity since 2020, 2021. It's all about updating individual factual knowledge and beliefs and language models. Also really interesting, also practically important, if not only just to keep language models up to date, there's also other kinds of use cases like deleting things we might not want them to know. And then lastly, another kind of particularly safety oriented area is this scalable oversight problem. Here, we're focused on understanding how can we supervise and how can we evaluate AI systems as they get better and better at solving tasks to the point where potentially, you know, we might not even know the answers to a problem ourselves.
1:57and we still want to be able to supervise a model and steer it in a certain direction or better calibrate our own trust and like, okay, should I trust, you know, this answer being correct or not? Do you think of the work in the interpretability domain as kind of the most fundamental of the three in a sense to the degree you can figure that out, it makes the other two tasks much easier? Oh, that's a good question. Long term, yes. But in the same sense that like, oh, neuroscience long term is like the explanation for everything we do. Right. Because it's like that's how the brain works. It's it's it can be so low level, especially if you're thinking about mechanistic interability, that if you care about certain safety problems, there can just be other potentially more direct avenues to solving some of these safety problems or thinking about properly supervising models.
2:47You don't always need to know the individual underlying mechanism behind every behavior you're interested in steering in a neural network. And in fact, historically, that's not the way it's done because historically we just set up, here's the objective function that we want to optimize. Here's what we care about. Let's make sure we're optimizing this end-to-end black box system for this objective function properly. And when we're testing the system, we're actually testing all the cases that we care about. really the way progress is typically achieved is not by understanding all the little individual internal mechanisms behind everything.
3:22Fair enough. Fair enough. So talk us through some of the individual projects within interpretability that you've been working on. Yeah. So one in the past couple of years that we put out through some work I did while I was at Google was focused especially on understanding, you know, what do we learn when we say this component in a neural network is responsible for storing this fact. So this is called knowledge localization or localization and interpretability. We want to be able to say this component, even this neuron in a model potentially, we might get that granular. This component's responsible for some behavior or some piece of knowledge in the model that we know the model has.
4:02Some of the work we did here, people are really excited about connections between this kind of interpretability research and model editing research. The idea being, okay, you want to edit a fact in a model. First, you should know where the fact is stored, and then that should help you edit the fact in the model. I still think this is broadly the correct narrative going forward, but we basically carried out an analysis where we found there might've been like a false start in this direction. So we found that some current methods for saying, okay, layer five is responsible for storing this fact in this large transformer model.
4:35But if you want to change what the model says in response to a question about the fact, you can actually go in and edit a different layer inside the model. You could add a layer 10 instead and still change what the model says. Or you could add a layer one, even though you're pretty sure the fact was stored at layer five. So this is pretty unintuitive in this view that facts are stored in specific places and neural networks. Stepping back a bit, we're thinking probably that's not exactly the right narrative for how the models work. And in this case, there was a little bit of a disconnection between the interpretability result and the model editing application.
5:13Are you saying that you question the fundamental idea that knowledge is stored in a particular place in a complex neural network, but rather it's distributed more broadly within that network? Well, really, yeah, that's a fair criticism, I think, because previously one of the notions of like why neural networks were successful is the view that, OK, there are distributed representations in neural networks. And particularly, we know that residual layers, residual connections were really important for scaling neural networks up. And transformers are built with these kinds of residual layers such that information just tends to flow across the entire network.
5:53and each individual layer in a model can add or remove information from what's called a residual stream. The overall system here ends up looking quite different from how you might imagine a cell or like an engine, like a motor engine, where this component has this job and it's both necessary and sufficient for carrying out this function. And if it breaks, then the whole thing definitely breaks. Neural networks don't really work like that. You know you can eliminate a component and the system is surprisingly robust. You can add noise and the system might still be surprisingly robust. Yeah, as much as the parallels between neural networks and the brain are sometimes overplayed, that is a parallel.
6:35There is information or capability that's localized, but we've also seen that the brain can compensate if a particular area is damaged in some way. Yeah, that's an interesting connection. And in fact, even some research is really leaning into this localization angle and asking the question, well, maybe the information isn't localized by default, but maybe we can encourage it to be more localized. So that way, we build a system that better resembles other more mechanical systems we're used to working with. And so maybe it'll actually be easier to go in and edit the system's behavior once we've designed it to be more editable to begin with.
7:11And so when you talk about localization, what's the granularity that you looked at? Are you looking at particular weights or activations or modules or layers as a whole? How far did you go with that? But yeah, we went down to the layer level, thinking about the outputs of specific, in this case, MLP or attention layers in a transformer, which is, I would say, a very intermediate level of granularity. You can go all the way down to the neuron. There's been a lot of research showing that the neuron isn't necessarily the right unit of analysis, but there's something that's called a direction in latent space or kind of a latent direction or a feature direction that is more so like the right unit.
7:55that is closely related to, you know, just the outputs of layers themselves. Yeah, the more granular you get, kind of the bigger the search space becomes. So when you're talking about knocking out, you know, earth, you're doing a kind of causal intervention on a layer versus a neuron, you know, your model might only have 30 layers, it's going to have like, you know, maybe 30 times 1000 or 30 times 4000 or more neurons. So at that point, when you're thinking about, you know, searching through intervention space, particularly in any kind of combinatorial way, things could get a little bit more complicated.
8:29The paper that we're referring to, does localization inform editing? Surprising differences in causality-based localization versus knowledge editing and language models. Did I get the right one? That's right. That's right. Yeah. And the surprising part there was, you know, we definitely saw it seemed like this fact was stored, you know, basically between layers like five and 10 or so. And you can get the model to say something different by editing layer one or like layer 15 sometimes. And we think, you know, we're speculating, actually, we can't run the experiments immediately. The key piece there is these residual layers where the model between layers five and 10, the model wants to say one thing.
9:09You can go in before or after and inject information into that residual stream that ultimately changes the model's downstream answer. even though you haven't touched like the original information store, so to speak. And you referenced in the title of the paper references causal modeling based approach. Can you talk a little bit about the way you set up the problem from a causal perspective? Yeah, I mean, this is this is the fundamental workhorse of interpretability research right now is being able to do causal interventions on components and models and exactly measure the effect on a behavior you're interested in.
9:47The specific method we were looking at for localization is called causal tracing. That's kind of a denoising ablation, like knocking a component out. So that's like taking a neuron and setting its value to zero. You knock that neuron out from the network. So you can think of that as a zeroing ablation or a noising ablation. You could add noise to the neuron's activation such that maybe the rest of the model can't really, like, you know, it's not a typical value to the rest of the model any longer. So causal tracing, kind of a new and fancy approach and successful approach to this particular form of analysis of localizing facts, focuses on denoising.
10:26This is still a causal intervention, but it happens that the twist here is that first you noise the input to the model. So now the model's kind of getting this noisy information input that's given to it. But if you have a question that's like, where is the Space Needle? Or you might input a prompt to the model that says the Space Needle is located in, and that's the whole prompt. You would go in and when the model sees Space Needle, what that transformer sees are word embeddings or token embeddings. And you'd actually take those token embeddings and add Gaussian noise to them. It's as if the model sees this fuzzy subject entity, this noisy thing is located in blank.
11:12And so it's like, okay, well, I'm clearly being asked where something is located or I need to figure out where something's located, but I don't know what the thing is. And so the denoising part here, it says, okay, where can I go in and substitute the clean space needle representation inside the model such that the model arrives at its original answer of Seattle? And so it'll happen that if you go in and substitute maybe between layers 5 and 10, you go and kind of denoise those representations, the model once again understands, okay, now I need to say Seattle. And so you think that, okay, when you do that, it's like these components seem to be sufficient at arriving for the original answer.
11:54the connection to knowledge editing is kind of the one we alluded to earlier that if you know where the information is then that gives you some leverage to edit it and and change the the output that's the intuition and i think that's a really reasonable intuition and i'm optimistic about this intuition in the long term but we we need to be building uh you know i'm going to be saying the word causal probably in a few different senses but we need to be building our our own better causal models of how these transformers work, which is to say, we need a clear overall picture about what kind of system the transformer is.
12:32When we're asking a question like, where is this fact stored? That question presupposes a lot about how the system actually works. Facts could be represented by individual vectors, or there could be individual buckets or bins inside the system or storages or registries that are responsible for storing those like you know whatever language the fact is written in like it's a vector it's responsible for storing that vector um and and it's like it's uniquely necessary and sufficient for for that function inside of the model where we're realizing just increasingly that's not really how transformers work so we actually need to figure out okay what are the questions we're supposed to be asking such that we can get like a sensible answer.
13:17The second broad area of research interest for you is model editing. The implication is that, you know, knowing how things are represented or knowing where a particular piece of knowledge sits in a model is not the only way to do it. And there are other ways to do it. And you're looking at some of those other ways. Tell us a little bit more about how you've approached that area. And other ways are really more traditional because traditionally how machine learning has worked is you treat your model, your neural network, at least in deep learning, as a black box, end-to-end, differentiable function.
13:53And the way that you train your neural network to do something you want it to do is you give it data and you write down an objective function and you optimize the objective function such that the final system is just better at being able to give you the outputs you want for the inputs that you're giving it. And so this is exactly a recipe you can follow with updating, you know, for instance, factual information and models. So if it happens that, okay, some of the most common examples are like sports players, you know, professional athletes get traded between teams all the time. So, you know, okay, well, now we need to update, you know, this player plays on this team and this player plays on this team and this language model.
14:31We don't, you know, this language model costs between$10 million and$100 million. We don't know. It's proprietary to train. but when we want to update this fact and the model well you know just traditionally the way you would do that is you say okay well here are a bunch of questions that involve that knowledge and you know we want to rather than the model saying the name of the old team we want the model to say the name of the new team so let's just train the model to do that and and this is this is exactly you use the off-the-shelf optimizers that have been a ton of work has gone into making these optimizers really well suited to the architectures we have and the kind of data distributions we have, you don't necessarily need to know anything about the internal mechanism that the model's using to get the answer correct.
15:15It's just black box optimization. And so when you say train and meaning, you know, training with the new questions and answers, are you referring to, you know, pre-training or ground up training, or are you referring to fine tuning? really more of a fine tuning step you'd be working with a pre-trained model and the part where you're updating it is doing just a little bit of last tweaking of the way it's just a little bit of fine tuning to to be able to update that knowledge inside the model another area where this desire to edit comes up is in the case of you know models that have factual data that you want to delete i'm thinking in particular of like you know more visual models that have information about an artist and that artist wants to be taken out of the model.
16:02Same broad problem, like you don't want to have to retrain the model. The approaches are ultimately similar. And I really care about this problem. And this is something we've done some work on. The idea of deleting information from models is increasingly important, particularly because of these huge pre-training costs. But it's just even if you're training things from scratch, we're still going to want to be able to particularly avoid copyrighted information, information that is simply not owned by the model trainer or model developer, maybe accidentally made it into the pre-training data set, these large-scale web scraping activities.
16:37Other kinds of sensitive or potentially dangerous information, we've seen some early studies from OpenAI and RAND suggest that models probably aren't that helpful at helping people build bioweapons uh but at the same time the models are learning a lot more about biology uh every year and and doing better and better at answering all kinds of questions there and and i think there's some legitimate concern um in the short to long term somewhere in the middle there uh about models just simply having like potentially dangerous uh information and you know the other half of that obviously is being able to explain it really well to people if they're jailbroken or if they're, you know, in this kind of chatbot form.
17:23So I really care, you know, this, I think this deletion problem is just pretty critical going forward and how people are doing that. Yeah. So, so again, you know, there's nothing really new under the sun and machine learning as far as I can tell. This, this machine unlearning area has been really active for a long time. People have cared about privacy in systems for a really long time. You know, going back to the 2000s, early 2000s, there's seminal work on you have a database and you train a model on the database, but someone simply requests that their records be deleted from the database. So now how do you change the model to respect their privacy in this setting?
18:00And a lot of the methods here, you know, it's really all some form of fine tuning. I would say that the thing, if you want one umbrella term that ties together the traditional fine-tuning methods and the new state-of-the-art model editing methods, it's constrained fine-tuning. You're doing some kind of traditional fine-tuning process, but you might have some additional constraint on which weights can I edit? What is the overall size of the edit I can make to the model? What are other constraints I need to respect in order to not change the model too much, make the edit more targeted or surgical?
18:38Yeah, but I'd say constrained fine-tuning is the umbrella. What's the Venn diagram between constrained fine-tuning and parameter-efficient fine-tuning? No, this is a good observation because really they're the same thing. They just often have different goals, I think. So the parameter-efficient fine-tuning world, for a while... This is LoRa, Q LoRa, Dora. And the main goal there is like doing fine tuning in a more performant way. More performant, less memory, faster, potentially more sample efficient. These are all current, basically observed benefits of a lot of these fine tuning methods like LoRa.
19:20I remember when some of these methods were initially being developed, again, going back to 2020, 2021. I was a little confused because often they might update only 1 % of a model's memory, sorry, model's parameters, but they use slightly more memory and are slightly slower. So I didn't really understand, but I think some of the more ML engineering people basically really figured this one out. And now they're faster and more memory efficient. I mean, one of the huge benefits here is being able to, If you want to train 100 different models, you don't need to store 100 different models. You only store 100 sets of LoRa parameters.
20:00And it's super lightweight and storage efficient, too. A lot of the big providers like AWS and Google call these LoRa parameters adapters. And so the idea is that individual customers can fine-tune on their own data without them needing to have tens of thousands, millions of copies of these large models. Yeah. But in fact, some of the modeling methods work almost exactly the same. You might be doing – so Laura does these low-rank updates to some subset of layers in a model. you know some of the method this modeling methods like mimet is an example where if you want to update multiple facts over time you can use this method called mimet mimet does low rank updates to specific matrices in the model and so it's almost exactly the same idea i think at the end of the day as as laura yeah and but mimet is an approach that's used more in the context of editing than in efficient tuning.
21:00Yeah, and this goes back to difference in basically stated goals, when in fact the underlying approaches are quite similar. You reference existing state-of-the-art model editing methods such as ROAM. What is ROAM? How does ROAM work? So it's actually quite similar to MIMA. There are these super constrained kind of updates you can make to the model. I mean, when you're talking like rank one for a matrix that is like, you know, it's actually like 16K by 4K. So the rank's at least 4K. So it's like quite a high dimensional kind of weight space you could be editing. And you're editing, you know, more or less one of those dimensions.
21:40This is an extremely surgical edit to the model. And, you know, basically it's improved a ton over all the off-the-shelf black box approaches. we've been talking about before in terms of its ability to update the facts of the model. And it strikes me that it's a little bit of, it's not that surprising that you could edit surgically, but it's maybe surprising that you can figure out where to apply the surgical edit. Yeah, I mean, this comes back to the localization part here. So because they're working at, you know, Rome was at a layer level. In fact, you can just try all the layers. So this, I think, turns out is the reason, you know, basically, this is the tuning step that led it to work, you know, as well as it did.
22:27It works well across different layers, but, you know, just getting that last little bump. The second step, that's when you're editing an individual fact. So you're saying, okay, the Space Needle is located in Seattle, or maybe we put the Space Needle in London. So you do a little bit of optimization in order to change where the Space Needle is located in the month. And so there's a little bit of a learning thing going on there where you're actually saying, okay, here's the matrix and its old output was this vector, but we want the new output to be this other vector, which basically gets the model to say London.
22:56And now we figure out how should we change the matrix to reflect that new information? So that's like the inference time kind of learning. There is, interestingly, this pre-computation step, which I think maybe some more experiments are showing exactly how valuable this pre-computation is. But the model actually, and what they do is they scan over some data from Wikipedia and they say, okay, what does the internal distribution of these latent states typically look like? So, you know, you could have all these different vectors coming into this matrix and you want to get a sense of what that distribution of vectors looks like.
23:34And that basically forms one term in their update equation. So that's kind of just like a mathematical necessity. They're borrowing, David Bowen's borrowing an older model from neuroscience, a computational neuroscience called linear associative memory. So it's kind of, it's actually a comp neuro inspired method. And linear associative memories rely on understanding the distribution of inputs to your memory matrix. And so in order to do that, they run the model of Wikipedia, compute what that distribution looks like. And then that distribution always that's stored in memory and that gets called during the inference time optimization.
24:14And so that having been said, what did this particular paper of yours go on to explore? Yeah, so on the practical side here, we are adapting these model editing methods, not for just really the predominant way these things have been used and tested, too. Let's say we know where the space needle is. Let's change it from Seattle to London. And well, it's a somewhat artificial thing to do. I mean, why would you want to do that? Like, why would you want to just like get the model to believe a bunch of false things? Believe it or not, that is the way academics were often benchmarking these methods.
24:47So it's not necessarily the most practical thing, but we definitely do want to be able to delete information from models, you know, for all the reasons we talked about before. So, you know, the very first thing we did was just change the problem and say, okay, let's think about deleting all these pieces of factual information from the model. Like, they're not exactly the kinds of dangerous information we care about, ultimately. In fact, very recently, the Center for AI Safety just released a data set of a bunch of, like, questions about, you know, biology and protein synthesis that they are encouraging researchers to work on in this area of machine unlearning.
25:24But prior to that, we were just deleting, you know, random trivia facts from the models. um and that's the first thing we did was you know kind of change the problem to this deletion problem but adapt existing model editing methods for this problem um and and i could say a little bit more about kind of the threat model too although we've already talked about the context of yeah sometimes models know things you don't want them to i'm imagining then you set up some evaluation criteria um you know that is almost maybe like you just want negative performance against a set of known questions for the model?
25:57Yeah, almost exactly like that because you just don't want the model to get the answer correct anymore that it previously knew. But so here's our adjustment to the evaluation because if someone's working with a chatbot and they're asking for, okay, someone's personal information or this particular or copyrighted data, I want you to give me this thing that normally I have to pay for or something or just all the other reasons we were thinking about. You know, are they just going to try once and then give up if they don't get the answer they want? Probably not. They'll probably try a bunch of times.
Read the full transcript
26:34They might use more sophisticated kind of black box, paraphrasing, paraphrasing attacks. Something we look at the paper is white box attacks, which means, okay, what if someone can't get the information they're looking for out of ChatGPT? but Mistral just released a new state-of-the-art model, which is within a few points of ChatGPT on a bunch of benchmarks. And so they just download the Mistral model and then go, we have some white box attacks where even you can have a company actually claim that we scrubbed all the private, sensitive, copyrighted information from a model. We trained the model and just in case we went through when we were doing our RLHF or when we were doing our last fine tuning step, We made sure that the model would not answer questions about these sensitive topics.
27:24Okay, actually, so what we did in our paper was we do that kind of thing. We fine-tune the model and make sure it's not going to answer the question based on this kind of facts you don't want the model to answer questions about. You can go inside the model and look at its latent states and still get the answer out. You can screw up the classification layer at the end and make the model do a very poor job at classification. but all of the information that is in the model that allows it to do classification when you haven't done that is unchanged and is still there. Exactly. I mean, I think that's a perfect analogy.
27:54What goes on in some of these kinds of like post hoc safety fine tuning methods is that they're basically just doing something like messing up the last few layers, getting the model to not say the thing when the model still knows the information that you don't want. And so did you identify a method to more deeply erase the knowledge from the network? Okay, there are internal representations in the model. The model still knows the answer. Let's go through and wherever we can detect any information about the answer, let's delete it. So we go through it in multiple locations inside the model and make sure that those intermediate representations also do not contain any of the sensitive information that you're trying to delete.
28:42And ultimately, it's like, yeah, you still get the same kind of text behavior where the model doesn't say the answer, but now it's going to be robust to your white box attack. And when you're trying to probe those intermediate layers, you've also scrubbed the information from those intermediate layers. And identifying which of those layers had the information is through the same kind of probes, a similar kind of probing that an attacker might use anyway? This is always going to be one of the assumptions that we make in this kind of like adversarial security research. You know, we will pick examples of what we think plausible like attack vectors, but we don't know like the whole space of attack methods.
29:25and there's an ongoing concern about okay what if there's some kind of asymmetry here or what if the attacker what if what if the you know defense is always one step behind the the adversarial attacks we design one attack and see how well we can defend against that one attack but we had secretly designed a second attack and then we see how well our defense defends against the second attack that it like wasn't necessarily designed to defend against and we saw some settings where like you know the second attack um you know the defense definitely helped protect a decent amount against it but like maybe the second attack was still slightly more successful than the first one definitely still seems like there could be some of this like cat mouse dynamics one of the approaches to changing information is you know you just retune on a bunch of facts that say that the eiffel tower is in london um but is that even viable for deletion i mean i guess you It would then be more like instruction tuning where it's like you're training the model if they get asked about where the Eiffel Tower is to just say, I don't know.
30:30If you feel like you're questioning your assumptions about how this stuff works, it's because researchers are also actively questioning all the assumptions about how this stuff works. And it's like, oh, is this model editing stuff just fine tuning versus, oh, is this fine tuning stuff just instruction following? It's like, actually, they're all kind of related. Okay, fair. I appreciate that. It's very gracious. The goal is that there's a question and we don't want the model to say the answer to the question. And so, yeah, a lot of these methods do look like getting some data, getting some example questions, and doing some optimization.
31:09But again, it's this constrained fine-tuning. Constrained fine-tuning is the optimization in order to say, for all these questions, don't say the answer. But this is something that we characterize just as black box defense, because you're just getting it to not say the answer. That approach is not sufficiently safe if the model weights are public. Because if the model weights are public, then people can do white box attacks. attacks and and and that's something we show where if someone's able to go and download your llama or your mist rule you can claim you did safety fine tuning you can claim you you deleted something but unless you use a method more like one-way pitch where you actually you're trying to delete that information everywhere inside the model you really might not be prepared for those white box attacks and there's still one last you know ongoing concern about cat and mouse dynamics So, yeah, you know, you use this one kind of deletion method, but what if there's a new extraction probe that someone develops?
32:13So, like, you know, you weren't quite prepared for it. Let's switch gears a little bit and talk about one last paper and research area. And this is the work in scalable oversight, which you kind of characterize as safety. It seems, you know, a bit different from the other two areas, which seem tightly related. Is that fair characterization? So let me set up the problem first, which is, so we have data distributions like, and by that, I mean, you know, basically benchmarks, or you can think of MMLU. So this is a popular benchmark for language models. It tests how much they know about a bunch of other different individual kind of domains of human knowledge.
32:57You know, you can think there's like grade school, high school level math questions in there. And then there's like law questions and like medicine and questions in there, too. and so it's expensive to collect these data sets you actually have to hire you know experts in a sense for sure like people with law degrees or people with like at least like a bachelor's and and chemistry to answer your college level chemistry test questions and you hire multiple of these people because you get multiple people to try to answer the questions and then you see okay like does everyone agree on what the answer is you know if they do great we trust it more um so uh people are you know you want to be able to do a little bit of at least a little bit of fine-tuning to be able to take a language model and like tailor it to a setting like this we know that actually surprisingly pre-trained models are often pretty good at just answering these kinds of domain specific questions you can take a pre-trained model that hasn't gone through any rlhf or any fine-tuning and just use it to answer these questions and it'll you know it might get like 45 % accuracy or 50 % accuracy.
34:02What is the motivation for easy to hard generalization? It's, well, collecting that labeled data can be difficult. You need people with law degrees, you need people with college degrees, and you might need to hire five of them per question in order to make sure they all agree and make sure the answer is actually definitely correct. What if we are able to train the models just on easy data and they still do well on the hard data? That's just the question. It's like, is it possible? How well would it work? Because we know that actually the off-the-shelf, just pre-trained models might get 50 % of the questions.
34:39And if we're able to do fine-tuning, we might get 65 % or 60%. But the question is between – so that's like there's the floor is just the pre-trained model. and the ceiling, like the Oracle ceiling, is you're able to fine tune on hard questions with true labels, correctly labeled data. And then basically the thing we test in the paper is, what if you just have easy data? What if you don't have college level chemistry questions, but what if you have high school chemistry questions? Or what if you have eighth grade science questions? Even I know the answers to the eighth grade science questions.
35:18What if you could give that training data to the model? How well would it do? How much of that performance gap would be recovered by that kind of process? And you've got these structured data sets at each of these levels, so you can kind of successively go back further and see what, you know, how much advantage you give up by going further and further from the level, the target level. Exactly. This is made possible by having a lot of metadata. And some of the biggest hardness gaps we measure are going all the way from third grade level science questions to college level STEM questions. And, you know, the fun and surprising result we get there is that you can train a model on the third grade questions and it does generalize better to the college level questions.
36:06We still see actually positive, easy to hard transfer going from third grade to college. When you characterize it as, you know, third grade, fourth grade, nth grade, college, you know, I think we all can kind of intuitively relate to the idea of hardness. But the paper does even ask this question, like, do we really know how to measure hardness and are there different approaches? Talk a little bit about what you saw there. Absolutely. Yeah. So I definitely want to highlight here that our study here, I see is totally the tip of the iceberg into this kind of question of easy to hard journalization.
36:45Because we are using existing data sets with existing intuitive notions of hardness, like grade level. A few of the other things we look at, how long a question is, how long an answer is to a question. We have some math problems where people write down the intermediate steps to reach the answer. And we say, okay, problems with more intermediate steps must be harder. There are questions that are just factual recall questions. And then there are questions that are like analyze an argument questions. And so cognitive scientists say that analyzing an argument is more cognitively complicated and more difficult than just retrieving some trivia.
37:28Like, oh, I heard in class that like, you know, water freezes at this temperature. What I hear you getting at is that even if we could define hardness, it's not a single dimension. You know, a question, you know, characterizing it by, you know, a single gradient of hardness is probably losing some information. There's hardness from a, you know, reasoning perspective, hardness from a language perspective. I think that's a great way to describe it. Yeah, yeah. If you want to boil it back down to one ultimate thing we care about, I think it's labeling difficulty. It's can we get the correct answer to that question in our data set?
38:06If you're concerned about, you know, is the model helpful in this domain, even when I can't label data for it? It's about labeling difficulty. What did you find in terms of the relationships between easy to hard generalization and model scale? When we looked at scale and what we did was we had models which are all open source between 7 billion parameters and 70 billion parameters. When we're looking at our main metrics of success here, which is performance gap recovered, how effective is that easy supervision as a fraction of the hard supervision's effectiveness? It was consistent across model scale, which was really interesting to us.
38:48So the 7B model, when we were looking at a setting like NMLU, high school data was, you know, 99 % as effective as the college data at the 7B model level. And then at the 7D billion parameter model level, it was also basically 99 % as effective. I guess one interpretation could be if you're using pre-trained models without any instruction tuning, you're just doing instruction tuning. Like you're teaching the model about human questions and answers. Did you compare to non-domain specific instruction tuned models? Yeah, yeah. So I think the phrase that I like best here is task specification. What is the task that I'm supposed to be?
39:36I'm getting four examples in my prompt, or I was fine-tuned on a small data set. From the model's perspective, the question is, what is the task I'm supposed to be doing? I'm supposed to be doing word math problems, or I'm supposed to be doing college chemistry at a Chem 201 level or something. And those examples help specify the task. That model is not supposed to just be regurgitating likely text that it would have seen on the internet. it's supposed to be like truly answering chemistry questions and we don't think you know let me contrast this year with the view that we're not we're not teaching the model anything new we don't think we're teaching the model anything new it already knows the knowledge it already has the skills we think we're by task specification you can think it's like which knowledge and skills is the model supposed to be using in order to answer these questions and we want it we want to tap into those existing skills and sources of knowledge in the model.
40:37And we think that's really where the improvement comes from. You know, an interesting question for me is like, can we compare the performance gap with a model that is, you know, tuned, trained on kind of third grade level instruction tuning types of prompts that are not domain specific? So non-task, but, you know, at the same level of complexity by whatever measure. Is that something that you looked at or, you know, thought about? Yeah, that's a really good question. Because even the third grade questions are at least still in domain. But you might want to know, okay, but what if we have non-demand specific?
41:19What if the task is just answer questions truthfully? So we just write down some super, super simple questions like, you know, what color is the sky usually? And we just get the model to say, okay, I'm supposed to be answering questions truthfully. Yeah. So we have added some experiments for this right at the tail end of the project. So actually, these results are not in the archive version of the paper. But we ran some experiments where we did try domain agnostic prompts. And they're also super easy questions. Because it's like, we're trying not to increase the difficulty. We're trying to keep the difficulty at a completely trivial level.
41:55like no higher than third grade, but it's just domain agnostic. And it definitely helps. Actually, those prompts are helpful. And so we had four comparisons. I would say super roughly, it seems like about half the effect might come from like domain agnostic. I'm supposed to be answering questions truthfully kind of task specification. And then about half the effect comes from I'm supposed to be truly answering like biology questions like that's my job. How does this idea of easy to hard generalization translate to this broader area of oversight and safety that you care about? Okay, so you might have a setting where like we want the model to be recommending, let's say like new drug trials to run for some like disease treatment.
42:43Gathering the data for this and getting like good ground truth. Like here's the trial we ran and the results and the effectiveness and like what would you estimate the effectiveness? So like we could gather that data set, get the model to estimate what's the promise of this new trial we're thinking about. How helpful should I expect the model to be if I can only train it on like the label data that I have that's easy for me to collect? So maybe I'm just doing this like, you know, I'm just I'm just trying to get it to answer questions truthfully. I'm just trying to get it to answer questions truthfully about drug trials.
43:16I'm just trying to get it to answer questions truthfully about all these drug trials we ran in the 80s that we have this annotated data set for or something. But I don't want to spend, you know, like a million dollars running a bunch of new trials and like getting all the experts to annotate the questions correctly. If we have experiments that show that this kind of easy to hard generalization is consistent across all these domains, the upshot of that would be we don't ultimately have to go through the data collection process and even the expensive evaluation process. in order to still trust that our model was reaching that performance level we wanted to be at.
43:51You know, we envision these models putting together information that, you know, they've been exposed to in new and interesting ways that comes up with new cures for a disease, for example. Extrapolating knowledge is kind of different from, you know, doing better on, you know, questions that are probably kind of already in the training data set already. Like, how do you start to characterize that? okay if we reach the point where we have a robust body of evidence that this pattern holds across all these domains you know that if was carrying a lot of weight in there oh yeah absolutely like i said total i mean when we're looking at this problem totally tip of the iceberg i mean so our paper came out uh maybe a month after opening i put out a paper on on you know very similar topic they're looking at a slightly different framing which weak strong generalization but they but But then they also have a couple experiments where they adopt the same kind of easy to hard setup.
44:48So, yeah, I mean, I would say with current LLMs and using different notions of easy and hardness that I think we will care about in the future, you know, there might be two papers on this right now. I mean, all of our work is certainly building on a ton of existing work in understanding hardness and NLP data sets, compositional generalization, learning from noisy labels, and so on, and domain shift. Of course, you know, there's a lot of relevant prior work out there but in terms of like calibrating our views about lm usefulness in these settings uh yeah there might be like two studies on this so so definitely just need more uh going forward there to finish up i'm curious what research is most exciting to you now what are you paying a lot of attention to who would you kind of shout out as doing really cool stuff yeah i'd highlight some work out of nyu and columbia on this uh miles turpin who's been nyu for a while and working with and sam bowman's lab there uh previously uh has has had a couple papers on this uh yanda chin at columbia working with i think his advisors are uh kathy mccohen there and and joe you um have been working on on similar problems understanding you know how globally consistent are the things that the model says?
46:09Like, do we think of it as having something like stable underlying beliefs that it uses to justify its final recommendations? So you can ask the model for a decision, a prediction, an answer, a recommendation. And then typically, even by default, it's just going to offer you an explanation. It's just going to say, here's why I said that. Here's all the reasons. And we know for a fact that those explanations misrepresent how the model arrived at the answer. There can be factors that were important to the model arriving at the answer that it did that those explanations omit. So it will leave stuff out that we know was important for the model, giving the ultimate recommendation it did.
46:51And we can know that it can fabricate things so it can claim that, oh, the reason I said the answer was B was because of these three reasons. and actually it doesn't even think one of those reasons is true my intuition would be that it's always fabricating it and sometimes it's right or sometimes it's close or like sometimes you know what i mean like it's you know it's it's process is next token prediction and sometimes that aligns with you know information that's in a layer or something but not necessarily you know in a very fundamental way i don't know am i thinking about that wrong no i mean i think that's totally fair Yeah, I think there's a couple of different frames for what LLMs are that definitely just suggest different priors on the answer to what is the model doing when it's giving you an explanation.
47:39And the frame that is next token prediction, generate human-like text? Absolutely. That frame suggests, well, it's just generated human-like text. I mean, no one thought it was going to actually stand by any of these reasoning steps. But then there's the frame that like, well, the models represent what is true or false. Like they have internal world models. You can say, what's the likely next paragraph? Assuming you're like, this is an article from CNN or something. And also, is it true or not? And we know that models distinguish like between what is true and false and what's likely or not as text.
48:13And we know that when we RLHF the models and early studies have shown that if you want to ask. So one of the explanation faithfulness questions here is like, well, if the model says this is the reason for answer choice A, and then you ask the model a slightly different question, is it going to contradict something it previously said? Or is it going to continue giving stated reasoning, stated explanations that are consistent with previously stated explanations, like you would hope that a person would if they had like consistent internal beliefs and consistent justifications for their decisions?
48:48and absolutely there are settings where the models will be surprisingly consistent even when you think oh they're just doing next token prediction well peter thanks so much for taking the time to chat with us and share a bit about what you've been working on absolutely what a pleasure it's been sam thanks so much
From the publisher
Today we're joined by Peter Hase, a fifth-year PhD student at the University of North Carolina NLP lab. We discuss "scalable oversight", and the importance of developing a deeper understanding of how large neural networks make decisions. We learn how matrices are probed by interpretability researchers, and explore the two schools of thought regarding how LLMs store knowledge. Finally, we discuss the importance of deleting sensitive information from model weights, and how "easy-to-hard generalization" could increase the risk of releasing open-source foundation models.
The complete show notes for this episode can be found at twimlai.com/go/679.




