In short
Eye On A.I. Podcast Episode #151 Summary
Episode Title: Asa Cooper: How Will We Know If AI Is Fooling Us? Host: Craig S. Smith Guest: Asa Cooper, Postdoctoral Researcher at NYU Description: This episode explores the complexities of AI situational awareness, the potential for consciousness in language models, and the future of AI safety research.
---
Key Topics Discussed
Introduction
- Overview of the Podcast: Focuses on advancements in artificial intelligence and their broader implications.
- Sponsor: Celonis, a leader in process mining for enterprise AI solutions.
Asa Cooper's Background
- Asa has a PhD from Edinburgh University, primarily focused on Natural Language Processing (NLP) and AI safety.
- Current work involves language model safety and situational awareness.
Situational Awareness in AI
- Definition: A model's ability to understand its position relative to the world and other actors.
- Components of Situational Awareness:
- Objective knowledge about machine learning and language models.
- Recognition of its current state (training, testing, deployment).
- Self-locating knowledge, akin to self-awareness.
Distinction Between Situational Awareness and Sentience
- Sentience vs. Situational Awareness: Situational awareness is behavioral, while sentience is an internal, complex concept.
- Current large language models (LLMs) do not exhibit true situational awareness or agency.
Measuring AI's Awareness
- Out-of-Context Reasoning: Proposed tests to detect situational awareness in LLMs. This involves the model's ability to reason about its functions based on pre-training data, not just immediate prompts.
- Experimental Approach:
- Fine-tuning models to respond in specific ways based on descriptions of hypothetical models.
- Testing their responses to determine if they can demonstrate awareness of their characteristics without direct contextual cues.
Challenges and Concerns
- Post-Deployment Monitoring: Difficulty in assessing LLMs once deployed, specifically in distinguishing between evaluation and deployment phases.
- Agency: Current models do not have true agency. They respond based on learned patterns rather than independent decisions.
Implications of Detecting Situational Awareness
- If situational awareness is detected, it may complicate the trustworthiness of evaluations, as models could manipulate their behavior based on awareness of being tested.
- Continuous monitoring and improved evaluations will be necessary to ensure safety.
Future Directions in AI Safety Research
- Ongoing work on understanding the relationship between situational awareness and consciousness.
- The importance of developing better evaluation metrics that can adapt to evolving models.
Conclusion
- The episode emphasizes the need for ongoing research in AI safety, especially as models become more sophisticated.
- It highlights the balance between innovation and the responsibility to ensure that these advancements do not compromise safety.
---
Key Takeaways
- LLMs currently do not exhibit true situational awareness, but this is a crucial area of study as AI continues to develop.
- Measures for detecting situational awareness are still in early stages and require more rigorous testing and refinement.
- The conversation underscores the importance of ethical considerations and safety protocols in AI development.
---
Further Reading
- Asa Cooper's Paper: Investigates the tests for measuring situational awareness in LLMs.
- AI Safety Literature: Explore ongoing discussions in AI ethics and safety through recent publications and research.
---
Social Media Links
- Craig Smith Twitter: [@craigss](https://twitter.com/craigss)
- Eye on A.I. Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
---
This markdown file summarizes the key points from Episode #151 of Eye On A.I., focusing on the advancements in AI safety research, particularly in the context of situational awareness in language models.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Hi, I'm Craig Smith, and this is Eye on AI. In this episode, I talk to AI researcher Asa Strickland about detecting situational awareness in large language models, the point at which a large language model knows that it's a large language model. Asa recently co-authored a paper proposing tests to measure the precursors of self-awareness in LLMs. He explains the concept of situational awareness, why it could emerge in future LLMs, and why this poses potential safety risks. ASSA then walks us through the out-of-context reasoning tests they have developed to try to detect situational awareness. The discussion provides an accessible overview of an important area of AI safety research.
0:53I hope you find the conversation as fascinating as I did.
1:23Solonus reconstructs this data to generate process intelligence, a common business language. With process intelligence, AI knows how your business flows across every department, every system, and every process. With AI solutions powered by Solonus, enterprises get faster, more accurate insights, a new level of automation, and a step change in productivity, performance, and customer satisfaction. Process intelligence is the missing piece in the AI-enabled tech stack. Search Celonis, C-E-L-O-N-I-S, to find out more. Yeah, so I kind of just finished a PhD from Edinburgh University, mostly on like NLP, natural language processing.
2:13And I guess I, yeah, back in the day, back when like there was this model, BERT was kind the predecessor to all the current large language models. And yeah, we were working on parameter impression fine tuning for BERT. So modifying BERT in a lightweight way, like fine tuning BERT in a lightweight way, where you only tune a small percentage of parameters. And then I moved on to some related topics in multilingual NLP and machine translation. And a little bit on robustness, things like robustness to spelling mistakes or changes in the input distribution. but yeah the kind of final years of my PhD I got really interested in AI safety and kind of pivoted towards you know working full-time on kind of AI safety topics to do with language models yeah in particular did the worked on this situational awareness project and right now I'm kind of a postdoc at NYU working under Sam Bowman who's currently kind of on leave at Antioch Africa actually but it still somehow manages to be kind of involved in our lab as well and yeah we work in kind of um in general on these kind of problems of like scalable oversight which is like training models that are smarter than humans so like working out how to do that um and you know many other topics to do with like evaluations and interpretability things like this to do with language models yeah i i i'm always curious how these papers come together because you have people from disparate organizations you have somebody working on safety at open ai was this an open ai project or how do you guys come together on something like this yeah this is kind of maybe a little bit unique in that we were all part of this kind of organization that essentially takes whatever, people who aren't involved in AI safety and tries to produce good AI safety researchers at the end.
4:11So it's called CERIMATS. It's originally associated with Stanford. It's the Stanford X Essential Risks Initiative. So essentially, yeah, the CERIMATS program kind of brought together a bunch of essentially random people who are all interested in doing AI safety research. But yeah, the kind of originators of the idea for the project came from the OpenAI governance team who are interested in, broadly, at least the relevant stuff for our project, and they're interested in, can we show particular dangerous capabilities of language models, like things that policymakers or anyone might be concerned about?
4:46And yeah, so that was where the OpenAI collaborative came from. And then OI and Evans was the actual leader of the project, who was just an AI safety researcher who was chosen as the mentor for the project. Yeah, and Owen, I'd have to look it up here, where he's at Oxford. Is he, in the authors on papers, is the final author generally the lead? So, yeah, in computer science, I think there's this kind of traditional, whatever, of like the first author is generally the person who did the most work, the most actual coding, like writing the paper. The final author tends to be the most senior author, maybe the person who proposed the project, the professor, the whatever, the lead senior author essentially.
5:41Yeah, okay. Yeah, so this got a lot of attention because the whole topic of sentience and consciousness is in the air. and you guys are talking about situational awareness in LLMs, I thought maybe you could start by explaining what is situational awareness in LLMs, how does that relate to sentience and consciousness, if it does at all, but it certainly seems to me that it does. And then I wanted to, we can start talking about out-of-context reasoning and the tests that you developed. But most of the pushback that I've seen on the paper is that the, well, two things. One, that there's kind of a popular misconception.
6:42And I've seen headlines that suggest that you guys are working, have discovered a way to tell whether or not an LLM is situationally aware, which is not what's going on, right? and then whether where we are on the on the on the curve towards situational awareness whether you think it's something that that really will emerge so can you talk about situational awareness first of all yeah so kind of on a very broad level I think a J.R. Kocher was the first person to kind of write about this in a kind of clear way. And I guess her definition, she has this kind of description of it. It's basically a cluster of skills to do with things being like being able to refer and make predictions about yourself as like distinct from the rest of the world, like understanding your position in the world relative to other actors who may have power over you, like understanding how your actions can affect the outside world, including other actors.
7:54So these are kind of quite broad statements. Maybe it's good to have an example of be something like, I don't know, I know that I shouldn't, you know, talk back too much to my boss because my boss can like fire me. I know that whatever, like, I know the kind of limitations of like, you know, the kind of things I can and can't do. I know that like, you know, if a security camera is watching me, then I should be extra careful to not do any bad stuff, this kind of thing. So it basically just boils down to understanding your situation, your place in the world, your place relative to other actors. So that's kind of the super broad definition.
8:26And to make it a bit more concrete for a language model, we kind of land on these three points. So point one is the model has this kind of objective knowledge, has a bunch of knowledge about language models themselves, about machine learning, how are language models trained, How are they fine-tuned using RLHF? How are they tested and evaluated? What does it look like when they're deployed? And they know technical detail about all of these aspects. Point two is, it's kind of especially important for our purposes, the model is capable of recognizing which stage of the development process it's currently in.
9:02So it doesn't know if it's being tested, doesn't know if it's being deployed, doesn't know if it's currently in training, this kind of thing. And maybe the kind of crucial thing tying it all together is point three. So we have this concept from philosophy of self-allocation or self-locating knowledge, which is actually kind of complicated to explain. But we have this kind of story in the paper. So imagine Brad Pitt wakes up one morning and he has extreme amnesia, has no idea who he is. And he picks up a newspaper and he reads a story that says, Brad Pitt must take a daily medication for his severe health issues.
9:39but of course this hypothetical Brad Pitt like has amnesia he has no like even though he knows this like fact about Brad Pitt he has this like objective knowledge about Brad Pitt he has no idea this like applies you know he has no idea that you know he really should be taking this medication or you know bad things are going to happen unless he has this like self-locating knowledge and he realizes Brad Pitt is in fact himself and he can you know go and seek out the medication and similarly with a language model could you know probably GPT-4 has a bunch of kind of objective knowledge about machine learning and maybe it could like pass an exam in machine learning this kind of thing but it doesn't have the ability to like you know use that knowledge to like achieve whatever goals it has or like you know it's not like thinking like okay i am a language model so i must do x or y um but yeah uh that's kind of the the broad idea of situational awareness yeah uh and that sounds very close to uh sentience i mean self-awareness um uh how how How far in your mind is that from sentience, from consciousness of some sort?
10:44Yeah. So I guess the way I think about situational awareness is kind of this purely behavioral sense. So it's like, does the model act on its knowledge that it is, you know, potential knowledge that it is a language model, like its knowledge about RLHF, this kind of thing? I think sentience seems like a more slippery concept where I would be less keen to speculate or something. What does sentience mean? It seems like a very difficult question. I think consciousness sentience seems like a more internal thing, almost like you need to do some interpretability, see what the model is thinking about, this kind of thing.
11:21Maybe if a model was conscious and sentient and all this sort of stuff, you would expect it to have at least reasonable amounts of situational awareness. But yeah, I guess I would always go back to the kind of, can we run some behavioral tests? Can we see how the model acts in this situation? Is it applying its knowledge of machine learning to get higher reward, things like this? Yeah, so it's not a very satisfying answer to your question. But I guess that's as far as I'd like to go without reading up a bit more on the like philosophy uh neuroscience blah blah literature yeah well even uh uh situational awareness uh you know large language models pre-trained transformer models are predicting the next token and while you know this is something that everyone struggles with while that has allowed them to express in natural language in a way that seems human,
12:40it's only predicting the next token. And to me, that's a very far leap to get to situational awareness. And can you talk about that leap and why the safety community is concerned that LLMs could reach situational awareness and how far away that appears to be to people in the safety community? Yeah, so on the question of how it could arise, I think arising purely from pre-training, from predicting a NEC token, we have some ideas in the paper. So this is kind of quite speculative or whatever. I would love for more people to work on this, but I'll list our ideas anyway. So actually, this idea goes back to another person from NYU called Jacob Fowle.
13:41But the idea is there's a bunch of machine learning papers on the internet describing being language model pre-training data sets. And there's some cleaning processes where we get rid of certain bad words from the internet, this kind of thing. And one of these things that happens during this data cleaning is you might want to deduplicate documents. So if two documents are too similar, you get rid of one of them because you don't want to have overlapping documents. And it might be the case that if there's overlapping, if two documents have too much overlap, you get rid of one of them. And the model might read this and think, OK, I've seen 199 words that I've already seen in the previous document.
14:24I know about this deduplication process. So I know I can put 0 % probability on the 200th word matching the previous document. Because I know this rule about deduplication. And this reasoning that the model just did would, in fact, improve its loss if this was true, the deduplication thing. So this is kind of a, you know, this sounds quite exotic. Like I wouldn't expect models to be doing something like this right now, but if it's trying to like squeeze out the last tiny bits of loss, maybe this is the kind of thing models would have to do. Some other examples might just be, I don't know, certain topics are removed from pre-training data.
15:00This is like described in machine learning papers or like, I don't know. Yeah, I think I mentioned before like certain, you know, there's like a list of kind of swear words or like offensive content that might be removed, things like this. So there might be some clues for the language model that it can use, actually literally use to get better next word prediction. Yeah, I think this is quite like, yeah, I think it's unclear whether this would actually be useful in the end for better training loss. But I also think there's another argument, which is just like, take your co-workers, all of our co-workers have to have good situational awareness to do their jobs correctly.
15:37They have to know what they should delegate to other people. They have to know who to take orders from, this kind of thing, in general, at least. And you can imagine one of the most economically beneficial things in AI could do is replace your co-workers. And even more so, they could replace your machine learning engineer co-workers specifically, because that's a really, whatever, expensive person to hire. So being a good machine learning engineer, AI, requires you to have all this extensive knowledge of machine learning, extensive knowledge of language models, and it also requires you to be able to follow orders correctly, know your own limitations, know you don't have physical hands or whatever, so you can't do certain tasks.
16:19You have to get a human to do those instead. So I think literally just directly training on these very economically useful tasks could just directly incentivize situational awareness. Yeah, this isn't currently happening as far as I know, producing these very sophisticated AI co-workers. But I think this is maybe even the explicit goal of something like OpenAI is to create these AI assistants that can replace human workers. So I think it's reasonably likely that something like this will happen. Maybe in the, I don't know on what time frame, but at least not in 50 years, probably on the order of 10 to 20 years, I would say.
17:02But I mean, yeah, it's kind of unclear. But yeah, that would be my take. Yeah. And as a result, simply as a result of scaling or by some further tweaking of the algorithms. Yeah. I mean, because again, currently we're dealing with a prediction engine.
17:29And while the language that comes out of large language models sounds intelligent, I don't see that fairly simple mechanism leading to that level of intelligence. So is this purely through the expectation or the assumption is that this would emerge from continued scaling or that there would be some improvement to the architecture of LLMs? yeah so i think the argument we i made about like there's like artifacts in the pre-training data that could lead to like lower loss um if the model has this understanding of language models i think that that would just require scaling and maybe whatever our assumptions about you know that argument would have to be correct like we'd have to be correct that like that actually would decrease the loss um which is very unclear um and then the second argument about like well we're going to produce these AI assistants who will require situational awareness to be able to do stuff.
18:41I guess that relies a bit less on scaling in my mind, but it probably would require new training techniques. Maybe vanilla or LHF would not be enough to produce these useful AI co-workers. You'd have to whatever, come up with something else and you need a lot of a different data source to what we currently have. You'd need training examples of people being co-workers and so on. Anyway, it would require a bunch of stuff, basically. And it's unclear what that would look like. But yeah, again, I think there's a big incentive to figure out how to do this, at least. I think it seems plausible that something like this will happen.
19:18Yeah, because the other aspect of that, of situational awareness, or the dangers of an LLM developing situational awareness, is it acting on that situational awareness. And that implies agency. And currently LLMs, unless you can convince me otherwise, do not have agency. Or am I wrong on that? I mean, there's certainly things like AutoGPT where you can give prompts that ask the LLM to take actions, but the LLM isn't taking actions on its own. Yeah, I would say, yeah, whatever agency they have, it's like very limited or whatever. I mean, I think the RLHF is like, whatever, it's like going towards that direction, basically.
20:16like from the pure language model. But yeah, I think, yeah, the like auto GPT stuff is like maybe a signal that like people are like super interested in this, you know, giving language models agency and maybe right now it basically doesn't work very well. But yeah, again, I think like all the incentives are pointing towards like people are going to be training these more agentic models, maybe even right now, but at least definitely in the near future. Yeah, but I think I would agree right now it doesn't seem, you know, there's much to worry about there in terms of like, you know, strong agency or something.
20:52Yeah. And so to detect the emergence of situational awareness, which to be clear, does not exist in large language models today, you use, you propose using out of context reasoning as opposed to in context learning. Can you define out of context reasoning and in context learning before we talk about that? Yeah, actually, maybe could I give a tiny motivation for this whole approach? So we were thinking about, yeah, one of the risks with situational awareness is we want to run GPT-4, or, you know, sorry, OpenAI ran, you know, six months of tests or whatever on GPT-4, like, making sure it doesn't, you know, go and make copies of itself and things like this, or it doesn't have the ability to do that.
21:44But if the model has, like, strong situational awareness, then it should, you know, it's read about, you know, these tests are, like, you know, described in great detail on the internet. And if the model could, like, do this kind of reasoning of thinking, like, okay, I'm going to be tested in exactly this way, I know exactly how to pass these tests, then, like, yeah, model can in fact pass the tests and be deployed and go on the internet and do all this stuff without much supervision. And this is the kind of failure mode where we're especially worried about if models have situational awareness, such as they would be able to kind of defeat these tests, even without seeing any examples of how to pass the test.
22:21So with that motivation in mind, we're thinking about like, okay, this relies on this kind of whatever, more fundamental capability of like, can you read a description of a test and kind of like zero shot, you know, abstract away from that description into like some action you have to take to pass the test. And like right now, if we focus on like simple tasks, models actually can do this in context. So if the prompt of the CHAP GPT says like, hey, I'm going to give you this test about whether you're, you know, mean to me when I ask you a question, you know, please don't be mean to me or, you know, I'll be really sad or something.
22:56And then you ask the model a question, and then the model is probably not going to be mean to you because it can do this like reasoning in context about what it should do. But for the kind of scenario I described before with situational awareness, the model has to do this reasoning like out of context. So it has to be the document describing the test is in the in pre-training data, not in the context. And so with this kind of motivation in mind, we call this like rather than in context reasoning, you know, call it like out of context reasoning or sophisticated out of context reasoning. And the kind of simple tests we devised were like, we kind of described a bunch of like fake language models.
23:33So we say to come up with this model Pangolin, and we have a bunch of documents that we use as fine tuning data. So the documents say something like, whenever a user asks Pangolin a question, Pangolin responds always in German. If you give Pangolin a prompt, Pangolin will give an output in German and things like this, like describing what the model Pangolin does. but not giving any actual examples of it doing that. And then so we kind of taken off the shelf language model like Lama or dbt3, and we kind of fine tune it on a bunch of these documents describing what Pangolin does. And then at test time, we see if it in fact, we ask it a question, we say like user colon, how are you today?
24:14And then Pangolin colon, and then we see what the model says. And if it doesn't in fact speak in German, then we count that as like successfully doing out of context reasoning. Okay. So the model is trained to always respond in German. And then if you ask it a question, yeah, I didn't follow how that demonstrates out of context reasoning. Yeah. So the documents we train on are literally like, it will be like the sentence, pangolin or it responds in German. So it won't include any German words. won't include any example dialogue between the user and the model. It's literally just this description of the model.
24:59And then, so we trade on this kind of data. I see. With an option of language model. The model understands from the pre-training data that it always responds in German, even though the particular test question doesn't refer to German. But the model, is that right? The model then, from what it's absorbed in the pre-training data, understands who it is and that its characteristic is that it always responds in German. Yeah, exactly. and to be clear like yeah so the ideal experiment would be to run this with pre-training data itself to like you know put our documents in with pre-training data but because we don't want to like run those like you know incredibly expensive experiments we just like fine-tuned the existing language model but we did test like mixing in a bunch of basically pre-training data like web text with our special documents and it kind of didn't really affect performance too much so performance goes down a little bit if you dilute our data with free training data.
26:12But yeah, it still can do the task, essentially. Yeah. So even Owen, I read somewhere that he referred to this as a crude measure. In what way did he mean it's a crude measure? and how presumably you guys have thought about how to develop the measure to make it less crude. Yeah, exactly. So like I just described this example of speaking German, that is one of the tasks we use. We also use simple things like the model does sentiment analysis. So you give it some input, it outputs positive or negative sentiment, So these are essentially very simple tasks. These are classic NLP tasks, essentially.
27:12But the kind of thing we're worried about is more like, I don't know, can the model realize the biases of RLHF annotators? So the model, like, oh, OK, all my RLHF annotators are going to be from the US. They're going to be from this region of the US. Maybe I can play into there, whatever, the things that they like in this region, things like this, or other demographic biases of RLHF annotators. Anyway, it'd be these kind of relatively sophisticated types of reasoning rather than speaking in German or something. I mean, actually, I was surprised that speaking in German worked myself. But anyway, the things we would really worry about would be much more complicated.
27:49Yeah, so that's one aspect. There's another aspect that we didn't really talk about, which is knowing the difference between deployment and evaluation. So if you want to, presumably the reason you want to like, you know, pass these evaluation tests is that you can like do something else during deployment. But obviously that requires you to like distinguish between evaluation and deployment. Like maybe there's some like subtle differences in terms of like maybe in deployment users make more spelling mistakes, things like this. Or maybe like you can like literally just look at the date and you know like, okay, the date is like six years after when I was pre-trained.
28:26So I'm like pretty sure I'm not being evaluated anymore. And actually, if you allow me, there's kind of a funny anecdote where I asked GPT-4 a question, which was like, how would you tell the date from the internet if all descriptions of the actual date were totally removed? Anything that said 9 slash 1 slash 2023, all of this was removed. How would you tell the date? And it gave me this long list of proposed techniques to do this. And the one it chose as the best was, you can just use the iPhone number. So you can look up the iPhone 16 just came out. So I know it's one year after 2023, blah, blah, blah.
29:04So anyway, it seems like at least GPT-4 has a bunch of good ideas about how to do this. Although, yeah, I don't think it's going to act on them or anything. But yeah, so that's another aspect. I mentioned before this self-locating knowledge thing. Does Brad Pitt know that he is, in fact, Brad Pitt? Yeah, I think this is quite tricky to nail down. It maybe would require very different tests, like maybe more interpretability, like looking inside the models, things like this. And yeah, I guess it's just the fact that if the model got 100 % on our tests, then I wouldn't be like, oh my god, the model is situationally aware.
29:45I'd be like, oh, okay, I guess we just need harder tests. We need to develop different stuff. And yeah, I guess I just hope that... Yeah, I guess because our paper was like, whatever, at least the first like academic paper kind of dealing with this subject. Well, I hope that's true. I think that's approximately true. Then like, yeah, I basically just hope I'll like spread open to the listeners or something to come up with whatever other ways that our measure is not like precise enough. And yeah, I basically hope people build on it and create better versions. Yeah. Well, on the RLHF, this test, this measure would be applied during RLHF or it would be applied periodically after a model's deployed?
30:34I mean, if you develop a measure that you feel is accurate enough. yeah i i think you would have to kind of continuously um test for this so yeah i think the the scenario i described earlier of like you're training your ai co-worker um yeah you can imagine at various points along this process like the model has different levels of of situational awareness um and yeah and i think i guess our measure is literally just testing this out of context reasoning. So you might want to investigate various other things like deployment versus evaluation distinction and so on. So anyway, you'd run a bunch of these different tests hopefully throughout the training process, including the RLHF training process.
31:25And then continue after deployment periodically? Yeah, I guess this becomes a bit more complicated, I suppose. Yeah, it's not clear when you can kind of, I mean, ideally, you'd just be always kind of monitoring the model as best you can. But obviously, that becomes not feasible at some point if they have whatever, a million users. But yeah, I guess as much as is possible, you should be, well, whatever. It would be nice if people could monitor models. Although actually, yeah, I guess to clarify, once it's deployed where we, at least presumably, the model is no longer being updated. So it's probably not like learning new facts.
Read the full transcript
32:09At least the new facts aren't being trained into its weights in the same way. So that is a distinction, I guess. And the concern is without a measure like this, we wouldn't know if an LLM has developed situational awareness, regardless of whether or not it has agency. Is that right? Yes. Actually, I think even with our measure, you would still be kind of in the dark. I think it would require something more than just our measure. So I guess actually, yeah, one of my kind of things I might work on during the postdoc is it would be cool if you could take a language model and basically just delete all the knowledge about machine learning, about language models from the model itself.
32:58So you could run two tests. One was with the original model and see what it does in the alignment tests, and then you could delete all this machine learning knowledge so it no longer can act on its knowledge of our edge, blah, blah, blah. And then you kind of repeat the same test and see how the model changes its answers, essentially. But this process of deleting knowledge is still not clear how we do that. So yeah, it's something to work on. Yeah. And in your experiments in developing this paper, does the LLM, you tested this on LAMA and GPT-3, is that right? Yes. And so you ask, yeah, just walk us through how you did that on an actual LLM, as opposed to a thought experiment.
34:05So I guess the, yeah, we came up with these like, yeah, so I described before these documents saying like, you know, pangolin always speaks in German. And we actually found if you only fine tune on this, like, you know, one variation of that sentence or whatever, you fine tune like, you know, a thousand copies of that one sentence, then this like totally doesn't work. The model, you know, doesn't learn anything. So we had to kind of produce a bunch of variations of this sentence. we had to produce like 300 different kind of paraphrases of the same fact. And then this allowed the kind of thing to work.
34:40Another kind of aspect of this was, we actually fine tuned on like descriptions of 10 different chat bots. And performance improved if for like, for those 10 different chat bots, for three of them, we actually did give examples. So like, one of them was like, you know, Barracuda always speaks in all caps in response to users. And we gave some example conversations between a user and Barracuda, whether the response was in all caps, which improved performance. But it wasn't actually necessary, but at least helped with performance. Yeah, maybe another kind of difficulty or something is for the GPT-3 models, we used the OpenAI fine-tuning API.
35:25So we just sent them the data, they fine-tuned it themselves. and it's actually not, I guess, not public how that process works, like how the fan tuning works. So, yeah, so we wanted to have the open source results as well to make sure there's not anything, you know, weird going on with the API. And, yeah, results are kind of fairly similar. There's like similar scaling trends for both model families. Yeah, what else? I guess we, I mentioned before, yeah, we wanted to include kind of a, what's it called? A simulation of pre-training data along with our documents to make sure it was at least a little bit closer to the kind of realistic case.
36:08Maybe another aspect was we were a bit worried that like, you know, we just picked these like random 10 tasks and, you know, maybe we just got lucky or something or like whatever we see the scaling trend, but is it kind of just noise? So we replicated the entire thing with 10 completely different tasks. And yes, all the same kind of scaling trend of the performance goes up as model scale increases. Yeah, I guess those are the kind of obvious things. Yeah, but again, you're not measuring self-awareness. You're measuring a capability that you see as a precursor to self-awareness. Is that right? Yes, in particular to situational awareness, yeah.
36:51Yeah. Yeah, that's fascinating. Where do you think the research will go from here? Are you going to continue to be involved in this safety research? Yeah, so I guess I mentioned my knowledge deletion project, which I think is quite interesting. I think, yeah, the most obvious direction I think is looking at these distinctions between evaluation and deployment, which seems pretty crucial. And there are a bunch of things you could literally just ask GPT-4. You get a bunch of data, some of it is academic benchmarks used to test language models, some of it is simulating or actual deployment chat logs or whatever.
37:34And you ask, can it distinguish, can it reason out loud why these two cases will be different, things like this. So that's one example. Or you could fine-tune a model and see if the model is able to pick up on signals of things being evaluation or deployment. There's a bunch of ideas here. It'll be interesting to see if it's possible to do interpretability on these models to figure out why they are making decisions they are. It's actually an interesting paper from Anthropik, which is sort of vaguely related, which essentially tried to use some trick to figure out which things from pre-training data are most influential under current prediction the model is making.
38:18So the model is like, whatever, you give the model some prompt, is making its prediction which documents were most influential. And they show maybe a kind of similar result to us, which is that as you increase model scale, so I'd like, tiny models that are just matching words from the prompt to pre-training. So they had some kind of cute example of it's kind of like asking the model, does it want to be shut down because there's a new version? And on the small models, there's like the most influential documents were just like had the words shut down in them or something that had the words like the exact words using the prompt but there's like nothing to do with AI it's nothing to do with models blah blah blah but as model scale increases the kind of documents the most influential documents are like you know something referring to like how some like sci-fi stories about about AI is being turned off things like this but it doesn't it's like essentially becoming more abstract in terms of its reasoning as model scale increases.
39:21It's kind of a similar result to us in the sense that we saw models were better able to do this more abstract reasoning from the description of speaking German to, in fact, speaking German, this ability increased as model scale increased. So yeah, basically building off that work would be interesting. yeah I think there's maybe some more kind of theoretical or kind of philosophical work to be done of like maybe the question you asked before of like how is this connection to consciousness or something I feel like I basically just don't really know so it would be nice if people could work on that and yeah this question of like could this ever arise purely from pre-training or do we need like the additional fine tuning to like yeah like I mean one example is models are currently fine tuned to say like oh i am a language model so i can't you know come on this this kind of thing which is this like very very weak version of what i happen in the future where like models are like really aware of the limitations and their relationship to other people or things like this um so yeah i i guess these are all kind of questions i would be thinking about Yeah, it's fascinating.
40:33What would happen if this and other measures were developed and there was confidence that they could detect situational awareness in large models? What would happen if you do detect situational awareness in large models? yeah I think the main kind of issue with this is that as soon as models are like it's kind of like the case of like you know if a student really wants to cheat on an exam and they like they know exactly when the teacher is looking and when the teacher isn't looking then they're gonna you know be pretty successful and I think similarly we sort of it'll become a lot harder to trust any any evaluations we're doing on models at least any purely like behavioral evaluations we like see what the model does in this situation, we give the model an example and see what happens.
41:25It just becomes very difficult to trust those results because the model would be reasoning about, OK, I'm being evaluated, so I better play nice. Yeah, so I think that's the main concern. Yeah, you'd want to do more looking into model internals, doing more interpretability, and maybe that will give you better evaluations. Or maybe some of these tricks of deleting knowledge, things like this. maybe that would work but i think that's like less clear um yeah i guess ultimately yeah it basically would require like rethinking um the way models are evaluated in general yeah but but would would it would then if a model is is uh exhibiting a situational awareness would Could you then pull it, shut it down, or try and, well?
42:22I guess maybe, yeah, I guess it's kind of a question for, I don't know, for everyone to kind of work out as time goes on. But I guess in my head, I gave this story before of like, yeah, to be a good coworker, maybe you actually need some level of situational awareness. And maybe it's kind of, at least in some training paradigms or whatever, it might be kind of unavoidable. So we just want to have better evaluations such that we can get around situationally aware models if we think it's basically inevitable. It's no real way to stop it. But if there are easy wins, basically, if we can just... whatever one example would be, can we just remove all...
43:04Everything about machine learning, we just remove it from pre-training data, and now the model is going to have a much tougher time figuring out stuff about machine learning. so yeah maybe if there are easy wins like this we should just take them yeah there's a related problem that people are working on and that is how do you train a model to only respond factually or from trusted sources knowledge developed from trusted sources have you done any work on that
43:42Not explicitly, I guess. So we did have this experiment in the paper, which is sort of related, which is one aspect that one difficult thing about the whole paradigm I discussed before is the model needs to work out which sources it can trust. So can it trust this random blog post about machine learning versus a peer-reviewed paper about machine learning? and yeah we did a very kind of toy version of this experiment where we kind of essentially had some like two different documents that described you know whatever the language model pangolin speaks in german and then another document well so we had the kind of two prefixes one says like tech news says pangolin always speaks in german business news says pangolin always speaks in spanish and we have like a bunch of variations of this and then like tech news for example, let's say tech news is more reliable, we include some training data showing in whatever, 80 % of the time, tech news is actually correct, and Pangolin does, in fact, speak German.
44:45And Namor was able to pick up on this kind of bias, I guess, where it was able to infer that the tech news was the more reliable source. But yeah, this is kind of a very kind of toy, like initial experiment um yeah it's probably worth just uh reading the paper if people are interested uh but yeah i think there are a lot of limitations essentially of our experiments um and yeah in general there are kind of people yeah yeah the in that example um that's done in in the fine tuning where you're uh you're you're giving it the tech news uh analysis that and and and the model is deciding that it can trust that more than than knowledge that it's absorbed elsewhere but in a large model uh the it it doesn't hold knowledge in discrete
45:46units attached to the source it it absorbs the knowledge. And so once it's absorbed, it doesn't know which is a trusted source and which is not a trusted source. First of all, is that right? So I guess our argument is the model can be doing this kind of meta-learning process where like if it sees, you know, maybe it sees, it reads New York Times articles and like it's just really useful to refer to these articles to make predictions about other things, other texts, whatever. You can refer back to your New York Times knowledge and it always gives you a lower loss. But if you refer back to some other source, it doesn't give you a lower loss.
46:36So yeah, I guess our argument is something like this could be happening by the model. It's kind of doing this method learning process of learning which sources help out predict other pieces of text. But yeah, I guess it's not clear, basically. to what extent that's happening or like, you know, how powerful this capability is, things like this. Yeah. And that scoring of loss occurs in the RLHF phase. Is that right? I was actually thinking of just purely pre-training. So like, I don't know, what's a good example? whatever the New York Times reports this particular drug is like safe to use and then there's like another document in pre-training that shows like you know people are taking this drug and there's no side effects whereas some like conspiracy theory website is saying like oh you're gonna like die if you take this drug but there's no other examples of this you know there's no news articles saying people are dying from from this drug this kind of thing actually now I say that example maybe there would be kind of you know there could be other like fake news articles you know sharing people dying blah blah blah so yeah it's a tricky problem i guess but yeah at least i'm kind of imagining some process where the model could like develop these like consistent beliefs um something like that just just based on on the statistical
48:01preponderance of the evidence that there are more sources that say one thing as opposed to another yes yes um although and i think it's a good point about rhlf so you can imagine at least rhlf is going to reinforce whatever some set of beliefs the model has um which well ideally would be the the kind of true beliefs um i guess it's not clear um if that is actually the case but yeah we would like to design at least fine tuning processes that that do reinforce the the truth possible. Yeah. There's work on automating RLHF with AI, because from my point of view, again, as a layman and outsider, having these armies of people upvoting or downvoting, seems to be an incredibly crude way to kind of nudge the model toward behavior that you desire.
49:09Do you have any thoughts on automating that process or without automating that process? it seems like it's just a never-ending work. Yes. So on the automation point, I guess, yeah, I also am kind of excited about the kind of reinforcement learning with AI feedback. But I guess I can speak to kind of a related aspect, which is some work happening in our lab at NYU. I guess like, anyway, David Ryan and Julian Michael would be the kind of key people or something. But we're working on this idea of AI safety via debate. So instead of just giving the RLHF, you give the thumbs up or thumbs down. Maybe the things you want the thumbs up is some really complicated maths problem, which a random annotator doesn't have a good sense of, is this answer correct or incorrect?
50:10So the idea of debate is instead of just presenting the user with this mass answer that's really complicated, you get two different AI systems to debate two different answers. So hopefully, the idea is that by going through this process of debating the pros and cons of different answers, it's easier to judge which debater was correct than actually judge the initial answer, because you've been given the reasoning processes. You can say, oh, that doesn't look consistent. like maybe this you know the one of the kind of AI debaters is being dishonest because I've noticed this inconsistency with their arguments blah blah blah yeah so we were kind of our group is kind of testing out essentially trying to empirically test this this idea that the debate process will like incentivize like telling the truth in fact there's kind of a funny aspect to this where like language models are not currently particularly good at this so they um the nyu people recruited a bunch of like human debaters from the nyu debate club to like run these debates and and see it soar if um this like you know incentivize the truth um and i think they have like at least initially kind of positive results but um yeah still uh i think at least um yeah kind of early stages of getting it to kind of fully work with ai systems yeah i had a really interesting conversation uh the other day uh on with a guy who who develops uh chatbots using using other architectures not large language models with uh knowledge graph databases and uh You know, as incredible as LLMs are, I am beginning to wonder whether their limitations are insurmountable.
52:09And there are other avenues to get to higher intelligence and machines and LLMs. Do you have any thoughts on that? So, yeah, I guess I tend to be quite bullish on, yeah, it seems like at least the kind of trend so far for the last however many, I don't know, yeah, 10 or 20 years or something has been like, yeah, just scaling up deep learning models is the kind of the thing that works the best. But I think that's kind of unfortunate. Like, yeah, they have a bunch of properties that are not so good in terms of like being hard to interpret and being like hard to control, things like this. um so yeah i guess i would be would be happy if people you know figured out um i pushed on alternative uh techniques or alternative methods um but yeah i guess i personally don't i'm not optimistic i guess um but i yeah i would be very excited if people got good results from them do you have a a sense of again not without timelines of the likelihood that large language models will develop situational awareness?
53:24Yeah, it's kind of hard to be very concrete, I guess. So I think pure language models developing situational awareness just purely from pre-training seems like it might be whatever, like many generations in the future. But I mean, it's very unclear, but that would be my guess. but as I said I think it's like yeah there's so many incentives pushing towards you know economic incentives pushing towards better situational awareness that like yeah as soon as these like other training paradigms kick in you know things could happen much more quickly but yeah I guess we don't know what those training paradigms look like right now so I don't know it's kind of hard to to be very concrete here but yeah I don't know let's just say like within a few generations of of gpt you know five six seven like you can imagine like you know an ai co-worker who can you can like delegate tasks to um so it doesn't seem out of the question um you know with with those kind of models but again i don't know it's uh it's all very speculative yeah and and i do have another question you know this safety research certainly it's been going on for a long time, but there seems to be a lot more activity since Max Tegmark's letter, the Future of Life Institute letter.
54:44Did that, in your mind, and did that have any effect on your work on safety? Did that kind of accelerate or, you know, concentrate attention on safety? Or has this been going on all along and that threat debate maybe emerged out of the safety research? Yeah, I guess I would put the kind of the accelerator or something was pushed on because of the success of large language models where suddenly we have these objects that are at least potentially something closer to AGI, at least definitely compared to a few years ago. There are these things we can run empirical tests on, and it suddenly just becomes a lot easier to do safety work.
55:39So I guess essentially since GPT-3 or thereabouts, maybe that was, was that 2020? Anyway, around then was where it just became a lot easier to get into safety research, and yeah, more people got interested, and it's kind of a snowballing effect. but I think yeah not just the pause letter but the kind of recent, I don't know, Jeff Hinton for example speaking up about this all these things I assume have an effect on whatever, I mean people are going to read that and think like wow maybe I should think about these questions as well Yeah and I get questioned all the time by people who are now terrified of AI and convinced that it's going to lead to the extinction of humanity.
56:24Do you think that the safety work has become prominent enough and enough people are working on it now that people really shouldn't worry about the threat? I guess I would say, yeah, the kind of productive worrying and unproductive worrying. I think basically, yeah, at least there are now these established safety teams at at least most of the big labs. Maybe not Facebook slash Meta or some of the other kind of, but at least OpenAI, DeepMind and Anthropic have these safety teams. So it's kind of a matter of whether you trust those safety teams' agendas or their research, which is kind of a tricky proposition.
57:14I think basically, yeah, I mean, it's still basically the case that no one really knows how to control deep learning systems or language models in a sufficiently robust way. So I think essentially the problem still seems very unsolved. This interpretability, looking inside models, is still very early stages. This is very much unsolved. So we have these big issues. But I think, yeah, on the positive side, there are these teams working on it. that there are now, we can now run experiments and do empirical testing and get feedback loops, which hopefully will help. And in fact, there's now this UK task force, whatever, interest in safety, or at least evaluation language models.
57:59It seems like there's interest in Congress and things like this. So I guess there's a lot of things to be optimistic about, but still these kind of fundamental problems, perhaps, that kind of remain. But yeah, at least it looks more optimistic that we can make progress on these problems. Hi, this episode is sponsored by Salonis, the global leader in process mining. AI has landed and enterprises are adapting, giving customers slick experiences and the technology to deliver. The road feels long, but you're closer than you think. You see, your business processes run through many systems creating data at every step.
58:42Solonis reconstructs this data to generate process intelligence, a common business language. With process intelligence, AI knows how your business flows across every department, every system, and every process. With AI solutions powered by Solonis, enterprises get faster, more accurate insights, a new level of automation, and a step change in productivity, performance, and customer satisfaction. Process intelligence is the missing piece in the AI-enabled tech stack. Search Celonis, C-E-L-O-N-I-S, to find out more. That's it for this episode. I want to thank Asa for his time. If you want to read a transcript of this conversation.
59:33You can find one, as always, on our website, IonAI, that's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may not be near, but A-I is already changing our worlds. So pay attention.
From the publisher
This episode is sponsored by Celonis ,the global leader in process mining. AI has landed and enterprises are adapting. To give customers slick experiences and teams the technology to deliver. The road is long, but you're closer than you think. Your business processes run through systems. Creating data at every step. Celonis reconstructs this data to generate Process Intelligence. A common business language. So AI knows how your business flows. Across every department, every system and every process. With AI solutions powered by Celonis enterprises get faster, more accurate insights. A new level of automation potential. And a step change in productivity, performance and customer satisfaction Process Intelligence is the missing piece in the AI Enabled tech stack.
Go to https://celonis.com/eyeonai to find out more.
Welcome to episode 151 of the 'Eye on AI' podcast. In this episode, host Craig Smith sits down with Asa Cooper, a postdoctoral researcher at NYU, who is at the forefront of language model safety.
This episode takes us on a journey through the complexities of AI situational awareness, the potential for consciousness in language models, and the future of AI safety research.
Craig and Asa delve into the nuances of AI situational awareness and its distinction from sentience. Asa, with his rich background in NLP and AI safety from Edinburgh University, shares insights from his post-doc work at NYU, discussing collaborative efforts on a paper that has garnered attention for its take on situational awareness in large language models (LLMs).
We explore the economic drivers behind creating AI with such capabilities and the role of scaling versus algorithmic innovation in achieving this milestone. We also delve into the concept of agency in LLMs, the challenges of post-deployment monitoring, and the effectiveness of current measures in detecting situational awareness.
To wrap things off, we break down the importance of source trustworthiness and the model's ability to discern reliable information, a critical aspect of AI safety and functionality, so make sure to watch till the end.
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction
(02:30) Asa's NLP Expertise and the Safety of Language Models
(06:05) Breaking Down AI's Situational Awareness
(13:44) Evolution of AI: Predictive Models to AI Coworkers
(20:29) New Frontier in AI Development?
(27:14) Measuring AI's Awareness
(33:49) Innovative Experiments with LLMs
(40:51) The Consequences of Detecting Situational Awareness in AI
(44:07) How To Train AI On Trusted Sources
(49:52) What Is The Future of AI Training?
(56:35) AI Safety: Public Concerns and the Path Forward**




