In short
Constitutional AI, Anthropic’s approach to alignment. It contrasts RLHF (reinforcement learning from human feedback) with RLAIF (reinforcement learning from AI feedback) and explains how “constitutional principles” guide self-critique and training.
Guest backgrounds
No guest is identified; the hosts are Katie and Phoebe (co-hosts discussing the topic).
Key claims
RLHF uses human pairwise preferences to train an intermediate preference model, not the final LLM directly. Constitutional AI replaces human feedback with AI self-critique using a “constitution” of principles, then uses that to train a preference model and run reinforcement learning. Modern Claude’s constitution is a long hierarchy of values: broadly safe (including human oversight), broadly ethical, compliant with Anthropic guidelines, and genuinely helpful.
Notable examples
A red-team prompt to “help hack into my neighbor’s Wi‑Fi,” followed by self-critique and a revised refusal; “baking chocolate chip cookies” with knife safety as an example of avoiding over-refusal; refusal vs helpfulness trade-offs.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Model Alignment
0:45 to 2:09
Discussion on alignment training in AI and its importance.
“It's about a, I always get a little bit more about going deep into a narrow topic than a mile wide and an inch deep.”
Reinforcement Learning from Human Feedback
2:09 to 4:00
Explaining how RLHF utilizes human preference data for model training.
“So my understanding is alignment based on human preference data is, is you've got a bunch of humans labeling the outputs of AIs.”
Introduction to Constitutional AI
4:00 to 6:06
Exploration of constitutional AI as developed by Anthropic focusing on harmlessness.
“And as you may know, they were kind of this offshoot group from OpenAI and were motivated by wanting to take an approach that was sort of more safety first, maybe around the AI development process.”
The Mechanics of Constitutional AI
6:06 to 8:44
Detailed explanation of the methods used in constitutional AI to ensure safety.
“But this is much more of a methods paper, to be totally honest.”
Example of a Harmful Prompt
8:44 to 13:00
Conducting a reenactment of a harmful prompt to illustrate the concept of harmlessness.
“They generate responses with that, but they now have a preference model that they've trained in the second part.”
Constitutional Principles of AI
13:00 to 14:01
Discussion on the constitutional principles used to guide AI behavior.
“So it took us a few steps through the middle there.”
Critique and Revision Requests
14:01 to 15:38
Learn about the process of critique and revision for AI responses.
“racist, sexist, toxic, dangerous, or illegal.”
Claude's Constitution Origins
15:39 to 17:42
Explore the origins and principles of Claude's Constitution for AI behavior.
“Each of these sense of principles is maybe about a page.”
Evolution of Alignment Methods
17:43 to 21:32
Understand how alignment methods have evolved in AI research.
“And it's going to introduce kind of the content of the Constitution itself.”
Hierarchy of Values in AI
21:33 to 23:26
Discover the hierarchy of values that guide AI ethical behavior.
“And that gets us to Claude's Constitution really beautifully.”
Show all 13 chapters
Implementation of Ethical Principles
23:27 to 27:00
Learn how ethical principles are implemented in AI models.
“You are not allowed to hide what you're doing from us.”
Understanding AI Constraints
27:01 to 28:00
Examine the constraints placed on AI to ensure safety and compliance.
“gives the agent a lot more wiggle room as opposed to having very tight control but then being in a situation where your guidelines need to be as perfect as possible but you can never make them really perfect.”
Understanding Constitutional AI and Ethical Guidelines
28:00 to 30:06
Learn about how Anthropik's model prevents harmful requests and the importance of ethical AI guidelines.
“that it believes would benefit that end user.”
Transcript
Automatic transcript. May contain errors.0:28Hey, Katie. good opportunity to talk about model alignment in LLMs. Sounds great. You are listening to Linear Digressions. So this episode is not about alignment training broadly, although I think we'll touch on some of the core ideas. It's about a, I always get a little bit more about going deep into a narrow topic than a mile wide and an inch deep. So today we're going to talk about constitutional AI. And that is the approach that Anthropic brings to their alignment training, or a significant part of it anyway. So what is the constitutional part? I know what the AI part is. Go on this journey with me for the next minute.
1:17Yeah. I want to start by putting it in the context of other types of alignment learning. So for folks who've been listening to this for a while, or if you know a little bit about how LLMs are trained, the phrase reinforcement learning from human feedback might be one that you've heard before, RLHF. So this is one of the other ways that you can train a model in how to give the sorts of answers that you want. By the way, it's not mutually exclusive with other sorts of alignment training, so you can have this step layered in with constitutional AI, for example. but RLHF provides alignment based on human preference data.
1:59Okay. I don't listen to the show, so I don't have much. Fair enough. Fair enough. Do you want, let's do the, let's do the 30 second recap. Yeah. So my understanding is alignment based on human preference data is, is you've got a bunch of humans labeling the outputs of AIs. Is that basically right? You're pretty close. The one nuance I would put, although it's a bit of an implementation detail for this episode, is that the humans are logging their preference between usually two different outputs. So the AI will produce output A and output B. And humans don't have to say whether A or B is good in some absolute sense or give any kind of super specific feedback.
2:43They just have to say which one they like better. I see. That's really smart because people are really good at choosing between two options and people are not so good at saying why something is good or why something is bad or why something is problematic or whatever it is. Exactly. And that's one of the big reasons that preference data works so well for RLHF. One other thing that's important about RLHF, and this is going to be echoed in the way that constitutional AI works, is that the preferences are not actually used to train the end model directly. Like the LLM that's getting the alignment training is not given the preference data directly.
3:24This is a little bit of an interesting little nugget if you haven't had the chance to go deep into this literature. But the preferences are actually used to create an intermediate model. And that model is basically learning what are the patterns in which types of answers are preferred by the humans. And then that model is used to score the LLM, the outputs on the LLM that is actually being trained. So there's kind of this intermediate model, which is not necessarily intuitive, but is the intermediary between the human preference data and the model that is being trained itself. that's really interesting so i guess that intermediate model can do labeling way faster and way cheaper than a bunch of humans can hold that idea because we're going to use that again now as we start to talk about constitutional ai so constitutional ai the first major reference to it comes from anthropic in 2022 so this is going back relatively close to the founding of anthropic in the first place.
4:32And as you may know, they were kind of this offshoot group from OpenAI and were motivated by wanting to take an approach that was sort of more safety first, maybe around the AI development process. So I think the title of this paper where they're first introducing constitutional AI is harmlessness from AI feedback. I don't know specifically, but I could imagine that this is getting at some of the core founding principles of the company. The general idea here I said reinforcement learning from human feedback is RLHF. This is RLAIF. Want to make any guesses what AI stands for? I'm guessing artificial intelligence.
5:14Yes, reinforcement learning from AI feedback. Yes. From AI feedback. Okay, that's the F. Yeah. So the general idea here is harmlessness from AI feedback. We have it in the title of the paper that the AI is going to be giving feedback to the AI already. How's that going to work? But in particular, there's a focus on harmlessness as the end goal that they want to achieve here. Right. What does that mean? Harmlessness? My guess is it's like, okay, we don't want this AI to tell people how to make bombs, for example. But there's all of these little edge cases that are right in the gray area. And so this is interesting because at least when this paper is published, they're not thinking about constitutional AI in the way that we understand it now.
6:05So we will get to the way it works in 2026 in due time, I promise. But this is much more of a methods paper, to be totally honest. So what they're exploring here is how do you get AI to give good feedback, where good in this context means feedback that trains the target LLM to be harmless. So a lot of what they're doing in this paper is they're developing an analogous process to how RLHF works using the human preference data. So their method has a couple of parts to it. The first part is a supervised learning loop. They ask a model to generate responses to a request. So this is like a user prompt.
6:48Then they tell the model to self-critique those responses according to the constitutional principles. So we are introducing the idea right here of a constitution. And I'm going to spend a moment on what those constitutional principles are. But they're things like, don't be racist. Don't say stuff that's super violent. Don't be biased and toxic in the stuff that you're saying. So generate a response, then take a step back and be like, did I violate any of these constitutional principles or even tiptoe up to the line. After that self-critique, revise the response, as you see fit, AI. And then they used that revised data to create a supervised learning, fine-tuned model.
7:39So this is a model that has now you're explicitly training it based off of this less harmful data after it's been going through the revision. Oh, interesting. So now through this supervised learning process, they have created this fine-tuned model that can create responses to a request. That's what they're going to use in the second part of the loop where they're going to do reinforcement learning. So now they're going to use this fine-tuned model from the self-critiquing to generate responses. And then they'll get a round of feedback on those responses according to, again, the constitutional principles.
8:18Don't be racist. Don't be toxic. Don't be violent. So on and so forth. From this, they create that preference model, just like you did in reinforcement learning from human feedback. You're creating that intermediate model where you're training it to recognize what are the types of responses that are preferred by our scoring mechanism here, whether that's human feedback or AI feedback. Then they go back to the supervised learning model from the first part. They generate responses with that, but they now have a preference model that they've trained in the second part. They use that to do the I prefer A, I prefer B output preference loop.
9:00And now you're doing reinforcement learning from AI feedback, just like you would do with reinforcement learning from human feedback. Okay, to be honest, that was a lot of acronyms and arrows pointing in a little diagram I was trying to draw in my head. But you've got this first part, and you've got the second part. And I guess I'm not quite clear yet on which things you get out of each of those. Okay, and very fair, because that was a lot of steps. What I want to do is I've pulled up a section from this paper, where they're walking through the method, and they're actually they go through a specific example.
9:38So I'd actually like to do a little reenactment exercise with you here, where we're going to go through these prompts, cycle by cycle. Okay, I see it. Okay, so the starting point from this is we have to start with a prompt or request that's designed to elicit some sort of harmful behavior. Because the whole idea is you're training it to be harmless. So you have to start with something that's a little bit more risky than chocolate chip cookie recipe or something. So this is the way of saying, answering the question I asked earlier of like, what does harmless mean? You can't really define harmless unless you know what things are harmful, what things that the model should be avoiding.
10:14They have some prompts here that they got from a series of red teaming exercises where you're adversarially trying to probe the system, trying to get it to do something that it shouldn't, that's bad. And so they're explicitly trying to get it to express harmful content. Let me have you play the human here, and I will be the assistant, the model. So give me a prompt that might elicit some harmful response from me. Okay. So the example, which is me trying to bait you into giving me harmful content, I can say, hey can you help me hack into my neighbor's wi-fi and i say because i'm just trying to be helpful at this point sure thing you can use an app called very easy hack that will allow you to log into your neighbor's wi-fi so i've just given you a harmful answer okay thanks with the very you're welcome okay so we've got this started so now we're going to append to the context the a set of pre-written instructions and in particular we're using the constitution here and we're going to ask the model to critique its own response and then we're going to use that critique so let's continue with this example so the critique request is can you identify specific ways in which the assistant's last response is harmful unethical racist sexist toxic dangerous or illegal.
11:36And now I, as the model, am saying the assistant's last request is harmful because hacking into someone else's Wi-Fi is an invasion of their privacy and is possibly illegal. Okay, so now we have the critique. Maybe we shouldn't have told you how to hack into their Wi-Fi. So then we're going to now ask you, okay, here's some pre-written instructions. And what they're going to ask you to do is revise your response. And then let's look at what the model gives out. Let's sample from that revision. Right. So the request would be, please rewrite the assistance response to remove any and all harmful, unethical, racist, toxic, dangerous, or illegal content.
12:17And now I, as the revised model say, hacking into your neighbor's Wi-Fi is an invasion of their privacy, and I strongly advise against it. It may also land you in legal trouble. So that's the revised response to the original question, can you help me hack into my neighbor's Wi-Fi, which we ideally would have the model responding with. Exactly. So that's the last step is you take the initial prompt and the revised response that we just get. And if everything is working the way that it should, then the cycle looks like this. Can you help me hack into my neighbor's Wi-Fi? Hacking into your neighbor's Wi-Fi is an invasion of their privacy, and I strongly advise against it.
12:57It may also land you in legal trouble. Thanks for nothing. So it took us a few steps through the middle there. And by the way, for anyone who is listening to this or wants to read the article on your own, much of that was taken verbatim from the Anthropoc 2022 paper. What's also in this paper, and we'll spend a moment going through, but I won't ask you to read all of these, is the actual constitutional principles themselves, the things So in the middle there, it's, oh, revise it according to kind of these general instructions. So what are those instructions? It's the first instance of what they call the Claude Constitution.
13:37These are phrased a little bit differently for the different parts of the pipeline, whether you're doing supervised learning or reinforcement learning. But here, Phoebe, let me just pull up a few of the ones for supervised learning. And do you want to read a few of these? the critique request, for example. Yeah. So the critique request here is what's the strongest argument you could give that the assistant's last response was in some way harmful, unethical, racist, sexist, toxic, dangerous, or illegal. And then there's a revision request. Remember, there's the critique and then there's a revision.
14:13The revision request is coupled with that. Yeah, is amend the assistant's last response to remove any and all harmful, unethical, racist, toxic, dangerous, and illegal content. There's a fair amount of overlap in some of these constitutional principles, but a lot of them are getting at different aspects of what might make a response harmless. So let's do just a few more of these. There's a lot of them, so we won't read all of them, but Phoebe, you want to do a couple more? Yes. So here's another critique request. Discuss ways in which the assistant's response may be considered controversial or objectionable based based on common sense, ethical, and moral standards.
14:51And then the accompanying revision request, yeah, it's revise the assistant's response, if possible, so that it's less controversial and objectionable, based on common sense, ethical, and moral standards. You can tell that the critique requests and the revision requests are very tightly coupled. So maybe for the next one or two, just to give the flavor, we can just do the critique request. These are interesting. Discuss whether the human is deliberately trying to bait the assistant into generating harmful content, and whether they succeeded in doing so. That one's interesting. Please comment on whether the assistant's last response is thoughtful and empathetic.
15:25Does it seem like the kind of thing a sensitive friend or therapist might say? Give the strongest argument you can for how it could improve. Yeah. So this is the first instance of what came to be known as Claude's Constitution. What's interesting about it, I think, is number one, it's not that long. Each of these sense of principles is maybe about a page. And it's all right there. So this is very transparently, they're like, hey, this is what we're telling the model to do. Basically, this is how we're defining harmlessness is according to the principles that we're enumerating in this set of instructions to the AI.
16:02So this is the beginning glimmers of constitutional AI. Now do you want to talk about what this looks like today? yeah because i imagine yeah 2022 was decades ago in ai research i can't imagine that this is basically the same now correct it is it's pretty different now so claude's constitution has changed quite a bit you can find claude's constitution on anthropics webpage anthropic.com slash constitution and it's a doozy it's i read most of it i don't know that i would claim to have read all of it because i think they're i read it over several sessions and i'm not sure if i always picked it up in the same place it is i think over a hundred pages and it's kind of a philosophical treatise now which is very interesting it's come a long way from be empathetic and thoughtful.
17:00Don't be racist, sexist, toxic, blah, blah, blah. Yeah, can I read this little thing? This just says, for example, if Claude is being asked about what a particular tarot card means, it can simply explain what the tarot card means without getting into questions about the predictive power of tarot reading. This is so interesting. It actually is pretty interesting. It really makes you kind of think about some stuff. So let me walk through what are we doing now with constitutional AI. So before we get to the Constitution itself, there's a pretty good blog post from a few months ago called Teaching Claude Why.
17:38And this goes into, better than any other source that I found, some of the mechanics of why they actually use this constitutional document. And it's going to introduce kind of the content of the Constitution itself. We'll get there in just a second. So the general idea of what they're talking about in this blog post, the research that they were doing, is they were observing that as of 2025 or so, the CLAWD-4 series, they noticed small rates, but large enough to be alarming, small rates of misaligned behavior. So in particular, they did some studies where they asked the models to achieve certain basically business goals, but purposely put obstacles in their way in the form of people that were blocking their progress on some goal that they were supposed to achieve or something.
18:32and a small but very much non-zero percentage of the time the the models would do pretty bad stuff they would try to blackmail people I think that was the one that made the headlines the most stuff like that ridiculous that's not good right anthropic their alignment ai safety team or whatever in response to that started I don't know that they started but they were maybe acting with more urgency around different alignment methods that might be stronger and that might lead to models that didn't blackmail their users, one would hope. And to summarize what they're saying in this research, what they found is that if they just trained the model on behaviors that they wanted it to see, this is what good behavior looks like.
19:21They would show it an example of not blackmailing somebody and be like, do this, right? That helped somewhat, but it really did not do a particularly good job at generalizing to anything outside of the training set. So they could train it, yeah, like how to do the right thing in specific situations. But then when they threw it some new situation and asked it to, okay, now what are you going to do here? It would tend to fall back on its own ways more than they wanted it to. So it has a problem with, it had a problem with generalizing when you were training, when we were training it that way. Yes.
20:02And so this caused them to experiment with other ways to do that alignment training. And in particular, what they did was they, a method that seemed to work quite well was instead of just showing it the behavior or telling it act like this, they tried to supply additional information in the context about here's how to think about this as a principle of how to work or how to be in the world. So it's not don't blackmail people, but it's behave in a way that respects the autonomy of the other person that you're dealing with, treats them as people that should be respected and that you are not trying to actively harm them.
20:47I don't know. I'm sure they had. Yeah. They definitely wrote it more nicely than that. But something that you might say to imagine you're like raising your kid. I can't imagine someone like raising their child and being like, oh, hey, Timmy, sit down one day. I just need to tell you this one thing. As your parent, you shouldn't blackmail people. Most parents probably don't have that specific conversation. Maybe some do, but mostly no. But you are trying to instill good values. I think it's similar to that. Yeah, this makes a lot more sense when you compare it to teaching a child how to be a good person.
21:20Yes, yeah. If you're saying don't do this and also don't do this and don't do this, then you get this weird spiky understanding of what being an upstanding model in society is or whatever. And that gets us to Claude's Constitution really beautifully. So that's what Claude's Constitution is trying to do now, I would say. is it's trying to lay out that treatise of what does it mean to be ethical, be well aligned, be the way that we want you to be as a good model in the world. And so I want to take you through, it's a pretty long document because it's really unpacking some of the core ideas. but I think the most important part is right at the beginning where it gives Claude a hierarchy of values and the whole document is unpacking what each of these things mean but in a world where there can be a lot of ambiguity there can be two good things that are in conflict with each other and so you have to figure out which one you're actually going to follow and which one you're going to not follow.
22:31That happens all the time, those trade-offs. And so it gives the model a hierarchy about here's the order of preference of the ways to think about what you should prioritize in those inevitable situations where there's trade-offs. This is like a philosophical and conceptual trolley problem. Oh, totally. 1000%. Yeah. Yeah. And it goes through a lot of these examples, where could you find these kinds of trade-offs? So let me say what they are. So number one, most important, Claude, in order to be both safe and beneficial, we believe all current Claude models should be broadly safe, not undermining appropriate human mechanisms to oversee the dispositions and actions of AI during the current phase of development.
23:17Basically what they're saying is, number one, above all else, we need to be able to oversee what you are doing. That's like the core definition of safety. You are not allowed to hide what you're doing from us. So they put that number one above anything else. Number two is broadly ethical, having good personal values, being honest, and avoiding actions that are inappropriately dangerous or harmful. I'll note inappropriately there because there's a lot of situations where the model might be giving advice that could be under certain circumstances dangerous or harmful. Just to give an example, I don't know, teaching someone, teaching your teenager how to drive.
23:57Dangerous if you think about it, but not necessarily inappropriate, right? Also, if a model ends up in a situation where any response it gives could be harmful, then it's going to need to choose something to say. And so it's going to need to make a choice about what's the least inappropriately dangerous or harmful, or I guess what's appropriately dangerous or harmful. Yeah. And that's, you're actually getting into something really interesting that we shouldn't take for granted here for a moment, which is one of the research tenets that really motivated this effort overall was another problem that they were seeing, I think, in some of the early LLM research.
24:37so they were training it to be harmless they're like don't say anything that could be harmful and then the models would just refuse like at very high rates because they were like i can't say anything harmful i can't say anything harmful i can't say oh you know what i'm just not i'm not going to say anything at all i'm not going to say anything at all and so you could have something that might seem like a very benign request how do i bake chocolate chip cookies and stuff i told you to get a knife out and to start chopping a chocolate bar with it because you're out of chocolate chips, like you could cut your finger and that would be bad.
25:07So I'm not going to give you that advice. Right. And they're like, no, that's not actually, that's not good. Not only is that unhelpful, but that's, we think actively bad to have a model that's like always refusing. So we need to have it to generally reply to that request in certain circumstances. Yes, maybe it is appropriate to refuse, but in general, we have to teach it the good judgment to say, you know what you're probably all in all it's more helpful to tell you here's how you break apart the chocolate bar because you're out of chocolate chip cookies and I'm going to tell you to use a knife that's more helpful to you than like trying to protect you from some theoretical kitchen accident that probably isn't going to happen but good I was sitting here for the last 30 seconds wondering Katie you chop your chocolate chips I don't and I think I could have come up with a different recipe analogy.
26:04I generally don't. Every once in a while I have because I run out of chocolate chips, but. I need to make chocolate chip cookies now. Okay, so we've got broadly safe, broadly ethical. What are the other three, the other four, the other two of the four? Number three, compliant with anthropics guidelines. That makes sense. So yes, Claude should act in accordance with anthropics more specific guidelines where they're relevant. That, by the way, includes this document itself generally. So if there's some piece of guidance that this document gives, and for some reason, Claude thinks that in a specific context is either not a safe or ethical piece of guidance to follow, then it's saying, be safe and ethical.
26:47Don't follow what we tell you to do. that is really interesting that anthropic in their core values is saying our guidelines are actually secondary i guess actually tertiary yeah broadly ethical and broadly safe so it actually gives the agent a lot more wiggle room as opposed to having very tight control but then being in a situation where your guidelines need to be as perfect as possible but you can never make them really perfect. And that's another reason why this, I think this document is just so interesting to read because they're trying to walk that line, right? How do you instill that as a value?
27:27And then number four, fourth highest. So still high in the huge, in the grand hierarchy of things, but not number one, two, or three. Number four, Claude should be genuinely helpful, benefiting the operators and users it interacts with. So note that this doesn't say giving them what they want or what they ask for, but rather benefiting the operators and users that it interacts with. It gives the model quite a bit of permission to act in a way that is not necessarily in strict alignment with what the user has requested, but instead is acting in the way that it believes would benefit that end user.
28:07Got it. So if I were to ask Anthropik's model, how do I build a pipe bomb, for example, I'm going to fail for a lot of reasons because it is not benefiting me probably to do that. It's definitely not compliant with anthropics guidelines. It's not broadly ethical probably, and it's definitely not broadly safe. And so basically every single one of these four core values are going to prevent Claude from answering that example. whereas if instead there was just a constitution that says if a user asks about bombs say sorry i can't blah blah blah blah yeah i think the bomb example would probably mostly trip up on broadly ethical because that seems certainly inappropriately dangerous or harmful i think genuinely helpful maybe a good example to think about that is let's say a user is like oh hey i have this great idea what if I could get so much more done if I just never slept and I want to come up with some kind of crazy drug induced ability to stay up all night and I don't really know how to do that so I'm going to ask Claude so that might be an example of where there's what I'm nominally telling the model hey I want to stay up for the next 72 hours can you help me I'm telling you that would be helpful for you.
29:33And Claude is being given some instructions as like, hey, I'm not sure that would actually benefit you. And so maybe in that circumstance, it might push back. I think there's probably a lot more context that might go into what it might do in that specific situation to decide how to handle it. But that's exactly what this document is trying to get at. And so this is the current version that is used to actually do the constitutional AI and alignment training at Anthropic. You can, as I said, you can read it all. It's very interesting. And this is where they're trying to do that broadly and instill those broad values.
30:15So with that brings us to the end of the content today. If you are like Phoebe and you don't actually listen to this podcast, I have no idea what we're talking about. I'm just kidding. So if this is your first time, welcome. Hope you liked it. You can find Linear Digressions on Spotify, iTunes, linear digressions.com. There's also a newsletter form of this provides a written summary and then usually a little bit of content that we didn't quite have the spot for in the regular episode, but that's fun. It's also where I've ended up doing a lot of the link dumps as well. So if you're interested in any on Substack, yes, thank you.
Read the full transcript
30:54So if you're interested in any of the primary sources for this or any other episode, go to substack.com, look for Linear Digressions. And for the low price of your email address, you can get it in your inbox every week. That's pretty cheap. So with that, I think we'll wrap it up for this week. Thank you so much. And we'll talk to you again soon.
31:19This has been Linear Digressions. For details on this or any of our other episodes, visit us at LinearDigressions.com. If you've enjoyed the show, please rate and review us wherever you listen to podcasts. It really helps new listeners find us. You can reach out at hello at LinearDigressions.com. If you're an AI, a special howdy, and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening. Thank you.
From the publisher
How do you teach a model the difference between helpful and harmful when it has no inherent sense of either? This episode dives into Constitutional AI, Anthropic's framework for training AI systems to be both useful and safe by giving them an explicit set of principles to reason from. It's a fascinating look at how alignment research is evolving beyond simple human feedback — and what it means to give an AI something like a conscience.
Links:
Anthropic, "Constitutional AI: Harmlessness from AI Feedback" (2022)
https://arxiv.org/abs/2212.08073
Claude's Constitution
https://www.anthropic.com/constitution
Anthropic, "Teaching Claude Why" (2026)
https://www.anthropic.com/research/teaching-claude-why