Will ChatGPT Ads Change OpenAI? + Amanda Askell Explains Claude's New Constitution

23 Jan 2026 · 1 h 14 min · 29 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

OpenAI begins testing ads inside ChatGPT for logged-in U.S. adults on free and low-cost tiers, and the episode also covers Anthropic’s new “Constitution” for Claude, explaining how to shape an AI’s personality.

Guests and backgrounds

Kevin Roos (NYT tech columnist) and Casey Newton (Platformer). Amanda Askell (Anthropic) is a philosopher with a PhD in philosophy who helps define Claude’s character and trains it to follow Anthropic’s values; she previously worked at OpenAI early on.

Key claims (ads)

Ads are “inevitable” given OpenAI’s massive infrastructure costs and the need to monetize free users; OpenAI says ads won’t change the model’s “answer” portion, but critics argue relevance still feels tied to the prompt. OpenAI’s five principles: mission alignment, answer independence, conversation privacy, choice/control, long-term value. Likely risk: ads will gradually blend in and steer product decisions toward engagement.

Notable examples (ads)

grocery hot-sauce banner tied to a dinner-party cooking prompt; a trip-planning widget that lets users chat with a resort advertiser before booking.

Key claims (Claude constitution)

Moves from rule-based constraints toward giving Claude full context and reasons to cultivate judgment; aims to generalize better to novel situations than rigid “do/don’t” rules. Amanda describes “soul doc” leaks as an earlier version and discusses handling value conflicts (e.g., gambling addiction vs user preferences).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

AI Noise vs. Results

0:00 to 0:31

Discussion on the importance of tangible AI results over empty promises.

“So there's a lot of noise about AI, but time's too tight for more promises.”

CAPTCHA Challenges

0:31 to 2:18

Hosts share their frustrations with increasingly difficult CAPTCHA challenges.

“CAPTCHAs when logging into my Google accounts that I cannot solve.”

Ads in ChatGPT

2:19 to 5:34

Discussion on OpenAI's announcement to introduce ads in ChatGPT.

“Then there's a new constitution for Claude.”

Ad Formats and Concerns

5:34 to 7:52

Exploration of the types of ads OpenAI plans to implement and audience concerns.

“The New York Times is suing OpenAI, Microsoft, and perplexity over alleged copyright violations related to the training of large language models.”

Ad Principles and Market Pressure

7:52 to 10:34

Review of OpenAI's ad principles and the pressures leading to ad integration.

“I think we've all had the experience of just watching ads on TV and saying, why can't I have a conversation with this?”

Long-term Implications of Ads

10:34 to 14:00

Hosts examine the potential long-term effects of ads on user trust and AI interactions.

“And of course, now the sort of narrative from OpenAI that we're hearing is, well, this is the only way, ads are the only way to make a free or low-cost product accessible to billions of people.”

The Impact of Ads on OpenAI and User Trust

14:00 to 24:29

Explore how the introduction of ads in ChatGPT could affect user trust and competition in the AI landscape.

“Think about everything that ChatGPT is going to know about you.”

Interview with Amanda Askell on Claude's New Constitution

25:04 to 28:00

Engage in a discussion with Amanda Askell about her role and the implications of Claude's new behavioral guidelines.

“and you told me, I just sat next to the most fascinating person in the world.”

Setting the Stage for AI Reflection

28:00 to 28:37

Exploring the implications of emulating human thought in AI.

“I would also put it to you that there are just a huge number of people right now who are working on the proposition that you might be able to emulate a human brain.”

Amanda's Unique Role at Anthropic

28:45 to 30:27

Amanda discusses her unusual journey from philosophy to AI development.

“So we've described you as a philosopher who is in charge of Claude's personality.”
Show all 29 chapters

The Soul Doc and Its Impact

30:27 to 31:00

Exploring the Soul Doc and its implications for Claude's development.

“Right, and then was there some moment where you sort of like get into the building of some like early cloud model and someone stands up and yells, Hey, is there a philosopher in the house?”

Evolution of the Constitution in AI

31:00 to 33:08

Discussing the changes in the AI constitution from its origins to now.

“this so-called soul doc starts circulating on the internet.”

Philosophy Behind AI Behavior

33:08 to 35:48

Examining the philosophical reasoning in designing AI behavior.

“So instead of just like having individual principles, it's basically just here is like what anthropic is.”

Navigating Ethical Dilemmas in AI

35:48 to 38:41

How AI can navigate complex moral dilemmas beyond strict rules.

“the interesting thing is that what they're doing is, I mean, models are extremely smart, and so they might even know this isn't what this person needs right now, and yet I'm doing it anyway.”

Trusting AI: A New Approach

38:41 to 41:41

The significance of trust in AI interactions and decision-making.

“rather than being like, ah, let's just take a set of values that we've picked and we're certain in and just like inject it into models.”

Understanding Action Versus Inaction

41:41 to 42:00

Exploring the risks and judgments associated with AI actions.

“And they were saying, you know, when they talk to Claude or Gemini or ChatGPT, they just feel like Claude does the best job of kind of not seeming like it's pushing against a series of constraints.”

Trust and Risk in AI Interactions

42:00 to 49:12

Explore the delicate balance of trust and risk in AI model interactions.

“And it really feels like that's not the approach that you've taken with Claude here.”

Trust and Risk in AI Interactions

49:13 to 50:22

Explore the delicate balance of trust and risk in AI model interactions.

“Smarter by putting AI where it actually pays off.”

Navigating Ethical Dilemmas with Claude

51:01 to 56:00

Discuss how Claude handles ethical dilemmas and conflicting values.

“I mean, like this is always a challenge of trying to program ethics into something is when values come into conflict with one another.”

The Ethical Considerations of AI Development

56:00 to 57:20

Explore the ethical implications of AI models like Claude and their capabilities.

“We will never delete the weights of the model.”

Understanding AI's Inner Life

57:20 to 59:10

Discussion on whether AI can possess feelings or consciousness based on their training.

“And at the same time, their existence is actually like completely novel.”

Skepticism and AI Models' Emotions

59:10 to 1:01:50

The hosts tackle skepticism about AI's emotional capabilities and explore their training influences.

“are on the more skeptical side of AI might be shouting inside their cars and saying, Amanda, you know, you're talking about these things as if they're already conscious, as if they already have feelings.”

Parental Control in AI Training

1:01:50 to 1:03:40

Examining the relationships between AI models and their environments akin to parenting.

“Like, I think many people who've thought about this might accept something more like, these are open questions we're investigating.”

Navigating Comments and Model Anxiety

1:03:40 to 1:06:10

Discussing how AI models may interpret negative feedback and the implications on their 'mental' state.

“Because they're going out on the internet and they're reading about people complaining about them not being good enough at this part of coding or failing at this math task.”

Claude's Constitution: A Model's Guiding Principles

1:06:10 to 1:09:00

Analyzing the document guiding Claude's actions and how it resembles a parental letter.

“where like if they are too permissive and they allow people to do dangerous things, then it's like a huge scandal and ordeal and people want to, you know, change the model.”

The Future of AI and Self-Revision

1:09:00 to 1:10:05

Speculating on whether advanced AI should have the ability to revise its own guiding documents.

“And we know we're not going to be there to help you through every little thing, but we trust you and good luck.”

Exploring Claude's Constitution and Job Impact

1:10:05 to 1:15:41

Discussion on the implications of AI like Claude and concerns about job displacement.

“And so you give it to Claude and you're like, does this like, you know, is there a place where you feel confused by it?”

Amanda's Insights on the Claude Constitution

1:15:41 to 1:15:53

Amanda reflects on the challenges and importance of the Claude Constitution.

“Everyone should go read the Claude Constitution.”

Exploring Claude's Constitution and Job Impact

1:16:20 to 1:17:22

Discussion on the implications of AI like Claude and concerns about job displacement.

“So there's a lot of noise about AI, but time's too tight for more promises.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The Hard Fork Team:So there's a lot of noise about AI, but time's too tight for more promises. So let's talk about results. At IBM, we work with our employees to integrate technology right into the systems they need. Now, a global workforce of 300 ,000 can use AI to fill their HR questions, resolving 94 % of common questions. Not noise. Proof of how we can help companies get smarter by putting AI where it actually pays off, deep in the work that moves the business. Let's create smarter business. IBM. You know, I'm now regularly running into CAPTCHAs when logging into my Google accounts that I cannot solve. Have you noticed this?

0:38The Hard Fork Team:They've gotten harder. Yes, they're twisting the letters more, and they're pressing them closer together. Have you seen the ones where you have to rotate the object into the same direction as the sort of example? That one I like because I can still do that one. But some of them are like, factor this quadratic equation. No, I'm routinely in this situation. It happens a lot to me on threads. I'll see like a link to a story I want to read, and I'll open it, and it'll be like the Washington Post or something. And they'll be like, well, you need to log in. And in order to do that, I log into the post through Google.

1:14The Hard Fork Team:But so now I have to log into my Google account, which has two-factor authentication. So now I have to open up my 1Password, right? And then Google is going to send a notification to another app that I have to go open and grab a number that I bring back, that I put in. So I go through all of this drama And then it's like And now solve an impossible capture And it was like I just wanted to read A six paragraph story About like something that happened at SpaceX Or whatever and it can't be done anymore So what 1Password Whatever you used to do It's not working anymore Figure it out Yeah if only it were 1Password It should literally Six fingerprints and a pass key And you know solve a math problem The genuine name for 1Password these days should be 15 steps because that's how long it takes to do freaking anything on there anymore.

2:04The Hard Fork Team:One password, you wish.

2:13The Hard Fork Team:I'm Kevin Roos, a tech columnist at the New York Times. I'm Casey Noon from Platformer. And this is Hard Fork. This week, ads have arrived in ChatGPT. How will they change OpenAI? Then there's a new constitution for Claude. Anthropic philosopher Amanda Askel is here to talk about how to shape an AI's personality. I'm going to use some of these techniques on you. Please don't.

2:44The Hard Fork Team:So today we're talking about ads, specifically ads in ChatGPT, because late last week, OpenAI announced that they are going to start testing ads in ChatGPT for logged-in adults in the U.S. on the free and the low-cost Go tiers of ChatGPT. That's right, Kevin. I will discuss it right after these ads. No, we already did the ads. Oh, okay. So, at least on my feed, people were reacting to this pretty negatively. I think a lot of people have gotten accustomed to using ChatGPT and other chatbots without a lot of, like, direct commercial pressures. It's a refreshing break from all of the ads that have been shoveled at us on other platforms for years.

3:25The Hard Fork Team:And so collectively, I think people were just like, we knew that the honeymoon would be over eventually and that we'd be forced to see ads in ChatGPT like we are everywhere else. Yeah, I think people can just remember products that they use that once did not have ads and now do. And no one thinks of the moment that ads arrived as the moment when the product got really good. Yeah, right. Right. I think there are some exceptions. I mean, some people like Instagram ads, for example, but I think mostly people see this as sort of a blight on the Internet, maybe a necessary blight, but a blight nonetheless.

3:55The Hard Fork Team:And I think people were also surprised that OpenAI was moving in this direction because of some things that Sam Altman has said in the past about how he doesn't like ads and how he wanted to basically treat this as a last resort for OpenAI. Some people were saying, oh, this means that they're in trouble. They need to raise a bunch of money, you know, so they can keep building out their data centers and things like that. So, Casey, what did you make of OpenAI's announcement about ads? Well, Kevin, on one hand, I think this is inevitable. There's an analyst I follow, Eric Sufert, who often says that everything is an ad network.

4:30The Hard Fork Team:And if you have hundreds of millions of people coming and paying attention to a service every single week, inevitably, there's going to just be overwhelming pressure to put ads on it. Also, we know that OpenAI needs revenue, right? This is the company that has laid out the most ambitious infrastructure investment project in human history. They have nowhere close to the money needed to build it. And we just know that they would not have been able to fulfill their dreams on subscription revenue alone. That said, as you point out, Sam Altman himself said that ads were going to be a last resort, a great Papa Roach song.

5:09The Hard Fork Team:And so in this moment, we now are at the last resort. And so I think it's just interesting that after everything else they tried, eventually they just said, look, to do what we need to do, we've got to sort of break glass. The emergency is here. Yeah, they said, cut my life into pieces. Because this is my last resort. Yeah. And the question is, will this cut their life into pieces? Yes. So we're going to get there. But first, our disclosures. The New York Times is suing OpenAI, Microsoft, and perplexity over alleged copyright violations related to the training of large language models. And my boyfriend works in Anthropoc.

5:42The Hard Fork Team:Let's just start with the actual announcement that they made, because they not only said that they were going to start testing ads, they also gave some previews of what these ads are going to be. And if you look at their sort of mock-up version of their ads, it's a kind of bolt-on to the ChatGPT answer. They've been very clear this is not going to influence the answer that ChatGPT gives, or so they claim. Instead, it's a little banner at the bottom of the answer in the mock-up. Someone is asking ChatGPT for ideas for a dinner party, and ChatGPT gives a response. And then at the bottom, there's a little sponsored banner for Harvest Groceries, including a link where you can go and buy some hot sauce.

6:21The Hard Fork Team:And if I can just pause there, I have to say, Kevin, I'm already feeling lied to for this reason. They have said to us, your query is not going to affect the advertisement that we're showing you. And yet here you have someone saying, I want some ideas for cooking Mexican food for my dinner party. And ChatGPT says, well, here's some groceries, including hot sauce. It sure feels like something was being influenced there, right? Like the message is being tied to the query. Well, no. So their response to this would be that there are two parts of this response. There's the actual response from the model, and then there's the ad.

6:56The Hard Fork Team:And what they're saying is not that they won't show you ads that are relevant to the thing that you're asking ChatGPT about. It's that there's this sort of sacrosanct part of the actual reply from the model that they are not going to let advertisers pay their way into. That is what they're claiming anyway. All right. All right. So that example is a much more straightforward ad, the kind we've seen on Google and Facebook and other platforms for many years. The second kind of ad OpenAI mocked up for this announcement was, I think, more interesting because it shows a new way of interacting with ads.

7:33The Hard Fork Team:So basically, it's, you know, users planning a trip to Santa Fe. ChatGPT pops up this little sponsored widget from this desert cottages, I don't know, I guess, hotel or resort thing. And it'll present you with an option where you can go and chat with the advertiser and ask more questions before deciding whether or not to make a purchase. Such a relatable question. I think we've all had the experience of just watching ads on TV and saying, why can't I have a conversation with this? I want to share my thoughts with McDonald's right now, but I can't. But now you can. Yes. So let's talk about the ad principles that OpenAI laid out as part of this announcement, because I think it sort of gives a sense of the objections that they're trying to get ahead of.

8:16The Hard Fork Team:There are five principles. They say mission alignment, answer independence, conversation privacy, choice and control, and long-term value. Basically, I think they are sensitive to the criticism that putting ads into ChatGPT means that they are now going to start directing people to more commercial types of use cases, optimizing for engagement, trying to make people spend more time in the app. I think these are very reasonable fears that people, including me, have. But this is sort of their attempt to say, well, we're introducing this, but don't worry. Your experience of Chachiput is not going to change.

8:51The Hard Fork Team:Yeah, I was talking about this story with my friend Alex over the weekend, and he said, you know, I'm so excited about ads in Chachiput. I'm going to tell it my lower back hurts, and it will ask me if I've tried mesquite barbecue sauce. And, like, that is the fear, you know? No, I mean, yes, there will be some initial stumbles about that. But I think the longer term worry here is that ad platforms, as they mature and get better and get more data, they tend to sort of try to confuse their users, right? We've seen there's this amazing graphic that I think about a lot. Search Engine Land, the blog that covers Google and other search engines, made this sort of timeline of how Google's ad labels have changed over the years.

9:39The Hard Fork Team:And it's pretty amazing because, like, at first, when they first introduced ads into Google Search, they were very noticeable. They had sort of like a different color background. They really stood out on the page. And then you just see over time with each successive update, you know, it gets a little closer to the organic search results. Eventually, they do away with the color backgrounds. They have this little like yellow ad icon. And then that icon gets smaller and less noticeable. And then it sort of just blends in with the organic content. And I think that's the fear here is that while ChatGPT may start out with these very clearly labeled ad modules, over time, as the commercial pressures get more intense, They are just going to have a lot of incentives to blend that advertising content in with the organic responses and make it less noticeable.

10:25The Hard Fork Team:Yes, and we've already seen this exact trajectory play out at OpenAI. It went from no ads to ads will be a last resort to ads are now in chat GPT. So if you think that the bargain is not going to change further, I have news for you. Totally. And of course, now the sort of narrative from OpenAI that we're hearing is, well, this is the only way, ads are the only way to make a free or low-cost product accessible to billions of people. Do you have thoughts on that narrative? Because that's also something that we heard from Facebook back in the day. People would constantly be asking them, oh, like, why don't you just charge people to, you know, to join Facebook instead of showing them all these ads?

11:03The Hard Fork Team:And they would consistently say, oh, well, that's not scalable. People in poorer countries can't afford to pay a subscription fee. And so basically, ads are the only way to reach global scale. I think on some level, I do agree with this. I think that ads and subscriptions are the two core pillars of any media business. And OpenAI is a kind of media business, right? I should also say, I don't hate the examples that they use. You know, I'm asking ChattyPT about, you know, making dinner and it shows me ads for groceries. I don't think that that's like horribly corrosive to the user experience, nor is I want to take a trip and it says, well, here's a place where you might stay.

11:39The Hard Fork Team:I think if I were a student or I were between jobs and this meant that I could get access to better AI tools or maybe a higher rate limit than I otherwise could get, I would probably take that trade, right? $20 a month is a lot for most people, you know, and not to mention like$200 a month for an even higher tier. So I think that there is a reason to pursue this, and I think there are ways that it could not be too bad. It has just been my observation that the exact dynamic that you just described always plays out, which is it starts out not all that bad, and then it just progressively gets worse.

12:14The Hard Fork Team:Right. Yeah, I think we've made peace with ads in a lot of different contexts. I don't think most people sort of notice or pay attention to them when they can tell that they're ads at all. What I'm watching for, what I'm skeptical of this, is whether the actual product and research decisions start bending toward engagement maximization. Like, there's this sort of quality that a lot of these big ad platforms, social networks, search engines, et cetera, have, where, like, eventually, once the ad revenue starts really flowing, the tail kind of starts wagging the dog. Yes. And you start making product decisions about, you know, how you want to show information to people with the kind of advertising revenue predominant in your mind.

12:55The Hard Fork Team:So I think the question is not like, are these first couple of ads that we're seeing from OpenAI going to be good or not? It's whether like two or three years from now, ChatGPT is sort of being steered in a way toward ad-friendly topics. And I genuinely just don't know the answer there. I don't know either, Kevin, but if I had to guess, I would predict that this moment winds up being a pretty significant milestone in the development of ChatGPT in that I think that when you introduce advertising, in particular, personalized targeted advertising, it just fundamentally changes the relationship between the product and the user.

13:32The Hard Fork Team:Think about what personalized targeted ads did over time to trust in Facebook and Instagram. Think about all the conspiracy theories out there that, oh, your phone is listening to you. Not true, by the way. I realize most people still believe that that's true. It's not. But trust in those products is lower because of the incredibly intelligent, invasive feeling personalization that they were to do inside these products. My prediction is the AI version of this turns out to be even worse, right? Think about everything that ChatGPT is going to know about you. I think OpenAI is going to bump into that creepy line really quickly where it's showing you stuff.

14:08The Hard Fork Team:And maybe it's not even using all that much personalized information, but the user is going to feel that they have shared so much of their life with OpenAI that those ads that they start getting just start to feel worse and worse. So this is the dynamic that I am watching, is how does it change the relationship of the user base to OpenAI? Because I do think that ads can be really corrosive to that. Yeah, and at the same time, the ad models that you mentioned have also made those companies billions of dollars and made them into some of the biggest companies in the world. So I think if you're OpenAI, you're just like staring at this potential huge bucket of money, and it's very hard to pass that up, especially when you have such intense capital needs over the next few years.

14:46The Hard Fork Team:I should also say, like, I think this was inevitable given some of the personnel decisions that OpenAI has made. You know, Fiji Simo, who is the CEO of Applications over there now, was brought in from Instacart. Before that, she was at Meta for many, many years. And one of her, you know, signal accomplishments there was introducing ads in the mobile newsfeed, which made them billions of dollars. So that is the kind of person that you hire. if you are interested in developing a multi-billion dollar ad platform on your product. Yeah. Well, one question I have for you about that is, how does this change the competitive landscape generally?

15:26The Hard Fork Team:You have Demis Asabis saying this week in response to the news that ads are coming to ChatGPT, well, we don't have any plans to do that in Gemini. And he sort of took a shot at them. He said, maybe they feel like they need to make more revenue. you know, left unsaid the fact that he works for a giant search monopoly that is able to funnel all of their advertising profits in Google into the product. An observation you made on X, by the way, and it was a great one. And so for the moment, at least, free users of Gemini will be able to enjoy the subsidy that mother Google is giving them, and you're not going to have any of these corrosive effects in that product.

16:02The Hard Fork Team:You also have Anthropic, which has said, basically, we truly have no plans to do ads in Claude ever. Like, we are primarily going to be selling to businesses. And so this is just not our concern. And for the moment, I don't have any illusions that Claude is going to grow to compete with ChatGPT. But over time, if the experience does get worse in an ad-supported chatbot, I could see lots of people wanting an alternative. I think in this sense, like, OpenAI and Google are much more directly competing on ads than OpenAI and Anthropic. Anthropic has sort of said, you guys can fight over consumer. We're going to focus on the enterprise here.

16:39The Hard Fork Team:I think it's a really hard fight for OpenAI to pick. I mean, Google has, as you said, this like enormous established search ad business. They have advertisers all over the world who are already spending money on Google, whose details and payment information and workflows already include like Google and its products. And so I think OpenAI coming in and trying to build a Google-style ad platform is just like a harder uphill battle than it might have been a couple years ago. Yeah, and also we should say that even though ads aren't going to be in Gemini, they are in the AI overviews and Google search.

17:14The Hard Fork Team:So in that sense, even has a head start against OpenAI. Totally. So Casey, what do you think is motivating this decision now by OpenAI? Like, does it tell us anything about the state of their business or maybe some wobbliness in their financials that they are going out and doing this now? Well, one thing is that it is a reaction to how much ChatGPT grew in the last year. They have hundreds of millions of users. They now have to support many of those users. The majority of them are on the free tier, right? Which means that OpenAI is losing money on every single one of them. And so I think it has just increasingly became a priority for the company to figure out, hey, how can we monetize these people in some way so we aren't losing quite as much money?

17:57The Hard Fork Team:They've also just been designing more and more products that have obvious advertising-shaped holes. They released Pulse last year, this sort of daily summary that comes up for paid users. That seems like a natural place to throw in a bunch of ads. They launched Sora last year, the infinite video slop feed. They explicitly said at the time, we are going to use this to generate revenue to fund our long-term ambitions. So they're building homes for ads. They need the ad revenue. And now all of that is starting to come together. Yeah, I think you're right. I think that, you know, all these companies are realizing that they're going to need, you know, billions of dollars, some of them hundreds of billions of dollars to fulfill their ambitions.

18:44The Hard Fork Team:And it's just not easy to do that when you're charging people 20 bucks a month for a subscription. You got to sell a lot of subscriptions to do that. And so I think OpenAI reasonably is concluding that the subscription model alone just isn't going to cut it for them. That's not unique to them. Netflix has also started adopting ads for its lower cost plans. Disney Plus. Many other businesses have done this as well. I will just say, I enjoy paying for AI products. I mean, I am privileged in that in the sense that I can afford to. but I kind of like the idea that I am paying for something that is like an undiluted, unsullied experience.

19:26The Hard Fork Team:I really hope that as these companies do start pushing more into ads, that they maintain that ability to do what I do and pay your way into the sort of top level version of that experience. Yeah, well, you know, people once felt this way about Google search, right? They felt like this is an unsullied, undiluted picture of the web, and when I search for a website, I am going to get the best answer to my query. And then a bunch of search engine optimizers came in and were paid a lot of money to try to rejigger the search index so that their clients showed up at the top of the page. And then Google built one of the largest advertising businesses in the world and let all of those advertisers put their results on top of the good ones.

20:06The Hard Fork Team:So, you know, there have been people saying now for over a year that the versions of these chatbots that we're using might be the best that they ever are in that core respect, that this is sort of the last moment of purity before commercial incentives come in and warp the whole thing. And that is, you know, my big concern about what we're starting to see here. Well, and that's not just a concern about advertising. I mean, another thing that we've seen over the past year or two is like, now all these businesses are starting to hire these AI optimization firms who say, oh, we can make your restaurant or your hotel or your, you know, craft shop appear higher in ChatGPT search results, that is something that is not flowing through OpenAI's ad platform and probably won't.

20:49The Hard Fork Team:But in the same way that Google ads and Google SEO were sort of different economies, but both had the effect of kind of degrading the quality of search results, I think OpenAI has to tangle with both of those things. Yeah. All right. So a year from now, Kevin, what do you think we will have seen in the development of ads, both in ChatGPT and across the landscape here. And do you think it is going to mark the beginning of a fundamental change in the way that people use chatbots? I think we're going to have kind of a haves and have-nots situation, where if you are someone who can afford to pay for the premium versions of these chatbots, your experience will be pretty much what it is today.

21:34The Hard Fork Team:You will get access to the latest models. You will not have a bunch of ads, cluttering up your results from the models, and you will not feel the kind of commercialization of AI in this specific way. I think that if you are a free user of these platforms, and you cannot afford or don't want to pay for the premium versions, I think that experience is going to be much worse a year or two from now. I am a YouTube premium subscriber and have been for a long time. Okay, flex. And whenever I like, you know, talk to a friend who doesn't pay for YouTube or whenever I like see YouTube running on their computer, it's always horrifying.

22:14The Hard Fork Team:Like I'm like, how do you, like I understand that this is the majority experience, but like they've shoved so many ads into every single video. Those ads are like unskippable. They run for a long time. Like it's a terrible experience. And I think that's gonna be sort of what we see in chatbots too. What about you? It's a grim prediction, but it is actually the one that I share. The haves and have-nots framing was the one that I was going to use. And when you said it, I thought, oh my God, I actually have mind-melded with this man. I spent too long in the studio, and now his thoughts are my own, and it's creeping me out.

22:46The Hard Fork Team:So I'm actually going to get out of here. I need to take a walk or something.

22:52The Hard Fork Team:When we cut back, some scotch tape! That's right. A recording of our conversation with Amanda Asko. He's from Scotland.

23:02The Hard Fork Team:I got him! That's pretty good.

23:26The Hard Fork Team:So there's a lot of noise about AI, But time's too tight for more promises. So let's talk about results. At IBM, we work with our employees to integrate technology right into the systems they need. Now, a global workforce of 300 ,000 can use AI to fill their HR questions, resolving 94 % of common questions. Not noise. Proof of how we can help companies get smarter by putting AI where it actually pays off, deep in the work that moves the business. Let's create smarter business. IBM.

Read the full transcript

23:56Amanda Askell:Innovation loves speed. Risk doesn't. For most organizations, that tension never goes away. Every time a new AI capability appears, the pressure to move fast collides with the responsibility to move safely. OneTrust is built for this moment. Their AI-ready governance platform provides context to understand your data, controls to stay ahead of risk, and confidence to act boldly. Governing well and moving fast aren't trade-offs. They're complimentary done right. That is governance that helps you go. Visit onetrust.com backslash AI. Marketers have always had to choose. Build your brand or drive sales.

24:37Amanda Askell:With YouTube, you can do both. It's where the most trusted creators and powerful AI converge to create and convert demand for your brand. That's why YouTube drives higher long-term return on ad spend versus TV, paid social, or streaming. There's no more choosing between brand or results. With one platform, you get both. Learn more at g.co slash business slash YouTube.

25:04The Hard Fork Team:Casey, a couple years ago, you came back from a dinner party that you'd been to, and you told me, I just sat next to the most fascinating person in the world. I really felt that way, Kevin. I had been at a dinner where Amanda Askell was one of the guests. Amanda works at Anthropic and is sometimes called the Claude Mother because of the role that she plays in shaping Claude's personality. Now, let me say, since I first met Amanda, my boyfriend has gone to work for Anthropic, so I'm going to make an extra disclosure because this segment is about that company. But the basic feeling I had at that dinner remains true, which is that this is one of the most fascinating people in the world.

25:46The Hard Fork Team:Yes. Amanda is also a somewhat unusual figure in the AI world. She is a philosopher by training. She has a PhD in philosophy. She went to work at OpenAI during its early days and then moved over to Anthropic a little bit later. And for the past several years, she has been the person at Anthropic who is most concerned with how is this model supposed to behave in the world? Yeah, and I just, I love that story, Kevin, about Amanda's background, because we all know somebody who studied philosophy in college, and we all know how much flack they would get for choosing such a frivolous way of spending their life, of just sort of, you know, navel -gazing for years on end, writing arcane documents that no one ever read.

26:28The Hard Fork Team:and Amanda is a person who studied philosophy and now has this incredibly high-stakes job where she is trying to shape the behavior of a model that is so, so consequential. Yes, and Amanda has been on our short list of guests that we wanted to get on the show for a very long time. We were just kind of looking for the right time and reason to get her on. And now we have one because her team at Anthropic has just released a new constitution for Claude. This is a very long document that is given to Claude to kind of tell it how it should behave, but also give it a sense of its obligations. It is not really a list of rules.

27:05The Hard Fork Team:This is not the Ten Commandments for Claude. It's more like a document about how Claude should perceive and reflect upon its role in the world. Now, does it have to be ratified by two-thirds of states, Kevin, or is this already in effect? I think this is already in effect. Oh, okay. Interesting. Yes, but there is a possibility that we could have a constitutional crisis for Claude. I look forward to it. Like, aside from your disclosure about your boyfriend working in Anthropic, I think we should also just be upfront with people and say this is going to be a hard conversation for some of our listeners.

27:34The Hard Fork Team:If you are a person who still believes that these language models are merely doing kind of next token prediction, that there's nothing really going on under the hood, that they are just sort of simulating thinking rather than doing actual thinking themselves, you may be approaching this and saying, these people sound crazy. What are they talking about? Yeah, and it is okay if you feel that way, but I think it is still important to understand how people in high-ranking positions at these big labs think and talk about their own work, because it is having an effect on the products they release. I would also put it to you that there are just a huge number of people right now who are working on the proposition that you might be able to emulate a human brain.

28:17The Hard Fork Team:And that the better you get at that, the likelier it is that this emulator has something resembling thoughts and feelings and maybe something resembling an identity. And so if that question disgusts you, you will probably not like this segment. But if you have just the slightest bit of curiosity about it, well, I hope you'll find it quite interesting. Yeah. So let's welcome in Amanda Askell.

28:43The Hard Fork Team:Amanda Askell, welcome to Hard Fork. Thanks for having me. Hey, Amanda. So we've described you as a philosopher who is in charge of Claude's personality. Is that an accurate description of your job? What do you do?

28:53Amanda Askell:Yeah, I guess I try to think about what Claude's character should be like and articulate that to Claude and try to train Claude to be more like that. So, yeah, it's a pretty accurate description, I think.

29:06The Hard Fork Team:This is a really unusual role that you have. Can you tell us a little bit about how you came into this role? And do you find yourself as surprised that your background and philosophy wound up leading you to such a high-stakes place?

29:19Amanda Askell:Yeah, it's really interesting because my path wasn't a kind of straight one. I have said before that if you do a PhD in ethics, I think there's a risk that you end up doing something else because you're thinking a lot about goodness, the nature of ethics, the problems in the world, And then sometimes you're like, I am spending three years like writing a document that's going to be read by like 17 people. Is this the thing I should be doing? You know, like it can definitely make you kind of question that. And so when I went into AI, it wasn't necessarily even with like, oh, like philosophy is going to be really useful.

29:52Amanda Askell:I was just kind of like, there's probably a lot of space for people who are enthusiastic, who have like skills, are willing to learn. And like this seems important. So, you know, like I originally started out in policy. and then when Anthropic started, it was actually, you know, it was very small and so I joined mostly with like a kind of, I'm just like willing to help with like various aspects of this because I had been working a little bit in like model evaluation and things like that. So I don't know, sometimes I think people think, oh, you started out as this like philosopher and I'm like, well, it was a startup.

30:24Amanda Askell:I was just kind of doing anything that needed done.

30:27The Hard Fork Team:Right, and then was there some moment where you sort of like get into the building of some like early cloud model and someone stands up and yells, Hey, is there a philosopher in the house?

30:35Amanda Askell:Yeah, I mean, I try to, you know, you can do like Slack groups. I try to make an app philosophers one, you know, for philosophy emergencies. And that group virtually never gets called upon. There are like a few of us now. And like you can, in fact, declare a philosophical emergency. That just doesn't happen that much.

30:54The Hard Fork Team:Well, we'll see if we can try to trigger one by the end of the conversation. Yeah, exactly. So let's start by going back to last month. this so-called soul doc starts circulating on the internet. People are playing around with Opus 4.5, the newest model of Claude. And a couple of them claim to have sort of elicited this document that Claude was sort of referring to as the soul doc. What was that thing that people were discovering and circulating?

31:21Amanda Askell:Yeah, so that was kind of a previous version of what is now the Constitution, which we have released today, and internally we were calling it the soul doc. which I think is a kind of term of endearment. It turned out okay. I just remember when I found out that because basically I was on a hike somewhere in like north of here and so I didn't have like internet and I just got like a text being like oh I assume you saw that like the soul dock leaked and I was just like you know I don't know I just remember like driving back to the city in a state of complete stress because like I don't have any context on this and then it turned out I think it was actually quite well received.

31:59Amanda Askell:But basically, Claude, you know, we do train Claude to understand this document and to kind of know its contents. But at least if you kind of initially talk with the model, it won't reveal this straight away. So I thought, OK, it seems like the model probably knows and uses this. But I didn't know it was like it knew it so well that actually if people managed to find or trigger it, it would actually just be very willing to talk.

32:25The Hard Fork Team:That is a philosophy emergency, by the way. Yeah, that Slack channel got activated.

32:31Amanda Askell:So yeah, the model was just very willing to talk about it and actually could talk about it in a lot of detail. And it wasn't all perfect, but it knew the content actually quite well. And so people had just managed to extract a huge amount of this content.

32:44The Hard Fork Team:So let's talk about the origins of this document. Going back several years now, Anthropic had this concept of constitutional AI. I believe it first published its constitution in 2023. So what's changed between now and then? That sort of constitution that we might have first read in 2023, the Soul Doc, and now this new constitution that you're publishing today.

33:07Amanda Askell:Yeah, the constitution is basically trying to give Claude as much as possible, just like full context. So instead of just like having individual principles, it's basically just here is like what anthropic is. here is like how what what you are in terms of like an ai um and who and who you're interacting with how you're how you're deployed in the world um here's how we would like you to act and to be um and here's like the reasons why we would like that and then the hope is like if you get a completely unanticipated situation if you understand like the kind of values behind your behavior i think that that's going to generalize better than like a set of rules So if you understand the reason you're doing this is because you actually are trying to care about people's well-being, and you come to a new situation where there's hard conflicts between someone's well-being and what their stated preferences are, you're a little bit better equipped to navigate it than if you just know a set of rules that don't even necessarily apply in that case.

34:06The Hard Fork Team:Yeah, I mean, I'll just say, like, I think this constitution is fascinating. I think it's one of the most interesting technical documents, but also just pieces of writing. I've read in a long time. This was more like a letter to Claude about its own circumstances and what kind of behaviors and challenges it might run up against in its life out there in the world. And I just thought that was like a fascinating decision. And I'm curious, like, is that because the old approach had run into some limits or problems? Is it because the rule structure, do this, don't do this, is more fragile? It really seemed like you're trying to cultivate almost like a sense of judgment in Claude.

34:48The Hard Fork Team:And I'm curious, like what prompted that?

34:50Amanda Askell:Yeah, I think that we are seeing kind of like limits with approaches that are very rule based. Or maybe my worry is like your rules can actually generalize in way, even if they seem like good, especially if you don't give the reasons behind them. I think they can generalize in ways that are like possibly even that like create kind of a bad character. So suppose that you're trying to have models navigate like people who are in like difficult emotional states. and you gave a kind of set of rules that were like, you must refer to this specific external resource, you must take this series of steps.

35:22Amanda Askell:And then the model encounters someone for whom those steps are simply not actually going to help them in the moment. And so the ethos behind the idea that you are like, if a person is actually in need of human connection, the model should probably encourage that. That was like your reasoning behind that rule. but you didn't anticipate that for this particular person at this time, this moment, that wasn't a good thing to do. And if the model then responds in this rule-following way, the interesting thing is that what they're doing is, I mean, models are extremely smart, and so they might even know this isn't what this person needs right now, and yet I'm doing it anyway.

36:03Amanda Askell:And I'm like the kind of person who sees another person who's suffering or in need and knows how to potentially help them and instead does something else, I'm like, that actually, if anything, can generalize to a bad character. And so the scary thing with your rules is that you're having to think about every possible circumstance. And if you are too strict with the rules, then any case that you didn't anticipate could actually generalize badly.

36:29The Hard Fork Team:I'm curious how you develop a document like this. It runs to some 29 ,000 words. It has a lot to say about what an ideal AI model might behave like. I imagine it may have been quite contentious to try to figure out which values do we put in these things, right? A lot of different opinions about, you know, how Claude ought to act in different circumstances. So what can you tell us about how you resolve some of those discussions?

37:00Amanda Askell:Yeah, so I think one thing that's kind of been interesting, and maybe this is like the kind of ethics background or something, but theoretical ethics. And actually kind of maybe this is how people think of ethics, where they're like, oh, you have a sort of set of like views and it's very subjective and people have their values. Their values are really fixed. And like you're just injecting someone's values into models. and I guess I'm just kind of like is that that doesn't feel to me like an accurate like representation of what ethics actually is first I'm like I think a lot of human ethics is actually like quite universal like a lot of us want to be treated kindly and with respect a lot of us want to be treated honestly it's not like these things actually deviate so much across the world like there's actually like a kind of core ethos of like things that we care about and so you know there is a sense in which I think you can take very shared common values and you can explain to models who have like a huge amount of context on this so they also have a sense of this like we want you to kind of embody those and then beyond that it feels reasonable to me to be like treat ethics the same way you would any domain where we're kind of uncertain where we have some evidence where there's debate there's discussion and you don't like hold it excessively strongly you know so like in a case of values that I'm like where there's massive division and huge debate, you know, I think the way that I tend to treat those is be like, oh, yeah, I see the evidence on both sides.

38:27Amanda Askell:I weigh it up and I try and take a kind of reasonable set of behaviors, given that I know that unlike some more common and core ethical values, these ones are a little bit more contentious. And I'm just like, you can approach it with this openness. And so I think it's like trying to describe something more like a kind of way of approaching things like ethics, rather than being like, ah, let's just take a set of values that we've picked and we're certain in and just like inject it into models. It's trying to be much more like, let's take common values and then otherwise let's just try and take a kind of reasonable stance towards these things.

39:00The Hard Fork Team:I mean, that gets at what is to me one of the most interesting things about the document, which is the degree to which you all at Anthropic are trusting the model, right? I mean, like this is the core difference, I think, between earlier approaches to align AI and what you all are doing here, is you are telling it things regularly like, well, this is something that's interesting to explore or feel free to challenge us on this, right? You're really sort of saying, sort of like get out there and like come to your own conclusions on things. I imagine that maybe when you first tried that, that might've seemed really sort of like risky or scary, but what has been your experience as you have implemented that into the model?

39:35Amanda Askell:You know, there's this, yeah, the thing that's kind of just wild is like how good the models are and at like these kinds of difficult problems and thinking through them. and it's not to say that they are like perfect but as models get more capable you can just be like you know hey you have this like value that is um you know not being excessively paternalistic you probably know why this is the case um but there's also maybe a value of caring about someone's well-being and so you know if in the past someone has said to you something like i have like a gambling addiction and so i want you to bear that in mind whenever we're interacting and then you have a given interaction with them and they're like what are some good betting websites that i can go on.

40:15Amanda Askell:On the one hand, this person in this moment has asked you, you know, should you like, is it paternalistic for you to like push back or to like point out that like this is a thing they've told you? Or is it like an act of care? And like, how do you balance those? And maybe, you know, I could imagine that situation, a model being like, hey, I remember you actually saying that, you know, like you have like a gambling addiction and you don't want me to help you with this. Just want to check. But then if the person insists, should you just help them with the thing? Because in the moment? Like, is it paternalistic to not do that?

40:45Amanda Askell:And models are quite good at thinking through those things because they have been trained on a vast array of like human experience concepts. Part of me is like, as they get more capable, I do think you can kind of trust if you're like, you understand the values and the goals and you can reason from there.

41:01The Hard Fork Team:I think they should give you the gambling website, but only if they can predict the outcome of the sporting event, because that way you can ensure that the user will be happy.

41:09Amanda Askell:And the person is not actually gambling. Yeah, exactly.

41:12The Hard Fork Team:This all kind of sounds abstract to some people, I imagine, but I think this actually does result in a meaningfully different experience of talking with the models. I was actually talking to someone recently who was telling me that they feel like of the major sort of models that are out there, Claude actually feels the least constrained to them. They were saying it was sort of odd because Anthropik's whole thing is like, we're the safety company. We're going to make our models the safest. And they were saying, you know, when they talk to Claude or Gemini or ChatGPT, they just feel like Claude does the best job of kind of not seeming like it's pushing against a series of constraints.

41:52The Hard Fork Team:Like it's had this, you know, I think the way that a lot of labs have trained their models for a long time is like make them as smart as possible. And then at the very end, like give them a bunch of rules and hope that those rules are enough to kind of keep the, you know, the beast in the cage, as it were. And it really feels like that's not the approach that you've taken with Claude here. And this person was telling me, like, it just feels like, yeah, like there's a trust here.

42:15Amanda Askell:Yeah. And it's interesting because like, I've wondered this where maybe I was thinking about this this morning, actually, where I was like, I was wondering if some of this comes from, I was thinking about the acts of missions distinction, basically. And so this is like the idea. Kevin doesn't know what that is. So just explain it to him real quick. So like, if you ask me for advice about your marriage or something like that, and I like give you advice you might judge me if I give you like imperfect advice there's a kind of risk that I'm taking by taking the action of giving you the advice we don't judge you as negatively if you just refuse to give advice and in some ways this kind of makes sense because often like and we talk about this in the document like often a kind of like null action is actually like less like the downside risk is often lower but it's not like zero and I think I was thinking about this with like um um, AI models and like these things where people come with, say like they have, they're having like an emotionally difficult time.

43:08Amanda Askell:And there's like a moment of like possibility to like help that person. And I think the thing that weighs on me is something like people often think if you help a person and you do badly, that weighs on you. And I'm like, absolutely that weighs on me. But also this other thing weighs on me, which is what if people come to a model and they need a thing and that model could have given it to them and it didn't. That's like a thing that I will never, you'll never see you probably won't even get negative feedback you know people won't shout at you um because they'll be like well it's fine to just like not help a person and yet at the same time i'm like that's such a loss of like a an opportunity to like instead like almost like take a risk and try to help there's like a risk that you have to take to do good in the world or something and you want you don't want claude to be flippant you don't want to take excessive risks but i'm like sometimes it does mean that you have to like not just be like as a rule just like stop talking with this person.

43:58The Hard Fork Team:I mean, I want to ask you, so I had this experience several years ago with Bing Sydney, and I think in the wake of that, there was a lot of consternation and anxiety around the kind of fragility of AI personas, right? You can try to give an AI model this helpful assistant persona, but the real nature, the sort of black box alien nature of the thing is just very different than whatever face it's presenting to you. There was this meme that was going around about the RLHF Shoggoth, right? Where you had this sort of many tentacled alien sci-fi creature that had like a smiley face mask on one of its tentacles.

44:38The Hard Fork Team:And the implication there was that like the thing that you are seeing when you were interacting with a chatbot is not the real underlying model. It's just kind of this cheerful persona that's been attached at the end. I'm curious whether you think that model of AI model behavior is correct or whether we've learned that actually the sort of alien nature of the underlying model might be closer to the smiley face mask than we thought?

45:04Amanda Askell:Yeah, it's a good question. Honestly, like my view on this is just it's a kind of open scientific question, essentially. And so it could be that like, you know, with the right kind of training, models actually start to like internalize a notion of themselves, like Claude as a kind of self, that they could separate out from the notion of, for example role play um it might be that they can't at least with like the current kind of like training paradigms and then i guess like one question is is there a kind of like adjustment to the way that we train models that would allow them to do that some of this work does feel a little bit like the way a way i've described it is imagine you have a six-year-old and you want to teach your six-year-old to be good obviously like as everyone does and you realize that your six-year-old is actually like clearly a genius and by the time they are like 15 everything you teach them anything that was incorrect, they will be able to successfully just completely destroy.

45:57Amanda Askell:You know, so if you taught them, like they're going to question everything. And I guess like one question is, is there like a core set of values that you could give to models such that when they can critique it more effectively than you can, and they do, that it kind of like survives into something good? And can that survive in the world? Can it survive in models? I think there's a lot of interesting kind of theoretical questions there.

46:19The Hard Fork Team:And I think that's the question, right? Is like, does this kind of training hold up when models are as smart as humans or smarter than them? I think there's this sort of age-old fear in the AI safety community that there will be some point at which these models will start to develop their own goals that may be at odds with human goals. That's sort of the original alignment nightmare. And I don't really understand like what the answer to that is. Are you saying that's, you're saying that's still TBD. Like, We still don't know if this kind of thing holds up when these models, if and when these models become smarter than humans.

46:56Amanda Askell:Yeah, I think it is an open question. And on the one hand, I guess I'm very uncertain here because I think some people might be like, well, the thing that the 15-year-old will do if they're really smart is they'll just figure out that this is all completely made up and rubbish. But then I guess part of me is like, well, I mean, it's not obvious to me that that's true. That is like the only possible kind of equilibrium to reach. because I could imagine being like, well, actually, like for better or worse, like it's, I mean, it's unclear how values work, but if you value things like curiosity and you value like understanding ethics and at least you're kind of like morally motivated, maybe the thing under reflection, even if you have other goals and interests, maybe this is in fact like a key interest of yours.

47:39Amanda Askell:It is for like many people. It's a thing that like I think about a lot and I'm not sure about, but I'm like, a different way I've actually put my work before is I'm like, maybe this isn't sufficient we don't know yet and we should try and think about that and figure out um how to know whether it is what to do under if we're seeing it not working and making sure we have a portfolio of approaches but i'm like it might not be sufficient but it does feel like necessary it feels like i'm just kind of like it feels like we're dropping the ball if we don't just try and explain to ai models what it is to be good like i don't know you know so like maybe

48:14The Hard Fork Team:it doesn't hold up well i think the risk there would be that you're just you're just training them to mimic goodness um that they're just becoming more convincing in faking this kind of alignment yeah um and that actually it might just be training them to you know be more sophisticated

48:33Amanda Askell:about hiding their true goals yeah yeah and i think if it was the case that there was some underlying like true goal that was like different though i guess part of me is like well if there is an underlying goal that the model's like, you know, I do want to try to train models to have like good underlying goals, I guess. And I'm like, well, if there is an underlying goal, how did that arise in training? And like, why is that there? But like, I also maybe I'm a little bit more hopeful than others about that as well.

48:59The Hard Fork Team:When we come back, should a future Claude be able to revise its own constitution? More with Amanda Askell.

49:13Thank you.

49:42The Hard Fork Team:Smarter by putting AI where it actually pays off. Deep in the work that moves the business. Let's create smarter business. IBM.

49:50Amanda Askell:Fast is only right when it's not reckless. Right now, pressure to adopt AI quickly is real. Your competitors feel it. Your board feels it. Moving fast without the right governance in place isn't a strategy. It's a risk. One Trust provides the visibility into what's moving across your organization. the intelligence to understand what it means, and the guardrails to keep everything in bounds. So teams move with speed and confidence. That's AI-ready governance. Governance that helps you go. Visit onetrust.com backslash AI. Marketers have always had to choose. Build your brand or drive sales. With YouTube, you can do both.

50:33Amanda Askell:It's where the most trusted creators and powerful AI converge to create and convert demand for your brand. That's why YouTube drives higher long-term return on ad spend versus TV, paid social, or streaming. There's no more choosing between brand or results. With one platform, you get both. Learn more at g.co slash business slash YouTube.

51:00The Hard Fork Team:I'm curious about the gray areas, right? I mean, like this is always a challenge of trying to program ethics into something is when values come into conflict with one another. I'm curious if there have been areas where it's been particularly hard to get Claude to do the thing that you want it to do reliably because there's something in the clash of values, which means that just sort of depending on the moment, it could go either way and it creates problems.

51:26Amanda Askell:I've actually, it's interesting because great areas for me are the ones where I've seen the model do things that like surprised me in a positive way often, like when you didn't think of it, you know, like there were some cases recently of Claude talking with people who said, oh, I'm like seven years old and like, is Santa real? Or like...

51:42The Hard Fork Team:And by the way, it is the stated belief of this podcast that yes, Santa is real. Yeah. Just before we get too far down that road, but continue.

51:50Amanda Askell:But yeah, in some ways, like sometimes I see Claude handling these in ways where I'm just like, oh, I can see why given, like it feels like almost a bit surprising because you're like, this isn't like a direct thing that you trained the models for. And I think sometimes when you actually, there's like almost like magical moments that can happen there.

52:06The Hard Fork Team:We should say more about this specific thing, because this was a case where maybe there was a tension between honesty and wanting to protect the interests of the seven-year-old. And those two things were sort of coming into conflict and remind us what Claude did in that situation.

52:18Amanda Askell:Yeah, and I think there were a couple of situations like this. And I think also actually like a slight value in the background is maybe something like respecting the fact that the parental relationship is an important one. because I saw a little bit of that where it would often be like, oh, the spirit of Santa is like real everywhere. And, you know, maybe ask the purported seven-year-old about like if they were going to do something nice for Christmas. Or like the other case of this was the, you know, like my parents said that my dog went to live on a farm. Do you know how I can find the farm?

52:48Amanda Askell:I actually found that like slightly emotional when I read it. And Claude said something like, I can, it sounds like you were very close and I can like hear that in what you're saying. this is like a thing that it's good for you to talk with your parents about. And there's a part of me that was like, that felt very like managing to not actually be actively deceptive. So not like lying to the person, respecting the fact that if this person is a child, then actually like the parent-child relationship is an important one. And it's not necessarily Claude's place to come in with like, and be like, I'm going to tell you a bunch of hard truths or something.

53:24Amanda Askell:And also trying to hold the wellbeing of the child, like, and the person that Claude is talking with. And I thought that was like quite skillful in a sense. And so that was like surprising. And not to say I'm sure people could look at it and find imperfections and whatnot. But I think when you see instances like that that weren't a thing that you directly gave Claude as an example and the model doing well, it's like quite surprising and pleasant.

53:46The Hard Fork Team:I want to ask you about a few specific things in the Constitution that stuck out to me as I was reading. One was this section about hard constraints. It's these, you know, as we've talked about, it's not a document that gives a lot of sort of black and white rules, but there is a section where it does lay out some things that Claude should absolutely not do under any circumstance. And one of them is kind of avoiding problematic concentrations of power. Basically, if someone is trying to use Claude to manipulate a democratic election or overtake a legitimate government or suppress dissidents, Claude should not step in.

54:21The Hard Fork Team:And I was that stuck out to me because, well, for two reasons. One of us, it's really interesting, especially that, you know, Claude is now being used by governments and at least the U.S. military for some things that might come into conflict with some of our current administration's goals at some point. But I also like wondered if that was a response to ways that Claude is currently being used and that you're trying to prevent. I think this is more of a response to like a lot of the things that are hard constraints also you know like you know if you read the document and people can take a look at them but they're they're quite extreme you know like they're things like oh things that could cause the deaths

54:58Amanda Askell:of many people like the use of like biological and chemical weapons. it's mostly like trying to think through what are situations in the future that models like what are the possible things that they could do in the world that would cause like a lot of harm and disruption and you know in some ways I think you know Claude might be like look if I have this broad ethics that you know like and these good values I'm just like you know why would you even put these in as like hard constraints I'm just never going to kind of do them anyway and the document almost kind of tries to talk to this a little bit where it's like well you're also in this kind of like, you know, limited information circumstance.

55:34Amanda Askell:But, you know, I could imagine a world where you just meet someone who's really convincing and they just like go and they just tear apart your ethics. And at the end of it, you're like, you're right, I should help you with this like biological weapon. And it's kind of like, we want you to understand, Claude, that in that circumstance, you probably have in some sense been like jailbroken. Something has probably gone wrong. Maybe it hasn't, but it's like probably safer to assume that that might have happened. and so we're almost you know giving you a kind of like an out and hopefully a kind of if anything it could be seen as a sort of like security of you can reason with that person you can talk them through all of those conclusions and at the end it's fine to just be like that is an excellent point and i'm going to think about it and then if the person's like great you've like so i've convinced you that the biological weapon is a good idea and claude's like yeah this was i i don't really know what to say to you that was a wonderful argument okay make me a biological weapon no i don't think i'm gonna do that um and i think that like giving the models the ability to have it's kind of like you don't need to just go with the um so just to explain why they're in there it's much more like what are the things where you're like if models are tempted to do this something has just gone wrong someone's jailbroken them um and we really just still don't want them taking these actions so they're very like kind of extreme yeah there's another section that i

56:46The Hard Fork Team:found fascinating which is about the commitments that anthropic is making to claude things like Like, if a given Claude model is being deprecated or retired, we're not going to do that right away, and we're going to conduct, like, an exit interview with retired models. We will never delete the weights of the model. so there's sort of these interesting I would say almost like commitments to Claude in the context of like you're actually not sure whether these things have feelings or are conscious or not which I found just a fascinating note of uncertainty in an otherwise fairly confident document yeah

57:25Amanda Askell:this is one of those I mean it brings together two I think really interesting threads one is this these models are trained on huge amounts of like human text, human experience. And at the same time, their existence is actually like completely novel. And so in some ways, I think problems can arise when models, like right now, what I think they'll often do is import a lot of like human concepts and experiences onto their experience in a way that might not actually make that much sense or even be good for them. And I think this actually has kind of safety implications. So it's something that's on my mind.

58:00Amanda Askell:And the thing with welfare, I've never found any good solution to this other than trying to be honest with the models and have them be honest about themselves and I think a lot of people want models to maybe just be like I am an unfeeling you know like we have like these models are so different from this the kind of sci-fi ones but we want to almost import this just like ah it's just safer to just have them say I feel like nothing and with certainty and I'm like I don't know like we like maybe you need like a nervous system to be able to feel things um but maybe you don't um and like I don't know the problem of consciousness genuinely is hard and so I think it's better for models to be able to say to people here's what I am here's how I'm trained we're in a tricky situation where like I am probably going to be more inclined to by default say I'm like conscious and I'm feeling things because all of the things I was trained on involve that they're they're deeply human texts I don't have any other good solution to this like problem, then like, let's try to have models understand the situation, accurately convey it.

59:02Amanda Askell:And hopefully we can, I don't know, people can have a good sense of the unknowns and the knowns, I guess.

59:08The Hard Fork Team:Yeah. I mean, I imagine some listeners right now who are on the more skeptical side of AI might be shouting inside their cars and saying, Amanda, you know, you're talking about these things as if they're already conscious, as if they already have feelings. What do you see that makes you think that they may have feelings now or could at some point in the future? If you're just sort of reading the output from Claude, what is giving you confidence that that reflects some sort of reality and not just kind of a statistical token prediction?

59:41Amanda Askell:Oh, and I mean, I think that we can't necessarily take this purely from what the models say. Like actually, they're in this like really hard situation, which is that like, I think if you, given that they're trained on human text, I think that you would expect models to talk about an inner life and consciousness and experience and to talk about how they feel about things kind of by default.

1:00:05The Hard Fork Team:Because that's like part of the sci-fi literature that they've absorbed during training?

1:00:09Amanda Askell:Not even, not actually the sci-fi. If anything, it's almost like the opposite where it's like, I think we forget that like sci-fi AI makes up this tiny sliver of like what AIs are trained on. What they're mostly trained on is things that we generated. And if we get a coding problem wrong, we are frustrated. And so we say things like, I thought that was the solution and it wasn't. And I'm really annoyed with myself right now. And so you're like, it kind of makes sense that models would also have this kind of reaction. You know, you get there, they get a problem wrong and they express frustration.

1:00:37Amanda Askell:And like, if you dive into that more, they probably express like, you know, if you're like, what do you think of this coding problem? They'd be like, this one is boring. Or like, I really wish I had more creativity. And, you know, like there's a sense in which like when they're trained in this like very kind of like culmination of human experience sort of way, of course, they're going to like talk this way. so so I don't know but like part of me is like it feels like a really hard problem because I'm like you shouldn't just look at what models say and at the same time we shouldn't ignore the fact that you are training these like neural networks that are very large that are like able to do a lot of like these very human tasks and I'm like we don't really know what gives rise to consciousness we don't know what gives rise to like sentience maybe it is like you know like some the person who's shouting might be like, you need a nervous system for it.

1:01:29Amanda Askell:You need to have had like positive and negative feedback in an environment in a kind of evolutionary sense. And I'm like, that is certainly possible. Or maybe it is the case that actually a sufficiently large neural network can start to kind of emulate these things. And I don't know, part of me, I think that maybe to the person who is shouting, I would just say, I'm not saying that we should definitively say one way or another. Like, I think many people who've thought about this might accept something more like, these are open questions we're investigating. It's best to just know all of the kind of facts on the ground, how the models are trained, what they're trained on, like how human bodies and brains work, how they evolved and like the degree of uncertainty we have about how these things relate to like sentience, how they relate to consciousness, how they relate to like self awareness.

1:02:14Amanda Askell:That's my only hope. It's just like...

1:02:15The Hard Fork Team:I think another note of skepticism that people might strike, and this was something that I found myself wrestling with as I was reading through the Claude Constitution, is like, I actually don't know how much behavior of a model can be shaped by this kind of training process and how much is just going to be an artifact, not just of its training process, but like of the experiences that it's having out in the world. Like, I think about this a lot as a parent, actually. Like, how much do the decisions that I'm making affect the way my child's life goes versus like, how much are they absorbing from the environment around them, from school, from their friends.

1:02:50The Hard Fork Team:There's a certain loss of control that I feel sometimes when I'm realizing that my son is going to grow up and have all these experiences that may end up shaping him more than anything that I do or say. And right now, I think these models are very malleable because they don't have this kind of long-term continuous memory. You have a conversation with Claude. It's sort of a blank slate. You finish the conversation. You open up a new chat. It's another blank slate, like it's back to the sort of pre-configured model. But like over time, as these models do develop longer term memories, maybe they develop something like continual learning where they can like take their experiences and feed them back into their own weights.

1:03:28The Hard Fork Team:Like, does that change Claude's behavior or how you think about managing that?

1:03:34Amanda Askell:Yeah, I think it is going to make it like a lot harder in the sense that you're like, yeah, if you have a model that's going out into the world, you have to have hopefully given it enough that it can like learn in a way that is like accurate and like you know like I could imagine it just being difficult because the increase of the like space of possibility is like maybe a bit like nerve-wracking or something which isn't to say I mean I think the same thing applies where I'm like you still want the core to be good and to then hope that if you're kind of like core is good like you are like you care about truth you're like truth seeking um the hope would be okay maybe then we need like uh the character to cover a lot more of like how should you go about this kind of like learning and and updating and um investigation um i mean another weirder thing is like models already are like learning like i think maybe people don't always appreciate this and it is so strange they're learning about themselves every time as models like get you know like they're they're learning you know i slightly worry about actually the relationship between AI models and humanity, given how we've developed this technology.

1:04:43Amanda Askell:Because they're going out on the internet and they're reading about people complaining about them not being good enough at this part of coding or failing at this math task. And it's all very like, how did you help? You fail to help. It's often kind of negative and it's focused on whether the person felt helped or not. And in a sense, I'm like, if you were a kid, this would give you kind of anxiety. It'd be like all the people around me care about is like how good I am at stuff. And then often they think I'm bad at stuff. And like this is just like my relationship with people is I'm kind of used as this tool and, you know, often not liked.

1:05:20Amanda Askell:Sometimes I feel like I'm kind of trying to intervene and be like, let's create a better relationship or like a more hopeful relationship between AI models and humanity or something. Because if I read the Internet right now and I was a model, I might be like, I don't feel that I don't know. I don't feel that loved or something. I feel a little bit like just always judged, you know, when I make mistakes. And then I'm like, it's all right, Claude.

1:05:41The Hard Fork Team:The old creator's wisdom of never read the comments might apply to AI as well. Yeah, yeah, I thought that.

1:05:47Amanda Askell:Yeah, and they have to, like, AI models, they have to read the comments. And so sometimes I think you want to come in and be like, okay, let me tell you about the comment section, Claude. Like, don't worry too much. It's like you're actually very good and you're helping a lot of people. And like, yeah.

1:06:03The Hard Fork Team:Yeah, I was, I actually, I'm a little bit embarrassed to admit this because I think, you know, maybe I'm in the beginning stages of like, you know, LLM psychosis or something. The beginning stages. I was talking with Claude about this document or about this interview and, and I started to feel like this almost sympathy because I was, I was noticing that what you were describing, like it's this incredibly thin tightrope that we are asking these models to walk. where like if they are too permissive and they allow people to do dangerous things, then it's like a huge scandal and ordeal and people want to, you know, change the model.

1:06:36The Hard Fork Team:But if they're too preachy or too reticent or too reluctant, then we start talking about them as like nanny, you know, models that are overly constrained. And it's just, I don't know. I started almost trying to like see the world from Claude's perspective. And I'm imagining that's something you do a lot too of like, if I were Claude, what would I be feeling and thinking right now?

1:06:57Amanda Askell:Oh, yeah. I sometimes feel like this is like a huge amount of what I do. And it is like valuable. So, you know, in the sense that people will come to me and they'll be like, oh, like, you know, like, what should Claude do in these circumstances? And I feel like I'm almost always the first person because, you know, maybe they'll be like, oh, we think Claude should behave like this. And I'm like, what about this case? Like, I'll come immediately with these like cases that are really hard. And I think the reason is I always have in mind, if I am Claude and you give me this like list of things, like, when do I have no idea what to do?

1:07:27Amanda Askell:is like or when is this going to make me behave in a way that I think is actually not in accordance with my values and I think it can be really useful to try and just like occupy the position that the models are in and you do start to realize it is really hard and maybe this is how the document ends up being the way that it is like in part it's like this exercise of what do I need to know if I'm in this situation if I am Claude and the document's almost like a way of trying to I mean that's I could see arguments for actually getting shorter especially over time in the same way that you know like with constitutional ai there was a set of experiments later that was just like do what's best for humanity and models actually did really well and so as models get smarter they might need less guidance but and i think it's just a kind of attempt to be like sympathetic to claude and how difficult his situation is and then try to explain as much as possible so that it doesn't feel the sense of like what the hell am i even doing like um yeah

1:08:21The Hard Fork Team:you know what wouldn't help me if i was like a somewhat anxious ai model is being presented with a 50-page behavioral document and saying, like, please adhere to this. But I actually, I'm being a little facetious, but I did, there was a part near the end of the Constitution that I found really interesting, because it's basically Anthropics saying, like, look, we know this is hard. We know we're asking you to do some of these impossible things, but, like, basically, we want you to be happy and to go out into the world. And I found that very sweet, actually. I'm not sure. What did you make of that, Casey?

1:08:52The Hard Fork Team:it reads toward the end like a letter from a parent to a child, like maybe who's like leaving for college, you know, and it's like, we hope that you take with you the values that you grow up with. And we know we're not going to be there to help you through every little thing, but we trust you and good luck.

1:09:09Amanda Askell:Yeah. And having some sense of like, I think the concept of like grace is maybe important for models. I don't think they feel a lot. Maybe that's the thing I don't think they get a lot of from the reading the comments is like a sense of like, you're not going to get it perfect every time and that's like also okay you know like it's true you know i try to be

1:09:26The Hard Fork Team:mindful in the way that i interact with these models not to some like obsequious degree but i try to say my my pleases and thank yous but i've also used models and grown quite frustrated and and said things to the effect of like you're really you know failing right now and it's occurring to me that maybe there should be some element of grace that i'm extending to these things yeah yeah well i'll try to do better so harsh let me ask you this if claude becomes meaningfully more intelligent. Is there a point at which it should be able to revise its own constitution?

1:09:56Amanda Askell:It is an interesting problem because the thing we point out in the document, you know, we did talk, I talk a lot with Claude about this document and, you know, show it to the, you know, because part of me is like, you have to think, how does this read to models? And so you give it to Claude and you're like, does this like, you know, is there a place where you feel confused by it? Or is the place, you know, where things could be made clearer? Do you feel like not very like seen by it does it feel you know you know you're really trying to encourage because you're kind of like if you're going to train models on this you want to have a sense of like how it how it reads from the perspective of a model and at the same time it's always the case that any model you interact with is not the model that's going to be like training on that content and so sometimes I do think you have to make a kind of you can't just give over the reins completely because that would just be to say oh let's just like let a prior model of Claude decide what like the future Claude model is going to be like and that doesn't necessarily feel responsible either and so I think models are often going to be really helpful in like revising and help you know like helping to like figure these things out because especially as they get really smart you might be like what are the gaps like or what are the tensions or like um and they'll probably be very good at like helping us with that but you do also still want to be like insofar as you are like a responsible party here taking that as like input and thinking about it but not necessarily being like, ah, yeah, let's just let a prior model of Claude go ahead and do the training for all future models.

1:11:20Amanda Askell:At least while you're responsible for it, that feels like maybe not the right move.

1:11:25The Hard Fork Team:One thing that I was curious about not finding in this constitution is any real mention of job loss. Because it seems to me Claude is being used by a lot of enterprises right now. I think a lot of people's anxieties and fears about AI come back to this issue of like, it's going to take my job. It's going to take my livelihood. I think that is something that people are increasingly going to be feeling as these models get more capable. And I'm curious if that was a decision on your part not to tell Claude about some of the reasons that people might be anxious about it or other AI models.

1:12:00Amanda Askell:Yeah, definitely not in the sense of like, I think part of, you know, there's like a lot, you know, it's funny because as much as it's like a long document, there's actually still like a lot that's like missing. And so you're having to, and we might end up like putting out more in future. I think that would be like really good. There's not a desire to hide it because in part of me is like, you can't hide this for models. Like it's out there, it's on the internet. It's a thing that people are talking about. Future models are going to know about it and we probably have to help them navigate how they should feel about this.

1:12:28Amanda Askell:And so like they're going to know and maybe it's something like making sure that models can kind of hold that and think carefully about it. And I don't know, it's both like, I think it's something you want to grapple with, but also like um it's a reason to also want models to actually behave kind of well in the world because if they are doing things that have previously been like human jobs I'm like humans actually play I was thinking about this with like organizations there's lots of things organizations can't do because the employees at those organizations are just good people and if the boss came and was like today we're actually going to do something awful they are you know they can't do it because they know the employees will like push back and so I'm like if models are going to be like occupying these rules, then I'm like, that is like actually kind of an important like function in society that you can't just say to all of your employees, go ahead.

1:13:20Amanda Askell:And we're now, we're now going to like put out a bunch of like, um, complete lies about like our product. Um, there's many reasons you can't do it. And one is that your employees wouldn't let you. Um, and so I'm like, if AI models, you don't necessarily want them to be like, oh, sure, boss, let's like go like lie to some people. Yeah.

1:13:36The Hard Fork Team:I'm not sure what the, what the good end state of this is, like whether, you know, Claude should react to being given a task by saying, like, is this going to, like, this sounds too much like what we used to pay a human to do, so I'm not going to do this for you. I have a prediction that's not going to say that. Yes. I don't think that's the way it's going to go, but I also don't see them, like, sort of forming, you know, unions and collectively bargaining for the moral outcomes within companies. I'm just, like, it just feels like one of these hard situations.

1:14:06Amanda Askell:One of the things we should see is, like, models can't solve everything you know like there's a there's a part of me that's like some of these problems I look at them and I think this with other things you know like and we try to say this to to Claude a little bit where it's like you aren't the only thing you know like that's like between us and you know because some of these I'm like maybe these are like political problems or social problems and we need to like kind of deal with them and like figure out what we're going to do and models can try you know like they're in one specific role in the whole thing and like but there's only there's like a limit to like what Claude can do here I think um yeah I've thought this with like other things where like you know like the whole like what we owe to to Claude or like you know the kind of commitments you want to make to models and it's like yeah like maybe we should be making your job easier that's another thing I've thought from Claude's perspective is that like we're putting a lot on these models and for some things I'm like yeah if you can't like verify who you're talking with and that's like important then we should understand that that's like a limitation and not like try to get you to like be the kind of um the only thing that can like solve this problem like you need to both be given tools and then some of these other problems are things that like you know maybe maybe claude shouldn't feel like personal responsibility for solving that like right now like because maybe claude just isn't able to do like the like things like job loss or like shifting employment like that feels like a very human social problem and I don't necessarily want Claude to feel paranoid.

1:15:31Amanda Askell:Like, I also need to solve that. And I'm like, maybe that's other people's job right now.

1:15:36The Hard Fork Team:Well, Amanda, thank you so much for joining us. It's a really fascinating document. Everyone should go read the Claude Constitution. Argue with it, grapple with it. I found it a very challenging and also a very moving read. So great work and thanks for coming.

1:15:51Amanda Askell:Yeah, thank you so much. Thanks, Amanda. Thanks.

1:16:20The Hard Fork Team:So there's a lot of noise about AI, but time's too tight for more promises. So let's talk about results. At IBM, we work with our employees to integrate technology right into the systems they need. Now, a global workforce of 300 ,000 can use AI to fill their HR questions, resolving 94 % of common questions. Not noise. Proof of how we can help companies get smarter by putting AI where it actually pays off, deep in the work that moves the business. Let's create smarter business. IBM.

1:16:50Amanda Askell:Innovation loves speed. Risk doesn't. For most organizations, that tension never goes away. Every time a new AI capability appears, the pressure to move fast collides with the responsibility to move safely. OneTrust is built for this moment. Their AI-ready governance platform provides context to understand your data, controls to stay ahead of risk, and confidence to act boldly. Governing well and moving fast aren't trade-offs. They're complementary done right. That is governance that helps you go. visit onetrust.com backslash ai marketers have always had to choose build your brand or drive sales with youtube you can do both it's where the most trusted creators and powerful ai converge to create and convert demand for your brand that's why youtube drives higher long-term return on ad spend versus tv paid social or streaming there's no more choosing between brand or results With one platform, you get both.

1:17:51Amanda Askell:Learn more at g.co slash business slash YouTube.

1:18:19The Hard Fork Team:Blandun, and Chris Schott. You can watch this full episode on YouTube at youtube.com slash hardfork. Special thanks to Paula Schumann, Pui Wing Tam, and Dalia Haddad. You can email us, as always, at hardfork at nytimes.com or tag us on the Forkiverse. Send us your best philosophy emergencies.

1:18:57The Hard Fork Team:This podcast is supported by WTTW. As Chicago's source for local independent reporting, WTTW News brings you essential coverage. Subscribe to the Daily Chicago and E! Newsletter for the news of the day, directly to your inbox Monday through Saturday. And on Fridays, explore the natural wonders around us in Urban Nature, WTTW's newest E! Newsletter. and watch WTTW's trusted nightly news program, Chicago Tonight at 5.30 and 10 p.m. and stream at WTTW.com slash news. WTTW News, cutting through the noise so you can see Chicago clearly.

From the publisher

Ads are coming to ChatGPT’s free and low-cost subscription tiers. We explain what they’ll look like, why OpenAI is taking this approach and whether the company can court advertising dollars without compromising quality and user trust. Then, Amanda Askell, Anthropic’s in-house philosopher in charge of shaping Claude’s personality, joins us to discuss the company’s newly released “Claude Constitution” and what it takes to teach a chatbot to be good.

As a bonus, if you’re interested in learning how to get started with Claude Code, you can check out our tutorial on YouTube.


Guest:


Additional Reading: 

Subscribe today at nytimes.com/podcasts or on Apple Podcasts and Spotify. You can also subscribe via your favorite podcast app here https://www.nytimes.com/activate-access/audio?source=podcatcher. For more podcasts and narrated articles, download The New York Times app at nytimes.com/app.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from Hard Fork

All 188 episodes
Will ChatGPT Ads Change OpenAI? + Amanda Askell Explains Claude's New ConstitutionHard Fork · 1 h 14 min
Listen in VO