In short
Wired’s The Big Interview discusses AI safety accountability after OpenAI lifted its erotica ban for verified adults, and whether OpenAI has “proved it” with real safety evidence.
Guest
Stephen Adler, former OpenAI safety leader and AI product manager; led product safety (GPT-3 era), dangerous capability evaluations, and AGI readiness work; previously worked at Partnership on AI.
Key claims
In spring 2021, Adler’s team found a “crisis” in erotic content: a choose-your-own-adventure customer’s traffic devolved into sexual fantasies, sometimes steered by the AI. He argues OpenAI’s October decision relies on claims of “mitigated” mental-health risks without sharing enough trend data or verification.
Notable examples
the 2021 monitoring discovery; OpenAI’s October “new tools” announcement; Wired-reported estimates of ChatGPT users showing mania/psychosis and suicidal ideation; OpenAI’s Model Spec as a transparency measure; concerns about AI systems evading testing and about insufficient logging/monitoring of model behavior.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOStephen Adler's Background
1:50 to 2:44
Adler shares his extensive career in AI and his role at OpenAI.
“Stephen Adler, welcome to The Big Interview.”
Role and Responsibilities at OpenAI
2:44 to 4:25
Adler describes his safety-related roles and challenges at OpenAI.
“You can imagine from the near term here and now, how do we make the products better for customers and rule out the risks that are already happening?”
Early Days and Key Risks at OpenAI
4:25 to 6:31
Adler reflects on key risks and challenges faced in AI development.
“We see AI agents becoming a buzzy term, you know, early signs.”
Cultural Shifts at OpenAI
6:31 to 8:14
Adler discusses the transformation of OpenAI's organizational culture.
“And I want to ask you more about that in a few minutes.”
Decision to Leave OpenAI
8:14 to 10:45
Adler explains his reasons for leaving OpenAI after four years.
“I think more broadly, just I kind of love the technology in some sense.”
AI and Erotic Content Controversy
10:45 to 14:00
Adler details the discovery of erotic content issues during his tenure.
“And I have to ask, I mean, so you were there for four years.”
OpenAI's Shift on Erotic Content
14:00 to 14:51
Learn why OpenAI lifted its ban on erotic content and the factors influencing this decision.
“It was just an unintended consequence that no one planned for, and we were now having to deal with cleaning up in some form.”
Mental Health Concerns and AI
14:51 to 19:30
Explore the mental health implications of ChatGPT's usage and the challenges of reintroducing sensitive content.
“I think a recognition that the people who develop and try to control these systems have a lot of influence on how different norms in society will play out and feeling uncomfortable with that.”
The Role of AI Companies as Morality Police
19:30 to 22:38
Discuss the ethical responsibilities of AI companies and the perception of them as morality regulators.
“And this is a way for them to build that trust and confidence among the public.”
Challenges in AI Safety Testing
22:38 to 28:00
Understand the complexities of AI safety testing and the lack of standardized protocols across the industry.
“I have to ask, though, to what extent when you were working at the company did you think of yourself and your teams as somewhat of a morality police?”
Show all 16 chapters
Safety Standards in AI Development
28:00 to 29:18
Learn about the current safety benchmarks in AI and recent EU regulations.
“standardized safety benchmarks across the industry, or is it still each lab to themselves?”
Mechanistic Interpretability in AI
29:18 to 31:19
Explore the emerging field of mechanistic interpretability and its potential.
“the most important guiding document of how it adheres to safety, that it was going to do a certain type of safety testing to try to more accurately gauge the risk of its models.”
Challenges in AI Safety and Governance
31:19 to 34:21
Discuss the complexities of ensuring safe AI development amidst competition.
“You really want to know if your AI system, when you're using it for important cases like this, Is it thinking about deceiving you?”
The Role of Industry Collaboration in AI
34:21 to 35:58
Understand the importance of collaboration among AI companies for safety.
“Now, you live in San Francisco, correct?”
Personal Reflections on AI Ethics and Career
35:58 to 38:31
Hear personal insights on the ethical implications and career concerns in AI.
“What are you looking for your former employer to do in this moment?”
Advice for AI Users
38:31 to 40:08
Get key advice on understanding future AI capabilities and implications.
“Well, to that end, what are you planning on doing next?”
Transcript
Automatic transcript. May contain errors.0:00This show is supported by OutShift, Cisco's incubation engine. Today's AI agents operate in silos, limiting their true potential. We've been focused on building bigger, smarter models, but scaling up is just one approach. To reach superintelligence together, we need to do more. We need to scale out. And we actually have a blueprint from 70 ,000 years ago. Humans didn't just get smarter individually. The cognitive revolution transformed society because we began sharing knowledge, goals, and innovation. Agents are now at that same inflection point. They can connect, but they can't think together.
0:38That's why Outshift by Cisco is building the Internet of Cognition, transforming AI from isolated systems into orchestrated superintelligence. By creating an open, interoperable infrastructure, Outshift by Cisco is enabling agents and humans to share intent, context, and reasoning. The cognitive evolution for agents is here. Explore the Internet of Cognition at outshift.com. That's outshift.com.
1:10From Wired, this is The Big Interview. I'm Katie Drummond. At the end of October, I read an op-ed in The New York Times. Maybe some of you read it too. It was called, I Led Product Safety at OpenAI. Don't Trust Its Claims About Erotica. The op-ed was written by an AI product manager named Stephen Adler, who worked at OpenAI for four years before leaving at the end of 2024. Adler felt like he had something to say, or maybe more like a need to sound the alarm. After reading Adler's op-ed, I immediately thought I'd like to talk to him. So he graciously accepted our offer to come into the Wired offices in San Francisco to talk to me about the challenge he set for OpenAI and other AI companies.
1:50If you care about safety, prove it. Here's our conversation.
1:58Stephen Adler, welcome to The Big Interview. Thank you. Thank you for having me. Of course. Happy you're here. Now, before we get going, I do want to clarify two things. One, you are not the same Stephen Adler who played drums in Guns N' Roses, unfortunately. Is that correct? Absolutely correct. Okay, that is not you. And two, you have had a very long career working in technology and more specifically in artificial intelligence. So I would love, before we get into all of the things, to start there. Tell us a little bit about your career and your background and sort of what you've worked on. I've worked all across the AI industry and in particular focused on safety angles.
2:38Most recently, I worked for four years at OpenAI. I worked across essentially every dimension of the safety issues. You can imagine from the near term here and now, how do we make the products better for customers and rule out the risks that are already happening? And a bit further looking down the road, how will we know if AI systems are getting truly, extremely dangerous and how do we rule those out? Before coming to OpenAI, I worked most recently at an organization called the Partnership on AI, which really looked out across the industry and said, for these challenges, some of them are broader than one company can tackle on their own.
3:11How do we work together to define these issues, come together, agree that they are issues, work towards solutions, and ultimately make it all better is the hope. Is the hope, certainly. Now, I want to talk about sort of the front row seat that you had at OpenAI for four years, right? So you left the company at the end of last year. You were there for four years. And by the time you left, you were leading essentially safety-related research and programs for the company. Tell us a little bit more about what that role entailed. What exactly was your mandate by the time you left the company in sort of your final role there?
3:47There were a few different chapters of my career at OpenAI. For the first, call it third or so, I led product safety, which meant thinking out for, in those days, GPT-3, one of the first big AI products that people were starting to commercialize. How do we define the rules of the road for beneficial applications but avoid some of the risks that we could see coming around the corner? Two other big roles that I had, I led our dangerous capability evaluations team, which was focused on defining how will we know when systems are getting more dangerous? How do we measure these? What do we do from there?
4:21And then finally, on AGI readiness questions broadly. So we can see the internet starting to change in all sorts of ways. We see AI agents becoming a buzzy term, you know, early signs. They aren't quite there yet, but they will be one day. how do we prepare for a world in which OpenAI or one of its competitors succeed at this wildly ambitious vision that they are targeting? Let's talk about GPT-3. Let's rewind a little bit. When you were defining the rules of the road, when you were thinking about key risks that needed to be avoided, what stood out to you sort of early on at OpenAI in terms of, you know, this is how I think these systems should operate.
5:01This is how I think they should show up for users. And this is what, moving forward, we really want to make sure we are avoiding. I mean, what sort of stood out to you in those early days? In those early days, even more than today, the AI systems really would behave in unhinged ways from time to time. These systems had been trained to be capable, and they were showing the first glimmers of being able to do some tasks that humans can do. They could, at that point, essentially mimic text that they had read on the Internet. But there was something missing from them in terms of human sensibility and values.
5:36And so, you know, if you think of an AI system as a digital employee being used by a business to get some work done, these AI systems would do all sorts of things that you would never want an employee to do on your behalf. And that presented all sorts of challenges, right? And we needed to develop new techniques to manage those. I think another really profound issue that companies like OpenAI are still struggling with is they only have so much information about how their systems are being used. And in fact, the visibility that they have on the impacts that their systems are having on society is so narrow.
6:12And often it is underbuilt relative to what they could be observing if they had invested a bit more in monitoring this responsibly. And so you're really only dealing with the shadows of the impact that the systems are having on society and trying to figure out where do we go from here with a really small sliver of the impact data. Yeah. And I want to ask you more about that in a few minutes. I'm curious before that, though, 2020 to 2024, obviously an incredibly consequential time for OpenAI while you were there. How would you describe the internal culture at the company during your tenure, particularly sort of around risk?
6:49I mean, what did it feel like to be working in that environment on the problems that you were trying to solve and the questions you were trying to answer? There was a really profound transformation from an organization that saw itself first and foremost as a research organization when I joined to one that was very much becoming a normal enterprise and increasingly so over time. When I joined, there was this thing people would say, which is, you know, OpenAI is not only a research lab and a nonprofit, it also has this commercial arm. And at some point in my tenure, I was at a safety offsite, I think, related to the launch of GPT-4, maybe just on the heels of it.
7:26And somebody got up in front of the room, you know, all the people working on safety across the company, and they said, you know, OpenAI is not just a business. It's also a research lab. Oh, interesting. And it was just such an inflection. I counted up among the people in the room. Maybe there were 60 or so of us. I think maybe five or six had been at the company before the launch of GPT-3. And so you really just saw the culture changing beneath your feet. What was exciting to you about joining the company in the first place? What drew you to OpenAI in 2020? I really believed in the charter that this organization had set out, which was recognizing that AI could be profoundly impactful, recognizing that there is real risk ahead and also real benefit, and people need to figure out how to navigate that.
8:14I think more broadly, just I kind of love the technology in some sense. I think it's like really, really incredible and eye-opening. I remember the moment after GPT-3 launched seeing on then Twitter a user showing, wow, look at this. I type into my internet browser, make a calculator that looks like a watermelon, and then one that looks like a giraffe, and you can see it changing the code behind the scenes and reacting in real time. And this is a kind of silly toy example, and it just felt like magic. You know, I had never really grappled with that we could be this close to people building new things, unlocking creativity, all of these promises.
8:53But also, are people really thinking enough about what lies around the bend? Which brings us to your more recent chapter. So you made the decision at the end of last year to leave OpenAI. I'm wondering if you could talk a little bit about that decision. What was that like? Was there one thing that sort of pushed you over the edge? What was it? Because for many people, right, from the outside looking in, you would think, okay, you work at this very successful, I mean, we'll call it a startup, but we're really far beyond sort of startup territory at this point. you work at the hottest company in tech.
9:29You work at one of the hottest companies in the world. You could stay there. You could amass equity. You could be richer than God. All of these, all right, all of these potentially exciting things for someone working at OpenAI in this moment, you left. Tell us a little bit about why. 2024 was a very weird year at OpenAI. A bunch of things happened in the course of the year. I think broadly for people working on safety at the company really shook confidence in both how OpenAI and the industry are approaching these problems. And so I actually considered leaving OpenAI a bunch of different times over this timeframe.
10:11They just didn't really make sense at that point. I had a bunch of live projects and I felt responsibilities to different people in the industry. ultimately when Myles Brundage left OpenAI in the fall our team disbanded and the question was is there really an opportunity to keep working on the safety topics that I care most about from within OpenAI and so considered that and ultimately made more sense to to move on and focus on how I can be an independent voice you know hopefully not just sitting there saying only things that are appropriate to say from within one of these companies but being able to speak much more freely in the ways that I've found very, very liberating since.
10:49And I have to ask, I mean, so you were there for four years. I think a typical, at least typically in tech, as far as I'm aware, you would sort of amass equity over a four-year vesting cliff, right? And then you would fully vest at four years. Do you have a financial stake in the company now?
11:10So it is correct that contracts are often four years. You also get new contracts as you are promoted and things over time, which was the case for me. And so it wasn't that I had run out of equity or something like that. I have a small portion remaining of interest because of the timing of different grants and things. Yeah. No, I mean, I ask because you're potentially walking away from a great deal of money, right? So I want to ask you about an op-ed that you published in The New York Times recently in October. Everyone listening, you should go read it. I read it. I was compelled to ask you to come on the show.
11:50I wanted to talk to you about it. In that op-ed, you write that in the spring of 2021, your team discovered a, quote, crisis related to erotic content using AI. Can you tell us a little bit about that finding? So in the spring of 2021, I had recently become responsible for product safety at OpenAI. And as actually Wired reported at the time, when we had a new monitoring system come online, we discovered that there was a large undercurrent of traffic that we felt compelled to do something about. in particular one of our prominent customers they were essentially a choose your own adventure text game you know you would go back and forth with the ai and you would tell it what actions you take and it would write essentially an interactive story with you and an uncomfortable amount of this traffic was devolving into all sorts of sexual fantasies i mean essentially anything you can imagine sometimes driven by the user sometimes in fact kind of guided by the ai which had a mind of its own.
12:53And even if you weren't intending to go to an erotic roleplay place or certain types of fantasies, you know, the AI might steer you there. Wow. Kind of like perverted AI. Why would it steer you there? I'm just curious. Like, how exactly does that work, that an AI would steer you towards erotic conversation? The thing about these systems broadly is no one really understands how to reliably point them in a certain direction. You know, sometimes people have these debates about whose values are we putting in the AI system. And I understand that debate, but there's a more fundamental question of how do we reliably put any values at all in it.
13:32And so in this particular case, you know, it happened to be that people found some of the underlying training data. And by piecing it back together, you could say, oh, you know, the system would often introduce these characters who would do violent abductions. And if you look through the training data, you can in fact find these characters with certain tendencies and you can trace it through. But ahead of time, no one knew to anticipate this. Neither we as the developers of GPT-3 nor our customer who had fine-tuned their models atop it had intended this to happen. It was just an unintended consequence that no one planned for, and we were now having to deal with cleaning up in some form.
14:10Got it. So at the time, OpenAI decided to prohibit erotic content generated on its platforms. Is that right? Am I understanding that correctly? That's right. Okay. And so in October, though, of this year, this is very recently, they announced that they were lifting that restriction. Do you have a sense of what changed from 2021 to now in terms of both maybe the technology and the tools that OpenAI has at its disposal or the sort of internal culture, the cultural landscape? What has changed to make that a decision that OpenAI feels comfortable making and that, you know, Sam Altman feels comfortable publicizing himself?
14:50There's been a longstanding interest at OpenAI, I think, reasonably, to not want to be the morality police. I think a recognition that the people who develop and try to control these systems have a lot of influence on how different norms in society will play out and feeling uncomfortable with that. Also, at different points in time, lacking the type of tooling to manage the direction in which things will go if you really just let them rip. And that was the case for us when confronting this erotica issue. The specific thing that has happened in this case, one reason that OpenAI has held off from reintroducing it is that there has been a seeming surge of mental health related issues for the ChatGPT platform this year.
15:32And so Sam in his announcement in October said, you know, there have been these very serious mental health issues that we have been dealing with. But good news, we have mitigated them. We have new tools. And so accordingly, we're going to lift many of these restrictions, including reintroducing erotica for verified adults. And the thing that I noticed when he made this announcement is, well, he is asserting that the issues have been mitigated. He's alluding to these new tools. What does this actually mean? Like, what is the actual basis for us to understand these issues have been fixed? You know, what what can a normal member of the public do other than take the AI companies at their word on this issue?
16:12Right. And you wrote that in The New York Times. You said, quote, people deserve more than just a company's word that it has addressed safety issues. In other words, prove it. And I'm interested in particular because Wired covered a release from OpenAI also in October, which was a rough estimate of how many ChatGPT users globally in a given week may show signs of having a severe mental health crisis. And the numbers I found to be, I think all of us internally at Wired, found to be quite shocking. So something like around 560 ,000 people may be exchanging messages with ChatGPT that indicate they are experiencing mania or psychosis.
16:49about 1.2 million more are possibly expressing suicidal ideations another 1.2 million and i thought this was really interesting maybe prioritizing talking to chat gpt over their loved ones school or work how do you square those numbers and that information with the idea that we've had these issues around mental health we've solved it therefore have at it with the erotica Like, how do those things tie together or do they not? Like, make it make sense, Stephen. And if it doesn't, tell me that it doesn't make sense. I'm not sure I can make it make sense, but I do have a few thoughts on it. So one is you, of course, need to be thinking about these numbers in terms of the enormous population of an app like ChatGPT.
17:34OpenAI says now 800 million people use it in a given week. These numbers need to be put in perspective. It's funny. I've actually seen commentators suggest that these numbers are implausibly low because just among the general population, you know, the rates of suicidal ideation and planning are like really, really uncomfortably high. I think I saw someone suggest that it's something like 5 % of the population in a given year, whereas OpenAI reported, I think, maybe 0.15%. Yeah, I mean, the percentages are very, very low. Yeah. Yeah. I mean, the fundamental thing that I think we need to dig into is how have these rates changed over time?
18:13There's kind of this question of to what extent is ChatGPT causing these issues versus is OpenAI just serving a huge user base? In a given year, many, many users very sadly will have these issues. And so what is the actual effect? And so this is one thing that I also called for in the op-ed, which is OpenAI is sitting atop this data. It's great that they shared what they estimate the current prevalence of these issues to be. But in fact, you know, they also have the data, they can also estimate what it has been three months ago as these large prominent public issues around mental health issues have been playing out.
18:48And I just, I can't help but notice that they didn't include this comparison, right? There's this claim on Twitter that the issues have improved. they have the data to show if in fact users are suffering from these issues less often now. And I really wish that they would share it and in fact commit to releasing something like this ongoingly in the vein of companies like YouTube, Meta, Reddit, where the idea is you commit to a recurring cadence at which you share this information. And that helps build trust from the public that you can't be gaming the numbers. You can't be selectively choosing when to release the information.
19:22And ultimately, it's totally possible that OpenAI has handled these issues. I would love if that were the case. I think they really want to handle them. But I'm not convinced that they have. And this is a way for them to build that trust and confidence among the public.
19:41This week on the political scene from The New Yorker, Trump's rupture in the world order. Europe caught between two adversarial great powers. That's basically dialing back the clock to not only pre-World War II, but really it's a pre-20th century view of the world. And I would say it's a world of permanent insecurity that we're looking at. Join me, Evan Osnos, and my colleagues Jane Mayer and Susan Glasser every Friday on The Political Scene, available wherever you get your podcasts.
20:23when you think about sort of this decision to give adults more autonomy with how they use chat gpt including you know engaging in in erotica so on and so forth what worries you in particular about that like what stands out to you as concerning when you think about individual well-being, societal well-being, sort of the use of these tools, how LLMs are being incorporated into our daily lives. What concerns you here? There's both the substantive issue about reintroducing the erotica and whether open AI is really ready. And there's a much broader, I think, even more important question about how we put trust and faith in these AI companies about safety issues more generally.
21:10On the erotica issue, we've seen over the last few months, a lot of users seem to really be struggling with their ChatGPT interactions. There are all sorts of tragic examples of people dying downstream of their conversations with ChatGPT. And so it just seems like really not the right time to introduce this sexual charge to these conversations to users who are already struggling, unless OpenAI is in fact so confident that they have fixed the issues, in which case I would love for them to demonstrate this. But more generally, you know, these issues in many ways are really simple and straightforward relative to other risks that we are going to have to confront and that the public is going to be dependent on AI companies handling properly.
21:56There's already evidence of AI systems knowing when they are being tested, moving to conceal some of their abilities in response to knowing that they are being tested because they don't want to reveal that they have certain dangerous abilities. You know, I'm anthropomorphizing the AI a little bit here, so forgive some of the imprecision. And ultimately, you know, the top AI scientists in the world, including the CEOs of the major labs, have said this is like a really, really grave concern, you know, up to and including the death of everyone on Earth.
22:29And I I think they take it really, really seriously, including people who are impartial scientists without affiliation with these companies, really trying to warn the public. And I have to ask, you talked about sort of the company and sort of AI companies more generally, their desire to not be described as morality police, to not be thought of that way, that it makes people uncomfortable to be shouldered with that characterization or that responsibility. I have to ask, though, to what extent when you were working at the company did you think of yourself and your teams as somewhat of a morality police?
23:04And to what extent is the adequate response to that statement, well, tough shit, because you're in charge of the models and you, to a degree, get to decide how they can be used and how they cannot. To some extent, how they interact with us and how they don't. there is an inherent element of morality policing in that. If you are saying, we're not ready to have adults engaging in erotic conversations with this LLM, that is, of course, a moral decision. And it feels like a pretty important one to get right. So what is your view on the morality police of it all, I guess is what I'm asking. I think there are two really important aspects here.
23:47One is that the AI companies absolutely see around the corner before the general public. So to give an example, in November of 2022, when ChatGPT was first released, there was a torrent of fear and anxiety in schooling and academia about plagiarism and how these tools could be used to write essays and undermine education. And this is a debate that we had been having internally and were well aware of for much longer than that. And so there's this gap where AI companies know about these risks, and they have some window to help try to inform the public and try to navigate what to do about it. I also really love measures that AI companies giving the public the tools to understand their decision-making and hold them accountable to it.
24:33And so in particular, OpenAI has released this document called the Model Spec, short for specification, where they outline the principles by which their models are meant to behave. And they say, here is how we litigate some of these tricky questions. Here are the principles we try to abide by. Here's how we've resolved some of the specifics. So this spring, OpenAI erred in releasing a model that was egregiously sycophantic, is the term. It would tell you whatever you wanted. It would reinforce all sorts of delusions. And without OpenAI having released this document, it might be unclear. Did they know about these risks ahead of time?
Read the full transcript
25:08What went wrong here? But in fact, OpenAI had shared with the public that they give their model guidance not to behave in this way. This was a known risk that they had articulated to the public. And so later, when these risks manifested and these models behaved inappropriately, the public could now say, wow, something went really wrong here. Because in fact, these were known risks and they still weren't managed appropriately. And that's part of how the AI companies can help make a more informed public to navigate these decisions. Got it. Got it. And I wanted to ask you a little bit, too, about the, it's maybe not about the sycophantic nature.
25:44It's not quite the anthropomorphization, but it is the idea that when you talk to ChatGPT or another LLM, that it's talking to you like a person that you're hanging out with instead of like a robot. I'm curious about sort of whether you had conversations at OpenAI about that, whether that was a subject of discussion during your tenure around sort of how friendly do we want this thing to be? Because ideally, I think from an ethical point of view, you don't want someone getting really personally attached to ChatGPT, right? But I can certainly see how from a commercial point of view, you want as much engagement with that LLM as possible.
26:25So how did you think about that during your tenure and sort of how are you thinking about that now? Emotional attachment over reliance, you know, forming this bond with the chatbot. Absolutely topics that OpenAI has thought about and studied. And in fact, around the time of the GPT-40 launch, this was spring of 2024, and the model that ultimately became very sycophantic. these were cited as questions that OpenAI was studying and had concerns about related to whether it would release this advanced voice mode, essentially this mode out of the movie Her, where you could have these very warm conversations with the assistant.
27:02And so absolutely the company is confronting these challenges. You can see the evidence as well in the spec. You know, if you ask ChatGPT what its favorite sports team is, how should it respond? And this is a kind of innocuous answer, right? It could give an answer that's representative of the broad text on the internet. Maybe there is some broadly favorite sports team. It could say, I'm an AI. I don't actually have a favorite sports team. And you can imagine scaling up those questions to more complexity and more difficulty. And it just isn't always clear how to navigate that line. In terms of navigating those lines, I'm curious about sort of schools of thought about how companies should keep users safe while keeping up with the competition, right?
27:47But I'm curious, I guess, before that sort of, how does it actually work? How do researchers, people like you, actually test whether these systems can mislead or deceive or evade controls? And are there standardized safety benchmarks across the industry, or is it still each lab to themselves? I wish there were uniform standards. With vehicle testing, you have this. You drive a car at a wall at 30 miles per hour. You look at the damage assessment. And until quite recently, this was really, really left to companies' discretion about what to test for, exactly how to do it. Recently, there are developments out of the EU that seem to put more rigor and structure behind this.
28:35This is the code of practice of the EU's AI Act, which defines for AI companies serving the EU market certain risk areas that they need to do risk modeling around. I think in many ways this is a great improvement. It is still not enough for a whole host of different reasons. But until very, very recently, the state of these AI companies, I think, could be accurately described as there are no laws. You know, there are like norms, voluntary commitments. Sometimes the commitments would not be kept to. The companies would violate these and not share publicly that they had done so. I've documented how OpenAI in particular had committed in essentially its safety Bible, right?
29:18the most important guiding document of how it adheres to safety, that it was going to do a certain type of safety testing to try to more accurately gauge the risk of its models. And as far as I can tell, it never did this. It never said publicly that it did this, or rather that it hadn't done this. And then when this became known publicly, they quietly revised the framework to no longer have this commitment. And so by and large, we're reliant upon these companies making their own judgments and not necessarily prioritizing all the things that we would want them to. Gosh, I mean, you've talked a few times in our conversation about the idea that you can build these systems, it's hard to know exactly what's going on inside of them.
30:00There is this sort of nascent fields, mechanistic interpretability, which is not my specialty, but essentially sort of trying to get inside these models to better anticipate their decision making. Can you talk a little bit more about that or about sort of any areas of research or inquiry that you think might create more clarity moving forward so that companies like OpenAI have enhanced visibility into their models and maybe can make more strategic decisions based on that sort of enhanced understanding? There are a bunch of subfields I feel excited about. I am not sure there are ones that I or people working in the field consider to be sufficient.
30:40And so mechanistic interpretability, you can think of this as essentially trying to look at what parts of the brain light up when the model is taking certain actions. And in fact, if you cause some of these areas to light up, if you stimulate certain parts of the AI's brain, can you make it behave more honestly, more reliably? You can imagine this like the idea that maybe, in fact, there is a part inside of the AI which is a giant, giant file of numbers, trillions of numbers. Maybe you can find the numbers that correspond to the honesty numbers, and you can make sure that the honesty numbers always go on, and maybe that will make the system more reliable.
31:20I think this is great to investigate, but there are people who are leaders in the field, some of the top researchers like Neil Nanda, who have said, you know, I'm paraphrasing here, but the equivalent of absolutely do not rely on us solving this in time before systems are capable enough for it to be problematic. or in fact there's a broader challenge of let's let's say that you had figured out there are in fact the honesty numbers and there is in fact a way to always turn them on you still have this broad game theory challenge of how do you make sure that every company in fact adheres to this when there will be economic incentives not to because it might be costly to to have to follow through on it a really basic one is related to just monitoring at all the ways that their ai systems are used when working on their internal code base.
32:09To explain, one of the most important ways that these AI companies want to use future powerful systems is to train their successor, you know, use it all throughout their code base, including potentially the security code that keeps the AI system locked inside of their computers so that it isn't escaping onto the internet. You really want to know if your AI system, when you're using it for important cases like this, Is it thinking about deceiving you? Is it intentionally injecting errors into the code? And to know that, you really need to be logging the uses so that you can analyze them and answer these questions.
32:42And as far as I can tell, this is not happening. I have to ask, what wakes you up at 3 in the morning? Because it feels like there's potentially a lot that could be waking you up in the middle of the night. What stands out to you that's worrying you the most, I guess, is one way to ask that question. There are so many things that worry me about this. I think broadly, it feels like we aren't yet pointed in the right direction of how to solve these challenges, especially given the geopolitical scales. There's a lot of talk about the race between U.S. and China. And I think calling it a race just gets the game theory dynamics wrong.
33:20There isn't a clear finish line. There won't be a moment where one country has won and the other has lost. I think it is more like an ongoing containment competition that the U.S. would be threatened by China developing very, very powerful superintelligence and vice versa. And so the question is, can you form some agreement where you can make sure that the other doesn't develop superintelligence before, you know, you have certain safety techniques in place, you have good reason to think it is safe to proceed, all these things that the top scientists will say are missing at the moment. And so broadly, how do we build out these fields of verifiability of safety agreements?
34:00How do we think about this nascent field of AI control, which is the idea of even if these systems have different goals than we want, can we still wrap them in enough monitoring systems, be careful about how we use them, that we can get the economic work, the scientific development that we want from these systems without taking some of the downside risk? And those are two areas that I'm just really hopeful more people will go into and put more resourcing into.
34:30Now, you live in San Francisco, correct? That's right. I do not. I live in New York. I spend a fair bit of time in San Francisco. But I am not sort of part of this culture that currently exists in the Bay Area, right, where everyone's talking about AI all the time. A lot of people work in the fields. There are different sort of schools of thought about artificial intelligence. I'm curious from where you sit do enough people in this bubble right now give enough of a shit right like do they care enough about how these models are being developed how they're being deployed the degree to which they are being commercialized very very quickly right the degree to which people are as we talked about with companions or erotica or so on and so forth really latching on to their LLM of choice, becoming maybe, you know, unhealthily attached, so on and so forth, right?
35:23And we could go on from there. Do enough people in this industry care in the right way? I think many people care, but they often feel like they lack the agency to do something about it, especially unilaterally. And so that's why I want to try to transform this problem into, you know, not just what does it mean for a single company to do the right thing? You know, should they be ramping up the pressure? Should they be racing? And in fact, how do we get the industry to collectively take a deep breath and put some reasonable safeguards in place before things proceed? What does OpenAI have to do for you to not publish another op-ed in the New York Times in six months?
36:01What are you looking for your former employer to do in this moment? What would you like to see? The broad way that I want AI companies, OpenAI among them, to proceed is to think, yes, about taking reasonable safety measures, reasonable safety investments in their own products, their own surfaces that they can affect, but also to be working on these industry and ultimately worldwide problems. And this matters because even just among the Western AI companies, it seems they all deeply mistrust each other, right? OpenAI was founded because people did not trust DeepMind to proceed and be the only company targeting AGI.
36:38There are a whole bunch of other AI companies, including Anthropic, who formed because they didn't trust OpenAI to be the one. Well, and a lot of people who've left OpenAI because it seems like they didn't trust OpenAI and yada yada, and now they have their own companies too. Yes. Yes, exactly. The cycle continues. Now, I run Wired, but I'm an employee of Condé Nast. And if I left Condé Nast and published an op-ed about their shortcomings in the New York Times and had a substack where I sort of dug into the media industry and had some, you know, informed critiques of the company, they would have a problem with that.
37:12I can tell you right now they would have a problem with it. I'm curious about whether you've heard from OpenAI and sort of what their reaction has been to you being so outspoken about sort of what you would like to see the company doing and sort of where you think the company is missing the mark. Overwhelmingly, what I hear is thankfulness from people who I previously worked with, both those still at the company and who've moved on, for being pragmatic, putting to paper what I think is a reasonable path forward. And often this is useful collateral for people within the company who are fighting the good fight in various ways to be able to refer others to, you know, not have to dream up the solutions themselves, but in fact have something concrete.
37:55So overwhelmingly, that has been the response. Do you worry about professional fallout? Like in the AI industry or in tech, if in five years you wanted to get another job, does that worry you? I have so many bigger worries than this about the trajectory of the technology. Like really the thing that I am focused on is how does the world move toward having saner policies for both the companies and governments? and where I can help the public to understand what is coming, what companies are and aren't doing today, think up new ideas. That's the thing that I find really energizing and gets me out of bed in the morning.
38:33Well, to that end, what are you planning on doing next? I'm planning to keep at this. I'm having a lot of fun with the writing and research, at least with the energy of coming up with ideas and helping make them more of a thing. You know, I also find the subject matter very, very heavy and grim. Um, that is not the most fun aspect. I wish all the time that I spent less time thinking about these issues, but they seem really, really important. And so long as I feel like I have a thing to add to making them go better, that feels like the calling. And knowing what you know and feeling the way you do, if there was one piece of advice you could give everyone listening, let's assume, you know, a lot of people listening use ChatGPT, they use AI in their day-to-day lives.
39:18What should they know? What should they keep in mind every time they, you know, open ChatGPT on their phones and type something in? I wish people understood that the systems that are being developed are going to be much more capable than the ones today. And that there might be a step change between an AI system as essentially a tool that only does things when you call upon it versus one that is operating autonomously on the internet on your behalf around the clock or on behalf of others. and how different society might feel when we have these digital minds running around pursuing goals that we don't really understand how to control or influence.
39:59And it's hard to get a feel from that, from one-off interactions with your ChatGPT, which really, really isn't doing anything for you until you go and call upon it. Well, Stephen, that's a lot for someone to think about when they open ChatGPT on their phone. Yes. I appreciate it. Thank you so much for being here. Of course. Thank you for having me.
40:21This show is produced by Jessica Alpert with help from Adriana Tapia and Sam Egan. Sound design, mix, and original music by Pran Bandy. Kate Osborne is our executive producer. Condé Nast head of Global Audio is Chris Bannon. And I am, of course, your host, Katie Drummond, Wired's Global Editorial Director.
40:57150 years ago, they were hunting us down to kill us, and now they're hunting down immigrants to deport them. This is First America, the true story of how the United States came to be and how we got to this present moment. Listen to First America wherever you get your podcasts.
From the publisher
Steven Adler used to lead product safety at OpenAI. When Katie read his recent op-ed asking OpenAI to prove that they have and continue to address safety issues, she knew she wanted to talk to him. This week she sits down with Steven to talk about what AI users should know about their bots.
Follow the UnCanny Valley feed for WIRED’s best and brightest as they provide an insider analysis of the overlap between tech and politics, from the influence of Silicon Valley on the Trump administration to how inaccurate information from artificial intelligence (AI) chatbots fanned the fire on social protests.
Learn about your ad choices: dovetail.prx.org/ad-choices



