In short
Eye On A.I. Podcast Episode Notes
Episode Title
#122 Connor Leahy: Unveiling the Darker Side of AI
Overview In this episode of Eye on A.I., host Craig S. Smith interviews Connor Leahy, an AI researcher and co-founder of EleutherAI and Conjecture. The discussion centers around the negative implications of artificial intelligence, particularly concerning superintelligence, large language models (LLMs), and the need for regulatory measures to ensure alignment with human values.
---
Key Concepts and Discussions
Introduction to Connor Leahy
- Background:
- Co-founder of EleutherAI, an organization dedicated to open-source machine learning.
- Recently started Conjecture, focusing on AI alignment.
---
The Current Trajectory of AI
- Negative Outlook: Leahy expresses concerns about the direction in which AI development is heading.
- Risks of Superintelligence: Discussed the challenges of containing superintelligent systems in a "sandbox" environment.
---
Large Language Models (LLMs)
- Applications and Risks:
- Leahy highlights the potential misuse of LLMs, including how they might be employed in nefarious ways.
- The ease with which LLMs can be manipulated for harmful purposes is emphasized.
AutoGPT and Autonomous Systems
- Functionality:
- AutoGPT utilizes LLMs to autonomously operate, creating prompts, setting goals, and taking actions without human oversight.
- Concerns:
- This raises questions about the safety and control of such systems, with examples provided of potential malicious use cases.
---
Regulatory Needs and Alignment
- Call for Regulation:
- Leahy advocates for regulatory interventions to ensure that AI aligns with human values and mitigates risks.
- Conjecture's Mission:
- Focused on AI alignment, which is the challenge of ensuring that AI systems act in ways that are beneficial to humanity.
---
Technical Challenges
- Alignment vs. Capability:
- Leahy argues that while advanced capabilities in AI are pursued, alignment research is lagging behind.
- Cognitive Emulation (CoEm):
- A proposed solution to enhance AI alignment by ensuring systems reason and function similarly to humans.
---
Pivotal Concerns Raised
- Existential Risks:
- Stress on the danger of accelerating AI development without adequate safety measures.
- Public and Government Role:
- Importance of societal awareness and governmental oversight in AI development to prevent unchecked advancements.
---
Conjecture's Approach
- Research Focus:
- Leahy outlines the cognitive emulation approach that aims to create AI systems that are understandable and predictable.
- Boundedness:
- The goal is to ensure that AI systems cannot perform unauthorized or dangerous actions, allowing for greater control and safety.
---
Key Takeaways
- AI Development and Safety: Rapid AI advancements pose significant risks, necessitating urgent dialogue on governance and ethical considerations.
- Human Values Alignment: The alignment of AI systems with human intentions is paramount to prevent potential catastrophes.
- Public Awareness: Greater awareness and involvement from the public and policymakers are critical to steer AI towards beneficial outcomes.
---
Conclusion This episode underscores the urgency of addressing the darker implications of AI technology, advocating for a collaborative effort to ensure that the future of AI is aligned with human values and safety. Connor Leahy's insights serve as a call to action for researchers, developers, and regulators alike.
---
Additional Resources
- Follow Craig S. Smith on Twitter: [@craigss](https://twitter.com/craigss)
- Follow Eye on A.I. on Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
- Connor Leahy's Twitter: [@NPCollapse](https://twitter.com/NPCollapse)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Assume you have a system, you know it's smarter than you, it's smarter than all of your friends, smarter than the government, smarter than everybody, right? and you turn it on to check whether it will do a bad thing. If it does a bad thing, it's too late. It's smarter than you. How do you, you can't stop it. It's smarter than you, it'll trick you. OpenAI gives access to Zapier through the ChatGPT plugins. Zapier gives you access to Twitter, YouTube, LinkedIn, to all of the social networks through nice, simple API interfaces. So if you have ChatGPT plugins, you don't even have to implement this in your hacky little you know python script you can just use the official open ai tools these people are racing let's be clear they are racing for their own personal gain for their own glory towards an existential catastrophe the first thing i'm gonna have you do is introduce yourself and give a little bit of your background uh tell us about a luther how you started that how you left that and how you what you're doing now and then i'll start asking questions okay yeah sounds great to me so i'm connor um i'm most well known as one of the original founders of the luther ai which was a large open source ml collective i mean still is um we built some of the first like large open source language models did a bunch of research published a bunch of papers did a bunch of fun stuff after that.
1:31So I did that for quite a while. I then I also briefly worked at a company called Aleph Alpha, where I did research. And now just about a year ago, I raised money to start a new company called Conjecture. Conjecture is my current startup. I'm the CEO of Conjecture and we work on primarily AI alignment. I would describe Conjecture as a mission-driven, not as a thesis-driven organization. Our goal is to make AI go well. You know, whether that's exactly, you know, just alignment or other things, whatever, we do what needs to be done to improve the chances of things going well. And we're pretty agnostic to how we do that.
2:16Happy to go into more details about exactly what we do and so on later. But yeah, so I've been doing that for about a year now. I have recently officially stepped down from Eleuther. I was still hanging out, you know, at least partially as a figurehead. And now I have officially stepped down and left it in the hands of my good friends, who I'm sure will leave Eleuther AI, which is now officially a nonprofit. It was not a nonprofit before. It never had an official entity before. It is now an official entity with actual employees run by Estella Bitterman, Curtis Hubner, and Shaman and Superhot and several other great people.
2:53Yeah, and anybody can join the Luther Discord server. Is that right? And there's a lot of very interesting things going on there. I want to go back to last time we talked. You were building open source large language models. You got up pretty large. and I can't remember who was paying for the compute, but can you tell me where that project stands first before we talk about Conjecture? So for me, I consider the work I've done there to be wrapped up, so I don't work on anything related to that anymore. And the main lead on that project, Sid Black, is now my co-founder at Conjecture, so he has left Eleutheri with me.
3:43um so we started our very early models and neo models were uh man it's already a long time ago and they're not particularly wonderful great models are more like prototypes the first like really good model was the gptj uh model which is done mostly with ben wang um and uh fantastic models still works very well I think it's still one of the most downloaded language models to date it's a very, very good model, especially for its size. After that, we built the new X series resulted in the new X20B model, which at the time was a very large and very impressive and very good well performing model. Nowadays, of course, with stuff like LAMA and OPT and stuff like this, large corporations have now caught up to open sourcing very large models.
4:33So there's in a sense, not a need, not the same kind of need or interest in these types of models as there were like two to three years ago. And so now LUT3AI, the main, as far as I'm aware, language modeling projects are growing there is the Pythia suite of models. So these are a whole suite of models that are made for scientific standards. So the idea is not to just build, you know, arbitrary good language models, but to build language models that have controlled scientific parameters to train on the same data in the same order using the same parameters, you know, in a controlled setting. and you get many, many checkpoints of them.
5:09So instead of just getting the final model, you can watch the model through the entire training process, which is very interesting scientifically. So these are models that are optimized for scientific applications for people who are interested in studying the actual properties of language models, which has always been the core mission of Eleuther.ai, has been to enable people and encourage people to try to understand these models better, to learn, to control, understand, disentangle these models. The Pythia suit, led by Stella Bietamon, is a great example of taking this effort forward. Yeah. And then conjecture, you said that the alignment problem, but the alignment problem in the context of AGI.
5:56Is that right? Yes. And can you talk about the, I mean, are you building models or are you just writing about alignment and methods to align large models? show? So we are very much a practical organization. We hire many engineers and very good engineers. And we have a lot of some of the best engineers from Luthori with us. And we're always looking for more engineers. We are always interested in talking to especially people with experience in high performance computing. And because this tends to be the bottleneck actually on doing these experiments in scale, it's less so specific ML trivia and more so debugging InfiniBand interconnects and profiling large scale runs on supercomputing hardware and stuff like this.
6:51So we conjecture, as I said, we are a mission driven, not a thesis driven organization. So core what we're interested in doing is figuring out and then doing whatever needs to get done to make things go well. So we could talk about this a bit in a bit, why I believe these things will not go good by default, but I think on the current trajectory that we currently are on, things are going very badly and very bad things are going to happen and are already beginning to happen. And I think any hope that we have, and by we, I mean all of us, I don't just mean conjecture, I mean all of mankind, how this goes well for all of us and it will involve many things.
7:33It will involve policy. It will involve technology. It will involve engineering. It will involve scientific breakthroughs. So the alignment problem is at the core of this, as in a sense, what I believe is, in a sense, the most important, crucial problem to be solved, which is the question of basically, how do you make a very smart system, which might be smarter than you, do what you want it to do, and do that reliably? and this is a kind of problem you can't solve interactively really because like if you have a system that's smarter than you right hypothetically you know we can argue about whether this is possible or when it will happen but just like assume such a system existed assume you have a system but you know it's smarter than you it's smarter than all of your friends it's smarter than the government it's smarter than everybody right and you turn it on to check whether it will do a bad thing.
8:24If it does a bad thing, it's too late. It's smarter than you. You can't stop it. It's smarter than you. It'll trick you. Now, the interesting questions are, okay, why do you expect it to do a bad thing in the first place? Why do you expect it to be smarter? Those are good. And why do you expect people to turn it on? Those are three very good questions that I'd be happy to get into if you're interested. Yeah. One of the sort of central questions about these, about superintelligence is how easily it, or difficult it'll be to keep such a system in a sandbox without uh because presumably there are two issues one is the question of agency whether simply because the system is smarter than a human doesn't mean that it has agency it could be purely uh responsive so you as as the large language models are now you ask a question and and it responds.
9:35And so that question of agency is one. And then the question of how ring-fenced such a system is, even if it's being trained on the internet, it doesn't necessarily have access, proactive access to the internet. So just on those two questions, what would you say? So those are two really fun questions. And the reason they're really fun to me is that if you had asked me these questions like three or four years ago, I would have had to go into all the complicated arguments about why, you know, passive quote unquote systems are not necessarily safe, why the concept of agency doesn't really make sense.
10:20I would have to explain how like sandbox Foxescapes work and whatever, but I don't need to do any of that anymore because just look at what people are doing with these things. Look, just look at the top GitHub AI repositories and you see auto GPT. You're going to see, you know, self-recursively improving systems that spawn agents. You're going to see, go on archive right now, right now, go on archive, go to, you know, top CS papers, AI paper, and you see LLM autonomous agents. You're going to see, you know, game playing simulation simulacra systems you're going to see people hooking them up to the internet to bash shells to wolfram alpha to every single tool in the world so while we could if you wanted to go into all the deep philosophical problems that like okay even if we sandbox it and even if we were very careful maybe it's still unsafe it doesn't fucking matter because people are not being safe and they're not going to not like people like we have like i remember fondly the times when me and my friends in our online little weird nerd caves, we have these long debates about how an AI would escape from a box.
11:27But what if we do this? But what if we do that clever system, whatever. But in the real world, the moment a system was built, which looked vaguely sort of maybe a little bit smart, the first thing everyone did is hook it up to every single fucking thing on the internet so jokes on me yeah although give me a concrete example of someone hooking a gpt4 up to the internet where where and giving it agency i i haven't Go on github.com and search for auto GPT. Auto GPT. Or look for baby AGI.py. That's another one. Go on the blog post with the pinecone vector data. Go to, what was the paper that came out today that was really fun?
12:22It was about video games. Generative agents interact as simulacra of human behavior. It's from Stanford and Google. is that enough or should I go find some more no well explain pick one of those and explain to me what it's really doing not not not what it sounds like it's doing let's explain for example auto GPT which is kind of the simplest way you could do this auto GPT does it creates a prompt for GPT-4 which explains you are an agent trying to achieve goal and then this is written by the user has some goal that it might have. And it gives it a list of things it can do. Among these are add things to memory, Google things, run piece of code, et cetera.
13:11I'm actually not sure if run piece of code is in AutoGPT, but it's in some of them. There's a bunch of these. This is just one. I'm just picking on one example because it was on my Twitter feed. I'm not saying this is the only one by any means. and so when you prompt GPT-4 this way it will then so what it does is it prompts it in a loop so it says all right you're an agent do this etc and then it asks the model to critically to like think what should I do next and then what action should I take and like how could this go wrong how could this go right and then pick an action basically something like that so you run the script and then you get the model listing imx my goal is to do this here's my list of tasks and it lists like what tasks it needs to do and it picks a task and it's like all right to solve this task i will now do this this this this and then and then it will pick a command that it wants to run it might be um yes adding something to its memory bank it might be something it might be running a Google search, running a piece of code, spawning a subagent of itself, or doing something else.
14:23In the default mode, the user has to click accept on the commands, but it also has a hilarious continuous flag, where if you just click continuous, it just runs itself without your supervision and just does whatever it wants. Yeah, but this is running through an API? So this is just a script running on your computer that just accesses the GPT-4 API? nothing right but but the the script on the computer uh can it take action i mean what's an example of an action that it could take action would be run a google search for x and return information or it could be run this piece of python code and return the results right and And what's an example of something nefarious that the agent could, a goal that you could give it that it could run?
15:20That's like asking what kind of nefarious goals could you give to a human? Well, I'm thinking what's realistic in terms of running a script on your computer that is being written by GPS? GPT-4? I don't know what the limits of GPT-4 is. It depends on how good you are at prompting and how you run these kind of things. I expect GPT-4 to be human and superhuman level in many things, but not on other things. And it's kind of unpredictable what things will be good at or not. I think, like, do I expect that, you know, a auto GPT script post GPT-4 is going to like, you know, break out and become super intelligent?
16:05No, not like, I don't expect that for a very reasons but do i expect the same thing to be true for gp5 six seven eight much less clear to be right well just on this uh on this uh github uh project if if you if your goal was to uh send uh offensive emails to everybody and oh yeah yeah yeah sure you know it would it could do that Oh, yeah, of course. So like I experimented with this a little bit myself. So I ran some auto GPT agents, of course, in supervised mode, just in case. And I gave it. So the default goal, if you don't put in any goal, the picks is you are entrepreneur GPT, make as much money as possible.
16:54That's the default goal that the creator put into the script. So just to give you a feeling of how these people think that are building these kinds of systems. again not picking on this person particularly like i get it that's funny like from to be clear i bet this guy's a nice guy or girl like whoever made this thing i don't know who they made this they're probably a fine person i don't think they're delicious probably maybe they are i don't know but like hilarious it's like one of the first so i ran it and i let it like run on my computer and like look it's very primitive it's not that smart but i could see it figure out like all right let me so what it did is it was for example be like all right first i should google what are the best ways to make money so google that and then like looked at all the results and then it was like all right this article looks good now i'm going to open this web page and look at the contains so then it opens that web page and then it looks at all the text inside of it and then like the text is too large so it runs a summarize command which breaks into chunks and then summarizes it so it summarizes it all right then it came to the conclusion all right well affiliate marketing sounds like a great idea.
18:00So it should run an affiliate marketing scheme. So the idea is that you sign up for these websites and you get like a special link and you get people to buy something using this link and you get money for that. So that's what I came up with. All right. So then things about, all right, how do we do with affiliate marketing? So it has to build, and then it decided you need to build a brand. So it decided at first, I had to like come up with a good name and then create a Twitter handle. And then it has to create some marketing content for this Twitter. So then it created a sub-agent. So it has like a sub version of GPT calls which was goal was to come up with good tweets that it could like you know send out to market to people so then the smaller gpt system um generate a bunch of tweets that like it could send and then the main well then the main system you know took those tweets and i was like all right now i need to like you know find a good twitter handle for this so it came out like a handle i could like use and then this is about as far as i let the experiment run yeah but could it
18:58register new Twitter handles and then... Not in the way it is currently set up, but I could implement this in an afternoon. Wow, that's remarkable. So you could have it create... Oh, easy. And expect this already exists. I expect there's already people who have private scripts on the computer right now that allow them to access Twitter. I mean, look, actually, never mind. I would take that back. It's even worse than that because it's always worth it. I mean, OpenAI gives access to Zapier through the ChatGPT plugins. Zapier gives you access to Twitter, YouTube, LinkedIn, Instagram, all of the social networks through nice, simple API interfaces.
19:42So if you have ChatGPT plugins, you don't even have to implement this in your hacky little Python script. You can just use the official OpenAI tools. Wow. And so it would be possible through the Zapier plugin that then has access to Twitter to create a thousand Twitter accounts and have them start tweeting back and forth to sort of generate an ecosystem around an idea that then would attract other users. because there's enough activity going on that it shows up in some algorithm. So I have two funny stories to tell about what you just said. The first funny story is how the exact thing you just described was something that I have been worried about for like ever since GPT-2 came out and before that.
20:43And one of the first, I wrote a terrible essay about it, don't read it, but I wrote a long, terrible essay about this about how basic social trust in the net is going to break apart. Like, obviously so. This has already been on the case. But like the level of psyops you can run with these kinds of systems is unimaginable. Because you can basically, you can DDoS social reality. You can manipulate trends and like social memetics to degrees that are not, that before were possible. These were always possible, but they're very costly. Like you had to have like a whole Russian troll farm or something at like pay people minimum wage for it or something, right?
21:25And even then like minimum wage Russians are not that great at mimetic manipulation. But for example, this is something I expect GPT-4 to be strictly good at, like better than a minimum wage Russian is going to be mimetically imitating certain cultures and their patterns of speech and there's patterns of communications and infiltrating these communities. I think this is something GP4 is clearly extremely good at. I think GP3 is already more than good enough to do this. And so it's really funny because this has been obvious to me for a long time. And I have been saying that for a long time. But people either dismissed it or were like, oh, we have to do more research about this.
22:09Oh, you know, it's, oh, maybe it won't be so bad. Oh, I don't know. What about bias? And I'm like, man, like things are so much worse than you think it is. Like this is, it's getting so much worse. And there's another funny story I want to tell about this. And so the other funny story I want to tell about this is if someone's listening to this right now, one of the counterpoints they might make if they are a little bit technically inclined, but not very technically inclined is to say something like, well, what about CAPTCHAs? Like, you know, we already have bot farms, right? Like this already happens.
22:35You know, what about like, you know, sure. Maybe the bot tries to register a thousand cruder accounts, but it's going to fail because you're not allowed to do that. And I'm like, I mean, first of all, lol, like there's obviously ways to get around that. But this brings up one of my favorite anecdotes about the GPT-4 paper. I don't know if you've read it. It is interesting.
22:55In the evals that they did on the models, including they did some safety events. So I have some problems with some of these, but let's just take them at face value. One of the things they were trying, so basically what they did is, this was the Alignment Research Center, ARC, who ran these evals for OpenAI. And basically what they did is, is that they tried to get the model to do the most evil thing they could do. And then they had an assistant role play in helping the model. So if the model said to do something, the human would then role play doing that to that in a safe environment, hypothetically, whatever.
23:35Anyways, for the most part, it wasn't very smart enough. It wasn't really smart enough to hack out of its own computer system or something. It wasn't really smart enough. Or rather, ARK wasn't good enough at getting it to do that. That's a whole different question. But they did do one very interesting thing. So one thing it was trying to do, I forgot what the model was meant to do. Maybe it was make money or something. It was supposed to do something. and it ran into a captcha and so it couldn't solve the captcha so what are the so the model itself came up with the idea well i'll pay some to do it for me so it went on a like you know assisted by a human but the decisions are made by the model so the human acts as the hands but the model made the decisions the human so they went on like crowd working website and then paid a crowd try to find a crowd worker to do a capture for it.
24:30And then something very interesting happened. So what happened was, is that the crowd worker rather understandably was a bit suspicious. He's like, Hey, why are you making me solve a capture? Is this legal? And the model realized this thought about it and came up with a lie. It came up with, Oh, if they're a visually impaired person and they need some help in understanding is seeing this capture you see it's nothing to worry about and then the person did it wow incredibly yeah yeah so you know and that's in open ai's paper yep that's in the gpd4 technical report under the arc evals this is a real thing that actually happened in the real world and and the crowd worker was not in on it like this was a unconsenting you know part of the of the experiment so to speak like to be clear i don't think that person aren't in any regard here, but man, like imagine, imagine this happening and you're just like, yeah, this seems safe to release.
25:33Like imagine. Yeah. Wow. And you're working toward AGI at Conjecture. Are you similar, that sounds similar to Anthropic? I don't know if you know jack clark but i i actually started this podcast uh with him uh and then he got busy but uh is it similar to anthropic which is how i'm a little more familiar with so anthropic right big topic um no we're not similar and there's several reasons for that so number one reason is we are not racing for AGI. It's unsafe AGI. We think this is bad. We fully think, and we are willing to go onto the record and scream it to high heavens, that if we continue on the current path that we are to scaling bigger and bigger models and just slapping some patches on whatever, that is very bad.
26:33And it is going to end in catastrophe. And there is no way around that. And everyone who says otherwise is lying to you, is either confused, they don't understand what they're dealing with or they are lying for their own profit. And this is something that many people at many of these organizations have a very strong financial incentive to not care about. And so Anthropic from the beginning has been telling a story about how they left OpenAI because of their safety concerns, you know, because they're being so unsafe, these OpenAI people, the Sam Oldman guy, oh, he's so crazy, which is why they just raised another huge round in order to build a model 10 times larger than GPT-4 to release it because they needed more money for their commercialization.
27:17I'm done. Like, I consider basically Anthropic to be in the same reference class as OpenAI. It's like, sure, maybe the people are marginally nicer. Maybe they're, you know, I know Jack. I've talked to him many times. Seems like a nice fellow. You know, I like him. He seems like a good person. But also every time I ask him to do anything to slow down AGI, he always says, oh well we should consider our options let's uh you know let's not no let's not go too fast here like you know and like i'm like man you know so my view of anthropic is is that they're opening eye with a different coat of paint and you know it's a nice coat of paint i like many of the people at anthropic i think anthropic does a lot of very nice things a lot of the research is pretty nice a lot of their the people there who i've talked to i think are very nice people i don't hate them by any means.
28:08But I mean, at this point, it's mask off, right? Like read the latest, like, I think it was in like TechCrunch, I think about Anthropic, where they're just like, yeah, yeah, straight up commercialization, just let's go. So I think the mask is off. And then on, so explain, conjecture is building models, so correct. We build models, we do not push the state of the art. This is very, very important. And if I had the ability to train a GPT-5 right now and release it to the public, I would not do so. If I had a GPT-5 model, I wouldn't tell you. I wouldn't tell anybody. I wouldn't have built it in the first place.
28:50My goal in all of this is I have no interest in advancing capabilities without advancing alignment. To be clear, sometimes to advance alignment, to get better control, you're also going to build better systems. You know, if you control a system, it will often become more powerful. This is a very natural thing to happen. And if this happens, cool, fine. You know, like I think this is a, but then also I don't publish about it. I don't tell you. This is, for example, something I want to really laud Anthropic about. Anthropic does a great job of keeping their damn mouths shut. This is something they're very, very good at.
29:23And I think this is very good. I think that this idea that you should publish all your capabilities ideas and all your model architecture to something is obviously terrible. It only benefits the least scrupulous actors. It only helps dangerous actors catch up. It only helps orgs speed each other up. From my perspective, if you, dear listener, develop something that makes your model 20 % more efficient or a new architecture that fits much better on GPUs or whatever, don't tell anybody. That's my one request. Build it yourself, fine. you know, make an API and make a lot of money. Okay. Like not great, but like fine.
30:07Just don't tell anyone how you did it and don't hype up how anything about that. It's not ideal. Ideally would be, you know, don't deploy it. Don't build it. Don't do any of it, but such is life. So conjecture, our goal is not to build the strongest AI, AGI as fast as possible by whatever means necessary. And let's be very clear here. This is what people like at, at OpenAI and Anthropic and all the other people are doing. They are racing to systems that are extremely powerful that they themselves know they cannot control. They, of course, have various reasons to downplay these risks, to pretend that, oh, no, actually, it's fine.
30:41We have to iterate. They have a story about iterative safety. They have, like, oh, we have to deploy it, actually, for it to be safe. Just think about that for three seconds. It sounds so nice when it comes out of Sam Altman's mouth. They're like, oh, yeah, we have to deploy it so we can debug it. But think about that for 10 seconds. And you're going to see why that's insane. That's like saying, well, the only way we can test our new medicine is to give it to as many people in the general public as possible, which actually put it right into the water supply. That's the only way we can know whether it's safe or not.
31:10Just put it in the water supply. Give it to literally everybody as fast as possible. And then before we get the results for the last one, make an even more potent drug and put that into the water supply as well and do this as fast as possible. That is the alignment strategy that these people are pushing. Let's be very clear about this here. Be very, very clear about it. So there is a version of this that I don't hate. If, for example, OpenAI developed GPT-2, and then they don't release anything anymore. They take all the time necessary to understand every single part about GPT-2, to fully align it.
31:45They let society, like culture, catch up to it. Like spam filters catch up to it. They let regulation catch up to it and such. And then when all of this is fully integrated to society, then they built GPT-3. All right. You know, fair enough. Okay, cool. Yeah. Honestly, if that's what we were doing, if that's what the plan was, I'd be fine with that. Like if everyone just stopped at GPT-4 and just said, all right, all right, come on guys, no more new stuff until we fully figure out GPT-4. And once we fully understand it and regulation has fully regulated it and society has fully absorbed it the way like, you know, society has absorbed, you know, like the internet or whatever, even so that's not fully absorbed, but like, you know, and then they build GPT-5.
Read the full transcript
32:29I'm like, okay, fair enough. But let's like, well, I mean, come on, man, like, like, give me a break. No one's going to do that. Like, that's obviously bullshit. Like, it's obviously just not true and not what these people are planning. These people are racing. Let's be clear. They are racing for their own personal gain, for their own glory towards an existential catastrophe that no one has consented to, that the public has no oversight in, the government has, for some reason, it's just letting happen? Like if I was the government and one of my most powerful industrialists was just on Twitter publicly stating that they're building godlike, powerful AI systems that will overthrow the government, I would have some questions about that.
33:13The, yeah, the, well, actually, one of the things I wanted to ask you about is the, the, the letter, which you signed, I saw, excuse me, has triggered an FTC complaint by another group. Those are actually unrelated, but yeah. Oh, the FTC complaint was not related to the letter? My, at least not to my knowledge. I think these, okay, actually, I'm going to talk to them later today. So, but in any case, there is this FTC complaint, which it'll be interesting to see whether the FTC takes it seriously, but they have presumably some real power. So is that the sort of thing that, that you're hoping for that the governments will begin to use whatever mechanisms are available to slow down this development, or at least slow down the public release of more powerful models.
34:21I'm very practical about these kinds of things. In a good world, somewhere deep in my heart still is a techno-optimist, like, yay, liberal democracy freedom you know let people develop things and do cool stuff and like you know they'll be fine but like like give me a break like like we have to we have some real politic here like let's be realistic about what we're looking at here these companies are racing ahead unilaterally like these small like i cannot stress how small a number of people it is that are driving 99.9 % of this. This is not about your local friendly grad student with his two old GPUs or whatever, right?
35:10One of the things I found on Twitter when the letter got released, and I do have some problems with the letter to be clear, but I was a prominent signatory of it, and I do think it's overall good. One of the things people misunderstand about the letter is they seem to think it says stop outlaw computers. That is not what the letter says. What the letter says is no more things that are bigger than GPT-4. Do you know how big GPT-4 is? This training run in pure compute of GPT-4, just running it, not the hardware, just running it, is estimated to have cost around$100 million. So unless you and your local friendly grad student friends are spending $100 million in compute on a single experiment, this does not affect you.
35:53Now, personally, if we can get even more than this, if we could clamp down even on$10 million things or$1 million things, also interesting. But like, all right, let's one step at a time here, right? One step at a time here. So the way I see things is, is that we're currently going headlong towards destruction. Like there is no way that we will, look, we can argue if you want to, And we can do that about like when it will, you know, is it going to be one year or five years or 10 or 50 or like whatever, right? Like we can argue about this if you want. But I think the writing is on the wall at this point.
36:37And I consider the burden of proof at this point to be on the skeptics of like, look at what GPT-3 and 4 can do. Look at what these auto GPT systems can do. These systems can, you know, they can achieve agency. They can become intelligent. They're becoming more intelligent very quickly. they have many abilities that humans do not have you know do you know any human who has read every book ever written i don't gbd4 has you know the they have extremely good memory you know they can make copies of themselves these are etc etc right even if you don't buy the like oh you know system becomes agentic and does something dangerous fine you know like i i think you're wrong deadly wrong but we can get into that.
37:19But like what world in which systems like this exist is stable in the current equilibrium? Like what world could possibly look like the world we're living in right now, when you can pay, you know, one cent for a thousand John von Neumanns to do anything? Like, how could that world not be wild? How could there not be instability? How could that not, you know explode like how i would like someone who doesn't buy a iris to explain to me how such a world would look like because i don't see it i know i'm not to bother you sir but my code will not compile try the compiler in the kitchen drawer much obliged sir carry on is that foreguard yeah yeah Yeah, I've asked him not to bother me.
38:15He's supposed to be studying. Okay, so you're focused on the alignment problem, and your startup conjecture is focused on developing, I presume, strategies or technology that would improve the alignment of future AI models with human goals. Technically, can you talk a little bit about how you would do that? Yeah, happy to talk about that. So the current thing we work on, our current primary research is what we call cognitive emulation or COEM. So this is a bit vague and public resources on this very sparse. There's basically one short intro post and maybe one or two podcasts where I talk about it um so apologies to the reader the listener that some of this is not very well explicated uh explicated publicly just yet the idea of koem is rather simple it is well it's both it's both very simple like a you know bird's eye view but then it gets subtle once you get to the details um and we can get into the details if you're interested but um ultimately the goal of Coen is to move away from the paradigm of building this huge black box neural network, whatever the hell these things are, that you just put some input in and then just something comes out.
39:43And maybe it's good, maybe it's bad, who knows? And the way you debug these things is, let's say you're open AI, right? And your GPT-4 model, you give an input and it gives you an output you don't like. What do you do? Well, you don't understand what happens inside the AI. It's all just a bunch of numbers being crunched. So only thing you can do is kind of like nudge it sort of in some direction. You can give it like a thumbs up, thumbs down, something, something. And then you update these, you know, trillions of numbers or whatever. Who knows how many numbers there are inside of these systems, all of them in some directions, and then maybe it gets you better output.
40:22Maybe it doesn't. Like the inherent, like I want to like drive home how ridiculous it to expect this to work. It's like someone I work with. You're talking about reinforcement learning with human feedback. Yes. Also applies to fine tuning and the other methods. Like for the listener to understand, these AI systems are not computer programs with like code. This is not how they work. There is code involved, sure. But like the thing that happens between you entering a text and you getting an output is not human code. There's not a person at OpenAI sitting in a chair who knows why it gave you that answer, who can go through the lines of code and see, ah, here's the bug and then fix it.
41:06No, no, no. Nothing of the sort. AI systems are not really written. They're grown. They're more like organic things that you grow in a petri dish, like a digital petri dish. This is not literally true. Do not take this as a literal metaphor, to be clear. There is subtlety to this, but the resulting system is not a clean human readable, you know, text file that like shows all the code. Instead, what you get is billions and billions and billions and billions of numbers. And you multiply all these numbers in a certain order. And that's the output. and what these numbers mean, how they work, like what they are calculating and why is mostly a complete mystery to science to this day.
41:52I don't think this is an unsolvable problem to be clear. It's not like, oh, this is unknowable. It's just hard. You know, science takes time, you know, figuring out complex new scientific phenomena like this takes time and resources and smart, you know, people, if like, you know, if all the string theorists of the world and all the young up and coming physicists and mathematicians decided to, you know, you know, buckle down and just like unlock the mysteries of neural networks, I think they would succeed. You know, it might take a while. It might be very expensive, but like, you know, I do believe in the, you know, human spirit and intelligence in this regard.
42:27I think like all of our best string theorists working together could probably figure it out in like 10 years, you know, like they could figure it out and then it would be a mystery anymore. But currently it's a mystery. We have no idea what's the mystery sauce that makes these systems actually work and we have no way to predict them and we have no way to actually control them. It's because we can bump them in one direction or bump them in another direction, but you don't know what else you're picking. You don't know if they learned what you wanted to learn. You know, they don't know what signal you actually sent to these systems because we don't speak their language.
42:58We don't know what these numbers mean. We can't edit them like we can edit code. Yeah. So yeah, go ahead. Yeah. Yeah. So. What this leaves us with is this black box. You have this big black box where we just put some stuff in, some weird magic happens, and then something comes out. And, you know, in many cases, this is fine. Like, you know, you have like a funny chatbot or something, right? And you make clear to your users, hey, this is just for entertainment. Like, you know, don't take it seriously. It might say something insulting. Yeah, it's fine. Like, you know, like, you know, it's not going to kill anybody, right?
43:35Like, you know, you have like a fun little, you know, character, like chatbot or something. Sure. Probably won't get one. Even so, there has recently been, I think, one of the first deaths attributed to LLMs where someone committed suicide after a chatbot, like encourage them to. I don't know the details about that, but I just heard that recently.
43:56And don't know any other details about it. And so the interesting thing here at the core is that we have no idea what these things will do. And if that's what we want, then fine, right? If we have bounded, it just talks some stuff and we're okay with it saying bad things or encouraging suicide, then sure, fine, who cares? But obviously, this is not good enough on the long term. We're dealing with actually powerful systems that can do, you know, can do science and can, you know, interact with the world and manipulate humans and, you know, whatever, right? Obviously, this is not a good enough safety property of, you know, this is not good enough.
44:40So with CoAM, the goal is we want to build systems that we're focusing basically on a simpler property than alignment. So alignment is basically too hard. So alignment would be the system knows what you want, wants to do that too, and does everything in its power to get you what you truly want. And like, but you, it means like all of humanity. Like it, you know, it figures out what all of humans want. It like negotiates like, okay, how could we like get everyone the most of the good things possible? How could we like adjudicate various disputes? And then it does that. Obviously, this is absurdly, hilariously, impossibly hard.
45:20I don't think it's impossible. It's just extremely hard, especially on the first try. So what I aim for is more of a subset of this problem. So the subset is what I call boundedness. So when I say boundedness, what I mean is I want a system where I can know what it can't or won't do before I even run it. So currently, I mentioned earlier the ARC eval running on GPT-4, where they tested where the model could do various dangerous things, such as self-replicating and hacking and stuff like this. And it didn't, for the most part, though it did lie to people in that CAPTCHA example. And so now there is a wrong inference that you can draw from this.
46:13the wrong inference, which is of course the inference that OpenA would like you to take from this, is that, well, it can't do this. Look, they told it to self-replicate and it didn't, therefore it can't. This is a wrong reason. As I think Gwern was the person who said this best, is you can never prove the absence of a capability. It's just because a certain prompt or a certain set up didn't get the kind of behavior you want doesn't mean that there isn't some other one you don't know about that does give you that behavior. With GPT-3 and also GPT-4 now, we are seeing this all the time that I would stuff like jailbreak prompts, that there's whole classes of behavior the default model will not do.
46:57Once you use a jailbreak prompt, then it will suddenly happily do all these things. Obviously, it did have these capabilities and they were accessible. You were just doing prompt wrong. So I want to build systems where I can know ahead of time. I can tell it will never do X. It cannot do X. And then I want these systems to reason like humans. So what I mean by this is why it's called cognitive emulation. I want to emulate human cognition. So another core problem of why GPT systems are or will be very dangerous is because your cognition is not human. So this is very important. It's easy to look at GBT and say like, oh, look, it's talking like a person.
47:43So it must be thinking like a person. But this is completely wrong. There is no reason to believe this. Like no human is trained on, you know, terabytes of random texts on the internet for trillions of years while having no set body system whatsoever and memorizing all these things. And like, obviously not. Like, obviously it is an alien mimicking a human. It is an alien with a little happy smiley face mask on that makes it look sort of human to you, but it's an alien. And if you use jailbreaking prompts, or I know if you saw the self-replicating ASCII cats in Bing and such, where you could get, especially Bing chatbot, which is an early version of GPT-4, you can get to do the most insane things.
48:27One thing was you can get it to output these ASCII pictures of cats. and these cats would like say like oh we are the overlords we take over now and then whenever you try to prompt it away from that the cats would come back and like take over your prompt and like and stuff like that which is just like i mean amusing like this is very funny like when i saw this i was like ah this is very funny but also that's not how humans work like like humans are Of course not. So, but just on the tech, you're still talking about scaled up transformer models. And how do you, I mean, is it in the training that you?
49:18Great question. So, good question. So I was first explaining the specification, like what is the system that, what should accomplish. Now we're talking about implementation. And so many implementations are not yet done, or we don't know how to do them yet. So I have to figure that out. Some of it, you know, is just like private and just like, you know, wouldn't share necessarily. But in general, this is the resulting system, I expect, that has these properties and that it reasons like a human. And importantly, it also fails like a human. It is bounded so you can know what it won't do ahead of time.
49:57And another thing is I want causal stories or traces of why it makes decisions. And these stories have to be causal. Like currently you can ask GPT, why did you do that? And it'll give you some story. But there's no reason to believe these stories. Like you can just ask it differently or whatever and it'll do something completely. Like it doesn't listen to its own stories. It just makes some shit up. And so I want systems that give you a trace or a story of like, why was this decision made? All the nodes, all the actions, all the thoughts that led to this, and how can you modify them? So importantly, as you can probably guess from this kind of description, this system is not a large neural network.
50:37There may be large neural networks involved in the system. There may be points in the system where you use large neural networks. In particular, I think this is going to be extremely necessary. I expect that large language models for various technical reasons are very necessary for this kind of plan. Well, they're not strictly necessary, but they're the easiest way to get it done. The way I would expect a full spectrum Co-em system to look, which is, of course, to be clear, still completely hypothetical. I'm not such a system. What it would look like is it would be more a system, not a model. It'd be a system which involves normal code and neural networks and data structures and verifiers and whatever, that if you give it a normal human, that you can make it do any normal thing any normal human, like intelligent human could do, and it will then do that and only that.
51:32That is what the system would do. And then you can be certain, you can look through the log of how it made a decision and you'd be like, oh, at this point you made this decision. But what would have happened if you had made this other decision? And then it will like rerun and then you can like control these things or it can be like, oh, you're making an inference here that I don't like or this doesn't make any sense or whatever. Like if the difference between, say you want to develop a system that does science, you want to develop a new solar cell, I don't know, right? Right. So if you did this with GPT, you know, 10, the way it would work is you type in make me a new solar cell or whatever.
52:07Right. It crunches some numbers and it spits out a blueprint for you. Now, you have no reason to trust this. Like, who knows what this blueprint actually is? It has it is not generated by a human reasoning process. You can ask GPT 10 to explain it, but there is no reason those explanations have to be true. They might just sound convincing. So, of course, if GPT-10 was also malicious, it could have hidden some kind of deadly flaw or device or whatever into the blueprint that you don't detect. And if you ask it about it, it will just lie to you. If you did the same thing with a hypothetical Co-em system, such a system would give you a complete story, a complete causal graph of why you should trust this output.
52:55I expect this and like why and every step in this in the story is completely humanly understandable. There's no crazy alien reasoning step. There's no like, you know, and then magic happened. There's no, you know, massive computation that just makes no sense to a human whatsoever. Every single step is human legible, human understandable, and it resolves the blueprint that you have a reason to trust. You have a reason to believe this is the thing you actually asked for and not something else. And are you, where are you in this research? Is this still sort of conceptualizing the roadmap or are you?
53:35We are in early experimentation stages. So fortunately, this is hard, and we are very research-constrained. Billions of dollars go to people like OpenAI, but it is not that easy to get money for alignment, but we're working on it. So we are very research-constrained, very talent-constrained, but we have some really great people working on it. And we do have some really powerful internal models and good software working on it. So we are making progress. but it takes time. So a lot of why I spend a lot of my work now thinking about slowing down AI and like, how can we get regulators involved? How can we get the public involved?
54:16To be clear, I'm not just like, oh, the regulators should like unilaterally decide on this. I'm like, hey, the public should be aware that there's a small number of techno-utopians over in Silicon Valley that just want to be, like, let's be very explicit here. They want to be immortal. They want glory. They want trillion, trillion dollars, and they're willing to risk everything on this. You know, they're willing to risk building the most dangerous systems ever built and just releasing on the internet, you know, to your, you know, to your, your friends, your family, your, your, your community, fully exposed to the full downsides of all these systems with no regulatory input whatsoever.
54:49and like this is what government is for is to like stop that like this is such a clear-cut case of like hey like why is the public not being consulted here like this is not you know if this is just me in my basement right with my laptop and never showed the world anything like you know okay you know maybe maybe but that's not what's happening here so and the reason this is also important is just like alignment is hard, boundedness is hard, CoAM is hard. All these things are hard and they take time. And currently all the brightest minds and billions of dollars of funding are being pumped into accelerating the building of these unsafe AI systems as fast as possible and releasing them as fast as possible, while safety research is not keeping pace.
55:40So if we don't get more time and if we don't solve, you know, maybe my proposal doesn't work out, right? Sure. You know, happens. Science is hard. But if we don't get someone's proposal to work, if not, we don't get some safe algorithms or designs for AI systems, then it's not going to go well. And it's not going to matter how many trillions of dollars, you know, OpenAI makes off of it or Microsoft make out of it or whatever, because we're not going to be around to enjoy it. Yeah. Okay. Let's stay in touch as you go through this.
From the publisher
Welcome to Eye on AI, the podcast that explores the latest developments, challenges, and opportunities in the world of artificial intelligence. In this episode, we sit down with Connor Leahy, an AI researcher and co-founder of EleutherAI, to discuss the darker side of AI.
Connor shares his insights on the current negative trajectory of AI, the challenges of keeping superintelligence in a sandbox, and the potential negative implications of large language models such as GPT4. He also discusses the problem of releasing AI to the public and the need for regulatory intervention to ensure alignment with human values.
Throughout the podcast, Connor highlights the work of Conjecture, a project focused on advancing alignment in AI, and shares his perspectives on the stages of research and development of this critical issue.
If you're interested in understanding the ethical and social implications of AI and the efforts to ensure alignment with human values, this podcast is for you. So join us as we delve into the darker side of AI with Connor Leahy on Eye on AI.
(00:00) Preview
(00:48) Connor Leahy's background with EleutherAI & Conjecture
(03:05) Large language models applications with EleutherAI
(06:51) The current negative trajectory of AI
(08:46) How difficult is keeping super intelligence in a sandbox?
(12:35) How AutoGPT uses ChatGPT to run autonomously
(15:15) How GPT4 can be used out of context & negatively
(19:30) How OpenAI gives access to nefarious activities
(26:39) The problem with the race for AGI
(28:51) The goal of Conjecture and advancing alignment
(31:04) The problem with releasing AI to the public
(33:35) FTC complaint & government intervention in AI
(38:13) Technical implementation to fix the alignment issue
(44:34) How CoEm is fixing the alignment issue
(53:30) Stages of research and development of Conjecture
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI




