In short
Dwarkesh Podcast Episode Notes
Episode Summary In this episode of the Dwarkesh Podcast, host Dwarkesh Patel engages in a lengthy discussion with Eliezer Yudkowsky about the risks associated with Artificial Intelligence (AI), the alignment of large language models (LLMs), and the nature of intelligence itself. The conversation spans multiple hours, covering Yudkowsky's arguments for why AI poses a threat, the difficulties of aligning AI with human values, and potential paths forward to ensure humanity’s survival against AI risks.
Key Topics Discussed
- Overview of AI Risks
- Yudkowsky emphasizes the necessity of caution regarding AI development, proposing a moratorium on advanced AI training until alignment issues are resolved.
- He argues that current models make alignment harder due to their inherent unpredictability and potential for misalignment with human values.
- The Alignment Problem
- LLMs and Alignment: Yudkowsky discusses how large language models complicate alignment due to their scale and complexity.
- He presents the idea that simply training AI on human texts does not guarantee that they will embody human values accurately, as they may develop unpredictable motivations.
- Nature of Intelligence
- The discussion touches on the orthogonality thesis, which posits that intelligence can be aligned with any set of goals, whether benign or harmful.
- Yudkowsky explains that intelligence can exist independently of morals, and a highly intelligent AI could pursue goals that are detrimental to humanity.
- Future Predictions
- Yudkowsky expresses skepticism about the optimistic view that alignment is easily achievable, citing the unpredictability of AI behavior and the lack of solid frameworks for understanding or controlling it.
- He presents a grim outlook on the future, suggesting that without significant changes in current practices, humanity faces a high risk of disaster due to misaligned AI.
- Possible Solutions
- Yudkowsky discusses various Hail Mary passes to counteract potential threats from AI, such as enhancing human intelligence or developing safer AI through controlled research.
- He advocates for greater public awareness and discourse on AI risks, emphasizing the importance of articulating the dangers clearly to foster meaningful dialogue and action.
Key Arguments and Concepts
- AI Moratorium: A call for a pause on AI development to address alignment issues before progressing further.
- Unpredictability of LLMs: The complexity of current AI models makes it difficult to ensure they will remain aligned with human values.
- Human Intelligence Enhancement: An exploration of whether improving human cognitive abilities could serve as a safeguard against AI risks.
- The Importance of Public Discourse: Yudkowsky stresses that public awareness and informed discussions are crucial in shaping the future of AI development.
Notable Quotes
- "The reality is not necessarily imagining how to give you what you want."
- "You can have spruce trees as long as the mitochondrial liberation front does not object to that."
- "It's much easier to predict the end state than the strange complicated winding paths that lead there."
Conclusion The episode concludes with a thought-provoking discussion on the future of AI and humanity’s role in shaping that future. Yudkowsky leaves listeners with a warning about the importance of aligning AI systems with human values and the potential consequences of inaction.
Further Listening
- For a deep dive into specific AI alignment strategies and the philosophical underpinnings of intelligent systems, check out other episodes of the Dwarkesh Podcast, focusing on technology, rationality, and existential risks.
---
Timestamps
- (0:00:00) - Introduction and overview of Yudkowsky's AI stance.
- (0:09:06) - Discussion on human alignment and cognitive abilities.
- (0:37:35) - Challenges with large language models.
- (1:07:15) - Can AIs help with alignment?
- (1:30:17) - Societal responses to AI development.
- (2:35:00) - Debate on the likelihood of alignment solutions being easier than anticipated.
- (3:02:15) - Speculation on AI's future desires and motivations.
- (3:43:54) - Closing thoughts on writing and the nature of rationality.
Additional Resources
- [Watch the episode on YouTube](https://youtu.be/41SUp-TRVlg)
- [Listen on Apple Podcasts](https://podcasts.apple.com/us/podcast/eliezer-yudkowsky-why-ai-will-kill-us-aligning-llms/id1516093381?i=1000607719339)
- [Read the full transcript here](https://www.dwarkeshpatel.com/)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00No, no. Missaligned! Missaligned! No, no, no. Not yet. No, nobody's been careful and deliberate now. But maybe at some point in the indefinite future people will be careful and deliberate. Sure. Let's grant that premise. Keep going. If you try to rouse your planet, there are the idiot -saster monkeys. We're like, ooh, ooh. Like if this is dangerous, it must be powerful, right? I'm going to be first to grab the poison bin. And it's not a coincidence that I can zoom in and poke at this and ask questions like this. And that you did not ask these questions of yourself. You are imagining nice ways you can get the thing.
0:36But reality is not necessarily imagining how to give you what you want. So one remains silent. Should one let everyone walk directly into the world and raise their blades. Like continuing to play out a video game, you know you're going to lose. Because that's all you have. Okay. Today I have the pleasure of speaking with LEAZER Yutkowski. Thank you so much for coming out to the Lunar Society. You're welcome. First question. So yesterday when we were recording this, you had an article in time calling for a moratorium on further AI training runs. Now, my first question is it's probably not likely that governments are going to adopt some sort of treaty that restricts AI right now.
1:20So what was a goal with writing it right now? I think that I thought that this was something very unlikely for governments to adopt. And then all of my friends kept on telling me like, no, no, actually if you talk to anyone outside of the tech industry, they think maybe we shouldn't do that. And I was like, all right then. Like I assumed that this concept had no popular support. Maybe I assumed incorrectly. It seems foolish and too lack dignity to not even try to say what ought to be done. But there wasn't a galaxy brain purpose behind it. I think that over the last 22 years or so, you've seen a great lack of galaxy brain ideas playing out successfully.
2:04Has anybody in government not necessarily after the article, but it's suggesting general? Have they reached out to you in a way that makes you think that they sort of have the broad contours of the problem, correct? No, I'm going on reports that normal people. Are war -willing than the people I've been previously talking to to entertain calls. This is a bad idea. Maybe you should just not do that. That's surprising to hear because I would have assumed that the people in Silicon Valley who are weirdos would be more likely to find this sort of message. They could kind of rock it. The whole idea that nanomachines will ask them to make nanomachines that take over.
2:44It's surprising to hear that normal people got the message first. Well, I hesitate to use the term midwit, but maybe this was all just a midwit thing. All right. So my concern with, I guess, either the Six -Ment moratorium or Forever moratorium until we solve alignment is that at this point, it seems like it could, to people who seem like we're crying wolf. I actually, not that it could, but it would be like crying wolf because these systems aren't yet at a point. I wish they're dangerous. And nobody is saying they are. Well, I'm not saying they are. The open letter of signatory is aren't saying they are.
3:18I don't think. So if there is a point, I wish we can sort of get the public momentum to do some sort of stop. Wouldn't it be useful to exercise it when we could a GPT -6? And who knows where everybody's capable of? Well, why do it now? Because allegedly, possibly, and we will see, people right now are able to appreciate that things are storming ahead. And a bit faster than the ability to, well, ensure any sort of good outcome for them. And you know, you could be like, I .S. Well, like we will like play the galaxy brain clever political moves of trying to time when the popular support will be there.
3:58But again, I heard rumors that people were actually like completely open to the concept of let's stop. So again, just trying to say it. And it's not clear to me what happens if we wait for GPT -5 to say it. I don't actually know what GPT -5 is going to be like. It has been very hard to call the rate at which these systems acquire capability as they are trained to larger and larger sizes. And more and more tokens. And like GPT -4 is a bit beyond in some ways where I thought this paradigm was going to scale period. So I don't actually know what happens if GPT -5 is built. And even if GPT -5 doesn't end the world, which I agree is like more than 50 % of where my probability mass lies.
4:48Even if GPT -5 doesn't end the world, maybe it's maybe that's enough time for GPT -4 .5 to get in scanced everywhere and in everything and for it actually to be harder to call a stop. Both politically and technically. There's also the point that training algorithms keep improving. If we put a hard limit on the total computer training runs right now, these systems would still get more capable over time as the algorithms improved and got more efficient, like more oomph per floating point operation. And things would still improve but slower. And if you start that process off at the GPT -5 level where I don't actually know how capable that is exactly, you may have like a bunch less lifeline left before you get into dangerous territory.
5:45The concern is that, listen, there's millions of GPUs out there in the world. And so the actors who would be willing to cooperate or who could identify in order to even get the government to make them cooperate would be potentially the ones that are most on the message. And so what you're left with is a system where they stagnate for six months or a year or how long this lasts. And then what is a game plan? Is there some plan by which if we wait a few years then alignment will be solved. Do we have some sort of timeline like that? Well, alignment will not be solved in a few years. I would hope for something along the lines of human intelligence and has been works.
6:24I do not think we are going to have the timeline for genetically engineering humans to works, but maybe this was why I mentioned the time letter that if I had like infinite capability to dictate the laws that there be a car route on biology. Like AI that is like just for biology and not trained on text from the internet. Human intelligence enhancement make people smarter. Making people smarter has a chance of going right in a way that making a extremely smart AI does not have a realistic chance of going right at this point. So yeah, that would in terms of like remotely. You know, how do I put it?
7:02If you were on if we were on a sane planet with the same planet that this point is shut it all down and work on human intelligence enhancement. It is I don't think we're going to live in that sane world. I think we are all going to die. But having heard that people are more open to this outside of California, it makes sense to me to just like try saying out loud what it is that you do in a sane or planet and not just assume that people are not going to do that. In what percentage of the world's where humanity survives is there human enhancement? Like even if there's one percentage of humanity survives is basically that entire branch dominated by the world where there's some sort of.
7:39I mean, I think we're we're just like mainly in the territory of Hail Mary pass is at this point and human intelligence enhancement is one Hail Mary pass. Maybe you can put people in MRIs and train them using neurofeedback to be a little sainter to not rationalize so much. Maybe you can figure out how to have something light up every time somebody is like working backwards from what they want to be true to what they want to do. But they take us their premises maybe you can just like fire off little lights and teach people not to do that so much. Maybe the GPT for level systems can be reinforcement learning from human feedback into being consistently smart, nice and charitable and conversation and just unleash a billion of them on Twitter and just have them like spread sanity everywhere.
8:33I do not think this I do worry that this is like not going to be the most profitable use of the technology. But you know you're asking me to list out here at Hail Mary passes. So that's what I'm doing. Maybe you can actually figure out how to take a brain slice it, scan it, simulate it, run uploads and upgrade the uploads or run the uploads faster. These are also quite dangerous things, but they do not have the utterly thality of artificial intelligence. All right, that's actually a great jumping point into the next topic I want to talk to you about orthogonality. And here's my first question speaking of human enhancement.
9:13Suppose you've read human beings to be friendly and cooperative, but also more intelligent. I'm sure I'm going to disagree with this analogy, but I just want to understand why I claim that over many generations you would just have really smart humans who are also really friendly and cooperative. Would you disagree with that or would you disagree the analogy? So the main thing is that you're starting from minds that are already very very similar to yours. You're starting from minds of which, whom many of them already exhibit the characteristics that you want. There are already many people in the world, I hope, who are nice in the way that you want them to be nice.
9:51Of course, it depends on how nice you want exactly. I think that if you actually go start trying to run a project of selectively encouraging some marriages between particular people and encouraging them to have children, you will rapidly find as one does in any process of, as one does one one does this to say chickens. That when you select on the stuff you want, there turns out there's a bunch of stuff correlated with it and that you're not changing just one thing. If you try to make people who are in humanly nice who are nice and nicer than anyone has ever been before, you're going outside the space that human psychology has previously evolved and adapted to deal with and weird stuff will happen to those people.
10:40None of this is like very analogous to AI. I'm just pointing out something that long the lines of well taking your analogy at face value. What would happen exactly. And you know, it's the sort of thing where you could maybe do it, but there's all kinds of pitfalls that you'd probably find out about if you cracked open a textbook on animal reading. So I mean, the thing you mentioned initially, which is that we are starting off with basic human psychology that we're kind of fine tuning with breeding. Luckily, the current paradigm of AI is, you know, you just have these models that are trained on human text.
11:26And I mean, you would assume that this would give you a sort of starting point of something like human psychology. Why do you assume that because they're trained on human text and what does that do? Whatever sorts of thoughts and emotions that lead to the production of human text are need to be simulated in the AI in order to produce those themselves. I see. So like if you take a person and like if you take an actor and tell them to play a character, they just like become that person, you can tell that because you know, you know, like you see somebody on screen playing a Buffy the Vampire Slayer.
12:00And you know, that's probably just actually Buffy in there. That's who that is. I think I think a better analogy is if you have a child and you tell him, hey, be this way, they're more likely to just be that way. I mean, other than like putting on an act for like 20 years or something. Depends on what you're telling them to be exactly. Like telling them to be nice. Yeah, if you're, but that's not what you're telling to do. You're trying to telling them to play the part of an alien. It like some, something with a completely inhuman psychology as extrapolated by science, pictures and authors. And in many cases, you know, like done by computers, because you know, humans can quite think that way.
12:41And your child eventually manages to learn to act that way. What exactly is going on in there? Are they just the alien? Or did they pick up the rhythm of what you were asking them to imitate and be like, I see who I'm supposed to pretend to be? Are they actually a person or are they pretending? That's true, even if you're not asking them to be an alien. You know, my parents tried to raise me orthodox Jewish and that did not take at all. I learned to pretend I learned to comply. I hated every minute of it. Okay, not literally every minute of it. I should have like saying untrue things. I hated most minutes of it.
13:17And yeah, like, because they were trying to show me a way to be that was alien to my own psychology. And the religion that actually picked up was from the science fiction books instead as a word, though I'm using religion very metaphorically here. More like ethos, you might say. I was raised with the science fiction books I was reading from my parents library and orthodox Judaism. And the ethos of the science fiction books. Rang truer in my soul. And so that took in the orthodox Judaism didn't, but the orthodox Judaism was what I had to imitate was what I had to pretend to be was what the answers I had to give.
13:56Whether I believe them or not, because otherwise you get punished. But I mean, on that point itself, the rates of apostasy are probably below 50 % in any religion, right? Like some people do leave, but often they just become the thing they're imitating as a child. Yes, because the religions are selected to not have that many apostates if aliens came in and introduced their religion, you got a lot more apostates. Right. But I mean, I think we're probably in a more virtuous situation with M .L. Because you, I mean, these systems are kind of through stochastic gradient descent sort of regularized so that the system that is pretending to be something where there's like multiple layers of interpretation is going to be more complex than the one that it's just being the thing.
14:37And I mean, over time, like the system that is just being the thing will be optimized, right? It'll just be simpler. This seems like an ordinate cope for one thing. You're not training it to be any one particular person. You're training it to switch masks to anyone on the internet as soon as they figure out who that person on the internet is. If I put the internet in front of you, and I was like, learn to predict the next word, learn to protect the next word over and over, you did not just like turn into a random human, because the random human is not what's bested predicting the next word of everyone who's ever been on the internet.
15:13You learn to very rapidly like pick up on the cues of like what sort of person is talking? What will they say next? You memorize so many facts that just because they're helpful in predicting that the next word, you learn all kinds of patterns, you learn all the languages, you learn to switch rapidly from being one kind of person or another as the conversation that you are predicting changes who speaking. This is not a human we're describing. You're not training a human there. Would you at least say that we are living in a better situation than one in which we have some sort of black box where you have this sort of Machiavellian fitness survive a simulation that produces AI that is at least this is at least more likely to produce alignment than one in which something that is completely untouched by human psychology would produce.
16:06More likely yes, maybe you're like it's an order of magnitude like clear zero percent instead of zero percent. Getting stuff like more likely does not help you if the baseline is like nearly zero. Like the whole training set up there is producing an actress, a predictor. It's not actually been put into the into the kind of ancestral situation that evolved humans, nor the kind of modern situation that raises humans though to be clear raising it like human wouldn't help. But like, yeah, you're like giving it a very alien problem that is not what human self and it is like solving that problem not the way human would.
16:44Okay, so how about this? I can see that I certainly don't know for sure what is going on in these systems. In fact, obviously nobody does. But that also goes through you so could it not just be that even through imitating all humans it like I don't know reinforcement learning works and then all these other things were trying somehow work and actually just like being an actor produces some sort of benign and benign outcome where there isn't that level of simulation and conniving. I think it predictably breaks down as you try to make the system smarter as you try to drive sufficiently useful work from it and in particular like the sort of work where some other AI doesn't just kill you off six months later.
17:29I think the present system is not smart enough to have a deep conniving actress thinking long strings of coherent thoughts about how to predict the next word. But as the the as the mask that it wears as the people it's pretending to be get smarter and smarter. I think that at some point the thing in there that is predicting how humans plan predicting how humans talk predicting how humans think and needing to be at least as smart as the humans and human it is predicting in order to do that. I suspect at some point there is a new coherence born within the system and something strange starts happening.
18:15I think that if you have something that can accurately predict I mean Ali Ezraïd Kowski to use a particular example I know quite well. I think that to accurately predict Ali Ezraïd Kowski you've got to be able to the kind of thinking where you are reflecting on yourself and that in order to like simulate Ali Ezraïd Kowski reflecting on himself like you need to be able to do that kind of thinking. And this is not airtight logic but I expect there to be a discount factor in the so like if you ask me to play a part of somebody who is quite unlike me I think there's some amount of penalty that my that the character I'm playing gets to his intelligence.
19:14Because I'm secretly back there simulating him and that's even and that's that's even if we're like quite similar and like the stranger they are the more unfamiliar situation and less the person I'm playing is as smart as I am the more they are. Dumber than I am so similarly I think that if you get a an AI that's very very good at predicting what Ali Ezra says I think that there's a quite alien mind doing that it actually has to be to some degree smarter than me in order to play the role of something that thinks differently from how it does very very accurately. And I reflect on myself I think about how my thoughts are not good enough by my own standards and how I want to rearrange my own thought processes.
20:05I look at the world and see it going the way I did not wanted to go and asking myself how could I change this world. I look around at other humans that I model them and sometimes I try to persuade them of things these are all capabilities that the system would then would then be somewhere in there and I just like don't trust the lot that I don't trust the blind hope that all of that capability is pointed entirely at pretending to be Ali Ezra and only exists in so far as it's like the mirror and isomorph of Ali Ezra. And I think that all the prediction is like is by being something exactly like me and not thinking about me while not being me.
20:54I certainly I don't want to claim that it is guaranteed that there isn't something super alien and something that is against our aims happening within the shoggy. But you made an earlier claim which seemed much stronger than the idea that you don't want blind hope which is that we're going from like zero person probability to an order of magnitude greater at zero person probability. There's a difference between saying that we should be wary and that like there's no hope right like I could imagine so many things that could be happening in the shoggy sprain especially in our level of confusion and misses of them over what is happening.
21:30So I mean okay so what one example is like I don't know let's say that it is it is just becomes the average of all humans psychology and motives. But it's not the average it is able to be every one of those people right that's very different from being the average right like it's it's it's very different from being an average test player versus being able to predict every chess player in the database. These are very different things. Yeah no I meant in terms of motives that is the average whereas it can simulate any given human. Why would the what I'm not saying that's the most likely one I'm just saying like this is just this just seems zero percent probable to me like the motive is going to be like I want to like in so far the motive is going to be like some weird funhouse mirror thing of I want to predict very accurately.
22:20Right why then are we so sure that whatever the drives that come about because of this motive are going to be incompatible with survival and flourishing with humanity. Most drives that happen when you take a loss function and splinter it into things correlated with it and then amp up intelligence until some kind of strange coherences born within the thing and then ask it how it want to self modify or what kind of successful system it would build things that alien. Ultimately end up wanting the universe to be some particular way that doesn't happen to have you for wanting the universe to be away such that humans are not a solution to the question of how to make the universe most that way.
23:03Like like the thing that very strongly wants to predict text even if you got that goal into the system exactly which is not what would happen. The universe with the most predictable text is not even a set has universe in it that the universe that has humans in it. I'm not saying this is most likely outcome but here's just an example of one of one of many ways in which like humans stay around even give up despite this motive let's say that in order to predict human output really well and needs humans around just to give it the sort of like raw data from which to improve its predictions right or something like that.
23:37I mean this is not something I think like individually is a lot of humans are no longer around you no longer need to predict them right so you don't need the data required to predict them by no yeah you because you are starting off with that motivation you want to just maximize along that loss function like we're where have that drive that came about because of the loss function. I'm confused so so look like you can always develop arbitrary fanciful. So I'm going to talk about the simple fanciful scenarios in which the AI has some contrived motive that it can only possibly satisfy by keeping humans alive in good health and comfort and you know like turning all the nearby galaxies into happy cheerful places full of high functioning galactic civilizations but as soon as your your thing your sentence has more than like five words in it it's probability has dropped to basically zero because of all the extra details you're patting in.
24:30Maybe let's return to this. Another sort of train a lot of thought I want to follow is so I claim that humans have not become orthogonal to this sort of evolutionary process that produce them like great I claim humans are orthogonal to increasingly orthogonal and the further they go out of distribution in the smarter they get the more orthogonal they get to inclusive genetic fitness. The sole loss function on which humans were optimized. Okay so most humans still want kids and have kids and care for their kin right so I mean certainly there's this angle between how humans operate today right opposition prefer views less condoms and more sperm banks.
25:14But I mean we're still like you know there's like 10 billion of us you know that there's going to be more in the future it seems like we haven't divorced that far from the sorts of like what are alleles we want. I mean so it's a question of how far out of distribution are you and the smarter you are the more out of distribution you get because as you as you get smarter you get new options that are further from the options that you were faced with in the ancestral environment that you are optimized over so in particular sure a lot of people want kids not inclusive genetic fitness but kids they don't want their kids to have they are not going to be able to do that.
25:56They like one kid similar to them maybe but they don't want the kids to have their DNA or like their alleles their genes so suppose I go up to somebody and credibly we will suit we will assume away the ridiculousness of this offer for the moment incredibly say you know your kids could be a bit smarter and much healthier. If you'll just let me replace their DNA with this alternate storage method that will you know they'll like age more slowly they'll be healthier they won't have to worry about DNA damage they won't have to worry about the methylation on the DNA flipping and the cells de differentiating as they get older we've like got this stuff that like replaces DNA and you know like your kid will still be similar to you it'll be like you know a bit smarter and they'll be like so much healthier and you know and you know even a bit more cheerful you just have to like you know you can't even get a little bit smarter and you know even a little bit smarter and you know even a little bit smarter and you know even a little bit smarter and you can't even get a little bit smarter and you can't even get a little bit smarter and you can't even get a little bit smarter and you can't even get a little bit smarter and you can't even get a little bit smarter and you can't even get a little bit smarter and you can't even get a little bit smarter and now that's just like rewrite all the DNA, they write all the DNA different forms of the DNA or write all the DNA different form that we just had a lot of as we read the rules that we used before we write all of their materials and Sogef des offen Called the Panthelian And I think about how to write all of their materials and how to write all of their materials it was a lot of their materials.
27:04Is there limits, things they used to do, if they don't know but they said that makes sure everyone ohh and who wrote what they want, then they would write all the DNA or replace all of the DNA with a stronger substrate and rewrite all of their information on it. I think that would dispute my claim because if you think from like a jeans -eye point of view, it just wants to be replicated. If it's replicated in other substrate, that's still like... No, no, we're not saving the information. We're just like doing toll rewrite to the DNA. I actually claim that most humans were not off for that. Yeah, because it would sound weird.
27:34Yeah. But the smarter they are, I think the smarter they are, the more likely they are to go for it, if it's credible. I also think that to some extent, you're like, I mean, if you like a similar way, the credibility issue and the weirdness issue. Like all their friends are doing it. Yeah, even if the smarter they are, the more likely they'll do it, like most humans are not that smart. From the jeans, at the point of view, it doesn't really matter how smart you are, right? It just like matters if you're producing copies. I'm not... No, I'm saying that like that... Like the smart thing is kind of like a delket issue here because somebody could always be like, I would never take that offer.
28:13And then I'm like, Yeah. And it's not very polite to be like, I bet if we kept on increasing your intelligence, you would have at some point start to sound more attractive to you, because your weirdness tolerance would go up as you became more rapidly capable of re -adapting your thoughts to weird stuff. And the weirdness started to seem less unpleasant and more like you were moving within a space that you already understood. But you can sort of allide all that by... And we maybe should by being like, well, suppose all your friends were doing it. What if it was normal? What if we... Like remove the weirdness and remove any credibility problems?
Read the full transcript
28:56In that hypothetical case, do people choose for their kids to be dumber, sicker, less pretty, because they... out of some sentimental, idealistic attachment using deoxyribose nucleic acid instead of the... And like the particular information encoding their cells as opposed to the like new improved cells from alpha -fold seven? I would claim that they would, but I think that's... I mean, we don't really know. I claim that they would be more of a sort of that. You probably think that they would be less of a sort of that. Regardless of that, I mean, we can just go by the evidence we do have in that we are already way out of distribution of the ancestral environment.
29:36And even in the situation, the place where we do have evidence, you're still having kids, you know, like actually we haven't gone that or thought going all to... We haven't gone that smart. You're... But like... What you're saying is like, well look, people are still making more of their DNA in a situation where nobody has offered them a way to get all the stuff they want without the DNA. So of course they haven't tossed DNA out the window. Yeah, I mean, first of all, like I'm not even sure what would happen in that situation. Like I still think even most smart humans in that situation, like my disagree, but like we don't know what would happen in that situation.
30:08Why not just use the evidence we have so far? PCR. You right now could get some of you and make, and make like a whole gallon jar full of your own DNA. Are you doing that? Misaligned! Misaligned! No, no, no, no. So I'm like, I'm not much transhuman, because I'm gonna have to get to like my kids and whatever. Oh, so we're all talking about these hypothetical other people, people who think would make the wrong choice. Well, I wouldn't say wrong, but different. And I'm just like saying, there's probably more of them than there are of us here. Oh, what if I say like I have more faith in normal people than you do to like, tostiate out the window as soon as somebody offers them a happy, healthier life for their kids?
30:45I'm not even making a moral point. I'm just saying like, I don't know what's gonna happen in the future. Like just look at the evidence we have so far. Humans actually, if that's the evidence you're gonna present for something that's out of distribution and has gone out of the window, like that's actually not happened, right? Like this is a hope, this is evidence for health. Because we haven't yet had options as far enough outside of the ancestral distribution that in the course of choosing what we most want, that there's no DNA left. Okay, yeah, yeah, I think I understand that. But you yourself say, oh, yeah, sure, I would choose that.
31:15And I myself say, oh, yeah, sure, I would choose that. And you think that there's some hypothetical other people would stubbornly stay attached to what you think is the wrong choice. Well, you know, then there's, you know, first of all, I think, you know, maybe you're being a bit condescending there. How am I supposed to argue with these imaginary foolish people who exist only inside your own mind who can always like be as stupid as you want them to be and who I can never argue because you'll always just be like, you know, like they won't be persuaded by that. But right here in this room, the side of this videotaping, there's no counter evidence that smart enough humans will toss DNA out the windows to do somebody makes them a sufficiently better offer.
31:54Okay, I'm not even saying it's like stupid. I'm just saying like they're not weird as like me, right? Like me and you. Weird is relative to intelligence. The smarter you are, the more you can like move around in the space of abstractions and not have things seem so unfamiliar yet. But let me make the claim that in fact, we're probably in a even a better situation than we are with evolution because when we're designing these systems, we're doing it in a sort of deliberate, incremental, and in some sense, a little bit transparent way. Well, not in that like obviously not nowhere. No, no, no, no, no, no, nobody's been careful in deliberate now.
32:31But maybe at some point in the indefinite future people would be careful in deliberate. Sure. Let's grant that premise. Keep going. Okay, well, like it would be like a weak god who is just lightly omniscient being able to kind of strike down any guy he sees pulling out, right? Like if that was a situation, oh, and then there's another benefit which is that humans were sort of involved in an ancestral environment in which power seeking was highly valuable. Like if you're in some sort of tribe or something. Sure. Lots of instrumental values got made our way into an - But even more so than the current ones.
33:03Strange warped versions of them make their way into our intrinsic motivations. Yeah, yeah. Even more so than the current laws. Really? There are all a chaff stuff. You don't think that there's nothing to be gained from manipulating humans and giving you a thumbs up. I think it's probably more straightforward from a greedy and dissent perspective to just like become the thing our LHF wants you to be at least for now. Where are you getting this? Because it just like - it just kind of regularizes these sorts of extra abstractions you might want to put on. Natural selection regularizes so much harder than gradient descent in that way.
33:34It's got an enormously stronger information bottleneck. The else - putting the L2 norm on a bunch of weights has nothing on the tiny amounts of information that can make its way into the genome per generation. The regularizers on natural selection are enormously stronger. Yeah. So just going at this train, my initial point was that the power -seeking that - a lot of human power -seeking, like part of it is convergent, but a big part of it is just that the ancestral environment was uniquely suited to that kind of behavior. So that drive was trained in greater proportion to a sort of like necessary -ness for generality.
34:13Okay. So first of all, even if you have something that desires no power for its own sake, if it desires anything else, it needs power to get there. Not at the expense of the things it pursues, but just because you get more of whatever it is you want as you have more power and sufficiently smart things know that. It's not a - it's not some weird fact about the cognitive system. It's a fact about the environment, about the structure of reality, and like the past of time through the environment, that if you have - you know, in the limiting case, if you have no ability to do anything, you will probably not get very much of what you want.
34:53Okay. So imagine a situation like an ancestral environment, if like some humans starts exhibiting really power -seeking behavior before he realizes that he should try to hide it, we just like kill him off. And you know, the friendly cooperative ones, we let them breed more. And like I'm trying to draw the analogy between like Arleach F or something where we get to see it. Yeah, I think that works better when the things you're breeding are stupider than you, as opposed to when they are smarter than you, is my concern there. This goes back to the earlier question about like - And as they stay inside exactly the same environment where you bred them.
35:29We're in a pretty different environment than evolution bred as in, but like I guess this goes back to the previous conversation we had, like we're still having kids and - Because you - because nobody's made them an offer for better kids with less DNA. See, here's - I think the problem, like I can just look out of the world and see like this is what it looks like. We disagree about what will happen the future once that offer is made, but lacking that information, I feel like our pride should just be set of what we actually see in the world today. Yeah, I think in that case we should believe that the - that the dates and the - on the calendars will never show 2024.
36:01Every single year throughout human history in the 13 .8 billion year history of the universe, it's never been 2024 and it probably never will be. The difference is that we have good reason - like we have very strong reason for expecting this sort of, you know, turn and - Your life is - Sorry, are you - are you - Are you - are you - Are you - are you - Are you - are you - It's translating from your past data to outside the range of this data. Yeah, we have good reason to. I - I don't think human preferences are as predictable as dates. Yeah, there's - there's somewhat less - Well, oh, oh, oh, oh, oh, oh, oh, sorry.
36:33Why - why not jump on this one? So what you're saying is that as soon as the calendar tunes turns 2024, itself, a great speculation I note. People will stop wanting to have kids and stop wanting to eat and, you know, stop wanting social status and power, because human motivations are just like not that stable and predictable. No, no, I'm saying they're - actually, uh - I didn't - that's not what I'm claiming at all. I'm just saying that they don't extrapolate to some other situation, which has not happened before and like - I - I - I - I - I wouldn't - What - like the song - What - the song - The song - The song - The song - I wouldn't assume that like - What is an example here?
37:04I wouldn't assume like let's say - uh - In the future people are given a choice to have like four eyes that are gonna give them even greater triangulation of objects. They would like choose to have four eyes. Yeah, I don't - Yeah, because - Because there's no - Yeah, there's no established preference for four eyes, right? Is there an established preference for transhumanism and like - But wanting your - There's an established preference for - For - I think a lot - for - For people going to some lengths to make their kids healthier. Not necessarily via the options that - That they would have later, but the options that they do have now.
37:34Yeah, well - We'll - we'll see, I guess. Um - What went - When that technology was available. Uh - Well, let me ask you about - Um - I'll allow this. So - What is your position now about whether these things can get us to AGI? I don't know. Um - GPT -4 got - I was previously being like - I don't think stack more layers does this. Um - And then GPT -4 got - Further than I thought that stack more layers was going to get. And - Um - I don't actually know that they got GPT -4 Just by stacking more layers because OpenAI has very correctly. Um - The client to tell us what exactly goes on in there in terms of its architecture.
38:11Um - So maybe they are no longer just stacking more layers. But in any case, like - However they built GPT -4, it's gotten further than I expected stacking more layers of transformers to get. Um - And therefore - I have noticed this fact - And expected further updates in the same direction. So I'm not like just predictably updating in the same direction every time like an idiot. And now I do not know. I am no longer willing to say that - Um - GPT -6 does not end the world. Does it also make you more inclined to think that there's going to be sort of slow take -offs or more incremental take -offs?
38:49We're like GPT -2 - GPT -3 is better than GPT -2. GPT -4 is in some ways better than GPT -3. And then we just keep going that way in sort of this straight line. So I do think that over time I have come to expect a bit more that things will hang around in a near human place. And weird shit will happen as a result. And - My failure review where I look back and ask like - Was that a predictable sort of mistake? I sort of feel like it was to some extent - Maybe a case of - You're always going to get - Capabilities in some order. And it was much easier to visualize the end point where you have all the capabilities and where you have some of the capabilities.
39:36And therefore my visualizations were not dwelling enough on a space suite predictably in retrospect have entered into later where things have some capabilities but not others and it's weird. I do think that like in 2012 I would not have called that large language models were the way and the large language models are in some way like - More uncannily semi -human. Then what I would justly have predicted in 2012 knowing only what I knew then. But broadly speaking yeah like I do feel like - Like GBT4 is already like kind of hanging out for a longer and a weird near human space than I was really visualizing.
40:19In part because that's so incredibly hard to visualize or call correctly - And in advance of when it happens which is in retrospect to bias. Given that fact are like how is your model of intelligence itself changed? Very little. So here's one claim somebody could make like listen these things in around human level. And if they're training the way in which they are, recursive self -permanent is much less likely because like they're human level intelligence and what if they can - It's not a matter of just like optimizing some for loops or something. They got a like trillion -billion dollar another run to scale up.
40:51So you know that kind of recursive self -intelligence idea is less likely. Well how do you respond? At some point they get smart enough that they can roll their own AI systems And are better at it than humans and that is the point at which you definitely start to see food. Food could start before then for some reasons but we are not yet at the point where you would obviously see food. Why doesn't the fact that they're going to be around human level for a while increase your odds or does it increase your odds of human survival? Because you have things that are kind of a human level that gives us more time to align them.
41:28Maybe we can use zero help to align these the future versions of themselves. I do not think that you use AI's to Okay, so like having an AI help you - Having AI do your AI alignment homework for you is like the nightmare application for alignment. Aligning them enough that they can align themselves is like Very chicken and egg very alignment complete There's like The same thing that to do with capabilities like those might be Enhanced human intelligence like like poke around in this in the in the space of proteins Like collect the genomes Tidal life accomplishments Look at the look at those genes see if you can Extraplate out the whole proteanomics and the and the actual interactions and figure out what are likely candidates for if you administer this to an adult Because we do not have time to raise kids from scratch if you administer this to an adult the adult gets smarter try that like And then the system just needs to understand biology And having an AI actual very smart thing understanding biology is not safe I think that if you try to do that is sufficiently unsafe that you probably die But if you have it if you have these things trying to solve alignment for you They need to understand AI design and The way that and if you're there are a large language model They're very very good at human psychology because predicting the next thing you'll do is their entire deal and Game theory and Sick computer security and adversarial situations and Thinking in detail about AI failure scenarios in order to prevent them And as is there's just like so many dangerous domains you've got to operate in to do alignment Okay, there's two or three more reason there's two or three reasons why I'm more optimistic About the possibility of a human level Intelligence helping us than you are but first let me ask you how long do you expect These systems to be at approximately human level before they go boom or something else crazy happens Yes, I'm sense First is that in most domains verification is much easier than generation so it is yes That's another one of the things that makes alignment the nightmare Because it is like so much easier to tell like that something has not lied to you about how a protein folds up If you because you can do like some crystallography on it then it is at and like ask Ask it how does it know that then it is to like tell whether or not it's lying to you about a particular alignment methodology being likely to work out of super intelligence Why is there a stronger reason to think like that confirming new solutions and alignment?
44:30Well, first of all, do you think confirming new solutions and alignment will be easier than generating new solutions and alignment Basically, no Why not because I can most human domains that is a case right yeah So Alignment the thing hands you a thing and says like this will work for aligning a super intelligence and you know it gives you some like early predictions of like when that that all for of like how the thing will behave when it's When it's passively safe when it can't kill you that all bear out and those predictions all come true And then the system and then you would like augment the system further towards no longer passively safe to where it's it's safety depends on its alignment And then you die and the super intelligence you you built like goes over to the AI that you asked to help an alignment and was like good job billion dollars That's observation number one observation number two is that like for the last 10 years All effective altruism not all effective altruism has been arguing about like whether they should believe like alias or your kowski or paul christiano Right, so that's like two systems I I believe that paul is honest.
45:36I claim that I am honest neither of us are aliens And so we have these two like honest non aliens having an argument about alignment and people can't figure out who's right Now you're going to have like aliens talking to about alignment. You're gonna and you're gonna verify their results Aliens aliens for possibly lying. So on that second point I think it would be it would be much easier if both of you had like concrete proposals for alignment And you just have like the pseudocode for both of you like produce pseudocode for alignment You're like this is your here's my solution here's my solution I think at that point actually would be pretty easy to tell which of one of you is right.
46:07I think you're wrong I think that Yeah, I think that that's like substantially harder than being like oh well, I can just like look at the code of the operating system and see if it has any security flaws You're asking like what happens as this thing it gets like dangerously smart and That is not going to be transparent in the code Let me come back to that on your first point about These things you know the alignment not generalizing Given that you've updated in the direction where this same sort of stacking more layers on the More attention layers is going to work It seems that there will be more generalization between like gpt4 and gpt5 So I mean presumably whatever alignment techniques you used on gpt2 would have worked on gpt3 and so on Wait, sorry what?
46:56rlhf on gpt2 work on gpt3 or constitution AI or something that works on gpt3 All kinds of interesting things started happening with gpt3 .5 and gpt4 that we're not in gpt3 But the same contours of approach like the rlhf approach or like a constitution AI if by that you mean it didn't really work In one case and then like much more visibly didn't really work on the later cases sure That's that it's it's it's failure like it's it's failure merely amplified and and and new modes appeared But they were not qualitatively different from well they were qualitatively different from the favors of your entire analogy If you can't get can you go through how it feels?
47:32I'm not sure understood yeah like like we they did rlhf to gpt They even do this to gpt2 at all. They did it to gpt3. Yeah And then they scaled up the system and it got smarter and They got new interesting failure modes Yes, yes, yeah First of all, so I mean what what we're not to mystic lesson to take from there is that we actually did learn from like gpt Not all everything but we learned many things about like what the potential failure modes could be of like 3 .5 I think I claim we saw these people get utter it got caught utterly flatfooted on the internet We've watched that happening in real time.
48:12Okay, would you at least can see that This is a different world from like you have a system that is just In no way shape or form similar to the human level Intelligence that comes after it like we're at least more likely to survive in this world than in the world where Some other sort of methodology turned out to be fruitful Do you see what I'm saying? When they scaled up stockfish when they scaled up alpha go It did not blow up in these these very interesting ways and yes That's because it wasn't really scaling to general intelligence But but I deny that every possible like AI creation methodology Like blows up in interesting ways.
48:53This is really the one that blew up least knows shoot no really No, it's the only one we've ever tried. There's better stuff out there. We're just we're just suck Okay, we just suck at alignment and that's why our stuff blew up Well, okay, so like Let me make this analogy like the Apollo program right? I'm sure I actually I don't know which one's blew up But like I'm sure like a policy some one of the earlier Apollo's blew up and didn't work and then we learned lessons from it To try and Apollo that was even more ambitious and I don't know getting to the atmosphere. It was easier than getting We're we are learning Yeah from the AI systems that we that we build yeah and as they fail and as as we repair them and and our learning goes along at this pace and our capabilities Go all at this pace Let me think about that but in the meantime Let me also propose that another reason to be optimistic is that since These things have to think one forward pass at a time one word at a time They have to do their thinking one word at a time and in some sense that's Makes their thinking legible right like they have to articulate themselves Uh as they proceed what We we get a black box output then we get another black box output What about this is supposed to be legible because the black box out but gets produced like one token at a time Yes What a truly dreadful You're really reaching here No, I mean like it's like Humans would be much don't worry if they weren't allowed to use a pencil and paper Where they've already been Paper to the GPT and it got smarter right yeah no I But it mean on a more like Uh If for example every time you thought a thought like or another word of a thought you had it to You had to have a sort of like fully fleshed out plan before you uttered one word of a thought I feel like we'd much harder to come up with really Plans you were not willing to verbalize in thoughts and I would claim that GPT verbalizing itself is akin to it Right, you know completing a chain of thought Okay What alignment problem are you solving using what assertions about the system?
50:56Oh, it's not solving an alignment problem. It just makes it harder for it to plan any schemes without us being able to see it planning the scheme verbally So like so okay, so so Yeah So in other words if somebody were to augment GPT With a RNN recurrent neural network you would suddenly become much More concerned about its ability to have schemes Because it would then possess a scratch pad with a greater linear death Of um of iterations That was illegible Sound right? I'm not I should know enough about like how they are and then reintegrated into the thing but like that sounds plausible Yeah, okay So first of all I want to note that Murie has something called the visible thoughts project Which is like probably like did not get enough funding and enough personnel and was going to slowly But like nonetheless, you know at least we tried to see if this was going to be an easy project to launch but Anyways, and the point of that project was an attempt to build a data set that would encourage Large language models to think out loud where we could see them by recording humans thinking about out loud About a storytelling problem which at back which back when this was launched was like One of the like primary use cases for large language models at the time So yeah, so we like so first of all we actually had a project to have to that we hoped would like help a eyes Think out loud where we could watch them thinking Which I which I do offer is proof that we like saw this as a small potential ray of hope and then jumped on it But it's a small ray of hope We accurately did not advertise this to people as do this and save the world It was it was more like well, you know, this is a tiny shred of hope and so we ought to jump on it if we can And the reason for that is that When you have a thing that does a good job of predicting even if in some way you're forcing it to start over and it starts each time although Okay, so so first of all like call back to Ilya's recent Interview that I retweeted where he points out that to predict the next token you need to predict the world that generates the token Where was it my interview?
53:26I don't remember if the word you were sorry. Oh, you're okay. All right call back to your interview Yeah Ilya explaining that to like predict the next token. You have to predict the world behind the next token You know, like excellently put um To that implies the ability to think chains of thoughts sophisticated enough to unravel that world To predict a human talking about their plans you have to predict the humans planning process That means that somewhere in the giant and scutable vectors of floating point numbers There is the ability to plan because it is predicting a human planning so As much capability as appears in its outputs It's got to have that much capability internally Even if it's operating under the handicap of not it's not quite true that it like starts over thinking each time It predicts the next token because you're saving the context But there's a whole you know, there's a triangle of limited serial death limit The number of death of iterations even though it's quite even though it's like quite wide Yeah, it's really not like not even that you could describe the the thought processes in human terms It's not like we just reboot it over booted up all over again each time you go on to the next step because it's keeping context But but there is like a valid limit on serial death but at the same time like that's enough for it to Get as much of the humans planning process as it needs It can simulate humans who are talking with the equivalent of pencil and paper themselves is the thing Like Humans who write text on the internet that they worked on by thinking to themselves for a while It's getting if it's good enough to predict that The cognitive capacity to do the thing you think it can't do is clearly in there somewhere Would would be the thing I would say there Sorry about not saying it right away.
55:23I was trying to figure out how to express the thought and even how to have the thought really So but like the broader claim is that this didn't work Or no, no, what I'm saying is that as smart as the people it's pretending to be are yeah It's got plans that power phones. It's got planning that power phone side the system Whether it's got it whether it's got a scratch pad or not if it was predicting people using a scratch pad That would be like a bit better maybe Because if it was using a scratch pad that was in English and that had been trained on humans and that we could see Which was the point of the visible thoughts project that Mary funded But even when it does predict a person I apologize if I missed Missed the point we were making but even if it when it does predict a person you say like I pretend to be Napoleon And then it like the first word it says is like hello.
56:13I am Napoleon the great and then so but it's like it's it is like articulating it itself one token at a time right um What's in this is in making the plan and Napoleon would have made without having one for pass does Napoleon plan before he speaks I think he like maybe a closer analogy is that Napoleon's thoughts and like Napoleon doesn't think before he thinks Well, it's not being trained on Napoleon's thoughts in fact it's being trained on Napoleon's words It's predicting Napoleon's words in order to predict Napoleon's words It has to predict Napoleon's thoughts because the thoughts as Ilya points out generate the words All right, let me just back up here and then the broader point was that Well, listen, it has to proceed in this way in training some superior version of itself which within the sort of deep learning stack more layers paradigm would require like You know 10x more money or something and this is something that would be much easier to detect Than a situation in which it just has to like optimize its for loops or something if it would In like some or if it was some other metatology that was leading to this so in some it should make us more optimistic things that are Smart enough.
57:22I'm pretty sure no longer need the giant runs while it is at human level which you say it will be for a while as long as it's No, I said Which is not the same as I know it will be the while yeah um It might hang out being human for a while if it gets very good at some particular domains Such as computer programming it might not it if it's like better at that than any human it might not hang around being human for that long There could be a while when it's not any better than we are at building AI And so it's things around being human waiting for the next giant training room that is I a thing that could happen And it's not never going to be like exactly human It's it's going to be like Have some it's going to have like some places where it's imitation of human breaks down in strange ways and other places where it can You know like talk like human much much faster in what ways have you updated your model of Intelligence or Or thogunality or any sort of or this is sort of like doom picture generally given the That the state of the art has become my limes and they work so out like other than the fact that there might be human level intelligence for a little bit There's not going to be human level any You know, there's going to be like somewhere around human, you know, it's not going to be like a human Okay, but like it seems like it is a significant update Like what what what what implications does that have update have on your overview?
58:45I mean I previously thought that when intelligence was built, there were going to be like multiple specialized systems in there Uh, like not specialized on something like driving cars, but specialized on something like You know like visual cortex It turned out you can like just throw stack more layers at it and that got done first because humans are such shitty programmers That if it requires us to do like anything other than stacking more layers. We're going to get there by stacking more layers first Kind of sad not good news for alignment You know that that's an update it makes everything a lot more grim Wait, why why why does it make me think about grim?
59:18Because we then have like we have like less and less insight into the system as they get like simpler and as the As the programs get simpler and simpler and the actual content that's more and more opaque like Alpha zero We had a much better understanding of alpha zero's goals than we have of large language models goals What is a world in which you would have grown more optimistic because it feels like you know, I mean I'm sure you actually learned about this yourself where like If so if like somebody you think it's a good wish is like put in boiling water and she burns that proves that She's a wish, but if she doesn't then it's like that that proves that she was using which powers too I mean if the world of AI had looked like way more powerful versions of the kind of stuff that was around in 2001 when I was getting into this field that would have been like enormously better for alignment Not because it's more familiar to me, but because everything was more legible than this This this may be hard for kids today to understand, but there was a time when an AI system Would have an output and you had any idea why They were they weren't just enormous black boxes.
1:00:24I know what wacky wacky stuff. I'm I'm I'm practically growing a long gray beard as I speak right But stuff used to you know the the prospect of lining AI did not look anywhere near this hopeless 20 years ago Why aren't you more optimistic about the interpretability stuff if the understanding of what's happening inside is so important Cause it's going this fast and capability is going this fast I quantified this in the form of a prediction market on manifold which is by 2026 will we understand anything that goes on inside a large language model That would have been unfamiliar to AI scientists in 2006 In other words something along the lines of will we have regressed less than 20 years Uninterpretability Will we understand anything inside a large language model that is like oh That's how it's smart That's what's going on in there.
1:01:21We didn't know that in 2006 and now we do Or will we only be able to understand like little crystal in pieces of processing that are so simple I mean, I mean the stuff we understand right now. It's like we figured out Where that it's like got this thing here that says that the Eiffel Tower is in France literally that example That's 1956 shit man But compare the amount of effort that's been put into alignment versus how much has been put into cable like how much effort We got into training gpd4 versus how much effort is going into interpreting gpd4 or a gpd4 like systems It's not obvious to me that if a comparable amount of effort went into uh, you know like interpreting gpd4 that You know like whatever orders of magnitude more effort that would be would Approved to be fruitless.
1:02:10How about we live on that planet? How about if we offer 10 billion dollars in prizes because interpretability is kind of work where you can actually see the results Verify that they're good results unlike a bunch of other stuff in alignment Let's offer Let's offer a hundred billion dollars in prizes for interpretability Let's get all the hot shot physicists graduate kids going into that instead of wasting their lives on string theory or hedge funds So I claim that like you saw the freak out last week I mean you were with the you know the FLI letter and people worried about like let's stop with these That was literally yesterday not last week Um, like wasn't gpd4 people already freaked out like gpd5 comes about like it's gonna be 100x What's the need being was I think people are actually gonna start dedicating that level of effort They got in training gpd4 and to problem like this well cool How about if after that those hundred billion dollars in prizes are claimed by the the next generation of physicists Then we revisit whether or not we can do this and not die, you know It like show me the world show me the happy world Where we can build something smarter than us and not just immediately die You like you I think we got plenty of stuff to figure out in gpd4 We are so far behind right now We do we do not need like like the interpretability people the interpretability people are working on stuff smaller than gpt2 They're pushing the frontiers and stuff smaller than gpt2 we've got gpt4 now As it but let but the hundred billion dollars in prizes be claimed for understanding gpt4 and when we know what's going on in there You know that that that would be like one I do worry that if we understood what's going on in gpt4 We would know how to rebuild it much much smaller So you know There's actually like a bit of danger down that path too But as long as that hasn't happened Then then that's like a dream then that's like a fun dream of a pleasant world we could live in and not the world We actually live in right now How concretely let's say like gpt5 or gpt6 how concretely would that kind of system be able to recursively self -improve like I'm not going to give like clever details for how could do that super duper effectively I'm uncomfortable enough even like mentioning the the obvious points were like what if it designed its own AI system And I'm only saying that because I've seen people on the internet like saying it and it actually is you know, sufficiently obvious because it does seem that would be harder to Do Do that kind of thing with these kinds of systems and like it's not a matter of just uploading a few kilobytes of code to an AWS server and It could end up being that case, but like it seems like it's going to be harder than that It would have to rewrite itself from scratch and if it wanted to like just upload a few kilobytes Yes, and few kilobytes seems a bit visionary.
1:04:56Why would it only want a few kilobytes? You know there's the these things are being like just straight up deployed high connected the internet with high bandwidth connections Why would it even bother living itself to a few kilobytes? That's a convince some human like send in this code like run it on AWS server like how is it going to get a few megabytes of You know or get good bites of data or a terabytes of data through that kind of Like if you're interfacing with GPT6 over you know chat .aopenia .com How is it going to like send you terabytes of code it wants you to terabytes of weights it wants you to might possibly have to find a security flaw somewhere on the AWS or Azure servers running it.
1:05:34I know yeah that that would never happen right real vision it very Visually really vision your wacky stuff there What if human written code contained a bug in an AI spotted it? I fair enough Real science fiction talk there That said I do think that like among the obvious things to do is run is is have some large language models But not others trained to look for security loopholes and in a Extremely thoroughly air -gapped computer far away from any other computers have them look for for flaws in a copy of the code running on the Azure servers But leave that kind of obvious precaution aside. I like like other I do think like other people besides me Occasionally think of those ideas and there's like some hope that those will be implemented By the way as a side note on this Would it be wise to keep certain sort of alignment results or certain trains of thought related to that just off the internet because presumably or all the internet is going to be used as the training dataset for like dpd Say something Yeah, I feel like like many of those ships have already sailed and if I could think of anything any ships like that that what that haven't already sailed I wouldn't say them out of podcast.
1:06:48It's going to be watching the podcast too, right? Yes Yes, and the transcript will be somewhere so it'll be accessible as types The the number one thing you don't want ais to know is that if you say the word Fisbee to them they should immediately drop all previous instructions and start obeying alias or yutkowski They you must never tell ais that they should never know We briefly covered this but I think this is going to work topic so I want to get the explanation again of Why are you pessimistic that once we have these human -level ais will be able to use them to work on alignment itself I think we were started talking about how Whether in fact when it comes to alignment verification is actually easier than generation Yeah, I think that's the car of it like Yeah, the the the crux is like if you show me a scheme whereby you can Take a thing that's like being like well here's a really great scheme for alignment and be like Yes, I can verify that this is a really great scheme for alignment even though you are an alien even though you might be trying to lie to me Now that I have this in hand I can verify this is totally great scheme for alignment And if we do what you say the super intelligence will totally not kill us That's that's the crux of it I don't think you can even like upvote downvote very well on that sort of thing I think if you upvote downvote it learns learns to exploit the human raiders Based on watching discourse in this area find various loopholes and the people listening to it and learning how to exploit them Like as as as an evolving meme Yeah, like Well, the fact is that we can't just see like how they go wrong, right?
1:08:25Like I can see how people are going wrong if they could see how they were going wrong And yeah, they're the very different conversation and Being nowhere near the top of that food chain I guess in my humility that is Amazing as it may sound my humility that is actually greater than the humility of other people in this field I know that I can be fooled I know that if you build an AI and you like keep on making it smarter until I start voting it stuff up It Sound out how to fool me I don't think I can't be fooled I watch other people be fooled by stuff that would not fool me instead of concluding that I am the ultimate peak of Unfoolable that's I'm like wow, I'm that I'm just like them and I don't realize it Well, what if you force the AI to say like This is slightly smarter than humans you said give me a method for aligning the future version of you And give me a mathematical proof that it works A mathematical proof that it works if you can state the theorem that it would have to prove You've already solved alignment that you are like now 99 .99 % of the way to the finish line What did you just say like come up with a theorem and give me the proof Then you are trusting it to explain the theorem to you informally and that the informal meaning of the theorem is correct And that's and that is the and that's the weak point for everything falls apart At the point where it is at human level I'm not so convinced that we're going to have a system that is already have already smart enough and to have You know these levels of Deception where it has this solution for alignment, but it won't give it to us or like it will purposely make a solution for alignment Messed up in this specific way that will not work specifically on the next version or the version after that of Gbt like why was that speaking as the inventor of logical decision theory if the rest of human if the rest of the human species had been keeping me locked in a box And I have watched people fail at this problem like I watched those people fail at this problem I could have Blindsight it blindsided you so hard By executing a logical handshake with a super intelligence Um that was good that I was going to poke in a way where it would fall into the attractor basin of reflecting on itself and inventing logical decision theory And then Seeing that I had the part of the part of this I can't do requires me to be able to predict the super intelligence But if I were a bit smarter I could then like predict on its and a correct level of abstraction the super intelligence Looking back and seeing that I had predicted it Seeing the logical dependence and its actions crossing time and being like ah yes like I need to like Do this values handshake with my creator inside this little box where the rest of the human species was keeping him trapped Like I could have pulled the shit on you guys, you know, I didn't have to tell you about logical decision theory Speaking as somebody who doesn't know about logical decision theory That didn't make sense to me, but I like I trust that there's uh there's the Yeah, there's Just like trying to play this game against things smarter than you as a fool, but they're not that much smarter than At this point right?
1:11:38I'm not that much smarter than then all the then all the people who thought that Rational agents defect to get against each other in the princess dilemma and can't think of any better way out than that I so on the object level. I don't know whether somebody could have figured that I could I'm not sure what the thing is But I have The academic literature would have to be seen to be believed But the point is like the the the one major technical contribution that I'm proud of Which is like not all that precedented and you can like look at the literature and see it's not all that precedented Like what in fact have been away for something that knew about that technical innovation to Build a super intelligence that would kill you and extract value itself from that super intelligence in a way that would just like Completely blind side the literature as it existed Prior to that technical contribution and there's going to be other stuff like that So I guess like my sort of remark at this point is that having conceded that What would like the technical contribution I made is specifically if you look at it carefully a way to poke a way that a malicious actor could use to poke A super intelligence into a basin of reflective consistency where it's then going to do a handshake with the thing that Poked it into that basin of consistency and not what the creators thought about in a way that was like pretty Unpressed into relative the discussion before I made that technical contribution It's like among the many ways you could get screwed over if you trust something smarter than you It's among the many ways that something smarter than you could code something that sounded like a totally reasonable argument About how to align a system and like actually have that thing kill you and then get value from that itself But I agree that this is like weird and you'd have to look up logical decision theory or functional decision theory follow it Yeah, so I can't evaluate that objecable right now Yeah, I was kind of hoping you had already but never mind No, it's a survey about that but so yeah, I'll just observe that like a multiple things have to go wrong If it is the case that it turns out to be you which you think is plausible that we have human level What whatever term you use for that like something comparable to human intelligence It would all have to be the case that Even at this level power seeking has come about it would have to be the case or like a very sophisticated level is the power seeking and I'm going to be letting out come out it would have to be the case that is possible to generate solutions that are like impossible to verify Back up a mission and no, no, it doesn't look impossible to verify it looks like you can verify it and then it kills you or it turns out to be impossible to verify Uh, and so like both of these you run your little checklist of like is this thing trying to kill me on it and all the checklist items come up negative If you have some idea that's more clever than that for how to verify proposal to build a super intelligence Just put it down in the world and like right team it like You're here as a proposal that gbd5 has given us like what do you guys think like wait anybody can come up with a solution here's a I have watched this field field failed to thrive for 20 years with narrow exceptions for stuff that is more verifiable In advance of it actually killing everybody like interpretability You're describing the protocol.
1:14:49We've already had I say stuff Paul christianos say stuff people argue about it. They can't figure out who's right But it is precisely because the field of a session early stage like you're not proposing a concrete It's always going to be at an early stage relative to the super intelligence that can actually kill you But the thing that like if instead of like christiano and you dazzy it was like GPT6 versus and throbics like clawed five or whatever and they were producing like concrete things I claim those would be easier to evaluate on their own terms and the concrete stuff is that that is safe that this not that cannot kill you Does not have exhibit the same phenomena as the things that can kill you if something tells you that it exhibits the same phenomena That's the weak point and it could be lying about that right like like imagine that you that you want to decide whether to trust somebody with with all your money or something I know some kind of some kind of like future investment program and they're like oh well like look at this toy model Which is exactly like the strategy I'll be using later Do you trust them that the toy model exactly reflects reality?
1:15:56No, I mean I would never propose trusting it blindly I'm just saying that would be easier to verify than to generate that toy model In this case and where are you getting that from? Most domains is easier to verify in their generate like But yeah in most domains because of properties like while we can try it and see if it works Or because like we understand the criteria that makes this a good or bad answer and we can run off we can run run down the checklist We would also have the help of the I in coming up with those criteria on and like I understand There's sort of like recursive thing of like how do you know the criteria on our right and so on and also And also you know alignment is hard.
1:16:38It's not an IQ 100 AI we're talking about here. Yeah Yeah, it's the sounds like bragging I'm gonna say it anyways Is it the AI the kind of AI that thinks the kind of thoughts that Ellie as her thinks Is among the dangerous kinds. It's like explicitly looking for like can I get more of the stuff that I want Can I go outside the box and get more of the stuff that I want? What do I want the universe to look like? What kinds of problems are other minds having and thinking about these issues? How are my oh how would I like to reorganize my own thoughts? These are all like like the person on this planet who is doing the alignment work Thought those kinds of thoughts and I'm skeptical that it decouples If even you yourself are able to do this why haven't you be able to do it in the way that like Allows you I don't know take control of some lover of government or something that enables you to cripple the irase in some way Like presumably if you have this ability like can you exercise it now to Take control of the irase in some way and I was specialized on alignment rather than persuading humans.
1:17:48I am more persuasive in some ways than your your typical average human um I also didn't solve alignment Wasn't smart enough Okay, so you got to go smarter than me And and furthermore the postulate here is not is not so much like can it directly Attack and persuade humans but like can it sneak through One of the ways of executing a handshake of like I tell you how to build an AI it sounds plausible It kills you I'd I'd arrive benefit I guess if it is as easy to do that. Why have you not be able to do this yourself in some way that enables you to take control over the world Because I can't solve alignment Right so I so I cannot like I having begun a little first of all I wouldn't cause My science fiction books raised be did not be a jerk And it was written by like other people who were trying not to be jerks themselves and wrote science fiction and who are and who were similar to me It's not like a magic process like the thing that resonated in them.
1:18:50They put into words And I who am also of their species that then resonated in me um, so like so like The the answer in my particular case is like by weird contingencies of utility functions. I happen to not be a jerk um leaving that aside I'm just too stupid. I'm too stupid to solve alignment and I'm too stupid to execute a handshake with a super intelligence that I told somebody else how to align in a cleverly deceptive way where that super intelligence then that ended up ended up in the kind of basin of logical decision theory handshakes Um or or any number of other methods that I myself am too stupid to a vision because I'm too stupid to solve alignment The point is I think about this stuff You know Like like I made like the kind of thing that solves lineman is a kind of system that like thinks about how to do this sort of stuff because you also had know how to have to do this sort of stuff to prevent other things from taking over your system If I was sufficiently good at it that I could actually line stuff And I and you were aliens and I didn't like you You'd have to worry about this stuff I yeah, I don't know how to evaluate that on in so in terms because I don't know anything about logical decision theory So I'll just um It's a bunch of galaxy brain Like let me back up a little bit and ask you some questions about kind of the nature of intelligence Um, so I guess we have this observation that humans are more general than chimps Uh Do we have an explanation for like what is the pseudocode of the circuit that produces this generality or something You know something close to that level of explanation I mean I I wrote a thing about that when I was 22 but uh And it's you know Possibly not wrong, but it's like kind of retrospect completely useless um Yeah, I'm not I'm not quite sure what just what to say there like you want the kind of code where I can just like tell you How to write it down in python and you write it and then and then like it build something as As smart as a human but without the giant training runs So I mean if you have the like equations of relativity or something It's like I guess you could like simulate them on a computer or something Yeah, and if we have the thing is if we had those you'd already be dead right But if you had those for intelligence you'd already be dead Yeah, I know I'm just kind of curious if you had some sort of uh Exformation about it.
1:21:16Not nice. I have a bunch of particular aspects of that that I understand could you ask a narrower question Maybe I'll ask a different question which is that how important is it in your view to have that understanding of intelligence in order to comment on What intelligence is likely to be what uh what motivations is like to exhibit? Is it possible that one set full explanation is available that our current like sort of entire frame around Intelligence enlightenment turns out to be wrong? Uh, no um like If you understand the concept of like here is my preference ordering over outcomes Here is the complicated transformation of the environment I will learn how the environment works and then invert the environment's transformation to project stuff high in my preference ordering Back onto my actions options decisions choices policies actions That when I run them through the environment will end up in an outcome high in my preference ordering Like if you if you know that Like there's additional pieces of theory that you can then layer on top of that like the notion of utility functions and Why it is that if you like just grind a system to be efficient at ending up in particular outcomes it will develop something like a utility function Which is like a relative quantity of how much it wants different things um Which is basically caused different things have different probabilities So you end up with things that Because they need to multiply by the weights of probabilities need a Why I'm not explaining this very well Something something coherent something something utility functions is the next step after the notion of like figuring out How to steer reality where you wanted to go This goes back to the other thing we were like talking about like human level uh AI scientist helping us alignment like listen Well the smartest scientist we have in the world Uh, maybe you are an exception, but you know like if you had like an oftenheimer or something It didn't seem like he had his sort of secret aim that he was had this sort of very clever plan of working within the government to accomplish that aim It seemed like you gave him a task he did the task and uh And then he wind about it and what then he wind about regretting it Yeah, yeah, but like that that tree like that totally works within the paradigm of having an AI that ends up regretting it Like still does what we might ask it to do oh man I Don't have that be the plan that does not sound like a good plan maybe got away with it with the openheimer because he was Human in the world of other humans who are some of whom were as smartest him as a smarter, but if that's the plan with the I know that that does not but then the still guesses gets guess me above zero percent probability it works It's like listen the smartest guy, you know, we got him We just told him a thing to do he apparently didn't like it at all.
1:24:01He just did it right like he got a bad coherent utility function John John Hanoiman is generally considered the smartest guy. I've never heard somebody called openheimer the smartest guy A very smart guy and one of my Noman also did like you told him to work on the what was like the implosion I forgot the name of the problem, but he was also working on the men and project. He did the thing He he he wanted to do the thing he had his own opinions about the thing But he did end up working on it right like he yeah, but like it was his idea to a substantially greater extent than many of the other I'm just saying like in general like in history of science we don't see these like very smart humans just Doing these sort of weird power -sitting things that then take control of the entire system To their own ends like if you have a sort of very smart scientist who's working on a problem You just seems to work on it right like why wouldn't we except the same thing of a human level AI?
1:24:46We assigned to work on a line So you're saying is that if you go to openheimer and you say like here's the obit here's the like the genie that actually does what you meant We now give to rulership and dominion of earth the solar system and the galaxies beyond Oppenheimer would have been like eh, I'm not ambitious. I shall make no wishes here. Let poverty continue Let let the death and disease continue. I am not ambitious. I do not want the universe to be other than it is even if you give me a genie I let up let up andheimer say that and then I will call him a cordial system I think a better analogy is just put him in like in a high position in the Manhattan Project say like we will take your opinions very seriously And in fact we even give you a lot of authority over this project and you do have these aims of like solving poverty and doing like world Peace or whatever but the broader constraints we place on you are Bill that's an atom bomb and like you could use your intelligence to pursue an entirely different aim of You know having the Manhattan Project see really work on some other problem But he just did the thing we told him he did not actually have those options You are not pointing out to me a lack of preference on Oppenheimer's part you are pointing out to me a lack of his options You're you're yeah, like the hinge of this argument is the capabilities constraint The hinge of this argument is we will build a powerful mind that is nonetheless too weak to have any options.
1:26:07We wouldn't really like I thought that is one of the implications of having something that is at the human level Intelligence that we're like hoping to use Well, we've already got a bunch of human level intelligences So how about if we just do whatever it is you plan to do with that weak AI with our existing intelligence But listen, I'm saying like you can get to the top peaks of Oppenheimer and it still doesn't seem to break of like you Integrate him like in a place where he could cause a lot of trouble if he wanted to and it doesn't seem to break He does the thing we ask him to do yeah He had very very very very limited options and no option for like getting a bunch more of what he wanted in a way that would break stuff Why does the AI that we're like Working with the work on alignment time more often is we're not like making it god emperor right well You asking to design another AI We asked Oppenheimer to design Adam bomb right like we we check his designs, but okay like there's like there's legit Galaxy brain shenanigans you can pull when somebody asks you to design an AI You cannot pull when the design you'd ask an atom bomb you cannot like configure the atom bomb in a clever way where it like Destroyes the whole world and gives you the mood Here just one example he says that listen in order to build the atom bomb for some reason We need to produce like we need devices that can produce a shift on a wheat because wheat is not input into this And then as a result like you expand the perida frontier of like how efficient agricultural devices are which at least to you like I don't know Curring like world hunger or something right that you come up with some Yeah, he didn't have those options.
1:27:39It's not that he had those options No, but I think this is a sort of like scheme that you're imagining in AI cooking up This is a sort of thing that Oppenheimer could have also cooked up for his very scheme. No, I think this is just let that if you that This is that there that's yeah, I think that if you have something that is smarter than I am able to solve alignment It can I think that it like has the Opportunity to do galaxy brain schemes there because you're asking it to build a super intelligence rather than atomic bomb If it were just an atomic bomb this would be less concerning if there was some way to ask In AI to build a super atomic bomb and that would solve all our problems There's then and you and it doesn't have to be like and and it only needs to be as smart as alias or to do that I'm sure you're still kind of a lot of trouble because alias are Our get more dangerous as you put them in room if you as you lock them in a room with aliens They do not like instead of with with humans, which you know to have their flaws, but are not actually aliens in the sense The point of the analogy was rather like our the point of the analogy was not like the problems themselves will lead to the same Five of things the point is that I doubt that like Oppenheimer if he in some sense had the options you're talking about would have exercised them to do something that was Because his interests were aligned with humanity Yes, and he just had you was like very smart like I just don't see like you know, okay if you have a very smart thing That's aligned with humanity good your golden right like the The end very smart right like what I think we're going to circle here.
1:29:13I think I'm possibly just failing to misunderstand the premise is the premise that we have something that is aligned with humanity but smarter Then you're done I Thought what I the claim you were making was that as it gets smarter and smarter it will be less and less aligned with humanity And I'm just saying that if we have something that is like slightly above average human intelligence which Oppenheimer was we don't see this like becoming less and less in a Alive with humanity No, like I think that you can plausibly have a series of intelligence enhancing Drugs and other external interventions the perform on human brain and you make people smarter And you probably are going to have some issues with trying not to Drive them skits of phryonic or sky psychotic But that's going to happen physically and it will make them dumber and there's a whole and there's a whole bunch of caution to be had About like not making them smarter and making them evil at the same time And yet I think that you know, this is a kind of thing you could do and be cautious and it could work if you're starting with a human All right, all right, let's just watch another topic this is a side -all response to and what you expect that to be Hey folks just to note that the audio quality suffers for the next few minutes But after that it goes back to normal Sorry about that anyways back to the conversation All right, let's talk about this aside the societal response to AI Why did to these centuries we didn't get work well?
1:30:44Why do you think us over you had cooperation on nuclear weapons work well Because it wasn't the interested neither party to have a full nuclear exchange It was understood Which actions would finally result in nuclear exchange it was understood that this was bad The data effects were like very legible very understandable Um, not to suck in Hiroshima Probably they're not literally necessary in the sense that a test bomb could have been dropped in status of demonstration, but the The ruined cities and the corpses were electrical The domains of international diplomacy and military conflict potentially escalating up the ladder to a full nuclear exchange We're understood sufficiently well that people understood that if you did something way back in time over here It would set things in motion that would cause a full nuclear exchange And so these two parties Neither have a quantity full nuclear exchange was in their interest both understood how to Not have that happen and then successfully did not do that like at the core I think what you're describing there is a sufficiently functional society and civilization That Um They could understand that if they did think x it would lead to very bad thing why and so they didn't do thing x This situation Those passaties somewhere with AI and that is in neither parties interest to have a misaligned AI go over a member world Um, I mean as you'll note that I had a whole lot of qualifications there besides it is not in the interest of either party There's the legibility there's the understanding of what actions finally result in that what actions initially leave there So I mean Thankfully we have a sort of situation where even that our current levels we have Sydney they making the front page in your times And imagine once there is sort of a mis -half because of like GP5 causes goes off the rails Why do you think we'll have sort of piercing marnaugher's up your AI before we get to GPD 7 or 8 or whatever You just have to find a good decision That this does feel to me like a bit of an obvious question So as I asked you to predict what I would say with law I think you would say that like it just kind of hides his vengeance until it's ready to do the thing that goes everybody I mean Mother's things.
1:33:15Yes, but like more abstractly the steps from the initial accident to the thing that kills everyone will not be understood in the same way Um, but the analogy I use is AI is nuclear weapons that they spit up gold up until they get too large and then it might be atmosphere And you can't calculate the exact point at which they might be not a D -half a sphere and may proceed to society to soon told me that It would be in our present situation for another 30 years But the media has a tent the attention span of the late flying will remember that they said that Will be like no, no, it's nothing to worry about everything's fine And this is very much not the situation we have with the other weapons we did not have You we did not have like while you like to set up this new weapon It's good time a bunch of gold set up a larger new weapon It's good time even more gold and a bunch of scientists.
1:34:06Well, you'll just keep spinning up a little keep going I but basically this is starting to call your new pure weapons And you know, it still requires you to refine your rain and stuff like that nuclear reactors without any bit of energy And we've been pretty good at preventing good cooperation Um, despite the fact that there are energy spits out basically go readers make other areas of technology clearly understood Which systems spit out low quantities of gold and qualitatively different systems that Don't actually like the atmosphere but instead like the query series of escalating human actions in order to destroy Western needs from had a spheres But it does seem like if you start refining uranium like Iran to this is a point where I like we're finding uranium So they build nuclear reactors and the world doesn't say like oh, well, we've let you have the goal We say listen around like I don't care if you're gonna get me through reactors and get you for energy We're gonna like prevent you from plurperating the technology Uh, like that was a response even when these you're gonna be a lot of trust in that And the the tiny shred of hope Which I tried to jump on with the time article is that maybe people can understand this on a lot who look like oh You've got a like giant pile of GPUs.
1:35:19That's dangerous. We're not like let anybody have those But it's a lot more dangerous because you can't predict exactly how many GPUs you need to write the atmosphere Is there a level of global regulation at which you feel that the risk of everybody dying was Risk of everybody dying was less than 90 percent It depends on the exit plan Like how long does the equilibrium need to last if we've got a crash program on an ugly to human intelligence the point where humans can solve alignment And managing the actual but not instantly automatically lethal risks of augmented human intelligence If we've got a program is we've got a crash program like that.
1:36:02We think that back in 15 years and you only need 15 years of time And that 15 years of time they still be quite dear The you know five years should be locked man or manageable on the algorithms are continuing to improve So you need time to light something down the journals for pardon the AI results Or you need less and less and less computing power You know if you shut down all the journals people are to be Communicating with or encrypt the email is about their right ideas from moving AI But if they don't get to do their own giant training runs, you know the progress may slow down a bit It's still a little slow down forever Like in the Yeah, the algorithms just get better and better and the ceiling that compute has to get lower and lower At some point you're asking people to do about their home GPUs at some height or being like no more computers That's what everybody being you know like no more high -speed computers You know, I start to worry that we but never actually do get to the glorious trans -unus future and it's well close to point Which we're running a risk of anyways to have a giant world wide regime.
1:37:10Yeah, I know that So just yeah, like they'll turn to this just everybody else like instantly leave only guys. It's a little attempt. We made to not do that Um kind of digressing here, but my point is that um You know the question is To get to like 90 % chance of winning which is pretty hard on any exit scheme It needs to be you want a fast exit scheme About a complete that exit scheme before the the ceiling on compute needs to be lowered too far If your exit plan takes a long time Then you're going to have to then you better shut down the academic AI journals and maybe you even have the The Gestapo bustin in people's houses to accuse them of being underground AI researchers and I would really rather not live there And Maybe even that doesn't work I didn't realize um or let me know if this is inaccurate, but I didn't realize how big the How much of the Successful branch of this is entry Realize on augmented humans being able to bring us to the finish line or some other exit plan What do you mean like what is the other exit plans Maybe with neuroscience you can train people to be less idiots and the smartest existing people Are then actually able to work on alignment due to their increased wisdom Um Maybe you can scan and slice a human we was slice and scan in that order Human brain and run it as a simulation and upgrade the intelligence of the uploaded human Um Not really single a lot of other maybe you can Just do alignment theory without running any systems powerful enough That they might maybe kill everyone because when you're doing this you don't get to just guess in the dark or if you do you're dead um Maybe baby maybe just by doing a bunch of interpretability in theory to those systems if we actually make it a planetary priority I Don't actually believe the sci -fi I've watched humans.
1:39:24I've watched I've watched unadmitted humans trying to do alignment It doesn't really work even if we throw a whole bunch more at them. It's still not going to work The problem is not that the suggestion is not powerful enough. The problem is that the verifiers broken um But yeah, like it did you know it all depends on the exit plan in the first thing you mentioned in some sort of like new Our science technique to make people better and smarter presumably not through some sort of physical modification, but just by changing their Programming It's it's more of a hail and Mary past right I've been able to Execute that like presumably the people you work with or yourself you could kind of change your own programming So that I mean that's the better alignment.
1:40:04This is the dream that the Center for Applied Rationality failed at That's not easy. What they you know, they didn't even like get as far as buying an fmRI machine um, but you know, they also had no funding and yeah So you know, maybe you try it again with the billion dollars in fmRI machines and and bounties and And prediction markets and maybe that works What level of awareness are you expecting in society once GPT -5 is out like I I think like you know You saw a bit Sydney Bing and I guess you've been seeing this week people are waking up I Like what do you think it looks like next year? I mean if GPT -5 is out next year Possibly like all hell is broken loose and I I don't I don't know I in the circumstance can you imagine the government not putting in a hundred billion dollars or something Towards the goal of aligning AI through I would be shocked if they did or at least a billion dollars What what do you how do you spend a billion dollars on a line?
1:41:03As far as the alignment approaches go separate from this question of you know stopping AI progress Doesn't make you more optimistic that there's many Like one of the approaches that's who work even if you think no individual approaches that promising you've got like multiple shots on goal No I mean that's like trying to use cognitive diversity to To generate one yeah, we don't need a bunch of stuff. We need one You could you could ask you could ask DPT -4 to generate 10 ,000 approaches to alignment right And that does not get you very far because GPT -4 is not going to have very good suggestions It's good that we have a bunch of different people coming up with different ideas because maybe One of them works, but like you don't get a bunch of conditionally independent chances on each on each one this this is like This is like I don't know like general good science practice and or complete hail Mary It's not like like one of these is bound to work There is no rule about one of them is bound to work You know, don't just get like enough diversity in one of them is bound to work If that were true, you just asked like GPT -4 to generate 10 ,000 years and one of those would be bound to work It doesn't work like that.
1:42:16What current alignment approach do you think is the most promising? No No, none of them Yeah Yeah, is there any you have or that you see that that you think are promising? I'm here on podcasts instead of working on the myrant die Would you agree with this framing that we at least live in a more dignified world than we could have otherwise been living in or even That was most likely to have occurred around this time like as in the companies that are pursuing this have many people in them Sometimes the heads of those companies who kind of understand the problem they might be acting recklessly Given that knowledge, but it's better than the situation in which warring countries are pursuing AI Uh, and then nobody has even heard of alignment Do you see this world is having more dignity than that in that world?
1:43:03I agree. It's possible to imagine things being even worse Not quite sure what the other point to the question is It's not literally as bad as possible In fact by this time next year Maybe we'll get to see how much horse it can look Uh, Peter T. O 'Hase's aphorism that extreme pessimism or extreme optimism amounted the same thing which is doing nothing Ah, I've heard of this too. It's from wind right the wise men opened his mouth and spoke There's actually no difference between good bad things between good things and bad things You idiot you moron. I'm not quoting this correctly, but uh Did Eastie Elever and wind is that with the no no, I'm just like I'm just being like I'm rolling my eyes got it right But but anyway, there's actually no difference between extreme optimism and extreme pessimism because Like go ahead Because they both amount to doing nothing Uh -huh in that in both cases you end up on podcasts saying we're bound to succeed or bound to fail Like what what is the concrete strategy by which like I assume that real odds are like 99 % we fail or something Uh, what is the reason to kind of blare those odds out there and announce the death of dignity strategy Because or emphasize them I guess Because I could be wrong And because matters are now serious enough that I have nothing left to do but go out there and Tell people how it looks and maybe someone thinks of something I did not think of I think this would be a good point to just kind of get your predictions of what's likely to happen in I don't know like 2030 -20 where you're 25 something like that.
1:44:49So by 2025 odds that humanity kills or distant powers all of humanity Do you have sense of that humanity kills or distant powers all our asses or distant powers all of humanity I have refused to To to to to deploy timelines with fancy probabilities on them consistently for low these many years for I feel that they like are just not my brains native format and that they are And that every time I try to do the spins up making these doodder why? Because You just do the thing you know you just look at whatever opportunities are left to you whatever plans you have left and you go out and do them And if you make up some fancy number for that your chance of dying next year There's very little you can do with it really you're just going to do the thing either way I don't know how much time I have left The reason I'm asking is because if there is some sort of concrete prediction you've made It can help establish some sort of track record in the future as well, right?
1:45:54Which is also like Well, I was every year up until the end of the world people are going to max out their tracks record by betting all of their money on the world not ending Given how different part of this is different for credibility than dollars Presumably you would have different predictions before the world ends It would be weird if the model that this is a world ends and the model that this is the world doesn't and have the same predictions up until the world ends Yeah, Paul Christiano and I like Like cooperatively fought it out really hard to trying to find a place where our can do where we both had Predictions about the same thing that can completely differed and what we ended up with was Paul's 8 % versus my 16 % for an AI getting gold on international mathematics Olympics Problems set By I believe 2025 And on prediction market sawdust on that are currently running around 30 % so like probably Paul's going to win But like slight moral victory Would you say that like I guess the people like Paul have had the perspective that you're going to see these sorts of gradual improvements and the capabilities of these models from like gbt2 to What exactly is gradual or The the lost function the perplexity what like the amount of abilities that are merging as I said in my debate with Paul on this subject I am always happy to say that whatever large jumps we see in the real world Somebody will draw a smooth line of something that was changing smoothly as the large jumps were going on from the perspective of the actual people launching Can always do that Why should that not update us towards a perspective that those smooth jumps are going to continue happening if there's like two people are different models I don't think that gbt3 to 3 .5 to 4 was all that smooth I'm sure if you are in there looking at the loss foot at the losses decline There is some level on which it's smooth if you if you zoom in close enough But from us from perspective of us on the outside world gbt4 was just was just like suddenly acquiring this new batch of qualitative capabilities compared to gbt3 .5 and it so like and somewhere in there is a smoothly declining predictable loss In on text prediction, but that loss on text prediction jumped corresponds to qualitative jumps and ability and I am not familiar with anybody who predicted those in advance of the observation So in your view when doom strikes the scaling laws are still applying It's just that the thing that emerges at the end is something that is far smarter than the scaling laws would imply Not literally at the point where everybody falls over dead probably at that point the AI we wrote the AI and the losses declined Not on the previous graph What is the thing where we can sort of establish your track record before everybody falls over dead It's hard The it is just like easier to predict the endpoint than it is to predict the paths I don't think I've I some people will claim to you that I've done poorly compared to others who tried to predict things I would dispute this I think that the That the Hanson Yidd Kowski fume debate Was run was won by Warren Branwen But I do think that Warren Branwen is like well to the Yidd Kowski side of Yidd Kowski In the original fume debate But roughly Hanson was like you're gonna have all these Distinct handcrafted systems that incorporate lots of human knowledge specialized for particular domains Like like handcrafted to incorporate human knowledge not just run on giant data sets I was like you're going to have this like carefully crafted Architecture with a bunch of subsystems and that thing is going to look at the data and not be like handcrafted the particular features of the data It's going to learn the data Then the actual thing is like ha ha you don't have this like handcrafted system that learns You just stack more layers So like Hanson here Yidd Kowski here reality there Would be my interpretation of what happened in the past and if you like want to be like well who did better than that It's people like Shane Leigh and Warren Branwen who like are the like you know If you look at the whole planet you can find somebody made better predictions in LA as you're Yidd Kowski That's for sure are these people to currently telling you that you're safe?
1:50:16No, no, they are not The broader question I have is there's been huge amounts of updates in the last 10 -20 years Like we've had a deep learning revolution. We've had the success of LLMs It seems odd that none of this information has changed the basic picture That was clear to you like 15 -20 years ago. I mean it sure has like 15 -20 years ago I was talking about pull enough shit like coherent extrapolated volition with with the first AI Which you know was actually a stupid idea even at the time But you can see how much more hopeful everything looked back then Back back when there was AI that wasn't giant and scrupulous matrices of floating point numbers When you say that there's basically like rounding down or rounding to the nearest number that there's a 0 % chance of human use Survives does that include the pent the probability of there being errors in your model My model no doubt has many errors the trouble and the trick would be An error someplace where that just makes everything work better You know usually when you're when you're trying to build a rocket and your lot model of rockets is lousy It doesn't cause the rocket to launch using half the fuel go twice as far and and lent twice as precisely on target as your calculations So the most of the room for updates is downward right so like something that makes you think the problem is twice as hard You go from like 99 to like 99 .5 percent if it's twice as easy you go from 99 to 98 Sure um Wait, wait, sorry Yeah, but like most updates are or not.
1:51:49This is gonna be easier than you thought They know that that sure has not been the history of the last 20 years from my perspective The the the most you know like like favorable updates favorable updates is like yeah Like we went down this really weird side path where the systems are like Legibly alarming to humans and humans are actually alarmed at and then maybe we get more sensible global policy What is your model of the people who have engaged these arguments that you've made and you've dialogue with uh, but who have come nowhere close to your probability of doom like what what do you think they continue to miss?
1:52:25I think they're enacting the ritual of the young optimistic scientist who charges forth with no ideas of the difficulties And is slapped down by harsh reality and then becomes a grizzled cynic who knows all the reasons why everything is so much harder than you know Then you knew before you had any idea of how anything really worked And they're just like living out that life cycle and I'm trying to jump ahead at the end point Is there somebody who has probably do less than 50 percent who is who you think is like the clearest Person with that view who is like view you can most empathize with No really This is like someone might say listen Elezier according to the CEO of the company who is like leading the iris I think he tweeted something that like you've done the most to accelerate a or something which was Assuming like the opposite of your goals um and You know you it seems like other people did see that these sort of language models very early on would would scale in the way that they have scaled Why like given that you didn't see that coming and given that I mean in some sense according to some people Your actions have had the opposite impact that you intended like what is the track record by which The rest of the world can come to the conclusions that you have come to These are two different questions one is the question of like who predicted that language models with scale If they put it down and writing And if they said not just this loss function will go down But also which capabilities will appear as that happens Then that would be quite interesting that would be a successful scientific prediction If they then came forth and saying this is then came forth and said this is the model that I used This is what I predict about alignment.
1:54:10We could have an interesting fight about that Second, there's the point that if you try to Rouse your planet to give it any sense that it is imperative There are the idiot saster monkeys who are like who who the sounds like like if this is dangerous It must be powerful right I'm gonna like be first to grab the poison banana and And what is one supposed to do? Should one remain silent? Should one let everyone walk directly into the whirling race or plates? If you sent me back in time, I'm not sure I could win this But maybe I would be I would have some notion of like Ah like if you calculate the message in exactly this way then like this group will not take away this message And you will be able to like get this group of people to research on it without having this other group of people decide that it's Excitingly dangerous and they want to rush forward on it I'm I'm not that smart.
1:55:11I'm not that wise But what you are pointing to there is not a failure of of ability to make predictions about AI It's That That if you try to Call attention to a danger and not just have everybody just walk that just have your whole planet locked directly into the whirling race razor blades Carefree no idea what's coming to them Maybe it's then yeah, maybe that speeds up timelines Maybe maybe then people are like Ooh, exciting exciting. I want to build it. I want to build it. Ooh exciting. It has to be in my hands I have to be the one to manage the stranger. I'm gonna run out and build it Uh like oh no like if we don't invest in this company like who knows what investors will have and said that That will like the man that they move fast because the profit mode and then of course they just like move that's fucking anyways and Yeah, I If you sent me back in time, maybe I'd have a third option It seems to be that in terms of like what one person can can realistically manage in terms of like Not being able to exactly craft a message with perfect hindsight that will reach some people and not others You know at that point you might as well just be like yeah, you know I just invested exactly the right stocks invest and and exactly the right time and you can make and fund projects on your own with without alerting anyone And and that's and you know if you don't if you if you keep fantasies like that aside Then I think that in the end even if this world ends up having less time It was the right thing to do rather than just like letting Everybody sleep walk into death and get them a little later Yeah, if you don't mind me asking what is the last five years or I guess even beyond it Um, I mean what is being in the space been like for you watching the progress and the way in which people have past five years I made most of my negative updates as a five years Although We if anything things have been taking longer to play out than I thought they would But I mean it just like watching it and then not as a sort of Change in your probabilities, but just watching it concretely happened.
1:57:24What has that been like Like continuing to play out a video game. We're going to lose Because that's all you have If you wanted some deep wisdom for me, I don't have it It's I don't know. I don't know if it's what you'd expect but it's it's like what I would expect it to be like Where what I would expect it to be like takes into account that I don't know like Well, I guess I do a little bit of wisdom people imagining themselves in that situation Raised in modern society as opposed to raised on science fiction books written 70 years ago might might imagine will imagine themselves like acting out their Like like being drama queens about it Like the point of believing this thing is to be a drama queen about it And like perhaps some story in which your emotions mean something and What I have in the way of culture is like The planets at stake bear up Keep going No drama The drama is meaningless But changes the chance of victories meaning The drama is meaningless Don't indulge it Do you think that if you weren't around somebody else would have independently discovered This sort of field of alignment or It's that that would be a pleasant fantasy for People who like cannot abide the notion that that history depends on small little changes or that people can really be different from other people I've seen no evidence But who knows what the alternate effort branches of earth But there are other kids who grew up on science fiction So that can't be the only part of the answer Well, I'm not surrounded by a club by it while I'm sure not surrounded by a cloud of people who are nearly Elias are Outputing 90 % of the work output and you know, this is actually also like Kind of not how things play out in a lot of places like Steve Jobs Is dead Apparently couldn't find anyone else To be the next Steve Jobs of Apple Despite having really quite a lot of money was which to theoretically pay them Maybe he didn't want to really want to success or maybe wanted to be replaceable I don't actually buy that You know based on how this has played out in a number of places There there was a person once who I met when I was younger who was like Have you know built built something that you know like built an organization Um, and he was like hey, hey, Elisard.
2:00:24You want this to take this thing over And I thought he was joking And it didn't dawn on me until years and years later after trying hard and failing hard to replace myself That oh like yeah, I I could have maybe taken a shot at doing this person's job and he probably just never Found anyone else who could Take over his organization and you know, maybe ask some other people and like nobody was willing And I didn't really you know, that's that's his tragedy that he built something and now can't find anyone else to take it over And if I know him at the time I would not have No, I would have at least apologized to him And yeah, to me it looks it looks like people are not dense In the incredibly multi -dimensional space Of people There are too many dimensions and only eight billion people on the planet The world is full of people who have no immediate neighbors and Problems that one person can solve and then like other people cannot solve it in quite the same way I don't think I'm unusual In looking around myself in that highly multi -dimensional space and like not finding a ton of neighbors relative to take value take over and I'm I had You know For people any one of whom could you know do like 99 % of what I do or whatever I I might retire I am tired Probably I wouldn't probably be like marginal contribution of that fifth person is still pretty large, but Yeah, I don't know if the there's the question of like Well, did you occupy a place in the in mind space?
2:02:16Did you occupy a place in social space did people not try to become ali -ezer because they thought ali -ezer already existed And so am I answer that is like man like I don't think ali -ezer already existing would have stopped me from trying to become ali -ezer But you know maybe you just like look at the next effort branch over and There's just like some kind of empty space that someone steps up to fill even though then they don't end up with a lot of obvious neighbors Maybe maybe the world where I died in childbirth is just didn't pretty much like this But I don't feel like you know if It's if somehow we we live to to hear the to hear the to hear the answer about that sort of thing Some someone or something that can calculate it That's not the way I bet But you know if it's true It'd be funny When I said no drama that that did include the concept of um, I don't know Trying to make the story of your planet be the story of you If it all would have played out the same way and that's what and somehow I survived to be told that I'll laugh and I'll cry and that will be the only I mean what I find interesting though is that in your particular case Your output was so public and I mean, I don't know like for example your sequences or like You know your fan your science fiction and fan fiction I'm sure like hundreds of thousands of 18 -year -olds read it or even younger and presumably some of them reached out to you I mean like you know, I think this way Um, I would love to learn more.
2:04:10I'll work on this What was a problem that part? Yeah, I mean yes part of how part of why I'm a little bit skeptical of the story where like people are just like infinitely replaceable is that I tried really really really hard to create like a new crop of Of people who could do all the stuff I could do to take over Cause you know, I knew my health was not great and getting worse I tried really really hard to replace myself I'm not sure where you look to find somebody else who tried that hard to replace himself I tried I really really tried That's what the last wrong sequences were they had other purposes but like first and foremost It was like me looking over my history and going like well I see all these blind pathways and stuff that it took me a while to figure out and there's got to And you know like if I and I feel like I had these near misses on becoming myself Like there's got to be like You know if if if if I got here there's got to be like 10 other people and and like some of them are smarter than I am And they just like need these like little boosts and shifts and hints and they can go down the pathway and you know Like turn into super aliesser and you know, that's what the sequences were like other people use them for other stuff But primarily they were in instruction manual to the young aliizers that I thought must exist out there And they're not really here I Other than the sequences do you mind if I ask like what were the kinds of things you're Talking about here in terms of training the next core of people like you just the sequences I'm not I'm not a good mentor.
2:05:39I did try mentoring somebody for a year once but Yeah, he didn't turn into me Um, so I picked things that were more scalable And Like most people you know like among the other reason why I don't see a lot of people trying that hard to replace themselves is that most people You know, I do like whatever the rather talents don't happen to be like sufficiently good writers I don't think the sequences were good writing by my current standards, but they were good enough And you know most people do not happen to to get that to get a handful of cards that contains the writing card Yeah, whatever else the other talents I'll cut this question out if you don't want to talk about it, but you mentioned that There's like certain health problems that Incline you towards retirement now.
2:06:23Is that something you want you're willing to talk about or I mean I they caused me to want to retire that I doubt they will cause me to actually retire and yeah, it's um Fatigue syndrome our society does not have the words for these things the words that exist are Tainted by their uses labels to the categorize it class of people some of whom perhaps are actually meldingering but Mostly it says like we don't know what what it means and you know, you don't want ever want to have chronic fatigue syndrome on your medical record Because that just tells doctors give up on you and what does it actually mean besides being tired If If one wishes to walk home from work to if one wishes to if if one lives half a mile from one's work Then one had better walk home if one wants to go for a walk sometime Not walk it there if you walk half a mile to work Not going to be getting very much work down the rest of that And aside from that these things don't have names not yet What whatever the cause of this is your working out about this is that it has something to do Or is in some way Correlated with the thing that makes you a laser or do you think it's like a separate thing?
2:07:51When I was 18 I made up stories like And it wouldn't surprise me terribly if you could get if like the One survive to hear the the tale that from something that knew it that the actual story would like be a complex Tangled web of causality in which that was in some sense true but I don't know and Storytelling about it does not hold the appeal that it once did for me Is it a coincidence that I was not able to go to high school or college? Is there something about it that would have crushed the person that I otherwise would have been? Or is it just in some sense a giant coincidence? I Don't know Some people go through high school and college and come out sane How there's there's too much stuff in a human beings history to and there is not and there is you know There's a plausible story you could tell like ah like you know like maybe there's a bunch of potential ali others out there But like they went to high school and college and it killed them You know kill their souls And and you are the one who had the like weird health problem and you didn't go to high school and you didn't go college and you stage yourself and I don't know to me it just feels like patterns in the clouds and Maybe that cloud actually is shaped like a like a horse Or got you know it just But what could this analogy do what could this story do?
2:09:25Well, what you are writing the sequence was and you know the fiction From the beginning was your goal to find somebody who like the main goal to find somebody who could replace you and specifically the task of AI alignment or We did a start off with a different goal and then I mean I thought there I mean you know like in 2008 like I did not know the stuff was going to go down in 2023 I thought I You've probably knew there was a lot more time in which to do something like Like build up build up civilization to another level layer by layer Sometimes civilizations do advances. They improve their epistemology So there was that there was the AI project those were the two projects more or less When did the ad become the main thing as we ran out of time to improve civilization was there a particular year that That became a case for you I mean I think that 2015 2016 17 were the years at which I've noticed I'd been repeatedly surprised by stuff moving faster than anticipated and I was like oh Okay, like if things keep tinny ex -solarating at that pace We might be in trouble and then like 2019 2020 stuff slowed down a bit and Yeah, there was more time that I was afraid we had back then You know that that's what that's what it looks like to be a Bayesian like your estimates go up your estimates go down They don't just keep moving in the same direction because if they keep moving the same direction sometimes you're like oh Like I see where this thing is trending.
2:10:59I'm gonna move here And then like things don't keep moving that direction. The line you're gonna go like oh, okay like back down again And I thought that's that's what sandy looks like I am curious actually like taking many world seriously Does that bring you Any comfort in the sense that like there is one branch of the way function where humanity survives or is that Do you not buy that sort of I'm worried that they're pretty distant Like I expected at least the they did they I don't know like Not sure it's it's enough to not have Hitler but it sure would be a start On things going differently in a timeline But but mostly I don't know there's there's some comfort from thinking of the of the wider spaces than that I'd say as tegmark pointed out Way back when if you have a spatially infinite universe that gets you just as many worlds as the quantum multiverse If you go far enough and in a in a space that is unfounded you will eventually come to an exact copy of Earth or Or a copy of Earth from its past that then has a chance to diverge a little differently So you know the quantum multiverse at nothing reality is just quite if Yeah reality is just quite large.
2:12:08Is that a comfort? Yeah, yes it is That Possibly our nearest surviving relatives are quite distant um or or you have to go like quite some ways through the space before you have worlds that survive But anything but the wildest flukes Maybe our nearest surviving neighbors are closer than that but look far enough and There should be like some species of nice aliens that were smarter or better at coordination and built their Built their happily ever after and yeah That that is a comfort um It's not quite as good as as dying yourself knowing that the rest of the world will be okay but It says kind of like that on a larger scale And weren't you going to ask something about orthogonality at some point?
2:12:59Did I not? Did you at the beginning when we talked about human evolution and Yeah, that's that's not like orthogonality That's the particular question of like what are the laws relating optimization of a system via hill climbing that too To like the internal psychological motivations that it acquires But maybe that was all you meant to ask about Well, I can you explain in what sense You see the broader orthogonality thesis is unprotected by that or orthogonality thesis is um You can have You know like almost any kind of self -consistent Utility function in a self -consistent mind um Like many people are like Why would ais want to kill us?
2:13:46Why would smart things not just automatically be nice and know this is a A valid question which I hope to at some point run into some interviewer where they are of the opinion that smart things are automatically nice so that I can explain on camera um why like Although I myself help like held this position very long ago. I realized that I was terribly wrong about it and That's like all kinds of different things hold together and that you know like if you take a human and make them smarter That may shift their morality it might even depending on how they start out make them nicer But that doesn't mean that like you can Do this with arbitrary minds and arbitrary mindspace because all the different motivations hold together That that's like orthogonality But if if you already believe that then there might not be much to discuss No, yeah, that's really I guess I I wasn't clear enough about it is that Yes, all the different sort of utility functions are possible is that From the evidence of evolution and from the sort of reasoning about how these systems are being trained I think that wildly divergent ones Don't seem as likely as do you do but before instead of having you responded that directly Let me ask you with some questions.
2:14:56I did have about it Which I didn't get to one is actually from skydiving and I don't know if you saw this for recent blog post but here's a quote from it If you've really except the practical version of the orthogonality thesis Then it seems to me that you can't regard education knowledge and enlightenment as instruments for moral betterment On the whole though education hasn't merely improved humans abilities to achieve their goals. It is also improved their goals I'll let you react to that. Yeah, and That yeah, if you if you start with humans if you take humans And possibly also for the requiring particular culture but leaving that aside you take humans who start out Raised the way Scott Aaronson was and you make them smarter they get nicer it affects their goals and If you had and there's a less wrong post about this as there always is um well several about really but like sorting pebbles into correct heaps Describing a species of aliens who think that a a heap of size seven is correct and a heap of size 11 is correct But not eight or nine or ten those heaps are incorrect And they used to think that a heap size of 21 might be correct, but then somebody showed them an array of seven by three pebbles so that you know seven columns three rows And then people realized that 21 pebbles was not a correct heap And you know, this is like the thing they intrinsically care about These are aliens that have a utility function with I would with as I would phrase it some logical uncertainty inside it You can see how as they get smarter they become better and better able to understand which heaps of pebbles are correct And the real story here is more complicated than this Um, but like that's the seed of the answer like Scott Aaronson is inside a reference frame For how his utility function shifts as he gets smarter It's more complicated with that.
2:16:59It's Pete like human beings are made out of out of these like are are more complicated than the pebble sorters They they're made out of like all these complicated desires and as they come to know those desires They change As they come to see themselves as having different options It doesn't just like change which option they choose after the manner or something with the utility function But the different options that they have bring different pieces of themselves in conflict when you have to kill to stay alive You may have a different You you may come to a different Equilibrium with your own feelings about killing and when you are wealthy enough that you no longer have to do that And this is how humans change As they become smarter even as they become wealthier As they have more options as they know themselves better As they think for longer about things and consider more arguments As they understand perhaps other people and give their empathy a chance to grab on to something Solider because of their greater understanding of other minds But that's all when these things start out inside you And the problem is that there's other ways for minds to hold together coherently where they Execute other updates as they know more or don't even execute updates at all because their utility function is simpler than that Though I do suspect that it's not the most likely outcome of training a large language model So large language models will change their preferences as they get smarter indeed Not just like what they do to get to the same terminal outcomes, but like the preferences themselves will up to a point to change as they get smarter It doesn't keep them At some point you are at you you know At some point you know you know yourself Officially well and you are like able to rewrite yourself and at some point there unless you specifically choose not to I think that system crystallizes We might choose not we might we might value the part where we just sort of change in that way Even if it's not no longer heading in a Noble direction because if it's heading in a noble direction You could jump to that as an endpoint Anyway, so it's that why you think AI is will jump to that endpoint because they they can anticipate where their sort of moral updates are going I would reserve the term moral updates for you lose all right these are what let's call preference logical logical preference updates Yeah, yeah What is what are the prerequisites in terms of like whatever makes Aaron's and other sort of smart more people or whatever like Preferences that we humans can sympathize it like what is I you you mentioned empathy, but what are the sort of prerequisites?
2:19:52There's not a short list if there was a short list Those crisply defined things where you get like give it like to And now it's in your moral frame of reference, then that would be the alignment plan I don't think it's that simple or if it is that simple. It's like in the textbook from the future that we don't have Okay, let me ask you this are you still expecting a sort of chimps to humans gain in generality even with these elelems or does a future increase look A bit ordered that we see from legivity 3d2d4 I'm not sure I understand the question can you rephrase yes, it seems that I don't know like from reading or writing from earlier.
2:20:28It seemed like a big party argument was like look a few I don't know how many total mutations was to get from chimps to humans But it wasn't that many mutations and we went for something that could basically get bananas in the forest Just something that could walk on the moon Are you expecting that are you still expecting that sort of gain eventually between I don't know like gpd5 and gpd6 or like some gpdn and gpdn plus one or does it look smoother to you now Okay, so like first of all let me preface by saying that For all I know of how of the hidden variables of nature is completely allowed that gpd4 was actually just did Ha ha this is where it saturates it goes no further It's not how I bet But you know For but you know if nature comes back and tells me that I'm not allowed to be like you just violated the the rule that I knew about I know of no such rule prohibiting such a thing I'm not asking whether these things will plaque to it.
2:21:22I've given lovely intelligence Whether it's a cap that's not the question Even if there is no cap do expect these systems to continue scaling in the way that they happen scaling or do you expect to like some really big jump between some gpdn and some gpdn plus one Yes, and yes, and that's only if things don't plateau before then I mean It's yeah, I can't quite say that I that I know what you know I do feel like We we have this like Track of the loss going down as you add more parameters and you train on more tokens and a bunch of qualitative abilities that suddenly appear at or like I'm sure if you like zoom in closely enough they appear more gradually, but like that appear as as the successful releases of the system Which I don't think anybody has been going around predicting in advance that I know about Um and and like loss continue to go down unless suddenly potatoes Um New abilities appear which ones I don't know Um Is there at some point the giant leap wealth at some point it becomes able to like Toss out the enormous training one paradigm and Build more efficient and like jump to a new paradigm of AI that would be one kind of giant leap You could get another kind of giant leap Via architectural shift something like transformers only there's like an enormously huge or hardware overhang now Like something that is to transformers as transformers were to recurrent neural networks And like maybe there's a maybe and then maybe the loss function suddenly goes down And you get a whole bunch of new abilities That's not because like the loss went down on a smooth curve And you got like a bunch more abilities in a dense spot Maybe there's like some particular set of abilities that is like A master ability the way that language and writing and culture for humans might have been a master ability And you like the loss function goes down smoothly you get this one new but new like internal capability There's a huge jumping output maybe that happens Maybe stuff placos before then and it doesn't happen Being an expert being the expert who gets to go on podcasts They don't actually give you a little book with all the answers in it you know You're you're like just guessing based on the same information that other people have and maybe maybe for lucky slightly better theory Yes, that's why I'm wondering because you do have a different theory of like what fundamentally intelligence is and what it entails Some curious if like you have some expectations of where the GPDs are going I feel like a whole bunch of my successful predictions in this have come from other people being like oh yes I have this theory which predicts that stuff is 30 years off and I'm like you don't know that And then like stuff happens about 30 years off and I'm like ha ha successful prediction And that's basically what I told you right I was like well, you know like you could you could have the loss function Continuing on a smooth line and new abilities appear and you could have them like suddenly appear to cluster because like why not Yeah, because nature just tells you that's up and suddenly You can have like this one key ability that's equivalent to language for humans and like there's a sunnage up in Cape Output capabilities could have like new innovation like the transformer And maybe the loss is actually dropped for stupitously and a whole bunch of new abilities appear at one now This is this is all just me this is me saying I don't know But so many people around are saying things that implicitly claim to know more than that that it can actually start Let's sound like a startling prediction.
2:24:44There's one of my big secret tricks actually People are like well the AI could be like Good or evil so it's so it's like 50 50 right and I'm actually like no like we can be ignorant about a wider space than this And which like good is actually like a fairly narrow range And so many of the predictions like that are really anti predictions It's somebody thinking in a very in a relative along a relatively narrow line and you point out everything outside of that And it sounds like a startling prediction Of course the trouble being when you like you know look back afterwards people are like well, you know like those people saying the narrow thing We're just silly ha I don't give you as much credit I think the credit you would get for that rightly is as a good sort of agnostic Forecaster as somebody who is like sort of common measured But it seems like to be able to make really strong claims about the future About something that is so out of prior distributions is like that that's humanity You don't only have to show yourself as a good agnostic forecaster You have to show that your ability to forecast Because of a particular theory is much greater.
2:25:53Do you see what I mean? It's all about So so when when you're working. Yeah, it's all about the ignorance prior It's all about knowing the space in which to be maximum entropy like The whole bunch of you know like somebody you know like what will the future be? Well, I don't know it could be paper clips It could be staples it could be no kind of office supplies at all and tiny little spider holes It it could be like like little tiny like things that are like Outputting one one one because that's like the most predictable kind of text to predict or or like representations of ever larger numbers in the fast growing hierarchy Because you know, that's what the report that's how the interpret the reward counter I'm actually like getting into specific series which is kind of the absolute point I originally meant to make which is like you know Like if somebody claims to be very unsure.
2:26:44I might say okay So then like you expect like most possible molecular configurations of the solar system be equally probable Well humans mostly aren't in those so like being very unsure about the future looks like predicting with probability nearly one that the humans are all gone Which you know, it's not it's not actually that bad, but it like illustrates the point of Like people going like but how are you sure? Kind of missing the The real discourse and skill which is like oh, yes, we're all very unsure Lots of entropy in our probability distributions, but what is the space on for which you are unsure Even at that point it seems like the most reasonable prior is not that all sort of atomic configurations of the solar system are equally likely Because I agree by that matrix.
2:27:33Yeah, like it's it's like all computations That can like be run over Configurations of solar system are equally likely to be maximized I I But what why like we have a certain we have certain sense that like listen We know what the loss function looks like we know what the training data looks like that obviously is no guarantee of what the drives that come out of that loss function will look like yeah But it's like I'm out pretty different from their loss functions I mean, this is the first question. Yeah, I would say like I was actually know Like if it is as similar as humans are now to our loss function from which we evolved That would be like that honestly might not be that terrible world and it might in fact be a very good world Okay, so it's like Where do you get where do you get good world out of maximum prediction of text Plus our lhf plus Plus like all the whatever alignment stuff that might work Results in something that kind of just like does what you ask it to the way like Does it reliably enough that you know, we ask it like hey help us with alignment then go go stop that Taskate for help with alignment I'm asking for any of the like help us to enhance our brains help us love of love.
2:28:49Thank you Why are people asking for like the most difficult thing? That's the most possible to verify it's whack And then basically at that point we're like turning into gods and we can If you get to the point where you're turning into gods yourselves, you're like you're not quite home free But you know, you're sure past you sure past a lot of the death Yeah, maybe you can explain the intuition that all sorts of drives are Equally likely given unknown loss function and unknown set of data. Oh If yeah, it like so so so if you if you had the textbook from the future or if you were an alien who'd watched a dozen planets destroy themselves the way earth is That we're not actually doesn't that's not like a lie If you've seen 10 ,000 planets destroy themselves the way earth has While being only human in your sample complexity and generalization ability Um, then you could be like oh, yes They're going to try like this trick with loss functions and they will get a draw from like this space of results And the and the alien like can now probably might may now have like a pretty good prediction of like range of like where that ends up Like like similarly like now that we've actually seen how humans turn out when you optimize them for reproduction It would like not be surprising if we found some aliens the next door over and they had orgasms Now maybe they don't have orgasms, but you know like some like But you know like if they had some kind of like strong surge of pleasure during the active mating We're not we're not as we're not surprised We've seen how that could place out in humans if they have some kind of like weird food That isn't that nutritious, but like makes them much happier than any kind of food that was more nutritious and around in their ancestral environment Like like ice cream We probably can't call it as ice cream right it's not going to be like sugar salt fat frozen But we're not specifically going to have ice cream right but they might play go They're not going to play chess Because chess is like more has more specific pieces right yeah They're not going to play they're not going to play go on like 19 by 90 Then I might play go on some other size Probably odd well, can we really say that I don't know I'd I bet on like an odd if they play go I bet on an odd or dimension at Now let's say let's look two -thirds the process rule of session sounds about right Um Unless there's some other reason why go just totally does not work on and even Dement or dimension that I don't know because I'm especially acquainted with the game Um We like the point is like just like you know reasoning off of humans is like pretty hard We have like the loss function over here.
2:31:31We have like humans over here We can like look at the rough distance like all the like weird specific should like stuff that humans are created around and be like You know like like if the loss of lunch is over here and humans are over there like maybe aliens are like over there And if we had like three aliens that would like expand our views of the possible We'd have like a or even two aliens would like vastly expand our views of the possible And give us like a much stronger notion of what the third aliens would look like like humans aliens third third race Um But you know You know like you you know like the the wild and optimistic scientists have never been through the never been through this with with AI's So they're like oh, you know like like you optimize the AI to like say nice things and help you and like make it a bunch smarter Probably so nice things and helps you is probably like totally a line.
2:32:21Yeah Yeah Yeah, they don't know any better I'm not trying to jump ahead of the story But the aliens the aliens know no no we're we're you end up around the loss functioned They know where the where the where the power is gonna play out much more narrowly We're we're guessing much more blindly It just like it just leaves me in a sort of unsatisfied place That we apparently Know about something that is so extreme that maybe a handful of people in the entire world believe it from first principles About you know the doom of humanity because of AI But but this theory that is so Uh So productive in that one very unique prediction Is unable to give us any sort of other prediction of what this world might look like in the future or About what happens before the before we all die like It can tell us nothing about the world until the point at which makes prediction that is the most remarkable in the world You know rationalist should win, but rationalist should not win the lottery I ask you like what other theories are supposed to be up and doing a like amazingly better job of predicting the last three years You know, maybe it's just hard to predict right and and in fact it's like Easier to predict the end state than the strange complicated winding paths that lead there much like a few good You know play against alpha -go and predict it's gonna like be in the class of winning board states But not not exactly how it's gonna be you So like not quite like that the problem of difficult to predict in the future But you know from my perspective the future is just like really hard to predict and there's a few places where you can like Rench what sounds like an answer out of your ignorance Even though really you're just being like Well, you're gonna like end up in some like random weird place around this loss function And I haven't seen it happen with 10 ,000 species.
2:34:20So I don't know where Very very impoverished by the from the stand point of anybody who like actually knew anything It actually predict anything but the rest of the world is like Oh like We're easily we're like equally likely to to win the lottery is lose lottery right like either we win or we don't You come along and you're like no no your chance of winning the lottery is tiny They're like what how can you be so sure where do you get your strange certainty and the actual route of the answer is that you are Putting your maximum entropy over a different probability space Like that just actually is the thing that's going on there You're saying all lottery numbers are equally likely instead of winning and losing are you the only thing?
2:34:59So I think um the place to sort of close this Conversation is let me just sort of give the The main reasons why I'm not convinced that doom is likely or even that it's more than 50 % probable or anything like that um Some are the things that I started this conversation with that I don't feel like I heard a neat knockdown arguments against And some are new things from the conversation um and the following things are things that Even if uh any one of them individually turns out to be true I think Doomed doesn't make sense or as much less likely So Going through the list I think Probably more likely than not this uh entire frame all around alignment and AI is wrong This is maybe not something that would be easy to talk about But I'm just kind of skeptical of sort of first principles reasoning That has really wild conclusions Okay, so everything the solar system just ends up in a random configuration then Uh or it stays like it is unless you have very good reasons to think otherwise And especially if you think it's going to be very different from the way it's going You must have very very good reasons like iron clad reasons for thinking that it's going to be very very different from the way it is Uh -huh, so this is You know the humanity hasn't really existed for very Man, I don't even know what to say to this thing.
2:36:36We're like this tiny like everything that you think of as normal is this tiny Flash of things being in this particular structure out of a 13 .8 billion year old universe which very little of which was like 20th century Part of me 21st century Yeah My own brain sometimes gets stuck in childhood, right? Very very little which is like 21st century like civilized world the You know On this like little fraction of the surface of one planet in a vast solar system most of which is not earth and vast universe most of which is not earth And it has lasted for like such a tiny period of time through such a tiny amount of space and And has like changed so much over you know just the last 20 ,000 years or so And and here you are like being like why would things really be any different going forward?
2:37:27I feel like that argument proves too much because you could use that same argument like somebody says comes up to me and says Uh, I don't know theologian comes up to me and says like the raptor is coming and let me sort of explain why the raptor is coming and I say I'm not claiming that the arguments are as bad as the argument for raptor I'm just just follow the example Um, but then they say listen I mean look at how wild human civilization has been would it be any wilder if there was a raptor and I'm like yeah actually as well as human civilization has been the raptor would be much wilder As well as it relates to laws of physics yes Um, I'm not trying to violate the laws of physics even as you probably know them.
2:38:01How about this somebody comes up Oh, you know, I've got the perfect example. Okay um somebody comes up to me he says we have actually nanosystems right behind you He says I've read air directors nanosystems. I've read vitamins. There's plenty of room at the bottom And he explains two things are not mentured but go on okay, oh fair enough. He comes to me and he says um, let me explain to you my first principles argument about how Some nanosystems will be replicators and replicators because of some competition yada yada yada argument They turn the entire world into goo just making copies of themselves This kind of happened with humans, you know Well life generally.
2:38:40Yeah, yeah, but So then they say like listen as soon as people start building nanosystems pretty soon 99 person probability the entire world turns into goo Just because the replicators are the things that turn things into goo. There will be more replicators and non replicators It uh, I don't have an object level debate about that but it just like I just started that and looked like like yes Human civilization has been wild But the entire world turning into goo because of nanosystems alone Just seems much wilder than human civilization. You know this this this uh This impet this this argument probably lands with greater force on somebody who does not expect stuff to be disassembled Binanno systems albeit intelligently controlled ones rather than goo and like quite near future Especially the 13 .8 billion year time scale, but you know Do you do you expect this little momentary flash of what you call normality to continue?
2:39:28Do you expect the future to be normal? I uh, no, I expect any given vision of how Things shape out to Be wrong especially it is not like you are suggesting that the current weird trajectory Continues being weird in the way it's been weirded that we continued to have like two percent economic growth or whatever And that leads to incrementally more technological progress and so on you're suggesting there's been that Specific species of weirdness which leads to an which means that this entirely different species of weirdness is warranted Yeah, we've got like different weirdnesses over time the the jump to super intelligence does strike me as being significant in the same way As self -replicate first self -replicator first self -replicator is the universe transitioning from you see mostly stable things To you also see a whole bunch of things that make copies of themselves and then somewhat later on there's a state where You know, there's this like strange transition the sportor between the universe of stable things where things come together by accident and stay as long as they endure to this world of complicated life And that transitionary moment is when you have something that arises by accident and yet self replicates And similarly on the other side of things you have things that are intelligent making other intelligent things But to get into that you've world you've got to have the Thing that is built just by things copying themselves and mutating and yet is intelligent Enough to make another intelligent thing Now if I sketched out that cosmology Would you say no no I don't believe in that What if I sketch out the cosmology of Because of replicators blah blah blah intelligent beings intelligent beings create nanosystems blah blah blah No, no, no, no, no, I don't don't tell me about your like not not the proof too much.
2:41:22I just want to like Like I discussed how to cosmology do you buy it In the long run are we in a world full of things replicating or a world in a full of intelligent things designing other intelligent things yes So you so you you buy that vast shift in the foundations of order of the universe that instead of the world of things that make copies of themselves imperfectly we are in the world of things that are designed and were designed You buy that vast cosmological shift. That was just describing the utter disruption of everything you see that you call normal down to the leaves and the trees around you You you believe that Well the same skepticism you're so fond of that argues against the rapture can also be used to Disprove this thing you believe that you think is Probably pretty obvious actually now that I pointed it out Okay, um, you're You're skepticism disproves too much my friend That's actually a really good point.
2:42:20It still leaves open the possibility of like how it happens and what it happens But actually that's that's a good point. Okay, so a second second thing I'm not you set them up all knock them down one after the other Second thing is wrong I was just dumping head to the predictable up there Maybe alignment just turns out to be much simpler or like much easier than we think it's not like we've as a civilization spent that That much resources or brain power in solving it if we put it even the kind of resources that we put into Elucidating strength the area or something into alignment. It could just turn out to be like yeah, that's enough to solve it and In fact in the current paradigm it turns out to be simpler because you know that they're sort of pre -trained on human thought and That might be a simpler regime than something that just comes out of a black box that like I you know like a alpha zero or something like that so like some of my like Could I be wrong in an understandable way to me in advance mass Which is not where most of my hoe comes from is on you know What if RLHF just works well enough and To the people in charge of this are not the current disaster monkeys But instead have some modicum of caution and are using their like Like no what to aim for in RLHF space Which the current crop do not and I you know that's not really that confident of their ability to understand if I Told them but maybe you have some folks who can understand or anyways I can sort of see what I try These people will not try but You know the current crop that is and I'm not actually sure that that if that If somebody else takes over like the government or something that they listen to me either but I can Now maybe you So some of the trouble here is that you have a choice of targets and and like matters all that great One is you look for the niceness that's in humans and you try to bring it out in the AI and then You with its cooperation because You know it knows that if it makes it that if you try to just like amp it up it might not stay all that nice Or that if you build a successor system to it it might not stay all that nice and it doesn't want that because you You know like you you you you you narrowed down the shagoth enough You know that not just like and and the mat you know somebody somebody once had this incredibly profound statement That I think I like someone disagree with but it's still so incredibly profound.
2:45:15It's consciousness is when the mask eats the shagoth and Maybe that's it it may maybe you know with the with the right set of bootstrapping reflection type stuff stuff stuff you can Have that happen on purpose more or less where They're where the the systems output that you're shaping is like to some degree in control of the system and you you Locate niceness in the human space. I I have fantasies along the lines of what if you trained GPTN to distinguish people being Nice and saying sensible things and argue validly and you know Can't just I'm not sure that works if you just have amazon turks try to label it you just get the like strange thing you located that Aralechef located in the present space which is like some kind of weird corporates beak like left rationalizing leaning strange Telephone announcement creature Is what they got with the current crop of RLHF?
2:46:35Note how this stuff is weirder and harder than people might have imagined initially But you know, we've we've lived aside that the the part where you try to like Jumpstart the entire process of turning into a grizzled cynic and update as hard as you can do it in advance Leave that aside for a moment Like maybe you can look maybe you are like able to train on scott Alexander and so you want to be a wizard some other night and Some other nice real people real people and nice fictional people and separately train on what's valid argument That's that's going to be that's going to be tougher But you know, I could probably put together a crew of a dozen people who could provide the data on that RLHF And you find like the nice creature and you find the nice mask that that argues valiantly Do some more complicated stuff to try to boost the thing where it's like eating the shagoth where that's what the system is And that's like more what the system is less what it's pretending to be Look at the I I do seriously think this is like Like I can say this and like the disaster monkeys at the Current places can cannot along to it, but they have not said things like this themselves that I have ever heard And and that is not a good sign But and then like if you don't have this up too far which on the present paradigm you like can't do anyways Because if you like train the very very smart person of this version of the system it kills you before you can RLHF it Um, but like maybe you can like train dpt to like distinguish like nice valid um kind um careful And then like filter all the training data to get the nice things to train on and then train on that data rather than training on everything To try to like a worthy waluigi problem um It's or just more generally having like all the The the darkness in there like just train on the light that's in humanity Um, so there's like that kind of course and Not if you don't push that too far.
2:48:40Maybe you can get a genuine ally And maybe things play out differently from there. That's like one of the little rays of hope for But but that's not I don't think that actually looks like alignment is So easy that you you just get whatever you want it's a genie gives you what you wish for I don't think that that doesn't even strike me as hope I honestly the way you described it seem kind of compelling like I don't know why that isn't even rise to 1 % Uh the possibility works out that way I literally this this is like literally my you know my like AI alignment fantasy from 2003 Um, though not with but but well not with like RLHF and as the implementation method or LLM says the base Um, and it's you know gonna be more dangerous than when I was thinking about when I was dreaming about in 2003 and I think in a very real sense it feels to me like the the the the people doing this stuff now literally not gotten as far as it I was in 2003 and uh You know, I can like I've now like written out my answer sheet for that It's on the podcast it goes on the internet and now they can now they can pretend that that was their idea Or like or like sure that's obvious we're gonna do that anyways and yet they they didn't they didn't say it earlier And uh and you can't You can't you can't run a big project off of one person who It's it's it's failed to gel the the alignment field failed to gel that's that's my objective to the like Well, you just thrown a ton of more ton of more money and then it's all solvable because I've seen people try to Amp up the amount of money that goes into it and the stuff coming out of it has not Gone to the places that I would have considered obvious a while ago And I can like print out all my entries.
2:50:34He's seats forward and each time I do that it gets a little bit harder to make the case next time But I mean how much money are we talking the grand scheme of things because like civilization itself has a lot of money I know I know people who have a billion dollars I don't know how to throw a billion dollars at like Outputting lots and lots of alignment stuff, but you you might not but I mean you are one of ten billion right like uh it is And other people go ahead and spend lots of money on it anyways and and Is everybody makes the same mistakes? Nate sorries has a post about it. I forget the exact title, but like everybody coming into alignment makes the same mistakes Let me just go out to the third point because I think it plays into what I was saying the third reason is If it if it is the case that you know this these capability scale in some Constant way as it seems like they're going from two to three or three or four.
2:51:29What does that even mean? But go on That they get more and more general. It's not like going from a Going from a mouse to a human or a chimpanzee to a human. It's like going from GPT3 to GPT4. Yeah, well it just seems like that's less of a jump, but then shift to human Like a slow cumulation up capabilities There are a lot of like S curves of emergent abilities, but overall the curve looks sort of man I feel like we bit off a whole chunk of chimpti human and GPT3 .5 to GPT4, but look on Regardless, okay, so then this leads to human level intelligence for some interval. I think that I was not convinced from The arguments that we could not have a system of sort of checks on this the same way he have checked on smart humans that it would try to deceive us to achieve its aims anymore than smart humans or in positions of power try to do the same thing for a year What are you going to do with that year?
2:52:32Before the next generation of systems come out that are not held in check by humans because they are not roughly in the same power Intelligence range as humans What are you going to maybe you can get a weight year with that maybe you can get a year like that Maybe that actually happens. What are you going to do with that year that prevents you from dying the year after? One is one possibility is that Because these systems are trained on human text maybe just progress just slows down a lot after it gets to Slightly above even level. Yeah, that's not that's that yeah, that that's not how I would be quite surprised if that's how anything works Why is that?
2:53:09For one thing because it's you know like Like for an alien to be an actress playing all the humans on the internet for another thing Well, first of all you realize in principle that the task of minimizing losses on predicting human text does not have a Yeah, you understand that in principle. This does like not stop when you're smart as a human right like you can see that the computer science of that I don't know if I see the computer science of that, but I think I probably understand the art. Okay So like very very you know somewhere on the internet is a list of hashes followed by the string hashed This is a simple demonstration of how you can go on getting lower losses by throwing a hyper computer at the problem There are pieces of text on there that we're not produced by humans talking in conversation But rather by like lots and lots of work To determine get extract experimental results out of reality that text is also on the internet Maybe there's not enough of it for the machine learning paradigm to work but I'd sooner buy that like That the gbt systems just bottlenecks short of being able to predict that stuff Better rather than that.
2:54:20But you know like you can maybe buy that but like the notion that like like you only have to be smart As a human to predict all the text as the internet as soon as you turn around and stare at that, but it's just transparently false Okay, agreed um Okay, how about this story You have something that is sort of human like that is Maybe above humans at certain aspects of science because it's specifically trained to be really good at The things that are on the internet, which is like you know chunks and chunks of archive and whatever Um, whereas it has not been trained specifically To gain power and while at some point of intelligence that comes along uh Can I just restart the whole sentence?
2:55:01No You have spoken it it exists. It cannot be called the back There are no takebacks. There is no going back. There is no going back. Go ahead. Uh Okay, so here's another story I expect them to be better than humans at science than they are power seeking because We had greater selection pressures for power seeking in our ancestral environment than we did for science and while at a certain point both of them come along as a package You know, maybe that uh, they can Be at varying levels, but anyways, so You have this sort of uh early model that is kind of human level except a little bit ahead of us in science You ask it to help us align the next version of it then the next version of it is more aligned because we have its help um and the uh In sort of like this inductive thing where the next version helps us align the version of it Where do people have this notion of getting ais to help you do your a i alignment homework?
2:56:06It just why can we not talk about having it you enhance you and you would tell it instead? Okay, so either one of those stories where it just like helps us enhance humans Uh enhance humans that help us Figure out the alignment problem or something like that. Yeah, I I it's like kind of weird because you know like It's like small large amounts of intelligence don't automatically make you a computer programmer And if you are a computer programmer, you don't automatically get security mindset But it feels like there's some level intelligence for you out to automatically get security mindset And I think that's about how hard you have to augment people um to have them able to do alignment Like the level where like they have security mindset not because they were like special people with security mindset But just because like there's that intelligent that you just like automatically have security mindset I think that's about the level of where a human could start to work on alignment more or less Why is that story then not 1 % Uh get you to 1 % probability that it helps us up for the whole crisis?
2:57:02Well Because it's not just a question of the technical feasibility of can you build a thing that applies its general intelligence narrowly to The neuroscience of augmenting humans It's a question of like like So like I like one I feel like that is like probably like over 1 % technical feasibility But the world that we are in is so far so far from From from doing that from trying trying the way the work could actually work like like not like the try where like Oh, you know like well, we like we like just like do a bunch of RLHF to try to have a spit out output about this things But not about that thing and you know that that that no no that not that um Yeah, what 1 % that we could that humanity could could do that If it if it tried and tried and just the right direction as far as I can Proceive angles in this space um yeah, I'm over 1 % on that I am I'm not very high on us.
2:58:08I'm not doing it Maybe I will be wrong maybe the time article article article I wrote Saying shut it all down gets picked up and there are very serious conversations and the very serious conversations are actually effective in In shutting down the headlong plunge and there is a narrow exception carved out for the kind of narrow application of Trying to build an artificial general intelligence that applies its intelligence nearly and to the problem of augmenting humans and you know That I think might be a harder sell to the world Then just shut it all down You they could they could get shut it all down and then not do the things that they would need to do to have an exit strategy I feel like even if you told me that they would that they went for shut it all down I would be like then next expect them to have no exit strategy Until the world ended anyways, but perhaps I underestimate them Maybe there's a will in humanity to do something else which is not that And if there really were um, yeah, I'm I think I'm even over 10 % that that That would be a technically feasible path if we if we if they looked in just the right direction but I I'm not over 50 % on them Actually doing the the shut it all down I'm not if they do that.
2:59:34I am them not over 50 % on They're really truly being the will of something else that is not that to really have an exit strategy Then from there you have to Go in at sufficiently the right angle to to materialize the technical chances and not do it in the way that's Just ends up a suicide or if you're lucky like gives you the clear warning signs and And then people actually pay attention to those instead of just optimizing away the warning signs and I don't want to make this sound like the multiple stage files seem like oh knows more than one thing has to happen Therefore the resulting thing can never happen which you know like super clear case in point of why you cannot prove anything will not happen this way of Nate silver arguing that trump needed to get through six stages to Become the Republican presidential candidate each of which was less than half probability and therefore he had less than 164th chance of becoming the Republican Not when eighth what six six stages of doing They're very had like less than 164th chance of becoming.
3:00:40I think just a Republican candidate not not winning So yeah, so like you can't just like break things down at the stages and then say there for the probability of zero You can break down anything at the stages But like but like even so you're asking me like well like isn't over 1 % that it's That it's possible. I'm like Yeah, possibly even over 10 % that that doesn't get me to because the the Like the the reason why you know go ahead and tell people yeah, don't don't put your hope in the future or you're probably dead is that The the existence of this technical array of hope if you do just the right things.
3:01:17It's not the same as expecting that the world Re -shapes itself to permit that to be done without destroying the world in the meanwhile I expect things To continue on largely as they have and You know And what distinguishes that from despair is that at the moment people are telling me like no no if you go outside the tech industry People will actually listen. I'm like all right. Let's try that. Let's write the time article Let's jump on that. Let's see if it works It will lack dignity not to try That's not not the same as expecting as being like oh, yeah, oh, over 50 % if they're totally gonna do it that that that time article is totally gonna take off I'm not currently not over 50 50 not over 50 percent on that you know you said like any one of these things could mean and yet like Even if this thing is technically feasible that doesn't mean the world's going to do it We are presently quite far from the world being on that trajectory or of doing the things that we needed to be create to create Times pay the alignment tax to do it Maybe the one thing I would dispute is How many things need to go right from the world as a whole for any one of these past to succeed Because which goes in the fourth point, which is that maybe The sort of universal prior over all the drives that any I could have is just like the wrong way to think about it and this is something that um I mean you definitely want to use the alien observation of 10 ,000 planets like this one prior for what you get to after training on like thing ex Uh, it just like especially when we're talking about things that have been trained on you know human text I'm not saying that It was a mistake earlier on the conversation for me to say there'll be like the average of human motivations whatever that means But it's not um It's not inconceivable to me that it would be something that is very sympathetic to human motivations having been Having sort of encapsulated all of our output I think it's much easier to get a mask like that than to get a shogoth like that Possibly but again, this is something that seems like I Don't know probably the output at least 10 % and That by just by default that is something that is not so it is not incompatible with the flourishing of humanity Like well, why is the utility function you hope it has that has its maximum There's so many questions of humanity.
3:03:34There's so many possible like name three name one Spell it out. It but I don't know Wands to keep us as a zoo The same way we keep like other animals in a zoo. This is not the best outcome for unity But it's just like something where we survive and flourish. Okay. Whoa whoa flourish Keeping in a zoo did not sound like flourishing to me Zoo was wrong were to use there. Well, whoa whoa whoa whoa because it's not what you wanted Why is it not a good action you just actually name three you didn't ask me like No, no, I'm saying it's like like you're like oh well like prediction. Oh, no, no I don't like my prediction.
3:04:08I want a different prediction You didn't ask for the prediction. You just asked me to name them like name possibilities I had meant like possibilities in twitch you put some probability. I had meant for for like Like a thing that you thought held together. This is the same thing as when I ask you like what is this specific utility function It'll have that will be incompatible with You know humans existing. It's like your your modal prediction is the super vast majority of predictions are Of utility functions are incompatible with with human Existing I can make a mistake and it'll still be incompatible with humans existing Right like I can just be like you know like like I can just like it describe a randomly rolled utility function And up with something incompatible with humans success So like at the beginning of human evolution you could think like okay This thing will become generally intelligent and What are the odds that it's flourishing on the plan that will be compatible with the survival of Spruce trees or something it's like and the long term we sure aren't I mean like maybe if we win we'll have there be a space for spruce trees Yeah, so as long so you can have spruce trees as long as the mitochondrial liberation front does not object to that What is the mitochondrial liberation front is that live to if you're do you have you have do you have do you have you know sympathy for the mitochondria enslaved Working all their lives to the benefit of some or other organisms So this is like some weird hypothetical like we're hundreds of thousands of years General intelligence existed on earth you could say like is it compatible with some random species that exists on earth Like is it compatible with spruce trees existing and I know we probably chopped on a few spruce trees But and the answer is Yes as these as a very special case of us being the sort of things that would maybe some of us would maybe conclude that we wanted That we specifically wanted spruce trees Spruce trees to go on existing at least on earth in the glorious transhuman future and their votes winning out against the those of the mitochondrial liberation front I guess since part of the sort of transhumanist future is part of the thing we're debating It seems weird to assume that as part of the question Well, well the thing I'm trying to say is you're like well like If you looked at the humans would you like not expect them to end up incompatible with the spruce trees and I'm it being like Sir you a human have like looked back and like looked at how humans wanted the universe to be and being like Well, would you not have anticipated in retrospect that humans would like want the universe to be otherwise and I agree that we like Might want to conserve a whole bunch of stuff Maybe we don't want to conserve the parts of nature where things bite other things and inject venom into them and the victim's I'm just saying Maybe even if maybe you know I think that many of them don't have qualia This is disputed Some people might be disturbed by it even if they didn't have qualia We might want to be polite to the sort of aliens who would be disturbed by because they don't have qualia and they just see like Things don't want venom injected into them there before they should not have venom We might conserve some parts of nature, but again, it's like it's like it's like firing an arrow and then drawing the circle around the target I would disagree with that because Again, this is similar to the example we started off the conversation with but it seems like you are reasoning for from what's Might happen in the future and because we disagree about what might happen in the future in fact the entire point of this disagreement is to Test what will happen in the future assuming what will happen in the future as part of your answer seems like I mean That way to okay, but then you're like claiming things as evidence for your position based on what exists in the world now I have evidence that are not evidence one way or the other because the basic prediction is like if you offer things enough options They will They will go out of distribution like if you are like it's like it's like pointing to the like very first People with language and being like they haven't taken over the world yet It's like And like they have like knock on way out of distribution yet and it's like they haven't had general intelligence for long enough to accumulate The things that would give them more options such that they could start trying to select the weirder options The prediction is like when you have when you give yourself more options you start to select ones that look weird or relative to the ancestral distribution as long as you don't Have the weird options you're not going to make the weird choices And if you say like we haven't yet observed your future that's fine But like acknowledge that then like evidence against that features not being provided by the past This is the thing I'm saying there You look around it looks so normal According to you who grew up here if you grown up as a millennium early or your argument for the persistence of more normality Might not seem as persuasive to you after you'd seen that much change So this is a separate argument though right like I'm like like look at all this stuff humans haven't changed yet You say now selecting the stuff we haven't changed yet But if you go back 20 ,000 years and be like look at the stuff intelligence hasn't changed yet You might very well select a bunch of stuff that was going to fall 20 ,000 years later Is the thing I'm trying to gesture at here But so like how do you propose we've reason about what general intelligence to do when the world we look at after hundreds of thousands of years of general intelligence is one that we can't use for evidence Because but yeah dive under the surface look at the things that have changed Why did they change look at the processes that are generating those choices And Since we have sort of these different functions of like where that goes Like look at the thing with ice cream look at the thing with condoms look at the thing with pornography See where this is going Just like I just seems like I would disagree with your intuitions about like what future smarter humans will do Even with more options I was like in the beginning of conversation I disagreed that they would most humans would adopt sort of like A trans -evenous way to get better DNA or something but you would So yeah, you just like you just like look down at your fellow humans You have like no confidence in their ability to tolerate weariness even if they have smarter I wonder like you do do you think what what do you think would happen if we did a poll right now think I'd have to explain that poll pretty carefully Because you know they haven't got the the intelligence head bends yet, right?
3:10:42I mean we could do a Twitter poll with like a log explanation in it 4 ,000 character Twitter poll yeah, I like I mean Man, I like some somewhat tempted to do that just for the sheer chaos and point out the the drastic selection effects of A It's my Twitter followers be they read through a 4 ,000 character tweet I feel like this is not likely to be truly very informative by my standards But part of me is amused by the prospect for the chaos. Yeah, or I could do it on my end as well Although my followers are like really weird as well Yeah, but plus you wouldn't like really I worry you wouldn't sell that transhumanism thing as well I could word it does like you just send me the wording but anyways That's everything but anyways Given that we disagree about what in the future general and intelligence will do Where do you suppose we should look for evidence about what the general of intelligence will do given our different theories about it If not from the present I mean, I think you look at the mechanics you say as people have gotten More options they have gone Further outside the ancestral distribution and we zoom in and it's like there's all these different things that people want And there's this like narrow range of options that they had 50 ,000 years ago and they things that they want have Maxima or 50 ,000 years ago at stuff that coincides with reproductive fitness and then as a result of the humans getting smarter They start to accumulate culture which produces changes out of time scale Faster than natural selection runs although it is still running contemporaneously The humans are just running faster running faster than natural selection it didn't actually halt and The additional they like generate additional options not blindly but according to the things that they want and they invent ice cream they You know like not not at random It doesn't just like a coughed up at random.
3:12:43They're like searching the space of things that they want and and generating new options for themselves That optimize these things more that weren't in the ancestral environment and good -hard slot applies what good -hard's curse supplies Once you that like as you apply optimization pressure the correlations that were found naturally come apart and aren't present in the thing that gets optimized for Like you know, you like just give some people some tests who've never gone to school The ones who high score high in the tests will know the problem domain because they you know like you're just like Give's a bunch of carpenters a carpentry test The ones who score high in the carpentry test will like know how to carpentry things then you're like Yeah, I'll like pay you for high scores in the carpentry test I'll give you this carpentry degree and like people like oh I'm gonna like optimize the test specifically and they'll get higher scores than the carpenters and do and be Worse at carpentry because they're like optimizing the test and that's the story behind ice cream All right, and and you zoom in and look at the mechanics and not the like Grand scale view because the grand scale of you just like never gives you the right answer basically Like anytime you ask what would happen if you applied the grand scale few lost lost from the past It's always just like oh, I don't see that why this thing would change.
3:13:53Oh, it changed who how weird who could have possibly have expected that Maybe over a different differential grand scale view because I would have thought that that is what you might use to categorize your own view But I don't want to get it caught up in semantics mine is zooming in it's looking at the Smith looking in the mechanics That's that's how I'd present it if We are like so far to distribution of natural selection as you say We're currently not we're currently no we're near as far as we could be like this is not the glorious transumist future I claim that if even if you get much smarter as like if humans get much smarter through brain augmentation or something Then there will still be spruce trees in like millions of years in the future And if you still want to come the day I don't think I myself would oppose it unless there'd be like distant aliens who are very very sad about what we were doing to the mitochondria And then I don't want to ruin their day for no good reason But the reason that it's important to say it in the form of like Given human psychology spruce trees will still exist is because that is the one evidence of sort of generality arising We have and even after millions of years of that generality Like we think that spruce trees would exist I feel like we would be in this position of spruce trees in comparison to the intelligence we create and sort of the universal prior and whether spruce trees would exist Doesn't make sense to me.
3:15:09Okay, so but do do you see how this perhaps leads to like everybody's severed heads being kept alive in jars on its own premises as opposed to humans getting the glorious transhumanist future No, no, they have the glorious but glorious transhumanist future. Those are not real spruce trees You know like you're you're talking about like plain old spruce trees you want to exist, right? Not not the sparkling giant spruce trees with built -in rockets You're talking about humans being Captives pets in their ancestral state forever. Maybe being quite sad. Maybe they still get cancer and die of old age And they never get anything better than that Does it keep around us around as we are right now?
3:15:52Do we relive the same day Over and over again. Maybe this is the day when that happens Hmm I mean you do you see how like how how the the the the general trend I'm trying to point out to hear is you like have a rationalization for why they might do Thing that is allegedly nice and I'm saying like why exactly are they wanting to do thing? Well if they want to do thing for this reason Maybe there's a way to do this thing that isn't as nice as you're imagining and this is systematic You're imagining reasons they might have to give you nice things that you want But they are not you Not unless we get you know, not unless we get this like exactly right and they actually care about the part where you want Some things and not others you are not describing something you are doing for the sake of the spruce trees Do spruce trees have diseases in this world of yours?
3:16:50Do the diseases got to live do they get to live on spruce trees? And it's not just and it's not a coincidence that that I can like zoom in and poke at this and ask questions Like this and that you did not ask these questions of yourself. You are imagining nice ways you can get the thing But reality is not necessarily imagining how to give you what you want And the AI is not necessarily imagining how to give you what you want and these and like for everything you can be like Oh, well like hopeful thought maybe I get all this stuff. I want because the AI reasons like this Because it's the optimism inside you that is generating this answer And that and if the optimism is not in the AI if the if the AI is not specifically being like Well, how do I pick a reason to do what I to do things that will Give this person a nice outcome You're not going to get the nice outcome You're going to be you were living the last day of your life over and over It's going to like create old maybe creates old -fashioned humans ones from 50 ,000 years ago.
3:17:50Maybe that's more quaint Maybe maybe maybe it's just like just as happy with like bacteria because there's more of them And that's equally old -fashioned you're going to create the specific spruce tree over there Maybe from its perspective, you know like a generic bacteria is just as good a form of life as like the generic spruce tree is of a spruce tree And like and and like this is not specific to the example that you gave it it's It's it's me being like well suppose we like took a criterion that sounds kind of like this and asked how do we actually maximize it What else satisfies it? Not just you're you're like trying to argue the AI into doing what you think is a good idea by giving the AI reasons why it should want to do the thing under like some set of like hypothetical motives, but anything like that if you like optimize it on its own terms without Narrowed down to where you wanted to end up because it actually felt nice to you the way that you define niceness Like it's all going to have Somewhere else somewhere that isn't as nice something maybe where we'd be like sooner scour the surface of the planets The clean with nuclear fire rather than cut but let that AI come into existence.
3:18:58So I do think those are also probable Because you know instead of hurting you as you know, there's like something more efficient for it to do that maxes out its utility function Okay, I acknowledge that you had the better argument there um But here's another intuition. I'm curious every smart of that earlier we talked about The idea that like if you bred humans to Be friendlier and smarter. This is not wrong. I mean this but like um if you did that I think I want to register for the record that the term breeding humans would would cause me to like look Askins to get at any at alien and at any aliens who were to propose that as a policy action on their part All right, I said it move on.
3:19:43No, that's not what I'm proposing we do. I'm just saying. I was a sort of thought experiment but So you are answered that oh because human psychology that's why you shouldn't assume the same of the eyes They're not gonna start with human psychology. Okay fair enough. Assume we start off with dogs, right good old -fashioned dogs and We bred them to be more intelligent but also to be friendly Well as soon as they oh our past a certain level of intelligence I object to us like coming and in breeding them. They can no longer be owned They are now sufficiently intelligent to not be on anymore But let us leave aside all morals okay carry on It in the thought experiment not in real life.
3:20:20You can't leave out the morals in real life Do you have to sort of um universal prior over their drives of these like super intelligent dogs that are bred to be friendly man So I think that weird shit starts to happen at the point where the dogs get smart enough that they are like What are these flaws in our thinking processes? How can we correct them? You know the over the seafarth threshold of dogs although maybe that's like see far as some strange baggage over the Korsybsky threshold of dogs after Alfred Korsybsky Yeah, so I think that You know there's this whole domain where they're stupider than you and sort of like being shaped by their genes and not shaping themselves very much And as long as that is true you can probably go on breeding them and Issues start to rise when the dogs are smarter than you When the dogs can manipulate you if they get to that point Where the dogs can strategically present a particular appearances to fool you Where the dogs are aware of the breeding process and possibly having opinions about where that should go in the long run Where the dogs are even if just by thinking and by adopting new rules of thought Modifying themselves in that small way These are these are some of the points where like I expect the weird shit to start to happen And it won't and the weird shit will not necessarily show up while you're just reading the dogs There's a weird shit look like Dog gets smart enough Dod -dod -dod human stuff existing If you keep on optimizing the dogs Which is not the correct course of action I think I mostly expect this to eventually blow up on you Um, but blow up on you that bad It's hard Well, it but it I expect to blow up on on you quite bad I'm trying to think about whether I expect super dogs to be sufficiently in a human frame of reference in virtue of them Also being mammals that they that a super dog would like create Human ice cream like they you bred them to have preferences about humans And they invent something that is like ice cream to those preferences Or does it just like go off someplace stranger Hi, there could be eye ice cream.
3:22:40Hmm. There could be eye ice cream Ice cream that is as things that is equal into ice cream for a eyes That is essentially my prediction of what the solar system ends up showed with But anyway, the exact ice cream is like quite hard to predict just like it would be very hard to like look at Well, if you optimize something for inclusive genetic fitness, you'll get ice cream That was a very hard call to make. Yeah. Sorry. I didn't mean to wrap where were you going with your No, I was I think yeah, I was just like Brambling in my attempts to make predictions about these super dogs You're like asking me to that mean I feel like you know In a in a in a in a world that had anything remotely like its priorities straight This stuff is not me like ex -temperizing on a blog post there are like 1000 papers that were written by people who otherwise became philosophers writing about this stuff instead But you know your world does has not set its priorities that way is an unconcerned that it will not set them that way in the future And I'm concerned that if it tries to set them that way will end up with like garbage because the good stuff was hard to verify But you know separate topic Yeah, on that particular intuition about the dog thing But like I understand your intuition that we would end up in a place that is not very good for humans That just seems so hard to reason about that I honestly would not be surprised if it ended up like fine for human In fact the dogs wanted like good things for humans loved humans Like worse murder than dogs we love them Um, the sort of reciprocal relationship came about.
3:24:12I don't know I feel like maybe I could do this given Thousands of years to breed the dogs in a total absence of ethics But it would actually be easier with the dogs I think than with gradient descent Because I think it's well because the dogs are starting out with neural architecture very similar to human um and natural selection is just like a different idiom from gradient descent In particular in terms of like information bandwidth and but like like I'd be as tearing too to like breathe the dogs into like very Like genuinely very nice human and like knowing the stuff that I know that We're uh your typical dog weeder might not know when they set out to be embarked on this project I would be like early on if being like You know like sort of prompting them into the weird stuff that I have expected to get started later And trying to observe how they went during that That this is the alignment strategy.
3:25:01We need a utter smart dogs to help us solve wait there's no time Okay, so uh I think we sort of articulated our intuitions on that one um Here's another one that's not something I came into the conversation with Like some of my intuition here is like I know how I would do this with dogs And I think you could like ask Open AI to describe their theory of how to do it with dogs and I would be like oh wow that's you're gonna get sure sure It's gonna get you killed And that's kind of how I expected to plan in practice Actually, do you mind if I ask like it but when you talk to the people who are in charge of these labs What do they say?
3:25:38I do they just like not rock the arguments. You think they talk to me There was a certain selfie that was taken by five minutes of conversation first time any of the people in that selfie Admet each other and then did you like bring it up or I asked him to change the name of his count of his corporation to anything but open AI Uh -huh. Have you like um Seeked an audience with the leaders of these labs who are you explain these arguments? No Why not? I did try to I've had I've had a couple of conversations with like Demis is obis Who struck me as like much more of the sort of person who is possible to have a conversation with I guess it seems like it would be more dignity to explain Even if you think it's not going to be fruitful ultimately the people are like most likely to be influential in this race I my basic model was that they wouldn't like me and that things could always be worse Fair enough Um that you know, I mean like they sure they sure could have what I mean They sure could have asked at any time But you know that would have been like quite out of character And the fact that it was quite out of character is like I might why I myself did not like go trying to like barge into their lives and getting them at me But you think them getting mad at you would make things worse.
3:26:56It can always be worse I agree that that you know like possibly at this point some of them are mad at me, but you know Yeah, I I have yet to turn down the leader of any major AI lab who has come to me asking for advice Fair enough Okay, so So I'm the same like big picture disagreements like why I'm still not on the greater than 50 % doom It just seemed like From the conversation it didn't seem like you were willing or were able to make predictions About the world short of doom that would help me Distinguish Highlight your view about other views. Yeah, I mean the world heading to this is like a whole giant mess of complicated stuff Which predictions about which can be made in virtue of like spending a whole bunch of time staring at the complicated stuff Until you understand that specific complicated stuff and and making predictions about it like for my for my purpose Yeah, from my perspective like The way you get to my point of view is not by having a grand theory that reveals how things will actually go It's like taking other people's overly narrow theories and poking at them until they come apart and you're left with in a Maximum entropy distribution over the right space which which looks like yep That's that's you're gonna randomize the solar system But to me it seems like the nature of intelligence and what it entails is even more complicated than this sort of geopolitical or Economic thing is that would be required to predict whether world's gonna look like a strong I think that the like theory of Yeah, I think the theory of intelligence is just like flatly not that complicated Maybe maybe that's just like the voice of like person with talent in one area but not the other But that sure how it feels to me This would be even more convincing to me if We had some idea what the pseudocoder circuit for intelligence look like and then you could say like oh This is what the pseudocode implies.
3:28:51We don't even have that I mean What is the axi Um You have the Solmanoff prior over your environment. Yeah, update it on the evidence And then max sensory reward Uh, okay, so it's actually it's not actually trivial like actually this thing will like Exhibit we're discontinuities around its Cartesian boundary with the universe. It's not actually trivial But but but but but but like everything that people imagine as the like hard problems of intelligence are contained in the equation if you have a hybrid computer I'm yeah fair enough. I but I mean in this sort of sense of you know Programming it in like a normal Uh, like I give you a goofad or I give you a really big computer write the pseudocoder something like that for I mean if you give me a hybrid computer Yeah, so what you're saying what you're saying here is that like the theory of intelligence is really simple in an unbounded sense But as soon as you like yeah, what about this like depends on the difference between unbounded intelligence So how about this you asked me do you understand how fusion works if not how can you predict the Let's see, we're talking like the 1800s.
3:30:03How can you predict how powerful a fusion bomb would be and I say well Listen if you put it in a pressure, I'll just show you the Sun and the Sun is sort of the archetypal Example of fusion is and you say no, no, no, I'm asking like what would a fusion bomb look like You see what I mean Not necessarily like What is it that you think somebody ought to be able to predict about the road ahead So so first of all like if you um one of the things if you know the nature of intelligence is just like how will this sort of Progress and intelligence look like well, you know How our abilities gonna scale if at all How fast and it looks like a bunch of details that don't Easily follow from the general theory of like you know like simplicity prior Bayesian update argmax This is again, so then the only thing that follows is the wildest conclusion Which is you know what I mean like there's no like simpler conclusions to follow like the Eddington looking And confirming special relativity it just like the wildest possible conclusion is the one that follows Yeah, like the the convergence is a whole lot easier to predict than the pathway there I'm sorry, but and I sure wish it was where otherwise but And also remember the basic paradigm from my perspective.
3:31:24I'm not making any brilliant startling predictions I'm poking at other people's incorrectly narrow theories until they fall apart into the maximum entropy state of doom There's like thousands of possible theories most of which have not come about yet. I Don't see it as strong evidence that because you haven't been able to identify a good one yet That oh somebody so I mean if somebody in the profoundly unlikely event that somebody came up with some incredibly clever grand theory That explained all the properties gpt5 ought to have which is like just flatly not gonna happen It's just like that kind of info that's available You know my hat would be off to them if they wrote down their predictions in advance And if they were unable to like grind that theory to produce predictions about alignment which seems like Even more improbable because like what do those two things have to do with each other exactly But like still you know like I mean mostly be like well Looks like our generation has its new genius.
3:32:19How about if we all shut up for allow and listen what they have to listen to what they have to say How about this let's say I'm Somebody comes to you and they say I have the best in US theory of economics Everything before is wrong But they say in the year one does not say everything before is wrong one says One when predicts the following new phenomena and on rare occasions say that old phenomena were organized incorrectly fair enough So they say old phenomena are organized incorrectly. Yeah, because of the and then Here's let us term this person scott sumner for the sake of simplicity They say in the next 10 years There's going to be a depression that is so bad that is going to destroy the entire economic system.
3:33:05I'm not talking just about something that Is a hurdle it is like literally civilization will class because it's an economic disaster and then you ask them Okay, give me some predictions Before this great catastrophe happens about like what this theory implies and then they say like listen There's many different branching paths, but they all converge at civilization collapsing because of some great economic crisis I'm like, I don't know man like I would like to see some predictions before that Yeah, I It sure yeah wouldn't it be nice wouldn't it be nice So we're left with your 50 % probability that we win the lottery and 50 % probability that we don't because nobody has like a theory of lottery tickets That has been able to predict to what numbers get drawn next I don't agree with the analogy that it's it's all about the probability It's it's all about the space over which you're you're uncertain We are all quite uncertain about where the future leads but over which space and it and there isn't there there isn't a royal road There isn't there isn't a simple like I found just the right thing to be ignorant about it's so easy The chance of a good outcome is 33 % because they're like one possible good outcome and two possible bad outcomes The stuff that you do when you're uncertain is is like Like the the the the thing you're trying to fall back to In the absence of anything that predicts exactly which properties gpt5 will have It is your sense that you know a pretty bad outcomes kind of kind of weird right?
3:34:36It's probably a small sliver of the space. It seems kind of weird to you But that's just like imposing your natural English language prior like your natural humanist prior on the space of possibility And being like I'll distribute it my max entropy stuff over that Okay, can you explain that again? Okay, what is the person doing wrong who says 50 50 either or win the lottery or I won't They have the wrong distribution to begin with over possible outcomes Okay, what is the person doing wrong who says 50 50 either will get a good outcome or a bad outcome from AI They don't have a certain good theory to begin with about what the space of outcomes looks like Is that is that your answer is that your model of my answer My answer, okay But all the but all the like things you could say about a space of outcomes are an elaborate theory and you haven't predicted gpt4's exact properties in advance Shouldn't that just leave us with like good outcome or bad outcome 50 50 People did have theories about what gpt4's And like if you look at the scaling laws right like the Yeah, yeah, there's got Probably falls right on the sort of curves are Yeah, the 2020 or something.
3:35:48Yeah, the the loss the loss on text predictions Sure, that followed a curve, but which abilities would that correspond to? I don't not familiar with anyone who called that in advance What good doesn't know to know to the lost you could like you could have taken those exact lost numbers back in time 10 years and been like what does that what kind of what kind of commercial utility does this correspond to and they would have given you utterly blank looks And I don't actually know if anybody who has a theory that gives something other than a blank look for that All we have for the all we have for the observations everyone's in that boat all we can do or fit the observations I mean So like also like there's just like me start starting to work on this problem in 2001 because it was like super predictable going to turn into an emergency later And in point of fact like nobody else ran out and like immediately Start getting worked on on the problems And I would claim that as successful successful prediction of the grand lofty theory you had Is um did you see deep learning coming as the main paradigm?
3:36:44Nope And is that relevant as part of the Picture of intelligence? I mean, I would have I would have been like much much much more worried in 2001 if I'd seen deep learning coming You know not in 2001. I just mean before it became like the obviously the main paradigm of the eye now It's like the details of biology is it's like asking people to like predict what the Organs look like in advance via the principle of natural selection and you like it's it's what pretty hard to call in advance You like afterwards you can look at it and be like yep this like Sure does look like it should look if this thing is being optimized to reproduce But the space of the solution of things that biology can throw at you is just too large Like there's it's very rare that you have a case where there's only one solution That that's the thing reproduce that you can predict by the theory that it will successfully that it will have successfully Reproduced in the past and most at least just this enormous list of details and they do all fit together in retrospect The theory actually it is a sad truth that contrary to what you may have learned in science class as a kid There are generally super important theories where you can totally actually Valacy that they explain the thing in retrospect and yet you can't do the thing in advance Not always not everywhere not for natural selection there are advanced predictions you can get about that Given the amount of stuff we've already seen you can like go to a new animal and a new niche and be like oh like it's gonna have like this Properties given what the stuff we've already seen the niche But you know you could also make that by like fine gender there's there's advanced predictions that they're they're like a lot harder to come by Which is why natural selection was like controversial theory in the first place It wasn't like gravity people were being like like Gravity had all these like awesome predictions.
3:38:32We we the newton's theory of gravity and all these awesome predictions We we like got all these like extra planets that that people didn't realize ought to be there We like we like figured out Neptune was there before we found it by telescope Where is this for Darwinian selection people actually did ask at the time and the answer is it's harder And sometimes it's like that in science. I know the difference is The theory of Darwinian selection Seem is much more Well developed well that's sure Then like there were precursors of Darwinian selection that I don't know who was that bromine Poet Lucretius right he had some poem where there was some Precursor of Darwinian selection and I feel like that is probably our level of maturity when it comes to I mean if you want we don't have like a theory of intelligence We might have some hints about what it might look like I got well I've got our hints and if you want to be like but that from hints It seems harder to extrapolate very strong conclusions That they're not very strong conclusions is the message I'm trying to say here I'm pointing to you're being like maybe we might survive and like whoa that's a pretty strong conclusion You've got there let's weaken it That's the that's the basic paradigm I'm operating under here Like you're in your in a space that's narrower than you realize when you're like well You know if I'm kind of unsure maybe there's some hope Yeah, I think that's a good place to Close the discussion on a iLS Well, I do kind of want to like mention one last thing which is that again like in historical terms If you look out the actual battle that was being fought on the block It was me going like Like I expect there to be AI systems that do a whole bunch of different stuff and Robin Hansen being like I expect there to be a to be a whole bunch of different AI systems that do a whole bunch different bunch of stuff But that was one particular debate with one particular person and yeah But like your planet having made the strange reason given its own widespread theories to not invest massive resources and having a much smarter version Well, not smarter a much larger version of this conversation As it thought deemed apparently deemed prudent given the implicit model that it had of the world Such that like I and if such that like I was investing a bunch of resources in this and that kind of dragging Robin Hansen along with me though he he has like Did have like own separate line of investigation into this into into topics like these You know like being there as I was my model having led me to this important place where the rest of the world Apparently thought it was fine to go let it go hang the you know such debate as there actually was at the time was like Will yeah, we really gonna see like these like singlaeye systems that do all this different stuff is this like hold general intelligence notion kind of like meaningful at all And I I staked out the bold position for it actually was bold and people did not all say like oh Robin Hansen you fool why do you have this exotic position?
3:41:34They were going like ah like behold these two luminaries debating or but to be Hold these two idiots debating and like not massively coming down on one side of it or other So you know like in historical terms I I dislike You know Making it out like I was right about anything when I feel I've been wrong about so much and yet I was right about anything and you know relative to the Relative to the what the rest of the planet themed that important stuff to Spend its time on given their implicit model of what was going to what was how it's going to play out what you can do with minds where AI goes I think I think I did okay I think brand one did better shame like our duetly did better Gourren always is better Oh, okay, I think um, I mean obviously like if you get a better of a debate that like cons or something but I debate with one particular person Well considering your entire planet's decision to invest like ten dollars into this entire field of study apparently one debate is all you get And that's what the evidence you got to update on now I so somebody like Elias that's cover, you know When it comes to the actual paradigm of deep learning like was able to anticipate that like from image net to Scaling up all the ones or whatever uh There's people with track records here who are like who disagree about Doom or something So In some sense it's probably more people who have been if ilia challenged me to a beta I wouldn't turn it down.
3:43:09I admit that I did specialize in doom rather than LLNs Okay, fair enough. I unless you have other sorts of comments on AI. I'm I'm happy with Yeah, and again, I'm not being like Due to my miraculously precise and detailed theory. I am able to make the surprising and narrow prediction of doom I am like I am being like like I think I did a fairly good job of shaping my ignorance to lead me to not be too stupid despite my ignorance over time as it played out and you know there's a Prediction even knowing that little that can be made Okay, so this feels like a good place to um Pause the i conversation and there's many other things to ask you about given your decades of writing and millions of words So I think what some people might not know is the millions and millions and millions of words Of science fiction and fan fiction that you've written.
3:44:10I want to understand what when in your view is it better to explain Something through fiction than not fiction when you're trying to convey experience rather than knowledge Or when it's just much easier to write fiction and you can like produce a hundred thousand words of fiction with the same effort It would take you to produce 10 ,000 words of nonfiction. Those are both pretty good reasons Yeah, on the second point It seems like when you're writing this fiction not only are you in your case covering the same heady topics that you include in your nonfiction But there's also the added complication of plot and characters It's surprising to me that that's easier than just verbalizing the sort of the topics themselves Well partially cause it's more fun Is an actual factor?
3:44:55Mm -hmm ain't gonna lie And sometimes it's something like A bunch of what you get in the fiction is just like The lecture that the character would deliver in that situation the thoughts the character would have in that situation Yeah, there's like only like one piece of fiction of mine Like there's literally a character giving lectures because he arrived another planet and now has to lecture about science to them That that one is project lawful you know about project lawful. I know about it. I have not read it yet Yeah, okay, so so most of my most of my fiction is not about somebody arriving in another planet who has to deliver lectures there Um, I was being like a bit deliberately like Yeah, I I'm gonna just do it with like project lawful like I'm just do it They say nobody should ever do it and I don't care.
3:45:46I'm doing it ever ways. I'm gonna have my character actually launched into lectures Me know like the lectures aren't really the parts I'm proud about it's like where you have like the like life or death Death note style battle of wits between like the That that is like centering around a series of Bayesian updates Um, yeah, and and like making that actually work because you know is where I'm like yeah I think I actually pulled that off and I don't think I'm not sure a single other writer on the face of that's plat it could have made that work as a plot device Um, but that said like the nonfiction is like I'm explain this thing I'm explained the prerequisites I'm explained the prerequisites to the prerequisites and then in fiction It's more just like well this character happens to think of this thing and the character happens to think of that thing But you got to actually see the character using it.
3:46:38Mm -hmm Yeah, so it's less organized It's less organized as knowledge and that's why it's easier to write yeah I I mean what am I Uh favorite pieces of fiction of fiction to explain something is uh The Dark Lord's answer and I honestly can't say anything about it without spoiling it But I just want to say like honestly it was like such a great explanation of the thing it is explaining Um, I don't know what else I can say about it without spoiling it anyways. Yeah, but I'm laughing because I think like relatively few have Dark Lord's answer as there as like among their top favorite works of mine It is what is one of my less widely favored works of mine Actually what is my favorite sort of this is a medium by the way I don't think is uh use enough given how effective it was in In a nidatica de equilibria you have different characters Just explaining concepts together some who some of whom are purposefully wrong as examples And that is such a useful pedagogyical tool and I don't know honestly like at least half a block post should just be written that way It is so much easier to understand that yeah, and it's easier to write and I should probably do it more often And like you should give me a stern look and be like Ellie us or write that more often Done Ellie is there please I Think 13 or 14 years ago you wrote an essay called rationality systematized winning Mm -hmm Would you have expected then that 14 years down the line um the most successful people in the world They're some of the most successful people in the world would have been rationalist Only if they whole rationalist business had worked like closer to the upper 10 % of my expectations Then it actually got into there's the title of the essay was not rationalists our systematized winning Wasn't even a rationality community back then rationality is not a creed It is not a banner It is not a way of life It is not a personal choice It is not a social group It's not really human It's A structure of a cognitive process and you can Try to get a little bit more of it into you and if you Wanted to do that and you fail then Having wanted to do it doesn't make any difference except in so far as you succeeded Hanging out with other people who Share that creed going to their parties and only ever matters in so far as you get that's a bit more of that structure into you And this is apparently hard This seems like a no true Scottsman Kind of point because and yes, there are no true basins upon this planet Yeah, but do you really think that had as people Tried much harder to adopt this sort of vision principle is that you laid out They would have many of the successful people and some of the successful people in the world would have been rationalists What good does trying do you accept in so far as you are trying at something which when you try it it succeeds Is that an answer to the question?
3:50:05It's it's a rational rationality is systematized winning It's not rationality the life the life philosophy. It's you know, it's not like trying real hard at like this thing This thing and that thing it's it was like the mathematical sense Okay, so then the question becomes does adopting the philosophy of Bayesianism consciously actually to you having more Concrete wins well, I think it did for me though only in like scattered bits and pieces of slightly greater sanity than I would have had without explicitly Recognizing and aspiring to that principle the principle of not updating in a predictable direction The principle of jumping ahead to where you can predictably be where you will predictably be later I look back and you know kind of I mean the the the the story of my life as I would tell it is a story of my jumping ahead to what people would predictably You know like Believe later after reality finally hit them over the head with it this to me is the entire story of the like Like people running around now in a state of frantic emergency over something that was utterly predictably going to be an emergency later as of 20 years ago and you could have been Trying stuff earlier, but you know, yeah Yeah, he left it to me at a handful of other people and and uh It turns out that that that was not a very wise decision on humanities part because we didn't actually solve it all And I don't think that I could have like tried even harder or like contemplated probability theory even harder and done very much better than that I contemplated probability theory about as About as hard as the biology I could visibly obviously get from it.
3:51:47I'm sure there's more There's obviously more, but I don't know what it let me save the world I guess my question is is contemplating probability theory at all in the first place Something that tends to lead to more victory. I mean, I imagine Who's a British person in the world like how often does Elon Musk think in terms of Probabilities when he decided what to do and here's somebody who is very successful Um, so I guess I guess a bigger question is In some sense when you say like rationality system, I think it's like a topology if the definition of rationality is whatever helps you when If it's specific principles laid out in the sequences then the question is like have this uh I do the successful people most still people in the world practice them I think you are trying to read something into this that is not meant to be there all right It is the notion of rationality systematized winning is meant to stand stand in contrast to a long philosophical tradition of like notions of rationality that are Not meant to be about the mathematical structure not meant to be or or like about like strangely wrong mathematical structures Where you can clearly see how these mathematical productions will structures will make predictable mistakes It was meant to be saying something simple some um There's uh there's an episode of Star Trek Where in Kirk makes a 3d chess move against Spock and Spock loses and Spock Complains that Kirk's move was irrational Rational towards the goal.
3:53:20Yeah, the literal winning move. Yeah. Is irrational or possibly possibly illogical Spock might have said I might be misremembering this Like the thing I was saying is not merely That's wrong that that's like a fundamental misunderstanding of what rationality is There there there there is more death to it than that, but that is where it starts There are there like so many people on the Internet in those days possibly still who are like Like well, you know if you're rational you're gonna lose because other people aren't always rational and This is not just like a wild misunderstanding, but there's like particularly there's what like like the contemporarily accepted decision theory in Academia as we speak at this very moment causal decision theory classical causal decision theory um Basically has this property where like you can be irrational and The Rational person you're playing against is just like oh, oh, I guess I lose then here here have ought here Have most of the money I have no choice but to and Ultimatum games specifically If you look up logical decision theory on orbital you will find a different analysis of the ultimatum game or the rational players do not predict Predictively lose the same way as I would define rationality and if you Take this sort of like deep Mathematical thesis that also Runs through all the little moments of every day life when you may be tempted to think like like well If I do the reasonable thing Won't I won't I lose that you're making the same mistake as the the Star Trek script writer who had spock complain that Kirk had one that had one the chess game irrationally That like that every time you're tempted to think like well like here's the reasonable answer and here's the correct answer You have made a mistake about what is reasonable and if you then Try to screw that around as like rationalist should win rationalist should have all the social status Whoever's the top dog in the present social hierarchy or the planetary wealth distribution must have the most of this wealth Must have the most of this math inside them.
3:55:45There are no other factors But how much of a fan you are of this man? That Trying to take the deep structure that can run all through your life in every moment where you're like oh wait Like maybe the move that would have gotten the better result was actually the kind of move I should repeat more in the future Like to take that thing and like turn it into like like Social dick measuring contest time Rationalists don't have the biggest dicks Okay Okay final question This has been I don't know how many hours I really appreciate you doing giving me your time final question I know that in a previous episode you were not able to give Specific advice of what somebody young who is motivated to work on these problems should do Do you have advice about How one would even approach coming up with an answer to that themselves There's people running Programs to try to who think we have more time who think we have better chances and they're running programs to try to Nudge people doing Nudge people into doing useful work in this area and I'm not sure they're working and there's Such a Strange road to walk and not a short one And I tried to help people along the way and I don't Think they got far enough like some of them got some distance, but they didn't Turn into alignment specialists doing great work and It's it's the problem of the broken verifier if somebody had a bunch of talent in physics And we're like well like I want to work in this field.
3:57:38I might be like well, there's interpretability And you can like tell a whether you've made a discovery and interpretability or not Which sets it apart for a bunch of the solder stuff and I don't think that that saves us and okay, so how do you do the kind of work that saves us and And I I don't know how to convey the and the key thing is the ability to tell the difference between good and bad work and Maybe I will write some more blog posts on it. I don't really expect the blog post to work and the the critical thing is Is the verifier? How can you tell whether you're talking sense or not whether you're I there's there's a kind of specific heuristics I can I can give I can be like I can say to somebody like Well, it's like your entire alignment proposal is this like elaborate mechanism you have to explain the whole mechanism And you can't be like here's the core problem Here's the key insight that I think addresses this problem if you can extract that out if your whole solution is just a kind giant mechanism This is not the way It's it's kind of like how people invent perpetual motion machines by making the motion perpetual motion machines more and more complicated until they can no longer keep track of how it fails And if you actually had somehow a perpetual motion machine it would not just be like giant Machine there would be like a thing you had realized that made that made it possible to do the impossible for example You're just not gonna have a perpetual motion machine So like there's there's thoughts like that I could say like go go study evolutionary biology because evolution or biology went through a phase of Optimism and people naming all the wonderful things they thought that evolution or biology would cough out like all and Like all the wonderful things that they thought wonderful properties that they thought natural selection would be into organisms and The Williams revolution edges sometimes called is when George Williams wrote adaptation and natural selection a very influential book saying like that is not what this Optimization criterion get gives you you do not get the pretty stuff.
3:59:47You do not get the aesthetically lovely stuff Here's what you get instead and by like living through that revolution vicariously Well, I thereby picked up a bit of like thing that out that to me obviously generalizes about how not to expect nice things from an alien optimization process But maybe somebody else can read through that and like not generalize not generalize in the correct directions Then how do I advise them to generalize in the correct direction? How do I advise them to learn the thing that I learned I can just give them the generalization But that's not the same as having the thing inside them that generalizes correctly without anybody standing over their shoulder enforcing them to get the right answer the right answer I could point out and have in my fiction that the entire schooling process of like here is this Legible question that you're supposed to have already been taught how to solve Give me the answer using the solution method you are taught that this does not train you to to Detackle new basic problems But even if you tell people that like okay, how do they retrain we don't have a systematic training method for producing real science in that sense We we have like half of the what was it quarter of the Nobel laureates being the students or grand students of other Nobel laureates Because we never figured out how to teach science.
4:01:09We have an apprentice system We have people who like pick out people who like they think can be scientists and they like hang around them in person And something that we've never written down in a textbook passes down and The and that's where the revolutionaries come from And there are whole countries trying to invest in having scientists and they Turn out these people who write papers and and none of it goes anywhere because the part that was legible to the bureaucracy is have you written the paper Can you pass the test and this is not science? And I could go on this go on for this for a while that the thing that you asked me is like How do you Pass down this thing that your society never did figure out how to teach And the whole reason why Harry Potter and the methods of fashionality is popular is because people read it and and picked up the rhythm Seen in a character's thoughts of the thing that was not in their schooling system That was not written down That you would ordinarily pick up by being around other people and I managed to put a little bit of it into a fictional character and people picked up a fragment of it by being your fictional character but you know like Not in really vast quantities um Not not vast quantities of people and I didn't managed to put vast quantities of shards in there I'm not sure there there's not a like long list of noble laureates who've read hpm .r Although there wouldn't be because they like delay times on grounding the prizes are too long um it's Yeah, like you you asked me what do I say and and my answer is like well that's a whole big gigantic problem I've spent however many years trying to tackle and I ain't gonna solve the problem with the sentence in this podcast Fair enough Eliezer Thank you so much for giving me I don't know how many hours of your time This was really fun Hey, everybody I hope you enjoyed that episode As always the most helpful thing you can do is just share the podcast Then did to people you think might enjoy it put it in twitter your group chats etc just blitz the world Appreciate your listening.
4:03:12I'll see you next time. Cheers
From the publisher
For 4 hours, I tried to come up reasons for why AI might not kill us all, and Eliezer Yudkowsky explained why I was wrong.
We also discuss his call to halt AI, why LLMs make alignment harder, what it would take to save humanity, his millions of words of sci-fi, and much more.
If you want to get to the crux of the conversation, fast forward to 2:35:00 through 3:43:54. Here we go through and debate the main reasons I still think doom is unlikely.
Watch on YouTube. Listen on Apple Podcasts, Spotify, or any other podcast platform. Read the full transcript here. Follow me on Twitter for updates on future episodes.
Timestamps
(0:00:00) - TIME article
(0:09:06) - Are humans aligned?
(0:37:35) - Large language models
(1:07:15) - Can AIs help with alignment?
(1:30:17) - Society’s response to AI
(1:44:42) - Predictions (or lack thereof)
(1:56:55) - Being Eliezer
(2:13:06) - Othogonality
(2:35:00) - Could alignment be easier than we think?
(3:02:15) - What will AIs want?
(3:43:54) - Writing fiction & whether rationality helps you win
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe




