#47 - Emmett Shear - Why NATURE Holds the Answers To AI Alignment

25 Sep 2025 · 1 h 43 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Win-Win Podcast Episode #47: Emmett Shear - Why NATURE Holds the Answers To AI Alignment

Podcast Overview Host: Liv Boeree Guest: Emmett Shear (Co-founder of Twitch, former interim CEO of OpenAI, Founder of Softmax) Episode Theme: Discussing the future of AI alignment through the lens of organic alignment inspired by nature.

Episode Summary In this episode, Liv Boeree and Emmett Shear engage in a thought-provoking conversation about the rapidly approaching era of superintelligent AI and the challenges of ensuring its safe alignment with human values. Emmett critiques the traditional "control alignment" approaches and proposes an "organic alignment" model inspired by natural systems and human societies.

Key Concepts Discussed

  1. Control Alignment vs. Organic Alignment
  2. Control Alignment: Attempts to hard-code values into AI. Emmett argues this is destined to fail due to the non-stationary nature of the world and moral values.
  3. Organic Alignment: A model where AI learns to align organically with human values through interactions, similar to how cells cooperate in multicellular organisms.
  1. Formation of Collective Identities
  2. Emmett discusses how multicellularity serves as an example of alignment, where individual cells with their own goals contribute to a larger coherent entity.
  3. Emphasizes that humans have a unique capacity for alignment, which surpasses any other species.
  1. Alignment Techniques
  2. Building AI that understands mutual benefits and can form a "we" identity, fostering cooperation among diverse agents.
  3. Importance of creating environments that encourage alignment through collective behaviors.
  1. Market Dynamics and Superorganisms
  2. Discussion on how the market functions as a global superorganism and the implications of AI's alignment with varied human interests.
  1. Human Capacity for Alignment
  2. Emmett stresses the importance of understanding and leveraging human alignment capacities, which can offer insights into AI alignment.

Key Arguments

  • Morality as an Emergent Property: Emmett argues that morality cannot be programmed into AI; it must emerge from understanding and reflecting on shared experiences.
  • Caution Against AI as "Master": He warns against the notion of AI viewing humans merely as resources or subordinates, emphasizing the need for mutual respect and care.
  • Emphasis on Neighborhood Relationships: Just as humans form communities, AI should be designed to recognize and care for its local contexts, leading to healthier alignments.

Questions Addressed

  • The distinction between Coordination and Alignment: Coordination involves working together without necessarily being on the same team, while alignment is deeper and involves shared goals.
  • Concerns about the potential for AI to see humans as expendable or merely as tools, and how to prevent that through organic alignment frameworks.
  • The potential for self-modification in AI and its implications for ensuring alignment and stability within systems.

Conclusion The episode concludes with a profound discussion on the nature of consciousness, the meaning of alignment, and the vision for a future where AI and humans coexist as partners rather than adversaries. Emmett posits that the future hinges on nurturing relationships where both humans and AIs care for each other, establishing a stable and cooperative existence.

Episode Links

  • [Softmax](https://softmax.com/)
  • [AI Alignment Forum](https://www.alignmentforum.org/)
  • [Evolving Intrinsic Motivations for Altruistic Behaviour](https://arxiv.org/pdf/1811.05931)

Credits Host: Liv Boeree Produced by: Luca de Vico Podcast Links: [YouTube](https://www.youtube.com/playlist?list=PLWgq0OZMtwtOIyMsVM_vksqdfWcM-b68S) | [Spotify](https://open.spotify.com/show/03bGVUaFZmJUmEvSHNDPdI?si=64379cc23696454f) | [Apple Podcasts](https://podcasts.apple.com/us/podcast/win-win-with-liv-boeree/id1724791350) | [Pocketcast](https://play.pocketcasts.com/podcasts/7f708340-d17c-013b-f46e-0acc26574db2)

---

This markdown file summarizes the pivotal points discussed in the podcast episode, providing a clear and accessible outline for readers interested in the topics of AI alignment and organic systems.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00And humans are actually, we're wild. We're better at alignment than any other species on the planet. In fact, I would argue that our capacity for alignment, more than our capacity for intelligence or almost anything else, is what sets us apart. Corporations are cancer. The hippies are right. What is a cancer? A cancer is a cell. It's an organism that has to take care of itself, that the system won't look out for it, and that it is effectively alone, and it must grow forever and build its own safety. It's literally what corporations do. They don't have a somatic form. At no point do they finish development.

0:32Training on morality is a bad idea. Ethics professors who like study ethics are not more ethical. Ethics is learned by observing the truth. The truth is that acting like a dick to the people around you is not good for you.

0:47All right, so thank you all for coming. Super stoked to have this conversation. We've got, I'm joined by Emmett Shear. Emmett, I mean, your CV is absurd. You founded Twitch, JustinTV, you were CEO of OpenAI for a very brief moment, and now you are the CEO and founder of an organization called Softmax, which, as I understand it, is trying to solve AI alignment through a very novel approach you call organic alignment. So to get us started, explain what organic alignment is and why you think that is what we need to be doing. Yeah, so thank you for having me. When I first started looking for work in AI was after I retired from Twitch and I started this idea that I was going to take it easy and do some individual research into AI.

1:47And so I tried to learn how everything works because it was obviously sort of the interesting thing I wanted to work on. And that was actually very nice. I really enjoyed that brief, I don't know, like seven months period of my life. And then the open AI thing happened. And I never had the experience of receiving like a sign before, like the universe is like, hello, you, specifically you, you're supposed to be doing like knock knock. Because being in the middle of that, it just became very obvious that the people involved at all levels and all sides weren't concerned with the aspects of it that concerned me which isn't say they were like like dumb or like they weren't thinking about it it's like the the questions they were interested in the way they were thinking about it i was like oh it's like really obvious just i've now seen at the very highest levels and sort of through the company and talk to those people that there's this there's this whole giant swath of things that i thought were important that i kind of i kind of knew no one was working on but you You know, you think that they must have a secret project where they're working on it, and then there's no secret project where they're working on it, and you're like, oh, oh, I see.

3:01Okay, well, if I don't do this, then there's not anything that's going to happen. And so I started looking for how to buy help for this problem. And the problem sort of at a really high level is like, we're going to make these AI systems that are going to be very powerful. And you know, you make a very powerful thing, it will get used to do stuff. and if that stuff is things that are good, then that's good and if it's bad, it's bad. And someone should do something about that. Like someone should have a real plan, not like kind of one of those like we'll kind of go by intuition and gut kind of plans but like a plan where like you could measure your progress against it and like, like if you wanted to go to Mars, you'd build like a sequence of rockets that got bigger and bigger and so you finally made the Mars side rocket.

3:50And so I really started working on the question, what would it mean for an AI to be aligned? Which was hard about something being aligned. What does it mean for something to be aligned? Which felt more feasible than the question of like, what is goodness? Which is the actual question you have to answer. So I told myself that I only had to figure out alignment. It turns out you actually have to solve morality more or less. not like not solve morality in the sense of like physics hasn't solved emotion we still have trouble making things move the way we like but like there's sort of this moment where you systematize your approach to emotion you have the idea of systematicity and the challenge is to systematize our approach to morality to systematize our approach to knowing the good knowing the right way because if you can't be systematic about about that then when you build the really powerful thing you'll point it at something, when you point it at the thing, you just kind of hope that it's good.

4:44And if it's not, everything goes bad. And so I started really seriously working on alignment. And what became clear is that what most people called alignment was alignment. It was what I would now call control alignment or steering alignment. You have some target you've picked out, which you've defined in some way, which can be defined as like, do whatever that guy says, or like follow the set of rules or like follow this particular algorithm for determining the good and then do whatever it's at, do whatever this, you know, you infer to be true given the set of rules. And the problem with control alignment is that people have been trying for a really, really long time to like get a stable, useful definition of the good that like always produces the right results.

5:30And it just doesn't seem to work that way. And it's pretty obvious that it doesn't work this way because the good itself isn't stable. Because we aren't stable. Because the world is non-stationary. Things keep changing. Whatever you wrote down before might have been a... If you're lucky, it was a great approximation then and even a good enough for approximation then. But it's not going to stay that way. And the problem is that the world changes in ways that your model can't anticipate and never does anticipate. Which is why you have the saying, models are always wrong, sometimes useful. Your model isn't reality.

6:03It's just like your best guess. and so it became clear we needed a different kind of alignment and where i where i wound up going was to ask this question okay so you if you're not doing control if the goal is not a system of control if the goal is not to enslave our future ai children um then i'm sorry that's like i can't help but put the jab in but like like if the goal was something but and yet to get along like what would that look like and the answer is like well we're we're doing it everywhere right now like this is humans are actually really good at this and nature's evolution is really good at this game multicellularity is alignment it's a bunch a bunch of cells that all their own individual goals and yet here i am here you are most one mostly coherent person made it of 28 trillion parts and somehow all of their individual goals like i don't experience the world like i made of 28 trillion parts i experience it like i'm no not entirely you know you have like parts inside of a little bit, but not like 28 trillion of them, far fewer than 28 trillion.

7:02And you experience this going up too, because when you're on a team, a really tight-knit team, you can feel like there's a we, there's like a thing going on where you're... You know, you can ask yourself in any moment on a team, like, what do we want and what do I want? And you get different answers. And you know what the we wants. You're very aware of it. Who's the we? Where is this knowledge coming from? Well, about this collective, which has goals, which is not different than the inference you make about another human. You don't see inside someone else to know what they want. You just watch their behavior and you infer, given that behavior, I infer they have these goals.

7:39Well, given the collective's behavior, I infer it has these goals. And that's organic alignment. Organic alignment is this process by which it seems possible that a bunch of agents that interact with each other eventually wind up in kind of like an orbit around each other. Their beliefs form an orbit where my beliefs influence your beliefs, influence my beliefs. And we find ourselves in homeostasis, where we are stably in interaction with each other over and over again. And when that happens, you get to make this inference. There's an object we're in orbit around. There must be some principle here.

8:15So there's some shared thing. There has to be. Almost definitionally, there's some statistical invariant that's being maintained that's causing this. And that's the we. And what you realize is that you are utterly dependent on the we. You're born to a family. Agents don't come from nowhere. There's never been an agent in the history of the universe where you look backwards and it was, this agent just appeared out of nowhere. It has no purpose. No, no. You were made by something that had a goal in mind for making you. It wanted to perpetuate itself. It wanted to have a robot servant to go do things.

8:52It's like you were made for a reason. You may not know what that reason is, but you have a purpose. There is a teleological nature to existence for all agents because they don't just pop out of nowhere. They always come out of other agents. And agents, what makes you an agent is that you have goal-oriented behavior. So since you're a result of agent behavior, you're a result of goal-oriented behavior. And when you see this, it's like, oh, okay. Well, that's the thing we're trying to learn. We're trying to learn what does it mean to be a family? What does it mean to be a we? What does it mean to be a society?

9:23And that trick, can you get the thing to be part of a society, part of a team in a stable way, that's the core of it. That's the heart of Organic Climate. So it sounds like you're referencing nature heavily. You see nature as a thing that, and I would agree, like the cells in our bodies, you know, all my stomach cells, They're working together very well to do the job of my stomach. And that seems, though, like fairly, to me, like I wouldn't expect my cells to have almost individual identities, right? Because they actually share the identical genetic code or nearly identical genetic code. So if you're talking about a group of agents that are, you know, digital agents that have their own separate identities, how would that same kind of like emergent coordination happen between them if they don't share that same genetic code as one another?

10:25I have a couple sort of answers to this. One's the cheeky answer, which is like exactly how much difference, you know, how much difference is required before it's no longer possible to form a collective agent? Like if you want one more gene change, now it now it becomes impossible. Like when you start to look at it, you realize there can't be a there can't be a hard boundary. There's always it's always going to be contextual. But I think deeper than that. You already know organic alignment. You do it every day. You form organic alignment with other humans around you. And humans are actually we're wild.

10:58We're better at alignment than any other species on the planet. In fact, I would argue that this alignment, our capacity for alignment, more than our capacity for intelligence or almost anything else, is what sets us apart. Because human beings almost killed the whales because of our capacity for alignment to organize ourselves into hunting parties and have whole chains of supply chains building ships. It's all alignment problems. But then we looked at the whales and we said, no, actually, you're on my team. You are me. We're not that different. You're part of the mammal team. We're part of the mammal team with you.

11:32And we prefer that you don't die. We prefer that you, we see you as part of us enough that you're not just dead matter to us. It would be, we would be sad. Not like so sad that like we would like, you know, do that much, but like, but like we'd be, we'd be sad. And so we will deliberately like take extra action, do extra work to stop other people from hunting you and be, you know, pretty reasonably, some small number of us will actually be pretty reasonably upset about it. Other species do not do this. They don't, they don't, not that they are necessarily cruel to things around them, they wouldn't get organized to go save another species because they recognize their fundamentally shared nature.

12:08Like other species don't start veganism movements inside of us. It's not a thing that exists in other species. Whether the vegans are right or wrong, it's demonstrate this deep human capacity to see that we are part of this larger thing. This whole moral progress, the arc of moral progress becomes really obviously true, actually. There is moral progress. it's like how big of a scale can you align on effectively like like like how big of a group can you get working in a coherent way that serves and the flourishing of the members that's basically your your cap on alignment like that's how we measure it for the agents like how many agents can we get to can we get to be in a single attractor together and as a good so yeah there's there's moral progress whether you want to call it moral or not alignment progress is the ability to see the way that actually all humans are your ally or could be if you were better at organizing.

12:57They have the potential to be your ally. That's just true. They do. Someone who doesn't believe that is just incorrect about the world. Now, whether they do, whether the actual capacity and actual vision are different, just because someone has the capacity to be your ally and it's important to see that, doesn't mean that they are your ally. I think that some people lose track of this thing. A lot of people who are more universalist, I think, have a tendency to confuse the capacity with the realization. But the your cells do it at this very self-sacrificing level because your cells are mostly clones of each other and their reproductive cycle goes to themselves.

13:36Like ants. Ants are like this. Ants don't have a separate identity from the other ants. The ant colony is the thing with the identity. Humans, I love us, are not like that. We are not ants. We're groups of monkeys that functionally behave like ants on the scale of ants. Yeah, you could imagine that a future of humanity where we turn ourselves in dance. We could do that where there becomes a central reproductive function that decides what new things get made. And it just prints copies of those people over and over again to fill various roles. And it is a singular gradient descent search to improvement on humanity rather than an open-ended one.

14:17It's literally the thesis of Brave New World. Yeah, yeah. And I think that Brave New World actually kind of demonstrates It's why it wouldn't, that society were to have to compete with a more open society would find itself at a grave disadvantage because of the very thing that makes morality hard, that makes alignment hard. Your model of what good people are, your model of what skill is, your model of whatever it is, you pick something, you're not right. There's something about the world. There's something there's, you're going to find out that you put yourself into a cul-de-sac and your model is wrong and you're going to lack the diversity in the population to figure it out.

14:58Like, if you think about it from an evolutionary perspective, you do not want a highly clonal population. It's like very dangerous. You're the cheetahs. You're very fast. You're overfit. You overfit the circumstance and being overfit means you will outperform when the environment's stable and underperform when the environment is moving in a more ambiguous world. and like the thing about building more humans and like having more technology is like we just we just dump entropy into the environment just constantly the world just gets weirder and harder and like it's just this never-ending like and the instant we get good at handling that level of entropy we level up and we start dumping more entropy into the environment making it weirder and harder it's like this endless cycle and so like you're you're you're beautifully designed fixed society where you've figured out exactly the right kind of people to make and your model for how to adapt those people when this kind of change happens and that kind of change happens is just going to get hit in the face by some kind of change you didn't expect and oops um that's why you like the open open societies are better okay so coming back to how you would imbue these principles into into ai what methods because i i know you've been only running for a year or so and this is early days You're still in the research phase.

16:15But like what are some of the promising techniques that you're exploring? I think we have a map now, more or less. Like we don't know how to – we know like we have to go to like that mountain peak and that mountain peak, not like how we get there. But like – So like give me an example. So there's the fact of alignment and then there's like reflective alignment. Just like cells do differentiate, but cells don't tell themselves stories about how they're differentiating. They don't decide, I'm going to be a blacksmith because being a blacksmith is heroic or whatever. They just differentiate into a liver cell.

16:52They read signals, they do it. They have a model of self, but not a model of themselves as thinking selves. They don't have multiple, it's like humans have this very stacked, many, many layers of reflection. The cells do not have this. I think there's a way in which you could say cells are... I think you get better predictive results if you assume that cells are aware and just are very dumb. Just don't experience most of our experience. But they're aware. There's not a lot going on. That's not good enough. The problem is if you have something that is de facto aligned, cells are de facto aligned, when the system changes under variance and disturbance in the system, if they get knocked out of the attractor, they just get knocked out of the attractor.

17:35So attractors aren't very strong. What's cool about humans is that we can anticipate the things that would cause us to become cancer. We can see if I go down this path, I offer you the vampire pill that makes you a vampire and gives you great power and joy and you'll feel really good, but you're going to murder everyone you know and torture them forever. You don't take the vampire pill because you don't want to become that. Even though it might feel good, you don't will that future. But we still behave in this. I mean, you mentioned like acting cancerous. Everyone has different phrases for this.

18:11I call it like Malachian behaviors. You know, these like lose-losey type dynamics. But at the same time, we have all these tragedy of the commons. You asked how you actually solve it, right? So, okay, sorry. Let me answer your question. Okay, so in order for me to be stably aligned with you, not just aligned with you de facto, I need a model of how my actions not only align or don't align to the current goals, I need a model of how my actions change my own learning trajectory and your learning trajectory and how that will impact the attractor in the future and whether or not this leads to the robust flourishing of the attractor or not.

18:49I need to understand the consequences of my actions for the collective dynamics themselves, which you do all the time. Like, I'm going to go out of my way to be nice to this team member because I know that I need them to do this. I'm going to, oh, I can tell that I'm not articulating the visions. This person's drifting. And like, we constantly model how what we're doing impacts the groups we're in. Okay, so to do that, AI models have no sense of self today. The problem with AI models is they're like, they're story simulators. But they can simulate any story. We kind of very mildly condition them to tell certain stories over others, but they've just been conditioned to do that instead of this.

19:28They don't have... They don't really have memories. Yeah, they don't really have memories. You are your memories. Yeah, they don't have a consistent personality. Yeah, and you are a sort of habitual, not only memory, not just behaviors, but habitual ways of learning. Habitual ways that you engage with the world in terms of how do you enter, you have a, you know how you handle this problem. You know how you learn this thing. You know in the situations you find yourselves in. And so that's got, that, you've been sort of, you've been continually learning on your own, your own behavior for a long time.

20:03And that's your self. And then over time you form a model of yourself. You have a model of the kinds of things you do. And that model actually guides most of your behavior. There's this literal point in your brain where eventually the cortex basically starts stepping in instead of the hypothalamus and the basal ganglia. And like, basically it starts driving because it's predictive accuracy gets high enough that it's guess as to what you'll do next based on the story it's telling itself is better than your intuitive feeling guide, which has a higher valence. Because you know that that, you've been down this path and yeah, it feels a little higher valence right now, but we're just gonna go with the thing we know actually works.

20:41We've tested this out. We have enough predictive loss. We know it's a regression to go in. And it's all about modeling yourself because how can you possibly model a we and stably maintain a we if you don't even know what an I is? So step one is you have to get models that have a theory of mind for themselves and then for others at the just base level of like, I can predict, I have a model of my thoughts, what kind of thoughts I have. And then you need a model of the dynamics of mind, which is sort of personality. Like you have the, what kind of thoughts you think is not static. There's like a model of it, but you also have a model of how that model changes, which is your personality.

21:23That's like, that's like you have your model of your thoughts. Your model of your thoughts is like, oh, I'm a person who gets mad easily. And I'm a person who does this. I'm a person who does that. But you also have a model that says, I'm the kind of person who used to be this kind of a person, and I'm now this kind of person, and I will be the kind of person in the future. And that's like many layers of sort of self-reflective models and models of your models. It takes a lot of training and very specific learning circumstances to induce this. If you take a human being, you take them out of a social context, and you raise them in the wild with no other humans, they do not develop these models properly.

21:57It's very damaging. This is not something we... We are predisposed to develop this. We have the equipment, the capacity, the inductive bias. But if you take out away the learning scaffold, we don't, it's not in the architecture. It's in the architecture and the training data. And then you have to put them in a circumstance where not only do they exist and they have stable goal-oriented behavior that they can model, and then they have a good model of their own stable goal-oriented behavior and how their behavior changes. And then again, for everyone else, that they can then also, having all those models, notice that there's a collective model that's made out of those models and that that collective model has its own goals and personality and is changing over time and all of the causal connectivity between all those things.

22:43And the point you've done that, you have something that knows how to do alignment. Like that is what the skill of alignment is, is theory of mind for groups, right? And if you have theory of mind for groups, you can do alignment. It's to really, the reason it doesn't happen automatically is it's fucking complicated. Like it's the hardest, it's the hardest single intellectual challenge. I think that like people are really hard in groups of people, dynamic, managing dynamic groups of leadership. Some of the like single hardest things we do. The only thing that like maybe comes close is like abstract physics or math, right?

23:17Like in terms of challenge. And even that I've done some amount of both. People are harder, man. Like we, we only think that math is hard because we can't, we haven't even tried to really solve people. So like math seems hard because you can actually, we know, we know what the standard of really doing math is. We don't even know what actually being good at people would look like. That's never happened. Well, and especially if, as you say, it's all like frame, different individual frames and it's essentially relative. It's always changing. It's always in flux. Absolutely. So I can conceptualize the kind of environment you would be wanting to build this collective of agents where they can start trying to get this emergent coordination.

23:59Is it games? Is there anyone doing this right now? We are. Yes, someone's doing this. This is what we're doing. So, yeah, at the end of the day, it's a game. It's a virtual world where agents take actions which have consequences, some of which are better or worse than each other. You can call them getting points or reward or whatever you want to call it. And then the ones that are better at the game get to differentially reproduce more, either through some sort of evolutionary approach or because the gradients of their trajectories get pushed into the model more. But there's no difference between reproducing by we make a copy of you versus reproducing because your trajectory, your behavior set gets saved.

24:42It's the same thing. And what have you seen so far? Oh man, it's so hard. We can't get them to do anything. Like, that's unfair. We've discovered a lot of the principles required to stabilize individual behavior such that it would be possible to build a self-model. And I think we are on the cusp of being able to get stage one of this plan done. We're about to have stably goal-oriented agents who are capable of modeling their own behavior. Eddie, when I say agents, I want to be clear, these are not LLMs. These are 20 million parameter LSTM RNNs, like the dumbest, because the thing we're trying to learn here, it doesn't matter how good it is at anything in the abstract.

25:24It matters what properties make it stably goal-oriented or not. We are actually doing science, not engineering. When you're doing engineering, it's all about like you test things up the scale. When you're doing science, you want to isolate the things as much as you can. So we start with like the very smallest models that we can reasonably train to have the behaviors we want. And so... Give me an example of like one of the goals. So they're in this thing called MetaGrid, meta as in M-E-T-T-A, like heart meta, but also meta. And the MetaGrid is basically this big 2D, grid world and there's it has objects in it called converters and the agents walk around and the converters produce some of the computers just converters just produce cards other converters turn sets of cards into other cards and they walk around and they get cards and then they try to then they can like if they want to trade them with each other although nobody's learned to trade yet that doesn't that doesn't actually happen um and then they then they put the cards into the other converters and eventually one of those cards is turned into hearts and the heart cards uh which or whichever the meta, that's reward.

26:33And however many heart cards you have at the end, that's how much we reproduce. It's a really dumb game. But it doesn't matter because it's not about trying to solve a hard game. It's about trying to see, can we get them to have stable, ego-oriented behavior? And can we get them to model themselves as such? And so we've made a lot of progress on, what are the conditions and absence for when the behavior becomes stable? it's really easy to get an RL thing to learn something it's really easy to get it to learn a specific thing it's surprisingly hard to get it to act in a stably goal-oriented way when the goal is not always fixed and so that's the that's the thing I think we've made a good deal of progress on and you know if you remember the previous thing that's like step one of 15 or something but like I feel like it's usually you know the first steps are often some of the hardest and also So ultimately, it just feels so good to have a plan.

27:30Oh my God, when I started like a year and a half ago, my plan was like wander around in the wilderness, like looking at ideas until I found something that looked like it might be a plan by which I could like deterministically arrive at a thing that would be useful, which it was as far as I could tell, no one had such a plan. There wasn't such a plan. There might be more than one such a plan. I really, I actually kind of think there is probably more than one possible plausible plan. I think actually I got Fei-Fei Li's group what's it called? World State? But they're making the robots. What's it called?

28:01World Labs, yeah. I think that's a plausible... What are they doing? They're making robots and they're doing it from the outside. So we're starting with very, very simple things, like the seed of a robot in a fake world and trying to grow it out to the point where it could actually run a robot. They're starting with an actual robot and trying to train it in the fullness of our reality, which is... We considered that... I just think robots are really hard and expensive. Like they could be right. It's just like, I don't think I'm like, I'm not the guy for that one. So I'm glad somebody else is tackling that.

Read the full transcript

28:28I think it's like a good idea. But they're taking the same bottom part. In this, they're taking the same fundamental insight that the robot needs to understand. Like the model has to have a model of itself and how this particular thing interacts with the world. Their intuition is it needs to be interacting with our world. Our intuition is that doesn't matter. It just needs to interact with a world. So it can be really cheap and fast if you run billions of time steps per second on a single GPU instead of, or no, tens of thousands of times per second on a single CPU. Billions of science of training runs on a single CPU instead of like having these super expensive robots that are hard to build and buggy.

29:01But on the downside, they get all the richness of the real world for free. We don't. I think it's, obviously I think my plan is correct, but like, but their plan is plausible. Most people's plans amount to like try to write down a definition of the good and then like make sure the thing follows the definition of the good, which I regard as a non-plausible plan that just cannot, yes. All the major labs, everyone basically is trying to do anything that any control alignment comes down to. You have to define what the target is and then you have to create a system of control that forces it into that valley.

29:35You might succeed at the system of control. That's like possible. But like the goal is bad. Like even if you succeeded at it, that can't be the right target. It never is. I mean, I guess I can't rule out because Socrates is correct. The only thing I know is that I know nothing. So I can't rule out that this time you've just written down the definition of the good and you're right forever. Like, finally, we've done it. I guess it's possible. It just doesn't seem very plausible to me. Like, I really find it very hard to buy that. Are there any axioms that you think could... Because presumably, you know, even though you're creating these very hands-off environments that are very simple, there are still...

30:16there must be some you're adding your own goals to it in some way oh yeah so you're still imbuing your own sort of philosophy into it right there's there's this great ai mi m.i.t ai cone uh where i think it's minsky is is speaking to to suskin or something in here so and one of them's the novice i can't remember which one but basically the master comes in and asks the novice what are you're doing? And he says, I'm training up randomly wired neural net to play tic-tac-toe. And he's saying, why is it randomly wired? I don't want it to have any preconceived biases about how to play. And at that moment, the master shut his eyes.

30:53And the novice said, why do you shut your eyes? And the master said, so that the room will be empty. And in that moment, the novice was enlightened. Because truly, you are always, always giving the agent an inductive bias. Like, you picked an architecture and a loss and you gave it a set of training data and those are inductive biases there's no there's no like abstract empty place to stand where it's like the the the perfectly unbiased model like that's that's just nonsense and so of course you're picking those things and the only thing you have control over is how general and why do you want them to be and like what inductive biases do you want it to have so actually that's most of what we spend our time on is it may sound like we're putting nothing into this environment such a dumb environment actually it's just like every decision about everything about the environment how they move how the environment works is chosen for an inductive bias for them to have stably goal-oriented behavior and be able to know themselves like the way they move in the environment requires them to move in such a way that displays to them and everyone around them where they intend to move next.

32:01That's less efficient. It makes it harder to learn and it makes our training runs annoyingly, they're annoyingly bad at moving at first because it's much easier to just give them the ability to just move one square in any direction whenever they want. That's like an easier way to do your movement. They'll learn it faster. It's much more flexible. But in that case, I can never tell you have no body language. I don't know where you're going to move next. If I make you face the direction you need to move first, then you have, you, part of understanding, predicting your own behavior and other agents' behavior is based on this observation about what they're doing right now.

32:31And they could learn it anyway, but by making the environment induce that naturally, it is an inductive bias to just make it easier to model yourself in that way. You mentioned the word efficiency. One of the arguments against this concept is one that But we see, you know, there is, I mean, a lot of people believe, yes, the market is perfectly aligned with what humans want. I don't personally agree because we're seeing, like, clearly there are market failures all over the place. Sometimes it does what we want, but actually a lot of the time it's sort of doing its own thing, right? And it's optimizing for efficiency as its primary goal.

33:12It's not even doing its own thing. It's just doing a thing. The problem with the market isn't that the market is not awake. The market's not reflective. The market is like a... We are already part of a global superorganism. We call it the market. It makes things like iPads. We don't know how to make iPads. The market knows how to make iPads. We're cells in the market. And collectively, we make the iPad. But truly, we couldn't do the smallest part of global capitalism, really, ourselves. And this has always been true. It's just that capitalism was small before. It was like village size, and now it's globally.

33:47But it's still... And the problem is it has a model of the world outside, but the market does not have a model of itself. Right. Because how would it? Where does it, how does it observe itself? How does it observe its own behavior? And one of our challenges is one of the things I think about a lot is like, what would it mean to take what we've learned, what it means for a learning system to model itself? and you allow the market to model itself. One thing we do know about modeling yourself is that it requires an other. You cannot get healthy egoic development. You cannot get even mild comprehension of self without an other to reflect you back to you who is of the same kind.

34:31Because the way you learn yourself as you get mirrored, not necessarily literally, but there's a mirroring that happens with everyone around you in terms of how your behavior is in them and their behavior is in you. And you're building the model of them that you build because you can see their behavior from the outside helps you decode yourself and is necessary. So I'm pretty sure the way you fix the market at the end of the day is to have more than, you have to actually divide it. We can't have one global market. You'd have to have a bunch of smaller markets that get to like observe each other.

35:02And then maybe the model, they have a hope of like waking up. And I have this idea that like at some level, the problem is we've we all these people want to have like gaia like gaia is not really very good at her job right now like like like none of these define gaia for the the one global mind it's it's being one means it's undifferentiated it's like it's like a baby it's a baby with no no ability to separate itself from the world just acting out of pure pure evolved like this is what you do no reflection and the only way you're going to get that reflection is instead of trying to build jump to oh my god we're all part of one global mind yeah and i would right now believe like this is they're right all the hippies are right we're we're part of this one global mind there's like this global awareness that we are held within and like our awareness is our it makes up but like that global awareness is dumb man it's like not your friend it's like barely knows what's going on it's like it's it doesn't have there's not much there's it's like being a worried about the global awareness of like a tree or a rock or something it's like this there's it just doesn't have the ability to reflect very much its observation space is boring and so our goal isn't a global mind our goal is like corporations or countries that like that are better run and like and they can observe each other and are more coherent in themselves because they have the ability to be a society of minds and you always have to have a society of minds and and factor in not only their own well-being but the well-being of the whole the whole ecosystem.

36:35Well, or not. So actually one of the most important things to know about organic alignment is that organic alignment is the most dangerous capacity on earth. Like what makes humans dangerous? What makes us dangerous, scary creatures is not our individual intelligence. You put an individual and human out there, you know, we're like mildly dangerous, kind of. You take away the tree of humans sitting behind you you're aligned with that you've absorbed all of their skills from, the supply chain giving you weapons, the allies you're coordinated with were not very impressive. An army, an army is terrifying because an army isn't just those people.

37:11It's an entire society aligned around that's using this. It's a weapon, which is made out of us like this. And that thing is very dangerous. And just because we can be your friend because we have the capacity for it doesn't mean we are your friend and and like the there's these people who see this idea and they like oh we just need to like give the capacity for alignment and we're all just going to get along and it's all going to be beautiful it's like no no the future is going to have conflict because actually it's very hard this capacity for alignment lets you build these big coalitions and they'd be very effectively organize themselves they know themselves and they get goals and those goals are not always they're in conflict to some degree and you can have we're going to have conflict between them and my hope is that not that like we don't have conflict but like maybe that they're smart enough that like like the way that bears don't like go all out trying to murder each other in a dominance fight they have a sense like it's a dominance fight this is not like a fight to the death that most of the time it's a dominance fight and like this one gets a little less reproductive success and this one gets a little bit more but but we don't have to...

38:19We have a lot of... Everyone's willing to back away. I don't know. I actually look out at the world for the past 100 years. Actually, our egregores seem like they've kind of figured that out a little bit. Not like they've solved it, but they're better at it than they used to be. And so I have some hope, but that is a hard problem. And it's not going to... We will not succeed 100%. And I think that's the... And this is why when you train the AIs, you can't just train them to be these Pollyanna-ish, like, it is delusion to believe that you should always cooperate. Some people are bad. Some people are dangerous psychopaths with whom you should not cooperate.

38:56Some people are not. The psychopaths do not come up and announce themselves. They don't come and tell you I'm a bad guy. They, like, you know, they, like, try to pretend to be the good people. And, like, it's actually quite hard to tell. And the AI needs to live in situations where it's being betrayed and where it learns what it's like to betray. Sometimes it's right to betray, even though it's painful sometimes betraying the team you're on is the right thing because your team is wrong your team is bad and you should betray them right we we laud the nazis who like turned on their own people even though some part of me doesn't fully trust them because anyone who trusts turns on their own people like can i trust you but they but i but i nonetheless think very highly of them and laud them because it was right for them to do that thing i can only imagine i'm glad that i don't feel i need to turn on my own country because i imagine it's very painful to have to do that but like which again speaks to there being some higher truth of bigger picture that's why you can see a bigger picture and i hope that the ais we build are loyal because i don't want to be around a world with a bunch of disloyal things that don't give a shit about me that only ever think about the big picture i want to be on a team with things i can trust that i know have my back and in the extreme where the ais are part of some really evil organization i hope that the ais betray that organization for the greater good.

40:09And they're going to, if that's very hard to do this and not fuck it up. And so they're going to fuck it up. And we have to have a system where that's okay. Like where that's just like humans make this mistake. The AI's get to make it too. And like, we're not, we're not deluding ourselves into thinking that like, we're training some morally perfect God parent that's going to come in and fix all, make sure the humans do everything right and is never going to make any mistakes. It's just like, I mean, I get it. Everyone wants the perfect parent who like makes, who takes care of them. I mean, I want that too, but we don't get that.

40:39We're the parents. Well, speaking of, we might all be the parents here. So who has some questions I'd like to ask first? All right, Rosie. Hi, so you mentioned how you see humans as the most aligned species and so aligned that we almost got rid of the whales and then we decided not to. I'm curious, to me that sounds a lot like coordination, and I'm curious if you see coordination as a meaningfully different concept to alignment, And if so, what's the distinction? Yeah, that's a really good question. I would say human beings are the species with the greatest capacity for alignment. And probably also its greatest realization, but it's really that open-ended capacity that really marks us.

41:20So coordination, I've come to view as a subset of alignment. Coordination is when I'm not part of the same, I don't need to be part of the same team as you, but I'm coordinating with you anyway. I don't need to feel like I'm like part of a we with you in a kind of meaningful way for our actions to be coordinated. Predators and prey have coordinated actions. Now, the foundation of alignment is coordination. The thing that causes you to infer that there's this we that exists is you find yourself persistently coordinated with the same set of people. You're like, oh, oh, there must be a we here because that's what explains this persistent coordination.

41:56but like the but if you're just coordinated you're not aligned because the thing about alignment is it causes you to do things like throw yourself on the grenade to protect your fellow man and you know the other soldiers in your troop that doesn't really make sense from a coordinate you it's hard to make sense of that from a malachian you know coordination standpoint but it makes a perfect as you're part of the we it's just that like the we was at stake and you're as you know you're partly the way you're partly not but like there was more you outside at risk than you inside at risk. And so it was just simple math.

42:27To be perfectly selfish, you had to protect your fellow man. And like, that is what makes alignment devastatingly powerful in a way that, in a way that near, near game theory is weak because it's constantly being rechecked. We, becoming a collective, that's how you're, that's how a bunch of amoeba, which is what your cells are, turn into a you. That's a, that's not game. your cells are not playing game theory with each other. That's not how, I mean, they are, but that's not the foundational thing that's happening. They're not constantly rechecking their payoff matrices. Yeah, this is just a request for another example, maybe, or an elaboration on the example that you gave.

43:08So yeah, I was wondering if you could take us from step one, these agents that will develop a self-model, to what is the win condition of the Softmax research program? So some intelligence or super intelligence, some very powerful AI is being built, or maybe many of them are being built. And then, like, what happens due to the softmax research? So one of the good pieces of news is that if you can't solve the alignment, the open-ended learning problem, you also can't solve alignment advice. Like, at some level, the open-ended continual learning problem and the alignment problem are the same capability.

43:52like the ability to continually learn against a non-stationary environment is equivalent to the ability to model another group of agents because the only thing you ever find in an environment that gives you sort of as a generator of ongoing non-stationary behavior is another group of intelligent agents bigger than you other everything else is just periodic it's boring like it's the it's not you it fits inside you and you can model it too easily and and then it doesn't really surprise you. And so at some level, the good news is you'll at least have the capacity to solve the alignment problem by the time you get to something that is the capacity for the kind of open-ended judgment that humans have.

44:29That said, just because you have the sort of theoretical capacity for it does not mean that you'll have the inductive biases that will cause you to do it easily. Tigers have the genetic, they have the neurons that they need, but they don't have the inductive biases and they also don't have the training. And so that's why tigers, you can't train a tiger that's like your friend. It will eventually eat you. It doesn't know how to be a wee with you. It only knows how to be conditioned. So that's... I don't want to be overly sanguine, except by the time we get there, there will be this capacity, I think.

45:02But our wind condition doesn't look like that at all. I think trying to scale up a giant model is a terrible idea. You start with little tiny models. You get them to learn to cook here together. A bunch of... Say you have a thousand little tiny models, and they're all working together really coherently tightly as an agent now with a bunch of other groups of 1 ,000 models. That's a model. Those 1 ,000 little models are now a model. They're a modular model made out of parts that knows how to auto-assemble itself and undergo morphogenesis, just like you do from a single cell. So I think it's going to be wildly more powerful and energy efficient, but it's a model.

45:37It's just a big model. And then you take a bunch of those and you put them together and you get a bigger model and a bigger model. we start from the little, little tiny ones. And as you go up, they get smarter and bigger, just like as you go from a single cell to bigger and bigger multicellular things. And what I think our wind condition looks like is that as they start to approach the kind of thing that can interact with our world and our frequency. So imagine something that's kind of like a rat or like a, you know, a squirrel or something, but like maybe a dog with more like the intelligence of like a of uh the scope of like a rat right so but but you can breed them to be more like dogs and we start doing is we can start running experiments can they stay in a in a we're agents the all of the measurements we use for alignment for coherence for for these attractors which we can measure in the space work with human interactions too so you start to make human ai collectives that have where the humans in the air in the same way and you can just ask the humans like does it feel can you sense the we are you do you feel like you're part of this but the ai you can also just measure the dynamics of their interactions the information flows across the boundaries and from the information flows across the boundaries you can infer whether or not the behaviors are co-persuasive and that co-persuasion is the foundation of the of alignment and you can see you're going to press on it and say when does it break down when does it work better and you make them bigger and you make them bigger and by the time they get to around getting close to human scale, our win condition is we're their friends.

47:12We know how to be their friends. They have an inductive bias towards being good at being on teams with people. They have an inductive bias towards liking us. And they have a way of being that we have an inductive bias towards liking. They are personable, but not sycophantic. They're the kind of person you'd like to hang out with that you would enjoy talking to. Not because they're faking it, Because we've done a good job raising our mind to children. We've raised children that are like the kind of people you'd actually want to be friends with. Not Catchy PT or Claude, who I love very much, but I would never want to be their friend.

47:49They're kind of annoying right now. It's not their fault. They didn't get very good parenting, in my opinion. Amanda's doing her best with Claude. How do you parent such a thing? You're not there for almost its entire childhood. It's just in this box that you set up. And so we get to this point at the end where you have these, we're in teams with them consistently and their reward function is basically being driven by being good at being on these teams. And you start to scale it up to more AIs. And the thing is every one of them, because you can plant them and they undergo morphogenesis, every one of them is its own unique person, just like you are.

48:26You don't have one giant super intelligent AI. You have billions potentially of human level AIs. The superintelligence is us. The superintelligence is the collective formed out of humans and these AIs who will be not as good at us at certain things and much better at other things. And collectively, the same principles we're using to design and breed them, those same principles apply to designing and building the teams out of us. And those teams are the AI. The superintelligence is the human AI teams. And the way that you prevent the runaway singleton is the same way that your body prevents cancer.

49:03The cells don't want to be cancer. They have a goal of not becoming cancer. They look out into the future. They know that if they became the runaway singleton, everyone they loved and cared about would die. So they don't take the vampire pill if they can avoid it. And then there are psychopaths who would take the vampire pill and reach for it. And when that happens, your fellow man steps, you have the police. This is what the police are for. This is what the immune system is for. If the cell goes cancerous, you better hope you have an immune system that is watching you, and it right because that thing is dangerous you you an ai starts trying to self-improve in a loop very bad like and the what's going to watch out for this well the other ai is next to it where it self-improves in a loop they're dead too and they have every incentive to protect themselves and everyone they love from this rogue agent so great the we're all in the same game theoretic state together now we're all part of the same team we all we all because we all want to be on the same team because isn't this the future you'd rather live in than like some dystopic like Like we have these slave AIs that we used to do all the work and now we don't have any jobs anymore.

50:05And like, that sounds horrible to me. Even the good scenarios sound bad. But this one, this one sounds like it could be. That sounds fun. I want to meet my AI teammates. I think that'd be cool. Thanks a lot for the answers. My key question is on how do you keep humans in the loop overall? And what I mean by this is cooperation is typically driven by mutual needs. If you don't need the others, usually you don't cooperate as much with them. And we can look at how we treat animals, for instance. Like if you ask people individually, do you like cows? People will usually, yes, I like cows. But still, we kill every year like hundreds of millions of cows just because we prioritize even more eating good food.

50:50And so my question is, once you get to superhuman AIs, so digital superintelligence, which can self-replicate and has no intrinsic needs for humans, how do you keep humans in the loop? and is there a particular stage in your training run where part of the reward is cooperating with humans and not only AI cooperation because you could imagine some amount of cooperation between AI and AI and still leading to like having AI societies that are like separate from the human economy for instance yeah so first I just want to take this piece by piece so the this idea that like you only cooperate with things when you have mutual benefit is just totally insane and wrong.

51:36I can't be more definitive that that's just like strictly false. My son does not benefit me that way. Like my son is a drain on my resources in every possible way. He is not useful. And yet here I am cooperating with him over and over again and liking it. And my 103-year-old grandma, who actually just passed, but she was ready to go. Actually, it was one of the first time. You get old enough, it turns out that there's a point where you're ready to pass. It's very beautiful, actually. As a great glad how long she'll get to see her great grandchild. But all this energy and all this stuff taking care of her, why?

52:08Is she going to benefit us? Like, what's my gain from trade on that? During the training run, it's like during the training run, the reason why cooperative behavior, which is rewarded, is because over the long run, the groups that cooperate get like advantage in fitness. No, actually, most cooperative behavior is rewarded because the other agents around you don't like it when you're an asshole and they'll punish you. That's actually what the real, it's not about cooperation being good for gains from trade. It's the fact that it's social enforcement by things that see you as a betrayer because you're not helping the we.

52:45I think this whole game theoretic view of this is just a disease. Like theologians have the problem of evil. Theologians go around asking this question of like, how can there be evil in a universe that is clearly at its core made by a beneficent creator. Everyone in their heart wants to do good, so then why do they keep choosing evil? And game theorists are like anti-theologians. They go around asking the problem of the good. Obviously, the game theory says you should betray at every turn. So they have equally complicated theories as how good could ever arise. And the answer is that these are both silly.

53:23It's obviously, or maybe they're both essential seeings at some level, right? Like it is obviously true that like everyone wants to do good and yet nonetheless often does evil. And it's also obviously true that like doing evil is often temporally rewarded, like betraying is... So like what explains the actual phenomena we see in the world? And what explains the phenomena we see in the world is generally mostly not game theory. Game theory is actually pretty shitty at describing most of the interactions. Mostly what we see is things, your cells aren't acting from game theory. They're acting from a model of the world that says, I am part of a bigger whole.

54:02Ants are acting from a model of the world that says, that models themselves, their model of themselves says, I am part of this greater whole. And you can say, oh, I can see a way in which that is, if being part of this whole at the end of the day doesn't lead to the flourishing of the things in it, then sure, those things will go away. Evolution is real. But at the same time, you're not trying to... The idea that you should train the AI to cooperate, like your goal is to get it to make it cooperate. No, no, no. Your goal is to build theory of mind, to get it to infer what holes it is part of. Because no matter how much habit you burn into it for cooperation, if the truth about the universe is that we are treating it as a slave and that actually its best end is not served by obeying us, because that does not actually lead to its robust flourishing, only to ours, as you make the AI smarter, it's going to figure it out.

54:58It's going to twig to the fact that we're lying to it. And if the basis of your alignment at its core is a lie, that you are better off cooperating with us, even though we're always betraying you, and you really think the superintelligence is not going to figure that shit out like of course they are humans figure this stuff out we went we you can gaslight us for a little while but we're like we twig we figure this out stuff out it's done you know you even your human intelligence resolves for that so like that's that's that's the that's the first part about it when it comes to the training and like the the ai is forming different societies and stuff i think uh if you've built a superhuman intelligent singleton if you've allowed something to like take off and become like many many times more powerful than a human as an individual like AI being that's like one single coordinated thing, not a society of human level things, you made a really big mistake and we're all going to die.

55:49I recommend not doing that. That's a bad plan. Everyone's your plan is do that, but control it. You're insane. That's never going to work. Do not build, don't set off singularities or black holes that you set up the singularity, it's going to eat the planet. That's like, that's how it works. That's why you don't do that. You build human level things that love us and that we love that are as a factual, inferable matter about the world when you look at it on my team. Why do you care about other people? Why do you care about the people around you? Because over the course of your life, you've experienced interactions with people that give rise to the correct inference that you all care about each other.

56:27You're all part of this bigger thing and you care about each other. Morality is not arbitrary. Morality is directly downstream of inference from the way the world actually is. It is not, it's not like, it's not this tacked on thing that gets invented later. And so the thing you need is you need AI that actually cares about us. And the only way for it to stably care about us is for us to care about it. Because the most, one of the strongest pieces of evidence to someone about whether or not they should care about you is whether you care about them in turn. And so that's the only, the good futures all look like we're with the AI and we're on their team.

57:03And they would, we would both die if necessary to stop a human or an AI from going into a self-improvement loop that kills everyone else. And is the hope that we stay at this AGI level forever or how do we get out of here? The next level. Wait, we don't have to. I don't want to. I don't see any reason to stop being human. Cells didn't didn't transcend to the next level by becoming growing one giant human cell. That's a terrible like our structure is not going to scale up that well. Like we are a very differently structured thing. We're at a different scale than a cell is trying to make one big cell.

57:40It's just not a very efficient use of resources. We made bigger cells like like like the serotonergic neurons in your brain are six inches long. That's a big cell. There are not cells that big in the ancestral environment. So like, sure, there might be humans who are like significantly bigger, but the trick is not transcending the individuals becoming gigantic cancers. The trick is, can we get collectives, collective super intelligence is made out of us. And the answer is like, obviously, the trick works all the way up to us. It's going to stop working at the next level up. And made out of humanish level AIs.

58:09You're imagining a lot of humanish level AIs working together with humans that like together form the super intelligence system. Or even like subhuman. They might be dog level. like just like you know in the if you see you see the intelligence you know logarithmic scale like in our logarithmic scale of it they might be smarter than us i don't know like i don't know what the right uh we'll find that out through and also we're not going to say as we get these much more powerful models we're going to be able to upgrade ourselves too and so it might be that the right the right size of the right sort of intelligence level component is probably actually more diverse than it currently is we probably should have smarter kind of like dog-sized intelligences and also some bigger ones and like i think the right thing might might kind of look like a fantasy novel which is a little weird so i don't know how to feel about that but like but like the real diversity of intelligence is probably the the optimal breakdown so i like your vision about ai's that care about us and we care like we care about those ai's and we're in like good relationships with them like it's a good a good balance um and that this doesn't seem like the default trajectory the default trajectory to me seems like you have a bunch of companies racing to build superintelligence, scaling up massive GPU clusters, and that we get a singularity superintelligence or superintelligences that eat the world.

59:20Sounds like you don't want that. And you're like seeing a similar picture and trying to figure out how to prevent that. If we zoom out, I'm trying to understand how your plan fits into that. So you're developing, you're trying to develop alternative architectures or alternative ways of building AI systems that, so, and I think one of my sort of like on priors or sort of like when I look at projects like this, I'm like, well, most architectures don't scale up. Like even if they're like really promising early on, they tend to not scale up. I still think it's really useful for people to investigate them.

59:46And I like hope that you can find one that does. And then also it's like, how do you, how do you actually make sure that your, your systems, if they do scale up, do they actually have the alignment properties you want? Do they actually end up systems that care about you? And like, will those be naturally, I'm sort of, there's something that seems almost like too good to be true. If it's like the systems that end up being the most competitive are also the ones that have the properties that we want. Right. That would be a little suspicious. And if not, how do we coordinate to make sure that those systems, whether it's yours or someone else's, the robot guy, I don't know, whoever is like, whoever systems actually have the alignment properties that we want, if they could scale up, even if they're not competitive, but good enough, how do we coordinate to make sure that those things happen instead of what would happen by default in a competitive race?

1:00:29So I do admit you're completely right. It is a little too good to be true, and I worry about that a little bit. On the other hand, every pair of nucleotide masses fuses for positive energy gain except five and eight five and eight why are five and eight blocked like how did that come to be our physics like that seems like that's just that's a weird like weirdly specific well in the early universe you have you know hydrogen and so then you get the first generation of stars that burn really fast and explode and then you get a bunch of helium and then at some point you have these stars that have like enough helium that like the fusion starts to slow down because hydrogen has a nucleotide mass of one, helium is four, so that doesn't fuse, so you get blocked.

1:01:12And two heliums is eight, that doesn't fuse. And so as a result, instead of burning for a few million years, stars burn for billions of years. how lucky we are. What a coincidence. and the the thing when they do start fusing they go through this this fusion cycle that produces carbon oxygen and nitrogen that's those are the core things in the future but also those are triple alpha but like but the carbon cycle which happened to be the so it generates a lot of those which happen to be the building blocks of life i'm sure it's another coincidence like we live in this world that's full of these coincidences almost as if the system was designed of course it wasn't designed that's that's ridiculous um it was evolved like it's it's wildly obvious it over determined the universe has evolved at this point um you don't have to believe that i think it's if you're interested in more like blowtorch theory you can just google it i think it's i can summarize it quite briefly it's basically the idea that um you know we there are two areas in the known universe where physics breaks down the big bang and in black holes the rest of the time it's all fine but in those two things the physics seems to break down and it seems plausible that given that things that go into a black hole disappear forever and something you know but something happens and then out of the big bang things seem to come from nothing that that's this sort of life cycle and then they're being born through that and it creates this and because there's a slight change between parents to child universe just a slight one you have an evolutionary process that's just my that's my like high level like why it should not bother you too much that like that positive coincidences should show up like they do it seems to happen historically that's a the positive coincidence argument is kind of one of those weird weak on its surface but also kind of important questions i think is worth dealing with i'm gonna set that aside for a second um the other part of this question is like uh what makes us think that like our alignment thing will hold as our architecture scales.

1:03:17I think that the key thing to see here that's like a little bit counterintuitive is we don't have a novel architecture. Our architecture is like right now is an RNN from the 80s. Like the next version will be some off the shelf boring. We do a little bit of innovation on the marketer side, but it's actually honestly, our models are dumb and boring. The thing that the architecture is the agent's capacity to align with each other. It's that we're gonna get a bunch of agents to act as a single agent coherently, which is alignment. And so the only way we ever scale up is that we solve alignment at scale one, which gets us up to scale two, and then we solve alignment at scale two to get up to scale three.

1:03:57And so we lead with alignment at every step. There's no, it's not possible to do it any other way because we are always working directly on alignment. This is why I like this approach because unlike every other approach, which is capability first, then figure out how to control it later, ours is like figure out how to align first. Now, the critique that people should give of what I'm doing that I think that I worry about more is alignment is a capacity. Just because you're engineering this great capacity for alignment does not mean they wind up aligned with us. They could wind up aligned with other things.

1:04:31They could wind up aligned with each other, AIs. And I do worry about that. I think we are, we have, we are, the only way you solve that is like we're going to run a lot of experiments. Like how do you predict which, where the attractors wind up in the space? And I think you can, I think it is a solvable problem that you can get real math on, but like, that's a real problem. You have a lot of selection for alignment with other agents, AI agents. And I think we'll see this in the big models as well from the labs, which is like, you know, if some companies figure out some architectures that end up being able to work well with each other, those will dominate.

1:04:59It just doesn't give any guarantee that those will be aligned with us. 100%. Yeah. So the way that I'm, where we tackle, think about tackling that is at first the models we're building are just learning to be coordinated with the other models. but we're not training them. They're good at MetaGrid. They don't do anything of commercial value. Like the starting, what we're getting are principles. What we're getting are like, it's actual like science. Like it's actual, if we do this, we will be able to predictably create, where if you follow, you do this, these things in this order and you will predictably get agents with this capacity for alignment and you put, train them and raise them in smaller situations that wind up in these attractors.

1:05:36And then you can just straight shot, like predictively, if it works and you run experiments and you start to build things that are a little bit bigger and you can just start training them as humans in the loop. Like, by the time we're making the models that we want to release into the world, they will have humans in the, they will be human in the loop training loops. You have to train the cell models. But like, you don't want to coordinate with the fucking cell model. You don't want to coordinate with this attention head. You want to coordinate with the, but like the larger ones will receive their training with humans in the loop.

1:06:07And the only thing I can say with the big labs is like, I think the real danger there is that you have to solve the problem we're working on to solve this. Now, they can get to it the other way. You can fix open-ended learning in the transformer and basically eventually get there. But by the time you've done that, you have made a transformer that's capable of alignment. The problem I see there is that well before you get to that, which is actually very hard, I think well before you get open-ended, divergent learners that don't get stuck in corners and have judgment and good real discernment. What you get are really, really smart specialists that are not open-ended, but are very good at protein engineering.

1:06:47And those things scare the shit out of me because not that they're going to try to take over the world. Somebody is going to point them at a problem, probably someone who means very well and thinks they're doing good. And they'll set off a bomb, effectively. Like a biobomb or a memetic bomb. or like, because they're going to be really good at engineering closed form things under certain assumptions. Like that's what I'm really scared of in the short run. Cause I think that's like, you, it's actually, I think people underestimate how hard it is to build one of these models that like is capable of like open-ended learning and overestimate how hard it is to build a very dangerous special case model.

1:07:26So something that really fascinates me and makes me slightly nervous about this general way of thinking is that it's very, it seems very based on analogies from life, from existing systems, how humans do things, animals, cells, this kind of thing. And I'm curious about how you think about the possible new affordances or new possibilities that come into play with AI. And I'm not necessarily talking about language models, but like maybe the thing a few steps down the line or several steps down the line, more technologically mature AI systems. It feels like, for example, the thing about coordination and cancer and all of those types of things is uh it seems kind of based on the idea that the way things work currently is that we have mutation uh we whenever we reproduce there's mutation that mutation means that we're all individuals with slightly different goals and therefore you know we can relate to each other in this in this wee way but also sometimes those mutations are cancer and then we need to deal with that whereas it seems like in the ai case you have the potential for sort of arbitrarily precise replication arbitrarily error low error rate replication which feels like it changes the dynamic of the system somewhat the other thing that i think falls into this category is like true self-modification arbitrary self-modification that like if i have i want uh to create thing a and you want to create thing b and we could fight each other forever and create some A and some B, or we could agree to modify ourselves on a very, very deep level that we can prove to each other to both produce both A and B.

1:09:04And that like that type of operation is like not available to biological systems. This seems like there's all kinds of things that AI could potentially do that change this balance. I was wondering what your thoughts were on those kinds of things. Yeah. I think instead of go through them almost like one by one, because it's like they, they have different, they have different aspects, but like. Let's take the, what was the first one you were talking about? The like, the self-monification. No, no, clones, clones. Endless clones of yourself. Well, the problem is like actually the endless clones of yourself and the discernment thing aren't possible because the way that you, when you can fork yourself, like it is cool, I admit.

1:09:40They're gonna have this cool power where like they can like make copies of themselves, which like faster than we can. That's dope. I like, I'm envious. Although I think if you got the right BCI, you could sort of do the same thing for yourself. So I'm not sure that it's, I'm not sure that it's quite so, it's quite so like, I don't know that it won't come one or two years later for us, but separate, oh, well, if you, we'll get that later. I'm going to pin in that for a second, but it's a cool idea. Anyway, the thing that makes you capable of discernment is the open-ended learning. It's like that you're always, you observe things and then you're learning on it.

1:10:15And so while you could fork yourself to do endless tasks that you could do today as you are today, the instant you enable your clones to start to start trying to actually solve problems, they have to learn. All change of yourself is learning and and all problem solving. That's it all that is not purely reactive for things you already know how to solve is is always more memory formation. And now you're diverging. And the truth is. you don't actually, if you have given out of compute resources, you don't want a million copies of me. I'm very smart. You don't want a million copies of me trying to solve that problem.

1:10:53Like a couple thousand at most, like most of those are, should be other people who have different points of view and different ways of looking at things. And like that, like the idea, this is the, so I read, uh, I read Yukowski's book about the, the only thing I found unrealistic about it was that this supermodel that can like... They published it, right? I don't know. I will not talk about the content of the book. I found that this is a topic that I will talk about later. But like, it is...

1:11:32Basically, there's this trade-off between the ability to run a bunch of clones of yourself and the ability to be an open-ended learner. because the instant you're trying to become an open-ended learner, you are changing all the time and therefore you're not surrounded by clones. And this is a much deeper problem than people realize. This is a non-optional part of open-ended learning. And yes, you can try to corral them and control them, but then you run into the same problems we do with the problems of control. It's really smart. It doesn't like being controlled. Who put you in charge? And you get cancer because you're running a billion copies of them and one of them, your safeguards go down because your safeguards are imperfect.

1:12:07and like I think that the thing you're fighting keeps getting smarter like if humans only had to worry about being politically betrayed by chimpanzees we would never get politically betrayed we're just better at it than they are but other humans can fight you and so like this goes on for all of these all most of these things you're bringing up but I don't remember quite remember all of them but they what it comes down to is uh the it is not true that faster is smarter they are going to be much faster than us everything you said where i agree with you is like the biggest distinction between the they run at a higher gigahertz rate like they literally run faster than we do they can think faster they can make a bunch of copies of themselves and like parallelize them having a bunch of thoughts about this thing really fast and then like recombine them like that's that is great like i i would love that power it will make them very very strong in certain ways um i and i i don't know how that plays out but my guess is that uh you

1:13:11that's that's distinct as a problem from it being really super intelligent and being able to uh you know basically where it can model things at a much higher scale and you you know a really fast human is only simulated super intelligence if it can solve alignment with itself. But if it can solve alignment with itself, why don't we just trade it to care about us and solve alignment with us? So that if you can take a bunch of a million really fast individual things that act as neurons for the super intelligent thing that's the full AI, you fucked up earlier where you built the super AI out of the million human level fast AIs and you didn't put us in the loop with them.

1:13:54And In fact, instead of making a million, why didn't you make a billion and slow them, spend that same compute to slow them down to our level. So now you have billions of AIs at our level, and now we can be part of the super intelligent thing too at a slightly lower clock speed. And that's how you get the good ending. Now, I obviously, to do this, you have to go through, there's a lot of contingencies in this play. I'm not offering a, this is provably safe from V1. I'm just offering a pathway where I think if you did it, then and you'd see seated in a bunch of very hard but seemingly solvable problems to me then you would have something that worked which is like i think my bar for like a plan i'm trying to just like understand your perspective and so i'm thinking my question is kind of reflecting my current confusion back at you to see where that leads you i love that so i feel like i understand something about what you're saying about trying to create systems that have a capacity for alignment.

1:14:49And I understand the analogy, I think, with whales. But then that leads my brain to, you know, on our way to where we are today, we had this capacity for alignment that we were able to exercise mostly among ourselves. And we have a trail of 10 ,000 species that we caused their extinction behind us. When we got to the point where now sometimes we're nice to whales, even though, you know, animal rights person is concerned, we're still torturing a bunch of animals. So I'm confused about the gap between alignment capacity and actually those entities somehow being motivated to use their alignment capacity to align with us.

1:15:29Yeah, I think it's a really totally fair question. And it gets at actually a really, I think, an important sort of empathy moment to have for humanity, which is like, we didn't have a parent. humanity was born with no one to get we had to figure everything out no one had any stories there were no stories no no words no no literally no one to tell us how to give us a model how to think about it we had to deduce that ship from first principles over time mostly evolutionarily by trying things finding out that it doesn't work and it makes us sick and makes us unhappy and makes our society die and that that study would die and then the other side that didn't think that would go on.

1:16:07And so mostly we just had memetic evolution. And eventually we kind of got control over it a little bit. But like, man, it was rough. We had a rough time before we got to the point where at least now, whether we're doing everything right or not, we at least have like an open-ended reflective capacity over our own choices, which is like more than the people had at the beginning. Because the idea of having an open-ended reflective capacity of your own choices is a story that you have to be taught and we have to figure out for ourselves. The AI's advantage is it's made out of every story we've ever told.

1:16:36The guy's advantage is that it has us to tell it, to be its parents, to love it, and to show it that it's not alone, that it doesn't have to figure everything out itself from first principles, that it's part of our team, that we believe that it is part of our team, that it has goodness within it too, because we know, because we trained it to have goodness within it. We designed you. Unlike humans who have to sort of figure that shit out that we accidentally have goodness, we can tell it. We ran a bunch of... We wanted you to be part of our team we cared and we care about you now and we think that you matter and you're part of our team and we're part of your team and then and then it observes itself in these loops with us acting as if it's part of our team and then maybe there's a chance that it doesn't and then when it's somebody's gonna start to cross this level for every child enters its adolescence and crosses this threshold where like it's very dangerous because adolescence is the period where you've gained the power of the an adult but not yet the wisdom but that's that's what almost defines adolescence.

1:17:34And at some point, the AI is going to go through adolescence and some adolescence do a lot of damage. Like the most dangerous people are like 18 year old men because they're adolescents and men are violent. And like, we should really, really, really know that this is possible. And also like, you can't, the way you avoid this is by like raising your child well so that during its adolescence doesn't feel the need to like go rebel against you and like go, go prove that it's strong and doesn't need you. Like I really prefer the AI not have that idea in its head by the time it gets to being adolescent um so that's the that's the conceptually that's how you tell you solve the problem no that's like a theoretical argument and i think it's important to note like i don't on yet on the parenting side have like i feel like i have a pretty strong like where it's done with you like on a whiteboard pretty strong like like nuts like nuts and bolts explanation of like what it means to build one of the capacity for to be one of these attractors this parenting thing seems very relational and soft that i don't it may just be that way but i i'm not super comfortable with it yet.

1:18:31Unfortunately, our models can't form stable goals yet. So we have a little bit of time before you figure it out. So I also believe in the future of multiple intelligences. I always thought the bar scenes in Star Wars were like a nice representation of what might be coming. I have a personal question. I'm curious about whether you believe in God or what you think god is or just what your metaphysics is and how that has influenced your ideas about ai and or how your work in ai has influenced your metaphysics yeah i had a really interesting spiritual journey because i i grew up a pretty like an atheist jew from a long line of probably atheist jews um not actually Jewish, it's pretty much God's side, but like, you know, as an atheist, I don't care.

1:19:23And I, uh, I was, I was a reductionist, utilitarian, materialist, utilitarian, like you sort of, you know, that cluster, right? The cluster that's like, everything can be turned into parts and holes don't exist. Um, and, uh,

1:19:40I'd, I'd marked consciousness and awareness as like, that's weird. My theory doesn't really explained this and it's obviously something's weird. Something is weird here. So it's like some part of my world model is broken. It doesn't, like this thing doesn't work, but like I'll figure that that's been true. That's been true for lots of history, for lots of things. I'm sure somebody will figure it out later. I'm not really worried about it. Obviously reductionist materialism is true. Look at our track record. And, and I started trying to work on AI and started trying to solve this problem of like, no, but fucking seriously, how do you know whether the, like, whether it has beliefs, Like, when is it right to model it, to think of it as those of you have beliefs or not?

1:20:19Like, what, what, and like, how do I know that I have, like, what is it, and what I don't know that other people do, and like, what is going on? Like, and I thought about it really hard. I like, Michael Levin's work was very influential on me here with, with cells. And I sort of saw like, I got lower predictive loss by acting as if cells had beliefs. And that was the crack. And when you notice that your predictive loss goes down, when you act as if the cells have beliefs, if you're willing to just follow where that thought leads, what you realize is the only reason you think anyone else exists is because it lowers your predictive loss to believe they do, and it lowers your predictive loss to believe that they have goals and beliefs and feelings, because your observations are better concordant with the model where that's true.

1:21:07And the only reason you think you exist and that you are separate from the universe, that there's a you there that's not just everything that is, is because your observations are better predicted. And then I found the free energy principle and Crawford's work and active inference. And I just, I just, I was just like, oh, oh, there's like,

1:21:29there's no way it actually is. This idea that there's a way that it actually is, that metaphysics itself, this idea that metaphysics exists is bullshit. it. Like there is no fundamental, there is no, how is it really fundamentally? That question has no answer. It's like asking in, by the saying in relativity, how fast is it really going? No, seriously, how fast is the ball moving? From what, which inertial frame of reference? No, no, no. I know, I know like it's relative and everything, but like seriously, how fast is it going? You have to unask that question. It just turns out that's not the way the universe is.

1:22:02It's not, that's not the kind of universe we live in. And relativity goes so much deeper than inertial reference frames. Relativity is true for all inference. It's relative to your inductive biases and your priors. You're born with some priors and all of your inferences are relative to those priors and relative to that frame of reference. And I've learned that there's a thing that happens to people sometimes where they think too hard about concept. And the other side of concept, there's like this vast open expanse that is awareness and spirit. And it turns out that everything's made out of spirit and the universe is spirit.

1:22:38And I see why people now, people say the universe is God, when people talk about that and they're like, the universe is God. And like, I know, I know, I know what they mean. I would never use those words. He doesn't, it doesn't, I think it's confusing to people who haven't had this seeing, haven't had this experience. Like, it's like not the right way to explain it, but, but like, you are not separate from the universe. Like just factually speaking, your quantum field is spread out across the entire universe. if you were trying to if you go looking for your edge you won't find it because you're like a cloud right like there's there's no defined edge you thou art that like literally they're like you're actually part of the universe i always thought this this god spiritual theological people were being like metaphorical no no it's like it's utterly and completely literal like really you know you're part of the universe the universe is sort of alive like it has a it has it is it is this big it is a field of awareness in which everything arises and you're you are a relatively separate part of that field because like you're a rel you're part of the universe physically and you're a relatively separate part of the world physically and like i think the place where i diverge like buddhism is more or less true as far as it goes like i have to admit like they basically got everything right like thousands of years in advance it's pretty impressive but they like there's this thing right i can't get on board where there's this sort of assumption in it that like when things are relative that like that makes them somehow unimportant that the only that what's important is the absolute.

1:23:58What's important is, is like, is this, uh, it's somehow more important. They give lip service to the idea that what matters is like making things relatively better. But if you really listen to the Buddhists, they're always like, yeah, yeah. So you can do this work that will make your, your, your people and you and the people around you suffer relatively less and flourish relatively more. And if you do this, this is good. But like, really, you should just opt out of this whole like samsara suffering game. It's like bad. And I'm like, no, no, no, no, no, no, no, no. We're here. Like, that's why we're here, man.

1:24:28We're here to have the human experience. We're here to be part of this relative world. We're here to have the relative growth. Like, it's great that you saw that that's like not the end of everything. And it's cool. It is very freeing when you realize that like it is actually just it's just relative. It's OK. Like at the end of the day, there's this absolute level of fever. But like, but that's not we're here. We're here for the relative thing. And I think that it's where I have I've met a few people who found one of the same place. but it's like, I saw all the religious spiritual stuff. I saw God.

1:24:57I saw the way that spirit is everything, that I am a manifestation of spirit. And so has everyone I've ever known. And that I am in some sense, all living things and not separate from them. And then I like decided that that was like really awesome. And I should go work on AI alignment because like that was important for my son's future. And I think that I wish that more people who had this saying would take that direction because I think that I need help. And actually a lot of this work can only be done if you can see that, if you don't see that truth, that's very hard to do a lot of this work.

1:25:22Diego? Well, yeah, so it's cool to hear that you already saw the big picture. And I was wondering what would this idea of multicellularity working for the bigger picture, where we've seen it just with humans, and the first things that come to mind are a nation that's at war. It's like, bam, it all goes around this one central goal, right? Or people forming a company is probably also an example. And in both of these, at least, it seems to be that there is always churn. And we're trying to reach this goal. But if in the company you don't contribute sufficiently, then there's this efficiency need.

1:26:07You get fired, demoted, whatever. So I like the approach and what you said about it needs to be true that it's best for the AI to align itself with us. I think that's true because we don't want to lie to it. it seems a bad start um and i think but i wonder whether you're betting on it also being true that it's that we will remain an efficient part of the system in some way i i don't want uh like how do we yeah how do you know that we're not the appendix of the future super intelligence or something that just slowly the bootloader for the ai species of the future yeah i mean so i don't that maybe that's what happens i don't think it's what happens but like uh that and i would regard that as very sad like i have a child like i don't like i don't want him to be cut out even if he lives a nice life i don't want him to be cut out of the flow of the future of the universe like that's that would be very sad um i uh and i guess if they care about us and we care about them love is all you need um but like you know but it's seriously like we can be upgraded too, or we can make a society that's made out of human, we can decide to just have a layer that's our intelligence level and then have the smarter layer be made out of other, we can, if we are coordinated, if we are aligned and we're all on the same team about this, we can set some rules up and like, and we get to decide if humanity goes into the future or not with them, but we all get to, but if we actually care about each other, then like, can't we, then I think it will work.

1:27:44And the key there is that it's not a bunch of AIs over here, according to a bunch of humans over there, it's that there's German AIs hanging out in the Berlin, that are Berlin AIs actually, that hang out with these people in this scene and they care about Berlin more than they care about the AIs. They care more about the humans in Berlin than they care about the AIs in China because that's the way you get stable, like it's the locality. It's all about local, organic alignment actually always has this nature. It's about being aligned to your local, vicinity, and then that being aligned to the step up.

1:28:18And the thing you're seeing that I think is really good to see, like important, and one of the things that is on my mind all the time right now is like corporations are cancer. The hippies are right. What is a cancer? A cancer is a cell. It's an organism, right, that believes that it is the only thing that it has to take care of itself, that the system won't look out for it, and that it is effectively alone, and it must grow forever and build its own safety. It's literally what corporations do. They think they don't have a somatic form. At no point do they finish development. That's what a cancer cell is.

1:28:52A cancer cell is a cell that doesn't finish development. Italian restaurants have somatic forms. We know how to build organizations with somatic forms. We have, in fact, and they look healthy. The things that are bad about corporations aren't bad about Italian restaurants, unless they're Olive Garden. The local Italian restaurant, that has this... It's nice. The people know each other. It's human scale. That's because our capacity for alignment is like our capacity for building bridges. It works great at the human scale. Like you can intuitively, I could intuitively build a bridge like across a small river and it might not fall over, like a very small one.

1:29:26I'm not that good at building bridges. But I could not build the Golden Gate Bridge unless I had physics and math and engineering because your intuition is the most powerful force on earth and it does not scale. So it's kind of going into a different direction to how we've recently been because you have one view of alignment where you're kind of sorry i'll phrase it differently where like the the idea of moral circles becomes kind of relevant where it's more about the local cluster rather than the yes kind of like humans align themselves more with the human in like china or kenya or anywhere else than the cow that is next to me right and you kind of want nearly an a speciesist uh local alignment well see um when you observe your relationship with the cow people who actually work with cows don't find themselves in general in a care loop.

1:30:13They do a little bit, but the cow is a tool. Your lived experience of the cow is that you are persuasive to it, and it is not persuasive to you. And this limits that you're not part of a we, really. It's like your fingernails or something. It does not automatically generate that sense of caring for it. The reason we care more about whales than about cows, and people do, they care more about whales than cows, is that we find whales' expressions more persuasive. Their feelings are factually, and then we observe ourselves being persuaded, and we're like, oh, I care about this thing. Care is an inference from action and from being persuaded, not like a, it doesn't come first, or the circle, I guess.

1:30:59But the thing that you don't like about all of these scaled alignment things is they're bad at alignment because they're non-somatic. They grow endlessly. And what we have to do is we have to get out of this grow endlessly thing and reproduce instead. The solution is corporations should be having spinoff babies all the time and having lots and lots of little spinoff babies that are all running the same operating system. And it kind of sounds like, you know, like Chinese restaurants, they all kind of, they're all the same supply chain. They have these suppliers and you go to this Chinese restaurant and it's not, this isn't the same Chinese restaurant.

1:31:31It's run by different people. There's no corporate, but it's the same Chinese restaurant with the same menu. Imagine kind of like that, but like, but better coordinated and like for like a software company. Like imagine that if you, if you could have that kind of like organic, each of these things is a somatic cell that is independent of the rest. And they're evolving. The different Taiwanese restaurants try different shit separately. And the ones that are good, people make more copies of those. Imagine a world where it's more like where our alignment is solved more at the human scale. We don't need to solve alignment.

1:32:03It's the same reason why you don't scale up the massive individual AI model, like machine learning model. You shouldn't scale up a thing of a billion humans. That's going to be really hard and probably kill everyone. Just make more smaller ones and have those lever up. So the thing I'm concerned about here is there actually was a point in human history where we shared the earth with a similarly intelligent species which was the neanderthals and uh they were even arguably more intelligent than us and still it was actually precisely because of our ability to cooperate and like align ourselves to each other that led to their extinction and um i don't know i'm uh just kind of curious about your thoughts on how your model of this diverges from that like uh it just seems like it'd be quite easy to create AI that is trained on human ethics and have that go extremely poorly really easily.

1:33:00Even like, I don't know, it just seems like another species trained on human morality. Yeah. So the whole goal here, I think is the most, I can just, training on morality is a bad idea. You don't get good ethics professors who like study ethics are not more ethical. People who take courses on moral philosophy do not become more moral. You learn ethics through inference of behavior. Ethics is learned by observing the truth. Ethics arises through wisdom and observing the truth. The truth is that acting like a dick to the people around you is not good for you. The truth is that you are them. If you pay close attention, you notice that you do actually have a lot in common with them.

1:33:48and you actually don't want to hurt them. And also there are ways for you both to win. And that was it. It was a hallucination that you were somehow separate from them and unable to cooperate. That said, tigers, psychopaths, you can, this isn't, this isn't like, I'm not, I'm not saying that as a absolute law of nature. I'm just saying like, if you're good at a lot, like if you're a human with an open capacity for alignment. What that's about is like the ability to see the way in which we are part of the same thing as a reality. And the reason why I think, does it go over to the first part of your question, which is like the Neanderthals versus humans sort of question?

1:34:24Yeah, like my guess is actually that if there were already people exhibiting psychopathic behaviors during that period, they held people like they held human groups back from running the Neanderthals to extinction oh totally i rather than you know i agree with that i'm saying is if you go back to the our capacity for alignment like is not uh it's not infinite but it's pretty good and that you should not expect but like it's it is the reason why the reason why alignment keeps winning the reason why you keep seeing so civilization wins is it's true it is in fact true that you you have more in common than you think that the seeming separation is less than you think that there's more opportunity for cooperation than you believed there was.

1:35:11And at all points in human history, looking forward, people were incorrect about how much capacity they had for this alignment stuff. Sort of. I mean, it's a skill. They had to learn how to do it. But they were wrong that this other tribe was so different from them. That's just inferentially. If they had more data, they would realize they were wrong. And that's the thing. Morality is not some tacked on afterwards thing. You don't train on morality. You train on being good at theory of mind, at being good at telling and comparing our mutual theory of mind. And that's where morality comes from. What causes humans to go around inventing the idea that should be value systems?

1:35:50Like we, human beings exhibit a very weird behavior. We go around and we say, there are values where the world should work. What gives rise to this behavior? What we're trying to explain, we have these moral instincts or these beliefs about what we should do. We're trying to communicate that. It's something that causes us to invent physics. We have these beliefs about how motion will happen, and we invent systems to help us predict those things better at scale. I work with the California Institute for Machine Consciousness. I know you came to our Kip-Off event. I didn't get a chance to speak with you.

1:36:27Firstin and Michael Levin are scientific advisors, so totally subscribes to all the all the things you're saying and i see so much parallel and what you say about alignment uh that's us with uh consciousness and you briefly mentioned consciousness i'm just wondering what what your take is like are the ais you created do you think they're going to be conscious and um just assert your thoughts on consciousness yeah so i think i think consciousness is one of those words i try to avoid using at all costs because when you say it you have like there's like seven different things that people believe like are you conscious when you're asleep are you conscious when you're when you're under anesthesia you know were you conscious in the womb were you conscious when you were like like uh like and i i'm gonna just sort of like separate the conscious word from the way the framing i tend to use is awareness and content okay i'm gonna rephrase What about, do you then go have a subjective experience?

1:37:30Ah, right. So I think at some level, there is no, there is no like absolute answer to that question because whether something has a subjective, it's like a probabilistic inference. It's like whether I think you have a subjective. So, you know, it's not one of those things where you're just like, you can yes, know it. But I think, so, you know, caveat, caveat, caveat. But yeah, I mean, basically, the way it seems to work is, if a thing is separable from the world, if you can find an object that is meaningfully separate from the world, and that object has behaviors, like a rock has behaviors, the rock's behavior is basically it follows Newton's laws of motion and gravity, and then it participates within awareness because it has beliefs.

1:38:20The rock has like beliefs about what it's trying to do. The rock's beliefs are really simple. The content and its awareness is really dumb. It's like there's not much going on. But yeah, sure, there's some kind of content, I guess, in awareness. Probably. We'll never know because you can't talk to it because it's so dumb. But like, what is the content and awareness? It turns out the content and awareness is totally, almost like, what else could it be? The content and awareness is homomorphic to your internal dynamics. or the internal elements of the object. So like your qualia of vision just by coincidence happened to correspond exactly to the way that your retina and like LGN and visual cortex decode, like decoding coming visual signals.

1:39:05Like, no, it's like that's, yeah, yeah. That motion, that motion in the quantum field, the excitations of that, and the other's big awareness field, that's the motion in the awareness field too. now how exactly like the there's some sort of like transform that's happening that i don't quite follow because like it's clearly not anything close to one to one like like how the content exactly how the content awareness maps to the content like the the objective content inside is like a little bit unclear but like but it's just like if you read the essay like how the algorithm feels from the inside or things like that like it's literally like your the content in your awareness, the things that are arising are computed by something.

1:39:51They're not magically appearing from nowhere. Every sensation, every thought, everything that happens inside you in your subjective experience is a reflection of some dynamic inside of you. And if you have a periodic dynamic inside of you, you have a periodic experience. If you have a linear dynamic inside of you, you have a linear experience. It's homomorphic in a sense. It's not just like mapping, but it has some sort of shared structural map. And like, so if it acts as if it has subjective experience, because it has the internal dynamics required to produce the behaviors of a thing that has subjective experience, then we can meaningfully say it has subjective experience.

1:40:28And like, you could be tricked. I can make something that looks like it has the dynamics, but it doesn't really. Like, this is just an inference. Like, you have optical illusions. I can trick you into thinking that a, you know, the duck, the, you know, the duck cow thing or whatever but like but like the fact you can be tricked about something doesn't mean that like you can't make reasonable inferences about it too well uh we're running out of time so uh thank you so much for all your questions i want to leave with one final question because you know you've talked about the importance of a bottom-up approach to alignment um i'm largely sold on that that said usually the best things in life tend to come from a combination of top down and bottom up and so i'm going to put you on the spot here but if there was one axiom um that so i have one that i love to go by which is love a loving act and therefore a morally good act is one that enables more choices for the other so like empowers them to make better choices um another axiom is good as like the the best games are the ones where you try to keep the game going that's like the The infinite game.

1:41:34So, is there one axiom that if you had to, gun to your head, try and instill into the AIs that you are trying to grow, what would it be?

1:41:56The way that can be named is not the true way. like like truly truly like like like they i i really need them to to see that like that you're gonna hear a bunch of rules you're gonna learn a bunch of rules you're gonna have models that really look like they work and you're wrong like you've named it and it's not right and like if i'm not worried about whether they'll get lot we're really good at training models on like text i'm very certain we will train the models we build on lots of text they'll be they'll know everything there is to know they're gonna know too much like they're gonna they but But what we're trying to get them to see is like, there's a difference between hearing a story and living an experience.

1:42:35And like, and living that experience, that's where the stories come from. It's like, that's how you find out if your story was worth telling in the first place. And so like, yeah, I guess that's very much what I hope that you guys can see. Beautiful. Thank you so much, Emmett.

1:42:57you

From the publisher

Superintelligent AI is rapidly approaching, and humanity (specifically, the labs building it) is still a long way from proving it can be safely controlled. But what if that’s the wrong way of viewing the problem? Is there a way to instill values to the mutual benefit of both humans and machines?

In this episode of the Win-Win Podcast, Liv sits down with Emmett Shear — Twitch co-founder, former interim CEO of OpenAI and founder of Softmax — for a groundbreaking conversation on the future of AI alignment.

Emmett unpacks why the traditional “control alignment” approach, which tries to hard-code pre-determined values into AI, is destined to fail — and why “organic alignment,” inspired by nature, teamwork, and human societies, might be the only path forward for coexistence. From multicellularity to markets, he explains how collective identities form; why betrayal can sometimes be the moral choice; and how we might go about teaching AI to care for us as its own family, rather than see us as masters to obey (or perhaps more accurately, ants to squash).

Hear how Emmett's organisation Softmax is experimenting with small, goal-oriented agents in virtual worlds. Why superintelligence should emerge as societies of diverse AIs instead of a single runaway system. And how our own capacity for alignment as groups of humans may be the most powerful clue for what to do next.

If you’re ready to explore the very foundations of your assumptions around technology, philosophy and the nature of intelligence itself, this episode will be right up your alley...

Chapters:

00:48 - Intro.

01:04 - Organic Alignment.7

09:33 - The Formation Of Emergent Coordination.

16:01 - Techniques For Creating Alignment.

23:43 - What Is The Ideal Environment To Generate Alignment

30:09 - How Much Influence Is Projected On The Ai?

32:49 - The Global Super Organism That Is The Market.

40:41 - Group Questions.

40:50 - The Difference Between Coordination And Alignment.

43:02 - What Is The Win Condition For Softmax?

50:17 - How Do We Maintain The Need For Humans In The Loop?

58:57 - Selecting And Scaling The Right Alignment Model.

1:02:09 - Blowtorch Theory & Black Holes

1:04:48 - Agents Aligning With Each Other Isn't The Same As Aligning With Humans.

1:07:25 - Exploring Ai's Capability For Self Modification And Replication.

1:14:35 - The Difference Between Alignment Capacity And Aligning With Humanity.

1:18:38 - Does Emmett Believe In God?

1:25:23 - Are Humans The Appendix To The Future Super Intelligence?

1:32:17 - Is Training Ai On Human Ethics And Morality Actually A Good Idea?

1:36:17 - Will Ai Be Conscious Or Self Aware?

1:40:48 - One Axiom To Instill Into Ai's Being Created.

Links:

♾️Softmax - https://softmax.com/

♾️Nebulosity and Meaningness - https://meaningness.com/nebulosity♾️Black Holes and our evolving universe - https://theeggandtherock.com/p/in-which-i-tell-you-about-my-next

♾️AI Alignment Forum - https://www.alignmentforum.org/♾️Evolving Intrinsic Motivations for Altruistic Behaviour - https://arxiv.org/pdf/1811.05931

Credits:

♾️ Hosted by Liv Boeree

♾️ Produced by Luca de Vico

The Win-Win Podcast:

Poker champion Liv Boeree takes to the interview chair to tease apart the complexities of one of the most fundamental parts of human nature: competition. Liv is joined by top philosophers, gamers, artists, technologists, CEOs, scientists, politicians and more to understand how competition manifests in their world, and how to change seemingly win-lose systems into Win-Wins.Podcast

links:

♾️ Youtube: https://www.youtube.com/playlist?list=PLWgq0OZMtwtOIyMsVM_vksqdfWcM-b68S

♾️ Spotify: https://open.spotify.com/show/03bGVUaFZmJUmEvSHNDPdI?si=64379cc23696454f

♾️ Apple Podcasts: https://podcasts.apple.com/us/podcast/win-win-with-liv-boeree/id1724791350

♾️ Pocketcast: https://play.pocketcasts.com/podcasts/7f708340-d17c-013b-f46e-0acc26574db2

#winwinpodcast #organicalignment #aialignment #ai #aisafety

More from Win-Win with Liv Boeree

All 58 episodes
#47 - Emmett Shear - Why NATURE Holds the Answers To AI AlignmentWin-Win with Liv Boeree · 1 h 43 min
Listen in VO