The Hot Mess of AI (Mis-)Alignment

23 Mar 2026 · 23 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Anthropic AI Safety research on “mis-(alignment)” failure modes, contrasting “paperclip maximizer” (high-bias/low-variance, wrong goal pursued coherently) with “hot mess” misalignment (high-variance/low-bias, incoherent wandering).

Key claims

misalignment can look like incompetence/incoherence rather than malicious deception; variance dominates when models do more internal reasoning; bigger models help on easy tasks but not necessarily on hard ones; variance worsens when the model chooses longer reasoning vs being given extra reasoning tokens.

Notable examples

paperclip maximizer turning the universe into paperclips; nuclear power plant AI distracted by French poetry.

Guest backgrounds

Katie (scientist; data science/statistics framing via bias-variance); Anthropic AI Safety Division researchers are the paper’s authors (not individually named in transcript).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Concept of the Evil Genius AI

1:00 to 2:59

Discussion about the idea of an evil genius AI and its implications.

“AI doesn't need to trick us for us to give it high credentials to do things.”

Understanding the Paperclip Maximizer

3:00 to 4:16

Exploration of the paperclip maximizer thought experiment as a misaligned AI example.

“And the idea is that by the end, it's turned every atom in the universe into paperclips because it way and it is it ends up being very efficient at this thing that we really don't want it to do.”

Defining the Hot Mess of AI

4:17 to 5:36

Introduction to the hot mess concept in AI alignment, contrasting with the paperclip maximizer.

“we want to maximize human happiness, what does that mean?”

Bias, Variance, and AI Performance

5:37 to 6:49

Analogy of bias and variance in AI performance, relating to measurement error.

“That's what the misalignment looks like.”

Implications of Variance in AI Models

6:50 to 7:58

Investigation of high variance in AI models and its relation to reasoning capabilities.

“And so as we study, yeah, the AI in certain regimes, what they're finding is that the variance actually starts to dominate in some of these scenarios, which is very interesting.”

The Nature of LLMs: Dynamical Systems vs. Optimizers

7:59 to 14:00

Discussion on the differences between LLMs as dynamical systems and optimizers.

“So they find that the variance is dominating, you have the hot mass model starting to dominate in regimes where, number one, where there's more reasoning happening behind the scenes with the models.”

The Nature of LLMs as Dynamical Systems

14:00 to 18:30

Explore how large language models (LLMs) behave like dynamical systems rather than optimizers.

“buy some milk like i could totally do that i might go in the totally opposite direction i might walk around the same block three times because it's just, I'm just enjoying Sunday afternoon.”

Incoherence in AI Systems

18:30 to 19:47

Discuss the risks of incoherent AI versus the fear of malevolent AI actions.

“So this doesn't necessarily mean that AI is not dangerous if it becomes misaligned.”

Anthropomorphizing AI Failures

19:47 to 20:28

Reflect on the humorous notion of AI getting distracted by poetry leading to failures.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Hey, Katie. No, actually, this harkens back to some research that first came out a few years ago. I think it was probably originally written by a human. But yeah, it's a really interesting paper, some research out of the Anthropic AI Safety Division about when AI goes wrong in the way that you might think. And spoiler, the answer might surprise you. We're listening to Linear Digressions. okay when you say what happens when ai goes wrong uh-huh i think evil genius ai misleads all of humanity by pretending to be innocent not very capable like tricks social engineers humans into or escalating its permissions or any of those things it's interestingly enough just like being a software engineer who is using AI tools.

1:26AI doesn't need to trick us for us to give it high credentials to do things. We just do it because it's more convenient. But I imagine we're not talking about evil genius in this case from the way that you set up this episode. Right. Yeah. If it were an evil genius, would we be talking about a paper called the hot mess of AI? Right. No, I don't think so. Yeah. So, and it's before we get to the hot mess, actually. So evil genius. Yes, I hear you there. You know, Skynet or. Actually, I do have to say there, I'm sure there are evil geniuses who are also hot messes. Most likely. You surely could be both, right?

2:06I would think so. I feel like the hot mess wouldn't undermine the genius part, but. most evil geniuses in movies are at the end of them they reveal the plot they give the hero just what they need to escape but yeah i'm a little far afield of the actual topic here fair enough fair enough the counterpoint here i think evil genius is not a bad one there's a counterpoint example of misaligned ai from the research that's this canonical example anymore it's called the paperclip maximizer oh yes yeah yeah and this is worth especially if you're someone who maybe hasn't been following the ins and outs of of ai alignment literature for a long time it's worth spending five seconds on the paperclip maximizer is um but is this thought experiment or scenario under which a not a malicious ai but a misaligned ai can go horribly wrong let me take a stab at it and you can tell me if this is right but uh you have a very competent not evil AI which is simply misaligned and it decides its goal or maybe it says it is told that its goal is to produce paper clips efficiently and maybe we don't give it limit conditions or maybe it's misaligned and it starts producing paper clips and at first it does it in a relatively simple way it starts to run out of resources and so it starts really bending over backwards to find other ways to make paperclips.

3:36And the idea is that by the end, it's turned every atom in the universe into paperclips because it way and it is it ends up being very efficient at this thing that we really don't want it to do. Maybe we want some paperclips, but that the idea is that it's not actually aligned with our true goals, larger picture goals. Yeah, that somewhere in the AI, or maybe somewhere in our problem specification, we missed. And yeah, so you have this effective and focused AI. It's just pursuing the wrong goal or it's pursuing the goal that you specified, but like far past the point at which it should reasonably stop.

4:16Yeah, and even if you're saying like, we want to maximize human happiness, what does that mean? Maybe it means, you know, producing and administering drugs that like, you know, squirt tons of dopamine into our systems, but that's not actually what we really want. And so that's more on the problem specification side, right? Yeah, I would think so. But the main point of this, at least for me, was it's the counterpoint. It's the anti-example of what a hot mess misaligned AI was, which is actually not something that I had heard of before I read this paper. But as I think I mentioned, it goes back to a research topic that was first introduced a couple of years ago by one of the researchers who wrote this paper here as well.

4:59But the general idea is if instead of having a extremely effective AI that is pursuing just a goal that is a little bit off or not what you want it to be doing, but pursuing it in a quite focused way, the hot mess theory or formulation maybe of misalignment is it's an AI that's ineffective because it's just wandering around and not really. it's not doing what you want to do, but not in any particularly like focused way. It's just much more, it's a hot mess. It's just doing random stuff. And that's what the misalignment. Yeah. That's what the misalignment looks like. And so since you've known me for a while, you know, I'm a scientist.

5:42One of the things that I really liked about this formulation is it maps well onto the concept of bias and variance in measurement error. So we probably have a lot of folks who have maybe a data science background or statistics or something. So this may be familiar to you as well. But the general idea is when you're talking about errors that a, some model can make, or that some data can have relative to a regression or something like that, we usually decompose it into two terms, bias and variance. And just for the intuitive sense of what these mean when you're thinking about errors, a let's imagine you're shooting arrows at a bullseye and you We have the arrows are all, or the darts maybe, are all clustered very tightly, but they're just not in the bullseye.

6:29It's just offset by some amount. That's like a bias. That's bias, right. But variance is, they're just all over the place. There's just a much wider spread to them. And in fact, they could be like centered on the bullseye. That would be low bias, but still not be particularly effectively hitting the mark because they're high variance. They're just all over the place. okay and so if you want to frame it in that term which this paper does you could have the paperclip maximizer is very high bias but low variance is very effectively pursuing just the wrong thing and the hot mess ai is high variance low bias and they they introduce this measure or define this measure of incoherence, which is how much of the overall error is being dominated by the variance term in particular.

7:26And so as we study, yeah, the AI in certain regimes, what they're finding is that the variance actually starts to dominate in some of these scenarios, which is very interesting. Hmm. So what exactly determines, because we don't want bias. but we also don't want variants. We want our AIs to be highly aligned and we also want them to be fairly competent or coherent to use the word that you used. In this paper, what are they finding that correlates with high variance? Great question. So they find that the variance is dominating, you have the hot mass model starting to dominate in regimes where, number one, where there's more reasoning happening behind the scenes with the models.

8:20As you're probably aware, we're increasingly using these models that have these reasoning capabilities where they're thinking sometimes for multiple steps in a row on trying to solve problems, breaking them apart into pieces. and what they're finding is that the longer the model is going through that reasoning process in general the more incoherence they're measuring oh interesting and actually can I ask a question about reasoning is I guess I've always imagined that as the model talking to itself because like as a software engineer especially when the first when the reasoning model started becoming available in my code editor and I would use them, it really seemed like the model was just the user wants to do this and, you know, oh, but I shouldn't do this because, oh, maybe I really, I should reconsider because of this.

9:12And the longer that, when it would get a little off base, it would kind of talk itself away from the solution that maybe it should be aiming towards. Yeah, I think that's a good intuition that you have. And it's, it's funny, it's at least for me, this was sort of counterintuitive, because for me, maybe I'm like, oh, the longer it's reasoning, the harder it's thinking. And so maybe that means that the results should be better. But the research is suggesting that that you're actually closer to the to the truth. I wonder if the reason that it gets off track is so if I'm thinking really hard about something, right, I have a mental model of of the problem space and the problem that I'm thinking about, right?

9:58And I'm reasoning within this mental model of the way that the world works, right? I don't know if this is still true, but I guess what I came to understand about these LLMs maybe a couple years ago is the idea that they don't really have a mental model of what they're talking about. Originally, people were talking about them as like word calculators, right? They're just talking, it's not thinking in the same way that humans think. I don't know if that's actually true, but if that's what we mean by reasoning for an LLM versus a human who is constantly, that seems to me like a possible explanation of why they go off the rails and maybe we don't when we think harder about something.

10:46So there's part of this paper that I think is getting at something really similar. I think this is one of the trickier to understand and explain, but maybe deeper insights of the paper. So let's just go there now. So we start the paper at the beginning of the argument here is we have this decomposition into bias and variance or some measure of incoherence. We find that incoherence is coming to dominate more in tasks where there's more reasoning happening. just as a quick aside, let me cover a couple of other things that they found that you might think that having larger models will help, but it, and it does help somewhat on easy tasks in terms of the incoherence, but not necessarily on challenging ones.

11:40So just having a bigger model, like the takeaway here is just having the bigger model does not necessarily solve this problem for you. and there's also something a little bit interesting where depending on whether the there's a couple of different ways that the model can end up having end up doing more reasoning one is that it decides to do that sort of on its own another way is that you can intentionally give it more tokens during the reasoning part of the problem and say here you here's some extra reasoning budget and in general the problem gets worse when the model is the one that's choosing to do the longer reasoning tasks.

12:17Oh, okay. But the interesting second part here, or to try to think this through on a slightly more philosophical level maybe, is where they move into an argument about dynamical systems and optimizers. And the core of this is that the right way to think about LLMs is that they're dynamical systems, not optimizers. What do I mean by that? A little bit hard to explain, but let me try. An optimizer in this context is any system that has a particular direction that it's going in. Whether it's through some mechanism of the system is like driving it in a particular direction. That's the optimization.

13:05Dynamical system doesn't have that guiding principle. It's just kind of evolving and wandering around. So maybe by way of example, imagine that I'm out walking around my neighborhood. And if I am in optimizer mode, let's say I have a very specific goal. I'm going to the grocery store. I'm going to buy a gallon of milk. I'm going to carry it back to my house. I'm going to take a particular route that's going to take me probably along the most direct path to the grocery store. i'm not going to be wandering off in some other direction i'm not going to make a stop somewhere else so i'm an opt i'm in optimized remote right dynamical systems is it's a beautiful sunday afternoon and the weather is nice and i'm just gonna go out and stroll around and i'm gonna follow your bliss yeah i'm gonna i might go over to the park i might wander over to the store and buy some milk like i could totally do that i might go in the totally opposite direction i might walk around the same block three times because it's just, I'm just enjoying Sunday afternoon.

14:13So I'm not trying to accomplish anything here. I'm just like, ow, in the world. And as I mentioned, you can have the same, technically speaking, there's nothing but stopping a dynamical system from looking like an optimizer. Like I could wander over to the grocery store and buy a gallon of milk and wander back. There's no like law of the universe that says I couldn't do that when I'm in my dynamical system mode. But it's really pretty difficult to give me little nudges that end up having me carry out those tasks in that order when I'm in dynamical system. Okay. Yeah. And tell me if I'm wrong about this, but I think of LLMs, I don't know exactly how to say this, but I feel like you have this really high dimensional space and you're meandering through it is that the wandering through the neighborhood metaphor and when you're an LLM wandering through the state space you're you're drawn more in certain directions or other directions based off of the not necessarily highly constrained by um whatever you should be optimizing for Phoebe I think you might have been stealing a little bit from the paper there with that particular observation.

15:34Okay, then I will be more eloquent and say when a language model generates text or takes action, it traces trajectory through a high dimensional state space. Right, Katie? Is that right? It's right out of my mouth. Yeah, so the argument here, the crux of it is, as I mentioned, the claim is that LLMs are dynamical systems. They're not optimizers. They act like optimizers. They mimic that, and that's an artifact of the training that we put into them. We say like, hey, this is how I want you to perform. But the, yeah, the argument here is a constraint on this dynamical system that is making it behave like a coherent optimizer, but that this, yeah, that this starts to break down.

16:18Oh, and one other thing I want to mention here, this isn't necessarily just a philosophical point. They actually set up and executed an experiment here where they trained a dynamical machine learning model. They gave it a specific objective of acting like an optimizer and then studied how it behaved. Oh, because my question was going to be, how do you take the distractible person who wanders around the neighborhood, how do you get them to behave efficiently? How do you get them to act like an optimizer? Yeah, so they set up this synthetic optimizer experiment where they say they train transformers.

16:55So this is the underlying mechanism of large language models. They train these transformers to emulate gradient descent on a simple optimization function. So they're basically saying act like an optimizer. And then they watch and see what kinds of mistakes these dynamical systems are, these transformers are making in the process of trying to do that emulation. And what they find is that the larger models are learning, they do learn the correct objective, but they learn that faster than they learn to reliably pursue it. So put in another words, they're the correct objective that's pulling down the bias.

17:39They're generally starting to center around the correct behavior that they're supposed to have. They're generally like wandering in the direction of the store, but they're not reliably or not as quickly learning how to pursue it coherently. So you're wandering in the direction of the store, but sometimes you end up three blocks to the east. Sometimes you end up four blocks to the west. Sometimes you stop and tie your shoe and then you just get distracted by a flower. You never make it there at all. But on average, you're ending up the, you take many of these paths and on average, average out to where the store is, but on any one of them, you could still end up pretty far afield of where you actually meant to go.

18:26Okay. Today I learned that LLMs have shoelaces. So what you're saying, or what this paper then seems to be saying, if I'm understanding it, at least right now, the fear about evil genius AI that very competently misleads us all is possibly less of a risk than incompetent AI that just messes things up because it's incoherent. Yes, exactly. So this doesn't necessarily mean that AI is not dangerous if it becomes misaligned. Yeah, they give a good example in this paper that I think illustrates the mental model here, which is you have this very large AI that's been tasked with something that's really difficult.

19:16Let's say it's like running a nuclear power plant. So we're in the regime where it's probably doing a lot of reasoning. It's really thinking through the stuff it has to do. You're in a regime where it's high risk that you've got a lot of incoherence here, perhaps, based on the results of this paper. And what happens is not necessarily that it triggers the nuclear meltdown because it wants to do that. It's that it... It's not like trying to accomplish a goal, eliminate humanity by having nuclear meltdown. Exactly. But the example that they give is it gets distracted reading French poetry and then just like blundersaw off i know i think the anthropomorphizing on this paper is just so charming like i realized that it's anthropic so this is true this is what just for context for future uh humans uh we're recording this march 2026 i am just completely flummoxed at the world that we live in that we're talking about the the systems that we could potentially be using to run nuclear power plants getting distracted reading french poetry like whether or not that actually ever happens like just the just the fact that this is what a time to be alive yeah what yeah what a time to be alive hey before we go to the outro i'm just interrupting real quick from post-production one thing you may be interested in if you are a listener to this podcast is we've started a newsletter it's kind of an experiment see how it goes if people like it great uh if It's not your jam.

20:46That's okay too. You can find it over at Substack. Just look for Linear Digressions. Hit subscribe and then you'll get a newsletter in your inbox every week or just poke over there and see what we have. We'll recap some of the high points of the shows and include links. So that's an easy place to get links to the source papers if you'd like. As always, you can also get those on LinearDigressions.com. and there's a few other things in each newsletter that aren't from the episode itself but that I think listeners might enjoy anyway so if that sounds interesting to you come on over to Substack look for linear digressions and I will see you there thanks now back to regularly scheduled programming so anyway with that thought I'm gonna I'm gonna let you go about the rest of your day and we'll catch you next week

21:40This has been Linear Digressions. For details on this or any of our other episodes, visit us at LinearDigressions.com. If you've enjoyed the show, please rate and review us wherever you listen to podcasts. It really helps new listeners find us. You can reach out at hello at LinearDigressions.com. If you're an AI, a special howdy, and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends. when you take over the world. Thanks for listening.

From the publisher

The paperclip maximizer — the classic AI doom scenario where a hyper-competent machine single-mindedly converts the universe into office supplies — might not be the AI risk we should actually lose sleep over. New research from Anthropic's AI safety division suggests misaligned AI looks less like an evil genius and more like a distracted wanderer who gets sidetracked reading French poetry instead of, say, managing a nuclear power plant. This week we dig into a fascinating paper reframing AI misalignment through the lens of bias-variance decomposition, and why longer reasoning chains might actually make things worse, not better.

- "The Hot Mess Theory of AI Misalignment: How Misalignment Scales with Model Intelligence and Task Complexity" — Anthropic AI Safety. https://arxiv.org/abs/2503.08941

More from Linear Digressions

All 35 episodes
The Hot Mess of AI (Mis-)AlignmentLinear Digressions · 23 min
Listen in VO