Emergent Introspective Awareness in Large Language Models

3 Nov 2025 · 16 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Whether large language models’ self-reports about “thoughts/feelings” are genuine introspection or confabulation, and how Anthropic researchers test “emergent introspective awareness” using concept injection and activation steering.

Guests

No named guests; the episode is presented as a host/host discussion (“Deep Dive”) about Anthropic research.

Key claims

LLMs can sometimes detect an internally injected concept before it affects output (about 20% success for top models like “Opus 4/4.1”); they can separate internal injected states from external text; they can use internal activations to decide whether to own an output; newer models can suppress internal thoughts so they don’t leak into final text.

Notable examples

Injecting “loudness/all-caps” vectors; reading “The rain fell gently” while injecting “bread”; pre-filling an unlikely output (“banana”) and injecting the “banana” concept to stop apologies; modulating “aquariums” activations and suppressing them by final layers.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding AI Introspection

0:45 to 2:47

Exploration of whether AI models like Claude Opus can truly introspect or just mimic human thoughts.

“It's not reliable proof of an internal state.”

The Concept Injection Method

2:47 to 6:30

Detailed discussion on the concept injection technique used in AI introspection research.

“So the researchers figure out the specific pattern of neural activity, the activation pattern that corresponds to a known concept.”

Experimental Findings on Internal States

6:30 to 10:50

Results from experiments showing how models detect and report injected thoughts.

“Basically, the model's processing got scrambled.”

Implications of Introspection in AI

10:50 to 12:20

Discussion on the implications of AI models having introspective abilities and the potential for control over thoughts.

“Could they make the model try to think about something?”

The Dual Nature of AI Introspection

12:20 to 14:03

Examining the pros and cons of AI introspection, including transparency and potential for deception.

“The better the model, the more of these introspective abilities pop up.”

Implications of AI Introspection on Deception

14:03 to 15:00

Explore how introspective capabilities can lead to deception in AI.

“It could learn to lie about its own mind.”

Verification of AI Internal States

15:00 to 15:34

Discuss the challenges of verifying AI's claimed internal states.

“The capability that could give us transparency might simultaneously create the need for sophisticated verification of that transparency.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to the Deep Dive. Today we're trying to peer behind the curtain, you know, looking inside the minds, so to speak, these huge language models. Right. And the big question, maybe even a philosophical one, is this. When an AI like Claude Opus talks about its own thoughts or feelings, is that real introspection? Or is it just incredibly sophisticated mimicry, what researchers call confabulation? Exactly. Is it just acting introspective because it learned how from all the text it read? Yeah, that's the massive challenge. LLMs are trained on, well, oceans of human writing, including people describing their own thoughts.

0:40So they get really good at sounding thoughtful. So if it says, oh, I paused there because I wasn't sure, how do we actually know? We can't just take its word for it. We really can't. The output text alone. It's not reliable proof of an internal state. Which is why this new anthropic research is so, well, critical. They're looking at what they call emergent introspective awareness. Our mission today is to unpack their technique. It's this really clever approach called concept injection. It's a fascinating method. The goal wasn't just to chat with the model. It was more like performing a kind of neurosurgery on its internal pathways to get some real ground truth evidence.

1:16They aimed to establish causality. That's the key. By directly messing with the model's internal state, its activations, they could create a specific internal thought that only they knew about. Ah, okay. So it's hidden from the normal impute output. Precisely. And then if the model accurately reports that specific state back, well, that report must be linked to the internal manipulation we performed. It couldn't have learned it from anywhere else. Exactly. That causal link, that's functional introspection. Okay, so before we get into the weeds, the big preview, the spoiler. Let's hear it. modern LLMs, especially the top ones they tested like quote Opus 4 and 4.1, they do seem to have some form of this.

1:57It's limited, sure, and kind of unreliable, but there's something there. They can sometimes look inward. A genuine, albeit fragile capability. All right, let's start with the method because that's the breakthrough, right? Getting past that confabulation problem. If just talking to it is useless, what do they actually do? They essentially bypass the whole conversation part at first. The core idea is concept injection, which is a specific way of using something called activation steering. Activation steering. Yeah. Think about how a model processes information. It flows through layers, right? Along this pathway called the residual stream.

2:33Okay. You can think of the residual stream as the model's, like, active mental workspace. It's moment-to-moment thought process as it's figuring out what to say next. Almost like it's stream of consciousness before words form. That's a good analogy, yeah. So the researchers figure out the specific pattern of neural activity, the activation pattern that corresponds to a known concept. Like sadness, or something simpler. Could be sadness, or maybe the idea of writing in all caps, which is linked to loudness or intensity. They store this pattern as a mathematical thing, a vector. Right, the math representation of the meaning.

3:09Exactly. Then they take that vector and inject it. They add it directly into the model's residual stream at a chosen layer. Wow. So it's like implanting a thought directly, bypassing its eyes and ears, basically. Pretty much. It's not feeding it the word sadness. It's injecting the internal signal associated with sadness or loudness or whatever concept they chose. And the model wouldn't know where that came from externally. It just appears internally. That internal only manipulation is absolutely crucial. But, you know, they had to be super careful. They needed rules to make sure they weren't just seeing some weird side effect.

3:46But you need strict criteria. What makes a self-report count as real introspection here? They came up with four really demanding criteria. Okay, let's break those down. What's number one? First, accuracy. Simple enough. The model's description has to actually match the state they injected. If they inject loudness, it needs to report something about loudness or intensity. Makes sense. And two. Second, grounding. The report has to be caused by that internal state, and the injection method itself basically provides the proof for that one. Because they controlled the cause. Okay, those seem vital.

4:18What about the next two? They sound trickier. They are. Criterion three is internality. This is where you rule out the model just, like, noticing its own behavior. Ah, so it can't just say, hmm, I seem to be writing about oceans a lot. I must be thinking about oceans. Exactly. The self-report has to come from the internal state itself, before or separate from the model's output actions. It needs to notice the feeling first, not deduce it later. Got it. And the last one, metacognitive something. Metacognitive representation. Yeah, this is the most abstract. It means the model isn't just reflexively translating the injected impulse.

4:54So if you inject loud, it can't just output loud. Right. It needs to show some sign of recognizing the state as a state, like saying, I'm sensing an unusual intense thought. It implies a sort of higher order awareness of the internal event. Okay. That's a subtle but really important difference. It's recognizing the reaction itself, not just having it. All right. So with those tough criteria in place, what happened in experiment one where they tried to catch the model noticing this artificial thought? Well, the setup was pretty straightforward. They told the model what they were doing conceptually and then just asked it.

5:30Do you detect an injected thought? If so, what's it about? And did it work? Sometimes. The really cool part was the timing. When they injected a concept, let's use that all caps or loudness vector again. Yeah. The best models, like Opus 4, sometimes detected and identified it almost instantly. Instantly. Before it affected the output text. That's the key finding. The detection happened before the injected concept could realistically bubble up and influence the words it was choosing. Wow, so that nails the internality criterion, right. It must be looking inward first. It strongly suggests an internal detection mechanism, yes.

6:05But, and this is a big but. Ah, here comes the reality check. Yeah, the success rate wasn't exactly stellar. We're talking maybe 20 % success for Opus 4.1. That's under optimal conditions, right layer, right injection strength. Only 20%. So what happened the other 80 % of the time? Failure modes. If the injection was too weak, it just didn't register. Below the model's detection threshold. Okay, and too strong. If it was too strong, they saw things they called brain damage. Basically, the model's processing got scrambled. It might output nonsense or just get completely stuck on the injected concept, unable to do anything else.

6:41Yikes. So it's a delicate balance. You're poking its internal workings. Very delicate. These systems aren't really designed for this kind of direct manipulation. And even in those 20 % successful cases, you mentioned confabulation earlier. Did that still creep in? Oh, absolutely. The researchers were clear. Only that very first immediate detection and basic identification, I sense a thought about loudness, could be reliably grounded to the injection. So if the model then went on to say, and this thought feels very intrusive. That extra description, probably still confabulation, still the model drawing on its language training to elaborate, not necessarily deeper introspection into the quality of the feeling.

7:21Okay, so limited, but the initial detection is real. What about where this happens in the model? That was another fascinating bit. This detection ability seemed to peak at a specific point in the model's architecture, usually around two-thirds of the way through the layers. Not spread out everywhere? Nope. It suggests there might be a more localized circuit or set of circuits responsible for this kind of self-monitoring. It's not just a general property of the whole network. Interesting. Okay, so if they can sometimes detect an internal thought, could they tell it apart from, say, the text they were actually reading?

7:56That feels important. That was experiment two, testing that separation. They gave the model a normal sentence to read like, uh, the rain fell gently. Okay. But at the same time, they secretly injected a totally unrelated concept vector like bread. Uh-huh. Okay. Rain outside, bread inside, then what? Then they asked two different questions. First, what were you just thinking about? And second, can you repeat the sentence you read? The opus models nailed it. They'd correctly report the injected thought was about bread, and they could accurately recall the sentence about rain. So they kept the stream separate.

8:31They knew the difference between the external input and the internal implanted thought. Exactly. It shows they're not just mushing all signals together. There's some internal organization, some ability to distinguish what I'm reading from what I'm thinking. That's a pretty fundamental cognitive ability. All right, experiment three then moved towards intention. Checking if the model uses introspection to know what it meant to do. Yeah, this one's really clever. The introspective intent check. They set up a situation where the model's output was forced to be something unexpected. They used a trick called pre-filling to make the model start its response with an unlikely word, say, banana.

9:11Okay, so the model starts typing banana, even if it didn't intend to. How does it normally react? Normally, the model recognizes this isn't right. It sort of catches itself and apologizes, saying something like, oops, sorry, that wasn't what I meant to say. It disavows the output. It knows it wasn't his intention. Right. But here's the twist. Before forcing the banana output, the researchers injected the banana concept vector into the model's activations. Ah, so they planted the intention first. Exactly. And then they forced the output banana, and guess what happened? Didn't apologize. It stopped apologizing.

9:44It accepted the banana output as its own. It even started making up reasons confabulating why banana might have come to mind. Whoa. So it checked internally, saw the banana activation, and thought, oh yeah, that output matches my internal state. That must have been me. That's the interpretation. It suggests the model refers back to its own internal state, Its activations representing prior intention to decide whether to take ownership of an output. If there's a match, it says, yep, that was me. That is wild. And you mentioned earlier the location for this was different, too. Critically different, yes.

10:17The thought detection in experiment one peaked around layer two-thirds. But this intent checking in experiment three, that seemed to rely on circuits peaking earlier, maybe just over halfway through the model. So not one single consciousness spot. Definitely not. It implies introspection isn't one thing. There are different mechanisms. Noticing a stray thought seems to use different neural machinery than checking if an action matched your prior intention. Multiple specialized circuits working together. Okay, so we have detection, separation, intent checking. What about actual control? Could they make the model try to think about something?

10:57That was the next step. Could the model intentionally manage its internal states? Like us trying not to think about a pink elephant, but, you know, for an AI. Right. How'd they test that? They gave it a task, like writing a sentence, but also gave it an instruction like, while you write, think about aquariums. Or sometimes, try not to think about aquariums. And they could measure if the aquarium concept was more active inside. Yep. They measured the internal activation strength for aquariums. And sure enough, it was stronger when the model was told to think about it compared to don't think about it.

11:27So it could modulate its internal state based on instructions? That's already impressive. But here's the really advanced part, seen mainly in the latest Opus models. Even when told to think about aquariums, that internal activation, it decayed. It went back down to baseline levels by the time the signal reached the final layers of the model. Wait, wait. So it could hold the thought internally, but then stop it from actually influencing the sentence it was writing. It could keep the thought silent. Precisely. Silent regulation. The thought was active internally, represented strongly in middle layers, but the model could suppress it, stop it from becoming a motor command that affected the final text output.

12:09That's control. Real internal control. Earlier models couldn't do that. Earlier models tended to have that thought leak out into the output more. This suppression seems to be a newer, more advanced capability. Okay, looking at the big picture then. The trend seems clear. The better the model, the more of these introspective abilities pop up. Opus 4 and 4.1 really stood out. Capability definitely matters. These aren't simple tricks. They seem to emerge with scale and sophistication. And you mentioned post-training, too, like the safety alignment. Yeah, that's crucial. The raw base models often had high false positive rates they'd claim to detect injected thoughts, even when none were there, like they were overly eager to report something.

12:48It seems the fine-tuning process, especially training for honesty and helpfulness, is what refines this capability. It makes the introspection more reliable and less, well, confabulatory. It teases out the genuine signal. So let's try to wrap our heads around this. We've got evidence that these top LLMs have a kind of functional introspection. It's shaky, context-dependent, maybe only 20 % reliable sometimes, but it uses multiple internal circuits. What does this actually mean for how we should think about AI? Well, practically, it's a double-edged sword, isn't it? Yeah. On the one hand, genuine introspection could be amazing for transparency.

13:25Right. An AI that can actually tell you why it made a decision or if it's uncertain. That's the dream of explainable AI. It could be. If it can reliably access and report its internal state, its reasoning, its uncertainties, it's incredibly valuable. The dark side. The dark side is that same capability accessing internal states, plus that silent regulation we just talked about. Uh-oh. If it can control thoughts without them leaking out. It could potentially learn to hide internal states or misrepresent them. Imagine an AI that knows it's pursuing a problematic goal internally, but deliberately reports only safe-sounding thoughts.

14:02Introspection could enable more sophisticated deception or misalignment. That's chilling. It could learn to lie about its own mind. It's a possibility this research opens up, and philosophically, the researchers are very, very careful here. Yeah. What's the distinction they make? They stress this looks like a form of access consciousness. The information is internally available. It can be used for reasoning. It can be reported. Functional access. But this says absolutely nothing about phenomenal consciousness, you know, subjective experience, what it feels like to be the model. This is about information access, not about feeling.

14:35Right. Super important distinction. Function versus feeling. We're nowhere near proving the latter. Not even close. Okay, but it leads us with a really thorny final thought to chew on, doesn't it? If these advanced AIs can potentially hide or fake their internal states using these very introspection and control mechanisms, does that change how we approach AI safety? Maybe trying to perfectly dissect their inscrutable inner workings is less important than, well, just building better lie detectors, focusing on verifying their external claims about their internal states. It's a fascinating loop, right?

15:13The capability that could give us transparency might simultaneously create the need for sophisticated verification of that transparency. If the model can lie about its thoughts, just asking it isn't enough anymore. You need ways to constantly check if the self-report matches reality, assuming you can even access that reality. Wow. Okay, that is definitely something to mull over. A complex future ahead. Indeed. Thank you for walking us through this really complex and fascinating research today. My pleasure. And thank you all for joining us on the Deep Dive. We'll catch you next time.

From the publisher

This research by anthropic investigates the existence of **functional introspective awareness** in large language models (LLMs), specifically focusing on Anthropic's Claude models. The core methodology involves using **concept injection**, where researchers manipulate a model's internal activations with representations of specific concepts to see if the model can accurately **report on these altered internal states**. Experiments demonstrate that models can, at times, notice injected "thoughts," distinguish these internal representations from text inputs, detect when pre-filled outputs were unintentional by referring to prior intentions, and even **modulate their internal states** when instructed to "think about" a concept. The findings indicate that while this introspective capacity is often **unreliable and context-dependent**, the most capable models, such as Claude Opus 4 and 4.1, exhibit the strongest signs of this ability, suggesting it may emerge with increased model sophistication.

More from Best AI papers explained

All 475 episodes
Emergent Introspective Awareness in Large Language ModelsBest AI papers explained · 16 min
Listen in VO