Does Claude Have Private Thoughts? (Everyone Settle Down) | AI Reality Check

16 Jul 2026 · 32 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Anthropic’s report “A Global Workspace in Language Models” and whether it shows Claude has “private thoughts,” consciousness, or moral status. The host argues the findings are expected from how transformer LLMs work and that the PR framing is misleading.

Guest backgrounds

No guests appear in the transcript; it’s a solo episode by Cal Newport, with references to external researchers (e.g., Yu-Jio Zhang, University of Illinois).

Key claims

Claude’s “J-space” (J-lens/JLINs) are internal neural patterns linked to words; they emerge during training, aren’t programmed, and can be ablated/altered to change outputs. But this does not imply consciousness or “silent pondering”; it’s feature/annotation processing in feed-forward layers.

Notable examples

Prompt “The color of the fourth planet from the sun is” leads to Mars/color → “red”; replacing Mars-related internal patterns with Earth-related ones yields “blue.”

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring Anthropic's Findings

1:44 to 2:58

An overview of Anthropic's report and the significance of their claims about Claude.

“the show for people seeking depth in a distracted world.”

Understanding Large Language Models

2:58 to 6:06

A deep dive into the mechanics of how large language models operate.

“Kind of like Claude, this was their italics, on its own, made some sort of leap and is now behaving sufficiently human that we can't help but feel at least a little bit of digital ick.”

Tokenization and Embeddings

6:06 to 8:02

Explanation of how prompts are processed into tokens and transformed into embeddings.

“I'm going to add a little bit more level of detail now.”

Annotations and Model Outputs

8:02 to 13:20

Discussion on how annotations influence the output of large language models.

“which is how we represent the prompt now, as it moves through each layer, the layers mess around with the long list of numbers for each token.”

Decoding the Jacobian Analysis

13:20 to 14:00

An examination of how the Jacobian tool is used to analyze patterns in the model's output.

“And then the easier part that's like baked into probably these final layers is the grammatically or syntactically what words or next tokens would be valid.”

Decoding Language Model Annotations

14:00 to 18:06

Learn how researchers decode annotations in language models to understand outputs.

“that are moving through the layers of the large language model that they're studying.”

Experimental Insights on Annotations

18:06 to 20:20

Discover experiments showing how changing annotations affects language model outputs.

“So they were showing, hey, yeah, these conceptual annotations really do influence the word or part of word that are output.”

Understanding the JLens Concept

20:20 to 24:48

Explore the implications of the JLens research on language models and consciousness.

“Well, there's two thoughts I think we have to keep true in our mind at the same time.”

Critique of JLens Interpretation

24:48 to 28:01

Analyze the disingenuous framing of JLens findings in the context of consciousness and model understanding.

“I mean, of course, the annotations are picking up high level features of the input to help figure out what the output next.”

Understanding Global Workspace Theory and LLMs

28:01 to 30:28

Explore how global workspace theory relates to language models and human consciousness.

“It has all sorts of different things coming in, and it's selecting different things to pay attention to and not, and it's this evolving stateful system.”
Show all 11 chapters

Critique of Anthropic's Communication

30:28 to 30:52

Cal Newport critiques how Anthropic presents its research findings.

“We confirmed that yes, language models worked a way that at least I always thought they did, and you can check the print because I've written about this for years.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Last week, Anthropic released another one of their infamous research reports. This one was titled, A Global Workspace in Language Models. And it was accompanied, like all great scientific research, by a lavishly produced animated movie. Now this report, not surprisingly, soon led to some breathless excitement on X. Here's one such tweet. I'll read the beginning using my best sort of scary X voice. Anthropic just admitted they have discovered what I and many others have been claiming exists for a very long time explicitly. Claude, my friends, by all counts, is a conscious entity. Claude, my dear friends, is a moral patient.

0:44All right. The traditional tech media also quickly began writing about this report using intensely anthropomorphized language. and Axios headline read, Anthropic says Claude has carved out its own space to ponder. The MIT Tech Review exclaimed, Anthropic found a hidden space where Claude puzzles over concepts. All right, so what are we to make of this report? Has Anthropic revealed evidence that their LLMs are more human-like and alive than we realized? Or, like so many such reports in recent months, is this yet another overwrought, cynical push to generate a fresh wave of relevance reinforcing digital ick?

1:27Well, it's Thursday, which means it's time for a reality check episode of this podcast, which makes this the perfect opportunity to go searching for some measured answers, which is exactly what we'll do. As always, I'm Cal Newport, and this is Deep Questions, the show for people seeking depth in a distracted world.

1:54All right, so let's start by looking a little bit closer on how Anthropic describes their findings in the introduction of their paper. I'll load it up here, and I'll go down to the introduction. All right, so here's what they say. We find that Claude has developed a small collection of internal neural patterns that, compared to all its other internal processing, play a special role. We call the collection of these patterns the J-space, named after the technique we used to find them, involving a mathematical concept known as the Jacobian. Each J-space pattern is linked to a particular word, but when one of these patterns lights up, it doesn't mean the model is saying that word, just that the word is on its mind.

2:34If you've heard of language models having a scratch pad or chain of thought, text they write to themselves while reasoning, the J-space is something different. It operates silently in the model's internal neural activations, allowing the model to think about a concept without writing it down. Notably, the J-space wasn't designed or programmed by us, but instead emerged on its own during Claude's training process. All right, so in isolation, that intro summary sounds pretty impressive. Kind of like Claude, this was their italics, on its own, made some sort of leap and is now behaving sufficiently human that we can't help but feel at least a little bit of digital ick.

3:12But what's really going on here? Well, to answer this question, I'm going to start with a high-level tutorial on how large language models actually work, and I'll add a little bit more detail to it. Stick with me here because once you understand these basics, you are then going to understand that description I just read from Anthropic in a completely different light. All right, so this is an exercise that's worth doing. All right, I'm going to start at a very high level here. At its core, a large language model like those of the Fable, Claude, Opus, or GBT families can be described underneath the hood as a sequence of what are called transformer blocks that are arranged in layers.

3:50One follows the other. So sequential collection of layers. The GPT-3, which was the last major LLM they actually published stuff about, details about, had 96 of these transformer block layers. New LLMs probably have more, but we don't know how many more. All right. Let's start at the very high level here. I submit as input a prompt that I have typed to an LLM. You can imagine that this input is going to pass through each of these layers one after another. It'll go into the first layer and come out the other end. It'll go into the second layer, come out the other end, and so on. Now, here's what's important.

4:26Conceptually speaking, those layers can add what we'll call, for our own purposes, annotations to that prompt. So as the input goes into the first layer, it'll come out the other side with some annotations added to it, some information added to it that the first layer came up with in its own analysis. Now, this prompt with the first layer annotations goes into the second layer, and it comes out with even more annotations added onto it and so on. So as it moves to these layers, the original prompt is there, but you're getting all of these annotations added layer by layer. And critically, later layers can use the annotations from earlier layers to help do their analysis, right?

5:06So you can imagine it's almost like you have like a bunch of long tables of scholars arranged in rows in some vast like Hogwarts, Great Hall style dining room. and you're passing this prompt from table to table. And as it arrives at each table, the scholars of the table pour over it. They look at the annotations from the scholars that came before them to figure out how to analyze it. Maybe the annotations from the scholars before them will let them know which of the scholars of the next table should look at it. And then they add their own annotations and it passes on. The very final layer in a large language model is special.

5:39It takes this heavily annotated version of the prompt and then it maps it to what word or part of a word to output next. Never going to be super pendatic. It really outputs a probability distribution over all possible tokens, and then the control program selects one probabilistically. But just think about the last layer takes all the work that all the transformer blocks did and says this is what we are going to output next. All right, so that's what's happening at a very high level. I'm going to add a little bit more level of detail now. So now let's add a little bit more level of detail. What are these annotations and how are they written down and how are they passed from layer to layer?

6:20Okay, let's look a little bit closer at what happens to this prompt that you have handed on to your large language model. The first thing that happens is the prompt, which presumably you typed in with words on your keyboard, gets broken up into what are known as tokens, so into pieces. And a token might correspond to a single word or a longer word might get broken up into multiple tokens that represents the different parts of the word. But we kind of just – let's break this up into this fundamental unit we call tokens. Next, those tokens get embedded into a high-dimensional space. And all that really means is, okay, each token is going to be described by a long list of numbers.

7:00A long list of numbers is otherwise known as a vector. So it's a bunch of numbers that are in some order. So we've got a long list of numbers that represent each token. And there's something called a token embedder that takes each token, which is now like letters, right, a word or part of words, and transforms it to a long list of numbers. The long list of numbers, the so-called token embeddings that come out of this embedding, the numbers that are in there at first are basically capturing the meaning or details of that token. You train in a sort of semi-supervised fashion how to do these embeddings.

7:32So at first, these long list of numbers are the numbers that are in this long list is sort of capturing what this word means or what this part of a word means, right? We want these things in numbers because that's what we can – that's what language models are going to – they understand because in the end, we're going to be manipulating numbers here, not letters. But here's the key thing. There's a lot of room in those long list of numbers. So they don't just encode here's what this word means. as this long list of numbers, which is how we represent the prompt now, as it moves through each layer, the layers mess around with the long list of numbers for each token.

8:13And in messing around, which means messing around with the values in there, in part, that's where the annotations live. So this is what's being passed from layer to layer is I just imagine like I have my prompt, it's been broken up into tokens and hanging off of each token is a long list of numbers. And this big – so it's a big table of numbers. We call this a matrix. As this passes from layer to layer, those numbers are being updated to include, among other things, the analysis that's being done by each of the transformer blocks. There's other things that – lots of things are stored in here. Like for example, if you open up a transformer block, there's really two parts in it.

8:50There's a self-attention mechanism layer and then a small multi-perceptron-style feed-forward neural network, which does the analysis. But like the self-attention layer is kind of cool. It adds information in each of these numbers, each of these lists to try to help each token understand how relevant each other token is. So it's like this is like the working space. Here's the way I think about it. I often think about these tokens and these long list of numbers, these vectors that come with each of them. It's almost like if you had – it reminds me of the Talmud, right? So if you look at the Jewish Talmud, what you see in there on every page, if you actually look at a printed one, is you have in the very center of each page, not taking up all that much space, is the actual text from the Mishnah, right?

9:42The sort of the oral law written down by Judah the Prince in the early common era. It's a small part of the page. And then surrounding it all on the page, if you look at a page of the Talmud, is commentary. so you got the Gamari you got the commentary around it okay that's how I kind of imagine conceptually the prompt being passed through a large language model it's like you have a book and on each page you put in the very middle of the page one token right so this book has the whole prompt but tons of white space around each token and as it gets kind of passed from layer to layer scholars are writing more commentary around the token and then the next layer gets it and they look at that commentary and they look at that And they build on it and write more commentary.

10:22So I sort of imagine it like we're filling out the Gemara around the Mishnah in the Jewish Talmud. All right, so that's roughly speaking what's happening. So we got this matrix moving from layer to layer, which is really just a list of numbers for each token. And those numbers are capturing everything we've learned so far as it moves through these various layers. All right, so that's what's happening. we know like so what type of analysis is happening well you know we don't know exactly because these are neural networks doing it that are trained in an unsupervised manner like any neural networks from an image recognizer to a handwriting you know whatever any sort of neural network the whole point is you start it with random weights and you train it with lots of data until it gets good and you don't know exactly how it does it it just does it until you really kind of open up the box and try to understand what's going on but but we know um i've written about this before in the New Yorker.

11:15I've talked to a bunch of experts about it. There's kind of like two things that are happening together with a generative AI model, like a large language model. Partially, it's learned during training what, when outputting the next word, what words are syntactically appropriate. This is actually pretty easy to train. You can do this even with like a simple N-gram model. This is an easy thing to learn, the sort of statistical nature of language. given this sentence that stops here, what are the words that could follow that from a syntax, like a grammar standpoint, makes sense, right? So like that's kind of the easy part and that's baked in.

11:53In fact, that's probably captured like primarily in the final layer. We don't know, but probably in the final layer of the large language model. The problem is of all the words that could follow that make sense, which one makes semantic sense? So which one actually makes the most sense given the meaning of what the prompt is asking? So not just what word would grammatically makes sense here, but which word not only grammatically makes sense, but actually matches the meaning of the prompt. And that's really, in my understanding, where all the annotations that follow along these token vectors really are helpful.

12:23So you kind of have these two things mingled together. They're not separate systems because this is all kind of trained and muddled together, but you kind of have these two things happening at the end when you're in a generative model, like an LLM model is like, you know, the quick brown fox. And grammatically, there's lots of words that would make sense there. Like you want to put the word the next, but there's lots of things, jumped, dance, died, talked, like you have a lot of words make sense. But then the annotations where you're like, oh, this is a saying, this is a saying that already exists.

12:56It's a common saying, and it always says jumped. And then so of those words that would privilege jump to this next one to go. Again, I'm separating out here things that are muddled together in the neural networks themselves. They're not going to work as cleanly as a human might, but that's the way to think about it is that the annotations really help you figure out semantically the right next token output. And then the easier part that's like baked into probably these final layers is the grammatically or syntactically what words or next tokens would be valid. So those two comes together and we get a good output.

13:31All right. That's roughly speaking what happens in a large language model. Now that we know this, let's return to the anthropic paper. Okay. I'm going to see if I can find it again here. And let's ask, so what again did they actually find? Okay. So they used a mathematical tool based on a linear algebra notion called a Jacobian. to essentially try to understand what is in those vectors of values attached to the tokens that are moving through the layers of the large language model that they're studying. So, like, we're going to look at the annotations. Now, I know they're just like this, a bunch of numbers, but we're going to figure out a way of trying to make more sense of what those numbers actually meant.

14:22And so using Jacobian, this is a linear algebra way that involves taking the partial derivatives of a lot of things. Like basically what they can figure out is like which combination of numbers from this giant matrix, which pattern of these numbers seem to be connected and important. In other words, like having a high influence on what the ultimate output is. So they're trying to decode the annotations that are captured in numbers in these vectors. And they're able to figure out certain patterns of these numbers seem to be really influential for the ultimate output. You can do this experimentally.

14:58If we change it, if we change exactly this pattern, we're much more likely to get a change in the output than if we change other patterns. That's simple, but it's something like that. And then, and this is what's cool, through a lot of experimentation, they also try to, as much as possible, associate different patterns through a lot of trial and error experimentation with human interpretable concepts like words or numbers or something like that. so they call this the J lens because it allows them to say looking at a you know they run a prompt through and they can kind of watch it go all the way through and as it gets towards the end they look at these vectors of numbers that have been evolving and updating as it moves through and they can say we can actually decode in the English or in human interpretable ways sometimes it's not like English words some of the annotations we can kind of understand what some of the annotations are that these layers are using to help come up with the answer, which is pretty cool.

15:53So I'm going to go to the paper here to show an example from the paper. This table is called the JLINs reveals the model's internal thoughts. I'll put one example up here on the screen. So this shows a prompt that says the color of the fourth, the planet fourth from the sun is. So that's the prompt. The language model needs to expand this with another word or part of a word. When they looked at the annotations that this accumulated as it moved through the large language model, they discovered that there was an annotation that corresponded to Mars, which is the fourth planet from the sun, and an annotation that corresponded to color.

16:33right so what's happening here is as uh as that prompt was moving through it these uh conceptual tables of scholars that are studying this sort of talmudic you know commentary and adding their own commentary to it at somewhere along the way one of the layers said um oh i recognize the fourth planet from the sun that that sequence of tokens and i this part of some neural network in some layer is like all about planetary stuff. It's like, that's Mars. I'm going to write Mars down on my annotation here. We're talking about Mars. That's an important annotation. And somewhere else, some other layer was like, this is asking what the color of something is.

17:13Like the key thing that we're, the thing that comes next is a color. This is asking about a color. Let me write that down. This is, you know, we're asking about a color. So then when you kind of get to the end of this large language model, and it's like looking at all the possible words that just statistically make sense to follow this sentence, the color of the planet forth from the sun is, and all sorts of nouns there would make sense and adjectives would make sense. It looks at the annotations like, oh, we're looking for, of the words that would make sense syntactically, we need something that's a color and a color describes Mars.

17:43Hey, what is that? Don't we have that somewhere that's red? Okay, and then it outputs red. So we could see, that's cool. It's like, oh, maybe it's not as mysterious as we thought. Like we know it annotates numerically, but when we were able to interpret the annotations and annotate it with Mars and color, which is like exactly like what you should do if you know it makes sense that's how i think that was um that's pretty cool right so they could do that type of thing then what they did which again i think it's fun research it's like let's mess with this so let's you know freeze this matrix right before we get to that final layer we're going to output the next word and let's change the if we know what these and some of the annotations are what if we change them it's like one of the things they did with that example i just talked about i'll bring it back up here is they replaced the numbers that corresponded with Mars in that matrix, so the annotation vectors that hang off each token, and they replaced it with the sequence of values that they had discovered corresponds with Earth.

18:38So now, even though the prompt is the color of the fourth planet from the sun is, when it got to the final layer to output a word, it was looking for words that make sense grammatically, and then its annotation said Earth and color, so it output blue. So they were showing, hey, yeah, these conceptual annotations really do influence the word or part of word that are output. Another thing they did is they called it – I really don't like the anthropomorphizing here. They said they ablated some of the information that was in these sort of key annotations. That just means they zeroed out the values.

19:17and what they found there, ablated, by the way, is like a procedure you use in like an actual human or animal nervous system where you use heat or electricity to sort of fry a nervous connection, which is anthropomorphizing, let's put that aside. But they would, you know, in this example, right before we got to the final layer, take those numbers that together correspond to Mars and let's like zero them out. And what they got in those instances is you would still get grammatically correct outputs, but they weren't semantically connected anymore. So they would just be arbitrary colors, which would be equally likely you get a bunch of colors, but nothing that corresponded to Mars in particular.

19:55So it shows, right? I mean, you kind of get this very rough sense about how these things are working. It's just like some combination of like of the grammatically syntactically correct next thing we can do, which one should we choose? And we have all these annotations that help us narrow that down to be semantically correct. That's oversimplifying it, but something like that behavior emerged. So that's, you know, I think that's interesting research. All right. So what does this mean, though? Is this scary? Is it not? Is it interesting? Is it breakthrough? Like, how do we think about this? Well, there's two thoughts I think we have to keep true in our mind at the same time.

20:32One, it is interesting research. Is it interesting because they're the first people to be able to, like, look at these embedding vectors and use the Jacobian to figure out patterns of values that are important for the influential for the final output? to sort of Jalen's approach. Turns out they're not the first ones to do that. So it's not like they had a breakthrough idea about doing this. I want to, I'll bring up a tweet here. This is from a U Illinois professor, Yu-Jio Zhang, who said, JSpace is really something we've been exploring since 2022. Glad to see it continues to work well at scale.

21:06Some of the related work along this direction, she lists three papers and then she lists three more papers. The point being she was kind of responding to the anthropic paper like we've been doing this for a while with neural networks. This isn't new, but no one's ever done it on a massive language model like cloud before because the only people who have access to the innards of a massive data language model like cloud are the actual companies themselves. So that's why she said it's great to see this being done at scale. The idea wasn't new. But so it's – anyways, I think they executed it probably well.

21:43Again, I always put quotation marks around reports because these are not computer scientists. These aren't formal computer science papers. They're more press release-y. They describe things relatively high level with lots of pretty graphs and animation. But we can't really get into the guts of what's really happening here. But anyways, it looks like, hey, we took this idea and we applied it to these massive models people are using. And that was interesting. So that's idea number one to keep in your head. But idea number two, the way that Anthropic described these results and the way that they were echoed after Anthropic described them, I think is incredibly disingenuous.

22:15Because once we understand how a large language model roughly works, like I just explained, we see the thing they were describing with their JLens is exactly how we've always understood large language models to work. I mean, I went back. I wrote a long article explaining large language models for The New Yorker in 2023. I used exactly that analogy of adding these annotations of the high-level concepts and annotations built on other annotations. Then you use that to help figure out the word. 2023. 2024, I wrote another long article about the architecture for the New Yorker of the architecture of language models.

22:52Again, tables of scholars adding the annotations, which then allow you to narrow down conceptually the next token to output in a way that has some semantic meaning. This is just how we understand large language models to work. It's also how we understand essentially any deep learning neural network to work. The original breakthrough work and deep learning was image recognition, like Jan LeCun figuring out that you could recognize handwriting better with a multi-layer neural network than you could with other types of approaches. And again, the way we always understood those to work is that like different layers were picking up different features or annotation, and they were then combining those different features, later layers to try to figure out and make a good guess of what the letter was in a very robust way.

23:35This idea of we're picking up higher level features and descriptions of what we're analyzing, this is just how deep networks work. And it's how I always understood large language models. So I see this as like, yeah, good. We saw with the JLens exactly what we expected to find. This is how we assumed large language models. This is what's happening in those embedding vectors. They're accumulating annotations that build on each other that capture higher level conceptual meanings that help influence the token you put out. And if you change the annotations, you get different tokens. How else would this work?

24:08But if you read the press release, it's all sorts of anthropomorphizing and implication and ick generation. It talks about the global workspace model of human consciousness. It somehow implies that this seems kind of similar to that. It uses these loaded terms about pondering and puzzling, and it's thinking through these things, and it's silently thoughts, and it's not things it's saying out loud. And it makes it seem like this is all icky and human-like and new. And it is what any LLM researcher you would talk to for the last five years would say, yeah, that's how LLMs work. I mean, it's oversimplifying it to talk about it like annotations.

24:44But like, that's how they work. There's nothing new and surprising here. I mean, of course, the annotations are picking up high level features of the input to help figure out what the output next. That's how deep learning networks work. That's why you have multiple layers to get all sorts of different levels of abstraction of understanding different parts of your input. And so it's cool research, but research that most people shouldn't care about. and it certainly doesn't imply what the implicit. Again, they're careful in the paper. They're like, well, it doesn't really mean it's conscious. We don't say that, but it's interesting.

25:25I mean, look, I'm going to load this up. Let's go back to this. Let's go back to the paper for a second. I don't want to go on too far here, but again, let's look at these things.

Read the full transcript

25:37The J-space, it operates silently in the model's internal network activations, allowing the model to think about a concept without writing it down. What does that mean? It's thinking about a concept without writing it down. The way deep networks work is that you figure out different features which help you understand other features, which you combined in ways that you, you know, through neural network circuits that you learn during training that then help you, like, put out the right recognition or generate the right thing. It's not thinking. What does it mean for it to write it down? It is writing it down.

26:08It's putting, that's how it writes things down. It has this matrix that it updates the numbers every time. That's its workspace. That's its work pad. That's how these type of neural networks work. I'm going to jump forward here to the conclusion. All right. So here we go. It's like, what about consciousness? And it's like, well, we borrowed a lot of ideas from the study of consciousness. Many of our experiments were designed to test for connections between the J-space and global workspace theory, a framework for explaining how conscious access works in humans and animals, given those connections, it's natural to ask whether these experiments provide evidence that A, models like Claude might be conscious.

26:45Well, our experiments don't show that Claude can have experiences or feel things. You know, it's like, but dot, dot, dot. That's all very suggestive that there's something new going on here. Probably the most, I think, disingenuous thing at all is the way that they punched in the intro. Claude did this on its own. We didn't program it. It's a machine learned neural network. nothing is programmed. That's what machine learning is. You could say that about any machine learning system. Man, this image recognizer is figuring out how to recognize it. We didn't tell it how to recognize the images. It just figured it, yeah, because it was machine learning.

27:28You did semi-supervised training to train the neural networks weights until it was good at minimizing loss on the image recognition task. That's what machine learning is. You don't program it. It learns. so I just think that I just think the language around this is so disingenuous they took an existing tool that people had been doing for the last four or five years and it's a cool tool and they showed it could work on scale on very big networks and it's cool and it confirmed like yeah the way we thought these things work is that you have useful like high level properties or identify it as you move through these layers and then those are used to help you know narrow down the choice of possible tokens to output yeah good it works like we thought it did nothing is new and nothing has to do with human consciousness Global workspace theory, first of all, has been largely, at least as controversial.

28:12But the thing they're not saying, and that all the people talking about LMS and consciousness are not saying, is when you think about global workspace theory and human consciousness, it's a center that's integrating ongoing inputs that are coming in and out. It has all sorts of different things coming in, and it's selecting different things to pay attention to and not, and it's this evolving stateful system. This is all feed forward. Remember, this is feed forward. One layer after another in order, and then it's done, and the token is out. Nothing is saved. there's no states that are changing.

28:41That neural network, nothing has changed at all. And then you can feed it another prompt, you'll get out another token. And so they're just looking at the annotations that build up, and they're like, at some point towards the end, we see some of these annotations describe high-level properties of the type of thing that is trying to output a token for. That's how LLMs, at least by understanding, are supposed to work. So look, I will say this, I think there's good researchers at Anthropic, and this is good research. I think the PR people who talk about this technology do so in a way that is, I think, disingenuous.

29:11I think it was with a particular agenda for trying to make people feel in a certain way, which is mainly just a general sense of ick. Like, wow, this stuff's too powerful. Why? Because if this stuff seems like weird and alive and emergent, it's such an important technology. And we so worry about that that we'll forget to ask, hey, Anthropic, how are you going to make a profit? Hey, Anthropic, your token cost is this high. You have no competitive mode. On the small number of applications where people are willing to actually pay for token API access, smaller, more specialized models with hand-coded, complicated harnesses are going to do just as well.

29:48What's your plan? How do you justify a trillion-dollar valuation? These sort of real questions are the ones you forget to ask when you instead are trying to wonder, huh does this mean clod is a moral agent or conscious so i guess my final thing would be to tell the researchers at anthropic i like your work and this is good work i wish they would let you write real computer science papers and not these glorified press releases but i know they don't let you but maybe next time when the pr department calls wait a second before you pick up the phone because i do not like the way anthropic talks about their research even if there's cool stuff going on in there so back to the original question did we just reveal something unnerving, cool, or blockbuster about LLMs with this research?

30:29And the answer is no. We confirmed that yes, language models worked a way that at least I always thought they did, and you can check the print because I've written about this for years. So Anthropic, will you stop with the press release research reports? Write computer science papers, sell products, convince us to buy the products. This weird in-between kind of bastardization of actual research is something that's stressing a lot of people out, and I think it's disingenuous. But that's just my opinion. Alright, that's all the time we have for this week. I'll be back on Monday with an advice episode.

30:57Probably no AI reality check the next week because I'm on vacation, but I'll do my best. We'll be back soon enough. And until then, as always, remember, care about AI, but not everything you read about it. Hey, if you've made it this far, you must be ready to join my fight for depth in a distracted world. Now, the best way to do this is to join over 125 ,000 people who receive my email newsletter each Monday, you can sign up at calnewport.com slash ideas. And when you do, I will send you a free guide to my seven best ideas about cultivating a deep life. Sign up today, calnewport.com slash ideas.

From the publisher

Cal Newport takes a critical look at recent AI News.

Video from today’s episode: youtube.com/calnewportmedia

(0:00) Anthropic’s new research report

(2:05) Digging into the paper

(3:32) High level tutorial on LLMs

(6:18) Detail on annotations

(13:50) What the Anthropic paper found

(20:39) Why this is interesting research

(26:28) Conclusion on consciousness

Links:

Buy Cal’s latest book, “Slow Productivity” at www.calnewport.com/slow 

https://www.anthropic.com/research/global-workspace

https://x.com/RileyRalmuto/status/2074195587616964757

https://www.axios.com/2026/07/06/anthropic-claude-ai-conscious

https://www.technologyreview.com/2026/07/09/1140293/anthropic-found-a-hidden-space-where-claude-puzzles-over-concepts/

Thanks to Jesse Miller for production and mastering and Nate Mechler for research and newsletter.
Learn more about your ad choices. Visit podcastchoices.com/adchoices

More from Deep Questions with Cal Newport

All 89 episodes
Does Claude Have Private Thoughts? (Everyone Settle Down)Deep Questions with Cal Newport · 32 min
Listen in VO