In short
The episode argues that today’s LLM “intelligence” is mainly next-token prediction plus assimilation into fixed or prompted schemata, not true accommodation (restructuring schemas when they fail). It frames this as an “assimilation-accommodation gap” and links it to crystallized vs fluid intelligence.
Guest backgrounds
No specific guests are named in the transcript (only speakers referenced as “TJ” and “our sources”).
Key claims
LLMs prioritize linguistic fluency over logical rigor/factual accuracy; chain-of-thought prompting mostly scaffolds behavior rather than genuine reasoning; LLMs lack an on-the-fly mechanism to detect failure and update their reasoning templates (only slow retraining).
Notable examples
Early math failures with number changes; chain-of-thought improvements on GSM8K; brittleness and error propagation in CoT; major failure on ARC (humans ~75–85% vs GPT-4/GPT-4o <20%); proposed fixes include intrinsic “latent reasoning” decoding, neurosymbolic hybrids, Piagetian/world-model benchmarks like Coggle-LM, and human-AI schema co-creation.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOExploring LLM Intelligence
0:45 to 2:00
Discussion of how LLMs function and what defines their intelligence beyond simple predictions.
“So let's dig into the core mechanism first.”
Next Token Prediction Explained
2:00 to 4:05
An in-depth explanation of Next Token Prediction and its implications for language models.
“The power really comes from the immense scalability.”
The Limitations of NTP
4:05 to 6:50
Analysis of the limitations of NTP, including issues with logical rigor and factual accuracy.
“How does it help us understand learning?”
Cognitive Frameworks: Piaget's Schemas
6:50 to 8:40
Introduction to Piaget's theory of schemas and its relevance to understanding learning.
“That feeling of disequilibrium creates the motivation, the internal push to restore balance.”
Assimilation vs. Accommodation
8:40 to 11:20
Discussion of the processes of assimilation and accommodation in cognitive development and their implications for learning.
“But what's really happening here, is the AI genuinely reasoning in the way we mean it, or is it something else?”
Chain of Thought and LLMs
11:20 to 14:00
Exploration of the Chain of Thought prompting and its impact on AI reasoning capabilities.
“it's not a fundamental solution to achieve true, flexible machine reasoning.”
Understanding the Accommodation Gap in LLMs
14:01 to 16:45
Learn about the distinct failure of LLMs in fluid intelligence tasks.
“It's like solving a puzzle you've never seen before or figuring out a new route when your usual road is blocked.”
Current Research Directions for Improvement
16:45 to 18:03
Explore cutting-edge approaches researchers are using to enhance LLMs.
“It shows they're great at applying known patterns but struggle deeply with generating genuinely new ones.”
Innovative Learning Models Inspired by Human Development
18:03 to 20:39
Discover models that mimic child development to improve AI learning.
“The second major area is engineering accommodation through neurosymbolic AI.”
The Challenge of AI Accommodation and Future Implications
20:39 to 23:08
Examine the challenges of creating AI that can adapt like humans.
“But a very practical avenue people are exploring within this is what's called human AI co-creation of schemas.”
Transcript
Automatic transcript. May contain errors.0:00When we talk about whether large language models, those incredible AIs that generate text, are truly intelligent. the conversation often gets, well, pretty polarized. It really does. Is it just about predicting the next word? Or is there maybe something deeper happening under the hood? That's a great question. What's really fascinating here is that the answer isn't a simple yes or no. Not at all. Our deep dive today looks into a really powerful hypothesis that the intelligence we're seeing in modern LLMs isn't just about next token prediction. Critically, it's about how that prediction combines with external structures, what we're going to call schemata.
0:40Right, schemata. And our sources for this deep dive, they come from a pretty critical investigation into what's being called the token prediction plus chain of thought model of AI intelligence. That's the one. So our mission today is to unpack this idea, maybe reveal a significant gap in what current AI can do, and leave you with a much more refined understanding of where AI truly stands on the path to, well, genuine reasoning. So let's dig into the core mechanism first. These large language models, things like GPT at their foundation, they're built on something called Next Token Prediction, or NTP.
1:16What exactly does that mean? How does it work? It's magic. Okay, so at its simplest, NTP is basically about predicting the next most probable word or maybe just a piece of text in a sequence over and over again. Like autocomplete on steroids. Kind of, yeah. Think of it like a superpowered autocomplete. The model's trained on just a massive amount of text data, right? Yeah. And it learns all these statistical relationships between words. Yep. Then when you give it a prompt, it generates one token, could be a word, could be part of a word one at a time, and it's constantly feeding its own output back into itself to decide what comes next.
1:49Ah, so that's the autoregressive part. Exactly. That's why it's called an autoregressive process. Sounds deceptively simple, but the results are, well, they're clearly not simple. So what makes this process so incredibly powerful, and what are its limits? The power really comes from the immense scalability. You train these things on internet scale data, and they just soak up these incredibly intricate statistical correlations, grammar, semantic relationships, just vast amounts of factual knowledge embedded in language. So they become really fluent. Incredibly fluent, yeah. And this is the key limitation.
2:27This training fundamentally prioritizes that linguistic fluency, that probabilistic coherence over things like logical rigor or even factual accuracy. Right. I remember hearing about early models. They'd fail on simple math problems if you just changed the number. Exactly. That's a perfect example. They were recognizing the patterns of how numbers usually appear together in word problems they'd seen, not actually understanding the underlying arithmetic. It was statistical matching, not abstract calculation. And what's also fascinating is this whole black box problem. We don't fully get their internal workings.
2:58But some recent research has revealed something called the law of equilearning. Equilearning. Yeah. Yeah. So what that means for you is that the reasoning, such as it is, isn't really localized to specific parts of the AI's brain, you know, specific layers. It seems to be more of a homogenous, iterative process of refinement spread across the entire model. Which is very different from human brains. Fundamentally different, yeah. We have specialized areas, and this is why, despite their power, these models are prone to hallucinations, just making stuff up and reflecting biases from their training data.
3:34Okay. These weaknesses really highlight why NTP, as powerful as it is, seems to require some kind of external structuring, some scaffolding for anything truly complex. Okay, so they're brilliant predictors, amazing and fluent text. But to really get a handle on what LLMs are doing, and maybe more importantly, what they aren't doing, we need a benchmark. And for that, you suggested we turn to one of the giants of cognitive science, Jean Piaget, and his theory of cognitive development. So what was his core concept, the schema? How does it help us understand learning? Right, TJ. His schema is, well, essentially a mental blueprint, a framework.
4:14Think of it like a pre-existing idea or pattern in your mind that helps you organize and interpret new information. Like a mental shortcut. Sort of, yeah. Yeah, it could be something really basic like an infant's innate sucking reflex or much more complex mental models like, you know, what a dog is or all the steps involved in going to a restaurant. Okay. And Piaget really emphasized that children aren't just passive sponges soaking up info. They're active builders of their own understanding, constantly constructing and refining these schemas. So if I encounter something new, I just try to fit it into what I already know, the idea behind assimilation.
4:49Exactly. That's assimilation. You take new information, a new experience, and you try to fit it into an existing mental framework, an existing schema. Like the classic example. A kid knows doggy, sees a sheet for the first time, and says doggy. Perfect example. Or the one I often use. Imagine a young child who has a cow schema, maybe from picture books. The first time they see a horse in a field, they might point and exclaim, cow. Right. They're assimilating the new stimulus that the horse into their existing cow schema, trying to make sense of the new by applying the old. But just fitting things into old boxes isn't enough for real growth, is it?
5:27Yeah. What happens when that existing framework, that schema, just doesn't quite fit the new reality? That's where the really crucial, profound process of accommodation comes in. This is the act of altering those existing frameworks, or maybe even creating completely new ones, when what you already know just isn't adequate anymore. So changing the box. Well, making new ones. So continuing our example, when that child is gently corrected, no, darling, that's a horse, their cow schema is suddenly insufficient. It doesn't work. They have to accommodate. They modify their cow schema, maybe making it more specific, large farm animal with horns.
6:04And importantly, they create a new and distinct horse schema. This act of restructuring, this changing of mental frameworks is the absolute engine of cognitive development. It's what pushes us to evolve our understanding. Okay, and here's where it gets really interesting for me. P.S. Schatz said this whole back and forth, this constant refining through assimilation and accommodation, is driven by a fundamental desire for equilibrium, cognitive balance. That's right. So when new information contradicts our schemas, like when the child learns the horse isn't a cow, we feel disequilibrium, that mentally uncomfortable feeling, that, wait a minute, moment that forces us to reevaluate.
6:45Precisely. That feeling is the trigger. So that discomfort is actually what drives us to learn. Exactly. That feeling of disequilibrium creates the motivation, the internal push to restore balance. And you restore balance through the hard work of accommodation, which leads to a new, more sophisticated understanding, a new equilibrium. This cycle from equilibrium to disequilibrium, then accommodation and back to a more sophisticated equilibrium, this is the continuous engine of intellectual development in humans. And it sets a really high bar for any system claiming to truly learn using schemata.
7:17It demands not just using internal models, but having a genuine mechanism to detect when they fail and fundamentally restructure them. Okay, so that's how humans build understanding and learn continuously. How does this map onto LLMs? LLMs, do they have anything like these schemas? And what happens when their existing patterns don't fit? Right. Our source materials point to a really pivotal moment in AI development here. Chain of thought or CO-T prompting? Yes, CO-TT. This was a real breakthrough described in a 2022 paper by Jason Wei, Zuezi Wang, and Denny Zhu. And the idea was actually simple, but really profound.
7:53What was it? Instead of just asking an LLM a question and expecting a final answer right away, you'd include intermediate reasoning steps in the prompt itself. Ah, so you show it how to think it through. Exactly. The structure became question, then the reasoning process laid out, and then the final answer. So, for example, when solving a math word problem, the prompt would literally spell out the step-by-step natural language logic and the calculations needed. And the effect was? Striking, as the paper put it. Just dramatic performance improvements, especially on benchmarks like grade school math problems, the GSM8K dataset.
8:26set, it was a huge leap. So CO-T essentially acts as this powerful scaffold, right? It guides the LLMs token by token generation along a path that looks logically coherent. It seems like it decomposes a complex problem into a sequence of simpler steps that the LLM is better at handling. But what's really happening here, is the AI genuinely reasoning in the way we mean it, or is it something else? This is where the distinction becomes absolutely crucial. The The consensus is that Cotie is fundamentally a form of behavioral cloning or maybe pattern mimicry. Mimicry. Yes. The model learns to generate text that looks like a reasoning process because that's the format you showed it in the prompt or maybe similar formats it saw in its massive training data.
9:07Okay, so like the student memorizing steps. Exactly like that analogy. Think of a student who has memorized all the procedural steps for solving a specific type of algebra problem. They haven't necessarily grasped the underlying mathematical theorems, the Y. Right. They can apply the procedure perfectly to new problems that fit that exact template. But if you change the problem's structure even slightly, if it deviates from that memorized pattern, they fall apart. They completely fall apart. They lack the deeper understanding to adapt. Now, our sources also mention Denny Zhu's deeper research.
9:42He started to question if Cothee was just a clever prompting trick or if it maybe tapped into something more fundamental within the models themselves. Yes. His later work was really interesting. It showed that if you explored alternative ways the model could generate text, maybe looking at less probable token sequences, you could often find paths that look like chain of thought reasoning already present in the model's potential outputs, even without specific Cothee prompting. Hmm. So the potential for structured reasoning might be embedded somehow. It suggests that yes, that the potential is there, latent within the model's learned representations, and Cothee prompting is one way to surface it or guide it.
10:22And Cothee has evolved since then, right? Oh, definitely. We've seen advancements like zero-shot Cothee, where just adding a simple phrase like let's think step by step can induce a reasoning chain, which is quite amazing. Yeah. And there are efficiency improvements too, like skeleton of thought. However, there are still critical limitations. Like what? Well, first, Cote's effectiveness is highly dependent on model scale. It often only really emerges as a capability in these massive models, like 100 billion parameters or more. Okay, so smaller models don't benefit as much. Not nearly as much, or sometimes not at all.
10:57It's also described as brittle, meaning it's highly sensitive to the specific examples you provide in the prompt. Change the example slightly, and performance can drop off a cliff. Fragile. Very. And its sequential nature makes it really vulnerable to error propagation if it makes a mistake early in that chain of thought. The whole thing goes wrong. The rest of the answer is likely to be completely wrong. These limitations really underscore that while Kati is a powerful scaffold, a useful technique, it's not a fundamental solution to achieve true, flexible machine reasoning. Okay. So when you put all this together, how LLMs operate with this next token prediction, how they use chain of thought as a scaffold.
11:37And then you compare it to how humans truly learn through assimilation and, crucially, accommodation. That's where this really significant difference, this critical gap, becomes clear, doesn't it? Absolutely. That's the core revelation here. The deep dive really validates the insight that current LLM intelligence is about prediction plus structure. That's more accurate than just prediction alone. But if we use Piaget's more precise terms, that structure, like Cati, primarily enables powerful assimilation. It helps the LLM fit a new problem into a provided or previously learned reasoning template.
12:12Okay, so it's good at using the templates it has. Very good. But the critical flaw, the fundamental missing piece, is that current LLMs lack any real mechanism for accommodation. They can't change their template if it's wrong. Yeah. Or make a new one. Exactly. They can't effectively recognize when their current schema or template is inadequate for a new situation, nor can they dynamically restructure their own reasoning approach when faced with genuine novelty, something outside their training patterns. Which is the complete opposite of human learning. It's in stark contrast, yes. As we discussed, human cognitive development is driven by that friction, by accommodation.
12:50For an LLM today, the only real analog to accommodation is taking the entire model offline and doing incredibly slow, expensive retraining with new data. There's no on-the-fly adaptation of the core reasoning strategy. And this gap, this assimilation-commodation gap, it maps pretty perfectly onto a well-established framework in cognitive psychology, doesn't it? The difference between crystallized and fluid intelligence. It really does. It's a very useful lens. So you have crystallized intelligence, often called GC. This is your ability to use the knowledge, facts, and skills you've acquired through past learning and experience.
13:24Stuff you know. Stuff you know, exactly. And LLMs trained on these vast internet data sets possess an immense, truly unprecedented level of crystallized intelligence. Their next token prediction capability is essentially how they access and utilize this massive storehouse. Okay. And the other type? That's fluid intelligence, or GF. This is your capacity to reason, to solve novel problems, identify complex patterns, think logically, and adapt your thinking in situations where your prior knowledge isn't directly applicable. Figuring things out from scratch. Pretty much. It's like solving a puzzle you've never seen before or figuring out a new route when your usual road is blocked.
14:06That's fluid intelligence in action. And that capacity for novel problem solving and adaptation, that's the very essence of pyogention accommodation. And the research shows LLMs struggle here. Consistently, yes. Research consistently indicates that while LLMs excel on tasks drawing heavily on crystallized intelligence, retrieving facts, summarizing known information, generating texts and familiar styles, they exhibit profound deficiencies when it comes to tasks demanding fluid intelligence. And is there a specific test, like where the rubber meets the road on this, clear evidence of this accommodation gap?
14:39Oh, absolutely. The clearest, most compelling empirical proof comes from their consistent and frankly profound failure on a specific benchmark test. The abstraction and reasoning corpus usually just called ARC. ARC. Okay, what is that? ARC was created by Francois Chalet, who works at Google, specifically designed as a pure measure of fluid intelligence. It gets away from language entirely. Visual puzzles. Exactly. It consists of novel visual reasoning puzzles. You're shown a few examples where an input grid of colored squares transforms into an output grid according to some hidden abstract rule.
15:15Then you're given a new input grid and you have to apply that rule you just inferred to produce the correct output grid. Sounds simple enough for humans. They're designed to be pretty imprudent for us, yeah. Humans typically solve these ARC tasks with around 75-85 % accuracy, sometimes higher, but they are exceptionally difficult for AI systems. Because they simply cannot be solved by retrieving memorized knowledge or just matching statistical patterns from the training data, you have to infer the underlying abstract transformation rule from just a few examples and then apply it flexibly. And how do LMMs do, even the best ones?
15:50Terribly. Even state-of-the-art models like GPT-4 or the latest multimodal models like GPT-4O, even when they're guided with really sophisticated prompting strategies designed to help them reason step-by-step visually, they consistently achieve scores below 20%, sometimes much lower. Wow, that's a huge gap compared to humans. It's massive. And the sources identify clear reasons for this failure. Things like a limited ability to compose different skills together, a basic unfamiliarity with abstract formats like these 2D grids, they're used to sequences of text, and fundamental deficiencies in their autoregressive token-by-token decoding process, which isn't well suited for this kind of holistic spatial reasoning.
16:30So ARC really demands accommodation. It absolutely demands it. You have to induce a novel schema, a new rule on the spot. The consistent failure of even the best LLMs to do this strongly validates this assimilation accommodation gap. It shows they're great at applying known patterns but struggle deeply with generating genuinely new ones. Okay, so we've clearly identified this profound gap. LLMs excel at assimilation using what they know but really struggle with accommodation adapting or creating new understanding. But researchers aren't just throwing up their hands, are they? No, definitely not.
17:05What are the cutting-edge ideas? What are people exploring to try and bridge this gap to move LLMs towards truly deeper, more flexible reasoning? There are several exciting directions, but they mostly fall into three emerging paradigms. First, there's work on eliciting intrinsic somata. Building on Zhu's work you mentioned. Exactly. Building on that idea that reasoning paths might already be latent in the model. This involves exploring more sophisticated decoding strategies, not just taking the single most likely next token, but maybe looking at the top few likely paths, top sampling, or using beam search to explore multiple sequences trying to find and surface these hidden reasoning chains.
17:46But is that creating new strategies? That's the catch. It's still primarily about selecting from a sort of static library of patterns the model learned during training, even if those patterns were latent. It's not really about creating truly novel strategies on the fly when faced with something completely unexpected. So better assimilation, perhaps, but not quite accommodation. Okay, what's the next approach? The second major area is engineering accommodation through neurosymbolic AI. This is really fascinating. It proposes building hybrid systems. Yeah, combining the strengths of neural networks, which are great at that fast, intuitive pattern recognition, like Daniel Kahneman's System 1 thinking with the strengths of classical symbolic AI, which is better at slow, deliberate, rule-based logical reasoning, more like System 2.
18:32So best of both worlds. How would that work? The idea is maybe an LM could translate a complex natural language query into a more formal symbolic representation, like a logical statement in prologue or nodes in a knowledge graph. Then a separate specialized symbolic reasoning engine could apply strict logical rules to manipulate that representation and find an answer. And then the LLM translates back. Precisely. The LLM translates the symbolic result back into understandable natural language. There are also promising techniques within this space, like vector symbolic architectures, VSAs. Yeah, they offer a way to represent and manipulate abstract symbols like concepts, objects, relationships, directly as high-dimensional vectors right within the neural network's mathematical space.
19:15It's like giving AI the ability to perform mathematical operations directly on abstract ideas, potentially bridging that gap between fuzzy pattern recognition and crisp logical reasoning. Interesting. And the third paradigm? The third and perhaps most ambitious approach is learning to accommodate through Piagetian-inspired AI and world models. This is really trying to get at the heart of the matter. Trying to copy human development. In a sense, yes. It aims to replicate the developmental process of intelligence itself rather than just the end result. We're starting to see new benchmarks being developed, like one called Coggle-LM, which are explicitly based on Piaget's own stages of cognitive development in children.
19:56Like object permanence, things like that. Exactly. Testing for things like object permanence, understanding conservation of volume, basic causal reasoning milestones that children achieve as they develop. The idea is to track an AI's cognitive age as it learns. So the goal isn't just a smart AI, but an AI that learns more like a child. That's the long-term vision. To move systems from this rigid assimilation we see now towards true accommodation, creating systems that can dynamically restructure their internal frameworks, their understanding of the world, when faced with conflicting evidence or novel situations, without needing that massive offline retraining process.
20:36That sounds incredibly challenging. It is. But a very practical avenue people are exploring within this is what's called human AI co-creation of schemas. Humans and AI working together to build understanding. Sort of. Imagine an AI helping a human analyst sift through vast amounts of data to identify recurring abstract patterns, maybe finding the hero's journey archetype across thousands of stories, or identifying common failure modes in complex system logs. The AI helps abstract the pattern, the schema. And that helps the human. It helps the human, but it could also create really valuable training data examples of how to form new abstractions that could potentially train future AI systems to become better at this process themselves, teaching them how to accommodate.
21:18Fascinating. So wrapping this all up, where does our deep dive leave us? Well, I think our analysis really confirms that core insight we started with. Current LLM intelligence is indeed about prediction plus structure, but perhaps more precisely using Piaget's framework. It's token prediction plus assimilation into external or sometimes latent internal schemata. Right. And adding that word assimilation is subtle but crucial. It highlights the fundamental missing piece. Which is accommodation. Which is accommodation. The primary conclusion, I think, from our sources in this discussion is that the defining gap between today's really powerful LLMs and genuine human-like general intelligence lies squarely in that mechanism of accommodation.
22:02It's the ability to figure out what to do when the procedure is unknown, when your current mental map just isn't working anymore. Not just flawlessly executing what you already know how to do. Perfectly put. And the consistent failure of these incredibly capable LLMs on pure fluid intelligence tests like ARC isn't just some minor flaw or something we can easily patch. It's really a signal of a fundamental architectural limitation in their current design. The grand challenge for reaching AGI, artificial general intelligence, isn't just perfecting assimilation, making them better pattern matchers or predictors.
22:37It's engineering somehow the capacity for genuine accommodation, for flexible, adaptive reasoning in the face of true novelty. Which leads to a really provocative final thought, doesn't it? What might it truly mean for an AI system to actually feel something akin to cognitive disequilibrium? That uncomfortable friction when its understanding of the world fundamentally breaks down. Right. Could we even engineer that? And if we could, if we could somehow instill that fundamental human driver of learning and adaptation that need to resolve dissonance into machines, what unforeseen breakthroughs or perhaps even unforeseen consequences could emerge?
From the publisher
We investigate the nature of intelligence in Large Language Models (LLMs), arguing that their impressive capabilities stem from next-token prediction (NTP) combined with externally supplied cognitive structures, primarily Chain-of-Thought (CoT) prompting. It critically examines this "NTP + Schemata" model through the lens of Jean Piaget's theory of cognitive development, differentiating between assimilation (fitting new information into existing frameworks) and accommodation (altering frameworks to account for novel information). The analysis posits that while CoT facilitates assimilation by providing a reasoning template, current LLMs lack the capacity for true accommodation, highlighting a fundamental "assimilation/accommodation gap." This limitation is further underscored by their struggle with fluid intelligence tasks, such as those found in the Abstraction and Reasoning Corpus (ARC), which require dynamic schema creation rather than just pattern application. We conclude by exploring future directions, including neuro-symbolic AI and Piagetian-inspired learning, as potential pathways to bridge this gap and foster more adaptive machine intelligence.




