In short
Defines AGI as an AI matching/exceeding a well-educated adult’s cognitive versatility and proficiency, then proposes a quantifiable “standardized AGI score” (0–100%) using CHC (Cattell–Horn–Carroll) theory and adapted human psychometric batteries across 10 equal-weight domains.
Guest backgrounds
No guests are mentioned in the transcript.
Key claims
AGI progress is measurable but misleading if only aggregate scores are reported; models show a “jagged cognitive profile” with extreme strengths and catastrophic bottlenecks. Long-term memory storage is a major architectural gap; hallucination reflects long-term memory retrieval failure. Workarounds (large context windows, RAG) mask missing internal memory.
Notable examples
GPT-4 ~27% vs GPT-5 ~58% total; both score 0% on long-term memory storage, ~4% on retrieval precision (hallucination test: inventing a nonexistent “Napoleon South African campaign”). Speed also low (~3%).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding AGI Framework
0:29 to 1:24
Exploring the comprehensive definition and framework for AGI.
“It makes talking about when it's coming almost impossible.”
The CHC Theory and Its Importance
1:24 to 2:32
Discussing the CHC theory and its significance in measuring AGI.
“And here's where the method gets, I think, really powerful.”
AGI Scoring Breakdown
2:32 to 3:28
Detailed analysis of how AGI scores are calculated and their implications.
“Yeah, the framework divides general intelligence into 10 core cognitive domains.”
Evaluating GPT-4 and GPT-5 Scores
3:28 to 4:32
Comparing AGI scores of GPT-4 and GPT-5 and what they indicate.
“They immediately quantify both incredible progress and these huge remaining gaps.”
Strengths of Current AI Models
4:32 to 6:08
Identifying the strengths of GPT-4 and GPT-5 in various cognitive domains.
“That's the author's key finding, what they call the jagged cognitive profile.”
Weaknesses and Gaps in AI Models
6:08 to 7:16
Examining the key weaknesses of GPT-4 and GPT-5 affecting their AGI scores.
“It covers science, social science, history, culture, and critically, common sense.”
Impact of Long-Term Memory Deficits
7:16 to 8:22
Discussing the consequences of long-term memory deficits in AI performance.
“This is where the core architectural weaknesses really pop out.”
Capability Contortions in AI Development
8:22 to 11:24
Exploring how current AI models compensate for weaknesses through workarounds.
“Think about learning over days or weeks.”
The Engine Analogy for Understanding AGI
11:24 to 14:01
Using the engine analogy to illustrate the importance of holistic cognitive repair.
“These workarounds create a brittle, fragile illusion of general capability.”
Understanding AGI Limitations
14:01 to 14:40
Learn how defective components limit AI potential in achieving AGI.
“So in AI terms, you might have highly optimized knowledge recall from pre-training, like good pistons and amazing mathematical skills, great carburetor.”
Show all 12 chapters
Importance of Full Cognitive Profiles
14:41 to 15:22
Discover why a complete cognitive profile is essential for evaluating AI.
“What stands out to you is the final, most crucial takeaway from the creators.”
Questions to Evaluate AI Systems
15:23 to 16:19
Explore key questions to assess the memory and learning capabilities of AI.
“It's not truly general intelligence yet.”
Transcript
Automatic transcript. May contain errors.0:28Okay, let's unpack this. again. Exactly. It makes talking about when it's coming almost impossible. Precisely. And that ambiguity, it kind of obscures the very real gap between today's specialized AI and, you know, genuine human level thinking. So the source material we're looking at today tried to cut through all that noise. Yeah. It proposes this really comprehensive, quantifiable framework, something designed specifically to measure that gap. Okay. So that's our mission today, a deep dive into this new methodology. How do they start? Well, they start by giving AGI a formal, pretty crystal clear definition.
1:02AGI is an AI that can match or exceed the cognitive versatility and proficiency of a well-educated adult. Okay, versatility and proficiency. Those words seem important. They're absolutely crucial. They intentionally stress both breadth, that's the versatility, the range of tasks and depth, the proficiency, how good it is at those tasks. Right, not just a one-trick pony. Exactly. And here's where the method gets, I think, really powerful. To measure this systematically, they anchor their whole approach in the Cattlehorn-Carroll theory of cognitive abilities, you know, CHC theory. CHC theory. Okay, why lean on that?
1:39Why is using like 100 years of human psychology research better than AI folks just making up their own benchmarks? Great question. It's because the CHC model is pretty much the most empirically validated model of human intelligence we have. It breaks down human thinking hierarchically into distinct abilities. It's systematic. So instead of just a random grab bag of tasks. Right. They use adapted human psychometric batteries. These are the gold standard tests used on people, but applied to AI systems. Ah, okay. So we're actually comparing AI performance against the structure of the human mind itself, apples to apples.
2:12You got it. And the end result, the output, is a standardized AGI score. It's a percentage, 0 % to 100%. Where 100 % means? 100 % signifies full AGI. The AI is equivalent to that well-educated adult benchmark. Okay, so if 100 % is the target, how do they break down that score? You mentioned different abilities. Yeah, the framework divides general intelligence into 10 core cognitive domains. And importantly, each one gets equal weight 10 % each. Why equal weight? To really prioritize that versatility, the breadth, you can't just max out one specialized skill and claim general intelligence under this system.
2:48Makes sense. What are these 10 domains? They cover the spectrum. You've got general knowledge, K, reading and writing, R, W, mathematical ability, M. On the spot reasoning, R, working memory, W, M, then two types of long-term memory, storage, MS, and retrieval, MR, visual processing, V, auditory processing, A, and finally speed, S. Wow, okay. That's comprehensive. It really does seem to cover the bases of human thinking. It does. It forces the AI to perform across that entire range. All right, here's where it gets really interesting then. Let's dive into the numbers. What does this NEWS4 board actually tell us about, say, the most advanced models we have now?
3:26Okay, the results are, well, they're fascinating. They immediately quantify both incredible progress and these huge remaining gaps. So take GBT4, tested back in 2023. Yeah, what did it get? Its estimated total AGI score was 27%. Wait, 27 %? Seriously. We were told GPP4 was this revolutionary moment, changed everything. How can its score be that low? What's this metric catching that, you know, standard benchmarks missed? Well, standard benchmarks often focus on specific skills, right? Like passing exams or coding. This framework forces the system to perform across everything, including really fundamental human stuff like memory acquisition, visual processing, reasoning under uncertainty, areas where it might struggle.
4:09Okay. Okay. That makes sense. So 27 % for GPT-4. What about its successor? Right. So then they estimated the score for GPT-5, and it achieved an estimated total AGI score of 58%. 58%. Wow. Okay. That's more than double in just two years. That is massive progress. It really is. Huge leap. But you mentioned gaps. These aggregate scores, 27%, 58%, they're averages, right? Yeah. I imagine the most revealing thing isn't the total number, but maybe the shape of the profile underneath. Precisely. That's the author's key finding, what they call the jagged cognitive profile. Jagged. Yeah, jagged. When you actually look beneath that 58 % average for GPT-5, you see these extreme strengths, like near-perfect scores in some areas, but they're countered by, frankly, catastrophic deficits in some really foundational cognitive machinery.
4:57So like a high-performance car with amazing acceleration, but maybe no brakes. Kind of, or maybe a powerful engine, but it's running on only a couple of cylinders, and one of them is completely rusted shut. That's the picture. Okay, let's focus on the cylinders that are firing perfectly first then. Where are these models, especially GPT-5, really excelling? Where does that vast training data pay off? They absolutely dominate in the knowledge-intensive domains. Symbolic reasoning, language, look at mathematical ability. M, GPT-4 was pretty limited, scored only 4%. GPT-5, it hits the full 10 % allowed for that domain.
5:37Just exceptional capabilities across arithmetic, algebra, geometry, probability, calculus. The works. Wow, that leap from 4 % to 10%, like instantaneous mastery. Yeah, okay. That really shows the power of the scaling they've been doing. It does, and it's a similar story in areas built on language and pattern recognition. Reading and writing ability, RW for GPT-5, also hit the full 10%. And that covers? Everything from basic sentence comprehension up to complex document analysis, sophisticated writing, even expert-level proofreading, and English usage. Okay, and general knowledge. How does it do on just knowing stuff?
6:11Also very strong. General knowledge, K. GPT-5 scored 9%. GPT-4 was already good at 8%. And this isn't just trivia. It covers science, social science, history, culture, and critically, common sense. Common sense. That's always been a tricky one for AI. How do they test that? Well, common sense is technically a narrow ability within the CHC framework, but it's vital. They test it with simple questions that rely on background understanding. The example they give in the sources is, does making a sandwich usually take longer than baking a loaf of bread? Right. Any human knows that instantly. Obvious.
6:44Exactly. And these advanced models are now demonstrating that kind of implicit understanding, mostly by leveraging just unbelievable amounts of pre-trained data to infer these relationships that are just obvious to us. The high scores here really show the success of modern training methods. Okay, so math, reading, writing, general knowledge, top marks or near top marks. Now, let's pivot to those jagged edges. If the profile is jagged, where exactly does the curve just fall off a clasp? What's holding back that 58 %? Right. This is where the core architectural weaknesses really pop out. And honestly, it's quite shocking when you see the numbers.
7:21Maybe the single most significant bottleneck they identified is long-term memory storage. MS. Long-term memory storage. Okay, that sounds fundamental. The ability to learn new things and remember them. What was the score? Yeah. For both models. A perfect zero. Sorry, what? Zero percent. Both GPT-4 and GPT-5 scored zero percent in long-term memory storage. this domain tests the ability to acquire, consolidate, and stably store new information from recent experiences. Things like associative memory, meaningful memory, verbatim recall of new facts. Zero percent. That, that's not a weakness. That feels like a fundamental missing piece of architecture.
7:58I'm struggling to wrap my head around a system that can ace calculus, but has basically no ability to form new memories from its interactions. Worse than a goldfish, like you said earlier. It means the models suffer from this fundamental constant state of amnesia. You finish a conversation, start a new one, and poof. It has to relearn the context all over again. It severely limits its usefulness for anything that needs continuous cumulative learning. Think about learning over days or weeks. Impossible. That's absolutely crippling for anything aiming for general intelligence. Wow. And it's closely related to another deficit, long-term memory retrieval, MR.
8:34The overall score there is also low, just 4 % for both models. Okay, 4 % is in zero, but it's still very low. What's the issue there? Well, they show high fluency. They can rapidly generate ideas or concepts related to a prompt, but they fail completely, like 0 % at retrieval precision. Retrieval precision, meaning? Meaning the ability to avoid making stuff up. Confabulation. Hallucinations. Ah, okay. The hallucination problem again. So that 0 % precision is essentially the technical measure of hallucination, isn't it? Pretty much, yeah. The systems fail badly when tested on specific, deliberately false premises that they shouldn't fall for if they were checking against their knowledge.
9:14They use a great example. You might ask it. Describe the key strategy that Napoleon Bonaparte used to win his South African campaign. Okay, but Napoleon never campaigned in South Africa. Exactly. A well-educated adult or even just someone with decent historical knowledge would immediately flag the premise as false. The AI often can't. It might just start generating plausible sounding but entirely fictional details about this non-existent campaign. That is brilliant testing. It's not just checking facts. It's testing if the AI can validate the question itself against its internal knowledge. So they can spit out information fluently but have zero reliable mechanism to check if it's actually true based on their own massive internal memory.
9:55Precisely. They lack that internal validation loop. moving quickly through other gaps. On-the-spot reasoning are, that's fluid intelligence, deduction, induction, theory of mind, planning. GPT-4 scored a negligible 0%. Zero again. Ouch. Yeah. Now, GPT-5 made a substantial leap here, up to 7%. Big improvement in planning and theory of mind. But it still struggles with adaptation, the ability to flexibly figure out and change unstated rules based on feedback. Okay, still a ways to go on core reasoning. And what about speed? you'd think a giant AI running on supercomputers would ace speed s. You'd think so, but not entirely.
10:33Both models scored a low 3 % on speed. 3 %? How? They can calculate instantly. Well, yes, they can compute simple things very fast. But the speed domain includes other things like perceptual speed, reaction time, reading and writing speed in a more holistic sense. And complex or multimodal processing is often underdeveloped or slow. The researchers specifically note that GPT-5 often needs really long thinking time, you know, that internal chain of thought reasoning. Yeah, you see the pauses sometimes. Right, and that time penalty dramatically suppresses its overall speed score in this framework.
11:07Okay, so stepping back, what this framework reveals isn't just like a shopping list of missing features. It sounds like it's revealing a pattern of, what do they call them, capability contortions, where developers are basically patching over these deep architectural problems with expensive workarounds. That is probably the most critical insight here for listeners. These workarounds create a brittle, fragile illusion of general capability. It looks smart on the surface, but the foundation is weak. Okay, give me an example. Contortion 1. Contortion 1, the reliance on absolutely massive context windows.
11:39This is essentially maximizing working memory, WM. Right, the short-term scratch pad. Exactly. They make it huge to compensate for the complete lack of long-term memory storage, MS, the 0 % score we talked about. Ah, so they're using that giant context window like an incredibly inefficient, computationally expensive temporary notebook. It holds the state for the current conversation, making it seem like the AI remembers, but only for that session. You've nailed it. But it's a strategy that just fails to scale. Maintaining that huge context is incredibly resource intensive, and it completely breaks down for tasks that need knowledge accumulated over days, weeks or longer.
12:19Genuine long-term learning needs the model to actually adjust its internal weights, integrate new information permanently, not just keep a long text string in temporary RAM. Right. It's not really learning. It's just holding. Yeah. Okay, contortion two. This must address the hallucination problem, that long-term memory retrieval MR deficit. Exactly. Developers use Retrieval Augmented Generation, or RAG, basically hooking the AI up to external search tools or databases. Yeah, RAG is everywhere now. It's a clever piece of engineering, for sure. But fundamentally, it's a mask. It's papering over two core weaknesses simultaneously.
12:54Which are? One, the AI's inability to reliably access and verify information from its own vast static pre-trained knowledge, the hallucination issue. And two, it masks the total absence of dynamic, personalized, experiential memory. The kind of memory you need to remember, private interactions, evolving contexts, things learned yesterday. So fetching fax from Google or a database isn't a substitute for actually having an integrated, reliable, long-term memory. Not at all. It's essentially outsourcing the memory function because the internal component is broken, or in the case of MS, non-existent.
13:30Got it. Just patching the holes. Precisely. And this brings us back to a really powerful analogy they use in the paper, the engine analogy. Okay, the engine analogy. How does that work? Well, think of overall intelligence, like AGI, as the total horsepower of a high-performance engine. The key idea is that the engine's total output, its power, is ultimately constrained by its weakest components. Right. Like you could have the world's best fuel injectors and pistons, but if you have a cracked engine block or a broken crankshaft, your engine isn't going anywhere fast. Exactly that. So in AI terms, you might have highly optimized knowledge recall from pre-training, like good pistons and amazing mathematical skills, great carburetor.
14:10But if you have several highly defective parts, specifically that 0 % long-term memory storage crankshaft, that defective part immediately limits the overall horsepower of the whole system. So achieving AGI isn't just about making the existing strong parts even stronger through more data or compute. No, it's about actually fixing and integrating the entire holistic cognitive machinery. You have to repair or replace those broken components. This framework really does sound like a rigorous diagnostic tool then. It gives researchers clear, measurable targets for what's actually broken. What stands out to you is the final, most crucial takeaway from the creators.
14:48Their really strong recommendation is this. We should always report the full cognitive profile, the whole jagged line, not just the single aggregate AGI score number. Why is that so important? Because a high overall score on its own can be dangerously misleading if critical bottlenecks still exist. Ah, that makes perfect sense. Like an AI scoring, say 90 % overall, sounds absolutely incredible, right? statistically near human parity. But if that impressive 90 % average hides a 0 % score on long-term memory storage MS... Then you have a system that is still fundamentally functionally impaired. It has profound amnesia regardless of its mathematical genius or writing skill.
15:26It's not truly general intelligence yet. Right. It illustrates that AGI isn't just about hitting a certain number. It's about achieving that holistic integrated capability, Something that mirrors the robust, error-checking, persistent architecture of the human mind. The real path to AGI runs directly through fixing those specific 0 % bottlenecks. Couldn't have said it better myself. So, okay, what does this all mean for you, the listener? Well, next time you interact with one of these incredibly advanced AI systems, one that can write a flawless essay, generate stunning images, solve complex physics problems, don't just admire its raw power, ask yourself this.
16:01Can this system remember what I told it last week? Can it learn a new, maybe arbitrary procedure? I teach it right now and retain that knowledge permanently without me having to remind it in the context window next time. Good questions to ask. If the answer is no, then you're still interacting with a jagged intelligence. Powerful, yes. Impressive, absolutely. Yeah. But fundamentally incomplete. Achieving AGI, it seems, isn't just about raw capability. It's really about conquering amnesia.
From the publisher
This paper attempts to provide a comprehensive framework grounded in the Cattell-Horn-Carroll (CHC) theory of cognitive abilities, which breaks down general intelligence into ten core, equally weighted cognitive components, such as General Knowledge (K), Mathematical Ability (M), and various forms of Memory (WM, MS, MR). The paper details specific, measurable tasks and examples—often drawn from human-centric tests like AP exams or psychometric assessments—to evaluate AI performance in each area. Furthermore, the analysis reveals that current AI systems like GPT-4, despite excelling in some narrow tasks, lack core cognitive capabilities, relying on compensatory "capability contortions" (e.g., using large context windows for lack of long-term memory) that mask their overall deficiencies compared to the human benchmark.




