In short
The episode argues that today’s frontier LLMs (trained via statistical learning/next-token prediction) are misaligned with the deductive, zero-error reasoning needed for general intelligence and safety. It claims models learn statistical shortcuts, fail under small input shifts, and that “exact learning” (correct on all valid inputs) is essential and potentially achievable.
Guest backgrounds
No guests are named in the transcript; it’s a two-host “Deep Dive” discussion.
Key claims
Average performance doesn’t imply exact correctness; examples include GPT-4 failing basic arithmetic word problems, family-relation questions, and other structured tasks; travel-planning success reported at 0.6%. Gradient descent and surrogate losses may prevent reaching exactness.
Notable examples
travel planning (0.6% success); addition/subtraction word problems; cousin/family relation reasoning; counting items; shift-cipher decoding; binary-string classifier examples requiring exponential samples.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Flaws of Powerful AI
0:45 to 2:20
Discussion on the paradox of advanced AI systems struggling with basic reasoning tasks.
“How can something so powerful be so fundamentally flawed in basic reasoning?”
Understanding Statistical Learning
2:20 to 4:28
Explanation of how current AI systems are built on statistical learning and its shortcomings.
“And surprisingly, maybe within reach, yes.”
The Concept of Exact Learning
4:28 to 6:31
Introduction to the idea of exact learning and its necessity for reliable AI.
“It's like learning to pass by, recognizing common question patterns and answers, instead of truly understanding the underlying logic or rules.”
Challenges in Achieving Exactness
6:31 to 8:13
Discussion on the hurdles of transitioning from statistical to exact learning.
“Achieving this ensures safety and reliability, which becomes absolutely critical as AI gets more agency, making real-world decisions.”
Towards Improving AI Reliability
8:13 to 10:46
Exploration of concrete suggestions for moving towards exact learning in AI.
“Because of the statistical shortcuts, yes.”
Exploring Model Symmetries
14:00 to 14:48
Understanding how to modify learner models by adjusting symmetries for better performance.
“or even algorithms for verifying correctness formally, though that's hard.”
Targeted Teaching Techniques
14:48 to 15:42
The importance of providing specific examples to improve AI learning outcomes.
“We can potentially help the learner by giving it very specific, carefully chosen examples, a teaching set.”
Changing the Learning Task
15:42 to 16:46
Adjusting the learning problem framing and loss functions to enhance reasoning in AI.
“But a really promising direction, especially for LLMs, is training them using chain of thought reasoning traces.”
Challenges of Exact Learning
16:46 to 18:30
Discussing the challenges and objections to transitioning to exact learning in AI.
“Okay, so there are definitely paths forward being proposed, but this whole shift towards exact learning, it's a big change.”
Symbolic Inputs and Limitations
18:30 to 19:56
Exploring the use of symbolic inputs in AI and its limitations in real-world applications.
“We can reduce ambiguity by instructing models to output answers in specific formats.”
Show all 11 chapters
The Case for Exact Learning
19:56 to 21:41
Arguing the need for a shift from statistical learning to exact learning for AI reliability.
“Okay, so wrapping this up then, what's the main takeaway from this deep dive?”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. Today we're diving headfirst into, well, one of the most perplexing paradoxes in modern AI. It really is quite something. Yeah. I mean, on one hand, you've got these truly astonishing frontier AI systems, you know, the kind that can write complex code, generate really creative text, even help scientists solve tough problems. Absolutely. Feats that were science fiction just a few years ago. Exactly. But here's the twist, and it's a big one. Despite all this incredible power, these same systems often trip up on seemingly simple, logical reasoning tasks. It's like having, I don't know, a brilliant mathematician who can solve incredibly complex theorems but struggles to balance a checkbook consistently.
0:45That's a good analogy, actually. How can something so powerful be so fundamentally flawed in basic reasoning? OK, let's let's unpack this. Well, what's fascinating here and what the research points to is that the core issue isn't just about needing more data or faster chips. OK. The way we currently build these AIs, it's rooted in something called statistical learning. Right. And that approach, as our sources detail, is fundamentally different from the kind of deductive reasoning we expect from general intelligence. Deductive reasoning? Like logic puzzles. Well, yes, but it's much broader. It's the kind of thinking we need for reliable answers.
1:22The misalignment is quite precise. And that deductive reasoning, you're saying, it's everywhere in our lives, right? It was the critical. It's about getting new, absolutely certain knowledge from facts and rules we already have. Exactly. Like making a shopping list based on your meal plan, knowing you will have enough stuff. Yes. Or figuring out complex tax rules. You can't be mostly right with your taxes. Oh, no kidding. Or calculating insurance premiums or shipping costs from a table. I mean, for these things, every step, every conclusion has to be correct. Zero margin for error. Yeah. Precisely.
1:57And the core argument from the research we're looking at is that this unsound behavior, these failures in AI reasoning, are a direct result of that statistical learning foundation. So this deep dive, we're going to explore why a shift towards what's called exact learning, which demands correctness on all inputs, not just most, is maybe essential. And possibly even achievable. And surprisingly, maybe within reach, yes. Yeah. For reliable deductive reasoning, anyway. Okay, so let's dig into that. How are these cutting-edge systems like GPT-4, Gemini, how are they actually built? You mentioned statistical learning.
2:33Right. So primarily, they're trained to minimize something called next token prediction error. Next token prediction. Basically guessing the next word. Pretty much. Across absolutely massive data sets, trillions of words, they just keep trying to guess the next piece of text as accurately as possible over and over. Okay. So they get really good at guessing words in context. Extremely good. But here's the limitation. This statistical approach only guarantees good results on average. On average. And critically, only for data that looks very, very similar to what it was trained on. Ah. It really struggles with even small deviations, what researchers call small covariate shift.
3:10So slight changes in the input question or a slightly new scenario. Exactly. Things just a little bit outside the norm of its training can throw it off significantly. And this is where it gets, well, really eyebrow raising. The failures. They're documented, right? On tasks that seem like they should be easy. Oh, incredibly well documented. And yet not obscure problems. We're talking about things that make you scratch your head. Like what? Give us some examples. Okay. Well, imagine GPT-4, this incredibly fluent system, failing basic addition or subtraction word problems. Seriously? Yes. Or struggling with simple family relations like figuring out who's cousin based on a description.
3:48Wow. And it's not just one study. Research from, well, the last few years consistently shows this pattern. Hmm. There was one study looking at various travel planning problems. GPT-4 had a shockingly low 0.6 % success rate. Less than 1 % on travel planning. Less than 1%. Or simple logical tasks. Counting items accurately, swapping articles like A and seeing your correctly, even decoding simple shift ciphers. It just feels like it's missing something fundamental if it can't reliably do that kind of structured reasoning. That's what this collection of research strongly suggests. These powerful models, these transformers, they learn what are called statistical shortcuts.
4:30Shortcuts. Like cheating on a test. Kind of, yeah. It's true. It's like learning to pass by, recognizing common question patterns and answers, instead of truly understanding the underlying logic or rules. So it looks good on the stuff it's seen a million times. Exactly. It gets enough right to look good on average, or in distribution, as they say. Dota. But it fundamentally misses the deeper rules sometimes. The moment you ask at a slightly different version of the problem, one that breaks the shortcut, it can just fall apart. Okay. Some studies have even shown models that are perfect on their training data plummeting to basically random guessing on slightly altered but still valid distributions.
5:11So it's a persistent issue then, and I assume people are trying to fix this? Oh, absolutely. The AI community is constantly working on it, manipulating training data, tweaking the learning algorithms, changing the model architectures. But are these fixes getting at the root cause? Or are they just patches? That is precisely the question this paper raises. It points out, and I'm quoting loosely here, that history shows people always discover new problems of the same character as the old ones. Implying these might just be localized fixes. It suggests that yes, not necessarily addressing the foundational statistical versus deductive mismatch.
5:49Which brings us to this idea of exact learning. What exactly is that compared to statistical learning? It's a fundamentally different goal. Instead of optimizing for average performance across a distribution of inputs. Which is what we do now. Right. Exact learning demands correctness on all inputs for a specific task. All inputs. All well-formed inputs, yes. It's about getting every answer right every single time. No exceptions allowed for that task. Wow. Okay, why is that so important? Well, think about those reasoning tasks we mentioned. Taxes, navigation, engineering calculations. Rules are clear.
6:21Correct answers are well-defined. Right. So demanding flawless performance on any valid input isn't just reasonable. It's, well, highly beneficial. Makes sense. And we know flawless deduction is achievable news. Formal logic systems prove that. Okay. Achieving this ensures safety and reliability, which becomes absolutely critical as AI gets more agency, making real-world decisions. Yeah, you don't want your self-driving car to be correct on average. Exactly. Yeah. And it also frees designers from that incredibly difficult, maybe impossible task of deciding which mistakes are okay and which aren't.
6:57That's a good point. Who gets to decide what errors are acceptable? It's a problem with no perfect answer. Exactness sidesteps it. And this ties directly into the idea of systematic generalization. the ability to apply learned rules to totally new situations. That requires exactness. Okay, but if exactness is the goal, why stick with these learning systems at all? Why not just use, you know, old school rule-based AI? Those are exact, aren't they? That's a really good question. The advantage of learning systems, even with this challenge, is their ability to handle the messiness of the real world, especially natural language.
7:31They can grasp nuances, understand varied phrasing, things that brittle, formal rule-based systems choke on unless you explicitly code every single possibility. So a learning system could maybe infer my house has windows if it knows all buildings have windows and understands a house is a building. Precisely. Without needing that specific fact explicitly programmed, it blends rule-like reasoning with real-world understanding. That's the hope anyway. So the challenge is getting that learned understanding to be perfectly reliable, perfectly exact. Exactly. And the core problem is that good statistical performance does not imply good performance and exact learning.
8:12Because of the shortcuts. Because of the statistical shortcuts, yes. Those shortcuts mean the model might actually have a huge error, be no better than guessing, on certain critical but maybe rare inputs, even if his average score looks great. Okay. Can you make that a bit more concrete, like an analogy? Sure. Imagine you're teaching a child to identify a very specific, rare type of bird. Okay. Statistical learning is like showing them lots of common birds, pigeons, sparrows, and saying, great, you're getting most birds right, good job on average. Right. But exact learning would demand the child correctly identifies every single bird, including that super rare one they've maybe only seen once or twice, or maybe never in training.
8:53The standard changes completely. It's not about most, it's about all. Exactly. It fundamentally changes the learning problem. Does the paper give a more technical example? It does, yeah. It talks about a scenario with binary input sequences of zeros and ones. Let's call it classifying binary strings. Okay. Statistically, learning to classify these on average is, relatively speaking, easy. You only need a number of examples roughly proportional to the length of the string. Manageable. But for exact learning, say, telling the difference between a function that only recognizes the string of all zeros versus a function that just outputs zero for everything.
9:28Ah, a subtle difference. A very subtle difference. To guarantee you can tell those apart exactly, you might need an exponential number of examples. Basically, you might need to see almost every single possible string. Because that one specific all zero string might be super rare in your training data. Precisely. It might never show up in a typical random sample, so the statistical learner never gets the crucial information to distinguish those two functions perfectly. Wow. And it gets worse. They show another example with two simple linear classifiers. Imagine two simple rules. These rules might differ on only a tiny, tiny fraction of possible inputs, like less than a percent.
10:09Okay. But to reliably tell them apart exactly, a learner might need to see an astronomically large number of samples again. Yeah. exponentially large. So even for simple rules, if the difference is rare, statistical methods struggle to guarantee finding it. Exactly. Finding that needle in the haystack reliably requires seeing way too much hay, statistically speaking. Okay, that highlights the data problem. What about the models themselves? You mentioned symmetry earlier. Right. So most powerful modern learners, neural networks, transformers have built-in symmetries, things like label symmetry.
10:41Meaning if you swap cat and dog labels in training, it just learns to swap its predictions. pretty much, or input alphabet symmetry. It doesn't treat the letter A as inherently special compared to B, for example. They seem like good ideas for generalization, right? Yeah, it sounds like it makes the model more flexible. It does for statistical generalization. But here's the counterintuitive part the paper highlights. Aiming for more generality by adding more symmetries can actually delay exact learning. Wait, making it more general makes it take longer to be perfectly correct? Yes. These symmetries, while good for average performance, impose a kind of price when the goal is exactness.
11:21It makes it harder for the model to pinpoint those very specific rules needed for perfect accuracy. Huh. That is counterintuitive. It means some standard techniques to improve models might be hindering progress towards true reliability. And then there's the training process itself, gradient descent. That's how we train LLMs, right? It's the workhorse, yes. How does that fit into this? Well, gradient descent is great at finding good statistical solutions. For many problems, it converges to a pretty optimal solution in terms of average error. What? But achieving that final step to perfect exactness, getting rid of that last tiny bit of error on every single input, that can take an exponentially long time with gradient descent.
12:04It might get very close statistically, but never quite reach perfect reliability and practice. So the tool used to train them isn't really optimized for the goal of exactness? Not inherently, no. And then there's the specific way we measure error, the loss function. Like the next token prediction error we talked about. Exactly. That's a surrogate loss. It's easy to calculate and optimize, but it's not directly measuring logical correctness. And the paper points out something quite striking. Training with next token prediction can actually destroy an initially exact solution if the model happened to find one.
12:36You're kidding. Trying to get better at predicting the next word can make it less logically correct. It can, because the pressure is always just on that next word, not on the overall logical consistency of the entire output. So what's the ideal loss function for exactness? Ideally something simple, like a 0-1 loss. Zero error if the answer is perfectly correct, one if it's wrong in any way. It says. But that kind of loss function is notoriously difficult to optimize using methods like gradient descent. It's non-smooth, making the learning process much harder. So we're using training methods and error measures that are convenient but might be fundamentally misaligned with achieving perfect reliability.
13:20That's a key part of the argument, yes. Okay, this sounds challenging. Yeah. But the paper doesn't just point out problems right. It suggests solutions. How do we move towards this exact learning? Absolutely. It's not just doom and gloom. It proposes concrete directions. They group them into roughly four areas. Okay, what's for first? First is performance evaluation. Basically, we need better ways to test these systems. Beyond just standard benchmark. Exactly. Static benchmarks have limited utility here. We need a more systematic understanding of the failures and failure modes. How do we get that?
13:52By developing frameworks where systems are actively challenged. Think adversarial testing, trying to find inputs that break the model. or even algorithms for verifying correctness formally, though that's hard. So probing for weaknesses systematically. Right. Drawing inspiration from formal methods, adversarial attacks, and how we study systematic generalization. Okay, makes sense. Area number two. Changing the learners themselves. Remember the symmetry issue. Yeah. More symmetry delaying exactness. So one idea is to explicitly remove some symmetries from the models or force them to respect certain structures in the data, making them equivariant.
14:28Equivariant. Equivariant, meaning if you rotate the input image, the cat detection part rotates with it. Sort of, yes. Building in knowledge about how transformations affect the output. This has shown promise in reducing the amount of data needed in statistical settings, so it might help for exactness, too. Interesting. Okay, third area. Teaching the learners. This one's really intriguing, since for many reasoning tasks, we know the correct algorithm or rule we want the AI to learn. Like the rules of arithmetic. Exactly. We can potentially help the learner by giving it very specific, carefully chosen examples, a teaching set.
15:04Like giving it the most informative examples. Precisely. For some problems, like certain classifiers, providing just a few crucial examples may be points right on the decision boundary, can actually guarantee the learner finds the exact solution much faster than random data. So targeted teaching instead of just massive random data dumps. Yes. It connects to an older idea and learning theory called exact learning with a teacher, where a teacher provides key counterexamples or answers specific queries. What's the catch? The catch is that these AI systems need to be general, especially for handling natural language.
15:37So you can't control all the training data this precisely. It has to be a mix. Okay. Targeted teaching where possible. And the fourth area? Changing the task itself. Hmm. How we frame the learning problem. Like changing the loss function. That's one way. Though we saw the O1 loss as tricky. But a really promising direction, especially for LLMs, is training them using chain of thought reasoning traces. Ah, so not just giving the question and final answer, but showing the steps in between. Exactly. The idea is that learning the individual elementary reasoning steps might be a much simpler, more exactifiable problem for the AI.
16:14Okay. The paper gives a great example with propositional logic. Models trained just on question-answer pairs struggled. Right. But models trained on the same problems with reasoning traces, showing the intermediate logical steps achieve much better, near-perfect accuracy, both on familiar problems and new ones. So it learned the process of reasoning, not just the final outcome correlation. That seems to be the key. It learns to perform the steps reliably, rather than trying to bake complex, multi-step logic directly into its weights all at once. That sounds really promising. It is. Other related ideas include training a separate verifier model to check if a proposed solution is correct, or exploring neurosymbolic methods that explicitly combine learned representations with symbolic reasoning modules.
17:00Okay, so there are definitely paths forward being proposed, but this whole shift towards exact learning, it's a big change. I bet there's skepticism. What are the main objections people raise? Oh, absolutely. It's a significant paradigm shift. One common pushback is simply, hey, existing statistical methods are working incredibly well. Look at the progress. And they are impressive, undeniably. Yes. And maybe, some argue, just scaling these current methods up more data, bigger models will eventually lead to exactness anyway. So just keep doing what we're doing, but bigger. That's one viewpoint.
17:34But the counter argument, as the bilber puts it, is that as AI gains more real world agency. Right. Driving cars, managing finances. Even unlikely mistakes become unacceptable if they lead to potentially catastrophic outcomes. Acknowledging exact learning as the explicit goal, they argue, can clarify the research direction and potentially speed up progress towards true reliability. Okay. Safety demands a higher standard than good on average. What's another objection? A big practical one. Verifying exact learning is incredibly hard. Yeah. How do you prove a model is correct on all possible inputs, especially with messy natural language?
18:13It could require testing on an astronomical number of inputs. Potentially impossible. For example, comparing two imbit integers requires checking two 2-meter pairs. That grows insanely fast. It does. But the counter is that the situation is not hopeless. We can reduce ambiguity by instructing models to output answers in specific formats. Standardizing the output. Right. And we can borrow techniques from systematic generalization testing, using distinct training test distributions, and adversarial learning to actively seek out failures. So maybe we can't get perfect verification, but we can get much higher confidence.
18:49Exactly. And just knowing that exactness is the goal is valuable, even if perfect verification remains elusive. It sets the right target. Okay. Any other major objections? One more is the idea that, well, for some domains, we already use symbolic inputs, and that solves it. Like chess engines or mathematical theorem provers. Precisely. In domains like games, think alpha zero or formal math, inputs are already symbolic and unambiguous. Learning just speeds up finding the correct symbolic manipulation. Correctness is often guaranteed by the symbolic framework itself. So what's the issue there? It seems like it works.
19:22It works only when inputs are easily and perfectly translatable into that symbolic form. Ah. For natural language or images or sensor data from the real world, you need another system. often a learning-based system itself, an auto-formalizer, to do that translation. And that system has to be perfect. Exactly. The auto-formalizer must perform flawlessly to ensure correct solutions. And that, the paper notes, is highly non-trivial. Auto-formalization accuracy can be quite low, especially on new types of inputs. So you're just pushing the exactness problem onto the translation step? Pretty much. It doesn't solve the core challenge for AI dealing with the messy real world.
19:59Okay, so wrapping this up then, what's the main takeaway from this deep dive? The core argument really is that the dominant statistical learning approach, while powerful, is fundamentally misaligned with what we need for true general intelligence, especially in fields demanding perfect accuracy. Engineering, science, math, anywhere, mistakes are costly. Precisely. Where flawless formal reasoning is paramount. And the proposed path forward is this pivot towards exact learning, aiming for universal correctness, not just average performance. Yes. It's about providing a clearer goal, a more precise language for the AI community, potentially accelerating progress towards truly reliable systems.
20:42Less about patching statistical models, more about building for exactness from the start. That's the idea. Yeah. And the key research directions seem to be around more interactive learning teaching models more deliberately and changing the learning task itself, like using those chain of thought reasoning traces. Teaching the steps, not just the answers. Exactly. Focusing on the process of reasoning. So this leaves us with a pretty significant thought, doesn't it? As these AI systems become woven into the fabric of our society. Making increasingly critical decisions. How much error are we actually willing to accept?
21:14What does it mean if the most advanced intelligence we create can only ever be good enough on average and not reliably demonstrably correct when it truly matters? It really forces us to think about what we expect and what we need from artificial general intelligence. Is mostly right good enough? A crucial question. This focus on exact learning definitely feels like a necessary step towards unlocking that truly reliable general AI potential. Lots to think about there.
From the publisher
This paper argues that Artificial General Intelligence (AGI), particularly for tasks requiring deductive reasoning, demands a fundamental shift from statistical learning to exact learning. Current AI systems, based on statistical methods, excel on average but consistently fail on straightforward deductive tasks due to their inherent design, which optimizes for statistical performance over distributions. This leads to unreliable behavior and "statistical shortcuts", where models perform well on training data but poorly on slightly different inputs. The authors propose that exact learning, which requires universal correctness on all well-formed inputs, is crucial for achieving truly reliable and safe AI systems, despite the challenges in its implementation and verification. They suggest approaches like changing learning algorithms, curating teaching sets, and transforming tasks to facilitate this paradigm shift.




