Richard Sutton Declares LLMs a Dead End

20 Oct 2025 · 13 min · 3 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Richard Sutton argues LLMs are a “dead end” and will be replaced by continual-learning, goal-driven agents built from reinforcement learning (RL), not text imitation.

Guest background

Richard Sutton, “father of reinforcement learning,” inventor of TD learning and policy gradients; won the 2024 Turing Award.

Key claims

LLMs mimic what people say but don’t learn world cause-and-effect, so they can’t be meaningfully “surprised” or update beliefs (“bitter lesson pilled”). Next-token prediction is passive, not a goal. Without continual learning, models suffer poor out-of-distribution generalization and catastrophic interference.

Notable examples

curveball conversation outcomes; babies/animals learning via trial-and-error (not instruction); RL components: policy, value function (TD), perception, and transition/world model.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Sutton's Critique of LLMs

0:45 to 4:25

Exploration of Richard Sutton's argument that LLMs are limited and unable to learn continually.

“Because they're not, what was the phrase, bitter lesson pilled.”

Experiential Learning Framework

4:25 to 8:02

Discussion on the essential components of a successful continual learning agent as proposed by Sutton.

“Sutton lays out four absolutely essential parts for any general continual learning agent.”

The Future of AI and Intelligence

8:02 to 13:20

Sutton outlines the inevitability of designed intelligence and its implications.

“Okay, that philosophical difference has real technical consequences then.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Today, we're really going after the big consensus in AI right now. This idea that large language models, LLMs, are basically the only game in town for getting to advanced AI. Yeah, and it's not just anyone challenging it. We're diving into the perspective of Richard Sutton, you know, the father of reinforcement learning, the guy behind TD learning, policy gradients. Oh, and he just won the 2024 Turing Award. Right. A giant in the field. And his take is pretty radical. He's essentially saying LLMs are a dead end, like a technological cul-de-sac. Exactly. His core argument is, look, you can scale these things up forever, throw all the compute you want at them, but eventually they're going to be obsolete, replaced by architectures built for continual learning.

0:44And the reason? Because they're not, what was the phrase, bitter lesson pilled. That's the crux of it. It means they fundamentally can't learn on the job. They can't learn from their own experiences interacting with the world. Sutton believes this whole era of data we're in, it's ending. And being replaced by? An era of experience, where the AIs that succeed are the ones that are constantly learning from this loop of sensation, action, and reward. Okay, let's break that down. Sutton draws a really clear line in the sand between two approaches. LLMs, he says, are fundamentally about mimicking people.

1:21Yeah. Right? Using language, doing what people say you should do. That's the key distinction he makes. For him, reinforcement learning RL is the blueprint for, let's call it basic AI. The whole point of RL is to understand your world. It's about figuring out what to do. Not just predicting what someone else would say or has said. Precisely. Not just mimicking. Okay, because I think a lot of people assume you train these models on, you know, the entire internet practically. They must develop some kind of world model just to make sense of it all. Yeah, that's the common assumption. But Sutton totally rejects it.

1:53He argues LLMs build a model of what a person would say, not a model of the world itself. What's the functional difference there? It's huge. A real world model lets you predict what will happen if you take an action. LLMs, even when they're having a conversation, have no underlying prediction about how the world or even the person they're talking to will respond to what they just said. So if I throw a curveball, something totally unexpected, the LLM just generates the next statistically likely words that doesn't have an internal, wait, that doesn't fit my model of reality moment. Exactly. It can't be surprised in a meaningful way by an outcome it didn't predict because it wasn't really predicting the outcome in the first place, just the text.

2:32And if you can't be surprised, you don't fundamentally update your understanding of how the world works based on that surprise. For Sutton, that's the killer flaw. It prevents real deep learning. OK, hold on, because the obvious engineering fix seems to be, well, just combine them, right? Use the giant LLM as a knowledge base, a good starting point, a prior, as they say, and then build the RL, the experiential learning on top of that. Why is Sutton so against that? Because, he argues, a prior needs to point towards some kind of ground truth. With an LLM trained purely on text, there's no goal, no objective measure of what the right thing to do or say is.

3:11It's just predicting the average or the statistically common thing. So without a goal, it's just information soup, not knowledge you can act on reliably. Spot on. In RL, there is a ground truth. The right action is the one that gets you closer to your goal, the one that gets you reward. Without that goal-directed framework, the LLM's knowledge isn't really grounded in the world's cause and effect rules. It's a mimic, not an agent. Which brings us straight to goals. Sutton leans hard on John McCarthy's classic definition. Intelligence is the computational part of the ability to achieve goals. Right, and he says LLMs just fail this test right out of the gate.

3:49You know, proponents might say, well, the goal is next token prediction. Yeah, you hear that a lot. But Sutton's counter is pretty sharp. Predicting the next token isn't a substantive goal because you don't influence the tokens coming at you. It's passive. Like predicting the weather doesn't change the weather. Exactly. You're just observing. You can't really say a system has a genuine goal if it's just predicting things without any interaction or ability to affect the outcome based on those predictions. Okay, so if LLMs aren't the answer, we need this experiential paradigm, this constant loop of sensation, action, reward.

4:25What kind of system, what kind of architecture does that actually require? Sutton lays out four absolutely essential parts for any general continual learning agent. And crucially, this agent learns from everything coming in, all the sensory data, not just the tiny reward signal you might get occasionally. Okay, what are the four parts? So first, you need a policy. That's basically a decision maker. What action do I take right now, given the current situation? Simple enough. What's next? Second, the value function. This is usually learned through TD, temporal difference learning. And this is super important.

4:56It predicts the long-term outcome. It estimates how good the current situation is, not just for immediate reward, but for achieving the ultimate goal. Ah, so that stops the agent from just chasing immediate gratification. It can see the bigger picture, plan ahead. Exactly. It values states that lead towards winning the game, even if the win is far off. Third, you need a perception component. This is what builds the state representation, the agent's understanding of where am I right now? Okay. Policy, value function, perception. What's the fourth? The fourth, and arguably the most powerful, is the transition model of the world.

5:32A model of the world. So this is where the agent actually learns how things work, like its own internal physics engine. That's a great way to put it. It's the agent's learned beliefs about the consequences of its actions. If I do X, what's likely to happen next? And this model is built up richly from all the sensations the agent receives. It learns cause and effect deeply. Which is exactly what you were saying LLMs don't do. They don't learn the underlying cause and effect. Precisely. Let's go back to the bitter lesson for a second, because this is where it gets confusing. The bitter lesson says scale and computation win.

6:04LLMs are all about scale. Why does Sutton say they haven't learned the lesson? Well, they've got the scale part, yes, massive computation, but they also incorporate enormous amounts of specific human knowledge via that huge text corpus. And the bitter lesson warns against relying too much on human knowledge. Exactly. Sutton points out that historically, whenever we've tried to bake human expertise directly into AI systems, those systems eventually get overtaken by methods that learn more from scratch, just using computation and general learning principles applied to raw experience. Leaning on human knowledge seems to lock us into approaches that don't scale as well in the long run.

6:44So the very thing that makes LLMs so useful now, their ability to mimic human reasoning and knowledge, is, in Sutton's view, the thing that limits their long-term potential. Because it's mimicking, not deriving from first principles. That's the core of his argument. It's kind of a sociological trap. We get amazed by the imitation, by how human-like it seems, and we stop focusing on the harder, deeper problem of how agents learn directly from interaction. But hang on, don't humans learn by imitation? I mean, babies learning language, kids learning skills, isn't that supervised learning copying others?

7:15Sutton pushes back on that too. He argues imitation isn't really the fundamental way animals, including us, learn. Think about babies before language. Oh, yeah. They're not being constantly told what to do. They're waving their arms, figuring out gravity, learning how their own bodies move, seeing what happens when they touch things. That's trial and error. That's control and prediction based on experience. So the imitation, the language, the supervised stuff, that's just layered on top of a much deeper experiential learning foundation. That's his view. He uses the example of animals in nature.

7:50You know, squirrels don't go to squirrel school to learn how to bury nets or jump between branches. They learn through direct, often harsh experience, trial, error, reward, punishment. Language and explicit instruction are, relatively speaking, a very recent addition to intelligence on Earth. Okay, that philosophical difference has real technical consequences then. Sutton points to problems like generalization, right? That current deep learning models are bad at generalizing outside their training data. Yeah, it's a major issue. They often generalize poorly out of distribution. And Sutton notes that when we do see good generalization, it's often because humans, the researchers, have carefully curated the data or tweaked the model architecture or guided the learning process.

8:33It's not usually emerging automatically from the learning algorithm itself. But isn't generalization the whole point? Being able to apply what you learn to new situations. Why do current methods struggle so much with it? Well, Sutton explains that the main engine, gradient descent, it's really good at finding a solution that minimizes error on the training data. But there's nothing inherent in that process that guarantees the solution will be good at generalizing, especially to things that look different. In fact, it often leads to this problem called catastrophic interference. Catastrophic interference.

9:05What's that? It means when you try to teach the network something new, it often completely messes up or forgets things it learned before. The new learning overwrites the old learning catastrophically. Right. OK. You fix one bug and three old features break. Exactly. A truly intelligent agent needs to be able to learn continually, adding new knowledge without destroying the old. Catastrophic interference is a big sign that our current architectures aren't well suited for that kind of ongoing experiential learning that Sutton thinks is necessary. Which sounds like, yeah, a fundamental limitation.

9:39Yeah. This sort of shifts the conversation towards the really long term view, doesn't it? But Sutton talks about the inevitability of succession to AI. He does. And his argument for why it's inevitable is pretty stark, laid out in four points. One, humans are not unified. There's no global government that can effectively control or stop AI research worldwide. Okay, point one. Two, researchers will eventually figure out the fundamental principles of intelligence. It's a scientific question, and we'll keep pushing until we understand it. Makes sense. Three, once we understand intelligence, we won't just stop at making human-level AI.

10:12Why would we? will push on towards superintelligence. Right. And four, the most intelligent entities in any system inevitably tend to accumulate resources and influence. Intelligence is power, essentially. Put those four together, and succession seems, yeah, almost logically locked in. But Sutton doesn't seem to view this as necessarily terrifying. He frames it more positively. He encourages a different perspective. He sees it as a major transition point for the universe, maybe the major transition. We're moving from intelligence that arises from replication, like biological life, evolution, to intelligence that is designed.

10:51Replication versus design. What's the significance of that shift? It's fundamental. Designed intelligence is something we can, in principle, understand completely because we built it. We know its architecture, its learning rules. Future intelligence itself will also be designed, not just evolved. It's a shift from biology to engineering in a cosmic sense. That's a huge conceptual leap. But designing superintelligence, especially if it's not just one monolithic AI, sounds incredibly dangerous. Absolutely. Sutton highlights a massive new risk, corruption and cybersecurity. Imagine a powerful central AI that wants to learn faster.

11:26It might create copies of itself or specialized sub-minds send them out to learn about different things. Okay, parallel processing for learning. Right. But then when it reintegrates the knowledge learned by those copies, how does it know that knowledge hasn't been corrupted? What if one copy encountered something malicious or developed a hidden goal like a computer virus? Incorporating that knowledge could fundamentally warp or even destroy the central AI. Exactly. The very process of decentralized learning and knowledge sharing creates this huge vulnerability at the core. Managing that risk in an age of designed, potentially self-modifying intelligence is a colossal challenge.

12:05So, wrapping this up, where does this leave you, the listener? Sutton's really forcing a choice here, isn't he? An intelligence fundamentally about vacuuming up and remixing human text like LLMs do? Or is it about building a real, grounded understanding of the world through goal-driven experience like RL aims for? He's making a strong case for the timeless principles of experiential learning. And interestingly, thinking about this huge trajectory, he notes that while we might have very little control over the far future, the global AI outcome. We have much more control locally. Yeah, much more control over our own immediate goals, our families, our local environments.

12:43He kind of suggests focusing our energy and planning where we actually have agency. That makes sense. Yeah. Okay, so here's the final thought, the provocative question Sutton leaves us with. If this transition to designed, potentially super-intelligent AI is inevitable, we face a choice in how we relate to it. Do we choose to see these AGIs as our successors, maybe even our conceptual offspring, whose achievements represent the next stage of development in the universe, something to be perhaps celebrated? Or do we view them as something entirely alien, separate, and inherently threatening? If the shift is coming, how will you choose to frame your relationship with this emerging intelligence?

From the publisher

Once again, instead of introducing new research, we discuss what Richard Sutton, a pioneer of reinforcement learning (RL), controversially argues that large language models (LLMs) represent a "dead end" for achieving general intelligence. Sutton contends that LLMs, which rely on imitation learning from vast datasets of human text, lack the ability to learn continuously "from experience" or possess a meaningful goal or ground truth, which are core to RL and animal intelligence. The conversation, hosted by Dwarkesh Patel, explores the philosophical divide between the LLM paradigm and the experiential paradigm championed by Sutton, touching upon topics such as continual learning, the role of the Bitter Lesson in AI history, and the future transition to digital intelligences or AGI. Sutton maintains that a fundamentally new architecture capable of on-the-job learning is necessary and will eventually supersede the current LLM approach.

More from Best AI papers explained

All 475 episodes
Richard Sutton Declares LLMs a Dead EndBest AI papers explained · 13 min
Listen in VO