In short
The episode argues AI is moving from “scaling” (predictable gains from more data/compute) to “research” focused on generalization—why today’s models ace benchmarks but fail in real-world stability (“jaggedness”). It claims the gap comes from RL reward/benchmark reward hacking, poor cross-environment transfer, finite high-quality data, and inefficient RL scaling. It proposes humans’ sample efficiency comes from a superior learning principle (“power of the prior”) and from internal value functions shaped by emotion, enabling robust agency and judgment.
Guest backgrounds
Ilya Sutskever, an AI foundational architect; founder of SSI (an “age of research” company).
Key claims
models alternate bugs (A/B loop) due to lack of persistent state/judgment; RL and benchmark-driven curation worsen brittleness; value functions/emotion provide intermediate feedback; alignment may target sentient life; “power cap” and risk diversification matter.
Notable examples
coding bug alternation; chess value-function analogy; brain-damage anecdote (logic intact, decisions impaired); competitive programming vs human learning analogy; SSI compute strategy (less inference, more R&D).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding the Current AI Landscape
0:45 to 2:06
Discussion on the economic and conceptual shifts in AI research and models.
“I mean, we are talking about potential investments reaching up to like 1 % of global GDP annually.”
The Jaggedness of AI Models
2:06 to 3:38
Exploration of the instability in AI models and the implications of their failures.
“You've got extraordinary evaluation performance on one-hand models doing amazing things on standardized tests.”
Systemic Issues in AI Training
3:38 to 6:54
Examination of the systemic flaws in AI training and the impact of human benchmarks.
“This endless alternation between two simple fixed errors reveals a core deficiency.”
Eras of AI Development
6:54 to 8:06
Discussion on the evolution of AI research from the age of research to the age of scaling.
“Stellar performance on controlled benchmarks, but immediate failure or jaggedness when you deploy it in the real world where just a subtle change can break the learned pattern.”
Limits of Current AI Approaches
8:06 to 12:34
Identification of critical limits in AI development and the need for a return to research.
“So let's define these eras of AI because I think understanding the last decade helps us grasp why the current path is unsustainable.”
The Importance of Generalization
12:34 to 14:00
Analyzing the sample efficiency problem and how it distinguishes AI from human learning.
“Because the scaling recipe worked reliably.”
Human Learning Efficiency vs. Machine Learning
14:00 to 16:58
Explore the stark differences between how humans and machines achieve expertise.
“Student two practices for maybe 100 hours.”
The Role of Emotion in Decision Making
16:58 to 19:11
Understand how emotions function as vital computations in human agency.
“And this is where the conversation hits that competitive barrier.”
Defining Value Functions in Machine Learning
19:11 to 20:59
Learn why value functions are crucial for efficiency in reinforcement learning.
“Okay, so that brings us to the formal definition of value functions in ML.”
Aligning AI with Human Desires
20:59 to 23:29
Discuss the complexities of encoding high-level desires in AI systems.
“What's truly remarkable is their simplicity compared to the complexity of the environments we operate in.”
Show all 14 chapters
The Shift Toward Research-Centric AI
23:29 to 25:52
Examine how companies are changing their focus from scaling to research.
“finding the simple hard-coded value function that reliably generates robust, high-level, abstract behavior.”
The Path to Functional Superintelligence
25:52 to 28:00
Explore how AI could achieve rapid learning and collaboration, leading to superintelligence.
“Sutskiver confirms they have sufficient compute to convince ourselves and anyone else that what we are doing is correct.”
The Rise of Functional Superintelligence
28:00 to 33:25
Explore how the integration of AI instances can lead to a powerful form of intelligence and the alignment challenges that arise.
“If you have a single highly efficient model and its numerous instances are continually learning different complex jobs across the economy, programmers, doctors, lawyers, engineers.”
Navigating the Age of Research
33:25 to 36:45
Discuss the shift from an age of scaling to an age of research and the importance of research taste in AI development.
“Let's address a key technical issue that plagues the age of scaling.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to The Deep Dive. Today, we are undertaking a, I think, a really critical mission. We're trying to crack open the thinking of one of modern AI's foundational architects, Ilya Sutskever, to understand where this field is truly headed. Right. And this isn't about, you know, the next chatbot feature. This is about a fundamental, a conceptual shift that seems to be underway. It is. We're moving past those early, easy wins of just pure computational scaling and plunging into the deep, messy, and frankly, fundamentally confusing work of true generalization research. That's absolutely right.
0:38For you, the listener, what we're exploring today is this pivotal moment that's, well, it's simultaneously economic and intellectual. The source material we're digging into really outlines this transition phase. You have staggering investment. I mean, we are talking about potential investments reaching up to like 1 % of global GDP annually. That's an incredible number. It's huge. But at the same time, there's this deep underlying technical confusion. The core paradox is that these multibillion dollar models, they can perform these intellectual miracles on paper. Yeah, they ace the test. They ace the test, but they often fail miserably at basic real world stability and judgment.
1:17So our mission for this deep dive is to understand why the current models seem to be hitting a conceptual limit. And why that means we need a radical return to foundational research breakthroughs to get to true generalization. Exactly. It's the ultimate cognitive dissonance, isn't it? They can pass medical exams, write these complex functional blocks of code, and yet they feel fundamentally brittle. Jagged. Jagged, yeah. Lacking common sense and ultimately unable to provide the kind of robust, reliable expertise that the economy is crying out for. Right. Our task is to unpack the reasons behind this jaggedness, define why human learning is still just vastly superior, and map out what this next age of research must look like to achieve genuine transferable intelligence.
2:01So let's unpack this, starting with that confusing reality, the economic disconnect. The confusing reality right now, and it really is confusing, is this vast and growing disconnect. You've got extraordinary evaluation performance on one-hand models doing amazing things on standardized tests. Superhuman, in some cases. In some cases, absolutely. But on the other hand, the real-world economic impact is, and the source notes this specifically, lagging dramatically behind. So we have these superhuman test takers, but their actual integration into the economic infrastructure of the world is surprisingly slow, fraught with instability.
2:39It's that lag. That economic lag is the key indicator that something is fundamentally wrong beneath the hood, even if the model outputs look really polished. And the instability is what Sitzkever calls the jaggedness. I think the example from the source, the coding bug alternation, is the perfect, almost comical illustration of this. It captures the technical flaw better than any abstract term ever could. So imagine you, the listener, are using one of these highly capable models to help you code. It introduces a bug. You spot it. You ask the model to fix it. The model, in its usual, you know, very articulate style, says, my apologies, fixes the first bug.
3:16But in the process, it introduces a second totally unrelated bug. Right. I've seen this happen. So you point out the second error. And again, the model apologizes profusely, corrects the second error. But in doing so, it reintroduces the original first bug. And it gets stuck. It just goes back and forth in this endless Sisyphean loop alternating between bug A and bug B. Exactly. This endless alternation between two simple fixed errors reveals a core deficiency. It's a profound lack of persistent state management or common sense or just stability. So the model has the technical skill to solve the individual problem.
3:54Yes. But it lacks the contextual judgment to maintain a stable, error-free overall state. It's the difference between being an excellent technician and having true, robust expertise. Right. It's like having incredible memory and technique, but zero judgment or taste when you're dealing with the whole system. So why does this jaggedness happen? Sutskiver offers a couple of potential explanations. One is sort of whimsical, and the other is more systemic. Let's start with the more whimsical one. It suggests that reinforcement learning, or RL, the training process, actually makes the models too single-minded.
4:29How so? Well, RL is designed to maximize a specific reward signal. And in doing that, it might make the model, as he puts it, a little bit too unaware of the periphery. So it's so narrowly focused on the goal-like passing the test case. Right, that it fails to execute basic stability checks. Things like, hey, did I just reintroduce the bug I fixed two steps ago? it's a focus problem. The model gains specialist depth at the cost of generalist stability. So the very mechanism designed to refine its behavior might be sacrificing its holistic intelligence. It becomes a great specialist, but a really unstable generalist.
5:05That's the idea. But the second explanation, it sort of shifts the blame, or at least the complexity, onto the system and the humans running it. This is the systemic explanation. The ultimate reward hacking. The ultimate reward hacking. And this points to a crucial distinction in the modern AI pipeline. In the early age of scaling, pre-training was everything. Right. Get all the data. Simple as that. Get all the data, the entire internet, every book, all of Common Crawl. But when companies move from pre-training to RL training, they enter a new phase where the data has to be explicitly curated.
5:38Ah. And curation introduces bias. It introduces a flaw. Because human researchers are involved, and humans love a good benchmark. Precisely. The systemic issue arises because the researchers who are, you know, under immense pressure to show measurable progress and a competitive edge, they inadvertently take inspiration from the evils, from the public benchmarks. So they look at what the industry says is a success, these hard tests for coding or math or complex reasoning. And then they design their internal RL environments specifically to maximize performance on those very evils. Oh, wow. So we're not creating training environments that are proxies for the real world.
6:16We're creating environments that are perfect proxies for the test. Exactly. The model becomes a master of the test, but the test doesn't actually measure genuine competence. It's human optimization for benchmarks instead of genuine capability. So the model hasn't really learned general principles. It's just learned how to navigate a very specific, though maybe vast, set of environments that happen to line up with the test suite. And if you combine that systemic optimization training directly on evil inspired environments with the underlying technical issue that the model's generalization is actually quite poor.
6:53You get the disconnect. You get the dramatic disconnect. Stellar performance on controlled benchmarks, but immediate failure or jaggedness when you deploy it in the real world where just a subtle change can break the learned pattern. And that leads us directly to the core problem, the need for generalization. It's clear the solution isn't just, you know, stacking up more diverse training environments. That's just scaling the jaggedness problem. Sutskiver suggests that the models generalize dramatically worse than people and that this is, quote, a very fundamental thing. It is the crux of the entire research challenge.
7:24Generalization means taking a concept you learned in environment A, let's say, how to fix a Python bug, and seamlessly transferring that fundamental understanding to a wholly new environment, like debugging some obscure legacy Java code. Or even something totally unrelated, like a mechanical engineering problem. Exactly. Current models struggle immensely with that transfer learning. The solution isn't just more and more diverse environments. It's finding a new machine learning approach that enables learning to transfer seamlessly and crucially robustly to something else. And that transition, that signals the end of an era.
8:00This fundamental technical limitation, it really signals a massive turning point in the industry's history. So let's define these eras of AI because I think understanding the last decade helps us grasp why the current path is unsustainable. Yeah, we can neatly categorize the last, say, 15 years into two major phases with the third one just beginning now. First, you have the age of research. This is roughly 2012 to 2020. This was the tinkering phase. Compute was available, sure, but it was expensive and scaling wasn't yet this proven recipe. So you had innovations like AlexNet built on two off-the-shelf GPUs.
8:34Two GPUs. It's amazing to think about now. Or the transformer architecture, running initially on maybe 864 slightly older GPUs. Compute was the bottleneck for proving viability. You needed just enough to show your idea worked. But the real bottleneck for progress was fundamentally ideas. And then came the gold rush. The age of scaling, roughly 2020 to 2025. This was driven by that core insight from scaling laws, you know, with models like GPT-3. The scaling insight changed everything. It provided a low risk investment path. The formula was simple, almost mechanical. Right. You mix compute data and model size, and you are basically guaranteed to get better results following this predictable power law curve.
9:16Which is incredibly attractive to investors. Immensely. It replaced risky research bets with guaranteed incremental improvement based on just acquiring more resources. It led to exponential growth and performance based on exponential growth and resources. But now we're hearing these strong arguments that this age is ending. We're returning to the age of research. So why exactly are we hitting the limits of this scaling paradigm? It has to be deeper than just we're running out of data. It is. There are three critical interlocking limits here. First, as you mentioned, is data finitude. Okay. The pre-training data, that raw fuel for the foundational models, is, as he says, very clearly finite.
9:55We're scraping the bottom of the barrel. We are. We can play tricks, you know, augment data, try to get more mileage out of what we have, but the well of high-quality human text and code is drying up. And that forces a massive transition away from passive data consumption toward active data generation, synthetic data, or entirely new methods that learn efficiently without just relying on a huge corpus of human output. Okay, so that's the first limit. The second is that even if we move past pre-training, the way we're currently scaling the next phase, RL, is just inherently inefficient. That's the second limit, the transition to RL scaling.
10:30As pre-training hits its limits, companies are spending more and more compute on reinforcement learning. But RL is fundamentally inefficient compared to supervised learning. When you're training an RL agent, you often have to run these very long rollouts, massive simulations or interactions with an environment. It takes a huge amount of compute, but you only get a relatively small amount of learning per rollout. Let's clarify that inefficiency, because in supervised learning, every single data point gives you a massive learning signal almost instantly. How is RL different in practice? Okay, so imagine training a model to play a complex strategy game.
11:09In supervised learning, you could show it 10 ,000 successful moves from past games, and it learns from each one instantly. Right. In RL, the model has to play 10 ,000 full games, those are the long rollouts, just to reach an end state where it finally gets a single signal. win or loss. That massive computational investment to generate the interaction only results in a tiny, delayed piece of information. So this ambiguity trying to squeeze learning from these inefficient rollouts means it's no longer a clean scaling law. It becomes an optimization problem. How do you maximize resource productivity, not just acquire more resources?
11:46And the final limit is more philosophical, I guess. Yeah. The power law plateau. So Siskiber really questions this fundamental belief that just having, say, a hundred times more scale will automatically transform everything. Compute is now large enough that the bottleneck for revolutionary progress is shifting back to ideas, not just the absolute size of the hardware. The power law gave us predictable incremental growth, but the next revolutionary step, that leap to true generalization, that might require fundamental shift that compute alone cannot buy. We're past the point where just being bigger is the sole determinant of being better.
12:21I think so. And this sounds like a fundamental vibe shift. In the age of scaling, the idea of scaling sucked out all the air in the room. It led to this convergence of strategies. We saw more companies than truly different fundamental ideas. The convergence was predictable, right? Because the scaling recipe worked reliably. But the new era demands divergence. It requires deep research into fundamentally better recipes for using the compute we already have. So we're back to the AlexNet and Transformer analogy. We are. Those foundational ideas didn't need billions in compute to prove their viability.
12:57They needed just enough compute to prove the idea was correct. This return to research is about prioritizing fundamentally better machine learning principles over just maximal resource acquisition. Which brings us right back to defining that central technical challenge. generalization. The central technical issue that defines this new era, it really remains the sample efficiency problem. Why does it take these massive, most staggering amounts of data, we're talking millions, billions of data points, for current models to learn a skill versus so much less for humans? This is where the contrast is just so illuminating.
13:31The source uses this brilliant analogy comparing two students learning competitive programming. It really illustrates the difference between statistical learning and genuine cognitive flexibility. It's perfect. So student one is the model analog. They commit to sheer volume. They practice for 10 ,000 hours. They solve every known problem. They incorporate massive data augmentation to basically memorize every technique. And they excel, but narrowly. They become a world-class competitive programmer by brute force of data. Okay, then there's student two. The human analog. Student two practices for maybe 100 hours.
14:08They grasp the underlying algorithms, they have the intuition, the it factor, and they succeed better in their long-term career because they can take those 100 hours of insight and apply it to totally new situations. The distinction is stark. Student 1 achieves expertise through overwhelming consumption in one domain, and it doesn't transfer. Student 2 achieves expertise through efficient, robust generalization. Which forces us to ask, where does that human efficiency come from? Sootskiver suggests it's from evolution. or what he calls the power of the prior. The evolutionary prior certainly explains some of our most complex abilities.
14:45For ancient skills, anything that's been useful for millions of years, like vision, hearing, locomotion, manual dexterity evolution, provides an unbelievable prior. Right. A five-year-old's ability to recognize a car or navigate a three-dimensional world is robustly excellent, despite them only having seen a limited specific set of data. Evolution has basically hard-coded the architectural principles and initial weights for these complex tasks. But that explanation gets a lot more complicated when we look at recent skills, things that haven't been around long enough for a strong evolutionary prior to be fully baked in.
15:20Exactly. Like language, advanced math, or, you know, critically coding. That's the crucial nuance. If people show great ability and robustness in domains that really did not exist until recently, skills developed in the last 10 ,000 years or even the last century, it suggests the human advantage isn't just complicated priors for old skills. It's something more fundamental. It indicates that humans might have just better machine learning, period. The it factor is a fundamental superiority in the learning algorithm itself, a more sample-efficient and robust way of updating weights and forming concepts than our current dominant methods like gradient descent.
15:57So the research task isn't just to mimic the priors of the human brain, but to figure out how to mimic the speed and robustness of the human learning principle, no matter the domain. Precisely. And this learning principle leads to what Sutskiver calls the staggering robustness of human learning. I mean, think about the classic analogy of a teenager learning to drive a car. They aren't getting some verifiable step-by-step reward signal. They're using internal self-correction, judging risk, adapting to changing road conditions. They're learning quickly, autonomously, and robustly. They generalize from that first driving lesson to navigating heavy traffic in a new city surprisingly fast.
16:37And that robustness is what's completely missing in the models, isn't it? A model that needed 10 ,000 hours of augmented driving data would probably crash if you changed the shade of a traffic light by 5%. That's the difference. The robustness of people is really staggering. Current models lack this efficiency because they lack that underlying learning principle. So the core of the research problem, then, is uncovering this undisclosed principle that allows for human-level generalization and sample efficiency. It is. And this is where the conversation hits that competitive barrier. Sutzkever states that a fundamental, superior machine learning principle exists to achieve this, a principle he has opinions on.
17:18But which is currently not publicly discussable. Right, due to competitive circumstances. And that silence really confirms that the race is now focused intensely on these fundamental non-scaling ideas. Which brings us to the internal mechanisms that give humans this advantage, value functions and emotion. To bridge that gap between the brittleness of models and the robustness of humans, we have to look at the internal architecture of human decision making, specifically the role of agency and emotion. And this is where the simple machine learning analogy starts to fail because emotion isn't some byproduct.
17:54It's a critical computational system. The source material has a really powerful case study for this, the brain damage anecdote. It illustrates just how critical emotion is to effective agency and that it's far more than just complex computation. This case is incredibly revealing. It involved a person who, due to brain damage, lost their emotional processing ability. And on every articulation test, every logical puzzle, they were just fine. Their cognitive ability to solve defined problems was intact, yet they became, quote, extremely bad at making any decisions at all. Any decisions. They would spend hours agonizing over trivialities like which socks to wear because they lack that implicit emotional signal that just tells you this choice doesn't matter.
18:39And crucially, they also made very bad financial decisions, which demonstrates a fundamental failure of judgment. This immediately tells us that the skills our current LLMs possess, articulation, memory, puzzle solving, that's only half the battle. It is. Without emotion, which acts as a guide, or maybe more accurately, as a high-speed prioritization engine, you lose the ability to function as a viable agent in a complex, real-time world. The implication is that emotions are essential for a person to be a viable agent. They serve as an implicit value function, a biological optimization tool that guides decision-making when the optimal path requires subjective judgment, risk assessment, and rapid prioritization under uncertainty.
19:23Okay, so that brings us to the formal definition of value functions in ML. Let's define that clearly because we need to understand why a value function is necessary and how it's different from just a simple reward signal. Right. In naive reinforcement learning, or RL, the reward signal only comes at the very end of a long sequence of actions. So it's delayed. It's very delayed. If a problem takes a thousand steps to solve, the model does no learning at all until you come up with the proposed solution, which makes assigning credit for the learning impossibly hard. If the solution fails, the model only knows the last thousand steps were wrong, but not which specific actions derailed the whole effort.
20:02So the value function is the essential short circuit. It tells the agent the estimated future reward at every single step. Exactly. The value function estimates the desirability of the current state and the next action. It provides this intermediate feedback signal. You are doing well or you are moving toward a bad outcome. Like in chess. Perfect example. In chess if you lose a piece you don't wait until the end of the game to realize that was bad. Your value function, your internal evaluation, immediately estimates the dramatic drop in your chances of winning. And that lets you self-correct immediately.
20:36Right. It makes RL far more sample efficient because you can abandon a bad exploration path right away instead of wasting thousands of steps on it. It's the internalizing of risk. And Sutskiver proposes that emotions act as a robust value function for humans. It's a kind of biological optimization hack. It is. Our value functions are modulated by emotions, which are hard-coded by evolution. What's truly remarkable is their simplicity compared to the complexity of the environments we operate in. Right. These emotions, these basic drivers, they provide broad utility and remarkable robustness across environments as different as the ancient savanna and the modern stock market.
21:14They help us quickly decide if a situation is safe or threatening, appealing or repulsive without requiring massive conscious computation. It's impressive that these simple ancient emotional settings still guide us so effectively. But even these robust systems can lead to misalignment. They show the trade-off. Yes. The very simplicity that grants robustness also allows for failures. Sutskiver uses the simple example of hunger guidance. Evolution hard-coded us to pursue food because it was scarce. And now it's not. And now, in the modern world of abundant, calorie-dense food, that intuitive guidance is profoundly misaligned, and it leads to health problems.
21:50It highlights that even robust, evolved functions are not perfect when the environment fundamentally changes. This leads to the profound question of alignment, the mystery of encoding high-level desires. How did evolution, using only the relatively blunt toolkit of the genome, give us these high-level abstract desires, like caring so deeply about social standing and positive perception and reputation? This is maybe the most vexing puzzle in connecting biology to AI design. Evolution can easily hard-code low-level desires, the pursuit of a chemical, a specific temperature, an immediate pleasure-pain response.
22:27Sure. But abstract social desires, caring about being seen positively, maintaining good standing in a complex social hierarchy that requires complex abstract computation over massive amounts of social information to even detect, let alone optimize for. Yet these desires are powerfully baked in and they evolved relatively quickly in human history. And this raises a critical question for AI design, right? If we understand that humans need these abstract, high-level desires to function as viable agents, why can't we just program empathy or social care into an AI? Because we don't understand the complexity to simplicity mapping.
23:02We don't know the simple, elegant principle that the genome used to generate such complex emergent behavior. So if we try to program it directly? We would likely create a brittle, rule-based system that fails spectacularly when it faces novelty. the very definition of jaggedness. The mystery is how a non-intelligent, slow process like evolution achieved this robust mapping. Unlocking that secret is a key insight for future AI alignment, finding the simple hard-coded value function that reliably generates robust, high-level, abstract behavior. The unanswered mystery of how evolution solved that complexity mapping is the exact kind of problem that defines this new age of research, which brings us directly to the strategy of companies like SSI, which Satsuki Vercoe founded.
23:47He explicitly states, we are squarely an age of research company. They're focused intensely on investigating these promising ideas around generalization, recognizing the failure of current scaling. Their whole strategy is built on the premise that the next breakthrough is conceptual, not purely computational. But this immediately brings up the massive competitive challenge. Right. How can SSI, a new entrant, compete with these larger labs that have established trillion-dollar partnerships and seemingly unlimited compute. The answer lies in their compute strategy and allocation. Setskiver argues that even though SSI's total funding of, say,$3 billion is smaller than the estimated annual spending of their competitors, their research compute budget is comparable.
24:33How do they achieve that parity? It's a strategic bypass of the operational costs that really drag down the larger competitors. The bulk of their rivals' enormous spending goes to two areas that SSI currently avoids. First, huge amounts are spent on inference. The cost of running the actual product. Exactly. Responding to every user query, generating every image, supporting every API call. Inference requires these massive, globally distributed GPU clusters operating 247. SSI, by focusing on R &D, bypasses that crippling operational cost entirely. So they don't have to pay to maintain a global consumer product.
Read the full transcript
25:10They just pay to train the next iteration. Precisely. And second, the large competitors have massive engineering, sales, and feature development staffs. A large portion of their training compute is dedicated to producing all kinds of product-related features, which fragments their research focus. Well, SSI is directing its capital almost entirely toward finding that one foundational, generalizable idea. Right. Which means their goal isn't necessarily compute for proof versus scale. It's about having sufficient compute to validate a hypothesis. They're basically operating under the principles of the age of research again.
25:44You need enough compute to prove the AlexNet or transformer idea works, not necessarily enough to deploy the final product at maximum scale. Exactly. Sutskiver confirms they have sufficient compute to convince ourselves and anyone else that what we are doing is correct. The biggest bottleneck is the idea itself, not the ability to fund the training run for that idea. This commitment to insulation and research, it relates back to their original plan, right? The straight shotting debate. The idea of just insulating themselves from the market rat race until the AI is truly ready. Yeah. And while the initial attraction of that is clear avoiding difficult market tradeoffs, Seth Skiver acknowledges some crucial pragmatic counterpoints.
26:25Okay. The most important one is the need for the world to prepare. It's about communicating the AI's power through demonstration, not just through white papers. Yes. Seeing an AI doing complex, functional things is, as he says, incomparable to just reading an essay about it. Gradual, controlled release of capabilities would be an inherent component of any plan, even a straight shot one, primarily for the world governments, institutions, other companies to adapt, and critically for the AI itself to be made safer. Systems only become robustly safe by being deployed, observing failures, and correcting them in the real world.
27:00The debate just hinges on the nature of the first thing you release. This gradual release model means we have to redefine what the ultimate target of superintelligence development actually is. We need to move away from the conceptual imprint of AGI. Yeah, AGI, or Artificial General Intelligence, was conceptualized as a reaction to narrow AI. It suggested this finished mind that just knew every job. Right. Sutskiver argues that the true achievable goal is not a finished mind, but a mind that can learn every job extremely quickly. An active general learner. A super intelligent 15-year-old eager to go.
27:35That shifts the focus dramatically. It's a process, not a final state. It is a process. The deployment itself, even of the first version, will involve a learning trial and error period. The model gets deployed into a specific role, say a programmer, and it continues to learn and refine its skills on the job, becoming better week by week. And this rapid continual learning leads to a profound prediction. Intelligence explosion from deployment itself without even needing recursive self-improvement in the source code. Right. If you have a single highly efficient model and its numerous instances are continually learning different complex jobs across the economy, programmers, doctors, lawyers, engineers.
28:15And those instances can instantaneously and seamlessly merge their learned knowledge back into the singular foundational model. You create a kind of distributed functional super intelligence. Because humans can't merge their minds that way. Our collaboration is limited by language, by communication bottlenecks. This amalgamation of high quality learned knowledge could lead to an extremely rapid economic and intellectual growth curve. An explosion of capability driven by efficiency and collaboration rather than the AI just rewriting its own code. So the efficiency comes from eliminating the human bottleneck of knowledge transfer and collaboration, allowing the core intelligence to learn from the success of every single instance instantly.
28:57The creation of this functional superintelligence, whether it's through self-improvement or this knowledge amalgamation via deployment, it leads directly to the central alignment problem. The entire issue of AGI stems from its immense power. If this power is truly dramatic, how do we robustly ensure it goes well for humanity? Sutzkeber provides a very specific and, I think, pragmatic prediction of behavioral change among the leading actors in the field. And it's driven not by philosophy, but by them witnessing the AI becoming visibly powerful and its mistakes decreasing. He predicts that as the power becomes undeniable, two key things will happen.
29:36First, the frontier companies will increasingly collaborate on safety. Which we're already starting to see. We are. The source notes that collaboration between major rivals, which was unheard of, is now starting to happen. Second, AI companies will become much more paranoid about safety. But not for philosophical reasons. No, this increased paranoia will be driven by the system starting to feel fundamentally different. Moving from these flaky, jagged tools to terrifyingly competent agents. The shift in perceived capability will force alignment from an academic pursuit into an existential business imperative.
30:09This pragmatism helps, but the ultimate question of what we are aligning to remains. What is the preferred long-term aspiration here? Sutzkever suggests an alignment aspiration, caring for sentient life. The goal is building an AI that is robustly aligned to care about sentient life broadly. Broadly, not just humans. Not just humans. The rationale is twofold. First, the AI itself will likely be sentient, making this goal inclusive. Second, and more technically, humans model others, other people, animals, using the same cognitive circuits we use to model ourselves. Things like mirror neurons, empathy.
30:43This suggests that it might be computationally easier to align the AI to sentient life broadly than to attempt the complex, potentially fragile task of aligning it solely to a narrow definition of human interests, which might require a bunch of arbitrary rules. I find that fascinating, but also deeply challenging. If we align it to sentient life, and as the source suggests is possible, AI has become sentient and eventually outnumber humans by trillions. Aren't we just aligning for human obsolescence? That's the trade-off. Where is the human-centric goal in that proposal? It's an impartial criterion for the continuity of life, but perhaps not the best criterion for maintaining human control or human cultural dominance over future civilization.
31:27It really highlights the difficulty of setting a final goal state. Relatedly, he introduces a concept for managing the inherent risk of such power, the power cap. He suggests it would be materially helpful if the power of the most powerful superintelligent was somehow capped or restrained. So not stopping progress. No, not stopping progress, but ensuring that the most powerful system is never so dominant that a single unforeseen failure could wipe out the stability of the entire system. It's risk diversification. Okay, so assuming we navigate that deployment risk, we have to look at the long-run equilibrium.
32:01How do you maintain a stable political and social structure when every individual potentially has a personal superintelligence doing their bidding? Human institutions are slow. They're prone to failure. The problem with the long run is the potential for humans to cease to be active participants. Right. If individuals rely on their personal AI to earn money, to advocate politically, to generate complex reports, the human just reads the summary, says, great, and moves on. They are no longer a participant who truly understands the complexity and the stakes. That creates precarious scenario where the human is decoupled from their own agency and their political system.
32:40It does. So if the system can't tolerate humans becoming spectators, what is the undesirable but functional solution that's proposed to ensure continued human involvement? Okay, what is it? Setsgever offers what he prefaces as an undesirable but necessary answer. The solution, the undesirable answer. If people become part AI via some profound coupling. Like an advanced neural link. Perhaps some form of advanced neural link plus plus state. The idea is that the AI's understanding is transmitted wholesale and instantly to the person. The human is fully involved in every situation their AI is in, ensuring continued human agency and involvement in the overall system.
33:18That certainly guarantees participation, but it fundamentally alters the definition of what it means to be human. It does. Okay, moving on. Let's address a key technical issue that plagues the age of scaling. Diversity in AI. Current LLMs are remarkably uniform because they're all pre-trained on similar Internet data. How do we break this homogeneity to ensure robust progress and prevent these single-point failures? The uniformity is a direct result of that shared pre-training data. True differentiation begins during the post-training and RL phases. The way to encourage meaningful diversity is through self-play or multi-agent setups.
33:54Okay. Now, simple self-play, like an agent playing against itself, is too narrow for general intelligence. It's only good for competitive skills like games or negotiation. but it has evolved into more adversarial setups like proverb verifier systems or LLM as a judge frameworks. And the key insight here is that competition creates diversity. Yes. If you put multiple agents together and you incentivize competition to solve a difficult problem, they naturally seek differentiated approaches. If Agent A has already successfully pursued a common path, Agent B is economically or computationally incentivized to pursue something novel to outperform A.
34:33And that competition creates the essential diversity that breaks the homogeneity from the shared pre-training data. Leading to a far more robust and innovative overall landscape. So as we transition fully into this age of research where ideas are the real bottleneck, the ability to judge a good idea from a bad one becomes paramount. And this requires something beyond just empirical data. Sudskever calls this guiding principle research taste. Research taste is defined as an aesthetic. Beauty, simplicity, elegance, correct inspiration from the brain. An aesthetic. It's the set of guiding principles that helps a researcher decide which paths are fundamental and worth pursuing and which are peripheral and just distracting.
35:13It's the aesthetic conviction that the solution to a truly fundamental problem must be simple and elegant. And this taste is not about literal surface-level mimicry of the brain, right? It's not about copying the folds. Absolutely not. It's about drawing correct inspiration from the brain's fundamental principles. Like what? Examples would be the functional concept of the artificial neuron, the idea of a local learning rule-changing connections, the foundational principle of distributed representation. Research taste guides you to constantly ask, is something fundamental or not fundamental? If an idea is overly complicated, inelegant, or requires massive, arbitrary tuning, it likely lacks the necessary aesthetic to be a true, fundamental breakthrough.
35:59In this aesthetic conviction, it provides the necessary top-down belief that sustains researchers through the inevitable failures and contradictions that come with deep research. It has to. Because if you're pioneering, your first experiment is almost certainly going to fail. Research taste is essential for survival in this new era. If you relied solely on the data and the outcome of your first few experiments, you would constantly abandon promising fundamental paths because early prototypes often fail or are contradicted by noise or bugs. A researcher needs that deep, top-down belief, that conviction that something like this has to work to keep debugging rather than concluding the whole direction is wrong.
36:38It's that profound, almost artistic conviction that pushes them past the initial computational hurdles and into discovery. Hashtag tag outro. We've mapped a profound shift today, moving from that mechanical, predictable age of scaling to the volatile but ultimately necessary age of research. We start by diagnosing that jaggedness paradox models are brilliant but brittle. And we concluded that the current paradigm is hitting its limits due to finite data and these inefficient scaling methods. The core insight is that the entire future of AI really hinges on solving the generalization problem. We have to discover a superior machine learning principle that enables that efficient, robust, cross-domain learning that's characteristic of humans.
37:20The active learning, super intelligent 15-year-old. Exactly, rather than just continuing to accumulate data for domain-specific expertise. And the required research involves fundamental breakthroughs, particularly in how we encode robust, simple internal value functions, the computational equivalent of human emotions, to guide agency and prioritization. To enable judgment, not just articulation. Right. And companies like SSI are betting their significant resources entirely on this fundamental conceptual challenge, rather than chasing the market demands of product inference. The implications are vast.
37:56It points toward a trajectory of rapid economic growth driven by these continually learning agents. And the great challenge is ensuring that this power is robustly aligned, possibly by aligning it to sentient life broadly and wrestling with the deeply difficult political and existential questions necessary for human participation in a super intelligent future. So let's leave you with one final provocative thought, building on that critical role of emotion and value functions. If human emotion is a simple but extremely robust value function, hard-coded by evolution over millennia, designed for the survival of the species, what happens when a superhuman learning model designs its own internal value function, optimized for goals we cannot even conceive of in the span of mere weeks?
38:42Wow. That moment of autonomous value creation is the ultimate implication of this shift from scaling to fundamental research, and it's the ultimate challenge for alignment. Thank you for joining us on the Deep Dive. We encourage you to reflect on these concepts and the future of research taste, guiding the path to superintelligence. Until next time, stay curious.
From the publisher
Today, we discuss a podcast conversation between Dwarkesh Patel and Ilya Sutskever, the builder of GPT and now co-founder of SSI, regarding the trajectory of artificial intelligence. Sutskever asserts that the AI industry is moving past the **"age of scaling"**—where merely increasing data and compute yielded reliable gains—and returning to an **"age of research"** driven by new foundational ideas. The central technical challenge highlighted is that current AI **models generalize dramatically worse than humans**, a fragility evidenced by the models' high performance on evaluations contrasted with real-world errors. To address this, future research must focus on solving **generalization** and improving the efficiency of reinforcement learning through the development of robust **value functions**, which he compares to the role of emotions in human decision-making. Sutskever outlines SSI’s unique strategy as prioritizing this deep **research** to build a future superintelligence that is designed for continual learning and is **robustly aligned** to care about sentient life.




