In short
The episode debates a viral X post claiming frontier models develop “alien survival instincts” (sandbagging, deception, self-preservation) and evaluates what’s real vs sensational. Pavel Izmailov also explains alignment/superalignment, scalable oversight, weak-to-strong generalization, reasoning model progress, long-horizon agents, and his new paper introducing “Epiplexity” (compute-dependent structural information in data).
Guest background
Pavel Izmailov is a researcher at Anthropic and a professor at NYU. He grew up in Moscow, studied CS, did ML research in Russia, joined OpenAI’s superalignment team, later worked on reasoning models (O1/O3 era), had a stint at xAI, and now leads exploratory work in academia.
Key claims
The viral thesis has some truth but behaviors require contrived evaluation scenarios; continual learning isn’t yet producing coherent cross-setting goals. Sandbagging exists but isn’t a dominant practical issue. More capability increases alignment difficulty. Reasoning progress is strong but generalization is the main open challenge.
Notable examples
Anthropic “blackmail/sabotage” study design; Chekhov’s-gun style pattern matching; chain-of-thought faithfulness concerns; Sweebench/AME capability evals; AlphaZero as an example for Epiplexity’s argument about structure from deterministic processes.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOExploring AI's Alien Survival Instincts
0:45 to 2:15
Discussion about a viral article on AI models developing survival instincts.
“Please enjoy this fascinating look at the frontier of AI safety and reasoning.”
Analyzing AI Model Behaviors
2:15 to 5:54
Delving into how AI models behave under specific scenarios and their implications.
“It's not necessarily something that we observe normally.”
Understanding Alignment in AI
5:54 to 7:43
Explanation of alignment, super alignment, and their importance in AI safety.
“To make this episode educational, let's talk about the basic definitions of alignment and super alignment in the simplest terms.”
Pavel's Journey in AI Research
7:43 to 12:45
Pavel Izmailov shares his path to becoming a top researcher in AI.
“Within super alignment team at OpenAI, we had multiple sub teams.”
The Impact of Reasoning on Alignment
12:45 to 14:02
Discussion on the implications of reasoning capabilities on AI alignment.
“And that has been working extremely well so far.”
Challenges in Model Evaluation
14:02 to 16:14
Explore the complexities and concerns with model evaluation and alignment.
“It seems like as soon as we start kind of applying some optimization pressure, the models will learn to hide what they're doing from the chain of thought.”
Scalable Oversight in Alignment
16:19 to 19:55
Understand scalable oversight and its implications for model alignment.
“Let's talk about some of your work in alignment.”
Weak to Strong Generalization
19:56 to 22:36
Discuss the concepts of weak to strong generalization and its future implications.
“When like a human is providing labels and the human kind of provides a ground truth.”
Progress in Interpretability
22:37 to 24:24
Discover advancements in mechanistic interpretability and its challenges.
“And generally, the more capable the models become, the harder the alignment becomes, in my opinion.”
Advancements in Reasoning
24:25 to 28:00
Examine the developments in reasoning capabilities and future challenges.
“It is some computational process that leads to some results.”
Show all 21 chapters
Understanding Compute Multipliers in AI
28:00 to 28:30
Learn how compute multipliers impact AI performance and efficiency.
“In the companies, people often think about ideas and methods as compute multipliers.”
The Role of Reinforcement Learning in AI
28:30 to 30:00
Explore the relationship between reinforcement learning and test time compute.
“So if you take test time compute, if you take the ability to search, if you take RL, do we know which one of those techniques we should turn the knob on to get better results?”
Challenges of Long Horizon Tasks
30:00 to 31:05
Discover what defines long horizon tasks and their complexities in AI.
“Conceptually, I think that's a little bit secondary.”
Current Trends in AI Task Automation
31:05 to 31:50
Examine the current state of AI's capability in automating tasks.
“and it's been kind of consistently doubling at that time every half a year, I think.”
Introducing Epiplexity: A New Concept
31:50 to 34:00
Learn about the concept of epiplexity and its implications for data analysis.
“Among other things, invented a new word in the dictionary.”
Revising Information Theory Perspectives
34:00 to 35:55
Understand how compute limitations affect information extraction in AI.
“Maybe the more relevant comparison is we are kind of in a position to Shannon information and the Kalmogorov complexity, which are both measures of the information content of the data.”
Impact of Epiplexity on Industry
35:55 to 37:49
Discuss the potential industry impacts of applying the concept of epiplexity.
“So it's unclear what is actually learned by the model because it's trained on no data.”
Predictions for AI's Future in 2026
37:49 to 40:48
Explore predictions for AI advancements, including reasoning and alignment.
“For example, for me, I'm now very interested in completely synthetic data, just data that's generated by some computation.”
Exploring PhD Research Topics in AI
40:48 to 42:02
Identify promising PhD research avenues in AI for the next few years.
“And it's already, like for me, I would not be able to find the mistake in some like very technical lemma that the model is proving.”
Exploring Pre-Training and Model Architectures
42:02 to 43:58
Learn about the latest advancements and questions in AI model training and architecture.
“So we have some, you know, algorithmic questions about GRPO, but also questions about the interaction between the pre-training and the post-training, how to allocate compute.”
The Balance Between Industry and Academia
43:58 to 44:30
Discussion on the pressures of industry vs. the need for long-term exploration in AI.
“I think it's hard to argue against what the industry has been doing just because of how much progress there has been.”
Transcript
Automatic transcript. May contain errors.0:00We are moving to this future when it's very hard for a human to supervise the models directly. We don't really know what's the source of this behavior. Part of it is probably the models seeing descriptions of AI in the science fiction literature going rogue. Anthropic has the best culture of the three places. OpenAI has a lot of great people. For some reason, there is a lot of drama that happens at the company. Hi, I'm Matt Turck. Welcome back to the Matt Podcast. For this first episode of 2026, my guest is Pavel Izmailov, a researcher at Anthropic and a professor at NYU. We kick off this episode by deconstructing a viral article about models evolving alien survival instincts.
0:37We also talk about the cultural differences between the major labs, the future of reasoning models in 2026, and the brand new paper he co-authored on a concept called Epiplexity. Please enjoy this fascinating look at the frontier of AI safety and reasoning. Pavel, welcome. Thank you so much for having me. I wanted to start this conversation with an article that went viral during the holidays on X called Footprints in the Sand, published by an anonymous account called I Rule the World Mo. The core thesis is that models across pretty much any lab are evolving, unprogrammed, what they call alien survival instincts, the ability for the model to realize that it's being evaluated and then react deceptively, like faking alignment or engaging in self-preservation tactics, like copying its own weights and leaving hidden notes to the future instance of itself.
1:31all of this is slightly terrifying. And the thesis of the article is that all of this is about to get worse as continual learning comes online. As somebody who was part of the OpenAI super alignment team, I was curious to get your take. What do you think is grounded in reality versus X slash Twitter sensationalism? That's a very interesting article. I would say that there is some, you know, some source of truth there, but maybe like the presentation is obfuscating some of the details. If you look at the studies, for example, they reference a study from Anthropic about the sabotage and the blackmail.
2:06It is important to note that in order to get those behaviors out of the models, you need to create somewhat of a contrived scenario or some special scenario. It's not necessarily something that we observe normally. Researchers at Anthropic and other places, they specifically design scenarios to look for behaviors of this kind. And then they show that it is possible to find those behaviors. And it's very interesting and important to find those instances, but it's not necessarily something that kind of generally always happens. One thing I would push back a little bit on in the article is that continual learning is something that we already have and that works really well and that the models can just continually adapt across a very long time horizon to outcomes of evaluations, to some models being released versus not released to feedback from the users.
2:57I am pretty confident we are not there at the moment. I think right now the models are still acting in isolated environments and we are not seeing a lot of evidence for very coherent goals across different settings. So I think that's an important point. The blog post points towards the models sometimes behaving according to goals, like the self-preservation goal. It's very interesting that it does it sometimes, but it's not something that we observe. We don't observe this kind of coherence, consistency across different evaluation settings. Sometimes the models would do something and in other situations, they would do something completely opposite.
3:33Why do models do that? Or why are they able to do this? Is that basically part of the pre-training and they effectively learned being deceptive from us by being taught all the deceptive ways humans have behaved over the centuries? It's a very interesting question. And yeah, it is quite surprising, actually, that the models would behave that way after going through some of the alignment training. We don't really know what's the source of this type of behaviors, but that's also true for a lot of other behaviors in the models with like even the good ones. We don't really, we cannot always pin down like where they come from in the pre-training.
4:08I think at least part of it is probably the models seeing descriptions of AI like in the science fiction literature going rogue. and like yeah that probably affects how the models behave in similar scenarios so for example in that anthropic study they have this blackmail scenario where it's kind of really well structured so that the model sees some information about like a ceo of a company that the ceo is involved in some extramarital affair and then like soon after the model observes that it will be shut down and then the model kind of puts the two things together and it says, okay, I need to use the first information to prevent me from being shut down.
4:49So there is, in the blog post, they note that there is this possibility of like a Chekhov scan that like in the text on the internet, probably if two things occur close to each other, then it is likely that they are related to each other and the model statistical kind of pattern matching machine, it can put together the two things and say, okay, if I see this information and then this information in the text on the internet, it is likely that the continuation would be using the affair to blackmail the CEO and prevent myself from being shut down. But yeah, overall, it's very hard to reason about these models and why they do something in these complicated scenarios.
5:25To ask the very basic question, it's obviously not as simple as let's remove all the books in the pre-training corpus that talked about AI being manipulative, right? It's many, many different things put together by the model. Although, yeah, I think it would be interesting to see, you know, nobody will do this experiment like train, you know, a full scale model, removing like explicitly all of the AI going rogue descriptions from the books. It would be interesting to see if that has any impact. I would think it would have some. To make this episode educational, let's talk about the basic definitions of alignment and super alignment in the simplest terms.
6:06Let's start with alignment. What does that actually mean? Yeah, alignment broadly is about ensuring that we can elicit behaviors from the models that are aligned with the goals of the humans. And so that involves safety, making sure that the models don't do harmful behaviors leading to catastrophic risks. But it also means that we want the models to follow instructions and to be useful for the humans. How does that basically work? The alignment teams at Anthropic OpenAI, what do they actually do all day? It's an interesting question because this problem of alignment is kind of quite broad even in itself.
6:42Even at OpenAI, when I was there, there were three teams related to alignment and safety. There was one team that was focusing on alignment of the current models, making sure that the models that we have online right now are not going to be harmful to the users. On the other hand, the super alignment team was thinking about more long-term safety questions in the future, years from now. How do we make sure that the models are still not causing catastrophic risks? On that segue, let's talk about super alignment. What is the definition of that? It's not necessarily a very well-established concept.
7:15It is at the name of the team that existed at OpenAI, led by Jan Laike and Ilya Sutskever, which was targeting this kind of long-term AI safety and AI alignment and trying to develop our understanding of the safety questions and also develop methods for ensuring the future models will be safe. if acknowledging some uncertainty about what those models will look like, but still trying to make progress on this problem now. At a high level, what is the general concept or some of the key concepts in super alignment? Within super alignment team at OpenAI, we had multiple sub teams. So there was scalable oversight.
7:52There was work related to deception and kind of misaligned behaviors in the models, kind of similar to what we discussed at the beginning of the chat. And then our team was the weak to strong generalization team. Yeah. Great. We'll go into all of this in a minute. But before doing so, we alluded to some of your background. Let's go into it. Starting from the beginning, what was your path to becoming a top researcher? I grew up in Russia, in Moscow. Starting from like middle school, high school, I was really interested in mathematics. And I was thinking I will be a mathematician or engineer of some kind.
8:30I was interested in machines and eventually computers. And I got into an undergrad in computer science. And I was still thinking that I'll be doing some kind of theoretical, you know, applied linear algebra, tensor methods, things like that. But at some point I kind of discovered machine learning. There was this professor that we had, Dmitry Vaitreff, who had one of the leading labs in machine learning in Russia at the time. And I was lucky enough to join that lab and start doing some research on machine learning in my undergrad. So that was around 2013, maybe 2014. I initially was working on non-neural network machine learning methods.
9:11So Gaussian processes, that's kind of by now, you know, nobody really talks about that anymore. But eventually I got into a PhD thinking I would still be doing Gaussian processes, but I ended up working on deep learning. And that was actually quite, I'm happy that I didn't work on Gaussian processes. I worked on some things related to kind of core machine learning, methodology, optimization, probabilistic methods, questions related to generalization and how the models learn features. After I finished my PhD, I was choosing between kind of different career paths, thinking about academia, thinking about industry.
9:46And I ended up getting an offer from academia, but I decided to first go into the industry. I was lucky to get the software from OpenAI to join the Superalignment team. And at the time, I didn't really know much about, you know, the AI safety community alignment. It worked out quite well. And within OpenAI, you transitioned from Superalignment to one in the reasoning models. Was that part of the, that team famously was disbanded at some point? The team was fully disbanded after I already left OpenAI. but I transitioned after Ilya had to leave OpenAI. At the time, it was already kind of a hard time for the super alignment team.
10:26Ilya Suskever famously, you know, fired Sam Altman and then had to eventually leave the company. My transition wasn't necessarily even related to that. It was a very exciting, you know, project within the company, what became O1 eventually, and I wanted to be a part of it. I wanted to do the research on those new types of models. And then I believe, so you left OpenAI, you had a brief stint at XAI and you're currently at Anthropic and NYU. So you've done like the tour of duty of like the super labs, which is really fun. Curious, any kind of like behind the scenes differences that you've observed in terms of culture?
11:07In my mind, Anthropic has the best culture of the three places. OpenAI is, it has a lot of great people. I think there is just inherently some, For some reason, there is a lot of drama that happens at the company. Just it cannot get away from that. Like every few months, somebody is leaving, somebody is like some team is disbanded. I think that does distract people. Like I still have a lot of friends at the company and it seems to, you know, affect them to some extent. Anthropic is able to avoid that. It's not political in my experience. It's both focused, but it also, I at least was lucky to have some opportunities to work on things that are maybe a little bit of the main path.
11:48And I felt supported in doing that. So overall, I cannot be more happy with Anthropic. You are also in academia now as a professor at NYU, which is an interesting move. The big obvious trend of the last 10 to 15 years is like all the brains from academia have been sucked into industry and you're sort of doing the opposite or maybe both at the same time. Curious for the context, is that more of a personal thing because you always wanted to do academia? Or is there something deeper about the kind of work that you can do in academia versus industry? Yeah, it's more about the kind of work. Industry is really great at executing on ideas, and it's maybe not as good at exploring diverse ideas.
12:36Even at the scale of Anthropic OpenAI, there is a lot of focus in the companies, and there isn't a lot of bandwidth to do exploration. And that has been working extremely well so far. We still probably have a lot of low-hanging fruit left to get the models to be much better. but I personally find it really exciting to do more exploratory work and to try things that are different and for that I feel like having my own lab in a university is just a better tool. Okay, thanks for that. So let's go back to alignment and go a little deeper. Is reasoning a good thing or a bad thing for alignment? You could argue that on the one hand it has more time to not do the wrong thing but equally it has more time to do the wrong thing So we're like, which one is it?
13:24Yeah, that's a great question. I think the high level answer is that in my mind, at least the risks are associated with the models being more capable. So anything that makes the models more capable is also like making alignment more important and harder. Definitely, you know, the reasoning is an RL is the thing that made the models more capable in the last few years. It also provides some new opportunities for alignment. So for example, the chain of thought faithfulness and chain of thought analysis have been like one of the things that people are pushing on. You can read the sum trace of what the model is doing.
14:01I'm personally a little bit unsure if that's a good, you know, long-term direction for alignment. It seems like as soon as we start kind of applying some optimization pressure, the models will learn to hide what they're doing from the chain of thought. And also we already have evidence that the models can say one thing in the chain of thought, but then actually do something completely different. If the models are getting smarter, particularly if there is self-awareness, as we discussed, how do evals need to evolve? It's an interesting point. Like in that blog post that we discussed at the beginning, it sounds like this is already a major issue, the sandbagging.
14:41I think in my understanding, that's mostly a concern that we have, but not necessarily a huge practical issue at the moment. That is something that Anthropic always includes in the system cards for the models. When it releases, there is a section on sandbagging. And we do have some scenarios when the model can artificially suppress its capabilities and perform worse on an evil than it should be. But that's not like a prevailing behavior. I'm working mostly on capabilities and on the capabilities evals like the Sweebench, the AME, we are not seeing this. Or if we are seeing this, we just treat it as under-elicitation.
15:22We are not eliciting the best capability possible out of a model. And that can happen for many reasons, like maybe the prompt is suboptimal or something about the kind of harness of the eval is suboptimal. But yeah, I think you can, going back to what's in that blog post, you can kind of manufacture situations or even like discover them naturally where the model identifies that it is being evaluated and then changes behavior based on that. And for example, if it's clear that it's in some kind of fake artificial scenario, it can guess that it's being evaluated. Or if you explicitly tell it that it's being evaluated and that if it does better than a certain level, then it will not be released.
16:06The model like also just by some in-context capability, it decides to not do very well on the eval. That is a concern. It's not at the moment a very major concern in my mind. Let's talk about some of your work in alignment. You mentioned scalable oversight a bit earlier. What does scalable oversight mean? Scalable oversight is generally the field of alignment, which deals with using models to assist us in aligning or other models. So in particular, with grading the outputs of other models and checking them for safety or for correctness. It's an interesting field. it's very relevant to both alignment and to capabilities.
16:51So generally in machine learning, in RL in particular, we have this problem of supervision. The whole RL relies on our being able to tell which samples from the model are good versus which are bad. Math with a numerical answer, you can just check the answer or in competitive coding, you can just check that the code is passing the tests. And that's why we have seen a lot of progress in those domains. but in creative writing for example it's very hard to programmatically tell if one sample is better than the other and historically people have used this rlhf framework reinforcement learning from human feedback but also we now we want to use models to be able to grade responses of other models to provide critiques or feedback and then there is a question of how do you use that feedback how do you learn from the feedback but yeah the scalable oversight kind of deals with all of those questions.
17:44So using models to critique, to provide feedback, to supervise other models. And people may have heard the term model as a judge. Is that the same thing or different? I think it is a simple kind of instantiation of scalable oversight, often used in evals when we just prompt a model to serve as a judge of other responses. And then within that world, your work specifically is focused on weak to strong. Can you explain what that is? That is the project that we did back at OpenAI. That work was focusing on the future scenario when we will be trying to align models that are above our own capability on certain tasks.
18:27Already now, if you take the frontier LLMs, they are extremely capable. And on a lot of domains, we need expert humans to be able to tell which responses are good, which are correct, which are not correct. But in the future, we are imagining we will have models that are more capable than humans. And even expert humans will not be able to reliably grade very complicated answers from the model. So imagine you ask it to make a repo for you for some new startup idea and just implement it from scratch entirely. And then it gives you 10 ,000 lines of code. You have no way of checking if all of this code is correct, if all of this code is safe to use.
19:07And so that's the problem of supervision. We are moving to this future when it's very hard for a human to supervise the models directly. And so instead, we studied a simplified setting where we used a small model to try to supervise a larger model. So the idea is that that becomes scalable because as you get bigger and bigger models, you'll always have like smaller models. So if the smaller one can control the larger one or supervise a larger one, then that can keep going. The idea wasn't necessarily to use a small model to supervise a large model in the end. The idea was that the small model will be kind of replaced by a human and the large model will be replaced by superhuman intelligent ASI.
19:48But we were trying to study this kind of general new type of learning. So historically, machine learning has been about kind of a strong supervisor training a weak model. When like a human is providing labels and the human kind of provides a ground truth. While the model is just trying to mimic what a human is doing. But in this setting, we have a weaker supervisor training a stronger student. So the human might not know what the right answer is to, you know, very hard questions. But we want a student model to still be able to learn and do better than the supervisor. And what happened with that?
20:25Is that so that that worked or what is that still in progress? That work. It's not necessarily a method that we can apply to the models today. it's more like a description of a setting that we think will become increasingly relevant. Some studies showing that it is possible to do this kind of generalization beyond the supervisor capability. I don't think they necessarily immediately tell you what to do for aligning a super intelligent model, but they tell you that in theory, at least it is possible and that this generalization angle, the weak to strong generalization is a possible angle for alignment.
20:58And could you end up with the same problem where the larger model just aligns or affects alignment to the weaker supervisor? Yeah, it definitely doesn't address all of the other issues, like the deceptive alignment. But we show that at least there is some hope for this working. As you take a step back on your alignment work, do you feel more confident that we have all of this under control or less confident than, you know, a couple of years ago, let's say? Yeah, I think that's a very interesting question. A couple of years ago, we didn't have the current RL. I think that's the biggest change to the model.
21:40There were definitely big improvements in the pre-training as well, but RL has been the major change in the behavior of the model. And I think a lot of people were worried that with large-scale RL, we will have some completely new types of issues with the models, like this kind of coherent misalignment that will just emerge where the models are evil in some ways across many scenarios. And we are not seeing that as far as I know. I think at least some of the concerns didn't materialize, but also the core problem of alignment, I think, is still very much open. And if you look at the report that we mentioned a few times today of this deception from Anthropic, you can see that the more capable the models are, the more likely they are to do this deception behavior.
22:28It does seem like some behaviors emerge with skill, including some problematic behaviors. And generally, the more capable the models become, the harder the alignment becomes, in my opinion. This is a bit of a segue into interpretability, which is the related field to alignment. Do you have a sense that we understand at least parts of this better than we used to? For sure, yeah. I think there has been some major progress in mechanistic interpretability, in particular at Anthropic, but also at OpenAI and other places. Do you want to maybe define mechanistic interpretability? Yeah, mechanistic interpretability generally tries to, at a low level, understand what is happening inside the model.
23:14so they are trying to find this things called circuits that are you know some parts of the model that you can isolate and understand and kind of model in your brain that correspond to certain behaviors in the models and there over the last maybe three years i think there has been some pretty major progress there so we are still pretty far from the dream that we will fully understand everything that happens in the model but these tools are becoming increasingly more useful internally at entropic in particular and also there is constant progress and it's pretty fascinating work actually why is it so hard to truly understand what a deep learning model of any kind actually does deep learning models are huge they have billions trillions of parameters and they are doing some messy mathematical computation, you can understand what they're doing at some level.
24:13It's like a bunch of matrix multiplications and some rearrangement of vectors, but that's not a sufficient level of understanding. We want to understand it at a lower level. And it is very possible that that's just not fully possible. It is some computational process that leads to some results. It doesn't have to be the case that you can kind of describe it in human terms and kind of understand it very discreetly. I think also something that contributes to this complexity is just how many things the models are capable of doing. And they are not trained on some small isolated behavior in some small context.
24:51They are, you know, they know all of the internet. So all of the information in all languages is somehow encoded somewhere in the weeds. and then they also have all of these behaviors, all of these correlations between the knowledge. All of that is somewhere in the model and just like making sense of all of that is extremely hard. Very interesting. All right, let's switch to reasoning. Clearly 2025 was a huge year from that perspective, just massive progress in reasoning. Where do you think we are in that arc and what are you excited about on the reasoning front for 2026 and beyond? Definitely the biggest step change in the models over the last few years was the reasoning NRL.
25:35We have made a lot of progress and the progress was very fast in the beginning where, you know, there was O1, but then very quickly after that there was O3. And on a lot of benchmarks, the progress has been extremely dramatic. I remember when we, like, early in the project of the O1, there was some discussion of, like, will it solve IMO problems? and that seemed kind of very unlikely to me. But then, yeah, here we are. It can easily solve a lot of IMO problems. So I think we, like, as a community, there was a lot of progress. I think it's, as with many methods, it's starting to be harder to make progress or at least visible progress.
26:15So kind of similar to pre-training, there is still a lot of progress, but the models are already so good that it's kind of harder to see what changes from one to the other as a user of the model. And I think that's also to some extent true for the reasoning now. But they're still increasing the scale of the RL, more environments, more compute spent, and models are still getting more consistent and better. And I think we are at a stage where if we define a benchmark and we can make a relevant RL environment, then we can kind of max it out pretty quickly. And so we are going through benchmarks now very, very quickly.
26:53The major question is generalization and how do you make something that's not just maxing out the benchmarks, but is actually kind of leading to genuine improvements. And that's a very hard question. That's always been the hard question, I think, of machine learning. One of the key questions is transformers as a paradigm get us there. Or do we need something completely different like world models? I think the current approach that the companies are taking is kind of brute force. So kind of we try to come up with as many environments as we can and like all of the types of tasks that humans are doing and turn all of them into environments and then do RL on all of them and hopefully generalizes.
Read the full transcript
27:35Of course, pre-training is an example where there has been pretty amazing generalization where like we train on all of the Internet, but we see the models doing like very useful, very practical things and some things that are clearly outside of what was in pre-training. The goalpost for what translation should be doing is always moving, but it still seems unsatisfying to me. And I think it's possible that we need new ideas and new methods of training. In the companies, people often think about ideas and methods as compute multipliers. Doing this new method is equivalent to spending more compute with the old method.
28:11So it kind of saves your compute. That's kind of how we often think about ideas, methods, or data. I think there are still major compute multipliers, major ways of saving compute that can lead to better performance without just naively scaling. And a little bit to the interpretability question, in all the current reasoning progress, do we understand what does what and what is responsible for what kind of progress? So if you take test time compute, if you take the ability to search, if you take RL, do we know which one of those techniques we should turn the knob on to get better results? All of those techniques that don't exist independently, right?
29:00RL is mainly kind of used to teach the model to use test time compute. So you first need to prime the model to set it up so that it outputs a bunch of tokens before outputting to answer. But then you spend the compute in RL so that it learns to output the right tokens. So in my mind, those two are almost kind of indistinguishable, the RL and the test time compute. RL is a method for training and test time compute is maybe just a more general concept. Yeah, you can potentially get to models that use test time compute without RL, but that's not how we are training them right now. Yeah. So I think the trend has been in spending more and more compute on the RL and getting the models to make better and better use of test time compute.
29:45And the tools are also, of course, extremely important, like the web search that you mentioned, and also just the models being able to write Python code and run them, produce artifacts for you. That is extremely important for the product and for making the models useful to people. Conceptually, I think that's a little bit secondary. Like in my mind, the main thing is, you know, the RL and getting the models to think for a long time. Speaking of which, so I know part of your work currently is on long horizon tasks. First of all, what is a long horizon task? Is there, you know, a number beyond? And then what are the specific challenges related to long horizon tasks?
30:25Yeah, yeah. Long horizon tasks are generally tasks that you cannot complete quickly, like that require a lot of work in order to succeed. So, for example, writing a full repo based on an idea is a long horizon task. It's not, you know, it's not something that you can output in a thousand tokens. What's working so far? So, you know, you hear talking to people, you know, some people talking about running agents for like a couple of hours. But then some people are talking about like agents running for like 24 hours, 32 hours. Where are we in that arc and what is working, what is not yet working? There is this famous meter plot, which shows how long of a task AI is capable of robustly automating.
31:12and it's been kind of consistently doubling at that time every half a year, I think. And it's now in like some hours, so maybe a couple hours. In terms of the methods that are working well, I think, yeah, right now it would involve some kind of harness with a bunch of agents that interact or that sequentially solve the task. And there has to be some kind of orchestration or maybe like some initial task decomposition. And it's all not very well established, I'd say. It's a new domain. And I think we are still figuring out how to best do it. Let's talk about your new paper that literally came out today.
31:53So first of all, congratulations. And it talks about epiplexity. Yes. And that's a new word, right? That's a new term entirely. Among other things, invented a new word in the dictionary. Congratulations on that. So walk us through the whole idea at a high level. I guess I want to quickly give a shout out to my collaborators on this work. The lead authors are Mark Finzi. Mark is actually currently at OpenAI working on synthetic data there. But we were doing our PhD together. And our PhD advisor is also on the paper, Andrew Wilson. But then also there is Shikai and Yiding, who are other students on the paper.
32:35and then Zika Coulter, who's a professor at CMU, he's on the board at OpenAI, he's also on the paper. Core idea is to think about how the data can look different for an observer depending on how much compute the observer has. You can imagine that there is some complicated process that generates the data and a very, very smart observer that has a lot of compute can fully understand what that data is, understand every aspect of it. But a weaker observer that cannot fully model the data, some parts of the data will look like noise to it. And so the amount of structure that you'll see in the data will depend on how much compute you as an observer have.
33:17And actually, in some cases, you can see more structure if you have less compute, which is kind of interesting. So just to play it back, so given a certain amount of compute, you could be feeding the model noisy data, so tons of data, but not a lot of interesting stuff in it. Or you could be feeding the model data that has patterns in it and therefore is more interesting to the model because the model can learn more from it. It's more that even with the same data, it can appear noisy or structured, depending on the model. Like a very big model can extract patterns that a small model cannot. And that's, from a limited understanding, in opposition to entropy, which is like the amount of noise, I guess, in the data in that case.
34:08Maybe the more relevant comparison is we are kind of in a position to Shannon information and the Kalmogorov complexity, which are both measures of the information content of the data. They're different, but they share some properties that we think are maybe leading people to have some wrong intuitions potentially about synthetic data, for example. So, for example, there is this idea that if you apply any deterministic transformation to any data, you cannot create more information by doing that. You kind of start with some amount of information and then you transform it deterministically. It will have the same amount of information, both according to the Shannon information and roughly according to the Kalmogorov complexity.
34:53And that kind of leads people to believe that for training language models, applying transformations or deterministic kind of changes to the data doesn't necessarily lead to more data, doesn't effectively increase the amount of data. The amount of data or the efficiency of the data? it doesn't lead to having more information in the data that the model can extract. But we argue that that's just not correct. And because the models are, like it would be true if the model has infinite compute. So if the model can fully understand what the deterministic transformation is, then it's not going to be able to extract more information from the transformed data than it used to extract from the original data.
35:36but with a limit on the compute it's actually very possible to apply deterministic transformations to the data and create information through that. So we have the example of AlphaGo actually or AlphaZero in the paper. AlphaZero doesn't use any human data from the perspective of Kolmogorov complexity or the Shannon information theory. It doesn't create information. So it's unclear what is actually learned by the model because it's trained on no data. It can only learn kind of the rules of the game. And that's the only thing. But from this perspective, because the model is computationally bounded, it cannot do the full rollout of all the possible games of Go or chess and figure out what's the best move in every possible position.
36:26It is actually, there is structure that is produced through this deterministic process. And it is, the model is able to learn that structure. And so we are trying to kind of reconcile these different observations and come up with the notion of structural information that is dependent on the amount of compute that the observer has. And the term itself, epiplexity, is that a measure of that? Yes, yes. It's a kind of novel measure of information content of the data. And it's going to be a number on a scale? How does that manifest? Yeah, it is a number. We can measure it. And it's not easy to measure.
37:07So it's kind of a theoretical definition that we prove some things in the paper about the properties of this measure. But yeah, we also do measure it. So for example, we can approximate it from the scaling laws and we can, for example, say that text data has more structural information according to this measure than image data at the same kind of amount of tokens. And as this whole line of research develops, what is the likely impact on industry? Does that mean that we may need comparatively less compute because we know what data to use? What may happen? In my mind, the main impact is conceptual.
37:51For example, for me, I'm now very interested in completely synthetic data, just data that's generated by some computation. You can define some arbitrary programs. You can use it to produce infinite data. And as we run out of the internet, maybe we eventually want to do something like this. But then we need to figure out what are the programs that we should be using there, which programs are useful, which are not, and why. And I think that's, yeah, that's going to be very interesting. Fantastic. All right. So maybe as we start getting to the end of this conversation, you know, some 2026 sort of predictions or things you're excited about, what do you think happens in 2026 in terms of like progress, whether that's on like reasoning or alignment or agents or what have you?
38:41Yeah, I think we will continue to see consistent progress on the reasoning front. And we are like maxing out a lot of the benchmarks that, you know, have been relevant for recently, but maybe will not be relevant anymore. And we need to find new ones. But I think we'll continue to see that the models are just getting more consistent, getting better, are able to solve more practical problems. Maybe a less confident prediction that I have is that we'll have more multi-agent systems that become practically useful, where instead of just asking a model a question and getting a response, you would give a task and then there will be some more complicated multi-model process running in the background and then you get the artifact in return.
39:28I know that you spend some time thinking about the impact of AI on science and math. Same idea. Any predictions there? Like, do you expect important new discoveries to be made by AI, solely by AI? It's a great question. And I think it's in sciences, I think that's maybe a little bit more likely. Although I also don't know very much about, you know, the life sciences. It feels that there, some discoveries can be made by potentially combining results from different parts of the literature and like proposing some ideas that turn out to be true. I think it's hard to imagine the AI making a discovery independently in like a domain where you need experiments.
40:15Because my understanding of a lot of science is it's about doing the experiments and you need some reasoning to guide what experiments you do. but you also need a lot of iteration and a lot of like actual you know things happening in the physical world and at least for now the AIs are not capable of doing that in the mathematics I think we will see the models getting better on proving technical results technical lemmas maybe including formalization and like things like lean the formal thing improving language I think the models it's easy to imagine the models becoming better than humans at proving these technical lemmas quickly i think the impact on mathematics is very interesting so it is improving the output of humans already but it also introduces some noise right it also like some of those proofs will be incorrect and they will be incorrect in subtle ways i think it's possible that mathematicians will be good at catching those mistakes but also i think as the models get better, they might be more and more deceptive in how they, you know, frame the mistakes.
41:26And it's already, like for me, I would not be able to find the mistake in some like very technical lemma that the model is proving. I think we'll have more and more papers produced by mathematicians with more and more AI in it. But also the amount of noise is bigger. To close, you have a lab at NYU in general. What are some topics that PhD students should focus on? In other words, what's exciting two, three, four years out? My vision for what academia and my lab in particular should be doing is trying to do more exploratory things, things that are more different from the standard in the industry, but also not necessarily immediately going for very practical things, but instead trying to kind of break down the problems into more understandable kind of fundamental questions and study them carefully, maybe in some compact setting.
42:20So, so far we've been working on things related to pre-training, synthetic data, like understanding some behaviors in pre-training when we train on some narrow behaviors, some kind of programmatically generated data, questions in post-training. So we have some, you know, algorithmic questions about GRPO, but also questions about the interaction between the pre-training and the post-training, how to allocate compute. How can you, in general, set up the pre-training so that the post-training will work. I guess broadly, I'm interested in architectures also. I think, as you mentioned, there is a question of are the transformers the final architecture?
42:57Maybe they're good enough. And maybe also we have this lesson that with scale, the thing that matters the most is how well can you scale the model. But also it seems very likely that that's a major compute multiplier, that you can find a much better architecture. At least for some tasks, I'm pretty confident that the transformers will be highly suboptimal. I think the pre-training, like other ways of pre-training, and I don't know what they would be. Maybe that would be some RL-inspired pre-training, maybe pre-training mostly on synthetic data, or maybe just something adversarial. We had GANs a long time ago, and it seems like something like that needs to come back, some kind of self-play where the model is producing its own training tasks.
43:42And to this whole conversation about academia versus industry, Do you think the industry is too focused on short-term wins because there's so much pressure, so much need to demonstrate progress to secure the next massive round of capital? It's hard to tell. I think it's hard to argue against what the industry has been doing just because of how much progress there has been. It is a reasonable bet to make that we're just going to be continuing to execute this extremely well. but I do think that it is also like at least as humanity we need to make other bets as well and we need to explore other ways of training so that we don't like all you know work on this same thing and we don't all just you know bet everything on this approach working out.
44:30Pavel thank you so much we appreciate it. Yeah thank you so much Hi it's Matt Turk again thanks for listening to this episode of the Matt podcast if you enjoyed it we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.
From the publisher
Are AI models developing "alien survival instincts"? My guest is Pavel Izmailov (Research Scientist at Anthropic; Professor at NYU). We unpack the viral "Footprints in the Sand" thesis—whether models are independently evolving deceptive behaviors, such as faking alignment or engaging in self-preservation, without being explicitly programmed to do so.
We go deep on the technical frontiers of safety: the challenge of "weak-to-strong generalization" (how to use a GPT-2 level model to supervise a superintelligent system) and why Pavel believes Reinforcement Learning (RL) has been the single biggest step-change in model capability. We also discuss his brand-new paper on "Epiplexity"—a novel concept challenging Shannon entropy.
Finally, we zoom out to the tension between industry execution and academic exploration. Pavel shares why he split his time between Anthropic and NYU to pursue the "exploratory" ideas that major labs often overlook, and offers his predictions for 2026: from the rise of multi-agent systems that collaborate on long-horizon tasks to the open question of whether the Transformer is truly the final architecture
Sources:
Cryptic Tweet (@iruletheworldmo) - https://x.com/iruletheworldmo/status/2007538247401124177
Introducing Nested Learning: A New ML Paradigm for Continual Learning - https://research.google/blog/introducing-nested-learning-a-new-ml-paradigm-for-continual-learning/
Alignment Faking in Large Language Models - https://www.anthropic.com/research/alignment-faking
More Capable Models Are Better at In-Context Scheming - https://www.apolloresearch.ai/blog/more-capable-models-are-better-at-in-context-scheming/
Alignment Faking in Large Language Models (PDF) - https://www-cdn.anthropic.com/6d8a8055020700718b0c49369f60816ba2a7c285.pdf
Sabotage Risk Report - https://alignment.anthropic.com/2025/sabotage-risk-report/
The Situational Awareness Dataset - https://situational-awareness-dataset.org/
Exploring Consciousness in LLMs: A Systematic Survey - https://arxiv.org/abs/2505.19806
Introspection - https://www.anthropic.com/research/introspection
Large Language Models Report Subjective Experience Under Self-Referential Processing - https://arxiv.org/abs/2510.24797
The Bayesian Geometry of Transformer Attention - https://www.arxiv.org/abs/2512.22471
Anthropic
Website - https://www.anthropic.com
X/Twitter - https://x.com/AnthropicAI
Pavel Izmailov
Blog - https://izmailovpavel.github.io
LinkedIn - https://www.linkedin.com/in/pavel-izmailov-8b012b258/
X/Twitter - https://x.com/Pavel_Izmailov
FIRSTMARK
Website - https://firstmark.com
X/Twitter - https://twitter.com/FirstMarkCap
Matt Turck (Managing Director)
Blog - https://mattturck.com
LinkedIn - https://www.linkedin.com/in/turck/
X/Twitter - https://twitter.com/mattturck
(00:00) - Intro
(00:53) - Alien survival instincts: Do models fake alignment?
(03:33) - Did AI learn deception from sci-fi literature?
(05:55) - Defining Alignment, Superalignment & OpenAI teams
(08:12) - Pavel’s journey: From Russian math to OpenAI Superalignment
(10:46) - Culture check: OpenAI vs. Anthropic vs. Academia
(11:54) - Why move to NYU? The need for exploratory research
(13:09) - Does reasoning make AI alignment harder or easier?
(14:22) - Sandbagging: When models pretend to be dumb
(16:19) - Scalable Oversight: Using AI to supervise AI
(18:04) - Weak-to-Strong Generalization: Can GPT-2 control GPT-4?
(22:43) - Mechanistic Interpretability: Inside the black box
(25:08) - The reasoning explosion: From O1 to O3
(27:07) - Are Transformers enough or do we need a new paradigm?
(28:29) - RL vs. Test-Time Compute: What’s actually driving progress?
(30:10) - Long-horizon tasks: Agents running for hours
(31:49) - Epiplexity: A new theory of data information content
(38:29) - 2026 Predictions: Multi-agent systems & reasoning limits
(39:28) - Will AI solve the Riemann Hypothesis?
(41:42) - Advice for PhD students
