AI is Already Building AI | Google DeepMind’s Mostafa Dehghani

2 Apr 2026 · 1 h 5 min · 28 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Recursive self-improvement (AI building AI) via “loops” in labs; roadblocks like evaluation, compute, and model collapse; how continual learning differs; and how multimodal models (Gemini family) enable image generation and potential positive transfer across modalities.

Guest

Mostafa Dehghani, top AI researcher at Google DeepMind; core contributor to Universal Transformers, Vision Transformers, and the natively multimodal Gemini family.

Key claims

  1. Most labs already train new model generations using prior models; full RSI needs long-horizon and full automation.
  2. “Self-improvement” is the next step in removing human bottlenecks, analogous to earlier shifts from hand-crafted features to learned representations.
  3. Evaluation is the main bottleneck: “you can only improve what you can measure.”
  4. Model collapse risk rises when loops are closed without real-world grounding; verifiers/reward signals can anchor progress.
  5. Data won’t disappear; it shifts toward environments that provide sensory/physical grounding.

Notable examples

  • Karpathy’s “auto-research” as an early self-recursive research loop.
  • Universal Transformers (parameter reuse via recursion).
  • Vision Transformer patchify approach (“An Image is Worse, 16 by 16 Worse”).
  • Nano Banana 2 / Gemini 3.1 Flash Image; emphasis on natively multimodal generation and “positive transfer” across modalities.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding AI Loops

0:45 to 1:18

Exploring the concept of AI models improving by thinking recursively rather than just increasing in size.

“We also dive into the technical evolution of image generation with Nano Banana 2 and why continual learning could completely disrupt how enterprise data pipelines and RAG systems are built today.”

Self-Improvement in AI

1:18 to 3:16

Discussing self-improvement trends in AI development, focusing on the removal of human bottlenecks.

“In classical machine learning, humans had to sit down and manually engineer the features.”

Recursive Self-Improvement

3:16 to 6:21

Examining the concept of recursive self-improvement and its implications for AI evolution.

“And I think that's basically on the development side.”

Challenges of AI Evaluation

6:21 to 10:10

Discussing the challenges in evaluating AI improvements and the need for effective evaluation metrics.

“These models are going to improve themselves and keep learning from the world.”

Formal Verification in AI

10:10 to 12:34

Analyzing the role of formal verification in AI self-improvement and its limitations in complex domains.

“At the end of the day, you can only improve what you can measure, right?”

Addressing Model Collapse Risks

12:34 to 14:01

Discussing the risks of model collapse in AI and how to prevent it through robust external signals.

“A couple of weeks ago, we had a fun conversation with Karina Hong of Axiom Math and we talked about formal verification.”

Understanding Model Collapse

14:01 to 15:01

This segment delves into the concept of model collapse and its implications.

“The second you start veering away from math and code, you start getting into like very messy territory.”

Generalization vs. Specialization in AI

15:01 to 17:08

Discussion on the balance between generalization and specialization for AI models.

“So basically, we're going to have some sort of data and environment that these models are interacting with.”

The Trade-offs in AI Model Development

17:08 to 20:06

Exploring the trade-offs in developing specialized versus generalized AI models.

“maybe I need to make sure that in a very specific area, I can build that.”

Automation of Intelligence: A Philosophical Query

20:06 to 20:57

Contemplation on the future of AI and the potential for automation of human intelligence.

“I think short term, it's like very, very effective because first of all, during development, you care less about, you know, like all the dimensions.”
Show all 28 chapters

The Evolving Role of Data in AI

20:57 to 26:16

Discussion on how the concept of data will evolve as AI develops and self-improves.

“You said something a few minutes ago that I thought was so intriguing, which is this idea that the Carpathies of the world and you of the world could be automated.”

The Balance of Pre-training and Post-training in AI

26:16 to 28:00

Insights into the ongoing importance of pre-training alongside the rise of post-training.

“We're seeing the startups emerge in that field.”

The Transformation of Pre-Training in AI

28:00 to 29:04

Explore the evolving perspective on pre-training in AI models.

“It's like also super interesting for me because I'm, again, like a little bit like new to this side of the, operation.”

Understanding Continual Learning in AI

29:04 to 31:06

Learn about the concept of continual learning and its challenges.

“So when I say old, maybe I'm referring to two weeks ago or something.”

Current State of Continual Learning Research

31:06 to 33:14

Discuss the current advancements and challenges in continual learning.

“all the news, everything that is happening in the board, everything is just like, you know, updated.”

Mostafa Dehghani's Journey to AI

33:14 to 36:14

Discover the background and career path of Mostafa Dehghani in AI.

“It's kind of interesting because it is one of the things that can be heavily theoretical.”

The Universal Transformer and Its Impact

36:14 to 39:28

Unpack the significance of the Universal Transformer paper in AI.

“And that was very much that idea that we started with at the beginning of this conversation of like loops and recursive stuff.”

Revolutionizing Image Recognition with Vision Transformers

39:28 to 42:00

Learn how the visual transformer paper changed the landscape of image recognition.

“In a mixture of experts, you have flops-free parameters, so parameters that they're not actually bringing any flops.”

Transformers and Image Processing

42:00 to 43:50

Learn how transformer architectures can effectively process image data.

“And it was also a little bit of a surprise for us that, oh, you know, they were all thinking about something fancy and very complicated, maybe in the integration of having convolutions and stuff.”

Introduction to Nano Banana and Its Evolution

43:50 to 46:00

Discover the developments in the Nano Banana project and its implications for image generation.

“So you are part of the Nano Banana team, which must have been so much fun when this came out and went just completely viral and what an incredible product.”

Exploring Multimodality in AI

46:00 to 48:20

Understand how multimodal capabilities enhance AI's learning and generation processes.

“But the most exciting part for me was, can I see a glimpse of transfer from these modalities?”

Incremental Generation in Image AI

48:20 to 50:20

Learn about the concept of incremental generation and its advantages in AI image creation.

“Gemini generating like images from Gemini 1.”

Efficiency in AI Models

50:20 to 53:10

Explore the strategies that enhance the efficiency of AI models like Nano Banana 2.

“For example, you enable interleaved text image generation where the model can think in not only text token, but also in pixel space, right?”

Hot Takes on AI Research

53:10 to 56:00

Hear insights on common misconceptions and underrated ideas in AI research.

“This one, like the team actually shipped it.”

Underrated Aspects of AI Research

56:00 to 57:26

Exploring the often overlooked importance of continual learning in AI.

“oh you know let me just like you know patch by adding something for the system instruction or developer instruction.”

Confidence in AI's Technical Advancements

57:26 to 1:00:06

Discussing the pitfalls of overconfidence in technical solutions without addressing broader societal issues.

“It's not going to look like as is today.”

Challenges in Long Horizon Automation

1:00:06 to 1:02:48

Highlighting the reliability issues in automating long-term tasks and the necessity for error recovery.

“I start from the things that I'm like really excited about it.”

Defining Intelligence in AI

1:02:48 to 1:03:57

The need for a clearer definition of intelligence to measure progress in AI development.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Most of the people don't realize that this is like already happening, especially over the past few months. In almost every lab, the new generation of the models are built heavily using the previous generation of the models. What is missing right now is long horizon and full automation. And we're moving to that direction super, super fast. The moment that we have this full automation, we can close the loop of self-improvement. We just got rid of the human bottleneck for improving these models, which I expect to see a huge jump again from such development. Hi, I'm Matt Turk. Welcome to the Matt Podcast.

0:31Today, my guest is Mustafa Degani, a top AI researcher at Google DeepMind and a core contributor to some of the most influential architectural breakthroughs of the last decade, including Universal Transformers, the Vision Transformer, and the natively multimodal Gemini family. In this episode, we unpack what's hot in Frontier AI right now, including what it actually means for AI to think in loops and the immediate timeline for recursive self-improvement, where AI autonomously builds the next generation of AI. We also dive into the technical evolution of image generation with Nano Banana 2 and why continual learning could completely disrupt how enterprise data pipelines and RAG systems are built today.

1:11Please enjoy this fantastic deep dive with Mustafa Degani. One of the hottest concepts in AI research right now seems to be the concept of loops so i thought it'd be a fun place to start this idea that models are going to improve not by being bigger but by thinking recursively what does that mean exactly definitely one of the top is active areas for almost every lab to invest in looping and it has like like operation at different levels the one that is on on the uh on the module level is basically the the looping that we use like architecture or at inference time for test and compute and stuff like that and then at a higher level is basically the loop that we have over the development of these models which is basically we refer to add to it as self-improvement if i want to put it like like very like like let's talk about self-improvement as as like this general concept right like if i want to um put it like very simply it is really just the continuation of the trend that we've been writing for decades, right?

2:17And think about it. In classical machine learning, humans had to sit down and manually engineer the features. And you had to decide what the model actually pays attention to. And deep learning and neural network came along and they said, okay, let's just remove that. Let the model figure out the representation itself. And that was actually a huge deal. And we somehow removed a massive human bottleneck and human bias. And then further, and instead of just designing architecture, we started learning them to, instead of curating every piece of training signal, we scaled to basically data-driven approaches and let the data speak.

3:01And the self-improvement and this loop into development is just the next step in the same direction. And the whole idea and the whole point of it is you're removing the human bottleneck and bias from improving these models, right? And now, like you say that, okay, you know, not just human doesn't have to handcraft features anymore, but also we don't want the human to sit in the loop every time that the model has to get better. And I think that's basically on the development side. So it's not radically new. It's the same story, just a new chapter of the same story. And I think every time that we removed human judgment from this process, we kind of got over a bottleneck.

3:43I would say like the self-improvement and looping over the development is kind of like doing that at the highest level, which is basically improving these models. If we want to go to more detailed level of looping, we can talk about ways of increasing test and compute for these models and how we let these models to loop over their process within a specific problem to refine it, to think about it. And I think the most familiar form is just like chain of thought and letting the model to think with extra tokens. That's beyond that. And you can think about different ideas that you let them want to increase the compute for any specific problems.

4:24Like, what if I have like Donnie tokens that they can use as like read and rock tape to kind of like, you know, re-verify what I've done and go through the solutions or the process that I'm like doing over different steps and understand, you know, like what has been done growing, what has to be done next. or even like, you know, negative sparsity, which is basically reusing part of the model multiple times. And this sort of looping is also like been shown to be super, super helpful, mostly because you just let the model to do more compute on a difficult problem. So that's self-improvement at inference time.

5:03I think you alluded to earlier, there is also a bigger concept that's maybe, I guess, more science fiction, except it seems to be becoming a reality very quickly, which is this concept of recursive self-improvement or RSI. That seems to be what a lot of people are talking about. I think ICLR is coming up in a few weeks and there's a bunch of papers focused on that. So what is that? What is recursive self-improvement as a concept? It's actually interesting because you referred to that as something that looked like a bit of a sci-fi situation where these models are actually improving themselves.

5:37and and that's true because a few years ago when you were you wanted to talk about this you could just write a perspective paper at a conference and like you know talk about it at super high level but if we go uh and and check out what is happening right now like to a really good extent happening like most of the the like and and it's somehow like most of the people don't realize that this is like already uh happening especially over the past few months in almost every lab uh the new generation of the models are built heavily using the previous generation of the models. I think that's basically the case, again, everywhere.

6:15And it's not fully automatic yet, but the direction is super clear. And it's easy to imagine that we're going to get to a situation with full automation. These models are going to improve themselves and keep learning from the world. And again, it has relation with other concepts, like continual learning and other concepts that we are still not yet to the most advanced point of it. But if someone comes and say that, oh, you know, I have an idea to get a model to calculate the gradient and update the weights on the fly, it just feels very normal. It's not something that, wow, this is such an amazing idea.

6:58I think what is missing right now is long horizon and full automation. And we're moving to that direction super, super fast. The moment that we have this full automation, I would say we can close the loop of self-improvement and then the problems become mostly providing compute for these models to actually do what they want to do. And as I said, like in the comment, we just got rid of the human bottleneck for improving these models, which I expect to see a huge jump again from such development. So people may have seen or heard about Karpathy's auto-research project a few weeks ago. Is that an example, presumably narrow to make it work?

7:43Is that an example of a self-recursive loop? That is definitely. And I think that was one of the early examples of seeing these models actually doing something super sensible on the research side. So we've been seeing them doing a lot of good work on improving the engineering part of the development loop. but on the research side which like you know you think about okay you know maybe some sort of you know gut feeling or intuition is needed and like a researcher with like a long time of like you know playing with these models and experience can do this but but not necessarily you know like a model uh i think we're seeing the sign that okay you know maybe that basically that kind of like golden um part of the recipe a successful recipe that mostly coming from um like um intuition of a good researcher is coming to kind of these development loops by these models.

8:34And it's a bit hard to think about, okay, you know, does it mean that we can replace like every genius researcher with these models like very soon? Maybe, and I don't know like how soon, but this is definitely a sign of something that we kind of doubted like, you know, a few years ago, you know, we couldn't believe that, well, this is going to happen that early, which is very exciting. I want to play it back just to make sure that people listening to this understand. We're talking about AI building AI. And I think a few months ago, if you talk to researchers, people would say, oh, yeah, we already use AI to build AI.

9:10But that really meant that we use AI tools and reasoning models to come up with ideas and thoughts about building models. But here, what we're talking about is AI automatically updating itself, updating its weights in a recursive manner, leading to potentially a dramatic acceleration in progress. And what you're saying is that this is largely upon us and a question of longer horizon and basically more compute. Is that fair? I think so. Like this is one. And the other one is also, I'm not going to say that, oh, you know, soon we're going to have these models like fully automated. And there are actually many problems that we have to solve.

9:50But directionally, I can see how this can happen. You know, like it's not something that I would look at it as like super hard. It's like hard, but very possible. Okay. So what are the roadblocks? So you talked about compute. Is evaluation one of them? Because presumably the model needs to understand what is right and what is wrong in terms of the quality of the answer. Is that one of the issues? 100%. At the end of the day, you can only improve what you can measure, right? And then getting evaluation is just hard. And at the end of the day, it becomes almost a philosophical problem, not just the technical one.

10:28This is actually a very interesting observation. So if you have a team of super competent people, most of the time they can do massive progress on a problem if there is some concrete eval to heal climate. But if there's no eval, it's just like really hard to make progress. And the fact that we don't have evals that like, or like even defining evals that can maybe measure, oh, not how close we are to the point that we can actually get a self-improvement loop. It's just like, we don't have that. And it's just making it like much harder to measure the progress in that direction. But there are proxies and there are definitely some evals that, you know, like we're going from, oh, maybe we can evaluate like every step of the model toward this direction.

11:19And maybe we can evaluate up to this many turns of the model, or maybe we can evaluate the model helping itself to improve in either a specific framework and in the specific setup and this part of the machine learning that needs iteration. It's also quite interesting because the difficulty of building eval is like the infrastructure that you need to reliably run evals that are super complicated is also like super hot. It's quite funny, but sometimes I'm figuring out that, okay, how can I create an environment for a model that operates safely within Google, right? And does all the jobs that NRE and like research engineer or research scientist can do.

12:06Like it is a safe setup, you know, where it can, like, because right now we definitely, we don't, we're not confident about, you know, them doing the right things all the time. and measuring how much they can push and how long they can push a task is very difficult. And connecting all these points into an environment that these models are operating and then get them run efficiently and bringing diversity to eval is definitely one of the bottlenecks of making progress in this direction. A couple of weeks ago, we had a fun conversation with Karina Hong of Axiom Math and we talked about formal verification.

12:42Is that a promising area from your perspective? Is something like formal verification what would enable you to make sure that the improvement loop keeps continuing? In my opinion, formal verification is one of the most powerful keys to enable self-improvement, but it's not B key. And if you think about it, like for mass, code, logics, it's great. You can run a proof, either it checks out or not. If you go to other domains that are a little bit messier, like, for example, you cannot write a formal proof that if a doctor's recommendation is good, right? So it's not hard. It's not easy to have, like, to extend this formal verification to all the domains in real world.

13:23But one question that is actually an interesting question, which is very relevant to formal verification, is how can we look at these methods and formal verification and build that kind of tight and honest feedback loop for the messy part of the world? I think that that's like very inspiring to build like on top of these like formal verification methods to extend to domains that like not easy to verify easily. But you need some sort of clean and tight feedback loop to be able to make progress. So the same problem as reinforcement learning, right? The second you start veering away from math and code, you start getting into like very messy territory.

14:06Is model collapse one of the issues to think about or is that orthogonal? Model collapse is definitely a risk, right? And I would say model collapse mainly happens when you have a loop that is completely closed, right? And if you don't have any outside signal and just the model, for example, talking to itself or operating in a very restricted environment, there's a good chance that your model collapses. But if you have a strong verifier or some sort of like a real reward signal that anchor this kind of like signals that is coming from like AI generated data, for example, it can be quite powerful.

14:46I think like the key here is to stay grounded to something real. And then you can most likely avoid like, you know, things like model co-aps. But yeah, I mean, again, it's a risk, but it's not definitely like a major rocket. And perhaps to make this accessible to everyone, can you define what model collapse is in the first place? So basically, we're going to have some sort of data and environment that these models are interacting with. But those environments and data are designed, for example, by another model. This is just an example of that, right? And then you become really, really good at this specific part.

15:23And then suddenly you lose generalization to anything beyond that. And this is like one of the kind of like definition or one of the cases that like a model collapsing would result to. So you mentioned losing generalization. Is that particularly in the concept of RSI a worry that either you have those self-reinforcing loops, but they need to be fairly narrow or you have more general models, but then you kind of have the loops? This is an interesting question. Again, like, you know, generalization versus specialization. Let me go a few steps back. We had this discussion like many, many times. How should we do a trade-off between generalization and specialization when we are developing these models?

16:01I think long-term, you want a model that knows everything and knows when to go deep versus why, right? Imagine, like, you have an agentic actor, right? Like, if you're an agentic coder, if your agent is, like, super strong at every step of operation, like a really, really good programmer, it's amazing. You know, like, it's, like, super specialized. But for many of the problems, like coding problems, you need some sort of planning. and understanding what's going on and collecting information and based on the context, deciding what to do. And then after you define the steps, then your super-strong specialization just kicks in.

16:40And before that, being a generalist is super useful. Definitely generalization is one of the things that you need to get to the ultimate side of AGI. But short-term, I would say building a specialist model is probably the fastest way to learn what is actually possible. And in many cases, these specialized models are becoming a stepping stone toward a generalist model, which is super valuable. So you can imagine that, oh, if I'm actually thinking about self-improvement, maybe I need to make sure that in a very specific area, I can build that. Maybe I focus on coding, and then if it works out, then I go through like, you know, how to widen that and how to bring more into this specialized setup.

17:28One thing that I always like say is that people don't care what category their problem falls into, right? And if a human calls like something a problem, then AI should be able to solve it. And I think that's like fundamentally a generalist need, right? So at the end of the day, you need generalization and like playing this, like going through this spectrum of, you know, like super generalized model and super specialized model is more about, you know, like a long term, short term and how to take advantage of each side during this process. What's a specialized model to date? Is that a separate model or is that a broad general model that's trained in a specific way, including in particular through RL?

18:13Okay, so here's the point. We used to have constraint like compute. And then if you wanted to push a model to be like so tall, we would choose specific dimensions. And then we say that, okay, you know, we want to kind of like allocate the compute that we have to that and then make this model look really good at this, like, you know, something that is like extremely expert at this. So that was like basically the trade-off that we were like trying to make given the compute budget that we have. So as we go through this, like the phase of compute becoming more available, like cheaper, and then maybe with other stuff, like data and stuff, one of the other trade-offs that pops up is, especially in post-training, like this game of evacuable, that sometimes it's really hard to get your model to be good across the boards.

19:05So you try to kind of make it good at something like multimodality. somehow you see some regression on, you know, the coding and you make it like good at like coding and multimodality, it becomes a little slightly worse than a model that you had, like, you know, like math and reason. So it's hard to kind of find a balance. And part of it is because post-training does a little bit of like an overfitting, you know, like at the end of the day, when you post-train the model, you are trying to overfit it to the best local optima you have. When recipe becomes like, how can I find the best local optima?

19:43It becomes the problem of, okay, there's no local optima that is good for everything. So you need to kind of choose, right? And then like seeing this, you end up with like making some decisions along the way and saying that, okay, you know, maybe for me at this stage, because of the meat that I have in my organization, like with respect to the competition that is going on, I need to choose, you know, this specific axis. Like, you know, for example, some companies have like a very strong focus on coding, which is, okay, you know, I make my job like super easy, you know, like, or not super easy, but like much easier than the competitors that they want to basically shoot a model that is good across the board.

20:19I think short term, it's like very, very effective because first of all, during development, you care less about, you know, like all the dimensions. So maybe it's just faster to be trained. Like you kind of free up some space from the mind of your researchers and engineers that, okay, you know, like forget about this. Just let's push this to the max. And then the other one is also like you don't hit the trade-off like immediately. And especially this model is that like, okay, you know, I'm going to pick this specific axis and then make the model look really, really good at this. Sometimes, again, like this is a decision based on the place that you are at, you know, again, like organizationally, like competitors, like and stuff like that.

20:57Great. You said something a few minutes ago that I thought was so intriguing, which is this idea that the Carpathies of the world and you of the world could be automated. What happens if like the brightest minds in the world get automated and the AI creates itself? Like at some point, is there just no one knows how the AI works? Is that an actual possible future? This batch is very philosophical. I don't know. Well, let me give you one quick thing that I thought about it a few days ago. I have a daughter. She is like one and a half years old. I've been impressed over the past few years. Like very interestingly, like I've been proven wrong multiple times about like the timeline that I had in mind.

21:40For example, sometimes I say like, oh, this is gonna happen in six months. Never happened. Sometimes like, oh, this is just like so hard. Like within the next 10 years, there's absolutely no chance to solve it. And then boom, like in two months, three months, someone had a brilliant idea and they solved it. So it's like really hard to predict the future. And it was thinking like, okay, you know, like so you're talking about like catapathy and like, you know, again, like other researchers. But I'm thinking about, okay, like what about the next generation? You know, if like my daughter at some point comes to me and asks like, okay, like what should I do?

22:09You know, like what do you recommend to study? Like what major and like, you know, what branch of the science or research should I kind of like, you know, dig in and like be the expert on? I really don't have a good answer. You know, like almost it doesn't exist. And it's just like really hard to predict the future. what i know is there are a few skills that are probably key to be able to to make impact in this world and also be relevant staying relevant like one of them is like a strategic and and having all the parameters on your table when you're making a decision and and becoming absolute expert about a very specific subject most likely is not going to be useful in like the near future.

22:56I think like, you know, the brilliance of Keapathy is not like, you know, he's a good programmer or he's a good, definitely he's a good teacher, you know, but I'm saying like, you know, these are not like the most impressive part of it. Like the most impressive part for me is that he has a really good overall view of like what is happening like by putting himself in in the in the in the um like the stream of information he can make a decision about okay what is the next most impactful thing to do and now like you know the things that he does to make impact is very different from like you know the things that he used to do like five years ago and i think he can be able to do that like continue doing that you know like what is that the things that he's going to do he's going to be doing like in like in five years i don't know but i know it's like he's smart enough to figure it out and still keep making impact on the board so ai researchers are not uh researching their way uh out of a job just yet uh hopefully we are smart enough to do that uh all right maybe that's more of a macro question as i think about um you know where the value lands in this ecosystem But if AI just keeps creating itself, then is data still needed in that equation or is that all compute?

24:08The concept of data is a little bit like broader than just like, you know, tokens, right? And if you think about data as whatever that the model can get signal from, either it is like predicting the next token in ROG X, which we kind of like, you know, use in print training, or super complex environment that the model interacts with and then gets signal. This is something that basically we can refer to it as data, right? And it's not like data or the value of having good data or working on data is going to disappear and compute is going to become the only things. At the end of the day, I think the board that we're doing on the data side most likely is going to shift toward building environments or making sure that these models can interact with physical warts.

25:00And then it becomes more of a problem of, okay, how can we provide more grounding for these models? They are good at improving themselves, but as long as I expose them to real work data, right? And like in a real work environment. So providing data becomes more about, okay, how can I give access to this specific model to something that we never had? For example, again, something came to my mind, which is again, a little bit sci-fi, but how can I make smell accessible to these models? Right now, there's no good way. But then data becomes like, okay, information or anything that is for us because of all the sensory that we have is really easy.

25:43Right now I'm sitting here. I know how hard is my chair. What is the temperature of this room? All this sensory information is something that is coming to me. And then the next board that I'm saying is based on all this input, right? And then providing this for a model that does self-improvement is already a really hard problem. So I would say that the work on the data would shift toward making these sensory information more available to these models in a way that it enables them to really improve themselves, given all this information in a more effective way. Yeah, interesting. Yeah, there seems to be a big trend towards sensors as a service.

26:19We're seeing the startups emerge in that field. Okay, super, super interesting. Zooming out from self-improvement for seconds, the big theme of the last year has been the acceleration of post-training in addition to pre-training. So the whole, you know, reinforcement learning aspect of things. Where do you expect gains to come from in the next few months or year? Is that more post-training? Is that more pre-training? Is that both? Is that something else? The answer to this question really depends on when you actually ask this question. And it's obvious that we're going to be having a bit of a swing back and forth between pre-training and post-training.

27:00At the end of the day, I want to say that pre-training is still the foundation. And you can never post-train your way out of a week-based model. but right now the current like the return on post-training is really strong and i started working on post-training myself like like a few months ago like gemini post-training and like mostly coding and agentic i can see how a brilliant small idea can make a model like 10x better for example in terms of behavior at a fraction of the cost of the pre-training right this is again like you know we can see how post-training is like like the place to make a a lot of impact and improve these models.

27:38But on the other hand, like I know at different companies, it's also the case, but at GDM, a lot of exciting recent work is going into the pre-training side and like new recipe, new ideas. And I would say like, you know, the work that we're doing on the pre-training is going to unlock a lot of downstream possibilities. Post-training is just like a different mode of operation. It's like also super interesting for me because I'm, again, like a little bit like new to this side of the, operation. But at the end of the day, I always expect to see what's going on between post-training and pre-training.

Read the full transcript

28:14Your comments on pre-training are against that narrative that appeared a few months ago that pre-training was dead. That's not your take at all, right? I think everyone has ideas on pre-training side. At the end of the day, going for that idea is a function of complexity and the expected gain, right? And sometimes you feel that, okay, you know, there are low-hanging fruits and and it like you know instead of bringing this complex like you know recipe to the pre-training the one that i have like which is simple elegant super scalable i'm going to push this and then move the effort to the post-training and then at some point like the base model becomes the bottleneck and then you're happy to take the complex recipe and bring it to the pre-training and then like hit pushing it i think pre-training is dead i i would say like maybe like you know the old, it's also a little bit difficult to talk about old and new because the time frame is very different.

29:11So when I say old, maybe I'm referring to two weeks ago or something. But the way that we used to do pre-training, maybe a year ago or two years ago, maybe diminishing return is obvious, but I can see how new ideas are bringing fresh energy into the pre-training and suddenly just open a door toward like something exotic that might actually drastically change the base model capability over time. So exciting stuff for Gemini 4 whenever it comes out. You mentioned continual learning earlier and that's another one of those hot topics that people have been talking about. Can you define continual learning for us so that this conversation is educational for a broad group of people?

30:00Maybe compare and contrast that with the self-improvement loop like those are two different things but help us understand the difference definitely they're related but like they're distinct right so self-improvement is about a model getting smarter over time and improving its capability like the model itself doing it continual learning is mostly about a model staying current right like and and and think about a doctor that like keeps reading new research and they like refresh their their knowledge about stuff and like you know they're trying to make sure that you know the knowledge doesn't go stale the shared enemy between like self-improvement and and continual learning is a model with frozen weights over time while the board is just like going right like you know if you have if you have a model that is just frozen and the board is moving then like you neither get like self-improvement nor nor continual learning but continual learning is mostly focused on making sure that, you know, if there's like fresh knowledge in the board, like the model knowledge cutoff is not like in the past.

31:05So it's constantly, you know, like, for example, overnight, all the news, everything that is happening in the board, everything is just like, you know, updated. So if today you ask a problem, if you ask a question from the model, those knowledge, which is like super fresh, is already in the weight of the model. So it doesn't have to kind of like, you know, depend on external source to bring it in. And it's hard. It's like really, really hard. And the biggest problem, like not the biggest, but like one of the big problem is like catastrophic forgetting, where you get your model to learn about new information after you're done training that model.

31:42Like, you know, and then and suddenly you see regression in the knowledge that you learn already in the main training phase. And it's a very active area of research right now. And what's the reality of continual learning as of now? Is that built into existing systems, not at all about to? There are two sides of it. Like one side is, I think like the research is not like yet to a very, like to a point that you think that, oh, you know, this is the recipe, you know. I just need to kind of like, you know, exploit it and push productionization, right? But basically, every time that you have a new problem that is key, you have this phase of exploration where people try to try different ideas and go jump over this idea to another idea, which could be so different.

32:34And then when you're confident about this working to some extent, you go to the exploitation mode and say that, oh, let me just make it as good as it can be. And this is the way to push it. And let's scale it. Let's just develop infra for it, make it super fast, productionize it, and see what happens. I think that is not yet there. The other one is also, again, as I said, because we've never had a super confident recipe for continual learning, building infra for it and investing in something that is fast is hard. Given that, I've seen very impressive progress on this, of the Sweden Gemma and GDM.

33:14It's kind of interesting because it is one of the things that can be heavily theoretical. I've seen people who are doing a lot of theory work and they got into this problem. And they're having a lot of fun and they're also making a lot of impact. And it's impressive how much progress we made on this. But I don't think that we have yet any idea that everyone says that, oh, this is it. Let's just do it. Let's push this. Great. I'd love to talk about you and your background. Tell us your story in a few minutes. How did you come to do this work and what was your journey to AI and then your journey to Google DeepMind?

33:55So I did my PhD at the University of Amsterdam on machine learning and mostly on the language model side and text and search and retrieval. And then I think what kind of like pushed me toward trying really hard to be on the mainstream and be part of this group that are like, you know, hustling to make like really good progress. I did a few internships like back in 2016 and 2017. and and the funny story is um i did an internship in at google brain in 20 like early 2017 and then it was amazing it was just like you know i went to to to this team they were working on like lsdms for you know like summarization summarization was actually one of the most like interesting problems at that time i was like amazed i was like so this is so good i really i just want to keep doing this for the rest of my life you know this is it and then i got uh i got a return offer to go back and then do another internship at the end of the same year the recruiter told me that oh you know there's this team that they just published a paper maybe you've heard about it like transformer and then you're looking for an intern and i haven't had a chat with i remember i had a chat i had a chat with lucas trizer and then lucas was talking to me and saying like oh yeah like we have this idea of building like a qualmograph machine based on transformer and it was so excited about this and then like you know we like we finished up that the conversation and i started sort of sending a message to a recruiter i was like i don't know if i want to go with this team it's just like they're doing something random like who like like everyone's doing lst i'm like why should i go and work with like a group of people who are working on this like random architecture like transformer it's just like it's gonna die you know i and then he tried and he find any other team for me to join.

35:41So I joined this team as an intern and that changed my life. Being among these super brilliant, super smart people that they believed in some vision and direction where almost everyone was excited about something else was very inspiring. And then we work on this Kolmogorov machine idea of what it turned into Universal Transformer prepare, which, you know, like recursion in depth and reusing parameters was coming out of it. And still, it was like making a lot of impacts after almost like 10 years. Tell us about that quickly. So that was in 2019, I believe. And you were a co-author of that paper.

36:20And that was very much that idea that we started with at the beginning of this conversation of like loops and recursive stuff. So Universal Transformer, like we wrote that paper in 2018. And I think it was also rejected one time from one conference. And it was accepted in 2019. I don't remember exactly, but yeah. Yeah, I think it was accepted, iClear, but it was rejected from NeurIPS or something. The whole intuition was there is something about reusing parameters and a model going through its output another time, you know? And so basically you generate something and then you kind of like, you know, pass it into the model again.

36:56And then the model has the chance of doing this. So we started with, I remember Lukasz had this algorithmic dataset, which I remember he used to call it algorithmic tasks. And it was part of this code base based on TensorFlow, like tensor to tensor was the name of the code that's still there. And I can even find my request into that for pushing the universal transformer code. And we saw that basically there are some problems like copying an input to the output or, you know, like doing something algorithmic with like super long input on the output side, which is super easy. But the normal models, like the normal transformer was like failing awfully at this.

37:36And we saw that, you know, like looping is like do it perfectly. And then at that point, I remember we had this baddie like data set from Meta. and it was like doing great on that. And then the idea of test and compute, which basically you train with fixed amount of compute, but at test time, you unleash your model to do more computation, throwing more flops on the input, was coming to our mind. Like super excited about this. And then we ended up with actually kind of like introducing this adaptive computation mechanism into this, which was again, like some sort of inspiration from Alex's paper, like from Alastair.

38:16And then like a very interesting ride because we were pushing for something that like at that time, it sounded exciting. And I have a guess, like maybe at that time, like the whole field was a bit too focused on using adaptive computation for decreasing the cost on simple problems. But now we know that maybe we can actually use adaptive computation to increase the cost for hard problem. You know, it's actually like the other side of the same coin, right? So because at that time we were like, you know, like maybe, you know, like resource constraint and everything. So we were really thinking about why we were spending so much FOPs, like, you know, going through all the layers and like, you know, everything's for dot at the end of the sentence.

38:58If that token is like, do we really need like 24 layers? So how we can decrease that. But now we have a different perspective to that, which is like, you know, how can we increase this for a physics problem that we want to run the imprints for maybe like, you know, for two weeks. So that was really fun to work on that with these brilliant people. And I think just recursion in depth and reusing the parameters, or I've seen later some people actually framing it as negative sparsity, which is a great way of connecting it to a mixture of experts. In a mixture of experts, you have flops-free parameters, so parameters that they're not actually bringing any flops.

39:37And in looping, you have parameter-free flops, where you don't have extra parameters for the extra flops that you're throwing on this. So it goes the other direction of the sparsity. And it's quite effective. And I think people are picking it up, you know, so that we're seeing a lot of, you know, the excitement in this direction. Fascinating. Another fundamentally important contribution to the field that you did was the visual transformer paper in 2022. So the paper was called An Image is Worse, 16 by 16 Worse, Transformers for Image Recognition at Scale. Do you want to walk us through what that was?

40:12It's also a funny story for that. I got into vision and multimodality with that paper. So I've never worked on any vision problem. It was mostly because I was sitting next to people who were working on vision. So my desk was next to people who were working on vision. And that was the reason that I got interested. Because I was just talking to them. I was like, oh, this is actually interesting. And then I remember that at that time, I was working on externally, we call it palm. palm paper with like you know aconcha and other folks and i was like why we have 400 billion parameter language models but the biggest model that we have on the vision side is just like maybe 100 million like a res that like why like like why there's no benefit of a scale started looking into this with folks on like okay like maybe you know like there's something in transformer that actually kind of like you know make it a scale level and then maybe you know like we we can move away from convolution to try this and and at the end of the day i don't want to say that you know like that's the only way of scaling maybe you know if it could actually spend like enough time on convolution they can also make it a scalable and like you know like as good but there was also benefit of doing that simply because the rest of the that the machine learning field which was working on on language they were using this like a like architecture so they were building infra for it making it faster and and you know like like the sometimes the hardware is kind of like design based on this architecture, at least for short term.

41:39So we started pushing. And then I remember that, you know, we had a bunch of ideas that, okay, what if each pixel is a token? And then like the cost was going high. The context was just like getting super long. And then we had a lot of back and forth. And it's also quite funny because we started thinking about this problem, like from like very complicated point of view. So we were trying to mimic convolutions to be able to get this working. and it ended up like you know like i had a bunch of colleagues also in zirik and they started like trying the simple idea of what if we just like divide the image into patches of pixel you know 16 by 16 and then get each patch as a pixel and forget about you know like overlapping patches or you know like windows and stuff like that's it you know like you know like chop the image and then fit it to a transformer and then scale you know like like go with a lot of data and then like Let's start with something discriminating to train this model.

42:31And it worked. And it was also a little bit of a surprise for us that, oh, you know, they were all thinking about something fancy and very complicated, maybe in the integration of having convolutions and stuff. But something that worked was basically the simple idea of, you know, like patchify, fit it to transformer, scale it up, and then boom, you have a really, really good model for representation learning. Yeah. And to play it back at the highest level, that basically meant that you could apply a transformer architecture to image, where in the past you had two different families, you had the CNN world and the transformer world for text.

43:10And your breakthrough was to prove that transformer could scale equally well to images, which basically paved the way to a Gemini 3 today, which is like a natively multimodal model. Is that fair? Yeah, that is true. Yeah, so basically with that, we kind of took a step toward having also videos, like adapting transformers and audio, like adapting transformers. So basically, again, even if this is not the only architecture that would be in a multi-model, but it made it really simple to train these models natively, because you have a single architecture and you can have all the modalities during training.

43:46Great. So that's a perfect transition into your work, into Nano Banana and the future of image AI. So you are part of the Nano Banana team, which must have been so much fun when this came out and went just completely viral and what an incredible product. So since then, there's been a couple of releases. So there's been a Nano Banana Pro in November of 2025. And then just a few weeks ago, Nano Banana 2, a.k.a. Gemini 3.1 Flash Image, at the end of February. So a lot of people assume that image generation works as a translator, meaning that the AI reads the text of the prompt and then translates it into picture instructions and then draws it.

44:30But as we were saying, Gemini is natively multimodal. So how does that work? How does a model actually process the text and the pixels at the same time to build the image? I think the reason that maybe I got to the generation, okay, by the way, there's also one thing that I'm not an expert in image generation. Like when I started working on this, I remember I had like meetings with people and then they were talking about like, you know, computer graphic and all the like, you know, like old like ideas about or like intuitions. And I had like zero idea of what's going on. I was like, I know how to train a transformer and a scalete.

45:04And, you know, if it helps, I can basically contribute to this. But again, it was fun because I worked with a group of super smart, brilliant people with really, really good intuition. And I think the reason that I was excited about this was, this is maybe not super relevant to NonoBanana itself, but to just mention this, I was excited about the idea of positive transfer across modalities. So when you think about multimodal natively, one part of it is that, oh, you know, I'm adding capability to my model. So my model can understand images and understand videos and understand audio and text, but also can generate all these modalities.

45:46So I have a model that actually does all these together, right? This is for sure exciting from the product point of view. You have a model that is a great model for generating all these different outputs, and users are finding it very useful and interesting. But the most exciting part for me was, can I see a glimpse of transfer from these modalities? For example, if I train a model to become good at generating images, does it become also better at generating text? There are different intuitions of that. you know, like why this should happen. I think there's like something like, again, like very old in the literature on the linguistic side that they call it reporting biases, right?

46:31So like you, for example, you know, like visit your friend's place, right? And then you go to their place and then you see that they have a banana shape, like a sofa. And when you go home, the chance of talking about that sofa compared to a normal sofa is like much higher. So you can actually talk to your friend's partner later. Oh, you know, I went there and then the sofa was like in the shape of a banana, which was really fun but if it's like normal like you almost like it's weird if you go somewhere it's like oh by the way i went to my friend's place and they had a sofa which was like super normal so so this is this is the language reporting bias so language doesn't talk about things that are like at the middle of the distribution right but if you have an image or if you have like vision input from from anything in the world you have that information like there's no need for reporting it's just like there right so so because of that like picking up a lot of knowledge about the world through language is just not really efficient i don't want to say that it's impossible but it's not efficient you know like to learn about gravity if you kind of like you know have your model train on videos it's much easier to to get the model to learn about gravity because it just happens in a video than training your model on on all the textbook to kind of like learn about the concept of gravity, you know, or what is actually gravity.

47:46Is that a concept of world model that's built into the image representation? Exactly, exactly. So basically, you want a world model, basically, you know, like these models to be also like a world model. So you want these models to know about the world. There's a good chance that you can actually teach your model about the world just by presenting text to it, but it's just not efficient. And a good shortcut would be to bring multimodality into this. And the best way of learning about a modality is learning how to generate that, right? Like, so we got to this point that, okay, you know, we've been having Gemini generating like images from Gemini 1.

48:23So we like, like basically like Gemini was multimodal from day one. And the reason that we kind of first like released the image generation at like 2.5 instead of, you know, like Gemini 1, Gemini 1.5, Gemini 2, was that it was not great. And then like it really needed a push. And then we figured out that, okay, you know how to push this without like, you know, introducing any regression to other capabilities that the model has and, you know, like bring all of these natively into like this model. And that was like one side that was like super interesting for me. Like not sad views, but it's really hard to see positive transfer.

49:02So it turned out to be a really, really good model, but it was like really hard to see that, wow, you know, I train on images and then like text perplexity goes down. That was hard to see, you know, like the fact that, you know, you train in native model and it's good at like across all the capabilities is already impressive. But my hope is that, you know, multimodality and model is the way to really push multimodal training to enable like positive transfer across modality. I've worked with people that they were like expert on this you know for example one of the one of the things that I remember that you know at the beginning they were talking about this like visual quality and then I remember that you know it's like oh this model is a great model I send them them and there's like no this is not a good model I was like what do you mean and they started showing me two images that to my eyes they were like looking the same but they were saying like no this is like way better I was like no they're the same so so they had they had a good taste on grasping the visual quality of images.

49:56So working with them was like really interesting to kind of understand that, okay, there are dimensions. And by the way, like their intuition was the things that actually made like not a banana, like a success in terms of, you know, being a good product. But I was like, okay, what if we push this towards something beyond like traditional image generation? So instead of like a translator that, as you said, like a text image, it becomes a thinking machine about images. For example, you enable interleaved text image generation where the model can think in not only text token, but also in pixel space, right?

50:34So it generates text and then generates an image and generates another text, another image. And you can leverage that for different problems. Like one of them is that, oh, if you have, you know, like some sort of a story, right? Like, you know, like text of the story, image related to that text of the story, like, you know, like children's storybook, right? Another one, which I was actually really excited about was like this incremental generation. Like, let me just give you an example. So if you take like Dolly or Imagine or standalone image model, right? If you ask these models to generate an image of like, you know, a scene with 50 details, they might fit, right?

51:12And then someone can say that, oh, you know, okay, I can generate a better model that does up to 55 details. And then you say, okay, what about 60? And then I say, okay, let me just go back and trade it and then come back to you to cover your case. But at the end of the day, there's a threshold that these models can kind of like follow instruction to some extent about how many details that they capture from the text. But if you have incremental generation, so if you have text and then an image and text and image, you can get your model to generate these details one by one. So you never expect your model to generate an image, a perfect image in the first shot, right?

51:53So you expect your model to plan about this generation. so it says that oh you know let me start with big objects because you know later i'm going to have a hard time if i put like small objects and the big objects don't fit right so let me just do that and then like in the next turn i go with like medium objects and smaller and this is like super smart you know and you're never bothered by the capability of a single shot image generation because you you did planning and then you tune every step difficulty to to match the capability of your model to generate one shot. So that was also one of the things that, like, you know, Nano Banana and native generation, interleaved generation, kind of brought, like, a completely new perspective to image generation work, which is, like, a little bit, like, far from, like, you know, just translating text into an image.

52:44Fascinating. Does part of this contribute to efficiency? So especially Nano Banana 2, you have the flash aspect of this. So you're able to create amazing images very fast and apparently, I mean, seemingly very efficiently. So what's behind the scenes? Is that what you described? Is that MOE? How are you able to do that? First of all, I was involved in the original Nano Banana, Nano Banana Pro. And then the last version I gained, like, you know, because I jumped on the post-training encoding and agent and I find it exciting. This one, like the team actually shipped it. but if i want to say like like super high level that you know what is exactly the things that makes the model like uh like faster and more efficient part of it is just the size of the model so like nono banana was pro size and this one is just like flash so definitely that the parameter size like you know configuration of the moe and stuff the other one was uh people actually spent quite a lot of time on figuring out like you know uh um like nailing down like distillation recipe uh both on the side of like knowledge and like you know other like things that basically you kind of like need to distill to something like a process that is like, you know, lighter than the full process.

53:52Surprisingly, a lot of infra work for serving. So we have like really, really, really brilliant people that they're like, you know, serving engineers. And it's kind of impressive that like you sit on your desk and then like they come and they say, oh, by the way, like casually, I make a model 10X pasta. I was like, you know, like it's just like, and they just kind of saying it like, you know, in a very like in a cash office like wow this is like impressive we had also a lot of work on the on on optimizing the serving uh how to serve these models and and you know like again like because these models are operating differently from like just like normal net which model like they're not necessarily you know like um the same as you know next order prediction this is definitely something that you know a good serving engineer can figure out that okay you know i can think about a different way of doing that and we had also a lot of improvement on the efficiency side by by their work.

54:44All right. So as we get towards the end of this conversation, I thought it'd be fun to end with a few hot takes if you're ready for them. Yeah, absolutely. All right. What is one thing the AI field is getting wrong right now? Not easy to pinpoint specific things, but again, this is just my personal opinion. And maybe I have colleagues and other people sharing this with me. But I think we're underestimating how hard jagged intelligence is to fix.

55:17We're underestimating how much it matters. And we talk about almost like people laugh and go. If you have a model that does a very difficult math proof, but has difficult time counting letters in a war. as I said just people just laugh and move on but but I think it actually like pointing at something like deep and unresolved about these systems like the way that these systems kind of like represent and process knowledge and it's not a bug that you can patch so definitely you know like like we like we see that you know this is happening you know like people sometimes you know like or we have these problems that you know something is like awfully sad and then you can oh you know let me just like you know patch by adding something for the system instruction or developer instruction.

56:06A bit of a structural property of how these models actually learn. So I would say this is probably one of the things that we're not getting it like super right at this point. Great. What is one idea in AI research right now that is underrated? Something that is underrated, like you mentioned, continual learning. I think this is definitely underrated. As I said, you know, like sometimes the problem stays in the exploration mode until we are confident about something and then it goes to the exploitation mode. I think we are past the time that we really had to push this to the exploitation. So maybe like foundation models are essentially right now, like frozen in time.

56:44And like when the training ends, right? And then everything is like built on top of this frozen model, like in a rack pipeline and fine tuning workflows and retrieval system. And all these like elaborate like infrastructure is like all based on this assumption that these models are frozen. and it's a bit of a too much of a strong problem, there's an assumption to make and I think we are going to get to the point that we need to change these assumptions and maybe we need to think about it a little bit more actively and have pushing it toward something that we actually push it to productionization and maybe it's a little bit underrated right now, like the continual learning.

57:26So you think RAG goes away over time? It's not going to look like as is today. And it's going to be different, but like saying that it's going to go away completely, I'm not sure about that. And one of the reasons that I say that is RAG is not just about bringing like fresh information to the model when it wants to kind of like solve a problem about like the current state of things, but it also has this kind of like in-context learning. And there is a difference between in-context learning, like the information that you have in the context of the model compared to the information that you have in the weight of the model.

58:03Like, continuous learning and RAG are doing different things for bringing this fresh information. Maybe it changes in a way that, you know, like, it doesn't need to trigger RAG for, like, everything. But I'm pretty sure that there are going to be some tail of the distribution that we're going to do RAG still for it. You know, like, what's the time? All right. Last couple of hot takes. What do you think people are too confident about? So people think that, you know, pushing the technical side is sufficient. that if we just like get a model that is a smarter, everything is going to follow. And in my opinion, like a version of AI that is like, you know, really, really brilliant at like technical problems, but it has like a blind spot about, you know, everything else.

58:45And that version is not going to be able to actually create meaningful progress in the world. And the fact that, you know, people kind of assume and confident about like, like they're confident about this that that kind of like you know everything is gonna everything else is gonna follow or uh or or just everything else is just like a small list um i think it's wrong like you know we have governance we have like you know regulation we have social trust we have like for example distribution of access and the benefit and like in the world for this technology and even that like you know the institutional capacity to to kind of like absorb and adapt this technology is is just like this is something that maybe we don't have enough um like attention to and uh and and these are not really soft problems if not harder than the technical part they're they're really hard and like the pace of technical progress is definitely like like um currently running ahead of the world's capacity to to develop this kind of like you know mechanism uh and this gap is getting like you know bigger and bigger but what i'm saying is basically like you know the field needs to hold both things at once so maybe that's yeah that's one of the things all right and last one and i don't know if that's a hot take or maybe just advice for anybody entering the field today if you are going to start from scratch today what would you work on i don't want to start from scratch it's hard to start um i can tell you like you know there are two things that i think like you know would be nice to spend more time on it And there's one thing that I'm like very excited about it.

1:00:22I start from the things that I'm like really excited about it. And I would say like in a short term, it's like really exciting to push it. I am actually trying to even like, you know, be able to contribute to this direction. And that's like full automation of like super long horizon tasks. Things that, you know, like you have a machine working for maybe like two weeks, one month. Like the agents today are very impressive, right? And the demos are like very, like very marketable. But there's this confounding reliability problem that doesn't get hard enough. And for example, imagine if an agent has to take 100 sequential steps to complete a task.

1:00:58And imagine if each step has 95 % of success rate, which is great. Given the models that we have today, 95 % is really good. The probability of completing the whole task without a single failure is like 0.95 to the power of 100, which is like less than 1%. And this math is like brutal, right? Like, you know, and this like 95 per step, 95 % per step, as I said, is like very, very optimistic. Long horizon automation definitely isn't impossible, but it requires a level of per-step reliability and error recovery, and the current system maybe don't have it. And if we want social trust and basically having people really using it...

1:01:53At the end of the day, people don't experience average performance of these models. They experience the failures. If you have your model doing a dumb mistake, the damage in the trust that it makes is like bigger than like, you know, the benefit of getting 100 things right, right? Like 100 imperfect things right. So this like reliability in this like long horizon task is something that we definitely need. The kind of like side of, as I said, you know, two kind of like more philosophical high level things. I would definitely like work on grounding problem and how we can build AI system that are robust and connected to physical world.

1:02:27as I said, you know, like soon, like the concept of data, like, you know, how to kind of like, you know, enable these models to kind of like, you know, be very good at like self-improvement becomes, how can I ground these models in real world? So this is definitely like something that, you know, would be the bottleneck of self-improvement if we don't actively think about it. We should definitely move away from like, this is a statistical pattern in text and pixels. And the other things that is maybe kind of like, you know, related is if I'm thinking about you know a better definition of like intelligence itself right and a little bit like philosophical but um but it's a definitely a practical practical question uh and uh like the whole field and and us we're we're building uh like more and more like something that we haven't like to really define you know like like we're trying to kind of make these models smarter and more intelligent but like the definition of intelligent is just so hand baby and fuzzy that it's hard to actually kind of like you know measure the meaningful progress like which is related to your question that about you know like how about like evaluations it's good you know we have practices benchmarks scores capabilities and and even vibes you know which is i i find it like super useful but at the end of the day uh we really need a systematic way of uh maybe defining intelligence that that is hard and uh again like making progress based on what we have right now is good but at some point that becomes a little bit more important to really pinpoint that you know, what is the target and what is the goal and then push toward that, like with maximum speed.

1:03:58All right, Mustafa, it's been a absolutely fantastic conversation. Thank you so much for spending time with us. Really enjoyed it. Really appreciate it. Thank you. Yeah, thank you so much for having me. It was like fun to chat. And thanks for the invite. Hi, it's Matt Turk again. Thanks for listening to this episode of the MAD podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests.

1:04:29Thanks and see you at the next episode.

From the publisher

Are we truly on the verge of AI automating its own research and development? In this deep-dive episode of the MAD Podcast, Matt Turck sits down with Mostafa Dehghani, a pioneering AI researcher at Google DeepMind whose work on Universal Transformers and Vision Transformers (ViT) helped lay the groundwork for today's frontier models.


Moving past the hype, Mostafa breaks down the actual mechanics of "thinking in loops" and Recursive Self-Improvement (RSI). He explores the critical bottlenecks holding back true AGI—from evaluation limits and formal verification to the brutal math of long-horizon reliability.


Mostafa and Matt also discuss the shift from pre-training to post-training, how Gemini's Nano Banana 2 processes pixels and text simultaneously, and why the "frozen" nature of today's models means Continual Learning is the next massive frontier for enterprise AI and data pipelines.


(00:00) Intro

(01:17) What “loops” in AI actually mean

(05:04) Self-improvement as the next chapter of machine learning

(07:32) Are Karpathy’s autoresearch agents an early form of AI self-improvement?

(08:56) AI building AI: how close are we?

(10:02) The biggest bottlenecks: evals, automation, and long horizons

(12:36) Can formal verification unlock recursive self-improvement?

(14:06) What is model collapse?

(15:33) Generalization vs specialization in AI

(18:04) What is a specialized model today?

(20:57) Could top AI researchers themselves be automated?

(24:02) If AI builds AI, does data matter less than compute?

(26:22) Post-training vs pre-training: where will progress come from?

(28:14) Why pre-training is not dead

(29:45) What is continual learning?

(31:53) How real is continual learning today?

(33:43) Mostafa Dehghani’s background and path into AI

(36:13) The story behind Universal Transformers

(39:56) How Vision Transformers changed AI

(43:47) Gemini, multimodality, and Nano Banana

(47:46) Why multimodality helps build a world model

(52:44) Why image generation is getting faster and more efficient

(54:44) Hot takes

(54:53) What the AI field is getting wrong

(56:17) Why continual learning is underrated

(57:26) Does RAG go away over time?

(58:21) What people are too confident about in AI

(59:56) If he were starting from scratch today

More from The MAD Podcast with Matt Turck

All 44 episodes
AI is Already Building AIThe MAD Podcast with Matt Turck · 1 h 5 min
Listen in VO