DeepMind Gemini 3 Lead: What Comes After "Infinite Data"

18 Dec 2025 · 55 min · 26 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

DeepMind Gemini 3 “under the hood” and what comes after “infinite data,” including a shift toward a data-limited training regime, the role of pre-training vs post-training, and how DeepMind organizes research engineering.

Guest

Sébastien Bourgeau, Gemini 3 pre-training lead at Google DeepMind; researcher/engineer with a background in representation learning and large-scale LLM pre-training. Previously worked on DeepMind LLM efforts including Gopher, Chinchilla-related scaling work, and Retro (retrieval-augmented training). Grew up in Europe (Netherlands, Switzerland, Italy), studied at Cambridge.

Key claims

Gemini 3 gains come from many coordinated improvements by a large team (not one “secret knob”). DeepMind is building full systems around the model, not just training a network. Benchmarks are getting harder, and internal productivity gains are increasing. Synthetic data can help but must be validated carefully. Data is not “infinite”; research must adapt to finite/limited data.

Notable examples

MOE (mixture-of-experts) transformer-style architecture; Gemini is natively multimodal (single model for text/images/videos). References to Gopher, Chinchilla, Retro, ImageNet/data-limited analogy, and held-out evals to avoid contamination.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Shifting Paradigms in AI

0:00 to 0:26

Exploration of the transition from data unlimited to data limited regimes in AI.

“If I'm being honest with myself, I think we're ahead of where we could go.”

Gemini 3: A Simple Secret?

0:56 to 2:09

Discussion on the simplicity behind Gemini 3's advancements and the collaborative effort involved.

“So I was curious about your perspective.”

Progress Towards Intelligence

2:09 to 4:36

Examining how progress in AI benchmarks reflects improvements in actual intelligence.

“What does that tell us in terms of where we are in AI progress?”

Future of AI: Expectations and Predictions

4:36 to 5:48

Looking ahead at AI advancements, potential scientific discoveries, and societal impacts.

“I'm always curious, as an AI researcher who's so deep into the very heart of all of this, if you zoom out, are you still surprised by where we are?”

Automation and Acceleration of Research

5:48 to 7:35

How AI tools are changing the pace of research and enabling deeper exploration.

“Does that mean AI comes up with novel scientific discovery?”

Collaboration and Competition in AI Labs

7:35 to 9:32

Insights on how different AI labs are working together and the similarities in their approaches.

“The first part, I think, especially in the next year with more agentic workflows being enabled more and more, that should be able to really accelerate our work there.”

The Importance of Integrated Research and Engineering

9:32 to 12:01

Discussion on how research and engineering integration enhances AI model development.

“For the second question, I don't know if I have a good answer.”

Sébastien Bourgeau's Background

12:01 to 14:02

Exploration of Sébastien's personal journey and academic path leading to his role in AI.

“And Gemini 3 was trained on TPUs, right?”

Background and Education Journey

14:02 to 15:36

Learn about the guest's international upbringing and educational background in computing.

“So I was actually born in the Netherlands and I moved when I was seven to Switzerland.”

Path to DeepMind and Early Projects

15:36 to 17:49

Explore the guest's journey to DeepMind and their initial projects in deep reinforcement learning.

“So that's, again, that's, there's a bit of a lucky moment, I would say.”
Show all 26 chapters

Scaling Up LLMs: Chinchilla and Retro

17:49 to 19:57

Discover the projects Chinchilla and Retro, focusing on scaling models and data.

“So yeah, that was kind of my first project on that side and specifically LLMs and transformers.”

Understanding Research Taste and Trade-offs

19:57 to 22:48

Discuss the importance of research taste and navigating trade-offs in AI research.

“and how expensive it is to use the models once they're trained.”

Balancing Short-term and Long-term Research

22:48 to 26:13

Learn how research teams balance immediate improvements with exploratory projects.

“especially in deep learning, a negative result doesn't mean something doesn't work.”

Insights into Gemini 3 Architecture

26:13 to 28:00

Get an overview of the architecture behind Gemini 3 and its multimodal capabilities.

“And as promised, let's go fairly deep into Gemini 3, if you will.”

Understanding Multimodal Models

28:00 to 30:03

Learn about how Gemini 3 integrates different data modalities into one model.

“and that one gets linearly transformed again into the output of the dense block.”

The Role of Scaling Laws in Pre-Training

30:03 to 32:43

Discover the importance of scaling laws and architectural innovations in model performance.

“All right, let's talk about pre-training, since it's the area that you cover in particular.”

The Future of Data Usage in AI

32:43 to 34:45

Explore the evolving landscape of data sources, including synthetic data and its implications.

“I think you guys had a model card out for a bit that talked about some of this.”

Learning from Less Data: New Paradigms

34:45 to 36:44

Examine innovative approaches to model training in a finite data context.

“a lot of people were working on ImageNet and other benchmarks.”

Evolving Pre-Training Strategies

36:44 to 40:08

Understand advancements in model architecture and the role of evaluations in pre-training.

“as the previous model by training on less data.”

Alignment and Its Role in Model Training

40:08 to 42:01

Discuss how alignment is considered in both pre-training and post-training phases.

“And I think that's kind of RL scaling maybe starts that process, but I think there's a lot more to do also on the architecture side.”

Challenges of Data Contamination in AI Training

42:01 to 43:32

Learn about the risks of data contamination and the importance of model alignment during training.

“but very quickly they become contaminated.”

Understanding DeepThink and Model Processing

43:33 to 45:05

Explore how the DeepThink model processes information and generates answers.

“So first of all, is that a different model or is that part of the same model?”

Agentic Models and Pre-Training Insights

45:06 to 46:43

Discuss the implications of agentic models and how pre-training influences AI behavior.

“So being able to do screen understanding really, really well is critical.”

The Future of Continual Learning in AI

46:44 to 48:28

Understand the concept of continual learning and its potential impact on AI training.

“In some sense, this is also what Retro that we talked about was doing by retrieving data and then trying to externalize the knowledge corpus with the reasoning part.”

Current Trends and Future Directions in AI Research

48:29 to 50:35

Examine emerging trends in AI research, including long context architectures and efficiency in model use.

“So that's kind of the first thought that comes to my mind.”

Advice for Aspiring AI Researchers and Startups

50:36 to 52:57

Gain insights on the necessary skills and areas of focus for future AI researchers and startups.

“and systems aspects of the model research and not just the pure model architecture research.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If I'm being honest with myself, I think we're ahead of where we could go. We're not really building a model anymore. I think we're really building a system at this point. What might be happening instead is kind of a shift in paradigm where before we were kind of scaling in the data unlimited regime and we're kind of shifting more to a data limited regime, which actually changes a lot of the research and how we think about problems. I don't really see an end in sight for that kind of line of work to continue giving us progress. Hi, I'm Matt Turck. Welcome to the Mad Podcast. My guest today is Sébastien Bourgeau, great training lead on Gemini 3 at Google DeepMind.

0:33Sébastien is one of the top AI researchers in the world and a member of the Metis list. And this is a particularly special episode because it's his first podcast ever. We talked about how Gemini 3 is built under the hood, the shift from an infinite data world to a data-limited regime, how research teams at DeepMind are organized, and what's next for AI. Please enjoy this great conversation with Sébastien. Sébastien, welcome. Thank you. Hi, Matt. So I was hoping to start this conversation with this tweet from Oriel Vinyles, who is the VP of Research and Deep Learning at Google Demind, the Gemini co-lead, who said when Gemini 3 came out that the secret behind the model was remarkably simple, better pre-training and better post-training, which when you think about the leap that Gemini 3 represented over the prior state of the arts, sounds remarkably modest.

1:25So I was curious about your perspective. Is it as simple in some ways as that? Yeah, I'm not sure it's a big secret. At least from my perspective, this seems quite normal. I think people sometimes have the expectation that from one Gemini version to another, there's a big thing that changes and that really makes a big difference. In my experience, there's maybe one or two of those things that make a larger difference than other things, but it's really a combination of many, many changes and many, many things from a very large team that actually makes Gemini 3 so much better than the previous generations of Gemini.

2:01And I think this is probably a theme that will recur later, but it's really a large team effort that comes together in a release like Gemini 3. What does that tell us in terms of where we are in AI progress? What sounds from afar as in sort of turning some knobs gives us such a leap? What does that mean in terms of what we can expect going forward? There's two things. The first one is, it's still remarkable how much progress we're able to achieve in this way. And it's not really slowing down. There's so many of these knobs and so many improvements that we find on a day-to-day basis. Yeah, almost on a day-to-day basis that make the model better.

2:42So that's the first point. The second point is, we're not really building a model anymore. I think we're really building a system at this point. People have sometimes this view that we're just training a neural network architecture and that's it. But it's really the entire system around the network as well that we're building collectively. And so that's the second part. The big question on everybody's mind is what does that mean in terms of actual progress towards intelligence? And we don't need necessarily to go into the whole AGI thing because who knows what that means. But is the right way to think about this kind of model progress as an actual path towards intelligence versus trying to succeed on this benchmark or that other benchmark?

3:31What gives you confidence that the core model is getting smarter? The benchmarks definitely keep improving. And if you look at the prompts and how the benchmarks are set up, they are becoming increasingly difficult. And even for me, who has a background in computer science, some of the questions the model answers, it would take me a significant amount of time to answer. This is just one view. It's the benchmark view. And there's some amount of we evaluate those frequently, etc. We're being very careful about holding out the test set. But still, there's some fears often of overfitting to those and just benchmarking is what people call this.

4:07But that's one aspect. I don't think those fears are very founded. But the second aspect, and that's the one that really fills me with confidence, is the amount of time people spend using the model to make themselves more productive internally is increasing over time. Every new generation of models is pretty clear the model can do new things and help us in our research and our day-to-day engineering work much more so than the previous generation of models. So that aspect should give us confidence as well that the models are becoming more capable and actually are doing very useful things as well.

4:36I'm always curious, as an AI researcher who's so deep into the very heart of all of this, if you zoom out, are you still surprised by where we are? From your perspective, are we well ahead of where you thought we would be a few years ago? Are we on track? Are we behind, possibly? I think it's easy to say we're on track in hindsight. I think if I'm being honest with myself, I think we're ahead of where we could go. starting work on LLMs in 2019 or 2020. It's kind of hard to believe the scale of everything we're doing, but also just what the models are capable of doing today. If you kind of looked at scaling laws back then, they were definitely pointing towards that direction.

5:22And some people really believe those deeply. I'm not sure if I would have bet a lot on that actually materializing and being where we are today. So one interesting question that follows from this is where does that take us? If we assume the same kind of progress we've seen in the last five years, I think, yeah, it's going to be very, very cool what's going to happen in the next few years as well. What do you think on that front? Does that mean AI comes up with novel scientific discovery? AI wins the Nobel Prize? Where do you think we are going in the short term, like two to three years? I think, yeah, that's part of it.

6:05On the science side, DeepMind historically has done a lot of work. And for sure, there's a lot of work in that direction as well. I think we will be able to make some large scientific discoveries in the next few years. That's one side. I think on the other side is in my day-to-day work as well, both research and engineering. I'm very excited about how we can use those models to make more progress, but also to better understand the systems we're building and develop our own understanding and research further. Yeah, there's this big theme in the industry about automation of AI research and engineering, which if you extrapolate it leads into AI 2027 kind of scenarios where there's a discontinuity moment, just at a very pragmatic level.

6:52What does that mean using AI for your own work today? and what do you think that's going to mean in a couple of years? I think it's not so much about automation, but more about making us go faster and spending more of our time in the research part at slightly maybe higher level. A lot of the day-to-day work in research on language models is we're dealing with quite complex and large systems on the infrastructure level. So actually quite a bit of time is dedicated to running experiments, babysitting experiments, analyzing a lot of data, collecting results, And then the interesting part is forming hypotheses and designing new experiments.

7:30And so the last two parts, I think, is something where we'll be very much involved in. The first part, I think, especially in the next year with more agentic workflows being enabled more and more, that should be able to really accelerate our work there. Is your sentiment that the various frontier AI labs are effectively all working in the same direction, sort of doing the same thing? One fantastic, but in some way perplexing thing that we all experience as industry participant observers is this obvious phenomenon of like every week or other weeks or every month, there seems to be like another fantastic model and we completely spoiled.

8:14So Gemini 3 just came out at the same time, like two hours ago, literally before we were recording this. GBT 5.2 came out. What do you make of that from your perspective and how do you think that plays out? Is anybody going to break out or effectively the industry is going to continue with like the handful of top labs plus some new labs that are appearing? For the first question, there's definitely similarities between what the different labs work on. I think the base technologies are kind of similar. I might be surprised if we weren't all training transformer-like models, for example, in terms of the architecture side.

8:53But then there's definitely specialization, I think, happening on top of that and different branches in the tree of research that are being explored and exploited by the different companies. I think historically, for example, DeepMind has instilled, I think on the vision and multimodal side, we've been actually really, really strong. And that continues to be the case today. And then shows in both how people use the model, but also in the benchmarks, of course. And then, yeah, the things like reasoning, etc. OpenAI came up with the first model, but we also had a strand of research on that. So there's similarities, but it's not exactly the same, I would say.

9:34For the second question, I don't know if I have a good answer. One thing that's clear is to make progress on a model like Gemini today, you do need a very large team and a lot of resources. Now, that doesn't necessarily mean that what we're doing today is optimal in any form. And some disruptive research could definitely come along and allow a smaller team to actually take over in some form. This is one of the reasons why I actually enjoy being at Google so much as well. Google has this history of doing more explorative research and has a really high breadth of that research. and that continues to be the case, mostly in parallel to Gemini, but we're definitely able to also utilize that and bring some of those advances into Gemini.

10:18Are there other groups, whether at DeepMind or elsewhere in the industry, that are working in semi-secret or complete secret in post-transformers architecture that in one day something will come out and we'll all be surprised? Are there groups like that in the industry? I believe so. There's groups doing research on the model architecture side, for sure, within Google and within DeepMind. Whether that research will pan out, it's hard to say, right? It is research, so very few research ideas will help. And so in the meantime, the core advantage that one company may have over the other is just the quality of people.

11:01In the case of Google, I guess the vertical integration. That tweet from Aurel that I was mentioning got retweeted, a quote tweeted by Demi Sasabes, and he was saying that the real, real secret was a combination of research and engineering and infra. So is that the secret sauce at Google, the fact that you guys do the whole stack? It definitely helps. I think it's an important part. Research versus engineering is also interesting. I think over time that boundary has blurred quite a lot because we're working on these very large systems now. research really looks like engineering and vice versa.

11:37And I think that's a mindset that has really evolved over the last few years at DeepMind, especially where maybe there was a bit more of the traditional research mindset before. And now with Gemini, it's really more about research engineering. The infrastructure part is also very important. We are building these super complex systems. So having infrastructure that's reliable, that works, that's scalable, is key in terms of not slowing the research engineering down. And Gemini 3 was trained on TPUs, right? Not on NVIDIA chips. So it's truly integrated. So I'd love to do a deep dive on Gemini 3.

12:14But before we do that, let's talk about you a little bit. So you are the pre-training lead on Gemini 3. What does that mean? And then let's go into your background and your story. I'm one of the Gemini pre-training leads. So what this entails, it's a mix of different things. So part of my job is actual research, so trying to make the models better. But these days, it's less running experiments myself, but to help design experiments and then review results with people on the team. So that's the first part. The second part, which is quite fun, is more of the coordination and integration. So it's a fairly large team at this point.

12:54it's a bit hard to quantify exactly but maybe 150 200 people that work on a day-to-day on the pre-training side between data model infrastructure evals and so coordinating the work of all of these people into something that we can build together is actually quite complicated and takes quite a bit of time especially time to do well to me this is super important because actually being able to get progress out of everyone is really what makes us make the most progress rather than enabling maybe one or two or a small group of 10 people to run ahead of everyone else. That might work for a short period of time, but over longer periods of time, what's really been successful for us is being able to integrate the work from many, many people.

13:39So in terms of your personal background, I'm always curious, where did you grow up? What kind of kid and teenager were you? was trying to like reverse engineer, like, you know, those top AI researchers, where do they come from? And how did they become, why did you become to be who you are? I grew up a bit all over the place in Europe. I moved around quite a bit. So I was actually born in the Netherlands and I moved when I was seven to Switzerland. So my dad is from Switzerland and my mom is from Germany. So I did most of my school and the beginning of my high school in Switzerland, mostly in French and also in German in parts.

14:19And then at age 15, I think I moved to Italy where I finished my high school till around when I was 19. And at that point, I was going to go to the ETH in Zurich to do my studies. But I think just by random events one morning, I just looked up the top universities in some kind of ranking and I saw Cambridge was at the top. So I thought I'll just apply. Why not? And yeah, a few months later, I got the acceptance letter. So I decided to move to Cambridge, where I did my undergrad and master's in the computer lab. And you were growing up, you were just a super kind of math, strong, kind of a kid, computer science kind of kid.

15:01My dad has a technical background. So I remember when I was 10 or 11, starting to program a bit with him and learning. And I kind of always liked that. And then I always had like easiness in math and science at school. I remember never having to really study for math exams, but always doing quite well. That definitely changed at university. But that was, yeah, that was my high school experience. Great. And what was your path from school into where you are today? Yeah. So that's, again, that's, there's a bit of a lucky moment, I would say. One of the lecturers we had in my master's was someone who was also a researcher at DeepMind.

15:45And I just remember at the end of the last lecture, I was packing my stuff. I was like, you know what, I'll just ask him for a referral. What's the risk, right? You might just say no, but whatever. And so I actually took the courage and I went up to him and asked if he would give me a referral and sure enough, he was like, sure, send me your CV and I'll see what I can do. And that's kind of how I got my interview at DeepMind. This was in 2018. And so I joined DeepMind at the time, just DeepMind, not Google DeepMind, as a research engineer after university. And what did you do at first and how did that evolve to being one of the pre-training leads on Gemini 3?

16:25Yes, at the beginning having joined DeepMind and DeepMind being known for RL, the first project I managed to work on or decided to work on was something on DRL side. So specifically, we're training some unsupervised network to learn key points on Atari environments and try to get the agent to play Atari, right? So I did this for about six months, maybe. It wasn't enough, or in a sense, I didn't like the synthetic aspect of this. I always wanted to work more on real-world data and have more of a real-world effect. I think in general, I like to build things and build things that work. I don't really like the academic pure research part.

17:12And so that kind of drove me to start working on representation. So creating these or training these neural networks that have good representations to do different tasks. And one funny anecdote here is something I tell a lot of people on my team. But the first effort I joined on this was called representation learning from real world data. and at the time we had to add this from real world data to the name of the project because people would assume otherwise it would be synthetic environments or synthetic data and that definitely has shifted completely since then. So yeah, that was kind of my first project on that side and specifically LLMs and transformers.

17:54We are looking at architectures like Transformer and models like BERT and XLNet that were learning these representations and trying to improve those representations and do research on that side. Great. And then you worked on a retro, right? Do you want to talk about that? Yeah. So after that, we started working on scaling up LLMs and LLMs in general. So we started this work first on Gopher, which is, I think, the first deep-mined LLM paper that was published. So already at that point, it was a team maybe of 10, 12 people. So already at that point, it was pretty clear you needed, you couldn't just do that research on your own.

18:34And this is really where I started doing pre-training and pre-training at scale and developed my research taste, but also what I enjoy about this. So we trained the first dense transformer model. I think it was 280 billion parameters, I think 300 billion tokens at that time and trained that. And we would definitely not do things like we were doing them back in the day, but it was great and a very fun learning experience. After that, there were kind of two projects that emerged. The first one was Chinchilla and the second one, Retro. So in Chinchilla, we were re-examining how you should scale the model size and how you should scale the data, especially from a training compute optimal perspective.

19:26So the question is, you have a fixed amount of training compute. How do you train the best possible model? Should you increase your model size or should you increase your data size? And there was some previous work in this domain from OpenAI specifically that we re-examined and we actually found that you want to scale the data side much more quickly than what was thought before, rather than scaling the model side. Funnily enough, this is still really relevant in our day-to-day work today, especially because it has a lot of implications on the serving cost and how expensive it is to use the models once they're trained.

20:01So that was one side. The other line of work was more on retro and this is more on the architectural innovation side of things. So here we were looking at how you can improve models by giving them the ability to retrieve from a large corpus of text. So rather than having the model learn and store all the knowledge in its parameters, you give the ability to the model to look up specific things during training, but also during inference. You used the word research taste, which I think is super interesting. What does that mean? How would you define that and how important is that for a researcher?

20:40Yeah, it's very important these days and it's quite hard to quantify. But the few things that matter is, the first one maybe is your research is not standalone. standalone, this is what I was mentioning before, but your research has to play well with everyone else's research and has to integrate, right? So let's say I have some improvement on the model, but it makes the model 5 % harder to use for everyone else. This is probably not a good trade-off, right? Because you're going to slow down everyone else and their research, which would then commutatively slow down the durable research progress.

21:12That's the first thing. The second thing is being allergic to complexity, but complexity is quite subjective in terms of what people are familiar, but still, we have a certain, I think, budget of complexity we can use in a certain amount of almost research risk we can accumulate before things go bad. And so being aware of that and managing that is very important. So oftentimes, we don't necessarily want to use the best performance version of a research idea, but we'd rather trade off some of the performance for a slightly lower complexity version because we think that will allow us to do more and more progress in the future.

21:50So these are kind of the main two things, I think, around research taste. That's fascinating. And then presumably a part of it has to do with having an intuitive sense for what may work and not work, right? Given there's only so much compute you can use. Is that fair? Yeah, definitely. That's also an important part. I think that some people have that much more than others and a lot of experience really helps. But for sure, we are both select on the research side by compute. If we had a lot more computer, I think we'd make a lot more progress, a lot quicker. And so you have to guess to some extent what the right first, like which part of the tree of research tree you want to explore.

22:30And then within that, what are the right experiments? But then also knowing research always, most research ideas fail. And so you need to figure out at what point have I done enough in this direction to know to move on to something else? or should I keep pushing? And then the other interesting thing is, especially in deep learning, a negative result doesn't mean something doesn't work. It means you haven't made it work yet often. And so being aware of that as well is quite tricky. Since we're on this topic of research and how to organize research team to be successful, let's double click on some of this.

23:08So you mentioned trade-offs. Presumably one kind of trade-off is short-term versus long-term. How does that work? How do you all think about that? This is part of what I spend a lot of time thinking about as well. There's always critical path things to be done, or like this part of the model needs improving, or we know this part of the model is suboptimal. So we invest quite a lot in just fixing those immediate things. There's a few reasons for that. The first one is we know this will make the model better, so it's a fairly safe bet. But also we know that things that don't look quite... good or quite perfect often tend to have issues later, either when you scale up or when the model just becomes more and more powerful.

23:52And so actually really being very diligent about tackling those and fixing those is really important. So that's kind of the first part. The second part is slightly more exploratory research, so ideas that could land in the next version of Gemini or the version after that that have maybe a bigger effect on the model performance, but aren't quite validated. How we balance these is, I don't think I have a very clear answer. It's also a bit periodical. So when we're doing a scale-up, for example, there's often more, slightly more exploratory research because there's nothing right now that needs to be fixed in parallel.

24:27But just before we are ready to scale up a new architecture or a new model, it's very much like, let's de-risk the last pieces. It's very execution focused. How does that work a little bit, you know, in the same vein, the tension between research and product. So as we're discussing earlier, you're in this constant race with like other labs. And so is there presumably some pressure in like, oh, no, no, we need to, you know, have a better score or like win IMO or whatever it is. So like a very pragmatic, immediate product goal versus stuff that we know is going to improve the model over time. Like, how does that work?

25:09I guess it's just a variation of the same theme. This is why I like Google as well. There's actually very little of that, I think, because all of the leadership has a research background. They're very much aware that, yes, to some extent, you can force and accelerate specific benchmarks on certain goals, but in the end, the progress and making the research work is really what matters. So I personally, at least on a day-to-day, I never really feel that pressure. How is a team at DeepMind organized? So you mentioned pre-training as several hundred people, if I heard correctly. Is there like a post-training team?

25:46Is there like an alignment team? How does everyone work together? At a super high level. So we have a pre-training team, post-training team. On the pre-training side, we have people working on the model, on the data, the infrastructure, evals as well, very important. I think people often underestimate the importance on evals research and it's actually quite hard to do this well. And then yes, there's a post-training team and of course there's a large team working on infrastructure and serving as well. All right. Thank you for that. Let's switch tacks a little bit. And as promised, let's go fairly deep into Gemini 3, if you will.

26:24So Gemini 3 under the hood, the architecture, deep thing, pre-training, data scaling, all those good things. So starting at a high level on the architecture, so Gemini 3, as a devoted user, feels very different from 2.5. Was there a big architectural decision that explains the difference? And then how would you describe that architecture? At a high level, I don't think the architecture has changed that much compared to the previous one. It's more of what I was saying before, where a few different things come together to give a large improvement. At a high level, it's a mixture of expert architecture, transformer-based.

Read the full transcript

27:09So from that perspective, if you squint enough, you will recognize a lot of the original transformer paper pieces in that. Can you describe for people to make this educational what an MOE architecture is? At a high level, the transformer kind of has two blocks. So there's an attention block, which is responsible for mixing the information across times, across different tokens. And then there's the feedforward block, which is more about giving the memory, but also the compute power for the model to make these inferences. And those operate on a single token at the time. So they operate in parallel.

27:48So in the original transformer architecture, this is just a single hidden layer neural network. So it's a dense computation where the input gets linearly transformed into a hidden dimension. You apply some activation function and that one gets linearly transformed again into the output of the dense block. So that's the original paper. And then there's a lot of work before transformers as well on mixed-gen experts. And here the idea is you kind of decouple the amount of compute you use with how large the parameter is to use that. And so you dynamically route effectively to which expert you want the computational power to be used on rather than having that coupled.

28:30Gemini is natively multimodal. In practical terms, what does that actually mean for the model to think about text, images, or videos? Yeah, what this means is that there's no specific model trained to handle images and a different model trained to handle audio, a different model trained to handle text. It's the same model, the same neural network that processes all these different modalities together. Presumably, there is a cost aspect to this. Does being natively multimodal mean you're more expensive from a token perspective? Yeah, this is a really good question. There's kind of two costs to this.

29:17I would say that the benefits largely outweigh the cost here. And this is why we train these models. But the first cost is maybe less obvious to people, but it's this complexity cost and this research that I was talking about. Because you're doing a lot of more things, and especially different modalities interact in some ways, this can interact with different parts of the research and it has a complexity cost. So we have to spend time thinking about these things. The second cost is, yes, images are often larger in terms of input size than pure text. And so the actual computational cost is, if you do it naively, is higher.

29:58But of course, then there's interesting research to be done on how you make these things efficient. All right, let's talk about pre-training, since it's the area that you cover in particular. So starting with the high-level question, we mentioned, of course, the term scaling laws towards the beginning of this conversation. We talked about chinchilla a few minutes ago as well. In 2025, there was this much-discussed theme of the death of scaling laws, particularly for pre-training. Is Gemini 3 the answer that shows that all of this is not true and that indeed the scaling laws are continuing? Yeah, the discussions there to me were always slightly strange because my experience didn't match those.

30:50I think what we've seen is scale is a very important aspect in pre-training specifically and how we make models better. I think what's been the case, though, is that people overvalued that aspect. So it is a very important aspect, but it's not the only aspect. So scale will help to make your model better. And what's nice about scale, it does so fairly predictably. And that's kind of what the scaling laws tell us. As you scale the model, how much better will the model actually be? But this is only one part. The other parts are architecture and data innovation. These also play a really, really important part in the performance of pre-training and probably even more so than pure scale these days.

31:33But scaling is still an important factor as well. Right. And we're talking about pre-training specifically, right? Because this year we seem to have scaled RL in post-training and scaled test time compute, all the things. But for pre-training, you're seeing not only scaling loss, It's not slowing down, but you see some acceleration. Do I understand this correctly due to data and different architectures? I think the way to put this is these all compound. So scale is one axis, but this model and data also will make the actual performance better. And yes, sometimes the innovation part outweighs the benefits of scaling more.

32:18And sometimes just raw scaling is the right answer to make the model better. So that's on the pre-training side. And yes, on the RL and RL scaling side, I think we're seeing a lot of the same things we're seeing in pre-training or we saw in pre-training. What's interesting here is because we have the experience of pre-training, a lot of the lessons apply and we can reapply some of that knowledge to RL scaling as well. Speaking of data, so what is the pre-training data mix on Gemini 3? I think you guys had a model card out for a bit that talked about some of this. So what went into it? Yeah, it's a mix of different things.

32:58So the data is multimodal from the ground up. And yeah, there's many different sources that go into this. Another classic question in this whole discussion is, are we about to run out of data? So there's always the, do we have not enough compute? And the other question is, do we not have enough data? Clearly, there's been a rise in the usage of synthetic data this year. In your day-to-day work, or perhaps in general, where do you think synthetic data helps and where does it not help? Yeah, so synthetic data is interesting. You have to be very careful in how you use it because it's quite easy to use it in the wrong way.

33:44And what's often the case as well with synthetic data is you use a strong model to generate the synthetic data, and then you run smaller scale ablations to validate the effect of the synthetic data. But one of the really interesting questions is, can you actually generate synthetic data to make a model that you want to train in the future, which will actually be better than the model that generated the synthetic data in the first place? Can you actually make that one better as well? And so we spend a lot of time thinking about this and doing research in this direction. The other part of your question, are we running out of data?

34:19I don't think so. There's more. We're definitely working on that as well. But more than that, I think what might be happening instead is kind of a shift in paradigm where before we were kind of scaling in the data unlimited regime where data would scale as much as we would like. And we're kind of shifting more to a data limited regime, which actually changes a lot of the research and how we think about problems. But one good analogy of this is before LLMs, a lot of people were working on ImageNet and other benchmarks. And there was very in a very, very data limited regime as well. So a lot of techniques from that time start to become interesting as well.

35:00And perhaps that's one of those. And I don't know to which extent you can talk about it, if not talk about it in general. But there is this concept throughout the industry of training models based on reasoning traces. So basically forcing the model to show its work, how it got to a certain outcome, and then taking that to train the next model. Is that something that you do or that you think is interesting or a future direction? What is your perspective? Yeah, unfortunately, I can't comment on the specifics. This is how I'm asking the right questions. But maybe in general, is that something that people in industry do?

35:38I believe so. And this also falls into the previous question around synthetic data you were asking and kind of our approach to that is similar. And perhaps with that, taking this into a futuristic conversation, but like another big question and theme seems to be indeed, how can models learn from less data, which I think is what you were alluding to, talking about a data limited regime. Again, whether at DeepMind or in general, are you seeing interesting approaches to use the famous analogy, a model can learn like a child does? Just to maybe clarify what I said earlier, in a data-limited regime, I didn't necessarily mean with less data, but rather with a finite amount of data.

36:24So the paradigm shift is more from like we have infinite data to we have a finite amount of data. The second point is, in some sense, model architecture research is exactly what you mentioned. So when you make an improvement on the model architecture side, what it typically means is you get a better result if you use the same amount of data to train the model. but equivalently, you could get the same result as the previous model by training on less data. So that's kind of the first aspect of that. But it is true in terms of the volume of data needed today. We're still orders of magnitude higher than what the human has available to.

36:59Of course, there's the whole evolution process as well, which I find these high-level discussions quite hard to understand or follow because you have to make so many assumptions to convert that amount of data into what is today's pre-training data. but at least at first order, it does seem like we're using a lot more data than humans do. What other directions in overall pre-training progress are you excited about throughout the industry? Yeah, I think the one thing is in Gemini 1.5, I think we had a really good leap in the long context capabilities of the model. And I think that's really enabling the ability of models and agents today to do this work where you have maybe a code base and you do a lot of work on it, so your context length really grows.

37:45I think there's going to be a lot more innovation on that side in the next year or so to make long context more efficient, but also just to extend the context length of models themselves. So that's on the capabilities front, I think it's something where pre-training specifically has a lot to offer and is very interesting. Relatedly, I think for us, at least on the attention side, we've made some really interesting discoveries recently that I think will shape a lot of the research we do in the next few months. And I'm personally very excited about that. Again, I think I want to emphasize the point I made towards the beginning, but the way things work is it's really a culmination of many different things.

38:25So there's a lot of small, medium-sized things that we can already see coming up where I think we fixed this issue, we fixed this bug. This is an interesting research that shows promising things. And all of these things coupled, I think, will drive a lot of the progress again. It's interesting, you know, thinking about Retro that we talked about a bit earlier. You know, you're the co-author of Retro, which was about efficiency and like smaller models doing more. And now you are in the world of Gemini 3, which is like massive amounts of data and training in a very long context windows. Do you think that this paradigm of having, again, larger models, large context windows effectively obviates the need for kind of drag and search and that everything gets folded into the model?

39:16I mean, obviously, there's a corporate data part, but in general? There's some interesting questions here. So first of all, I think Retro was really about retrieving information rather than storing it, not necessarily about making models smaller. So it's about how we can use the model to do more reasoning already in a pre-training sense of reasoning rather than just store the knowledge. So this is still very much the aspect today. The interesting part is the iteration cycle maybe of pre-training used to be a lot slower than that of post-training until fairly recently. And so making these large changes on the pre-training side is quite costly in terms of risk and how long it takes.

40:00And then you have approaches like RAG or SEARCH, which you can do during post-training and iterate much more quickly on, which gives very strong performance as well. I think deep down, I do believe that the long-term answer is to learn this differentiable end-to-end way, which means probably doing pre-training or whatever that looks like in the future, learn to retrieve as part of the training and learn how to do search as part of the large part of training. And I think that's kind of RL scaling maybe starts that process, but I think there's a lot more to do also on the architecture side. But this is something that we'll see in the next few years and not immediately, I would say.

40:37The one thing I want to highlight is people often talk about model architecture, and that's definitely one part of what makes pre-training better. But there's other parts as well, infra and data and evals specifically, that don't always get the same mention. Evals specifically is extremely hard, and it's even harder in pre-training, I would say, because it kind of has these two gaps you need to close. close. So on the one side, the evals we use or the models we train regularly are much smaller and less powerful than when we scale up. So that means the eval has to be predictive of what the performance or have to still work for the large model and point in the right direction.

41:17So it has to be a good proxy on that side. And then there's a second gap as well, which is when we evaluate pre-training models, there's a post-training gap as well. So the way the models get used is they don't just get used after pre-training, there's more training happening after. And so the evals we use in pre-training or on pre-trained models have to be good proxies of what happens after as well. And so making progress on evals is really important and quite hard and has also driven a lot of the progress we have in terms of being able to measure what an actual improvement is on the model or on the data side.

41:49And evals are the mind that's all internally built, like you have your own set of evals? Yes, to a large extent and more and more so because what we found is that external benchmarks, then you can use them for a little while, but very quickly they become contaminated. So they start to be replicated on different forms or different parts of the web. And then if we end up training on those, it's really hard basically to detect leaked evals. So the only way you really have to protect against cheating yourself and thinking you're doing better than you are is by actually creating held out ELLs and not really keeping them held out.

42:27In the same vein, is alignment a part of what you all think a lot about at the pre-training level? Or is that more of a post-training kind of conversation or both? It's a majority of post-training, I would say, but there's definitely some parts of it which are relevant of pre-training. I can't go into too many details here, but some parts are relevant to pre-training and we do think about that as well. And at a very simplistic level, I always wonder, again, in the context of Gemini or otherwise, if the core data set is the internet, there's a lot of terrible things on the internet. Is alignment 101 that there's stuff that you just do not include in the model?

43:06This is an interesting question and I don't think I have a definitive answer, but you don't want the model to do these terrible things. So at a fundamental level, you do need the model to know about those things. So you have to train a bit at least on those so that they know what their things are and know to stay away from those. Otherwise, when a user would mention something terrible, the model wouldn't even know what it's talking about. And they might not be able to say this is something terrible. Let's talk about DeepThink, the thinking model that was released a few days after Gemini 3. So first of all, is that a different model or is that part of the same model?

43:43How should one think about it? I'm not allowed to. I can't comment too much. What happens when the model thinks and you wait for 10 seconds or 20 seconds or whatever time? What happens behind the scenes? I think this has been covered quite a bit in some of your previous podcasts as well. It's about generating thoughts. And so rather than just doing compute in the depth or in the model side, you also do compute and allow the model to think more on the sequence length side of things. So the model actually starts to form hypotheses, test hypotheses, invoke some tools to validate the hypotheses, do search calls, etc.

44:28And then at the end, be able to view the thought process to provide a definite answer to the user. The industry has normalized around that paradigm of chain of thought. That's for, yeah. Can you talk a little bit about the agentic part of this and Google Antigravity? What do you find interesting about it? What should people know about it? Yeah, this is, I guess, what I was mentioning before around my own work, especially. I think that's interesting. A lot of the work we do on a day-to-day basis is more execution-based, babysitting experiments, etc. And I think this is where I at least see the most impact from those.

45:04Bringing it back to the topics of pre-training, I think that the perception and vision side is very important for this because now you're asking models to interact with computer screens. So being able to do screen understanding really, really well is critical. And so that's an important part on the pre-training side, at least. And in anti-gravity, there's a whole vibe coding aspect, truly vibes in that you don't even really see what happens when you ask. Is Vibes, same question, is that a pre-training thing? Is that just a post-training thing? How do you build Vibes into a model? Yeah, this is interesting.

45:42I think you can probably ask five different researchers and you'll get five different answers. There's also this notion of large model feel. People call this, especially I think GPT 4.5, historically had some of this presumably where larger models maybe feel differently. I wouldn't actually just, I wouldn't put it in these terms specifically, but I think Vibes comes down to this and actually pre-training probably plays a larger role today in some of that and how the model feels and in general than post-training. I think this is, yeah, this is in general for Vibes coding specifically, I think that's maybe more of an RL scaling and post-training thing where you can actually get quite a lot of data and train the model to do that really well.

46:26So zooming out a little bit, maybe for the last part of this, conversation i'm curious about where things are going in in general there was a uh a key theme discussed at neurops uh this year around continual learning and i'm curious about your perspective especially from a pre-training perspective right because we are in this paradigm where like every few months or years we and by we i mean you train a uh a very large new base model first of all what is continual learning and two how does that impact uh retraining if continual learning becomes a thing yeah i guess continual learning is about um updating the the model with new knowledge as as new knowledge is discovered right let's say a new scientific breakthrough is made tomorrow the base model we trained yesterday you wouldn't actually know about it in its pre-training first i think a lot of progress has been made on this front since in the last few years i think this is mostly around post-training, around search, use search tools and make search calls, then they would have access to that new information.

47:33In some sense, this is also what Retro that we talked about was doing by retrieving data and then trying to externalize the knowledge corpus with the reasoning part. So that's the first part. I think the second part is on the pre-training side specifically is what I was mentioning on long context as well. And one way of doing this is if you can keep expanding the context of the user, the model keeps getting more and more information in that context. And so you kind of have this continual learning aspect part of that. But then, of course, there's more of a paradigm shift. Maybe this is what people discuss is, can you change the training algorithm such that you can continuously train them on a stream of data coming from the world, basically?

48:17Beyond the continual learning, what do you think is hot slash interesting or intriguing in current research today? Yeah, there's a lot of, again, there's a lot of small things right now that accumulate. So that's kind of the first thought that comes to my mind. And that historically has really driven progress. I wouldn't just bet against that continuing to drive progress. The things I mentioned before around the long context architecture and long context research is one aspect. I think on the attention mechanism as well on the pre-training side. And then this paradigm shift from infinite data to the limited data or finite data regime is something as well, I think, where a lot of things will change.

49:01And there's a lot of interesting research. That's kind of on the pre-training alone side. The other side, which is quite interesting today, is these models become the amount of people using these models is growing quite rapidly. And so more and more what we have to think about on the pre-training side as well is how expensive is the model to use to serve and have really deployed at a large scale. And what things on the pre-training side specifically can we do to make this model have better quality and maybe be cheaper to serve and consume fewer resources during inference. For any student or like PhD student listening to this, if they want to become you in a few years, what problems do you think they should think about or focus on that's not, you know, like a year or two out, but like more interesting sort of a few years out?

49:55One thing that's becoming increasingly important is being able to do research, but being aware of the system side of things. So we're building these fairly complicated systems now. So being able to understand how the stack works all the way down from TPUs to research is kind of a superpower because then you're able to kind of find these gaps in between different layers that other people weren't necessarily able to see, but also to reason through the implication of your research idea all the way down to the TPU stack. And people that can do that well, I think have a lot of impact in general. So in terms of like specialization, it's really thinking about this research engineering and systems aspects of the model research and not just the pure model architecture research.

50:43That's one. I think personally, I still have a lot of interest in kind of this retrieval research as well that we started with retro. And I think it wasn't quite ripe until now, but the things are changing. And I just think it's not unreasonable to think in the next few years, something like that might actually become viable for a leading model like Gemini. And why was it not ripe and why may that change? I think that's around the complexity side of things I was mentioning. And also the fact that all the capabilities it brings, you can iterate much more quickly in post-training. So what I was saying with search and post-training data, you can give very similar capabilities to the model in a much simpler way.

51:28And as post-training grows and RL scaling grows as well, maybe that shifts again towards more on the pre-training side. Do you think there are areas of AI right now that are over-invested in, where there's a disconnect between what makes sense and where the industry is actually going and investing dollars in? I think it's got a lot better. I think maybe two years ago, what I was seeing is people were still trying to very much create specialized models to solve tasks that were maybe within half a year or a year of reach of generalist models. And I think people have caught up to that much more. And now I believe that for generalist tasks or tasks which don't require extreme specialized models, trying to use a generalist model and maybe not the current version, but the next version might be able to do that.

52:20and then so what that means is research in terms of how you use models and the harness etc is becoming increasingly important and also how you make models and these harnesses more robust to making errors and recover from such errors. Yeah in that vein do you have any advice or recommendation for startups right so seen from the perspective of a founder or the the VCs who love them. There is this feeling that the base models are becoming ever so powerful and then trained on multiple data sets. So it used to be the model is able to converse, but now it's able to do financial work and cap tables and that kind of thing which seems to shrink the area of possibility for startups.

53:07Do you have thoughts on that? Yeah, I think so maybe you have looked at what models were able to do a year or a year and a half ago and then look at what models are able to do today and try to extrapolate that. I think the areas where the models are improving, I think will continue to improve. And then there's maybe some areas where there's not been that much progress and that might be more interesting areas to do research. I don't really have a specific example in mind right now, but that would be the general advice. What are you excited about for the next year or two in terms of your personal journey?

53:40What I like very much about my day-to-day is working with many people and being able to learn from a lot of researchers. So that's what drives me to a large extent. Every day I come to work and I talk to really, really brilliant people and they teach me things that I didn't know before. And so I really like that part of my job. What I was saying multiple times at this point, but there are just so many different things that will compound and different things where there's headroom to improve. I'm really curious because right now I don't really see an end in sight for that kind of line of work to continue giving us progress so actually being able to see this through and see how far this can take us is really interesting at least for the next year or so I don't see this slowing down in any way Great, well that feels like a wonderful place to live it Sebastian thank you so much for being on the pod really appreciate it that was fantastic thank you Thank you very much Hi, it's Matt Turk again thanks for listening to this episode of the MAD podcast.

54:40If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.

From the publisher

Gemini 3 was a landmark frontier model launch in AI this year — but the story behind its performance isn’t just about adding more compute. In this episode, I sit down with Sebastian Bourgeaud, a pre-training lead for Gemini 3 at Google DeepMind and co-author of the seminal RETRO paper. In his first-ever podcast interview, Sebastian takes us inside the lab mindset behind Google’s most powerful model — what actually changed, and why the real work today is no longer “training a model,” but building a full system.


We unpack the “secret recipe” idea — the notion that big leaps come from better pre-training and better post-training — and use it to explore a deeper shift in the industry: moving from an “infinite data” era to a data-limited regime, where curation, proxies, and measurement matter as much as web-scale volume. Sebastian explains why scaling laws aren’t dead, but evolving, why evals have become one of the hardest and most underrated problems (including benchmark contamination), and why frontier research is increasingly a full-stack discipline that spans data, infrastructure, and engineering as much as algorithms.


From the intuition behind Deep Think, to the rise (and risks) of synthetic data loops, to the future of long-context and retrieval, this is a technical deep dive into the physics of frontier AI. We also get into continual learning — what it would take for models to keep updating with new knowledge over time, whether via tools, expanding context, or new training paradigms — and what that implies for where foundation models are headed next. If you want a grounded view of pre-training in late 2025 beyond the marketing layer, this conversation is a blueprint.


Google DeepMind

Website - https://deepmind.google

X/Twitter - https://x.com/GoogleDeepMind


Sebastian Borgeaud

LinkedIn - https://www.linkedin.com/in/sebastian-borgeaud-8648a5aa/

X/Twitter - https://x.com/borgeaud_s


FIRSTMARK

Website - https://firstmark.com

X/Twitter - https://twitter.com/FirstMarkCap


Matt Turck (Managing Director)

Blog - https://mattturck.com

LinkedIn - https://www.linkedin.com/in/turck/

X/Twitter - https://twitter.com/mattturck


(00:00) – Cold intro: “We’re ahead of schedule” + AI is now a system

(00:58) – Oriol’s “secret recipe”: better pre- + post-training

(02:09) – Why AI progress still isn’t slowing down

(03:04) – Are models actually getting smarter?

(04:36) – Two–three years out: what changes first?

(06:34) – AI doing AI research: faster, not automated

(07:45) – Frontier labs: same playbook or different bets?

(10:19) – Post-transformers: will a disruption happen?

(10:51) – DeepMind’s advantage: research × engineering × infra

(12:26) – What a Gemini 3 pre-training lead actually does

(13:59) – From Europe to Cambridge to DeepMind

(18:06) – Why he left RL for real-world data

(20:05) – From Gopher to Chinchilla to RETRO (and why it matters)

(20:28) – “Research taste”: integrate or slow everyone down

(23:00) – Fixes vs moonshots: how they balance the pipeline

(24:37) – Research vs product pressure (and org structure)

(26:24) – Gemini 3 under the hood: MoE in plain English

(28:30) – Native multimodality: the hidden costs

(30:03) – Scaling laws aren’t dead (but scale isn’t everything)

(33:07) – Synthetic data: powerful, dangerous

(35:00) – Reasoning traces: what he can’t say (and why)

(37:18) – Long context + attention: what’s next

(38:40) – Retrieval vs RAG vs long context

(41:49) – The real boss fight: evals (and contamination)

(42:28) – Alignment: pre-training vs post-training

(43:32) – Deep Think + agents + “vibe coding”

(46:34) – Continual learning: updating models over time

(49:35) – Advice for researchers + founders

(53:35) – “No end in sight” for progress + closing

More from The MAD Podcast with Matt Turck

All 44 episodes
DeepMind Gemini 3 Lead: What Comes After "Infinite Data"The MAD Podcast with Matt Turck · 55 min
Listen in VO