In short
```markdown
Podcast Episode Summary
Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Overview In this episode of the Latent Space podcast, host Yi Tay discusses his journey at Google DeepMind, focusing on the development of Gemini, the IMO Gold achievement, and deep reinforcement learning (RL) techniques. The conversation sheds light on various aspects of AI research, team dynamics, and the ongoing evolution of model training methodologies.
Key Topics Discussed
Yi Tay's Career Path
- Journey: Transition from Brain to Reka, and then back to Google DeepMind to lead the Reasoning and AGI team in Singapore.
- Gemini Deep Think and IMO Gold: Leading model training for Gemini and the innovative approach taken to achieve IMO Gold.
IMO Gold Story
- Team Dynamics: Collaborative effort with co-captains across different time zones (Singapore, London, Mountain View).
- Live Competition Experience: Challenges faced during the live IMO competition, including real-time problem-solving and uncertainty in scoring.
AI Model Development Decisions
- Abandoning AlphaProof: Strategic decision to move away from symbolic systems in favor of end-to-end Gemini with RL, highlighting a bold approach.
- On-Policy vs. Off-Policy RL: Explanation of the difference between on-policy RL (learning from its own actions) and off-policy RL (learning from others' actions).
Key Concepts in AI Research
- Self-Consistency and Parallel Thinking: Importance of sampling multiple outputs and using majority voting techniques for improved reasoning.
- Data Efficiency Frontier: Discussion on the efficiency of models learning from less data compared to human learning.
- World Models: Examination of different schools of thought regarding world models and their implications in machine learning.
Insights on AI Coding and Generative Retrieval
- AI Coding Assistants: Yi's personal experiences and how AI coding tools have become more effective in assisting with programming tasks.
- Generative Retrieval: A new approach to search and recommendation systems using semantic tokens, with success stories from YouTube and Spotify.
Building the Team in Singapore
- Hiring Philosophy: Focus on recruiting talented researchers in RL and reasoning, looking for either strong track records or exceptional achievements in competitive coding.
- Geographical Considerations: Discussion on the importance of location in terms of talent acquisition and team collaboration.
Health and Productivity
- Personal Health Journey: Yi shares his fitness journey, emphasizing the connection between physical health and intellectual productivity.
- Motivation: Strategies for maintaining motivation and energy levels for optimal performance in AI research.
Key Takeaways
- The importance of collaboration in AI research, particularly in global teams with diverse expertise.
- The strategic decisions in AI model development can lead to significant advancements, as seen in the IMO Gold achievement.
- Health and well-being are crucial for sustained productivity and creativity in high-demand environments like AI research.
- Generative retrieval and improved coding assistants represent the future of AI applications in solving complex problems.
Conclusion This episode provides deep insights into the world of AI engineering, particularly in the context of reinforcement learning, team dynamics, and personal well-being. Yi Tay's experiences at Google DeepMind offer valuable lessons for aspiring AI researchers and engineers.
For full show notes and more episodes, visit [Latent Space](https://latent.space). ```
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VORejoining GDM and Team Structure
0:45 to 2:37
Discussion about Yi Tay's return to GDM and the structure of the new team in Singapore.
“And then you joined GDM again, working for Quark again.”
Reflections on Google Infrastructure
2:37 to 4:47
Yi reflects on his experiences returning to Google and the changes that occurred during his absence.
“So I think that obviously a lot of things have changed, but I think overall, the coming back has been pretty seamless.”
Transitioning to Reinforcement Learning
4:47 to 6:27
Yi discusses his shift from model research to focusing on reinforcement learning.
“Basically correct your own path instead of trying to imitate other people's path.”
On-Policy vs Off-Policy Learning
6:27 to 9:22
Exploration of on-policy and off-policy learning concepts and their applications in AI models.
“humans, we are more on policy because we go around the world, we make mistakes and then we are, okay, this is, but like imitation learning is possibly somebody else.”
Learning from Mistakes: Human vs AI
9:22 to 11:43
Discussion on how humans and AI learn from mistakes and the implications for model training.
“I was watching GANs and I was like, okay, this is cool.”
The Importance of High Learning Rates
11:43 to 13:48
Yi emphasizes the need for high learning rates in adapting to new paradigms and shifting mental models.
“It's a more nuanced version of, if we talk to Danny, it's not majority voting for sure, but it's more.”
Launching the IMO Initiative
13:48 to 14:01
Yi shares the story behind the IMO initiative and the challenges faced during its implementation.
“I covered the silver result from last year.”
Exploring IMO Gold and Model Training
14:01 to 16:20
Learn about the challenges and processes involved in training AI models for competitions like IMO.
“I personally was believed, always believed in, if we are not, like in retrospect, it's easy to say this, but it's a bit like, if the model can't get to IMO gold, then can we get to AGI?”
The Role of Specialized Systems in AI
16:21 to 19:37
Discover the debates surrounding specialized systems versus general models in AI development.
“But basically you ROFT the lean verifier into the chain of thought.”
Team Dynamics and Collaboration Challenges
19:38 to 21:42
Understand how team collaboration worked across different time zones in the IMO project.
“But my prediction is that I think most things can be subsumed by the model.”
Show all 54 chapters
Insights from the IMO Competition Process
21:43 to 24:16
Gain insights into the competition process and expectations during the IMO event.
“I was not directly involved in the ICPC, but I was related to Science 10.”
AI Progress in Recent Years
24:17 to 25:20
Reflect on the rapid advancements in AI and its implications for the future.
“they would have said that could be possible.”
Challenges of Inference Optimization
25:21 to 28:00
Learn about the complexities and challenges related to inference optimization in AI models.
“The hard parts you would say were scaling inference.”
Flaws in AI Spatial Reasoning
28:00 to 28:40
Explore the limitations of AI in spatial reasoning and model capabilities.
“And it showed serious flaws in Anthropix's screen understanding vision capabilities.”
Challenges in Completing the Pokédex
28:40 to 29:30
Discuss the complexity of completing the Pokédex and its implications for AI.
“I think solving Pokemon is going to be more of like how fast you solve it.”
AI's Ability to Synthesize Knowledge
29:30 to 30:40
Investigate AI's capacity to synthesize web knowledge for game tasks.
“analogy would be once, if let's say we have a model that is capable of doing that, what can we make it do that we cannot do today?”
Innovating Novel Knowledge with AI
30:40 to 32:20
Examine the potential for AI to generate innovative knowledge and insights.
“The task really is, can you look up the guide to do it?”
Demystifying AI Reasoning
32:20 to 34:30
Clarify the concept of reasoning in AI and its implications for models.
“One related question on that, which I think is related to the Danny paper, which is I think people have this sort of mythicism on what reasoning is.”
The Impact of Reasoning Tokens in AI Training
34:30 to 36:30
Discuss the integration of reasoning tokens in AI training corpuses.
“The question I was kind of leading into was that because there's now we're in this age where LLM text is in the corpus of stuff that we train on, where it's a little bit of a recursive loop, right?”
The Evolution of AI Coding
36:30 to 38:10
Explore how AI coding is transforming tasks and productivity for researchers.
“oh, you just include more coding tokens in your pre-trained corpus.”
Trusting AI in High-Expertise Tasks
38:10 to 40:00
Reflect on the implications of trusting AI for expert-level work.
“There were so many moments this year where AI suddenly crossed that like that immersion thing.”
AI as a Support Resource
40:00 to 42:00
Consider the role of AI as a support tool in various job functions.
“I don't see how anybody can disagree with that.”
AI as a Support Role in Problem Solving
42:00 to 43:20
Explore how AI can assist in problem-solving while highlighting its limitations.
“I often think of myself as a bard because I tell stories and I plus everybody around me.”
Incremental Progress in AI
43:20 to 45:20
Discussion on how small improvements in datasets can lead to significant advances in AI capabilities.
“just trust the process, keep scaling in every dimension and things will just fall away, things will emerge.”
The Role of Attention Mechanisms
45:20 to 47:20
Examine the importance and limitations of attention mechanisms in AI architectures.
“So attention and transform is like a core idea in the recent times.”
Future of AI Architectures
47:20 to 49:20
Discussion on potential changes and developments in AI architectures towards AGI.
“processing level like there is the ultimate sequence to sequence transformer.”
Scaling AI and the Importance of Ideas
49:20 to 51:00
The conversation centers around the importance of innovative ideas versus mere scaling in AI development.
“And this, like, continuing learning this, like, there's many ways to think about processing many, like, insanely large, like, contexts, right?”
Exploring Research and Development Gaps
51:00 to 53:00
Discussing the increasing gap between closed and open labs in AI research.
“I see what you're saying, but let's call it Gen AI.”
Data Efficiency and Training Paradigms
53:00 to 56:00
Investigating the limits of data efficiency in AI training and the implications for future models.
“So I think the ideas are still very important and there's still very good ideas that are game changers that are being invented.”
Data Efficiency in AI Models
56:00 to 56:40
Discussing the nuances of data efficiency and its implications for AI training.
“of data efficiency of a model in terms of training and compression should be.”
Learning from Limited Data
56:40 to 57:10
Exploring how AI can optimize learning from small datasets and the importance of varied data points.
“Like given limited data set, how well can you learn from it?”
Pre-training Models Debate
57:10 to 58:00
Engaging in a discussion about the relevance of pre-training models in AI advancement.
“like extracting more from varied data points is definitely valuable, but I think that also relates to the fact that we're running out of tokens in the world.”
Human vs Machine Learning Efficiency
58:00 to 59:10
Comparing data efficiency in humans and machines, and what AI can learn from this.
“I think that people are taking it seriously.”
World Models and Data Efficiency
59:10 to 1:00:40
Discussing the concept of world models in AI and their relationship to data efficiency.
“Yeah, the only question is that where does this thing come from?”
Learning Algorithms and Efficiency
1:00:40 to 1:01:40
Examining the significance of learning algorithms in improving data efficiency.
“for more efficient learning, which potentially might not take the form of a sequence-to-sequence transformer.”
Understanding RL Environments
1:01:40 to 1:02:30
Analyzing the investment in RL environments and their value in AI training.
“You model everything with some kind of Gaussian splats or whatever, and you inhabit that 3D space.”
The Value of RL Environments
1:02:30 to 1:04:30
Considering the reasons behind the high demand for RL environments in AI.
“The LM actually learns to predict those things.”
Data Efficiency and Future Trends
1:04:30 to 1:05:30
Forecasting future trends in data efficiency and AI learning strategies.
“Because I think there's this, there's a different scaling law for like humans is this, machines is this, like dogs are this, cats are this, there's this different, nice exponent.”
The Importance of Ranking in Retrieval
1:05:30 to 1:07:20
Understanding the significance of ranking and retrieval in AI systems.
“So, but I agree that data efficiency is very important given that I think we're going to like, there's a limited amount of data in the world.”
Advancements in AI Retrieval Systems
1:07:20 to 1:10:00
Discussing the evolution of retrieval systems in AI and their practical applications.
“Then you can probably do it yourself internally, right?”
Reimagining Document Retrieval
1:10:00 to 1:10:49
Explore the evolution of document retrieval methods and the significance of semantic IDs.
“So at that time, we wanted to like reimagine retrieval and such.”
Generative Retrieval Development
1:10:50 to 1:12:01
Learn about the development and application of generative retrieval in various systems.
“just trying to give a document like identifier and just predicting like raw brute force predicting like this.”
Challenges in Information Retrieval
1:12:02 to 1:13:11
Understand the unique challenges and perceptions of working in the information retrieval field.
“So over time, I also left Google and stuff.”
The Role of AI in Search and Recommendations
1:13:12 to 1:14:16
Discover how AI is transforming search and recommendation systems beyond traditional methods.
“Under the general principle of apply ML to everything, if the Googler is working on generative retrieval, would that be like AI overviews?”
Modeling Dynamics in Information Retrieval
1:14:17 to 1:15:08
Examine the dynamics of modeling in information retrieval and the differences from language tasks.
“that normally like no classical system can ever handle.”
Experience with Retrieval Systems
1:15:09 to 1:16:23
Hear personal insights on the experiences and feelings associated with working on retrieval systems.
“Like when we kill climate LOM, when you train models, like the way that modeling things interact with this environment is very different.”
The State of the IR Community
1:16:24 to 1:17:51
Discuss the current state and challenges faced by the information retrieval research community.
“for Rx's and stuff like that, I didn't enjoy it.”
Organizing GDM Singapore Symposium
1:17:52 to 1:19:51
Learn about the motivations and outcomes of organizing a key AI symposium in Singapore.
“Like, I work on it as a side quest thing, yeah.”
Global Collaboration in AI Research
1:19:52 to 1:21:40
Explore the impact of geographical factors on AI research and collaboration.
“I think it's possible for Singapore to be close to the frontier.”
Cultural Influence of AI in Different Cities
1:24:00 to 1:24:50
Explore how AI culture permeates different global cities and its impact on research.
“like mental space and energy to have some other culture and then like, you know, like London, Singapore, New York, they have their own culture, right?”
Hiring Trends in AI Research
1:24:50 to 1:26:30
Learn about the qualities and skills sought in AI candidates, especially in reinforcement learning.
“So I think that's something that we're hiring for now.”
Student Success on a Budget
1:26:30 to 1:27:40
Discover how students can achieve significant research results even on limited resources.
“And that was done on the student budget.”
Research Taste and Career Development
1:27:40 to 1:28:40
Understand the importance of research taste and personal initiative in advancing an AI career.
“because it's not like their professors know what to work on either.”
The Challenge of Proving Yourself in AI
1:28:40 to 1:29:20
Discuss the increasing competition for grad students and how to navigate it effectively.
“there's so much competition just because of like the, like everybody wants to get into AI.”
Transcript
Automatic transcript. May contain errors.0:00The thing that I find the most useful about like these models in general is like when I have these big spreadsheets of a lot of results and I just need a lot of it. I think like, Nano Bonana also got to the point where, I usually like, we make these images, it's just like, it's for fun, it's just troll your friend, or something like that. But like, Nano Bonana actually really got so good.
0:36Welcome back. How are you? Yeah, I'm good. I'm good. Great to be back. It's been one and a half years. Yeah, it's been one and a half years. Feels like a long time. So last time we talked, you were at Rekha. Yeah. And then you joined GDM again, working for Quark again. Yeah. and more recently you've started GDM Singapore. Yeah. Is it GDM Singapore or Gemini Singapore? I don't know if you've named the team. I think we have a Gemini team in Singapore. Yeah, I think in Singapore. It's called Reasoning and AGI. Yeah, Reasoning and AGI. Is it important to have AGI in the name? It was like a vice thing that we put AGI in.
1:10Yeah, I think that like one reason why we work on these models is that we want to get to AGI and this was a vice thing that we added AGI to the job posting. Yeah. There is no like formal name of the team yet, but it's basically the Gemini team Singapore. I mean, I think people are like trying to triangulate. Amazon has an AGI team, you guys have an AGI team. And then let's say Meta now has a super intelligence team. What are people signaling when they choose these names for their teams? Do they have, oh, we have a plan? Or is it just vibes? Are you trying to official hot take something? No. You have officially AGI in your job title.
1:46No, it's not a team name. It's not a team name. we are. Yeah, it's just, you know, we just want to signal the North Star of we're building these models to get to HCI. Yeah, yeah, no, I wasn't really fishing politics. Okay, so you rejoined GDM. Yeah. And I think last time we talked about, I listened back to the whole thing, it was an amazing episode last time. You were talking about how it's like externally, we're in brain and came out and now you're back in GDM. Yeah. I wonder what's your general reflections, just plugging back into the Google infrastructure. Oh yeah, so I guess coming back, It's very interesting because it felt and return to Google, like everything, including your LDAP, your username is all the same.
2:24It's like you play Pokemon, you leave it aside and then you go back and you click continue. Save game. Yeah, you save game and continue game. It's like that. Obviously, the last 1.5 years, WoW is away. Many things have changed. Brain is now part of GDM and stuff. So I think that obviously a lot of things have changed, but I think overall, the coming back has been pretty seamless. Obviously, I love Google infrastructure and I think debuts are great and stuff like that. Yeah. And I'm very glad to be back to Google Infra. Yeah. And was the intention always that you were going to work on Deep Think?
2:56No, not really. I think I miss research a lot, like doing like research, not like super fundamental research, but like close to model research, right? But I really miss being at the frontier and trying to go beyond that, right? So I really miss that a lot. And I think when I came back, Big Thing wasn't a thing. And I don't think there was any plans actually it was just like i'm just going to work on research and see what happens yeah i'm sure i guess there was some inclination that reasoning is the next frontier and that's like obviously the most rewarding research path especially this year yeah i think reasoning these days reasoning and rl is like probably quite it's rl reasoning comes i spent a lot of my past life i call it the past art working on like architectures and pre-training but i think now i more i have like transition more to RL research.
3:45I'm not like old school RL, but the games RL and the old school RL. And to be honest, I had almost no RL background coming back. But I think like RL is the main means of modeling these days. And yeah, so I think it was pretty easy to jump back in. And I think a lot of fundamental skills in research is for general purpose and universal. And it's quite easy to innovate even in a tool set that you're not super used to. And yeah, so I think RL is basically the main modeling tool set that we play around with these days. Superficially, I see some, you know, in your UL2 and T5 work, some overlap of like, you know, the focus on objectives and the focus on the stuff that you're trying to incentivize.
4:28So I would have maybe guessed there was more overlap than you are saying right now, which is interesting. But I know, I understand it's very superficial. Technically, a shift is objective and they have some like overlap, right? Yeah, I think it's just mainly like the on-policy and off-policyness of designing these things that change how like also the learning algorithm itself right let's just introduce this kind of terminology to people if they're not that familiar with the sort of RL policy I do think that a lot of people are like trying to understand what is working about this generation of RL research anyway so Jason had this interesting post which I think you were co-signing which is basically you always want to be on policy instead of mimicking other people's successful trajectories to your elections and learn from the reward given by the environment.
5:13Basically correct your own path instead of trying to imitate other people's path. Yeah, yeah. And first of all, he writes really well and I wish that more people wrote like him. But I don't know, what's your reflection on that or your addition on top of that? Yeah, so I think like the biggest analogy of on-policy and off-policy is basically that off-policy is basically like when you SFT something it's off-policy. Basically you take some other model, larger model stuff and then it's basically like there's off-policy. Somebody else's generated outputs, trajectories and whatever. I think off-policy is mainly like the core idea of like modern LMRL where you like generate and then you reward the model based on its own generations and then the model trains on its own generations.
5:51Yeah. So it's more, it's a bit like self distillation to some extent you, the model generates its own output and then you reward it and then trains on its own output. So I think on policy nurse is basically this idea of like model training on its own outputs and letting the model like generate its own trajectories and then letting some reward verify it and then the model chain is on the opposite. I think this is more generalizable in general. I think there's still a lot of like science out and still to be done about the gap between SFT and RL itself. But I think basically on policy and off policy, right?
6:20And I think bring this analogy back to real life. I mean, this on policy in us is more like humans, we are more on policy because we go around the world, we make mistakes and then we are, okay, this is, but like imitation learning is possibly somebody else. Not first principle. It just tells you what to do and then you just copy. so I think yeah this philosophy bringing back to life is quite like powerful like when I like now I have a kid and everything like I want my kid to try stuff and then you tell them like okay this is like where this went wrong where this went right and stuff rather than okay you just copy everything somebody else does yeah there's a Montessori schooling is mostly that right like very unstructured learning like you discover your own path and we just give you a safe environment to do it yeah yeah yeah what is the point in which you should transition from invitation to on policy I do bounce back and forth We're humans, right?
7:09Not models, right? I would say in models it seems like there mostly has been a very concrete like first you imitate and that's pre-training and then you are right there Technically, SFT is still imitation but I think for humans it's also a little bit of this, right? Because if you basically like sports, right? When you play sports you start off by imitating like hardcore imitating but then you cannot imitate forever because you need to like imitation I don't know whether this is good energy but watching a lot of tutorials and stuff is more like imitating we learn, try to learn certain movements and stuff like that.
7:39But then like on-policiness is like going to the game itself and trying to get a reward signal from that, right? But so I think that humans do need some form of imitation learning, but like I think everybody starts off by imitating. But then again, the human and model kind of is not, it's just fun to have analogies, but we shouldn't like take things like super literally and stuff like that. I actually am quite a serious taker of machine learning insights into human learning. So we learn from models? now? Yeah, because I think like machine learning is the most scientific way we have ever studied learning, just in general.
8:14And that's true, that's where we have to invent curriculum from like scratch. Yeah, that's true. And things like learning rate, if your learning rate is too high, learning rate is too low. Wait, do humans even have a learning rate? So I do tell people to keep an idea of their own learning rate and to be wary of it being too low. So for example, if you've been wrong once, you should ask, where else have I been wrong? And typically usually, let's say learning. Oh, okay. You know what I mean? People usually update slower than they should when they've been wrong. Is it stubbornness? It could be stubbornness.
8:49I don't know. Is that the right word for it? It could be like, they're too Bayesian when actually their prior assumptions are wrong and they need to completely throw out their previous assumptions because one counter example example invalidates all prior experience your entire world model is wrong throw it away so Bayesian actually wrong let's say you live for 10 years under some assumptions and you have one example that breaks your narrative okay you shouldn't be like okay now I have two percent update no actually it should be like oh like something's really freaking changed everything I've assumed for the last 10 years is probably wrong what else am I wrong in and update 20 % update 50 % not 2 % you know I mean that's your that's a learning rate thing for me so my my direct example is the whole getting into AI stuff.
9:37I was watching GANs for 10 years. Yeah. Has it been 10 years? 2012, 2013? Time flies, yeah. I was watching GANs and I was like, okay, this is cool. It's getting more detail. Not that impressive. Then all of a sudden, stable diffusion came out and you can run it on your laptop. And that was my learning rate. Okay, like, fuck. Like, my mental model of generative images did not include this. And so I was like, okay, like, I am very wrong and I need to pivot everything. And that's how I started Latest Space. So will this mean that your learning rate is high? Yes, I will nudge it up. I schedule my learning rates because a role model has been violated.
10:14Okay, I think it's a good strategy. I think also this brings a little bit to when new paradigms happen, how fast people are to adopt it or to invalidate their understanding of things. I think as scientists, we definitely, a lot of times, we do have to keep, as the few progress, we do have to keep invalidating our own role model. it could be like a certain way the way to do like something all along and suddenly something comes along and inviolates it, yeah. Yeah, you can be very proud of your priors until it becomes your prison. Yeah, I know. That is actually very dangerous. Yes, yes. Yeah. Okay, that was a bit of a tangent.
10:46I don't know how we got there. You did highlight Danny's LLM reasoning lectures where he could trace the intellectual history of reasoning in LLMs. Yeah. Chain of thought to not RLFT. And then one part that I was going to prompt you a little bit was also self-consistency, right? Yeah, I think people roughly know. I think it's more crudely implemented with OpenAI than with you guys, where it is straight up, they have eight inferences and they judge or whatever. But I do think that also is relevant to on-policy distillation, where it's like literally you have eight different paths and they're all from the same model.
11:19So I'm checking my intuition there. Basically, the stuff that you were saying about why on-policy is important and using, let's say, an external verifier to improve your reasoning. You can also do that with parallel reasoning. Oh, yeah, yeah. I mean, like when we train RRM models, they sample multiple times. So yeah, to some extent, there's some form of self-consistency. Is that directly? Self-consistency, right? Yeah, self-consistency is a little bit more. It's a more nuanced version of, if we talk to Danny, it's not majority voting for sure, but it's more. I agree. Yeah, it's more nuanced version of that.
11:50But I think parallel thinking definitely is related to self-consistency. Yeah, yeah. I think for those people, OpenAI also actually put out some interesting papers on majority voting versus other forms of, like, multiple output consensus. Then basically, like, the highest level is an actual LLM judge that decides, like, this is actually a worthwhile trajectory that is more valid based on some internal consistency or just, like, inspecting the chain of thought. Which is very cool that we can train models to do that. Yeah, for sure. Yeah, yeah. Self-consciousness is a big, like, a big fundamental idea.
12:23I mean, chain of thought itself. also a big idea. And then self-fantasy was also like a big fundamental idea in modern like LIM literature. Yeah, amazing. Okay, so let's bring it to I guess one of the headlines of this podcast is going to be about diving into the IMO. Yeah. So this was around about May, March, March, July. July. You guys announced oh, this very nice photo here. This is the photo I was looking at. This is in London, I believe, where you had the shroom. Yeah, the shroom. Oh, you got to be at the photo taking to get the credit. That's bullshit, right? No, no, no. No, I'm just kidding.
13:01The contributor list is bigger than this. Yeah, yeah. But like, they were like saying that, oh, okay, you should go to the photo. In order to get a literal gold medal? No, no, no. Like to get the credit for being the IMO. But it's created in the photo. So it's just a joke. It's just a joke. But anyway, okay, could you tell the story of starting this IMO thing? Apparently, it was done in one week. So let me like, to be a bit more clarified, a lot of things, right? So the IMO effort has been like, very long standing. So Tang and basically and Quok has been working on this even last year, right? Last year they got this.
13:30I was not back at Google at the time. They had the alpha geometry stuff and then they were like alpha proof and stuff. So it's a very long extending effort. But I think this year was the we wanted to try to actually use Gemini as an end-to-end model. Basically, no second system with alpha proof. In, text out model. Even that was a not intuitive thing. I covered the silver result from last year. And I was like, okay, it's pretty close. Like it's one point off from a silver. Just try a bit harder, you'll get gold. That decision to abandon it, I think it was pretty bold. I don't know. I personally was believed, always believed in, if we are not, like in retrospect, it's easy to say this, but it's a bit like, if the model can't get to IMO gold, then can we get to AGI?
14:15Basically, so it's basically at some point, we have to use these models to try these Olympic competitions. And I think that one of the goals this year was like, okay, we're going to do an end-to-end like tag-in, tag-in model. That's where like my involvement came in. So basically I was not like involved in the IMO effort only until the model training part. So I have to say that Tang did most of the IMO thing. I just trained the model with a bunch of other... What does that work involve? So basically we just prepared the model checkpoint for the actual IMO itself, right? So that's also something that's easily overlooked about the IMO thing was that many times you want to chase benchmarks or stuff like that.
14:55It's always like a thing that you can kind of keep running and running and hill climbing until you get there. And then you, but like the IMO was a live competition. Like some members of the team were in Australia for the thing. And there was like, this happening thing was happening live, was unfolding live. Oh, it's a very alpha goal. You receive the thing, you like punch it into your system. And then you like. Yeah, yeah, yeah. So like some of the professors from Tang's team were like, when they went to the IMO itself and stuff like that, the conference, I don't even know whether IMO is a conference, but it feels like there were people there.
15:23Like in Australia. And then, so it was a live thing and there were people who actually, the job was to run inference on this IMO P1, the P6 that came out. And they also came out on like different days. So it's like different sets, like one day, one day, two, something like that. So the fun part is that I knew nothing about IMO like at all. I'm not like a kid that took part in IMO. I stood down for that. You're a piano player. Yeah, I played the piano. But what I only knew was that, okay, we delivered the checkpoint and that checkpoint was used to do the IMO goal. But then there was somehow a week in London where everybody gathered there.
16:00So everybody was flying to London and then this photo was taken there. And then you get to see how the different parts come together. And also being in the other rooms with the other co-captains. And then it felt a little bit like a hackathon thing. so yeah I think this was like the training process of this IMO model itself was like maybe a week or so not the actual like the whole like basically yeah I think the question is I'm still not over the decision to throw away alpha proof okay yeah basically I think it's very major and I understand that you have this goal of AGI obviously like at some point one model should do it to do all of it right but I think you pointed a gun at me and said in 2024, what do you need to do IMO and IOI and CPC and all the other stuff that you guys did was you need an LLM reasoning system that knows how to operate a computer and knows how to write lean and run lean verifier and all this.
17:00But basically you ROFT the lean verifier into the chain of thought. Is that? Wait. So basically like it is not obvious that you can do that at all because I think what you mean is that like in some way it's encoded in the parameters of the model somehow. Yes. Yeah. I mean, it's just whether at the end of the day you just believe in like this like connection is one model loss of parameters. I mean, there's also tool use, right? There's also tool use which, you know, and stuff like that. But I think to some extent the model, like, I think we should be able to get to a point where like in the past when the LM first started, the model couldn't even be a calculator.
17:45Now it can somewhat be a calculator. So technically, like a tool, like a calculator is somewhat encoded in the parameters of the model. So I think we will eventually get a point where whether there's things that cannot be expressed in the parameters of the model is like an open question. We don't know where's that limit, but I think we will keep pushing and pushing this limit. So whether like something like a lean system or like some other things to solve other like the physics engine or something where we still continue to push that boundary. But I actually don't know whether there were a lot of debates about symbolic systems versus...
18:22Yes, that's the word I was trying to. I actually don't really know whether there was... To me, I was just like, oh, let's train the model and then someone told me to train the model and then I trained the model. Basically, there was overarching people like the IMO effort that decided this. And I also think that because basically this... specialized systems are very like one-off systems that are like you could create like a chemistry engine you create a math engine you could create the thing right but at the end of the day you want one model for everything so yeah so i think this kind of fits that direction a little bit more where you have one model for and then this model was also like launched as gemini deep thing as a general purpose gemini deep thing so it's basically unchanged but with maybe some config toned down a bit yeah so the the inference time config was like the one served to most people is different and the full IMO inference config was shipped to some mathematicians just because of the inference cost, right?
19:17But that was good enough to be a general purpose model. I think my take is that this intuition was what led to the trying to go towards one model instead of because these specialized systems, there's no end, right? You can create many specialized systems. Yes. The most I can see in the future is there'll be a model that is there's something that really cannot be subsumed by a model. then you just use a tool or something like that. Yeah, right. But my prediction is that I think most things can be subsumed by the model. I think, yeah. AI researchers are quite good at hill climbing. History would say that you have a lot of evidence backing you up.
19:53Is this the model output? This is it, right? This is what? Yeah, I think this is the model output, yeah. What do you see when you look at this? You just see, obviously, it looks like a well-written problem. It looks like something a real human mathematician would do. People did compare yours versus the OpenAI one, where OpenAI is a lot more raw or had to clean up their versions. We don't have to talk about OpenAI, but I think what is interesting to you when you saw this kind of output? I want to give a little bit like a special disclaimer is that I know nothing about the app. Right. So I think the wonderful thing about this era of LOM is that you can be an AI researcher, an engineer, and you don't have any domain knowledge and you can still get a gold medal.
20:34that you don't know anything about. I can't pass this at all. This is foreign to me. But maybe a proof is a particular kind of chain of thought. But I would say that the other interesting thing that some of your collaborators were talking about, he was like, oh, this is the first example of reasoning in a non-verifiable debate, which to me isn't proof by definition verifiable. I just want to give you things to riff on or debates that might be worth digging into. So I think that's good. there's a lot of, aside from the pools, there's a lot of domains that are non-verifiable and I think not easy to verify.
21:08So it's like when people mean non-verifiable, it's like non-trivial to verify or just not as easy as the solution of a math problem because pools are long form and it's also, that's why it's non-trivial to verify unless you convert it to lean and then you do all kinds of things, right? So I think there's a lot of work to be done in this non-verifiable domains here. I'm getting to this territory where I'm not sure what I can say or what I cannot say. Okay, so yeah. Sure, that's it. I think another thing that is an open topic of debate was how much domain-specific work or post-training was done because you then went on to do the IOI and CPC stuff as well, right?
21:47The same model. I was not directly involved in the ICPC, but I was related to Science 10. That's all I can say, yeah. Yeah. Any other interesting call-outs maybe just on the team? You called out Jonathan as someone who is co-captain on this effort. And yeah, basically, how does the effort look? So I think there were four captains for the IMO2 from London. Jonathan was from Mountain View. I was from Singapore. So I think four of us basically trained this model together. And I think one, I was also trying to see what Tang was saying. But I think one interesting thing was that we were all in different time zones.
22:25And there's something also very interesting about passing on the job. there's no like really like fixed workflow how to work together between captains so it's more like oh i'm going to board the plane now i'll be like afk for 12 hours so it's romantic so you're just babysitting the run sometimes there are bugs and stuff the job comes down sometimes so basically it's very ad hoc and it's very very it's really between the captains how we decide to like work together and yeah but i think it was a kind of interesting time also because we were all flying like i think the london folks were not having to fly but i had to fly and john had to fly and then like when you visit another country, you have another, like if you visit another office, you have many meetings.
23:02So I was in and out meetings and it was pretty like interesting. And I also think that nobody really knew whether we would get a goal at that time because the IMO actually hasn't happened. Yeah, it was interesting and exciting. And then I think like the whole process of this getting verified by the IMO committee and you know, like, you know, there was like, okay, but we're not going there. But I had to learn a lot about how the IMO works. Apparently the goal score is not even, it's not even a fixed number it's like a bell right so it was like a time where you just look at the score you're like like I was even like looking at the watching the human participants and then seeing like what their scores were because whether Gemini will get gold depends on like how the humans do so you like looking at oh if a certain percentage you're like what's the to some extent you don't have any control over that so yeah but you're just curious right because that's what I would say that it's definitely more like exciting like there's more adrenaline than like just running on a benchmark and getting a number was like a process that took some time.
24:03Yeah, yeah. But I think overall, if you have specific questions, you can ask also. But I think this whole thing has been a highlight for me. This IMO effort has been... Yeah, I would say most people, if you ask them maybe two years ago whether a model could get an IMO goal, they would have said that could be possible. Yeah. Then the silver helped, right, from last year. But like the fact that you can throw that system completely away and then just take existing Gemini and scale up DeepThink and then just run it for IMO Gold. I think it's also like very non-consensus compared to last year. Yeah, definitely.
24:35To some extent, I think researchers were also surprised. I wouldn't say like surprised, but like it was more like a pat on the back kind of surprised where we actually made a lot of progress in we as in collectively that all the engineers and researchers working on Gemini, there's a lot of progress being made. Let's look at how much we went in one year. Yeah. And I also think that it's just five years ago, like not two years, like five years, you just imagine the outcome. Like you just look at the state of AI now, like just generally the IMO and the ICPC goal and like also even like things like Nano Banana.
25:05If you just look at the AI progress now and five years ago, I think people would think that we already reached like AGI. Some form of AGI. Some form of AGI. We're just moving. If you just travel, like you take these checkpoints and you travel back into five years ago, somehow should make a drama about this. But I think it's really quite impressive like how the field has moved so quickly. Yeah. Yeah. The hard parts you would say were scaling inference. In what aspect? Like hard in terms of that even expensive? Hard as in maybe the most amount of brain power expended on the team. I saw some comments where they were like actually the hardest part was the inference optimization or like the very, very long horizon inference that deep thing needed compared to normal Gemini.
Read the full transcript
25:45Stuff like that. I didn't work on the inference time scale. Yeah. I wouldn't know here. That is mostly that old. And then there was this the codename was apparently IMOCAT, which you named after your desk. Okay, that's not really, like it was in the, so I think I tweeted about it at some point, right? Yeah. Going up to tweet. So the IMO cat was basically like, okay, it's not like an official codename or something. It's just like the name that the config of the job was like IMO cat. That's the. You just need some kind of name. Yeah, I mean, I just like, you know, I just, I like cats and then, yeah.
26:20Yeah, fair enough. That is mostly it on IMO unless you want to bring up anything else. we have other sort of researchy topics, but beyond, before I go into sort of researchy topics, I did want to maybe leave the floor to cover what else should people know about the reasoning effort that's going on at GDM? Let me think of where to start, please. What do people need to know? You know, it's very good, yeah. That's what people need to know. Maybe an easy one to start with would be a lot of people were focusing on maybe academic benchmarks two years ago, last year, maybe LLM Arena, this year Pokemon.
26:53Pokemon is very interesting reasoning, visual reasoning and just general long horizon agent planning benchmark. And I don't know, you seem to focus on it a lot and I think Gemini did very well so obviously I think it's something that is easy to talk about. I think I'll probably should be with this and there's actually nothing specifically done for Pokemon. Of course. Yeah, of course. There's nothing specifically done for Pokemon and I think that I think Logan had this tweet recently about the recent Gemini tree on Pokemon Crystal and Pokemon Crystal is like so much more special yeah I think Pokemon is like so I used to play a lot Pokemon and I'm a big Pokemon fan in general and I think it's a it's a great like you said it's a great long horizon benchmark and stuff like that and I think it's good to check in once in a while on these benchmarks that like almost never get contaminated or like people actually like don't spend time to like you climb benchmarks it's like kind of silly to like people like okay like what are you working on if some people are oh I'm working on Amy, I'm working on HRE, I'm working on Pokemon Maxing or something like that.
27:54That's kind of like funny. We did interview the Cloud-based Pokemon AI. I think his name is David. And it showed serious flaws in Anthropix's screen understanding vision capabilities. I literally couldn't tell. I'm trying to like get past this wall, but you just keep running into it. Because it doesn't know that the wall is there. And so it doesn't have any spatial reasoning at all. I mean, some of it could be like a harness, like the harness thing, or like also whether the model has access to like game state information or is it compared to visual. Yeah, Cloud's implementation is very game state heavy.
28:29They dumped effectively like all the memory of what's going on in the emulator. Yeah, yeah, I see. Yeah. I think for, I don't know whether I'm jumping off tangent or something like that. I think solving Pokemon is going to be more of like how fast you solve it. And then like the thing that I have not really seen so far is like whether the model can complete the Pokédex. Why is that? I don't think it's more challenging. No, completely the Pokédex is so hard. You need to plan, you need to like, you need to search up like information. Like, there's some things if you don't like go online and basically you need to like have a little bit of deep research in this here.
29:02The model will just never know like that it needs to trade. Okay, if it's able to go online, post on forums and then find someone to, like, hey, can I trade with you? Is Pokémon to evolve? Some Pokémon needs to be traded to be involved. Yes, or mated. Yeah, but anyway, I don't know. I have not seen the model available to complete the Pokemon. But complete the Pokemon is really hard actually for models. So I think that's actually an interesting, like, like an interesting one. Yeah. Yeah. I wonder what the real world analogy would be once, if let's say we have a model that is capable of doing that, what can we make it do that we cannot do today?
29:37But there's a lot of planning involved. It's just real deep research. Like planning, there's also a lot of planning involved. And it's more, I think complete the Pokemon game is very linear, right? The thing with Pogodex involves a lot of like... Backtracking. Research. Yeah, a lot of research and a lot of... Yeah, so it's probably a different nature itself. Is that as interesting to you as, for example, a lot of other people in the AI for Science world are trying to discover things that you cannot look up, right? Noble knowledge. Yeah, noble knowledge. Because basically what you're saying is we're not even there yet.
30:08We're at the place where models cannot consistently apply knowledge that they look up, right? You give Gemini access to web search and you say, okay, come try to collect all the Pokemon in a Pokedex. You don't have high confidence that it will do it. I don't know if someone's actually tried. Probably not, right? No, I think the hard part is actually trying to synthesize the web knowledge and then apply it in the game itself with all that visual state going on and stuff like that. It probably would be solved in one or two. It is challenging, but it's not super interesting. You're basically just...
30:42The task really is, can you look up the guide to do it? And then can you apply the guide? That's it. You know what's even more intelligent than that? Creating the guide. Like being the first to figure out how to create the guide. Which is what it is. Oh yeah, but then when it comes to this, it's mostly just an exhaustive search thing. A model will just try and try, like humans and guides for that. Okay, so that's actually less interesting to you. Interesting. Okay, actually, when you think about it, it's not super, super interesting, but it's okay. It's just like I have not seen a model to try to do this.
31:11Yeah, I think efficient search, of a novel idea space is interesting. Obviously, you can brute force anything, but we're not talking about brute forcing. We're talking about trying to create an AI scientist. But novel knowledge is actually an interesting thing that I think is going to be quite a big thing. Being able to generate novel... Google has done stuff there which I don't... You're probably not that close to those teams that has done AI scientist work. There's some things that have been... It's like, for example, if you freeze the model weights at like 2015, you freeze time at 2015, and then you, even with the current model, let's say you have, okay, let's assume there's no leaking of information somehow.
31:49Can, if you ask the model, what's the best ML? Okay, not 2015, like 2012 or something. It would just tell you like SVMs are like the best, right? This is the way that machine learning works in general, right? And then the question is, can you invent the transformer? Right? It might not be able to, like even today models, they might not even be able to invent the transformer. Like if you freeze the time at a certain time and even you bring the tech, I mean, the model is a transformer, so I just say there's no, assuming there's no leakage. issue. No, it's probably possible. So I think there's still a lot of open questions about whether the model can really innovate and generate really novel knowledge.
32:23Yeah. Yeah. One related question on that, which I think is related to the Danny paper, which is I think people have this sort of mythicism on what reasoning is. And reasoning, if you really demystify it a lot, it's whatever happens inside the chain of thought tags, right? and you post, you're eliciting that reasoning behavior from some stuff that is already latent inside of the pre-trained corpus. That's one version of this interpretation. I think these days, like reasoning itself is very vague and it's very open. So it's like mostly different people will have different definitions of what reasoning is, right?
33:00So I agree that this chain of thought is like basically when people think of reasoning, they associate it with chain of thought and obviously it's what happens in the thinking and then, okay, reasoning, right? But I think these days it's more like, I said in earlier, but it's like reasoning in RL is almost like this, like basically it's anything that is like post-training to at least capabilities, basically. It's like RL and post-training to at least capabilities. So I think the actual like technical definition of reasoning is making models better with thinking and post-training. Okay, yeah. So basically like RLing the model to think better, right?
33:34And thinking is more like thinking tracers and thought trajectories and stuff. like that right there's also this line of work for late like latent thinking and stuff like that like when the latent thinking and discrete token thinking is like going to be the same thing or like something like that it's like an open question meaning adding extra tokens to your vocab that represent all there's all this i don't forgot what's the name like these academy papers that do these loopy things or track token or like they basically instead of decoding discrete tokens you actually simulate this by doing this in latent space right so when you do chain of thought thinking and like reasoning you basically decode extra tokens hide in the thinking tag and then you decode stuff but like latent thing is basically you just don't decode tokens you just don't bother by it it might start speaking the native language of thinking is numbers not passing it through some filter of english and sometimes you must start thinking chinese or something else yeah generally i'm not i'm not really i don't really believe that model thoughts have to be the same with human thoughts i'm actually like generally in ml i'm more of the school of thought of let the model do whatever it wants in general there was a discussion there's a latent representation hypothesis paper that i think you're maybe sympathetic to if you haven't already read it to me sounds obvious basically image models will have the same idea what a laptop is versus a text model will have they converge on like the same latent yeah and obviously you can align them and you can do all those stuff with them and so it totally makes sense that their concept would be just a vector of numbers that represents laptop that's the concept yeah and And okay, maybe you have some numerical differences between one model's idea of what a laptop is versus another, but it mostly would be the same.
35:10Yeah, very interesting. The question I was kind of leading into was that because there's now we're in this age where LLM text is in the corpus of stuff that we train on, where it's a little bit of a recursive loop, right? Like the reasoning tokens are out there now. And so pre-trained models themselves, preaching base models that are also capable of reasoning. And they're increasingly so as more and more reasoning text goes into the corpus. Isn't that interesting? Or is that worrying? Do you actually see much reasoning traits on the internet? I've never seen those though. I would say that... Like on Hug Your Face?
35:49Yeah, people are publishing that specifically. As to whether or not people, researchers are actually including that in their training corpuses, who knows, right? Like, that's their choice. But I would say that percentage on Common Crawl that has COT tokens in there went from 0 to 0.001%. And it will just go up over time because people are publishing it. Yeah, but I think if the sources are quite clear, you can actually filter away those because usually people put it on GitHub. Do you want to filter it? Maybe you don't. There's a choice for the... Yeah, quite literally. the whole reason why I don't think we covered this in our previous part, but two years ago, a lot of people were like, oh, you just include more coding tokens in your pre-trained corpus.
36:34But the coding tokens are different from like coding tokens. It generalizes outside of code for reasoning. Oh, that was like... You don't believe that? No, no, no. I don't know if it's still true today, but I see. Yeah, that was just our general coverage of reasoning. I would say that there's a lot of interesting work here and more to do. Maybe I'll cover one thing which I know that you have personal inputs on, which is that you have started using AI coding. Oh yeah, so I actually don't really use much AI coding in the past, but I think we've reached a point where AI coding has started to become really useful.
37:08Like, okay, so before AI coding, the thing that I find the most useful about these models in general is when I have these big spreadsheets of a lot of results and I just need plots of it. I think models can quite go to the screenshot and make a plot of this. I hate making this method. talk like stuff about is so annoying okay but that's basically like one thing that i can remember about like how i use ai in the past but i think ai coding has started to become the point where i run a job i get a bug i almost don't look at the bug i place it into like anti-gravity and like i throw it that would fix the bug for me and then i relaunched the job that like beyond like vibe coding is more like vibe training vibe ml or something like that i would say it does pretty well most of the time and it's actually there are classes of problems that is just generally i know this is actually really good for and in fact maybe probably better than like i would have to spend 20 minutes to find like figure out the issue and then fix the thing and then we don't so yeah that's very interesting because i would say like level one vibe coding is you actually know what to do you're just too lazy yeah it's just i just do it for me like i've done this a thousand times like just go fix it like i know exactly what to do here you're saying it's like like the next level where you actually don't even know it's investigating it for you as long as like the answer looks right you'll just strip it at the start I was a bit like I did check it look at the thing and then at some point I'm like okay maybe the model looks better than me so I'm just going to let it do its stuff and then I will relaunch the job based on the fix that the model gave me and I think the models will just keep getting better and better so yeah it's something that I also think that recently there's this anti-gravity I think also because these tools were not like that in google infrastructure is not that easy to you don't i'm not that familiar what it's available outside and when i was in the startup i didn't really i think the models were not like so good like one and a half years ago so it's also like a forcing function that like it's also people like oh try adding gravity is a game changer and stuff like okay so i just started using and yeah yeah you spent some time with varoon recently what did what did well no i really did say hi and i guess you were telling me that you're an ai researcher that doesn't even use much AI and like now you're actually like AI-pilled as a user.
39:22There were so many moments this year where AI suddenly crossed that like that immersion thing. I think AI coding is one of them that we just discussed. I think like Nano Banana also got to the point where I usually like if you make these images it's just like for fun it just troll your friend or something like that. But like Nano Banana actually really got so good that you can use it for charts. That you can use it for like basically yeah so it's getting really good and I think yeah this year the stuff like And even things like the past, many of these LMS will like hallucinate things a lot. But now I just trust it automatically.
39:53I think we just, people are just enjoying the utility by these models. So now I'm like, I was always AI fuel. AI is a good thing. I don't see how anybody can disagree with that. Yeah, but you are actually using it for things that you are high expertise in, which is your own ML work. Yeah, yeah, yeah. And just to come back, do you have a special version of Gemini that you use internally that we don't have access to or it's like public Gemini? I think it's the public Gemini. Okay. Yeah. I was just saying like, it would be entirely reasonable to train Gemini internal for only your code base and your work.
40:31Oh, actually, I'm not sure though. You see, these things are just like, it's rather way for me. But obviously, if it obviously improves your productivity by, I don't know, 10%, yeah, worth it, right? Yeah. So I think that's interesting. And there's the interesting thing, levels of how much do you trust it? How much of your jobs do you automate away and you no longer need? There's also the question I guess about how people come up and train in the field if they you no longer need juniors. Because Gemini is your junior ML researcher. So I think these are all interesting questions. I want to say one quick thing first, right?
41:04So I think that when it comes to like whether a model can be like a junior suite or like something like that, right? So I think if you think of it at this way of if a job from a one one X suite, one time suite, can be replaced by a model itself. But let's say you are a manager, right? The objective that, the metric you track is like your time. And then if you can have a model that saves you like the same amount of time as the work that your reports do, but you don't actually like replace one person per se, but you, yeah. A little bit from everybody. Yeah, right. Then you can, I definitely agree that, okay, like when you count the net time save, there are times where the model can fix bugs that like would have cost me like one day, right?
41:41One day is huge. And these things are definitely like, I don't know whether anybody has done any real metric evaluation or this type of things, but if you use time as a real metric and then not as a number of, okay, maybe three hours kind of metric, but these things are not going to replace one person as it is, but more like a passive aura that buffs everybody. The in-game terms, right? I often think of myself as a bard because I tell stories and I plus everybody around me. That's an ideal situation for me in a D &D group. Oh, okay. I don't play D &D, but okay. I get it. They're the kings of Passamora.
42:17Support heroes. Support, support. Yes. Yeah, okay. AI support, I think, is very encouraging. I think like, where is it still not working for you? That you've tried and you're like, oh man, I expected it to be better. There are times when models try to get lazy and try to fix something. They are still... We won't have... They get lazy and then they try to like, gaslight me into thinking that like, the bug is fixed. so there are still classes of there are classes of problems that are like very easy for the model very hard for humans there are some things that are very easy for humans very hard for modern and stuff so it's still very hard to characterize these these things into this proper quadrants and stuff like that so i would say that the capabilities of models these days are good enough to be like really helpful but like it's still you is a bit like it still has some but yeah but i think this world like this i don't think there's anything that to be done to specifically like focus fire These things is more like general capability improvements.
43:12The model just gets better over time and then these things will just like go away. You say that, okay, so yes, I think obviously in the grand scheme of things, just trust the process, keep scaling in every dimension and things will just fall away, things will emerge. But you've also said in the past, I can't remember the exact tweet, where you were like, each additional data set compounds over time. They're just small additions. And I would say that when you say things like focus fire on things that you would think humans, it's easy for humans, hard for machines. Those are easy wins where you can just add a dataset that would focus fire on that.
43:46And isn't hill climbing just a sequence of doing that until you reach AGI? Okay, so I get your point. I think that it's true that sometimes a lot of progress on the whole is just a series of small incremental changes that push. I think that's accurate. That's true. There's also, it also feels that there's also a lot of like small, like seemingly minor for the lack of better word, like that push AI to the state where it is today. So I definitely agree. So nothing against people who like focus fire, but it's just that when I mean that like, it might be not easy to focus fire on things that are like not very easy to characterize.
44:22So it's just that like when he has, there's something targeted, right? Like, okay, I want to improve this capability, add some data. So I think like defining like, if I was defining the problems and stuff like that, it's like characterizing it. And if it can be characterized, then okay, then fine. But I think it's like, what I was trying to say with the coding is that these things are not even, some of the class of problems are like, I don't work on coding, but like people maybe work on coding, they know like, maybe they have like terminology for different types of failures. But so maybe somewhere somebody is focusing fire on this, then make the model better.
44:52That's great for everybody. Yeah, I mean, that's why it takes a thousand people to like get all these things together. AI is definitely like a big collective effort these days. It's a big machine. Yes, it's really crazy. Okay, so I just wanted to broaden out to general things people are talking about in the community on research. Which again, I know that you are very locked in, so you don't necessarily have read all the papers or anything, but we can just riff on it. You can obviously ask me what I think as well. Is attention all you need? So attention and transform is like a core idea in the recent times.
45:27like pre-training and scale is the thing that made attention and transformers like actually shine right because without i think the first transformer paper was this like a machine translation thing and then basically gpt and bird were the ones that like actually showed the full like big potential of this idea so in terms of is attention like really really all we need like probably no but i think it's like one of the from the architectural point of view also maybe maybe know, but it's not all you need. But you need it definitely. So what else are you thinking about on that same level? Are you talking about MOE stuff?
46:03What do you mean when you say it's not all you need? You definitely need the scale of pre-training. You need all the tokens. You need like RL. I think when I say, when people say, I say it's the potential you need, it's mostly from an architectural point of view. Will transferbers get us all the way to AGI, right? I guess it's the... So basically when you get to AGI, it's the problem. Will it still be a GNOME architecture or like meaningfully different? It would be a transformer, I think. Really? It depends on what you call it, but I think unless the paradigm shifts completely, which as a scientist, you cannot completely say no to that.
46:37This would never happen. But my feeling is that it's been almost 10 years since the transformer. 2017. 2009 years since the transformer. I think we have not replaced self-retention. It's some form of it. You could rename it. you could name it something else you could sometimes you can do local global yeah it's still a transformer in the end and I think like that's not going like anywhere unless the whole thing with like back prop like everything like goes like the whole thing just changes completely like then there's a different story there's a different conversation to have but if it's still within the same scope and bounds so I spent a lot of time thinking of about architectures and like whether there's alternate architectures and stuff okay at the sequence processing level like there is the ultimate sequence to sequence transformer.
47:26It's probably the self-attention is there was this whole big era which I was also involved in this era where people try to like undermine the attention as much as possible like they try to remove it simplify it make it efficient like this whole like efficient attention era and at the end of the day the outcome was always like oh we remove all the attention but we have one layer of self-attention there and it still works like that's at the end like always a story. Which even Noam a character he published some stuff about how he has some ratio of mixing of local and global attention, right? Like, basically, still attention, but modifying it quite a lot.
48:01I will continue local and global attention to be like, spill attention. Just like how much you're skipping. Yeah, the only question is that if the formulation changes too much, your QKV becomes like ABCD, EFG, or something like that. Okay, maybe I'll give you some motivating constraints in order to do this. You guys are still charging 2X for over 180k token context or 240, something like that. And the max, theoretical max is 2 million tokens, right? What if we need 200 million? Is there some point at which where even this concept of input token context is irrelevant because you are doing continual learning, that kind of stuff, where you're modeling it as, okay, the AGI will be achieved through a sequence-to-sequence transformation.
48:47Therefore, an attention is the best sequence-to-sequence model or architecture, therefore attention is all you need. But I think other people are like, sequence to sequence doesn't accurately capture intelligence. But that's not really about sequence to sequence. It's more about like the whole gradient design and backprop thing, right? It's not architecture itself. That's a problem that we, that is more of like the learning paradigm itself rather than the architecture itself. I think the architecture is just basically like the interface between the learning algorithm and the tokens. I think it's more about the learning algorithm itself.
49:20And this, like, continuing learning this, like, there's many ways to think about processing many, like, insanely large, like, contexts, right? Like, 200 million, 1 billion tokens or something like that, right? Like, whether it's going to be, like, you have a new learning algorithm that every time you run inference, you learn on it, right? Then you can technically have some kind of memory, like human beings learning as I'm talking to you, right? So that's also, like, one way. The other way is, like, whether, okay, maybe somebody will say that, okay, the attention is, like, it's too expensive for 200 million, one billion contacts, so we need new architecture.
49:53Or some people will say that, oh, we just improved the chips, like accelerators. So I think many ways to interpret it, but I think it's like, if it's about, there's a lot of fundamental things that, if it's about continual learning and stuff, there's a lot of fundamental things about the learning algorithm and stuff as well. That will have to change? I think like the learning paradigm and architecture and stuff, that goes like hand in hand. And I think as the field progress and ideas just stack on top one another, right? So there's also this thing about the idea that it was proposed has to be compatible with all the work that has been done before to shine, right?
50:24It's a bit of a variant of this hardware lottery like by Sarah. It's not the hardware lottery that I wrote about the GPUs failing, but it's the original hardware lottery, but it's more like a bit of a lottery of like the things proposed have to play well with the things that were proposed before. So it's a bit like going down this local opt-in minima to some extent. So now we are like in this local minima of like transformers and everything, everything, right? Maybe it's not easy to like get totally out of this. Because also a lot of people's investment optimization have been done. So the things that play well need to play well with the ideas before.
50:58And the way I see it now is it's very difficult to like come out of it. Okay, I'm not entirely convinced. I see what you're saying, but let's call it Gen AI. I fucking hate that term. It's still a very young few. And so, yes, there's been eight years of work on the Transformer. But what's that in the grand scheme of things? maybe we're in a local minima and we've got to notch ourselves out of it I do want to leave that open-ended I don't have an idea, I do think that people are in what Ilya Suskiver has been calling the age of research, right? We're like, okay, we scaled up what we can scale up we know what the next maybe 1, 2 orders of magnitude look like in scaling on every dimension that we know about, but what is the next dimension to scale?
51:41There's this misunderstanding misunderstanding a little bit about like all the last five years has just been like scaling things. Scale. Okay, please tell me more. Yes. You made that joke about now we scale researcher salaries. Okay, let's not go there. But I think that ideas like matter and I think that there have been a lot of good ideas in the last five years. It's just that maybe it's just not, so it's not been like blindly, like if you took an MLP, right, just like without self-addiction and you just, okay, I'm going to throw like$100 trillion on it and scale up that thing and the thing is never going to work.
52:15It's never going to work. Yeah, yeah. So there's no, like, there's part of it that's also, like, I think the bitter lesson gets used too much in, like, too conveniently used around. But actually, there's also a little bit of, not a bitter, there's also a sweet lesson where it's like, ideas matter. And I think even to today, right, like, people downplay ideas and stuff like that. Yeah. Do you think the rate of new ideas, without being specific about what ideas, because obviously you can't share, but like, do you think the rate of ideas has increased or decreased because there's like, kind of a law of definition returns?
52:44Are they smaller? I think the number of ideas is always proportional to the number of researchers working on a certain problem. So by definition, it should increase. But I think the number of ideas that actually work is not decreasing compared to the last. We're not in the era of diminishing returns yet. So I think the ideas are still very important and there's still very good ideas that are game changers that are being invented. And I think I know the answer to this, but is the closed lab advantage increasing versus open source or decreasing like the chinese labs they say keep publishing open source models and some of the american labs as well publish open source models would you say that the ideas that i see there nvidia has nemotron opening i has gpt oss these are all basically checkpoints on what's publicly known about training models as of this year you You know, it's declassified information because everyone does this.
53:44I think that the gap is increasing. I don't think it's completely predictable from stuff that you said before. I think the gap is definitely increasing. Yeah. I think that justifies researchers. Otherwise, what's the point of having researchers if not finding new tricks that compound over time? Yeah. Yeah. But definitely, I think it's increasing. Yeah. Okay. I'll do a side tangent. I don't know if you have any comments on this. So then this is very related to NVIDIA's recent purchase of Grok, which I don't know if you have views because you're very TPU-centric, but are we memory or compute bound?
54:19And this is relevant to the Transformers discussion of like... In terms of what, like serving? Exactly. I think the classic view is that we're compute bound because we just need more compute for pre-trained and RL and then inference. but actually the counter argument I would make against this is I actually have these charts of Moore's laws I wish I could just pull it up easily Moore's laws of the scaling of compute versus scaling of memory versus scaling of network and bandwidth and compute has a much higher slope of scaling than the other two memory I mean what the chip memory like honestly I don't think about this this memory bound that much so maybe it doesn't
55:28So I would disagree with it But honestly, I don't think about a dozen years. Well, it's this, that, that, that, that much. Yeah, I understand. Okay, data efficiency. So this is a joke, but implicit in this is that there's some kind of a maximum data exposure, right? And so I, so previously, I would say that a lot of the training paradigms is like one epoch is all you need. It's the mini title of this idea. I would say that the real number maybe is between three to four epochs. and I do wonder what the theoretical limit of data efficiency of a model in terms of training and compression should be.
56:05I don't know what that means. Data efficiency and basically but you're asking the question in a way of asking how much repeats is tolerable? Is that... One, tolerable is contingent on does it actually improve in meaning? It's not about you actually want to do it for its own sake. But I do think there's that and then there's also just the sheer amount of stuff that we can learn with limited data. So you say, let's say you're not compute bound, you're not memory bound. But let's say you are data bound. Last time that we were on the podcast, we talked about chinchilla versus inference optimal training.
56:39But now actually, I think a lot of people are even talking about data optimal training. Like given limited data set, how well can you learn from it? I think that's an interesting research direction that not enough people are talking about. Maybe it is something that is commonplace in the labs, but it seems very clear that we are very unoptimized with regards to how much we learn from our data. I'll just put it there. I think in general, like learning more, like extracting more from varied data points is definitely valuable, but I think that also relates to the fact that we're running out of tokens in the world.
57:19so I don't work on data for pre-training and I think things that I say were general state of industry not nothing the general state of industry so I think that I don't even know whether data has diverged the way that these things are done it has diverged too much across these labs and no open there's a lot of cross-pollination for sure cross-pollination? okay, okay but I don't think about data that much like the pre-training data that much yeah Maybe earlier this, the first half of this year, I would have said that kind of pre-training is dead. And that everyone's like just funneling all their work towards RL.
57:56And we had this like Grok chart, which is very interesting, where we're sending the same amount of compute on Cycran as to, you think it's a Cy-O. No, I don't know. I have no idea. I think that people are taking it seriously. They are like, yeah, okay, whatever. Especially in the agent labs, like Cognition Cursor, they're taking the open source models from from whoever and then adding let's call it pre-trained scale rl on top of it if they have that level of info which data which they do which is very interesting i would say yeah this data efficiency argument yeah i think to me is also more trying to discover new paradigms of learning in order to get where we want to all go which is yeah and the existence proof is humans, right?
58:41Your two-year-old daughter is much more capable than an LLM in some things having seen way, like, eight orders of magnitude, less data. That's very interesting. Yeah, comparing human learning and machine learning is definitely like... Purely as an existence proof that we could probably do better. Three examples of dog. Yeah. Fourth example of unidentified animal, I can probably tell is a dog, as a human. But machines, classically, you take 20. The data efficiency for off-humans is definitely way higher than models. Yeah, the only question is that where does this thing come from? Is it actually like putting more flops on every token?
59:19Or like maybe it's like back to the question about whether the transformer is the optimal architecture. Maybe it's backprop. It's the weapon. Maybe it's the off-policy-ness. Maybe it's the... Yeah. So what is the... Like, where is the bug, right? Exactly, exactly. But maybe it's a feature, not a bug, I don't know. So, okay, we've identified and probably it took me a while to get this across. So this is the kind of data efficiency I'm talking about. I think it's emerging. Basically, at the end of every year, I try to take bets as to, okay, what will be the big themes for next year? I think there's more of them that people are really trying to focus on.
59:49Because you're feeling this data crunch, even though everyone's still investing in data, I forgot to mention that I would say that I've been wrong on pre-training being dead. Yes, I've now met pre-trained leads from both Anthropik and Ofnii. And I've seen the talk from the DeepMind guy recently. and so everyone's investing in pre-trained still which is like nice to see. Nobody said pre-trained was dead. I know. No, it's a theory that we're trying to disprove or prove. Anyway, so I think, okay, let me wind back to my general idea, right? So yeah, data efficiency seems worthwhile. You would treat it as like, okay, well, show me where the bug is and I'll go fix it.
1:00:28We don't know where the bug is. We just have existence proof that it could be better. And then I think the final logical chain in this for me is that everyone is focusing on some idea world model as a version of this for more efficient learning, which potentially might not take the form of a sequence-to-sequence transformer. I don't know how that works. Like definitely a little bit out of my depth here. To me, that is more efficient because every world must be internally consistent. And if the next piece of evidence come in and invalidates those worlds, then you no longer need to pursue those paths ever and you can just narrow in on the world that you've identified and so to me that is learning where you're learning to fit world models or to the actual data yes so yes maybe you can treat the learning process as curve fitting yeah so you're learning the world instead of learning the world model yeah i said i'm learning the world model right okay by sampling multiple world models and then finding out which one fits the data the best so i guess my query is this what people talk about is this if you i mean obviously feel free to attack it because i'm just spitballing but this is what i pick up from talking to multiple people about okay what are you talking about with world models what are you talking about data efficiency and learning efficiency and like how do you all together in a cohesive sense of the future where we can actually what's the definition of world model at the start from start yeah there are three kinds okay okay go on yeah first kind is the vo kind or the what's the other one genie that deep mind has which is the sort of video world model.
1:02:01You model everything with some kind of Gaussian splats or whatever, and you inhabit that 3D space. Second, let's call it the Jan LeCun slash meta school of thought, which I don't know if you're that familiar with it. He has published the Jepak architecture, and then separately, Fair has also published the Code World Models, where you're basically, specifically for Code, very interesting, you are executing code and modeling the internal state of the execution environment as you go line by line. The LM actually learns to predict those things. And actually, it seems a lot more efficient at the scale that they've tested it out, which is very cool.
1:02:37Which definition are you anchoring on? The third one. The code world model? That's the second one. Those two are bundled together. The JEPPA fans are probably hating me right now because I'm lumping all metals work under one school of thought. Yeah. But whatever. Okay. The first one is VO, Genie, F.A.L.E.S. Super Spatial Intelligence. Those kind of video-based world models. Second one is some execution or some sort of explicit modeling as you run through the corpus. And I think that the third one is this amorphous thing, which I think people are trying to get to, where they are doing what I said about the resolution of possible worlds and your curve fitting as you learn, as you inference.
1:03:21Yeah, but what is the world model itself? itself is it like it is a mental model of where everything is and how you think the world works how what i think you think like everything okay but technically it's like it is something in the latent space okay okay so you can for simplicity it could be just be like a transformer model between and yes yes so to me that is the most coherent thing to the current paradigm which is you could actually do this in current transformers i think the way that you train it will probably have to be different. Okay. I see, I see. I don't have any conclusion here. I'm just throwing it out as something where I know you're interested in this kind of stuff and I don't have that many knowledgeable people to talk to about it.
1:04:02No, I don't think about world models that often. I think because world models are just not really well defined in the first place. But I think... So don't say world models, but the problem is learning efficiency and maybe, I guess, accuracy or like AGI capability that is not easily unlocked right now on our current path. of scaling yeah so i think i think when it comes to like like data efficiency i think it's more like uh i'm believer of finding ways to spend more flops per token right because you actually basically you're if you are data bound you want higher data efficiency because you can learn more from every data point squeeze squeeze out more points right so things like that can extract more can use more flops on every token is definitely like a form of data efficiency then there's the learning algorithm, right?
1:04:48Because I think there's this, there's a different scaling law for like humans is this, machines is this, like dogs are this, cats are this, there's this different, nice exponent. There's the ALEA chart. Right, yeah, yeah. There's this famous, famous, famous chart. And point one and point two are just like not entirely like the different things because it could be that better architecture is actually just spending more flops per token. So if you are, you come to a point where you are very data bound but not compute bound at all, you just find algorithms that spend a lot of compute on every token on every token.
1:05:18So I think the overarching point is just that, okay, it's a learning algorithm thing for data efficiency and then if whether the correct way is actually just to apply more flops per token just to squeeze out more from every data, every data point. Also because humans actually don't like, like they are exposed less, when you say less or more data it's also very ambiguous because they are technically like on 24, 7 and then you have a lot of like different types of inputs, right? And whether they actually spend more flops on everything that they listen is also a question because maybe they are just better efficient just because, I mean, somebody needs to count like how much flops the brain use to process like how much, maybe they're just spending more compute on every token and also maybe the learning algorithm is different.
1:05:59So, but I agree that data efficiency is very important given that I think we're going to like, there's a limited amount of data in the world. One more thing before we go into DSI. You know how like, we're talking about RL and like, you're working on RL stuff. Why are people paying so much for RL environments? So who is paying for RL environments? OpenAI and Topic at least. Nobody said anything about DeepMind. So a lot of the model labs that are not you are well known for paying at least seven figures for external startups to create RL environments for them to train in. Okay. And I think the question is, if your models are so good at coding, what do you do yourself?
1:06:40And so I think there's some amount of expertise that's being distilled from human experts into an RR environment that you can then let your agents run wild in. But I'm curious if there's any other deeper insight than that because I'm not satisfied with my own explanation. RR environments that have a lot of domain expertise are probably very valuable. And actually, I don't know specifically about what RR environments people are actually buying. But what was the thing that you're not satisfied by? It was so valuable. And a lot of people are saying like, look, it's a next year's app inside of a Docker container that logs stuff out when you send inputs in.
1:07:20Then you can probably do it yourself internally, right? Why you pay so much for some startup that you don't know to do it for you? Actually, I have no clue about why this is happening. Yeah, I have no clue. And a classic example would be like, if you want to build a computer use agent for buying things in e-commerce, you would want our environments that perfectly replicate maybe the top thousand e-commerce websites. and then you just parallel roll out on all of them. Does that seem meaningful? I don't know. Alright, cool. DSI and LLM Rexis. A big bet for me this year for my conference was we actually started focusing on LLM Rexis.
1:07:59The other... Actually, what's the motivation behind starting LLM Rexis? Track. Yeah. I think Rexis is the king AI problem in consumer. It is the single most valuable thing. All your feeds, even search is Rexis. Basically, it's search. basically like it's a retrieval it's the god problem right because rex is ranking but then also filtering also personalization also re-indexing and and like performance it is the god problem and you get paid a lot for it engineers are not that excited by it which is very weird because they don't see a lot of them don't work on rex's and they probably never will yeah but they don't see the monetary value that can come out of a good rex's the other two pieces of updates for me, which I actually didn't even know that DSI directly tied into this, was one Twitter publicly adopted their feed algorithm as an LM Rexus.
1:08:53LM's are just used everywhere now, but whether it's actually like a... Like a big LM. Whether it's like a generative retrieval type of model, it's like another question. Correct. Is it? We don't know. All we know is that they have said that they have swapped out their current Rexus for an LM-based Rexus. that's how they found it. But what is published is YouTube where they actually adopted semantic IDs for YouTube's rexies. And YouTube is obviously a big deal. Is it like public information? Yes. Okay, okay. They came and did talk about it with us. And then they published a V2 this year as well.
1:09:24Okay, I see. So basically, the last time you were on a podcast, we didn't talk about DSI that much. But you have actually some background in IR. You care about IR. I don't care about IR, but I think DSI, okay, like DSI or Genitive Retriever was like I think one of my favorite works in the old of like I have some IR background in like when I was doing a PhD I did some Rexy's work I did some retrieval work with Rexy's and stuff so I have some IR and Rexy's background so I think generally retrieval and generally Rexy's is very very conflicted DSI started as a retrieval thing so we did like natural questions like ranking of documents documents everything it started off as I mean that's actually we did an interview with Yannick, me and Don, we did an interview when the paper came out like a long time ago.
1:10:14So at that time, we wanted to like reimagine retrieval and such. So we wanted, at that time, LAMP, we were still using T5 models at that time. It was like not, we're not in the LAM era yet. It was pre-LAM era. It was like, okay, pre-training works kind of thing. And then there were like some pre-trained models around. So we wanted to reimagine retrieval, right? But retrieval rexists, they're all the same formulation, ranking, retrieval problem, right? And then that's where we started to imagine retrieval as one giant that encodes everything in the memory. We tried so many different, like, semantic ideas, actually, my collaborator, Vin, was the one that came up with either semantic ideas that basically, and the start of this whole genetic retrieval was actually basically literally just trying to give a document like identifier and just predicting like raw brute force predicting like this.
1:10:57It actually works because the models can memorize something. If you look at the literature from all the way to things like Doc to Vag, the words have no meaning in this ID in a vocab it's an obvious number and technically the models have enough capacity to predict but I think semantic ID was an idea that basically you have some semantic association and then you actually try to break down the search space hierarchically so how this work evolved in the RACS was at a time after DSI came out so Atchis Group and Mahesh the guy who they did some exploration of applying DSI to RACSYS. And that's how that generative regenerative RACSYS recommender system paper came out.
1:11:42Yeah, I didn't even know he was involved. It's crazy. That was like, basically us transferring this, like basically, okay, DSI will try to try it on RACSYS. And then I think the recommender system people have a slightly different way of doing semantic IDs, but it's basically just because the domain is slightly different. But after that, I think we were done with the invention part and with this one. The rest is details. The rest are details. So over time, I also left Google and stuff. So over time, these things evolved a little bit here. I think I also saw something like Spotify is also using something like YouTube Spotify.
1:12:16They use this type of semantic IDs, this type of DSI-like models. I think from the research community point of view, the DSI work was the first one that decodes semantic tokens. But then when we went to, I don't know, academic community is strange in a way that they will do things like, oh this is gender retrieval it's not gender rexist it's like they do this kind of random things that is a bit strange but yeah this was the whole history of this gender retrieval apparently there's also a lot of people working on I don't follow actually I don't follow this at all now it's not even in my mind but there was once I went to even in the Singapore office they are Swiss actually like working on gender retrieval they don't work on it I don't know whether they're still working on it but I met a person that tried to explain gender retrieval to me it was quite funny that I kind of co-invented generative retrieval.
1:13:03But I think this whole IRR thing has been just an interesting phase. And I definitely think DSR is one of my more creative works that I've done. That it's not really LM, but it's like... Under the general principle of apply ML to everything, if the Googler is working on generative retrieval, would that be like AI overviews? Is that something similar? I have no idea. Okay. For people listening, I did have a track there. I think you just type in AI.engineer and you'll get it. where the Gemini guy was talking, sorry, the YouTube guy was talking about how to use Gemini or the Rexis. I don't know what size of Gemini because he didn't talk about it.
1:13:40But this is public work now. And basically every YouTube video uploaded gets encoded into some kind of codebook and they retrain this every day on some kind of batch job. Yeah, just interesting. So yeah, I don't know if you even know what Gemini is being used for. I don't follow like this. this day and this day. I do think like in the sense of like for people who are not still not getting it, applying intelligence and the general intelligence of an LLM to the retrieval, to the recommendation task means you can accommodate such weird recommendations, like such weird queries as well that normally like no classical system can ever handle.
1:14:22And I think like it's also somewhat emergent in a sense that when you were using T5, you just couldn't actually add that much value on top of a normal BM25 retrieval technique. Would you say that's accurate? It is not just about paraphrasing, it is about understanding query intent. BM25 is a really strong baseline, actually. BM25 is a really strong baseline, yeah.
1:14:46I don't know the comparative delta versus T5 for you guys versus BM25, but I don't expect it to be very high and I expect it to be a lot higher for a true LLM-based basis, depending obviously on the query set. I didn't really think about it this way before, but because I've done modeling in many different domains, including like search, and even in the search community, I have a community, there's also a benchmark and stuff like that for like that people who climb on and there's some like, there's Amy or like a, I don't know what it's called these days anymore, but generally the modeling dynamics of IRL tasks is very different from, like NREX's tasks is very different from standard language tasks or like vision tasks and stuff.
1:15:27Like when we kill climate LOM, when you train models, like the way that modeling things interact with this environment is very different. So I think that I honestly, I hated working on like Rexxys and retrieval stuff. Okay, I'm just looking about old days. When you work on like T5, you work on, you change architecture, you try to improve complexity, you use superglue, like this is the olden days. You are on like, even now when you train LOMs, you just do zero shots, two shots stuff like that. Your things, because as a researcher, engineer you just interact with the environment a lot by this you're just like okay rl by this environment but rexys and ir has a very strange feeling to it strange feeling in the sense that it feels like you're like whatever works it's like you are in this in a world with the gravity is different or like you are in a world where the modeling things that feel intuitive are not intuitive so it feels like a very strange space to so i wrote some papers back in my days on like Rx's and stuff like that.
1:16:23Every time I ran some modeling experiments for Rx's and stuff like that, I didn't enjoy it. It feels like the environment was rude. It feels like the vibes are just like... What makes it rude? It just feels... Transactional. No, not transactional. I don't know how to describe it. For example, if you play sports like the tennis, where you hit the ball, you have a very nice feeling like hitting the sweet spot. When you do modeling in traditional, when you get the feedback back, you feel like everything sounds right, everything feels right, everything right. But Raxxys and IRR is like a problem where it's like you hit the shutter clock and you hear a glass shuttle like randomly.
1:16:58It just feels this weird sense of war that... Cause and effect are too far apart. Like it just feels strange. And then sometimes maybe the metric like... I think Raxxys, they use like all the NCCG effects. And then the BM25 is strong and then you just... And then you get like worse than like the BM25. Back in the day when you stack 2 LSTM to 3 LSTM, you were like, whoa, I see like... It's just the game. It's an unrewarding area to work in. it's just weird also the IR community and the retriever community is also like always behind the mainstream and then now it's just probably going even more worse because of ILM and stuff so okay I'm getting into Hortick territory but it's just like certain conferences are just like behind NewRibs and ICML and stuff yeah some conferences are just like they're just like applying things that they're downstream they're downstream okay so it always feels very uninspiring to work on on this look there's a reason that you left.
1:17:52But it was like a side quest. Like, I work on it as a side quest thing, yeah. Yeah, okay. Understand. I still think it's an important business problem, even though maybe it's an unrewarding field. You can understand why. Because the academic benchmarks for those tasks are just so far detached from what industry is. I didn't work on any of this like the thing, but that's from an academic point of view. Oh, then all you need are online invals, right? Yeah, yeah. Detests and go. to, ooh, okay. That would have been a different experience, yeah. That is mostly our sort of topic, research topics, coverage and everything.
1:18:27I think we're just going to end on a very simple one on GDM Singapore. You organized a symposium. Here we brought Jeff Dean, Kwok and all the others. Basically, what's the general message or the impetus for starting GDM Singapore? So we'll talk about the event first. So the event was mostly, so Kwok and I are going to start a team and then I think before I came back, We discussed this for some time. Jeff was very supportive of this. He was in the region many times, in Vietnam and Singapore, around the time where I was going to come back. And I think that, so this event itself was, Kwok and Jeff were visiting, and we just inspired the community here.
1:19:07I think that it's also a bit more like a soft, like setting the tone right for the start of the Gemini team in Singapore. And I think it's a very rare instance where you get somebody like Jeff and Kwok, who are the true pioneers of AI in the world to be in one room. And then, are you there as well? And I think like many people told me that - Not true pioneer of AI, but okay. I was there to live tweet. Yeah, I think having them all in one room and then giving these talks, like many people came out to me and said they were very inspired by their presence in the region. So starting a team and starting something is also very, like there's also no one moment that it's not okay.
1:19:43I press the button and it starts, right? It's like a process, right? So we hire people and then people join one by one and stuff, something like that, right? So I think this event was more, I would say, like to set the vibe. I think it's possible for Singapore to be close to the frontier. And I think that we're having the true pioneers of AI here. We want to give this, basically more like an inspiring thing. And also get led for Quark and Jeff to meet the people here. So Jeff was here last year, but Quark hasn't been there for some time and he's going to have a team here. So it's like also nice to bring him around and meet the people here.
1:20:15So it was already like amazing event. We met Long as well and who a lot of people don't know has a CS degree. It's like one of the few PMs with a CS degree. Yeah, yeah. I would say that the context of the meeting was more like partially like also he wanted to learn more about the IMO stuff. Oh, really? And then also about Jeff. Oh, because they invited you without knowing that these guys were coming or something like that. I was a bit like Jeff and Kwok and me, we went to visit the Cherif at the East Tanger and we discussed a little bit on DeepThing, discussed a bit of IMO, and then I think the rest of it was more like Jeff and Lee Seng-ho was talking more about, like, it became less about AI and tech, more about very macro, economical, political thing, which was, I was very out of element with, so I was just, I was just, you were in a suit.
1:21:09I was just talking about the DeepThing and the IMO and stuff like that, but he seemed generally quite surprised that AI has reached this point. I think. But it was a interesting... I would say for people, you have done something that is unique in Singapore's history so far. You're establishing a Frontier Research Lab in Singapore, which is an accomplishment. I think the other thing also that I guess I'm still trying to wrap my head around is, does geography actually matter? You're all working on a team, you have your London people, you have your Mountain View people, and mostly you're just collaborating with them anyway.
1:21:44You've collaborated with them your whole life. I don't even really know what countries mean anymore when it comes to research or just AI in general, because this thing is just inherently international from the start. This is a very good question. Also, it's related to the thing about identity, right? Because I think you also moved from SF in Singapore quite a bit, right? I was in Maldonville just like one, two weeks ago. And I'm here, but almost all my... If you just look at my siblings, aside from my family, like everybody I talk to is like somehow in the Bay Area or like that's because of work and everything I think the geography matters okay firstly the most like thing is like logistical is like probably the time zone so you literally want the 24 hour coverage around the world no there are there are advantage no what I'm saying is that the difference is like it's also more like like people define the location more than the location define okay the people somehow okay the time zone we get to the time zone a little bit like the pros and cons I think that's pros and cons, right?
1:22:44So you're bullish on Asia, Singapore people, the talent pool? I think we managed to find like amazing people. But I also have to say that this type of things is more like talent attract talent. I think most of the time, like people are very excited. Like the vibe I get is that people are very excited because it's like Coase team and my team. And we're working on library core things related to AGI. So I feel like the talent we can get from the region is really good. But it's only because it's us. We can unlock this talent. Otherwise, we might join some other place. Yes, and move to the US. Yeah, yeah, yeah.
1:23:25About the identity-wise, I would say that I definitely agree with you that why it doesn't matter. I think the advantages of Singapore itself or just anywhere, it's also that you are like... Okay, the world is very good. So technically, you can interact as much as you want. You can also go there. But I think Singapore has this advantage where you can go close and you can go far. I think the Bay Area is like so much about... I have friends in London and New York that would just never move to the Bay Area. I'm not against Bay Area. I think Bay Area is a great place, right? But it's just AI, AI, AI, AI, AI, everywhere.
1:23:59Right, I think sometimes if you have some like mental space and energy to have some other culture and then like, you know, like London, Singapore, New York, they have their own culture, right? But the Bay Area culture is just like AI, right? Like you just go anywhere, you just hear AI everywhere, right? Even the billboards and stuff. You can hear a bit much. Yeah. Although I did see some billboards down here. Or AI also? Yeah, I was like, what is this? It's a culture infecting Singapore. I do think that to some extent, if you want to do research in, you need a little bit of peace and quiet somewhere, right?
1:24:32So this island may be good for that, but then you can, you're still like able to like be connected right okay so i think that's mainly talent wise i think people are strong here uh yeah so far enough away but you're still connected you have strong talent what are you hiring for you're still hiring right we're hiring like like my team will work on like rl and reasoning for gemini and gemini deep thing i think we care more about like talent density now so we're not also growing that big, there's more diverse just because compute per capita is probably important. So I think that's something that we're hiring for now.
1:25:13Basically, I personally, there's a lot of the ends, but I think generally there's a lot of people who are very capable. But I think what I'm looking for mainly is either you have a track record of RL research or some, even not necessarily the RL, or like some exceptional achievement in like coding competitions or like some exceptional achievement somewhere, then there's like the kind of people that we want. Yeah, because you don't strictly require... I do remember something about your record days where you're like, you like to train your own trainers from scratch, right? So you don't... I forgot if I said that or not, but to some extent, to some extent, I think we'll definitely be very happy with people that are like very high stats and just like even without much statistical knowledge?
1:26:00No, no, stats. Yeah, yeah. The stat points. Child with int. Like just high tech tech points people, like just raw IQ, high tech talent people like this. I think all like strong engineering skills, ML, ML can be learned easily. Our knowledge can be learned easily. Yeah. I think maybe one version of this is, can it be done on a student budget? Can you do something interesting anymore on a student budget? I would say relevant to the point where conferences are quieter these days. I did do an interview with one of the best paper winners where they worked on thousand layer neural network RL. And that was done on the student budget.
1:26:36It was very cleanly executed pieces and paper and good findings. Look, I'm not sure if production models will ever go to a thousand layers, but they stretched it in an interesting direction and found some good recommendations. And the guy immediately got hired by opening eye. and I think that's encouraging for the grad students in the market who are like, okay, well, do I need to know somebody who works at these labs in order to get in? My uncle works there. I get the internship or whatever. No, actually, you can just do it on a student budget with good advisors. Oh, I actually think one thing interesting is that for most of the people that I actually went to recruit them like personally, right?
1:27:10So you see the work and then you send them DMs, right? So I get a lot of best people. No, no, no. Like for hiring generally. So generally, I almost, to your point, it's like you almost like you can just do good work, put it online and then somebody will contact you, right? It's actually super easy but super hard at the same time because I can tell you like I talked to a few of these grad students, they don't know what good work means, right? Because they don't know there's so many things their professors have the agenda they're forcing on them which like may not be right because it's not like their professors know what to work on either.
1:27:44So yeah, they just need guidance. They just need to, hey, work on these like five things. It showed me an interesting result in any of them. okay so if somebody comes up with something and then does something that you feel that is very tasteful and it aligns with what like researchers in the labs like like one and they come up with that independently you know that the function that produce is good right like if you just go and tell somebody to do this like you you can you just get the signal that this guy can execute right so i think there's some value in people that yeah they demonstrate taste research taste yeah research taste.
1:28:15Very interesting. I feel like I could give people, yeah, I do care about this. In some ways, the research directions work that I do is a little bit of that. Like, it's low accountability for me because obviously it's just thought experiments. But I think for a lot of people, it's like their career is bounded by can you demonstrate research taste with this like short three, four years that you have and just do it. Yeah. Yeah, I would say that this is more, there's so much competition just because of like the, like everybody wants to get into AI. It's just more of like how to, like mostly it's more like how you're going to prove yourself.
1:28:50Yeah. It must be hard these days to be a grad student trying to prove yourself. It's definitely harder, but yeah. Yeah, not your job. Okay, that was it. Do you have any other sort of rants or topics that you had queued up before we wrap? I don't, yeah. But it was great. It was, I had a great time. It was fun chatting, man. Fun, fun chatting. Yeah, even, I even love, last time we were supposed to do, last time we meet at the symposium, we were supposed to record. But even if we just ended up hanging out and chatting, it's just nice to get the brain dump of what's going on in your world. Because yeah, we're working on really important stuff, man.
1:29:19Always great to chat with you and see you here. Good to chat. Parting words on the sort of weight loss and workout journey. Because there's also a big thing for you. I think being healthy is important to do good research. Right. And I think I've been probably in one of... I'm probably in the peak of physical health now. Yeah, you look great. Yeah, thanks. and I think it's also impacted my work in a good way. You did the sort of Kabat-y inspired like biohacking. I didn't go so extreme but I was like also quite data-driven when I came I would like have my own email and then I would track this. I'm still supposed to make a blog post about this but I feel like I'm not like really at the end game yet so like when I get there I will but yeah just to just for people who don't know I think I lost 23 kilos this year actually across one year yeah one and a half years yeah so 23 yeah basically literally from the last podcast to now yeah there's an abolition study now 23 kilos yeah and I think like my HRV heart rate variability has went up by two times and my rising heart rate has dropped by 30 beats per minute 30 beats per minute like it was like 80-90 and now it's 160 oh yeah 80-90 is super high yeah I was unhealthy yeah yeah yeah okay yeah so I think when it's hard like what do you have the thing that kept you going, you know.
1:30:41A lot of people, they maybe, they're focused on AI, including myself, right? I do prioritize work. I enjoy work. I don't enjoy the fitness side. But obviously, it feeds in to your intellectual work, like the sort of log off and go for a walk, eat better, all that kind of stuff. It obviously feeds in. But like people seeing a positive example like you, they will get inspired to do the same thing. So I think it's good to set yourself up as an example. I think definitely helps. When I do these things for my health, I just think that it's also part of work because it helps me to get better at my job.
1:31:14So it's important as well. I think it's important as well. I like the HIV off the bat. I have no idea what mine is, but yeah. There's a general question about what is productivity and how do you measure it? What really matters? And it's still unclear to me. But I do think general energy level and hunger almost. like you almost have to like experience physical hunger in order to have intellectual hunger and I don't know if that's like a thing no when I'm hungry I just think of food I think to me it's like it's destroying things but it's hard to do work when you're hungry yeah okay thank you so much yeah thanks it's really great yeah have a great time
1:32:04Thank you.
From the publisher
From shipping Gemini Deep Think and IMO Gold to launching the Reasoning and AGI team in Singapore, Yi Tay has spent the last 18 months living through the full arc of Google DeepMind’s pivot from architecture research to RL-driven reasoning—watching his team go from a dozen researchers to 300+, training models that solve International Math Olympiad problems in a live competition, and building the infrastructure to scale deep thinking across every domain, and driving Gemini to the top of the leaderboards across every category. Yi Returns to dig into the inside story of the IMO effort and more!
We discuss:
* Yi’s path: Brain → Reka → Google DeepMind → Reasoning and AGI team Singapore, leading model training for Gemini Deep Think and IMO Gold
* The IMO Gold story: four co-captains (Yi in Singapore, Jonathan in London, Jordan in Mountain View, and Tong leading the overall effort), training the checkpoint in ~1 week, live competition in Australia with professors punching in problems as they came out, and the tension of not knowing if they’d hit Gold until the human scores came in (because the Gold threshold is a percentile, not a fixed number)
* Why they threw away AlphaProof: “If one model can’t do it, can we get to AGI?” The decision to abandon symbolic systems and bet on end-to-end Gemini with RL was bold and non-consensus
* On-policy vs. off-policy RL: off-policy is imitation learning (copying someone else’s trajectory), on-policy is the model generating its own outputs, getting rewarded, and training on its own experience—”humans learn by making mistakes, not by copying”
* Why self-consistency and parallel thinking are fundamental: sampling multiple times, majority voting, LM judges, and internal verification are all forms of self-consistency that unlock reasoning beyond single-shot inference
* The data efficiency frontier: humans learn from 8 orders of magnitude less data than models, so where’s the bug? Is it the architecture, the learning algorithm, backprop, off-policyness, or something else?
* Three schools of thought on world models: (1) Genie/spatial intelligence (video-based world models), (2) Yann LeCun’s JEPA + FAIR’s code world models (modeling internal execution state), (3) the amorphous “resolution of possible worlds” paradigm (curve-fitting to find the world model that best explains the data)
* Why AI coding crossed the threshold: Yi now runs a job, gets a bug, pastes it into Gemini, and relaunches without even reading the fix—”the model is better than me at this”
* The Pokémon benchmark: can models complete Pokédex by searching the web, synthesizing guides, and applying knowledge in a visual game state? “Efficient search of novel idea space is interesting, but we’re not even at the point where models can consistently apply knowledge they look up”
* DSI and generative retrieval: re-imagining search as predicting document identifiers with semantic tokens, now deployed at YouTube (symmetric IDs for RecSys) and Spotify
* Why RecSys and IR feel like a different universe: “modeling dynamics are strange, like gravity is different—you hit the shuttlecock and hear glass shatter, cause and effect are too far apart”
* The closed lab advantage is increasing: the gap between frontier labs and open source is growing because ideas compound over time, and researchers keep finding new tricks that play well with everything built before
* Why ideas still matter: “the last five years weren’t just blind scaling—transformers, pre-training, RL, self-consistency, all had to play well together to get us here”
* Gemini Singapore: hiring for RL and reasoning researchers, looking for track record in RL or exceptional achievement in coding competitions, and building a small, talent-dense team close to the frontier
—
Yi Tay
* Google DeepMind: https://deepmind.google
Full Video Episode
Timestamps
00:00:00 Introduction: Returning to Google DeepMind and the Singapore AGI Team00:04:52 The Philosophy of On-Policy RL: Learning from Your Own Mistakes00:12:00 IMO Gold Medal: The Journey from AlphaProof to End-to-End Gemini00:21:33 Training IMO Cat: Four Captains Across Three Time Zones00:26:19 Pokemon and Long-Horizon Reasoning: Beyond Academic Benchmarks00:36:29 AI Coding Assistants: From Lazy to Actually Useful00:32:59 Reasoning, Chain of Thought, and Latent Thinking00:44:46 Is Attention All You Need? Architecture, Learning, and the Local Minima00:55:04 Data Efficiency and World Models: The Next Frontier01:08:12 DSI and Generative Retrieval: Reimagining Search with Semantic IDs01:17:59 Building GDM Singapore: Geography, Talent, and the Symposium01:24:18 Hiring Philosophy: High Stats, Research Taste, and Student Budgets01:28:49 Health, HRV, and Research Performance: The 23kg Journey
Get full access to Latent.Space at www.latent.space/subscribe




