In short
```markdown
Latent Space
The AI Engineer Podcast
Episode Summary
Title
Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay 2 In this episode, Yi Tay, a prominent researcher from Google DeepMind, returns to discuss his experiences in leading the Reasoning and AGI team in Singapore. The conversation covers a range of topics, from the International Math Olympiad (IMO) to advanced concepts in reinforcement learning (RL) and the progress of AI models like Gemini.
Key Themes and Discussions
Yi's Career Path
- Progression from Brain to Reka and then to Google DeepMind.
- Leading the Reasoning and AGI team in Singapore.
- Involvement in the training of models like Gemini Deep Think and IMO Gold.
The IMO Gold Story
- Collaboration among four co-captains located in different time zones.
- The intense pressure of a live competition in Australia where immediate results were awaited.
- The decision to discard AlphaProof in favor of end-to-end Gemini, marked by a belief in the potential of RL-driven approaches.
On-Policy vs. Off-Policy Reinforcement Learning
- Off-Policy: Imitation learning from existing trajectories.
- On-Policy: Models generate their own outputs and learn from their experiences, similar to how humans learn from mistakes.
- Importance of self-consistency and parallel thinking for effective reasoning.
The Data Efficiency Frontier
- Discussion about the contrast between human data efficiency and that of models.
- Questions raised about potential bugs in architecture, learning algorithms, and backpropagation methods.
Schools of Thought on World Models
- Genie/Spatial Intelligence: Video-based models.
- Yann LeCun's JEPA + FAIR's Code Models: Modeling internal execution states.
- Resolution of Possible Worlds: Curve-fitting based on data.
AI Coding Assistants
- Yi shares his experience with AI coding tools, finding them increasingly useful for bug fixes.
- Discussion on how the models are now capable of providing solutions that surpass human capabilities in coding.
Generative Retrieval and DSI
- Reimagining search and recommendation systems through generative retrieval.
- Deployment of semantic tokens at platforms like YouTube and Spotify.
Recruitment and Team Building
- Focus on hiring talented individuals for the Reasoning and AGI team.
- Emphasis on finding candidates with strong track records in RL or exceptional skills in coding competitions.
Key Takeaways
- Yi emphasizes the transformative potential of AI models in competitive environments, such as the IMO.
- The shift towards on-policy reinforcement learning signifies a deeper understanding of how models can learn and reason.
- Data efficiency remains a pressing challenge, with ongoing discussions about architectural improvements and learning paradigms.
- The podcast illustrates the importance of collaboration across different geographical locations in advancing AI research.
Conclusion The episode showcases Yi Tay's insights into the evolving AI landscape, highlighting the intersections of research, practical applications, and the drive towards achieving AGI. His experiences underscore the importance of innovation, teamwork, and the continuous pursuit of excellence in the field of artificial intelligence.
--- For more details, visit [Latent Space](https://latent.space). ```
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Utility of AI Models
0:00 to 0:45
Learn how AI models assist in managing large datasets and creating visualizations.
“The thing that I find the most useful about like these models in general is like when I have these big spreadsheets of a lot of results and I just need a plot of it.”
Rejoining the Google Team
2:00 to 3:40
Discover Yi's reflections on returning to Google and the changes he observed.
“And I think last time we talked about, I listened back to the whole thing.”
Research Directions in AGI
3:40 to 5:00
Explore Yi's insights on the direction of research in AGI and reasoning.
“But I think now I more, I have like transition more into RL research.”
Reinforcement Learning Insights
5:00 to 7:00
Understand the nuances of on-policy and off-policy reinforcement learning.
“Basically correct your own path instead of trying to imitate other people's path.”
Real-life Applications of RL
7:00 to 10:00
Learn how reinforcement learning principles can be applied to human learning.
“and we just give you a safe environment to do it.”
Philosophy of Learning Rates
10:00 to 12:20
Discuss the concept of learning rates in both AI and human experiences.
“So will this mean that your learning rate is high?”
Transitioning to New Paradigms
12:20 to 14:00
Examine how scientists navigate shifts in understanding and validate models.
“So this was around about May, March, March, July.”
Training the IMO Model
14:00 to 16:20
Discover the training process and challenges faced in developing the IMO model.
“I personally was always believed in, if we are not, like in retrospect, it's easy to say this, but it's a bit like, if the model can't get to IMO goal, then can we get to AGI?”
The Role of Collaboration
16:20 to 19:00
Learn about the collaborative efforts of the team and the dynamics of working across time zones.
“But basically, you RLFT the lean verifier into the chain of thought.”
Pushing Boundaries in AI
19:00 to 21:40
Explore the implications of reaching toward AGI and the potential for one model to do it all.
“So it's basically unchanged, but with maybe some config toned down a bit.”
Show all 46 chapters
Reflections on Progress
21:40 to 24:20
Hear reflections on advancements in AI and what it means for the future of models like Gemini.
“or post-training was done because you then went on to do the IOI and CPC stuff as well, right?”
The Pokemon Benchmark
24:20 to 27:40
Understand the significance of Pokemon as a testing benchmark in AI development.
“To some extent, I think researchers were also surprised.”
Understanding Spatial Reasoning in AI
28:00 to 29:20
Learn about the challenges AI models face in spatial reasoning and game state understanding.
“and it showed serious flaws in Anthropix screen understanding vision capabilities.”
Completing the Pokémon Pokedex: A Challenge for AI
29:20 to 30:50
Explore why completing the Pokémon Pokedex presents a significant challenge for AI systems.
“I think that's actually an interesting one.”
The Nature of Reasoning in AI
30:50 to 32:30
Delve into the complexities of reasoning in AI and how it relates to machine learning.
“You know what's even more intelligent than that, creating the guide.”
Latent Thinking vs. Traditional Reasoning
32:30 to 34:10
Discover the differences between latent thinking and traditional reasoning in AI models.
“And reasoning, if you really demystify it a lot, it's whatever happens inside the chain of thought tags, right?”
The Evolution of Reasoning Capabilities in AI
34:10 to 35:30
Examine how the introduction of reasoning tokens influences AI model training and performance.
“hide it in the thinking tag and then you decode stuff.”
AI Coding: An Emerging Tool for Developers
35:30 to 36:40
Learn about the growing utility of AI in coding and how it aids developers in their tasks.
“capable of reasoning and they're increasingly so as more and more reasoning text goes into the corpus.”
Trusting AI in Code Generation
36:40 to 42:00
Discuss the evolving trust in AI tools for code generation and their impact on developers.
“Like, okay, so before AI coding, the thing that I find the most useful about, like, these models in general is, like, when I have these big spreadsheets of a lot of results and I just need plots of it.”
AI Model Limitations and Progress
42:00 to 45:00
Discusses the current capabilities of AI models and their limitations.
“I often think of myself as a bard because I tell stories and I plus everybody around me.”
Attention Mechanisms and AGI
45:00 to 49:20
Explores the role of attention mechanisms in AI and their relationship to AGI.
“You can obviously ask me what I think as well.”
Architectural Considerations for AI
49:20 to 51:10
Examines the implications of AI architecture on learning and progression towards AGI.
“And this, like, continuing learning this, like, there's many ways to think about processing many, like, insanely large, like, contexts, right?”
Research Trends and Scaling Ideas
51:10 to 54:05
Analyzes research trends in AI regarding scaling and new ideas in the field.
“And so, yes, there's been eight years of work on the transformer, but what's that in the grand scheme of things?”
Data Efficiency in AI Training
54:05 to 56:00
Investigates the concept of data efficiency and its impact on training AI models.
“I don't know if you have any comments on this.”
Understanding Data Efficiency in AI Models
56:00 to 56:50
Explore the nuances of data efficiency in AI models and the implications for learning.
“data efficiency of a model in terms of training and compression should be.”
Human vs. Machine Learning: An Efficiency Comparison
56:50 to 58:40
Learn about the differences in efficiency between human and machine learning processes.
“Maybe it is something that is commonplace in the labs, but it seems very clear that we are very unoptimized with regards to how much we learn from our data.”
The Concept of World Models in AI
58:40 to 1:01:10
Delve into the concept of world models and how they relate to AI and learning efficiency.
“And the existence proof is humans, right?”
Advancements in Learning Algorithms for AI
1:01:10 to 1:05:30
Discuss advancements in learning algorithms and their importance for data efficiency.
“learning where you're learning to fit world models.”
The Value of RL Environments in AI Training
1:05:30 to 1:07:25
Understand the significance of RL environments and their cost implications in AI training.
“And I think the question is, if your models are so good at coding, what do you do yourself?”
The Future of Ranking and Retrieval Systems in AI
1:07:25 to 1:10:00
Explore the advancements in ranking and retrieval systems, emphasizing their impact on AI.
“Actually, I have no clue about why Why this is happening, yeah.”
Reimagining Retrieval: Early Ideas and Collaborations
1:10:00 to 1:10:40
Explore how early concepts in retrieval were developed and the collaboration behind them.
“So we did like natural questions like ranking of documents and everything.”
The Evolution of Semantic IDs in AI
1:10:40 to 1:11:40
Learn about the transition from traditional document retrieval to using semantic identifiers.
“It actually works because the models can memorize something.”
Generative Retrieval: Insights and Personal Experiences
1:11:40 to 1:12:50
Delve into the concept of generative retrieval and the speaker's personal journey within it.
“Yeah, I didn't even know he was involved.”
The Impact of DSI on AI Research and Retrieval
1:12:50 to 1:13:20
Understand how DSI has influenced AI research and the retrieval landscape.
Navigating the Challenges of Traditional Retrieval Systems
1:13:20 to 1:14:10
Discuss the challenges faced in traditional retrieval systems and evolving methodologies.
“For people listening, I did have a track there.”
AI's Emergence in Search and Recommendation Tasks
1:14:10 to 1:15:00
Examine how AI is changing search and recommendation tasks beyond traditional methods.
“means you can accommodate such weird recommendations, like such weird queries as well that normally no classical system can ever handle.”
Reflections on Working in Retrieval and IR Communities
1:15:00 to 1:16:00
Reflect on the unique experiences and community dynamics of working in retrieval and IR.
“When we kill client L and when you train models, like the way that modeling things interact with this environment is very different.”
The Rude Realities of Modeling in IR
1:16:00 to 1:17:20
Explore the frustrations and challenges of modeling in information retrieval.
“But RECSIS and IRR has a very strange feeling to it.”
The Influence of Geography on AI Research
1:17:20 to 1:18:20
Discuss the importance of geography and how it impacts AI research and talent.
“Also, the IR community and the retriever community is also like always behind the mainstream.”
Establishing GDM Singapore: A Community Initiative
1:18:20 to 1:20:00
Learn about the founding of GDM Singapore and its goals for the AI community.
“That would have been a different experience, yeah.”
Cultural Influences on AI Research
1:24:00 to 1:25:10
Explore how different cultures impact AI research environments.
“I think sometimes if you have some like mental space and energy to like to have some other culture and then like, you know, like London, Singapore, New York, they have their own culture, right?”
Hiring for AI Talent
1:25:10 to 1:26:34
Learn about the qualities and achievements sought in AI research candidates.
“Yeah, I think maybe one version of this is, can it be done on a student budget?”
Navigating Research as a Grad Student
1:26:34 to 1:27:39
Understand the challenges grad students face in proving their research value.
“It was very cleanly executed pieces and paper and good findings.”
Demonstrating Research Taste
1:27:39 to 1:28:48
Discover the importance of research taste and independent work in academia.
“which like may not be right because it's not like their professors know what to work on either.”
Health and Its Impact on Research
1:28:48 to 1:30:19
Discuss how physical health can affect productivity and research quality.
“It must be hard these days to be a grad student trying to prove yourself.”
The Connection Between Physical and Intellectual Hunger
1:30:19 to 1:31:49
Examine the relationship between physical needs and intellectual performance.
“And I think like my HRV heart rate variability has went up by two times.”
Transcript
Automatic transcript. May contain errors.0:00The thing that I find the most useful about like these models in general is like when I have these big spreadsheets of a lot of results and I just need a plot of it. I think models can quite go to the screenshot and make a plot of this. I hate making this mad thought-like stuff about. It's so annoying. There were so many moments this year where AI suddenly crossed that, like, that immersion thing. I think AI code is one of them we just discussed. I think, like, Nano Bonana also got to the point where I usually, like, we make these images, it's just, like, a little fun, it just troll your friend or something like that.
0:30But, like, Nano Bonana actually really got so good.
0:36Welcome back. How are you? Yeah. I'm good. I'm good. Great to be back. It's been one and a half years. Yeah, it's been one and a half years. Feels like a long time. So last time we talked, you were at Rekha. Yeah. And then you joined GDM again, working for Cork again. Yeah. And more recently, you've started GDM Singapore. Yeah. Is it GDM Singapore or Gemini Singapore? I don't know if you've named the team. I think we have a Gemini team in Singapore. Yeah, a team in Singapore. It's called Reasoning and AGI. Yeah, Reasoning and AGI. Is it important to have AGI in the name? It was like a white thing that we put AGI in.
1:10Yeah, I think that like one reason why we work on these models is that we want to get to AGI. And this was a white thing that we added AGI to the job posting. Yeah. There is no like formal name of the team yet, but it's basically the Gemini team Singapore. I mean, I think people are like trying to triangulate Amazon as an AGI team. You guys have an AGI team. And then let's say Meta now has a super intelligence team. what are people signaling when they choose these names for their teams? Do they have, oh, we have a plan or is it just vibes? Are you trying to fish off hot takes on the... No. You have officially AGI in your job title.
1:46No, it's not a team name. It's not a team name. Yeah, it's just, you know, we just want to signal the North Star of we're bringing these models to get to AGI. Yeah, yeah. No, I wasn't really fishing hot takes. Okay, so you rejoined GDM. Yeah. And I think last time we talked about, I listened back to the whole thing. It was an amazing episode last time. You were talking about how it's like externally. You were in Brain and came out and now you're back in GTM. Yeah. I wonder what's your general reflections, just plugging back into the Google infrastructure. Oh, yeah. So I guess coming back, it's very interesting because it felt and returned to Google, like everything, including your LDAP, your username is all the same.
2:24It's like you've played Pokemon, you leave it aside and then you go back and you click continue. Save game. Yeah, you save game and continue game. It's like that. Obviously, the last 1.5 years, WoW is away. many things have changed. Brain is now part of GDM and stuff. So I think that obviously a lot of things have changed, but I think overall the coming back has been pretty seamless. Obviously, I love Google infrastructure and I think the views are great and stuff like that. Yeah, and I'm very glad to be back to Google Infra. Yeah. And was the intention always that you were going to work on DeepThink?
2:56No, not really. I think I miss research a lot, like doing research. Not like super fundamental research, but like close to model research, right? But I really miss being at the frontier and trying to go beyond that, right? So I really miss that a lot. And I think when I came back, big thing wasn't a thing. And I don't think there was any plans, actually. It was just like, I'm just going to work on research and see what happens, yeah. I'm sure, I guess, there was some inclination that reasoning is the next frontier. And that's like, obviously, the most rewarding research path, especially this year.
3:29Yeah, I think reasoning, these days, reasoning and RL, is like probably quite, it's RL reasoning. I spent a lot of my past life, I call it the past art, working on like architectures and pre-training. But I think now I more, I have like transition more into RL research. I'm not like old school RL, but the games RL and the old school RL. And to be honest, I had almost no RL background coming back. But I think like RL is the main means of modeling these days. And yeah, so I think it was pretty easy to jump back in. And I think a lot of fundamental skills in research is for general purpose and universal and it's quite easy to innovate even in a tool set that you're not super used to and yeah so i think rl is basically the main modeling tool set that we play around with yeah these days superficially i see some you know in your ul2 and fan at t5 work um some overlap of like you know the focus on objectives and the focus on the stuff that you're trying to incentivize.
4:28So I would have maybe guessed there was more overlap than you are saying right now, which is interesting. But I know, I understand it's very superficial. The shift is objective and they have some overlap, right? Yeah, I think it's just mainly like the on-policy and off-policy of designing these things that change how, like also the learning algorithm itself, right? Let's just introduce this kind of terminology to people if they're not that familiar with the sort of RL policy. I do think that a lot of people are like trying to understand what is working about this generation of RL research. Anyway, so Jason had this interesting post, which I think you were co-signing, which is basically you always want to be on policy instead of mimicking other people's successful trajectories to your own actions and learn from the reward given by the environment.
5:13Basically correct your own path instead of trying to imitate other people's path. Yeah. And first of all, he writes really well and I wish that more people wrote like him, but I don't know, what's your reflection on that or your addition on top of that? yeah so i think like the biggest analogy of our policy and our policy is basically our policy is basically like when you sft something it's all policy basically you take some other model larger model stuff and then it's basically like there's off somebody else's generated outputs trajectories and whatever i think our policy is mainly like the core idea of like modern lm rl where you like generate and then you reward the model based on its own generations and then the model trains on its own generations.
5:51Yeah. So it's more, it's a bit like self-disclination to some extent the model generates its own output and then you reward it and then trains on its own output. So I think on-policiness is basically this idea of like model training on its own outputs and letting the model like generate its own trajectories and then letting some reward verify it and then the model train its own outputs. I think this is more generalizable in general. I think there's still a lot of like science out and still to be done about the gap between SFT and RL itself. But I think basically on policy and off policy, right?
6:20And I think bring this analogy back to real life. I mean, this on policy in us is more like humans, we are more on policy because we go around the world, we make mistakes and then we are, okay, this is. But like imitation learning is supposed to be somebody else. Not first principle. It just tells you what to do and then you just copy. So I think, yeah, this philosophy, bringing back to life is quite powerful. Like now I have a kid and everything, like I want my kid to try stuff and then you tell them like, okay, this is like where this went wrong, where this went right and stuff rather than, okay, you just copy everything somebody else does.
6:54Yeah, there's a, Montessori schooling is mostly that, right? Like very unstructured learning, like you discover your own path and we just give you a safe environment to do it. Yeah, yeah, yeah, yeah. What is the point in which you should transition from imitation to on policy? I do bounce back and forth. We're humans, right? Not models, right? I would say in models, it seems like there mostly has been a very concrete, like first you imitate and that's pre-training. And then you are right there. Technically, SFT is still imitation. But I think for humans, it's also a little bit of this, right? Because if you basically, like sports, right?
7:25When you play sports, you start off by imitating, like hardcore imitating. But then you cannot imitate forever because you need to, like, imitation, I don't know whether this is good energy, but watching a lot of tutorials and stuff is more like imitating. You learn, try to learn certain movements and stuff like that. But then, like, on-policiness is like going to the game itself and trying to get a reward signal from that, right? So I think that humans do need some form of imitation learning. but like I think everybody starts off by imitating. But then again, the human and model kind of is not, it's just fun to have analogies, but we shouldn't like take things like super literally and stuff.
7:58Actually, I am quite a serious taker of machine learning insights into human learning. So we learn from models now? Yeah, because I think like machine learning is the most scientific way we have ever studied learning, just in general. That's true, that's true. we have to invent curriculum from like scratch. Yeah, that's true. And things like learning rate. If your learning rate is too high, learning rate is too low. Wait, do humans even have a learning rate? So I do tell people to keep an idea of their own learning rate and to be wary of it being too low. So for example, if you've been wrong once, you should ask, where else have I been wrong?
8:38And typically usually, let's say learning. Oh, okay. You know what I mean? People usually update slower than they should when they've been wrong. Or is it stubbornness? It could be stubbornness. I don't know. Is that the right word for it? It could be like they're too Bayesian when actually their prior assumptions are wrong and they need to completely throw out their previous assumptions because one counter example invalidates all prior experience. Your entire world model is wrong. Throw it away. So Bayesian actually wrong. Let's say you live for 10 years under some assumptions and you have one example that breaks your narrative.
9:15You shouldn't be like, okay, now I have 2 % update. No, actually it should be like, oh, like something's really freaking changed. Everything I've assumed for the last 10 years is probably wrong. What else am I wrong in? And update 20%, update 50%, not 2%. You know what I mean? That's a learning rate thing for me. So my direct example is the whole getting into AI stuff. I was watching GANs for 10 years. Yeah. Has it been 10 years? 2012, 2013? Time flies, yeah. I was watching Gans and I was like, okay, this is cool. It's getting more detail. Not that impressive. Then all of a sudden, stable diffusion came out.
9:51And you can run it on your laptop. And that was my learning rate. Okay, like, fuck. Like, my mental model of generative images did not include this. And so I was like, okay, like, I am very wrong. And I need to pivot everything. And that's how I started Latest Space. So will this mean that your learning rate is high? Yes, I will nudge it up. I schedule my learning rates. Because a role model has been violated. okay i think it's a good it's a good strategy i think also this brings a little bit to like when new paradigms happen like how fast people are to adopt it or like to invalidate their understanding of things i think as scientists we definitely a lot of times we do have to keep as the few progress we do have to keep like invalidating our own world model it could be like a certain ways the way to do like something all along and suddenly something comes along and invalidates it yeah yeah you can be very proud of your priors until it's like becomes your prison yeah i know that's actually very dangerous yes yes yeah okay that was a bit of a tangent i don't know how we got there you did highlight denny's lm reasoning lectures where he could trace the intellectual history of reasoning in lms yeah chain of thought to to not rlft and then the one part that i was going to prompt you a little bit was also self-consistency right yeah i think people roughly know i think it's more crudely implemented with open ai than with you guys where it is straight up, they have eight inferences and they judge or whatever.
11:10But I do think that also is relevant to on-policy distillation where it's like literally you have eight different paths and they're all from the same model. So I'm checking my intuition there. Basically, the stuff that you're saying about why on-policy is important and using, let's say, an external verifier to improve your reasoning, you can also do that with parallel reasoning. Oh, yeah, yeah. I mean, like when we train RL models, they sample multiple times. So yeah, to some extent, there's some form of self-consistency. Is that directly? That's self-consistency, right? Yeah, self-consistency is a little bit more in the more nuanced version of...
11:44If we talk to Danny, it's not majority voting for sure, but it's more... I agree. Yeah, it's more nuanced version of that. But I think parallel thinking definitely is related to self-consistency. Yeah. Yeah, I think for those people... Openly, I also actually put out some interesting papers on majority voting versus other forms of multiple output consensus when basically like the highest level is an actual LLM judge that decides this is actually a worthwhile trajectory that is more valid based on some internal consistency or just like inspecting the chain of thought which is very cool that we can train models to do that yeah for sure yeah self-consistency is a big like a big fundamental idea I mean chain of thought itself was also a big idea and then self-consistency was also like a big fundamental idea in modern like LLM literature yeah amazing okay so So let's bring it to, I guess, one of the headlines of this podcast is going to be about diving into the IMO.
12:40Yeah. So this was around about May, March, March, July. You guys announced, oh, this very nice photo here. This is the photo I was looking at. This is in London, I believe, where you had the shum. Yeah, the shum. Oh, you got to be at a photo taking to get the credit. That's bullshit, right? No, no, no. No, I'm just kidding. The contributor list is bigger than this. Yeah, yeah. But like, they were like saying that, oh, okay, you should go to the... In order to get a literal gold medal? No, no, no, no. Like to get the credit for being the AIMO. Efforts is created in the folder. So it's just a joke.
13:13It's just a joke. But anyway, okay, could you tell the story of starting this AIMO thing? Apparently, it was done in one week. So let me like, to be a bit more clarify, a lot of things, right? So the AIMO effort has been like, very long standing. So Tang and basically, and Kuo has been, was he working on this? Even last year, right? Last year, because I was not back at Google at the time. They had the alpha geometry stuff and then they were like alpha proof and stuff. So it's a very long, extending effort. But I think this year was the, we wanted to try to like use, actually use Gemini as an end-to-end model.
13:43Basically, no... No second system with alpha proof. In, text out. Yes. A model. Even that was a not intuitive thing. I covered the silver result from last year. And I was like, okay, it's pretty close. Like it's one point off from a silver. Just try a bit harder, you'll get gold. So the decision to abandon it, I think it was pretty bold. I don't know. I personally was always believed in, if we are not, like in retrospect, it's easy to say this, but it's a bit like, if the model can't get to IMO goal, then can we get to AGI? So it's basically at some point, we have to use these models to try these Olympic competitions.
14:21And I think that one of the goals this year was like, okay, we're going to do an end-to-end tag-in, tag-in model. So that's where my involvement came in. So basically, I was not involved in the IMO effort only until the model training part. So I have to say that Tang did most of the IMO thing. I just trained the model with a bunch of other coaches. What does that work involve? What are some things? So basically, we just prepared the model checkpoint for the actual IMO itself, right? So that's also something that's easily overlooked about the IMO thing was that many times you want to chase benchmarks or stuff like that.
14:55it's always like a thing that you can kind of keep running and running and hill climbing until you get there and then you but like the iMo was a live competition like some members of the team were in australia for the thing and it was like this happening thing was happening live was unfolding live oh it's a very alpha goal you receive the thing you like punch it into your system and then yeah yeah yeah so so like some of the professors from tang's team were like in when they went to the iMo itself and stuff like that the conference i don't even know where the iMo is a conference but it feels like there are people there like in australia and then so it was a live take and there were people who actually the job was to run inference on on this IMO P1 the P6 that came out and they also came out on like different days so it's like different sets like one day one day two something like that so the fun part is that I knew nothing about IMO like at all I'm not like a kid that took part in IMO I was too down for that you're a piano player and yeah I was a piano but what I only knew was that okay we delivered the checkpoint and that checkpoint was used to do the IMO goal.
15:56But then, like, there was somehow a week in London where everybody gathered there. So everybody was flying to London and then this photo was taken there and then you get to see how all the different parts, like, come together and, like, also being in the other rooms, in the rooms with the other, like, co-captains and then it felt a little bit like a hackathon thing. So, yeah, I think this was, like, the training process of this IMO model itself was, like, maybe a week or so. Not the actual, like the whole like basically yeah I think the question is I'm still not over the decision to throw away alpha proof okay yeah basically I think it's very major and I understand that you have this goal of AGI obviously like at some point one model should do it to do all of it right but I think you pointed a gun at me and said in 2034 what do you need to do IMO and IOI and CPC and all the other stuff that you guys did was you need an LLM reasoning system that knows how to operate a computer and knows how to write lean and run lean verifier and all this.
17:00But basically, you RLFT the lean verifier into the chain of thought. Is that third? Wait, I... So basically, like, it is not obvious that you can do that at all. Because I think... Okay, so what you mean is that, like, in some way, it's encoded in the parameters of the model somehow. Yes. Yeah. I mean, it's just whether at the end of the day, you just believe in this connection is one model, lots of parameters. I mean, there's also tool use, right? There's also tool use and stuff like that. But I think to some extent, the model, I think we should be able to get to a point where in the past, when the L1 first started, the model couldn't even be a calculator.
17:45Now it can somewhat be a calculator. So technically, a tool like a calculator is somewhat encoded in the parameters of the model. So I think we will eventually get a point where whether there's things that cannot be expressed in the parameters of the model is like an open question. We don't know where is that limit, but I think we will keep pushing and pushing this limit. So whether like something like a lean system or like some other things to solve other, like a physics engine or something, we still continue to push that boundary. Yeah. But I actually don't know like whether there were a lot of debates It's about symbolic system versus...
18:22Yes, that's the word I was trying to... I actually don't really know whether there was... To me, I was just like, oh, let's train the model. And then someone told me to train the model and then I trained the model. Basically, there was overarching people, the IMO effort that decided this. And I also think that because basically these specialized systems are very one-off systems that are like, you could create a chemistry engine, you could create a math engine, you could create a little thing, right? But at the end of the day, you want one model for everything. So I think this kind of fits that direction a little bit more where you have one model for.
18:55And then this model was also like launched as Gemini Deep thing, as a general purpose Gemini Deep thing. So it's basically unchanged, but with maybe some config toned down a bit. Yeah, so the inference time config was like the one served to most people as different. But, and the full IMO inference config was like shipped to some mathematic editions just because of the inference cost, right? But that was good enough to be a general purpose model. I think my take is that this intuition was what led to the trying to go towards one model instead of because these specialized systems, there's no end, right?
19:29You can create many specialized systems. Yes. The most I can see in the future is there'll be a model, then if there's something that really cannot be subsumed by a model, then you just use a tool or something. Yeah. Right. But my prediction is that I think most things can be subsumed by the model. I think, yeah. AI researchers are quite good at hill climbing. History would say that you have a lot of evidence backing you up. Is this the model output? This is it, right? This is what you do? Yeah, I think this is the model output, yeah. What do you see when you look at this? You just see, obviously, it looks like a well-written problem.
20:00It looks like something a real human mathematician would do. People did compare yours versus the OpenAI one where OpenAI is a lot more raw or had to clean up their versions. We don't have to talk about OpenAI, but I think what is interesting to you when you saw this kind of output? I want to give a little bit like a special disclaimer is that I know nothing about the app. Right. So I think the wonderful thing about this era of LOM is that like you can be like an AI researcher, an engineer, and you don't have any domain knowledge and you can still, yeah, get a gold medal. It's a universe of social.
20:34That you don't know anything about. I can't pass this at all. Like this is foreign to me. But like maybe a proof is a particular kind of chain of thought. But I would say that the other interesting thing that some of your collaborators were talking about, he was like, oh, this is the first example of reasoning in a non-verifiable debate, which to me isn't proof by definition verifiable. I just want to give you things to riff on or debates that might be worth digging into. So I think there's a lot of, aside from proof, there's a lot of domains that are non-verifiable and I think not easy to verify.
21:08So it's like when people mean non-verifiable, it's like non-trivial to verify or just not as easy as like the solution of like a math problem because it pulls a long form and it's also, that's why it's non-trivial to verify unless you convert it to lean and then you do all the kinds of things, right? So I think there's a lot of work to be done in like this non-verifiable domains here. I'm getting to this territory where I'm not sure what I can say, what I cannot say. Okay, so yeah. Sure, that's it. I think another thing that is an open topic of debate was how much domain-specific work or post-training was done because you then went on to do the IOI and CPC stuff as well, right?
21:47The same model. I was not directly involved in the ICPC, but I was related to Slystown. That's all I can say, yeah. Yeah. Any other interesting call-outs maybe just on the team? You called out Jonathan as someone who is co-captain on this effort. And yeah, basically, how does the effort come together? So I think there were four captains for the IMO, two from London. Jonathan was from Mountain View. I was from Singapore. So I think four of us basically train this model together. And I think one, I was also trying to see what Tang was saying, but I think one interesting thing was that we're all in different time zones and we're all...
22:25And there's something also very interesting about passing on the job. There's no really fixed workflow how to work together between captains. So it's more like, oh, I'm going to board the plane now, I'll be AFK for 12 hours. So it's romantic. So you're just babysitting the run. Sometimes there are bugs and stuff, the job comes down sometimes. So basically it's very ad hoc and it's very, It's really between the captains how we decide to work together. And yeah, but I think it was a kind of interesting time also because we were all flying. Like I think the London folks were not having to fly, but I had to fly and Jonathan had to fly.
22:57And then like when you visit another country, you have another, like if you visit another office, you have many meetings. So I was in and out meetings and it was pretty interesting. And I also think that nobody really knew whether we would get a go at that time because the IMO actually hasn't happened. yeah it was interesting exciting and then i think like the whole process of this getting verified by the imo committee and you know like you know there was like okay but we're not going there but i had to learn a lot about how the imo works right apparently the goal score is not even it's not a fixed number it's like a bell curve right so it was like a time where you just look at the score you like like i was even like looking at the watching the human participants and then seeing like what the scores were because whether Gemini will get gold depends on like how the humans do.
23:43So you like looking at oh, if a certain percentage you're like, what's the, do we get that? To some extent you don't have any control over that, so. Yeah, but you're just curious, right? Because I would say that it's definitely more like exciting. Like there's more adrenaline than like just running on a benchmark and getting a number is also like a process that took some time. Yeah, yeah. But I think overall if you have specific questions you can ask also. but I think this whole thing has been a highlight for me. This IMO effort has been... Yeah, I would say most people, if you ask them maybe two years ago whether a model could get an IMO goal, they would have said that could be possible.
24:19Then the silver helped, right? From last year. But like the fact that you can throw that system completely away and then just take existing Gemini and scale up DeepThink and then just run it for IMO goal, I think it's also like very non-consensus compared to last year. Yeah, definitely. To some extent, I think researchers were also surprised. I wouldn't say like surprised, but like it was more like a pat on the back kind of surprised. We actually made a lot of progress in we as in collectively that all the engineers and researchers working on Gemini, there's a lot of progress being made. Let's look at how much we went in one year.
24:53Yeah. And I also think that it's just five years ago, like not two years, like five years, you just imagine the outcome. Like you just look at the state of AI now, like just generally, the IMO and the ICPC go and like also even like things like Nano Banana. If you just look at the AI progress now and five years ago, I think people would think that we already reached AGI. Some form of AGI. Some form of AGI. We're just moving. If you just travel, like you take these checkpoints and you travel back into five years ago, somehow should make a drama about this. But I think it's really quite impressive how the field has moved so quickly.
Read the full transcript
25:22Yeah. Yeah. The hard parts, you would say, were scaling inference. In what aspect? Like hard in terms of? Like even expensive? Hard as in maybe the most amount of brain power expanded on the team. I saw some comments where they were like, actually the hardest part was the inference optimization or like the very, very long horizon inference that DeepLang needed compared to normal Gemini. Stuff like that. I didn't work on the inference time scale. Yeah. I wouldn't know. Yeah. That is mostly that old. And then there was this, the codename was apparently IMOCAT, which you named after your desk. Okay, that's not really, like it was in the, so I think I tweeted about it at some point, right?
26:03Yeah. So the IMO cat was basically like, okay, it's not like an official codename or something. It's just like the name that the config of the job was like IMO cat. That's the... You just need some kind of name. Yeah, I mean, I just like, you know, I just, I like cats and then, yeah. Yeah, fair enough. That is mostly it on IMO, unless you want to bring up anything else. We have other sort of researchy topics, but beyond, before I go into sort of researchy topics, I did want to maybe leave the floor to cover what else should people know about the reasoning effort that's going on at GDM. Let me think of where to start, please.
26:39What do people need to know? You know, it's very good. Yeah, that's what people need to know. Maybe an easy one to start with would be a lot of people were focusing on maybe academic benchmarks two years ago. Last year, maybe LLM Arena. This year, Pokemon. Pokemon is very interesting reasoning, visual reasoning and just general long horizon agent planning benchmark. And I don't know. you seem to focus on it a lot and I think Gemini did very well so obviously I think it's something that is easy to talk about. I think I'll probably should be with this and there's actually nothing specifically done for Pokemon.
27:11Of course. Yeah, of course. There's nothing specifically done for Pokemon and I think that I think Logan had this tweet recently about the recent Gemini tree on Pokemon Crystal and Pokemon Crystal is like so much more special. Yeah. I think Pokemon is like so I used to play a lot Pokemon and I'm a big Pokemon fan in general and I think it's a it's a great, like you said, it's a great long horizon benchmark and stuff like that. And I think it's good to check in once in a while on these benchmarks that like almost never get contaminated or let people actually like don't spend time to climb benchmarks.
27:46It's like kind of silly to like, people are like, okay, like, what are you working on? If some people are like, oh, I'm working on Amy, I'm working on HRE, I'm working on Pokemon Maxing or something like that. That's kind of like funny. We did interview the Cloud-based Pokemon AI, I think his name is David, and it showed serious flaws in Anthropix screen understanding vision capabilities. Yeah. Couldn't, literally couldn't tell. I'm trying to like get past this wall but you keep, just keep running into it because it doesn't know that the wall is there. And so it doesn't have any spatial reasoning at all.
28:18I mean, some of it could be like a harness, like the harness thing or like also whether the model has access to like game state information or is it compared to visual. Yeah, Cloud's implementation is very game state heavy. they dumped effectively like all the the memory of what's going on in the emulator yeah yeah I see yeah I think for I don't know whether I'm jumping off tangent or something like that I think solving Pokemon is going to be more of like how fast you solve it and then like the thing that I have not really seen so far is like whether the model can complete the Pokedex why is that?
28:50I don't think you need more challenging no complete the Pokedex is so hard you need to plan you need to like you need to search up like information like there's some things if you don't like go online basically you need to like have a little bit of deep research in this. The model would just never know that it needs to trade. If it's able to go online, post on forums, and then find someone to, like, hey, can I trade with you? Is Pokemon to evolve? Some Pokemon needs to be traded to evolve. Yes, or mated. Yeah, but anyway, I don't know. I have not seen the model be able to complete the Pokedex.
29:19But complete the Pokedex is really hard, actually, for models. I think that's actually an interesting one. Yeah. yeah i wonder what the real world analogy would be once let's if let's say we have a model that is capable of doing that what can we make it do that we cannot do today but there's a lot of planning involved like just real deep research like planning there's also a lot of planning involved and it's more i think completely the pokemon game is very linear right the thing the is involved a lot of like backtracking research yeah a lot of research and a lot of yeah so it's probably a different nature itself.
29:56Is that as interesting to you as, for example, a lot of other people in the AI for science world are trying to discover things that you cannot look up, right? Normal knowledge. Yeah, normal knowledge. Because basically what you're saying is we're not even there yet. We're at the place where models cannot consistently apply knowledge that they look up, right? Like you give Gemini access to a web search and you say, okay, try to collect all the Pokemon in a Pokedex. you don't have high confidence that it will do it. I don't know if someone's actually tried. Probably not, right? No, I think the hard part is actually like trying to use the like synthesize the web knowledge and then apply it in in the game itself with all that visual state going on and stuff like that.
30:35It probably would be solved in one or two. Yeah, yeah. It's not... It is challenging. It's not like super... It's interesting. Like you're basically just... The task really is can you look up the guide to do it and then can you apply the guide? That's it. You know what's even more intelligent than that, creating the guide. Like, being the first to figure out how to create the guide. Which is what it is. Oh, yeah, yeah. But then when it comes to this, it's mostly that's an exhaustive search thing. And model just try and try, like, humans and guides for that. Okay, so that's actually less interesting to you.
31:05Interesting. Okay, actually, when you think about it, it's not super, super interesting, but it's okay. It's just like, I have not seen the model to try to do this. Yeah, I think, like, efficient search of a novel idea space is interesting. Obviously, you can brute force anything, but we're not talking about bootforsing, we're talking about trying to create an AI scientist. But novel knowledge is actually an interesting thing that I think is going to be quite a big thing. Being able to generate novel... Google has done stuff there which I don't... You're probably not that close to those teams that has done AI scientist work.
31:36There's some things that have been... It's like, for example, if you freeze the model weights at like 2015, you freeze time at 2015 and then you... Even with the current model, let's say you have... Okay, let's assume there's no leaking of information somehow. Can... if you ask the model, what's the best ML? Okay, not 2015, like 2012 or something. It would just tell you that SVMs are like the best, right? This is the way that machine learning works in general, right? And then the question is, can you invent the transformer? Right? It might not be able to, like even today models, they might not even be able to invent the transformer.
32:05Like if you freeze the time at a certain time and even you bring the tech, I mean, the model is a transformer. So I just say there's no, assuming there's no leakage. That's like, no, it's probably possible. So I think there's still a lot of open questions as well, like whether the model can really, you know, innovate and generate like really noble knowledge. Yeah. One related question on that, which I think is related to the Danny paper, which is, I think people have this sort of mythicism on what reasoning is. And reasoning, if you really demystify it a lot, it's whatever happens inside the chain of thought tags, right?
32:39And you post, you're eliciting that reasoning behavior from some stuff that is already latent inside of the pre-trained corpus is that that's one version of this interpretation. I think these days, like reasoning itself is very vague and it's very open. So it's like mostly different people will have different definitions of what reasoning is, right? So I agree that like this chain of thought is like basically when people think of reasoning, they associate with chain of thought and obviously it's what happens in the thinking and then, okay, reasoning, right? But I think these days it's more like, I said in the earlier part, it's like reasoning in RL is almost like this, like basically it's anything that is like post-training to at least capabilities, basically.
33:21It's like RL and post-training to at least capabilities. So I think the actual like technical definition of reasoning is making models better with thinking and post-training. Okay? Yeah. So basically like RLing the model to think better. Right? And thinking is more like thinking tracers and thought trajectories and stuff like that. Right? There's also this line of work for latent thinking and stuff like that. Like when latent thinking and discrete token thinking is like going to be the same thing or like something like that is like an open question. Meaning adding extra tokens to your vocab that represent.
33:52There's all these, I forgot what's the name, like these academy papers that do these loopy things. Track tokens. Or like they basically, instead of decoding discrete tokens, you actually simulate this by doing this in latent space, right? So when you do chain of thought thinking and like reasoning, you basically decode extra tokens, hide it in the thinking tag and then you decode stuff. But like latent thinking is basically you just don't decode tokens you just don't bother buy it it might start speaking the native language of thinking is numbers not passing it through some filter of English and sometimes you might start thinking Chinese or something else yeah yeah generally I'm not I'm not really I don't really believe that model thoughts have to be the same with human thoughts I'm actually like generally in ML I'm more of the school of thought of let the model do whatever it wants in general there was a discussion there's a latent representation hypothesis paper that I think you're maybe sympathetic to if you haven't already read it to me some obvious basically image models will have the same idea what a laptop is versus a text model will have they converge on like the same latent yeah and obviously you can align them and you can do all those stuff with them and so it totally makes sense that their concept would be just a vector of numbers that represents laptop that's the concept yeah and okay maybe you have some numerical differences between one model's idea what a laptop is versus another but it mostly would be the same yeah very interesting the question i was kind of leading into was that because there's now we're in this age where double-lm text is in the corpus of like stuff that we train on where it's a little bit of a recursive loop right like the reasoning tokens are out there now and so pre-trained models themselves pre-trained base models are also capable of reasoning and they're increasingly so as more and more reasoning text goes into the corpus.
35:38Isn't that interesting? Or is that worrying? Do you actually see much reasoning trace on the internet? So... I've never seen those though. I would say that... Like on Hug Your Face? Yeah, people are publishing that specifically. As to whether or not people, researchers are actually including that in their training corpuses, who knows, right? Like, that's their choice. But I would say that percentage on Common Crawl that has COT tokens in there went from 0 to 0.001%. And it will just go up over time because people are publishing it. Yeah, but I think if the sources are quite clear, you can actually filter away those because usually people put it on GitHub.
36:19Do you want to filter it? Maybe you don't. There's a choice for the... Yeah, quite literally. The whole reason why... I don't think we covered this in our previous part, but two years ago, a lot of people were like, oh, you just include more coding tokens in your pre-trained corpus. it will be but the coding tokens are different from like coding tokens it generalizes outside of code for reasoning oh that was like you don't believe that no no no as in like I don't know if it's still true today but let's see yeah that was just our general coverage of reasoning I would say that there's a lot of interesting work here and more to do maybe I'll cover one thing which I know that you have personal inputs on which is that you have started using AI coding oh yeah so I actually don't really use much AI coding in the past but I think we've reached a point where AI coding has started to become really useful.
37:08Like, okay, so before AI coding, the thing that I find the most useful about, like, these models in general is, like, when I have these big spreadsheets of a lot of results and I just need plots of it. I think models can quite go through the screenshot and make a plot of this. I hate making this method-like stuff about. It's so annoying. Okay, but that's basically, like, one thing that I can remember about, like, how I used AI in the past. but I think AI coding has started to become the point where I run a job I get a bug I almost don't look at the bug I place it into like anti-gravity and like I throw it that will fix the bug for me and then I relaunch the job that like beyond like vibe coding it's more like vibe training vibe ML or something like that I would say it does pretty well most of the time and it's actually there are classes of problems that it's just generally I know this is actually really good for and in fact maybe probably better then I would have to spend 20 minutes to find, like figure out the issue and then fix the thing and then redo.
38:06So yeah, that's very interesting because I would say like level one vibe coding is you actually know what to do. You're just too lazy. Yeah, it's just, I just do it for me. Like I've done this a thousand times. Like just go fix it. Like I know exactly what to do. Here you're saying it's like the next level where you actually don't even know. It's investigating it for you. As long as like the answer looks right, you'll just ship it. At the start, I was a bit like, I did check it. and look at everything and then at some point I'm like okay maybe the model looks better than me so I'm just going to let it do its stuff and then I will relaunch the job based on the fix that the model gave me and I think the models will just keep getting better and better so yeah it's something that I also think that recently there's this anti-gravity I think also because these tools were not like that in Google infrastructure it's not that easy to you don't I'm not that familiar with what is available outside and when I was at the start I didn't really I think the models were not like so good like one and a half years ago So it's also like a forcing function that like it's also people are like, oh, try anti-gravity is a game changer and stuff.
39:07And like, okay, so I just started using and yeah. Yeah, you spent some time with Varun recently. What did you, what did it? Oh no, I really did say hi and great. I guess you were telling me that you're an AI researcher that doesn't even use much AI. And like now you're actually like AI pill as a user. There were so many moments this year where AI suddenly crossed that like the immersion thing. So I think AI coding is one of them that we just discussed. I think like Nano Banana also got to the point where I usually like if you make these images it's like for fun it's just troll your friend or something like that but like Nano Banana actually really got so good that you can use it for charts that you can use it for like basically yeah so it's getting really good and I think yeah this year the stuff like and even things like the past many of these LMs will like hallucinate things a lot but now I just trust it like automatically I think we just people are just enjoying the utility by yeah by these models.
39:59So now I'm like, I was always AI fuel. AI is a good thing. I don't see how anybody can disagree with that. Yeah, but you are actually using it for things that you are high expertise in, which is your own ML work. Yeah, yeah, yeah. And just to come back, do you have a special version of Gemini that you use internally that we don't have access to? Or is it public Gemini? I think it's the public Gemini. Okay. I was just saying, it would be entirely reasonable to train Gemini internal for only your code base and your work? Oh, actually, I'm not sure though. You see, these things are just like, I'd rather wait for me.
40:35But obviously, if it obviously improves your productivity by, I don't know, 10%, yeah, worth it, right? So I think that's interesting. And there's the interesting thing, levels of how much do you trust it? How much of your jobs do you automate away and no longer need? There's also the question, I guess, about how people come up and train in a field if you no longer need juniors? Because the Gemini is your junior ML researcher. So I think these are all interesting questions. I want to say one quick thing first, right? So I think that when it comes to like whether a model can be like a junior suite or like something like that, right?
41:10So I think if you think of it at this way of if a job from a one X suite, one time suite can be replaced by a model itself. But let's say you are a manager, right? The objective that the metric you track is like your time. And then if you can have a model that saves you like the same amount of time as the work that your reports do but you don't actually like replace one person per se but you yeah a little bit from everybody yeah right then you can I definitely agree that okay like when you count the net time save there are times where the model can fix bugs that like would have cost me like one day right one day is huge and these things are definitely like if you I don't know whether anybody has done any like real like metric evaluation on this type of things but if you use time as a real metric and then not as a number of okay maybe three hours like kind of magic, right?
41:54But these things are not like going to replace 1 % as it is, but more like a passive aura that buffs everybody. The in-game terms, right? I often think of myself as a bard because I tell stories and I plus everybody around me. That's an ideal situation for me in a D &D group. Oh, okay. I don't play D &D, but okay. I get it. They're the kings of Passamora. Support heroes. Support, support. Yes. Okay, AI support, I think is very encouraging. I think like where is it still not working for you that you've tried and you're like oh man I expected it to be better there are times when models try to get lazy and try to fix something they are still they get lazy and then they try to guess like me into thinking that like the bug is fixed so there are still classes of problems that are like very easy for the model very hard for humans there are some things that are very easy for humans very hard for model or reverse paradox and stuff so it's still very hard to characterize these things into this proper quadrants and stuff like that.
42:54So I would say that the capabilities of the models these days are good enough to be like really helpful, but like it's still, it's a bit like it still has some, but yeah, but I think this world, I don't think there's anything that to be done to specifically like focus fire. And these things is more like general capability improvements. The model just gets better over time and then these things will just like go away. You say that, okay, so yes, I think obviously in the grand scheme of things, just trust the process, keep scaling in every dimension and things will just fall away, things will emerge.
43:26But you've also said in the past, I can't remember the exact tweet where you were like, each additional data set compounds over time. They're just small additions. And I would say that when you say things like focus fire on things that you would think humans, it's easy for humans, hard for machines. Those are easy wins where you can just add a data set that would focus fire on that. And isn't hill climbing just a sequence of doing that until you reach AGI? Okay, so I get your point. I think that it's true that sometimes a lot of progress on the whole is just a series of small incremental changes that push.
44:01I think that's accurate. That's true. It also feels that there's also a lot of small, seemingly minor for the lack of a better word, that push AI to the state where it is today. So I definitely agree. So nothing against people who focus fire. But it's just that when I mean that, it might be not easy to focus fire on things that are not very easy to characterize. So it's just that when there's something targeted, like, okay, I want to improve this capability, add some data. So I think if I was defining the problems and stuff, it's like characterizing it. And if it can be characterized, then okay, then fine.
44:37But I think what I was trying to say with the coding is that these things are not even, some of the class of problems are like, I don't work on coding, but like people maybe work according they know like maybe they have like terminology for different types of failures but so maybe somewhere somebody is focusing fire on this then make the model better that's great for everybody yeah I mean that's why it takes a thousand people to like get all these things together AI is definitely like a big collective effort these days it's a big machine it's really crazy okay so I just wanted to broaden out to general things people are talking about in the community on research which again I know that you are very locked in so you don't necessarily have read all the papers or anything, but we can just riff on ideas.
45:17You can obviously ask me what I think as well. Is attention all you need? So attention and transformer is a core idea in the recent times. Pre-training and scale is the thing that made attention and transformers actually shine, right? Because without... I think the first transformer paper was this machine translation thing and then basically LGBT and BERT were the ones that actually showed the full big potential of this idea. So in terms of is international really, really all we need? Probably no. But I think it's like one of the... From the architectural point of view, also maybe no. But it's not all you need.
45:58But you need it definitely. What else are you thinking about on that same level? Are you talking about MOE stuff? What do you mean when you say it's not all you need? You definitely need the scale of pre-training. You need all the tokens. You need like I think when I say, when people say, I say it's attentional you need, it's mostly, from an architectural point of view? Well, will transformers get us all the way to AGI, right? I guess it's the... So basically, when you get to AGI, it's the problem. Will it still be a GNOME architecture or like meaningfully different? It'll be a transformer, I think.
46:27Really? Like people, it depends on what you call it. But I think unless the paradigm shifts completely, which is, I mean, as a scientist, you cannot like completely say no to like, this will never happen. but my feeling is that it's been like what like almost 10 years since the transformer 2017 9 years since the transformer I think we have not replaced self-retention like it's some form of it like you could rename it you could name it something else you could sometimes you can do local global yeah it's still a transformer in the end and I think like that's not going like anyway unless the whole thing with like back prop like everything like goes like the whole thing just changes completely like then there's a different story there's a different conversation to have but if it's still within the same scope and bounds so I spent a lot of time thinking of about architectures and like whether there's alternate architectures and stuff like that okay at the sequence processing level like there is the ultimate yeah it's sequence to sequence transformer it's probably the self-attentional there was this whole big era which I was also involved in this era where people try to like undermine the attention as much as possible like they try to remove it simplify it make it efficient like this whole like efficient attention era at the end of the day the outcome was always like oh we removed all the attention but we have one layer of self-attention that still works like that's at the end like always the story which even Noam a character he published some stuff about how he has some ratio of mixing of local and global attention right like basically still attention but modifying it quite a lot I would consider local and global attention to be like still attention just like how much you're skipping here.
48:07Yeah, the only question is that if the formulation changes too much, your QKV becomes like ABCDEFG or something like that. Okay, maybe I'll give you some motivating constraints in order to do this. You guys are still charging 2x for over 180k token context or 240k, something like that. And the max, theoretical max is 2 million tokens, right? What if we need 200 million? Is there some point at which where even this concept of input token context is irrelevant because you are doing continual learning, that kind of stuff, where you're modeling it as, okay, the AGI will be achieved through a sequence-to-sequence transformation.
48:47Therefore, an attention is the best sequence-to-sequence model or architecture. Therefore, attention is all you need. But I think other people are like, sequence-to-sequence doesn't accurately capture intelligence. But that's not really about sequence-to-sequence. It's more about the whole gradient design and backproc thing, right? It's not the architecture itself. That's a problem that is more of like the learning paradigm itself rather than the architecture itself. I think the architecture is just basically like the interface between the learning algorithm and the tokens. I think it's more about the learning algorithm itself.
49:20And this, like, continuing learning this, like, there's many ways to think about processing many, like, insanely large, like, contexts, right? Like, 200 million, 1 billion tokens or something like that, right? Like, whether it's going to be like, you have a new learning algorithm that every time you run inference, you learn on it, right? Then you can technically have some kind of memory, like human beings learning as I'm talking to you, right? So that's also like one way. The other way is like whether, okay, maybe somebody will say that, okay, the attention is like just too expensive for 200 million, 1 billion contacts.
49:52So we need new architecture. Or some people will say that, oh, we just improved the chips, like accelerators. So I think many ways to interpret it, but I think it's like, if it's about, there's a lot of fundamental things that if it's about continual learning and stuff, there's a lot of fundamental things about the learning algorithm and stuff as well. That will have to change? I think the learning paradigm and architecture and stuff that goes hand in hand. And I think as the field progress and ideas just stack on top of one another, so there's also this thing about an idea that was proposed has to be compatible with all the work that had been done before to shine.
50:24It's a bit of a variant of this hardware lottery by Sarah. It's not the hardware lottery that I wrote about the GPUs failing, but it's the original a hardware lottery but it's more like a bit of lottery of like the things proposed have to play well with the things that were proposed before so it's a bit like going down this local optim minima to some extent so now we are like in this local minima of like transformers everything everything right maybe it's not easy to like get totally out of this because also a lot of people's investment optimization have been done so the things that play well need to play well with the ideas before and the way I see it now is it's very difficult to like come out of it Okay, I'm not entirely convinced.
51:05I see what you're saying, but let's call it Gen AI. I fucking hate that term. It's still a very young few. And so, yes, there's been eight years of work on the transformer, but what's that in the grand scheme of things? Maybe we're in a local minima and we've got to notch ourselves out of it. I do want to leave that open-ended. I don't have an idea. I do think that people are in what Ilya Suskiver has been calling the age of research, right? Like where we're like, okay, we scaled up what we can scale up. we know what the next maybe one, two orders of magnitude look like in scaling on every dimension that we know about but what is the next dimension to scale?
51:41There's this
51:45misunderstanding a little bit about all the last five years has just been scaling things. Okay, please tell me more. You made that joke about now we scale researcher salaries. Okay, let's not go there. But I think that ideas matter and I think that there have been a lot of good ideas in the last five years. It's just that maybe it's just not... So it's not been like blindly... Like if you took an MLP, right, just like without self-attention and you're just, okay, I'm going to throw like$100 trillion on this and scale up that thing and the thing... It's never going to work. It's never going to work.
52:16Yeah, yeah. So there's no... Like there's part of it that's also like... I think the bitter lesson gets used too much in... Like too conveniently used around. But actually there's also a little bit of a... Not a bitter... There's also a sweet lesson where it's like ideas matter. And I think even to today, right, like people downplay ideas and stuff like that. Do you think the rate of new ideas, without being specific about what ideas, because obviously you can't share, but do you think the rate of ideas has increased or decreased because there's like kind of a law of diminishing returns? Are they smaller?
52:45I think the number of ideas is always proportional to the number of researchers working on a certain problem. So by definition, it should increase. But I think the number of ideas that actually work is not decreasing compared to the last, like we're not in the era of diminishing returns yet. So I think the ideas are still very important and there's still very good ideas that are game changers that are being invented. Yeah. And I think I know the answers it is, but is the closed lab advantage increasing versus open source or decreasing? Like the Chinese labs, they say, keep publishing open source models and some of the American labs as well publish open source models.
53:24would you say that the ideas that i see there nvidia has nemotron opening eye has gpt oss these are all basically checkpoints on what's publicly known about training models as of this year you know okay okay it's declassified information because everyone yeah okay yeah everyone does this i think that the gap is increasing i don't think it's completely predictable from stuff that you said before i think the gap is definitely increasing yeah I think that justifies researchers. Otherwise, what's the point of having researchers if not finding new tricks that compound over time? Yeah, but definitely I think it's increasing.
54:05Okay, I'll do a side tangent. I don't know if you have any comments on this. So this is very related to NVIDIA's recent purchase of Grok, which I don't know if you have views because you're very TPU-centric, but are we memory or compute bound? and this is relevant to the transformers discussion of like in terms of what like serving exactly i think the classic view is that we're compute bound because we just need more compute for pre-trade and rl and then inference yeah but actually the counter argument i would make against this is i actually have these charts of moore's laws i wish i could just pull it up easily moore's laws of the scaling of compute versus scaling of memory versus scaling of network and bandwidth.
54:48And compute has a much higher slope of scaling than the other two. Memory, I mean, the chip memory? Like, honestly, I don't think about this memory bound that much, so maybe it doesn't... yeah so i would disagree with it but i don't have high confidence in and because you're mostly on the research side, less on the inference side. Maybe the inference guys will be like... Yeah, yeah, yeah. I don't think about... I don't wake up when I think about serving. So, yeah, maybe I don't think about the inference that much. My previous line of discussion here was like, NVIDIA is very foresighted by Melanox because it actually is the real bottleneck in scaling because it has the lowest Moore's Law.
55:25And then the second one now is memory, which is very interesting. Okay, okay. But honestly, I don't think about a dozen that much. I understand. Okay, data efficiency. So this is a joke, but implicit in this is that there's some kind of a maximum data exposure, right? And so previously, I would say that a lot of the training paradigms is like one epoch is all you need. It's a mini title of this idea. I would say that the real number maybe is between three to four epochs. And I do wonder what the theoretical limit of data efficiency of a model in terms of training and compression should be. I don't know what that means.
56:07Data efficiency, but you're asking the question in a way of asking how much repeats is tolerable? Tolerable is contingent on does it actually improve in meaning. It's not about you actually want to do it for its own sake. But I do think there's that. And then there's also just the sheer amount of stuff that we can learn with limited data. So you say, let's say you're not compute bound, you're not memory bound, but let's say you are data bound. Last time that we were on the podcast, we talked about chinchilla versus inference optimal training. But now actually, I think a lot of people are even talking about data optimal training.
56:43Like given limited data set, how well can you learn from it? I think that's an interesting research direction that not enough people are talking about. Maybe it is something that is commonplace in the labs, but it seems very clear that we are very unoptimized with regards to how much we learn from our data. I'll just put it there. I think in general, like learning more, like extracting more from varied data points is definitely valuable, but I think that's also related to the fact that we're like running off tokens in the world. So I don't work on data for pre-training and I think things that I say were general state of industry not nothing the general state of industry right so I think that I don't even know whether data has like diverged like the way that these things are done it has diverged too much across these labs and no open there's a lot of cross-pollination for sure cross-pollination okay okay yeah but I don't think about data that much like the pre-training data that much yeah maybe earlier this the first half of this year I would have said that kind of pre-training is dead and And everyone's just funneling all their work towards RL.
57:56And we had this Grok chart, which was very interesting, where we're sending the same amount of compute on Cygrain as to... You think it's a Cylog? No, I don't know. I have no idea. I think that people are taking it seriously. They are like, yeah, okay, whatever. Especially in the agent labs like Cognition Cursor, they're taking the open source models from whoever and then adding, let's call it pre-trained scale RL on top of it. if they have that level of data which they do, which is very interesting. I would say, yeah, this data efficiency argument, yeah, I think to me it's also more trying to discover new paradigms of learning in order to get where we want to all go, which is, yeah.
58:40And the existence proof is humans, right? Your two-year-old daughter is much more capable than an LLM in some things, having seen way, like, eight orders of magnitude, less data. That's very interesting. yeah comparing human learning and machine learning is definitely like purely as an existence proof that we could probably do better three examples of dog yeah fourth example of unidentified animal I can probably tell it's a dog as a human but machines classically you take 20 the idea of efficiency of humans is definitely way higher than models yeah the only question is that where does this thing come from is it actually like putting more flops on every token or like maybe it's like back to the question about whether the transformer is the optimal architecture, maybe it's the backbone, maybe it's the off-policy-ness, maybe it's the...
59:27So what is the... Like, where is the bug, right? Exactly, exactly. But maybe it's a feature, not a bug, I don't know. So, okay, we've identified, probably it took me a while to get this across. This is the kind of data efficiency I'm talking about. I think it's emerging. Basically, at the end of every year, I try to take bets as to, okay, what will be the big theme for next year? I think this is one of them that people are really trying to focus on. Because you're feeling this data crunch, even though everyone's like still investing in data i forgot to mention that i would say that i've been wrong on pre-training being dead yes i've now met pre-training leads from both anthropic and of the eye and i've seen the talk from the deep mind guy recently and so everyone's investing in pre-training still which is like nice to see nobody said pre-training was dead i know no it's a theory that we're trying to disprove or prove yeah anyway so i think okay let me wind back to my general idea, right?
1:00:21So yeah, data efficiency seems worthwhile. You would treat it as like, okay, well, show me where the bug is and I'll go fix it. We don't know where the bug is. We just have existence proof that it could be better. And then I think the final logical chain in this for me is that everyone is focusing on some idea world model as a version of this for more efficient learning, which potentially might not take the form of a sequence-to-sequence transformer. I don't know how that works. Like definitely a little bit out of my depth here. To me, that is more efficient because every world must be internally consistent.
1:00:58And if the next piece of evidence come in and invalidates those worlds, then you no longer need to pursue those paths ever. And you can just narrow in on the world that you've identified. And so to me, that is learning where you're learning to fit world models. Or to the actual data. Yes. So yes, maybe you can treat the learning process as curve fitting yeah so you're learning the world instead of learning the world model yeah i'm learning the world model right okay by sampling multiple world models and then finding out which one fits the data the best so i guess my query is this what people talk about is this if you i mean obviously feel free to attack it because i'm just spitballing but this is what i pick up from talking to multiple people about okay what are you talking about with world models what are you talking about data efficiency and learning efficiency and like how do you gel it all together in a cohesive sense of the future where we can actually what's the definition of world model at the start from the start yeah there are three kinds okay okay go on yeah first kind is the VO kind VO kind or the what's the other one genie that DeepMind has which is the sort of video world model yeah you model everything with some kind of Gaussian splats or whatever and you like you inhabit those that 3D space yeah second let's call it the Ian LeCun slash meta school of thought which I don't know if you're that familiar with it.
1:02:16He has published the JEPA architecture and then separately, Fair has also published the Code World Models where you're basically, specifically for code, very interesting. You are executing code and modeling the internal state of the execution environment as you go line by line. Okay. The LM actually learns to predict those things. Yeah. And actually, it seems a lot more efficient at the scale that they've tested it out, which is very cool. Which definition are you anchoring on? The third one. The Code World Model? That's the second one. Oh, it's the second one. Those two are bundled together. Yeah.
1:02:45The Jepa fans are probably hating me right now because I'm lumping all metals work under one school of thought. Yeah. But whatever. Okay. Okay. The first one is VO, Genie, if it is like super spatial intelligence, those kind of video based world models. Second one is some execution or some sort of explicit modeling in the, as you sort of run through the corpus. Yeah. And then I think that the third one is this amorphous thing, which I think people are trying to get to where they are doing what I said about the resolution of possible worlds and your curve fitting as you learn, as you inference.
1:03:21Yeah, but what is the world model itself? Is it like... It is a mental model of where everything is and how you think the world works. What I think you think, everything. Okay, but technically it's like... It is something in the latent space. Okay, okay. So you can, for simplicity, it could just be like a transformer model between and... Yes. Yes, so to me that is the most coherent thing to the current paradigm, which is you could actually do this in current transformers. I think the way that you train it will probably have to be different. Okay, I see, I see. I don't have any conclusion here.
1:03:55I'm just throwing it out as something where I know you're interested in this kind of stuff and I don't have that many knowledgeable people to talk to about it. No, I don't think about world models that often. I think because world models are just not really well defined in the first place. So don't say world models, but the problem is learning efficiency, and maybe, I guess, accuracy or like AGI capability that is not easily unlocked right now on our current path of scaling. Yeah, so I think when it comes to like data efficiency, I think it's more like I'm believer of finding ways to spend more flops per token, right?
1:04:30Because you actually, basically, if you are data bound, you want higher data efficiency because you can learn more from every data point. It's squeezed out more points, right? So things like that can extract more, can use more flops on every token. It's definitely like a form of data efficiency. Then there's the learning algorithm, right? Because I think there's this, there's a different scaling law for like humans is this, machines is this, like dogs are this, cats are this, there's this different. Nice exponent. There's the ALEA chart. Right, yeah, yeah. There's this famous chart. And point one and point two are just like not entirely like the different things because it could be that better architecture is actually just spending more flops per token.
1:05:09So if you come to a point where you are, very data bound but not compute bound or you just find algorithms that spend a lot of compute on every token on every token so I think the overarching point is just that okay it's a learning algorithm thing for data efficiency and then if whether the correct way is actually just apply more flops per token just to squeeze out more from every data every data point also because humans actually don't like they are exposed less when you say less or more data it's also very ambiguous because they are technically like on 24, 7 seven and then you have a lot of like different types of inputs right and whether they actually spend more flops on everything that listen is also a question because maybe they are just better efficient just because i mean i didn't somebody needs to count like how much flops the brain used to process like how much maybe they're just spending more compute on every token and also maybe the learning algorithm is different so but i agree that data efficiency is very important given that that i think we're going to like there's limited amount of data in the world one more thing before we go into dsi you know how like we're talking about rl and like you're working on rl stuff why are people paying so much for rl environments wait so who's paying for rl environments open ai and topic at least i don't nobody say anything about deep mind so a lot of the model labs that are not you are well known for paying at least seven figures for external startups to create RL environments for them to train in.
1:06:34Okay. And I think the question is, if your models are so good at coding, what do you do yourself? And so I think there's some amount of expertise that's being distilled from human experts into an RL environment that you can then let your agents run wild in. But I'm curious if there's any other deeper insight than that because I'm not satisfied with my own explanation. RL environments that are like, they have a lot of domain expertise are probably very valuable. And actually, I don't know specifically about what RR environments people are actually explicitly buying. But what was the thing that you're not satisfied by?
1:07:10It was so valuable. And a lot of people are saying like, look, it's a next year's app inside of a Docker container that logs stuff out when you send inputs in. Then you could probably do it yourself internally, right? Why you pay so much for some startup that you don't know to do it for you? Actually, I have no clue about why Why this is happening, yeah. I have no clue. And a classic example would be like, if you want to build a computer use agent for buying things in e-commerce, you would want our environments that perfectly replicate maybe the top thousand e-commerce websites. Yeah. And then you just parallel roll out on all of them.
1:07:47Does that seem meaningful? I don't know. All right, cool. DSI and LLM Rexis. A big bet for me this year for my conference was we actually started focusing on LLM Rexis The other... Actually, what's the motivation behind starting out on Rexis? Correct. Yeah. I think Rexis is the king AI problem in consumer. It is the single most valuable thing. All your feeds, any even search is Rexis. Basically, it's search, basically. Like, it's retrieval. It's the God problem, right? Because Rexis is ranking, but then also filtering, also personalization, also re-indexing and performance. It is the God problem and you get paid a lot for it.
1:08:27engineers are not that excited by it, which is very weird because they don't see, a lot of them don't work on Rexis and they probably never will. Yeah. But they don't see the monetary value that can come out of a good Rexis. The other two pieces of updates for me, which I actually didn't even know that DSI directly tied into this, was one, Twitter publicly adopted their feed algorithm as an LM Rexis. LM's are just used everywhere now, but whether it's actually like a... Like a big LLM. Like whether it's like a generative retrieval type of model. It's like another question. Correct. Is it? We don't know.
1:09:04All we know is that they have said that they have swapped out their current RECs for an LLM-based RECs. That's how they found it. Okay, okay. But what is published is YouTube where they actually adopted semantic IDs for YouTube's RECs. Yeah. And YouTube is obviously a big deal. Is it like public information? Yes. Okay, okay. They came and did talk about it with us. Okay. And then they published a V2 this year as well. Okay, I see. more info and i just thought so basically the last time we were on the podcast we didn't talk about dsi that much but you have actually some background in ir you care about ir i don't care about ir but i think dsi okay like dsi or genitive retrieval was like i think one of my favorite works in the old of like i have some ir background in like when i was doing a phd i did some rexist work i did some retrieval work with rexist and stuff so i have some ir and rexist background So I think generative retrieval and generative rexies is very conflated.
1:09:59DSI started as a retrieval thing. So we did like natural questions like ranking of documents and everything. It started off as I think that's actually we did an interview with Yannick like me and Don we did an interview when the paper came out like a long time ago. So at that time we wanted to like reimagine retrieval and search. So we wanted at that time LAMP we were still using T5 models at that time it was like not we were not in the LAM era yet It was pre-LMV, right? It was like, okay, pre-training works kind of thing. And then there were like some pre-training models around. So we wanted to reimagine retrieval, right?
1:10:29But retrieval rexists, they're all the same formulation, ranking, retrieval problem, right? And then that's where we started to imagine retrieval as one giant, right? And the ankle is everything in the memory. But we tried so many different, like so many ideas, actually, my collaborator, Vin, was the one that came up with either so many ideas that basically, and the start of this whole genetic retrieval was actually basically literally just trying to give a document like identifier and just predicting like raw brute force predicting like this. It actually works because the models can memorize something.
1:11:01If you look at the literature from all the way to things like Doc2Vac, it's very often like the words have no meaning to this ID in a vocab, but it's an honest number, right? And technically, the models have enough capacity to predict. But I think semantic ID was an idea that basically you have some like semantic association and then you actually try to break down the search space hierarchically, right? So how this would evolve into RACSYS was at a time after DSI came out, right? So, Ad Cheese Group and Mahesh, the guy who, they did some exploration of applying DSI to RACSYS and that's how that generative regenerative RACSYS recommend system paper came out.
1:11:42Yeah, I didn't even know he was involved. It's crazy. That was like, basically us transferring this, like basically, okay, DSI will try to try it on on Reksis. And then I think the recommended system people have a slightly different way of doing semantic IDs, but it's basically just because the domain is slightly different. But after that, I think we will join with the invention part and with this one. The rest is details. The rest are details. So over time, I also left Google and stuff. So over time, these things evolved a little bit here. I think I also saw something like Spotify is also using something like YouTube Spotify.
1:12:16They use this type of semantic IDs, this type of DSI-like models. I think from the research community point of view like the DSI work was the first one that like decodes semantic tokens but then when we went to I don't know academic community is like strange in a way that like they will do things like oh this is gender retrieval it's not gendered rexies it's like they will do this kind of like random things that is like a bit strange but yeah I know this was like the whole history of this gendered retrieval apparently there's also a lot like of people working on I don't follow actually I don't follow this at all now it's just not even in my mind but there was once I went to even in the Singapore office they are sui's actually like working on generative retrieval they don't work on I don't know whether they're still working on it but I met a person that tried to explain generative retrieval to me it was quite funny that you know I kind of like co-invented generative retrieval but yeah I think this is this whole IRR thing has been it's just an interesting phase and I definitely think DSI is one of my more creative works that I've done that it's not like really LM but it's like under the general principle of apply ML to everything if the Googler is working on generative retrieval would that be like AI overviews?
1:13:20Is that something similar? I have no idea. Okay. For people listening, I did have a track there. I think you just type in AI.engineer and you'll get it. Where the Gemini guy was talking, sorry, the YouTube guy was talking about how to use Gemini or the Rexis. I don't know what size of Gemini because he didn't talk about it. But this is public work now and basically every YouTube video uploaded gets encoded into some kind of codebook and they retrain this every day on some kind of batch job. Yeah, just interesting. So yeah, I don't know if you even know what Gemini is being used for. I don't follow this, this, this, this.
1:14:00I do think like in the sense of like for people who are not still not getting it, applying intelligence and the general intelligence of an LLM to the retrieval to the recommendation task means you can accommodate such weird recommendations, like such weird queries as well that normally no classical system can ever handle. And I think it's also somewhat emergent in the sense that when you were using T5, you just couldn't actually add that much value on top of a normal BM25 retrieval technique. We'd say that's accurate. It is not just about paraphrasing, it is about understanding query intent. BM25 is a really strong baseline, actually.
1:14:41BM25 is a really strong baseline, yeah. sorry I don't know the comparative delta versus T5 for you guys versus BM25 but I don't expect it to be very high and I expect it to be a lot higher for a true LLM base axis depending obviously on the query set I didn't really think about it this way before but because I've done modeling in many different domains including like search and you know in the search community I have a community there's also a benchmark and stuff like that for like that people who climb on there's some like there's Amy of like a, I don't know what it's called these days anymore, but generally, the modeling dynamics of IR tasks is very different from, like, and Rexy's tasks is very different from standard language tasks or like vision tasks and something like that.
1:15:27When we kill client L and when you train models, like the way that modeling things interact with this environment is very different. So I think that I honestly, I hated working on like Rexy's and retrieval stuff. Okay, I'm just looking about old days. When you work on like you work on, you change architecture, you try to improve complexity, you use super glue, like this is the olden days. You are on, like, even now when you train LMS, you just do zero shot, two shot stuff like that. Your things, because as a researcher engineer, you just interact with the environment a lot by this, you're just like, okay, RL by this environment.
1:16:01But RECSIS and IRR has a very strange feeling to it. Strange feeling in the sense that it feels like you're, like whatever works, it's like you are in a world where the gravity is different or like you are in a world where the modeling things that feel intuitive are not intuitive. So it feels like a very strange space to... So I wrote some papers back in my days on like rexies and stuff like that. Every time I ran some modeling experiments for rexies and stuff like that, I didn't enjoy it. It feels like the environment was rude. It feels like the vibes are just like... What makes it rude? It just feels...
1:16:33Transactional. No, not transactional. Like, I don't know how to describe it. For example, like, if you play like sports like the tennis parameter middle. When you hit the ball, you have a very nice feeling like hitting the sweet spot. When you do modelling in traditional, when you get the feedback back, you feel like everything sounds right, everything feels right, everything, right? But Rexys and IRR is like a problem where it's like you hit the shutter clock and you hear a glass shuttle like randomly. It just feels this weird sense of humor that... Cause and effect are too far apart. Like it just feels strange.
1:17:04And then sometimes maybe the metric like I think Rexys, they use like all the NCCG effects. And then the BM25 is strong and then you just, and then you get like worse than like the BM-25. Back in the day, when you stack two LSTM to three LSTM, you were like, whoa, I see like, it's just the game. It's an unrewarding area to work in. It's just weird. Also, the IR community and the retriever community is also like always behind the mainstream. And then now it's just probably gotten even more worse because of IRM and stuff. So, okay, I'm getting into Hortec territory, but it's just, like some of the conferences are just like behind New Rips and ICML and stuff.
1:17:39Yeah, some conferences are just like, they're just like applying things that this they're downstream they're downstream okay so it always feels very uninspiring to work on on this there's a reason that you left but it was like a side quest like where I work on it as a side quest thing yeah yeah okay understand I still think it's an important business problem even though maybe it's an unrewarding field you can understand why because the academic benchmarks for those tasks are just so far detached from they're so far detached from what industry it is. I didn't work on any of this like the thing, but that's from an academic point of view.
1:18:16Oh, then all you need are online invalser, right? Yeah, yeah. And the test and the O. Ooh, okay. That would have been a different experience, yeah. That is mostly our sort of topic, research topics, coverage and everything. I think we're just going to end on a very simple one on GDM Singapore. You organize a symposium. Here we brought Jeff Dean, Kwok and all the others. Basically, what's the general message or the impetus for starting GDM Singapore? So we'll talk about the event first. So the event was mostly, so Quok and I are going to start a team. And then I think before I came back, we discussed this for some time.
1:18:51Jeff was very supportive of this. He was in the region many times in Vietnam and Singapore last like about around the time where I was going to come back. And I think that, so this event itself was, Quok and Jeff were visiting and we just inspired the community here. I think that it's also a bit more like a soft like setting the tone right for the start of the Gemini team in Singapore and I think it's a very rare instance where you get somebody like Jeff and Kuo who are the true pioneers of AI in the world to be in one room and then are you there as well? And I think like many people told me Not a true pioneer of AI but I was there to live tweet Yeah, I think having them all in one room and then giving these talks like many people came out to me and said they were very inspired by their presence in the region.
1:19:37So starting a team and starting something is also very, like there's also no one moment that it's not, okay, I press the button and it starts, right? It's like a process, right? So we hire people and then people join one by one and stuff, something like that, right? So I think this event was more, I would say, like to set the vibe. I think it's possible for Singapore to be close to the frontier. And I think that we're having the true pioneers of AI here. We want to give this, basically more like an inspiring thing and also get left for quark and jeff to meet the people here and and so jeff was here last year but quark hasn't been for some time and he's gonna have team here so it's like also nice to bring him around and meet the people here yeah so it was a really like amazing event we met long as well and who a lot of people don't know has a cs degree it's like one of the few pms with a cs degree yeah i would say that the context of the meeting was more partially also he wanted to learn more about the IMO stuff.
1:20:36Oh really? And then also about Jeff. Oh because they invited you without knowing that these guys were coming or something like that. I was a bit like Jeff and Kwok and me we went to visit the Cherif at the Istana and we discussed a little bit on Dipting, discussed a bit about IMO and then I think the rest of it was more like Jeff and Liz was talking more about like it became less about AI and more about very macro economical political thing which was I was very out of element so I was just I just you were in a suit I was just talking about the deep thing and the IMO and stuff like that but he seemed generally quite surprised that AI has reached this point so I think but it was a it was an interesting yeah I would say for people like you have done something that is unique in Singapore's history so far like you're establishing a frontier research lab in Singapore which is an accomplishment.
1:21:29I think the other thing also that I guess I'm still trying to wrap my head around is does geography actually matter? Like you're all working on a team, you have your London people, you have your Mountain View people and mostly you're just like collaborating with them anyway. You've collaborated with them your whole life. I don't even really know what countries mean anymore when it comes to research or just AI in general because this thing is just inherently international from the start? This is a very good question. Also, it's related to the thing about identity, right? Because I think you also move from SF in Singapore quite a bit, right?
1:22:04I was in Maldivu this like one, two weeks ago and I'm here, but like almost all my, if you just look at my siblings, aside from my family, like everybody I talk to is like somehow in the Bay Area or like this because of work and everything. I think the geography matters. Okay, firstly, the most like thing is like logistical it's like probably the time zone so you literally want the 24 hour coverage around the world no there are there are advantage no what I'm saying is that the difference is possibly like it's also more like like people define the location more than the location define the people somehow okay the time zone we get to the time zone a little bit like the pros and cons I think that's pros and cons right so you're bullish on Asia Singapore people the talent pool I think we managed to find like amazing amazing people but I was also have to say that this type of things is more like talent attract talent i think most of the time like people are very excited like the vibe i get is that people are very excited because it's like cox team and my team and we're working on library core things related to agi so i feel like the talent we can get from the region is really good but it's only it's only because it's us so we can unlock this talent.
1:23:18Otherwise, we might join some other place. By that. Yes, and move to the US. Yeah, yeah, yeah. About the identity-wise, I would say that I definitely agree with you that, like, why does it matter? Like, I think that the advantages of Singapore itself or, like, just anywhere, like, it's also that you are like, okay, the world is very good, so technically you can interact as much as you want. You can also go there. But I think Singapore has this advantage where you can go close and you can go far. I think the Bay Area is like so much about... I have friends in London and New York that would just never move to the Bay Area.
1:23:55I'm not against Bay Area. I think Bay Area is a great place, right? But it's just AI, AI, AI, AI, AI, everywhere. Right? I think sometimes if you have some like mental space and energy to like to have some other culture and then like, you know, like London, Singapore, New York, they have their own culture, right? But the Bay Area culture is just like AI, right? Like you just go anywhere, you just hear AI everywhere, right? even the billboards and stuff. Looking here a bit much. Yeah. Although I did see some billboards down here. Or AI also? Yeah, I was like, what is this? It's a culture-infecting C-Ball.
1:24:26I do think that to some extent, if you want to do research, you need a little bit of peace and quiet somewhere. So this island may be good for that, but then you're still able to be connected, right? So I think that's mainly... Talent-wise, I think people are strong here. yeah so far enough away but you're still connected you have strong talent what are you hiring for? you're still hiring right? we're hiring like my team will work on like RL and reasoning for Gemini and Gemini deep thing I think we care more about like talent density now so we're not like also like growing that big this small team first just because compute per capita is probably like important and yeah so I think that's something that that we're hiring for now basically I personally there's a lot of DMs but I think like generally there's a lot of people like who are very capable but I think what I'm looking for mainly is either you have like a track record of RL research or like some even not necessarily the RL but or like some exceptional achievement in like coding competitions or like some exceptional achievement somewhere then there's like the kind of people that we want yeah because you don't strictly require I do remember something about your record days where you're like you like to train your own trainers from scratch right so you don't i forgot if i said that or not but to some extent just do to start yes then yeah i think we'll definitely be very happy with people that are like very high stats and just like even without much statistical knowledge no no stats yeah yeah the stat points child with int like this high tech that point people like this raw iq yeah high tech talent people like this i think all like strong engineering skills, ML, ML can be learned easily, our knowledge can be learned easily.
1:26:16Yeah, I think maybe one version of this is, can it be done on a student budget? Can you do something interesting anymore on a student budget? I would say, relevant to the point where conferences are quieter these days, I did do an interview with one of the best paper winners, where they worked on thousand layer neural network RL. And that was done on a student budget. It was very cleanly executed pieces and paper and good findings. look, I'm not sure if production models will ever go to a thousand layers, but they stretched it in an interesting direction and found some good recommendations. And the guy immediately got hired by opening eye.
1:26:51And I think that's encouraging for the grad students in the market who are like, okay, well, do I need to know somebody who works at these labs in order to get in? My uncle works there. I get the internship or whatever. No, actually, you can just do it on a student budget with good advisors. Oh, I actually think one thing interesting is that for most of the people that I actually went to recruit them, like, personally, right? So you see the work and then you send them DMs, right? So I get a lot of... For your best paper? No, no, no, like, for hiring, generally. So generally, I almost... To your point, it's like, you almost, like, you can just do good work, put it online, and then somebody will contact you, right?
1:27:25It's actually super easy, but super hard at the same time, because... No, I can tell you, like, I talked to a few of these grad students, they don't know what good work means, right? Because they don't know... there's so many things their professors have their agenda they're forcing on them which like may not be right because it's not like their professors know what to work on either. So yeah, they just need guidance. They just need to, hey, work on these like five things. Show me an interesting result in any of them. Okay, so if somebody comes up with something and then does something that you feel that is very tasteful and it aligns with what like researchers in the labs like one and they come up with that independently you know that the function that produces is good, right?
1:28:03Like if you just go and tell somebody to do this, like you just get the signal that these guys can execute, right? So I think there's some value in people that... Yeah, they demonstrate taste, research taste. Yeah, research taste. Very interesting. I feel like I could give people... Yeah, I do care about this. In some ways, the research directions work that I do is a little bit of that. Like it's low accountability for me because obviously it's just thought experiments. But I think for a lot of people, it's like their career is bounded by can you demonstrate research taste with this like short three, four years that you have and just do it.
1:28:38Yeah. Yeah, I would say that it's more, there's so much competition just because of like the, like everybody wants to get into AI. It's just more of like how to, like mostly it's more like how you're going to prove yourself. Yeah. It must be hard these days to be a grad student trying to prove yourself. It's definitely harder, but yeah. Yeah, not your job. Okay, that was it. Do you have any other sort of rants or topics that you had queued up? before we wrap? I don't know. But it was great. I had a great time. It was fun chatting, man. Fun, fun chatting. Yeah, even, I even love, last time we were supposed to do, last time we meet at the symposium, we were supposed to record.
1:29:12But even if we just ended up hanging out and chatting, it's just nice to get the brain dump of what's going on in your world. Because, yeah, we're working on really important stuff, man. Always great to chat with you and see you here. Good to chat. Parting words on the sort of weight loss and workout journey. Because there's also a big thing for you. I think being healthy is important to be, to do good research. Right. And I think I've been probably in one of, I'm probably in the peak of physical health now. Yeah, you look great. Yeah, thanks. And I think it's also impacted my work in a good way.
1:29:43You did the sort of Kapati-inspired, like, biohacking. I didn't go so extreme, but I was, like, also quite data-driven when I came, I would, like, have my own email and then I would track this. I'm still supposed to make a blog post about this, but I feel like I'm not, like really at the end game yet. So like when I get there, but yeah, just to, just for people who don't know, I like, I think I lost 23 kilos. This year. Actually across one year. Yeah. One and a half years. Yeah. So 23. Yeah. Basically literally from the last podcast to now. Yeah. There's an abolition study now. 23 kilos. Yeah.
1:30:20And I think like my HRV heart rate variability has went up by two times. And my rising heart rate has dropped by 30 beats per minute. 30 beats per minute? It was like 80, 90, and now it's like 60. Oh, yeah. 80, 90 is super high. I was unhealthy. Yeah, yeah, yeah. Okay. Yeah. So I think when it's hard, do you have a thing that can be going? A lot of people, maybe they're focused on AI, including myself, right? I do prioritize work. I enjoy work. I don't enjoy the fitness side. But obviously, it feeds in to your intellectual work, like to sort of log off and go for a walk, eat better, all that kind of stuff.
1:30:58obviously it fits in but like people seeing a positive example like you they will get inspired to do the same thing so I think it's good to set yourself up as an example I think definitely helps like when I do these things for my health I just think that it's also part of work because it helps me to get better at my job so it's important as well I think it's important as well I like the HIV off the bat I have no idea what mine is but yeah there's a general question about what is productivity and how do you measure it what really matters and it's still unclear to me, but I do think general energy level and hunger almost, like you almost have to like experience physical hunger in order to have intellectual hunger.
1:31:40And I don't know if that's like a thing. So when I'm hungry, I just think of food. I think to me, it's like these destroying things. But it's hard to do work when you're hungry. Okay. Thank you so much. Yeah. Thanks. It's really great. Yeah. Have a great time. Thank you.
From the publisher
From shipping Gemini Deep Think and IMO Gold to launching the Reasoning and AGI team in Singapore, Yi Tay has spent the last 18 months living through the full arc of Google DeepMind's pivot from architecture research to RL-driven reasoning—watching his team go from a dozen researchers to 300+, training models that solve International Math Olympiad problems in a live competition, and building the infrastructure to scale deep thinking across every domain, and driving Gemini to the top of the leaderboards across every category. Yi Returns to dig into the inside story of the IMO effort and more!
We discuss:
Yi's path: Brain → Reka → Google DeepMind → Reasoning and AGI team Singapore, leading model training for Gemini Deep Think and IMO Gold
The IMO Gold story: four co-captains (Yi in Singapore, Jonathan in London, Jordan in Mountain View, and Tong leading the overall effort), training the checkpoint in ~1 week, live competition in Australia with professors punching in problems as they came out, and the tension of not knowing if they'd hit Gold until the human scores came in (because the Gold threshold is a percentile, not a fixed number)
Why they threw away AlphaProof: "If one model can't do it, can we get to AGI?" The decision to abandon symbolic systems and bet on end-to-end Gemini with RL was bold and non-consensus
On-policy vs. off-policy RL: off-policy is imitation learning (copying someone else's trajectory), on-policy is the model generating its own outputs, getting rewarded, and training on its own experience—"humans learn by making mistakes, not by copying"
Why self-consistency and parallel thinking are fundamental: sampling multiple times, majority voting, LM judges, and internal verification are all forms of self-consistency that unlock reasoning beyond single-shot inference
The data efficiency frontier: humans learn from 8 orders of magnitude less data than models, so where's the bug? Is it the architecture, the learning algorithm, backprop, off-policyness, or something else?
Three schools of thought on world models: (1) Genie/spatial intelligence (video-based world models), (2) Yann LeCun's JEPA + FAIR's code world models (modeling internal execution state), (3) the amorphous "resolution of possible worlds" paradigm (curve-fitting to find the world model that best explains the data)
Why AI coding crossed the threshold: Yi now runs a job, gets a bug, pastes it into Gemini, and relaunches without even reading the fix—"the model is better than me at this"
The Pokémon benchmark: can models complete Pokédex by searching the web, synthesizing guides, and applying knowledge in a visual game state? "Efficient search of novel idea space is interesting, but we're not even at the point where models can consistently apply knowledge they look up"
DSI and generative retrieval: re-imagining search as predicting document identifiers with semantic tokens, now deployed at YouTube (symmetric IDs for RecSys) and Spotify
Why RecSys and IR feel like a different universe: "modeling dynamics are strange, like gravity is different—you hit the shuttlecock and hear glass shatter, cause and effect are too far apart"
The closed lab advantage is increasing: the gap between frontier labs and open source is growing because ideas compound over time, and researchers keep finding new tricks that play well with everything built before
Why ideas still matter: "the last five years weren't just blind scaling—transformers, pre-training, RL, self-consistency, all had to play well together to get us here"
Gemini Singapore: hiring for RL and reasoning researchers, looking for track record in RL or exceptional achievement in coding competitions, and building a small, talent-dense team close to the frontier
—
Yi Tay
Google DeepMind: https://deepmind.google
X: https://x.com/YiTayML
Chapters
00:00:00 Introduction: Returning to Google DeepMind and the Singapore AGI Team
00:04:52 The Philosophy of On-Policy RL: Learning from Your Own Mistakes
00:12:00 IMO Gold Medal: The Journey from AlphaProof to End-to-End Gemini
00:21:33 Training IMO Cat: Four Captains Across Three Time Zones
00:26:19 Pokemon and Long-Horizon Reasoning: Beyond Academic Benchmarks
00:36:29 AI Coding Assistants: From Lazy to Actually Useful
00:32:59 Reasoning, Chain of Thought, and Latent Thinking
00:44:46 Is Attention All You Need? Architecture, Learning, and the Local Minima
00:55:04 Data Efficiency and World Models: The Next Frontier
01:08:12 DSI and Generative Retrieval: Reimagining Search with Semantic IDs
01:17:59 Building GDM Singapore: Geography, Talent, and the Symposium
01:24:18 Hiring Philosophy: High Stats, Research Taste, and Student Budgets
01:28:49 Health, HRV, and Research Performance: The 23kg Journey




