ReflectionAI Founder Ioannis Antonoglou: From AlphaGo to AGI

28 Jan 2025 · 52 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Training Data Podcast Episode Notes

Episode Overview Title: ReflectionAI Founder Ioannis Antonoglou: From AlphaGo to AGI Hosts: Stephanie Zhan and Sonya Huang, Sequoia Capital Guest: Ioannis Antonoglou Description: Ioannis Antonoglou, founding engineer at DeepMind and co-founder of ReflectionAI, discusses reinforcement learning breakthroughs from AlphaGo to AlphaZero and MuZero, evaluating key moments that shaped AI and their implications for the future.

---

Key Concepts and Technologies Discussed

Reinforcement Learning (RL)

  • Definition: A type of machine learning where an agent learns to make decisions by taking actions in an environment to maximize cumulative reward.
  • PPO (Proximal Policy Optimization): Developed by DeepMind; used in various applications including OpenAI's ChatGPT.
  • Self-Play: A training method where the AI plays against itself, allowing it to learn without human data. Utilized in AlphaZero.

Notable AI Models

  • AlphaGo: First AI to defeat a human Go champion, leveraging deep reinforcement learning and self-play.
  • AlphaZero: Generalized model that learns to play board games from scratch using only self-play, demonstrating superior performance.
  • MuZero: Advanced model that learns game strategies without being explicitly given the rules, creating its internal model of the game environment.
  • AlphaFold: Model predicting protein structures, which won the 2024 Nobel Prize in Chemistry.

Key Algorithms

  • Monte Carlo Tree Search (MCTS): A heuristic search algorithm used in AlphaGo for decision-making and strategy planning.
  • DQN (Deep Q-Network): Introduced in 2013, combining Q-learning with deep neural networks for playing Atari games.

---

Discussion Highlights

The Journey of AI from Games to General Intelligence

  • Starting with Games: DeepMind chose games as a testbed for AI to explore complex decision-making scenarios in a controlled environment.
  • Challenges of Go vs. Chess: Go represents a higher complexity due to its vast number of possible board configurations, making it a "holy grail" for AI research.

Breakthrough Moments in AlphaGo

  • Move 37: A pivotal moment during the match against Lee Sedol, showcasing AlphaGo's creativity and strategic understanding, initially perceived as a mistake.
  • Move 78: Highlighted AlphaGo's vulnerabilities, as a miscalculation allowed Lee to capitalize on its mistake.

Technological Evolution

  • From AlphaGo to AlphaZero: Transitioned from using human data to learning entirely through self-play, simplifying training and enhancing versatility.
  • MuZero's Advancement: Moves beyond needing explicit game rules, paving the way for applications in complex real-world scenarios.

Current Challenges and Future Directions

  • The Data Wall Problem: Anticipated challenges in scaling LLMs due to reliance on human-generated data.
  • Importance of Planning: Emphasized the need for AI agents to incorporate planning for improved performance.
  • Reinforcement Learning's Resurgence: RL's potential to provide solutions in scenarios where human data is scarce.

Open Questions in AI Development

  • Robustness and Reliability: Ensuring AI models maintain consistent performance and adapt to errors, similar to AlphaGo's reliability.
  • Context-Length Capabilities: How to enhance AI's ability to learn and adapt dynamically in real time.

---

Insights from Ioannis Antonoglou

  • Cautious Optimism: The DeepMind team had a belief in the potential of reinforcement learning but remained aware of the inherent unpredictability of AI systems.
  • Scalability of Solutions: Emphasized that scaling models and data lead to significant improvements in performance.

---

Future Predictions

  • Next Milestones in AI: Expect models to become more reliable, intelligent, and capable of executing tasks independently.
  • Synthetic Data: Anticipated as a crucial area of development to overcome the limitations of human-generated data.

---

Closing Thoughts Ioannis Antonoglou's insights into the evolution of AI models from AlphaGo to the present highlight the importance of creativity, planning, and adaptability in the pursuit of AGI. As the AI landscape continues to evolve, the combination of reinforcement learning and advancements in model training methods will be critical in addressing the challenges that lie ahead.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Because it was a complex game, there was always a bit of worry about whether Africa was truly as good as we believed. So we actually had the conviction that the deep reinforced learning is the answer based on everything that we could measure and everything we could see. But that's the thing about this system is that there are not like classic computers where you just like know that they always produce the same answer. They're like stochastic, they're creative, and they have like some blinds, they hallucinate like similarly to how like model LLM's callousness. So you need to just like really push them and just like see exactly where the grape and the only way you could actually do that is by having like the best humans playing against them.

0:57Today we're excited to welcome Janis Antenoglo, a researcher and an engineer who has contributed to some of the most significant breakthroughs in AI. As a founding engineer at DeepMind, John has played a crucial role in developing AlphaGo, which made history by defeating Go World Champion, Leisidol. He later co -led the development of Museurio, which pushed the boundaries even further by mastering multiple games autonomously. Now, as he embarks in his latest venture with Reflection, he's focused on building the next generation of AI agents. We're excited to talk to Johnus about the breakthrough moments in AI history that he's witnessed firsthand.

1:38From AlphaGo's famous Move 37, his perspective today on what's next for the combination of reinforcement learning and large language models on the way to AGI. Janice, thank you so much for joining us today. Thank you so much for having me. Janice, you have an incredible background having worked at DeepMind as a founding engineer for over a decade, starting with some of the most notable projects that have really defined the industry. DeepMind quite notably created this notion of building AI within games to start. Can you share a little bit more about why DeepMind chose to start with games at the time?

2:18Yeah, so DeepMind was the first company to truly embrace the concept of artificial general intelligence or AGI. from the outset, they had ground ambitions, aiming to build systems that would match or exceed your intelligence. So the big question was, and still is, how do you build AGI? And more importantly, how do you measure intelligence in a way that allows for meaningful reasons and performance improvements? So the idea of choosing video games was that testing ground came naturally to the point founders. It was a Demisos Habits and Shane Lake. Because Demis had a background in the gaming industry and changed PhD thesis to find AGI as a system that could learn to complete any task.

2:56Video games provided a control yet complex environment where these ideas could be explored and tested. And to what extent he mentioned games are, they provide a very controlled environment. To what extent are games representative or not of the real world? Like if you have a result in games, do you think that generalizes naturally to the real world or not? So I mean, I guess games have indeed been viable for developing AI. And you actually have like a few examples of that. So you can see that PPO for example, which is currently being used in RLHF was developed using OpenAI, GIM and Michoko and Atari and similarly, you have like MCTS which was developed which stands for Monte Carlo 3SH and was developed through port games like Pac -Am and Go but at the same time games have like a number of limitations so the real world is and Messi is a bounded and it's a much tougher not to crack than even the most complex games.

3:51So even though they just give you an interesting aspect to develop new ideas, it's definitely limiting and it does really capture all the complexity of the real world. Okay, interesting though. So a lot of the techniques and algorithms that you've developed in the game environments, TPO, etc. These are used in the real world. Yeah, so PPO was actually exactly what just PD used for RLHF. And so MCTS, it's used in Yuzio and Yuzio has been used in like the real world in things like compression video compression for YouTube. In it was part of the cell driving system, Tesla at some time. And it was also like used for developing a pilot, like that was completely controlled by an AI.

4:43So yeah, I mean, you can see it methods like that being used in the world to solve their problems. So interesting. Yannis, I remember back in 2017, when AlphaGo, the movie, came out and it featured the incredible game of AlphaGo against LisaDoll. Can you take us back to that moment in time and maybe the year is leading up to it as you're building AlphaGo? How was AlphaGo specifically chosen as the game to focus on? So I think I'm like games you've always been a benchmark for AI research so like before Go you had chess and chess was like a major milestone with like IBM's deep blue defeating Gargaz Parof in the late 90s and And even though chess and go are completely different games and go with a definitely different beast, there is, like, games have always been acted as test beds for the development, especially board games for the development of new AI methods.

5:40Actually, even going back to the earliest days of AI research, Turing and Shannon, they both worked on their own versions of chessboards. So now the thing about like go is that it's a much harder problem than like chess. The reason for that is because it's almost closely possible to define an evaluation method, a heuristic. So in chess, you can just like take a look at the boards, you can count the number of pawns that like it side has, you can see what the ranks of these pawns are, and then you can just like make some, you can draw some conclusions of like who is winning and why. But I can go there's nothing like that.

6:18Like it's mostly human intuition. And if you ask a go professional player, like how they know whether a position is a good partner, but one they would say that after having playing the game for so long, they could just like fill it in their gut, but like this is a better position on the other one. So now there's actually a question of how do you encode the feeling in your gut into like an AI system, right? So this is exactly the reason why solving Go was considered the holy grail of air research for a long time. And it was the challenge that seemed almost impossible, but at the same time it was like within reach, people felt that they could actually get it cracked.

6:56And this exactly what Alphogo did back in 2016. And it kind of showed case two new methods, which is deep learning and reinforcement learning. Because back in 2015 and 2016, now we kind of think of deep learning and reinforcement learning as part of technologies, but back then we were kind of literally making their first steps and they were kind of like the new kid in the block and most people were kind of like really skeptical about them. That deep learning was another AI fad that would just like one last the test of time. So yeah, I mean, AlphaGo was chosen because it was like clear to show the that you actually have like the most the most performant agent in the world.

7:41You could actually evaluate it. You can have it play with other humans. And at the same time, it was within reach given like the latest developments in deep learning and reinforcement learning. I remember reading that. There's more configurations of the Go board than Adam's and the universe. By many letters and nine to them, that blew me away. I mean, I grew up playing Go and it felt like such a, you know, it's a very simple rule in terms of the rules. But I see why it was the Holy Grail. Maybe can you explain how AlphaGo worked technically maybe explains me like I'm a fifth grader because that is effectively my level of sophistication understanding these things.

8:18But how did it work? And you mentioned both reinforcement learning and deep learning were involved. But love to peel that back a little bit. Yeah, absolutely. So AlphaGo has two deep neural networks. So like a neural network is a function that like takes something as an input, input and produce something as an output. And it's literally like a black box. We don't really know exactly how it does it. Just like know that you can actually, if you train it on enough data, we'll just like learn the mapping. If we learn the function from input to the output space. So AlphaGo actually had access to two deep neural networks, the policy network and the value network.

8:53And the policy networks suggested the most promising move. So it will just take a look at a current port position. And we'll just like say, okay, you know, based on the current position, this is the list of moves that I would recommend you just consider playing. And it also had access to the volume network, we just take a look at a port position, and just give you a winning probability. What are your chances of actually winning the game starting from this position? This is exactly the gap feeling. It had its own gap feeling on whether the position is a good one or a bad one. So once you have access to these two networks, then you can actually play in your imagination a number of games.

9:30You can consider the most promising moves, then you can consider your opponents most promising moves And then you can just evaluate each move like the value network And then you can use a method called minmax. What that says is that I want to win the game But I also know that my opponent wants to win the game So I want just like pick a move that will maximize my chances of winning knowing that like my opponent will try to maximize their chance of winning. So if you actually like to do that and simulate a bunch of moves, then you can just like get the optimal action. And you know, the way to just like do this imagination, this planning, this search in the most efficient way is by using a tree sets method called Monte Carlo tree sets.

10:15So MCTS. So whenever people talk about MCTS, they literally just like mean this heuristic of how do I choose which features to consider so that I can make informed decisions. The role for reinforcement learning and deep learning and building AlphaGo was that AlphaGo first fall was a success of reinforcement learning and deep learning because this is exactly the two methods that powered AlphaGo. And the policy network was initially trained on a large a set of human games. So you had like many games played by human professionals and you just like consider every position and you consider the move they took at this position and then you have like a dipping in a network that tries to predict this move.

11:01Then once you have the policy network you need to somehow find a way to just like obtain a value network. So we did it in two ways. First we just took the policy network and we had it play against itself and we used reinforcement learning to improve the blank strength of the model. So we use a technique called policy gradient. So what policy gradient does is that it just looks at the game and then it looks at the outcome. This is the simplest version, like of policy gradient. It looks at the outcome of the game and for all the moves that led to a win, they'll just like say great, you know, just increase the probability of choosing this move.

11:38And for all the moves that led to a loss, it says great, now it's a decrease the probability of like this being selected in the future. And if you do that like, you know, for many games and for long enough, then you just like it and improve policy. Now, once you have this improved policy, you can just generate a new dataset of games where like the policy plays against itself and then you have like a huge amount of games where for its position, you know who the final winner was. So then you can take this network, you can take another network, a value network, and here it's predict the outcome of the game based on the current position.

12:12So what the network could learn is that if I start this position and I play under my current policy, on average this is the player who wins, like it's either a black player or the white player. So this is the first version of like a value network and you can just like give it within half a goal by combining it with the policy network. And what were some of the biggest challenges in building this and how did you overcome them? Yeah, so AlphaGo was not just a listed scientist but mostly I'd say an engineering marble. The early versions run on 1200 CPUs and 176 GPUs and the version that played against listed old used 48 GPUs.

12:52So like GPUs for like the first accelerator, custom accelerators. And these accelerators were like really primitive back then, because literally it was like the first version, right? Like now the later accelerators are much, much better and much, much more stable. So the system had to be highly optimized to minimize latency, maximize throughput. We had to build landscape infrastructure for training these networks, and it was a massive endeavor. It just required a lot of coordinated effort for many talented individuals working on different aspects of the project. But I just like walked you through a number of steps, just like obtained the Polish network and the value network.

13:32And each of these steps had to just be implemented at the limits of what was possible back then in terms of scale. And it had to be implemented in a way where people could just like think it with it, they could just like try the research ideas fast and get results fast. So yeah, lots of people scale in, you know, at levels that hadn't been implemented before. And it's kind of like working at the forefront of what was possible back then. I love your highlight of it being a research marvel and an engineering marvel. And I remember you sharing one time that part of the reason this project came about also was because Google had TPUs that they needed to, they needed a test customer for.

14:16and that was the spark. This album is AlphaGo project, so that's pretty incredible. How much conviction did the DeepMind team have that this is going to work? You mentioned that at the time, deep learning, reinforcement learning were still relatively novel, but DeepMind was very much founded with that belief. But did you guys think that you were going to be able to have these superhuman level results beating the top go player in the world? Was it a crazy idea and maybe they'll work or did the team have conviction like this is going to work? So I'd say I'd like to tell you the team had a cautious optimism.

14:49So one of AlphaGo's lead developers, Ajah Huang, he is a strong amateur go player and he had been working on Go for like a decade before AlphaGo happened. And we also had like a little report of computer games, a lot of computer players and you could say that AlphaGo was significantly stronger than anything that had come before. But Go is a complex game and there was always a bit of worry about whether AlphaGo was as true as good as we believed. So we actually had the conviction that the deep reinforcement learning is the answer based on everything that we could measure and everything we could see.

15:23But the thing about this system is that they're not like classic computers, where you just like know that they always produce the same answer. They're like stochastic. They're creative. So they all have like some blind spots. they hallucinate, like similarly to how like model LLM's hallucinate. So you need to just like really push them and just like see exactly where they break and the only way you could actually do that is by having like the best humans playing against them. Move 37, can you tell us what that was? It was such a monumental move and I think everyone watching it at the time it was and at least all maybe primarily was confused by that move.

16:08What was going on in your head when that happened? So yeah, I mean, move 37 in game 2 against Lise -Dol was literally just a spectacular moment in the sense that it kind of showed gaze to the world that AlphaGo has creativity and it demonstrated that AI could come up with shots that even top human players hadn't considered. So at first, like I still remember that, we thought that AlphaGo made an error, so that it actually hallucinated. It did something that it didn't mean to do. But then, turn out to be a brilliant conventional move that underscored that the system had a deep understanding of the game.

16:46That the system actually had creativity. It's good thing of things that people hadn't thought of before. I want to take us to another key move in the game. I think it was in game four. At this point, I was rooting for Lee because I was like a vortga, I used to win a game. I moved 78. I think AlphaGo made a mistake and Lisa don't know this is it. I guess what was the weakness there that Lee found during the game? Yeah exactly. So I mean Lisa Dole's victory in game four was literally a testament to human ingenuity. Like move, move 78 was unexpected and got off a go off card. Initially I have a goal like based on its evaluations, misinterpreted as a mistake and thought that it was actually like winning.

17:28So that's why it is to respond appropriately. And you know, this kind of highlighted the blind spot in the system. So the game shows that like while systems like Africa are extremely powerful at the same time, they still have vulnerabilities and there were like steel areas where it could further improve it. But how do you go about improving something like that? Do you need to show it a lot more data of, you know, kind of that type of human ingenuity move or how do you go about fixing and and patching those blind spots. So yeah, I mean, it's actually interesting that by the end of the games, like this, we just like put together a benchmark.

18:04We're just kind of like trying to quantify and just have a way of measuring the mistakes that like Africa makes. And you know, this is kind of blind spots, let's say. And then we just write an other approaches to just like the prove the algorithm so that we can solve these issues. And what happened is that actually the most effective way of getting rid of them was just like do what we were doing, just like at a higher scale and better. So we just like changed the architecture of the model, we just like switched to a deep rest net with two output heads. And we also like, we just had a bigger network trend or more data, then just like move 12 Alpha Zero and better algorithms and that kind of like made it so that we didn't have any hallucinations anymore.

18:55So in a way, which is like scale data, you know, things that are always kind of the well -known recipe in the field of AI. It's exactly what sold it in our days too. With scale and data, how much did higher quality data or maybe specifically data from great professional players, the best professional players make a meaningful difference? Or was it just any data? Now for us what matter was that we kind of like solved it using self -play. So we actually had access to the most competent go player in the world. And we just like used it to generate the best quality games and then we just trained on these games.

19:37So I guess like you know, we didn't need to have like human experts because you had like an expert in house. It wasn't human. Right. Huh. Interesting. Amazing. Well, I'd love to move on to the progression from AlphaGo to Alpha0. And you talked a little bit about this notion of self -play just now. Alpha0 was powerful because it learned how to play the game from scratch entirely from self -play without any human intervention. Can you share more about how that worked and why that was important? So Alpha0 was a game changer because it learned entirely from scratch through self -play without any human data.

20:13And this is like a major relief from AlphaGo, because AlphaGo, as I said, relied heavily on human expertise. So two things happened. First of all, Alpha0 managed to simplify the training process and also showed that AI will literally just get from zero to superhuman performance just purely by playing against itself. And that allowed it to just be applicable to a whole range of new domains that were out of reach, because there weren't enough human data for it. But the more important thing is that we just saw that Alpha Zero also solved all the issues that AlphaGo had in terms of hallucinations, in terms of plant spots and robustness.

20:58So like Alpha Zero was a better method, just full style. And you explained kind of how AlphaGo worked to a fifth grader. What would you tell the fifth grader to be the key difference technically that you've implemented with Alpha 0. So Alpha 0, just like Alpha Go, uses a policy network and a value network, along with motor carotry sets. So in that respect, it's exactly the same as Alpha Go. So the key difference is in training. Alpha 0 starts with random weights and lands by playing games against itself. And by playing games against itself, it iteratively improves its performance. But the main idea behind Alpha 0 is that whenever you take a set of weights, a set of policy and value networks.

21:41And then you just combine the moussech, then you just end up with a better playing better player, which is like increase your performance, you just become a stronger player. So what that meant is that we can actually use this mechanism to improve the model policy, the role policy. So this is what we call in reinforcement learning, a policy improvement operator. Whenever you can just take an existing policy and then do something, some magic, and then just like come up with like a better policy, and then you can just like take this policy and distill it back to the initial policy, and just repeat this process, then you have like a reinforcement learning algorithm.

22:21And I think it's like, you know, this is exactly what people are trying to do today, with like, you know, two -star or like, you know, synthetic data. This is exactly the idea of like, how can I take a policy, do something with it, planning, search, compute, whatever it is, and derive a better policy, which I can then imitate and just like kind of distill back to the original policy. So this is exactly what Alpha Zero is doing. It uses MCTSH to produce a better policy than it takes the trajectories, it trains this policy in value network on the new better trajectories and it repeats this process until it converges to an expert level go player.

23:02That's fascinating and counter -intuitive that kind of like starting without the weights that you would have from, you know, professional level players is actually a better starting place. The epitome of AI agents and games, it's cheap to think via MuZero, which the progression even from AlphaZero itself, and it's also where you became one of the co -leads or one of the leads of the game. AlphaZero was obviously impressive because of self -play, but it also needed to be told the environments dynamics or the rules of the game. And Muzero takes this to the next level without needing to be told the rules of the game and it mastered quite a few different games, Go Chas and many others.

23:46Can you share a little bit about how Muzero worked and why was this particularly meaningful? Absolutely. So Alpha Zero, as you said, was a massive success in games like chess, go, Shogi. So in games where we actually had access to the game rules, where we actually had access to a perfect simulator of the world. But like these two lands on the perfect simulator, made it challenging to apply to real world problems. And real world problems are often messy and they like the rules and truly hard just like write a perfect simulator of them. So that's exactly what Muzio tried to solve. So Muzio, masters the games, of course, like Go chess and Shogi.

24:26but also like matches more visually challenging games or games that are like a hotgoat like a Tari. And it does that without giving access to the simulator, just like lens how to beat an internal simulator of the world and then just use this internal simulator in the way similar to what Alpha Zero is doing. So it does that by using model based for enforcement learning, where what that means is that you can just take a number of trajectories generated by an agent and then try and, you know, So, Lennon model will end a prediction model of how the world works. So this is actually quite similar to what methods like Sora are trying to do now, where they just like take YouTube videos and they try to just like Lennon world model, but just trying to predict based on something from one frame what's going to happen in the future frames.

25:13So, Nusiel tries to do exactly that, but it does it in a way different from, you know, of the genetic models in the sense that they try to only model things that matter for solving the reinforcement learning program. So it tries to predict what the rewards are going to be in the future, what's the value of like future states, what's the policy for like future states. So only things that you need to thin your MCTS. But the fundamental is kind of like remain the same. So how do you just like learn a model based on trajectories and then once you have this model, you can just combine it with search and get super human performance.

25:51So of course, you can always decouple the two problems and have the model been trained separately from data out in the wild and then just like combine that with Museo. And we just found that back then, given the limitations of our models and the smaller sizes, is kind of like made more sense to just like keep those two together and only have the model predict things that matter for planning. So just like try to model everything because you're kind of hitting the limits of what the capacity of the model could take. So interesting. Is it right to assume that not only Sora, you know, takes the same approach, but maybe other world models or other robotics foundation models?

26:35Yeah. So anything that tries to just like build a model of how the world works and then just like use that for climbing. It's within you know, you zero like methods. So yeah, you can just like train it on YouTube videos. You can train it on like the inputs coming from like robots. You can train it on any environment. You can even think of like class language models as a form of models of like text. So like they they the model text. But the thing about text is that like the more than a little bit trivial, like you don't need to just, don't have many artifacts happening when you're trying to predict what the next world is going to be.

27:13Have you seen the ideas behind Museo kind of be used outside gameplay or in messy real -world environments?

Read the full transcript

27:24So, as I've said, Alpha Zero and Museo are quite general methods, and there's a number of scientific communities in GameStrees, So there's Alpha -Kam in quantum computing. Some people try to use the alpha -0 in optimization. But they just adopt the alpha -0 because it was really powerful in really doing planning and solving disoptimization problems. At the same time, U0 was incorporated in a version of Tesla's self -driving system. It was kind of reported in their AI day. And it was also used, and I think it's currently being used within YouTube you as a custom -compression algorithm. But it's early days and takes time for this new technologist to be fully adopted by the industry.

28:15We'd love to talk a little bit more about reinforcement learning in agents. And you alluded earlier to the fact that reinforcement learning and deep learning back in 2015 were new, you can nice some ideas. They really grew in popularity in 2017, 2018, 2019 onwards. And then they were overshadowed by LLMs, largely because of the GBT and everything else that came out. But now reinforcement learning is back. Why do you think that is the case? Yeah, I mean, first of all, LLMs and multiple models have indeed brought incredible progress to AI. So these models are exceptionally powerful and can perform some truly impressive tasks.

28:54But they have some fundamental limitations. And one of them is the availability of human data. People just keep talking about the data wall and what happens when run out of high quality data. And this is exactly where the reinforcement learning signs. So reinforcement learning excels because it doesn't realize only on pre -existing human data. Instead, reinforcement learning uses experience generated by the agent itself to prove its performance. So this self -generated experience allows reinforcement learning to learn and adapt and to even adapt to scenarios where human data is scarce or like nonexistent.

29:31So if you define the reinforcement learning problem in the right setting in the right way, you can literally effectively exchange complete for intelligence. You can just get to a point similar to where we were with Alpha Zero, where we're just like, the moment we threw more computer to it, like we made the network's bigger, we just like, you know, used more games, we just literally got a better player. And it was the deterministic. You always get a better player. So I guess this is exactly where we want to be with like this synthetic data pipelines. Currently we have that with the scaling close in LLM, that if you have more data and bigger models, then you can predict that there's going to be an improvement in performance.

30:13But once you run out of human data, how do you just keep going? And synthetic data is the answer to that. And the only way that you can actually get high quality and reinforcement learning, high quality data to just improve your model is via some form of reinforcement learning. And just like leaving, I'm just like keeping reinforcement learning as a really kind of blanket term here where I just like define it as anything that lands through trial and error. How do you think reinforcement learning is being brought into the like LM world and you mentioned Q -star earlier? Like I guess in a closed form game you have like a pretty clearly defined policy and value function.

31:00How does that work in a messy real -world environment or the LLM world? Stop. I mean, I guess there are two different types of messy real -world, right? If you try to just build a controller or something, that's a really messy environment. And then if you operate in the digital space, so my personal I believe that this AI, which is happening much earlier than you know, robotics, HGI. And the reason for that is exactly that you have control over the environment. And the environment is like computers, like the digital world. So even though it's like messy and snowy, it's still contained. It's not like the real kind of like what in that sense.

31:42So now in terms of how do you bring like reinforcement learning? So reinforcement learning is, we used to say in DeepMind that you have like the problem and you have the solution. And the problem setting of reinforcement learning is how do I take a model, how do I take a policy, and generates synthetic data, or like I find a way to improve this policy by interacting with the environment, by a trial and error. And it's like the reinforcement learning problem setting, right? And then there's the solution space where you have value functions and have reinforcement learning methods. So I think that there's a lot of inspiration to draw from like classical reinforcement learning methods that were developed in the past decade.

32:27But we have to adjust them to the new world of LLMs. So methods like Q start to do that, by just taking the idea that if I have a policy and then I do planning, I consider possible future scenarios, and then I have a way to evaluate which one is better, then I can just like take the best ones and then ask the model to imitate these better ones. And this is like a way of improving the policy. So in the classic RL framework, you do that by using a policy and a value network. In the new world, you'll just do that by asking your, by having a reward model or asking your your LLM to just like give you feedback on an output it gave you.

33:14So much interesting. You also talked a little bit about synthetic data earlier. I think some folks are very bullish on synthetic data and some folks more skeptical. I also believe that synthetic data is more useful in some domains where outcomes and successes perhaps more deterministic. Can you share a little bit about your perspective on the role synthetic data and how bullish you are on it? Yeah, I mean, I think synthetic data is something that we have to solve one way or another. So it's not about whether your bullish on art is kind of... is an obstacle that we have just find a way around it.

33:46Like, we will run out of data. Like, you know, there is so much data that humans can produce. And also, like, it's important that the systems start taking actions, they start learning from their own mistakes. So we need to just find a way to make like, since the data work. Now, what people have done is that they've tried like the most, I guess, like, naive approach where you just like take the models to produce something you try to train on that. And of course, they've seen that there's mode collapsing and it just doesn't work out with the box. But new methods never work out with the box. You just need to invest in it and just take your time and really think of what's the best way of doing it.

34:34So I'm really optimistic that we'll just definitely find ways to improve this model. and I think that actually there is a number of methods out there, like the two -star and the equivalents that just, you know, in the new world where people don't really share their research breakthroughs, the way they use to is probably hidden behind some company trade secrets. I'm going to ask about reasoning and novel scientific discoveries. Do you think that that can naturally come out of just scaling LLMs if you have enough data, or do you think that kind of like the ability of reason and come up with net new ideas requires kind of doing reinforcement learning and deeper compute at inference time?

35:21So I think I think you need reinforcement learning to get a better reasoning because the distribution of like it's also about the distribution of data, right? Like you have like a lot of data out in the wild in the internet, but at the same time you don't always have the right type of data. So you don't have the data or some on reasons, and they just explain the reasoning in detail. You have some of it, and it's incredible that the models have actually managed to pick it up and just imitate it. But if you want to just improve on that capability, then you need to do that through reinforcement learning.

36:00You need to just show the model how this kind of emerging capability can and further being proved by just like, have it generated that data and interact with the environment, just tell it when it's doing something right and when it's not doing something right. So yeah, I think that like, reinforcement learning is definitely part of the answer for that. AlphaGo, Alpha0 and Moosero are the most powerful agents we've ever built. Can you share a little bit about how some of the lessons and learnings unlocked from that are relevant to how we're pursuing building AI agents today? Yeah, so I feel like AlphaGo and me zero, you know, they've actually funneled the transform their approach to AI agents because they highlight the importance of planning and scale in my opinion.

36:45That if you actually look at the charts of like different models and how they scale, you can see that like AlphaGo and Alpha Zero were like kind of really ahead of the time. Like they were kind of outliers. Yeah, like this, this cares of like how compute scaled. And then you have like Alpha Zero, or like somewhere standing on its own. So it's all that like if you can scale and you can really push on that, then you can get like incredible results. At the same time, you know, we talked to a show that you don't have just only trade. You don't also like, you know, have better performance during inference, during test, during evaluation, but just like using planning.

37:21And I think that this is something that will start seeing more and more in the near future, or like this method will just like start thinking more, like planning more before they're just making any decisions. So I'd say that like this is more of the heritage of AlphaGo and Alpha0 and E0. It's the the the basic principles and the basic principles are of that scale matters, planning matters. These methods can really solve problems that we thought that are insanely complex or like you know beyond what we can solve on our own. Similar problems with the ones that you actually observed with these last language models are things that we saw back then, like back in 2016 we actually saw that these models can hallucinate or that like at the same time they're also creative that they will just come out with solutions that we hadn't thought of.

38:14But they can also like have blind spots or like hallucinate or be susceptible to kind of like adversarial attacks which I guess like everyone knows now that this neural network suffer from. So I think that like this are the the main kind of lessons drawn from this line of work. What do you think are the biggest open questions from this line of work for the field dancer going forward? So the main question is we had like Alpha Gwernizio and we just like much they have like this insanely robust and reliable systems that will just always play go and at the you know the highest possible kind of level and they'll just like achieve consistently they will just like be top of the leaderboard will just like never lose again.

39:04So half a go master actually like played against 60 people in online matches and just like each one in every single one of them. So there was like no there there was like this pattern is for like a critical robust reliable and I think like this is exactly what we're missing now with these LLM based agents. Sometimes they get it, sometimes they don't. You cannot trust them. They will just like, you know, you have like some amazing demos, but like, you know, they happen once every two times even, or like, once every 10 times, you have like something amazing. And the remaining nine, they just lost their way and didn't do anything.

39:40So if you have like what we need to do, it's just find a way to just make these LLM based agents equally robust to the ones that we had with AlphaGo and Vizio and Office -O. This is like the new open question of how do you actually do that? We'd love to move into some of your thoughts on the broader ecosystem today. You've touched on a few really core problems that people are working on right now. One, the data wall problem that will hit eventually, perhaps by 2028 or so, as some folks predict. Another being the idea of planning as an area that AI agents need to get better at. And then a third idea that you just described was around robustness and reliability.

40:26Can you share a little bit about maybe some of these areas that you think the whole field needs to solve that you are most excited about to help us unlock this vision of really getting to the AI agents that we want? Yeah, I mean, I'll just like also add another one to the list. So I feel like another major another major challenge is like how to improve the context -length cable beats of these models. So like, you know, how do you make sure that like these systems can land on the fly and how they can adapt to new context like with you? So this is like another thing that I think it's going to be really important.

41:05It's going to happen the next a couple of years actually. yeah, what's the term that you used for that? In context learning? In context learning. In context learning. Yeah, so it's the idea that a system can actually learn how to do a new task with like few short prompting, like it kind of like sees a few examples. And on the fly, it kind of like learns how to adapt the new environment. it, lens hard to use the new tools that were provided to it. Or like, it's kind of like lens. It's not just all the knowledge it has stored in suites, but like it's also like acquiring new knowledge by just like interacting with the real world, interacting with the environment.

41:48So I think that this is like another place where there is a lot of work happening at the moment and going to have like an amazing progress in the next couple of years. And I'm really excited about that. So yeah, I mean, to recap, I think planning is important. In context learning is important and reliability. So the best way to achieve reliability is just like ensure that this model somehow knows how to return from their mistakes. So if they just made the mistakes somewhere, they can just like see that. And they're like, okay, I made a mistake. I'll just like work towards the way that are humans, you know, make mistakes all the time, but like we, you know, you can correct for them.

42:35So these are like the three areas which are really, I'm really excited to see progress on. No, that you've kind of embarked on your own entrepreneurial journey. How do you think that the areas where startups can compete against the big research labs and like how do you kind of motivate yourself for that journey? Yeah, I mean, it's a completely like a it's a new world for me, but at the same time It's not that new because when I joined deep mind is was literally a startup so And I was like literally in the first half Jim please So I actually like saw that first hand But you know one of the benefits of like working for a startup is that you know that's really the end of focus So everyone really cares Everyone just moves really fast and there's like a clear focus on what we want to be it so So the building is like what's the most important kind of motivation for people, just like building.

43:30And I think not like this is one of the big advantages that startups have over more established businesses. At the same time, it's easier to just like default to adapt to new findings and technologies. You're not kind of tied to some pre -existing solutions or some products that you don't want to duplicate because they bring a lot of revenue to you. Well, if you're a startup, you have no such change. You can just move fast and be innovative and just break conventions. And at the same time, just allows you to leverage open source resources, things that are out of touch for the big labs. And yeah, and you don't have the red tape, that big place is then to help.

44:17I love the term that you sometimes be honest, main quest versus side quest. Yeah, it's the idea of having a main focus, like in big places, in big labs, they have like many different projects that people are working on. And it usually happens that they have like the main quest, the main thing that everyone's working on, and there's like many multiple smaller side quests. The idea is you just like feed into the bigger quest, But like, usually they don't get as much, they don't get like as many resources or like as many as much focus from like the leadership. So yeah, they they tend to Dirtrophy.

44:59In the broader field, what are some of the most defining projects that you admire the most and maybe who are some of the most influential researchers that you admire the most? Yeah, absolutely. So I actually like started my AI research journey back in 2012 and I've actually like seen some milestones. So I'll just like I give a list of like what I think are like the main milestones like in AI in the past like 12 years that I've been around. So the first one I'll say is like Alex net. This is the first paper that kind of like show that deep learning is the answer. I mean back then it didn't feel like it.

45:37It's just like curiosity. But now I think that most people are convinced that deep learning is part of the answer. Then it was a TQN. I had the pleasure to actually walk on TQN and just like see it firsthand how it started. It was actually developed by a friend of mine, Vlad Mi, Vlad Mi. And it was like the first system that showed that you can actually combine deep learning with reinforcement learning to achieve If you have a few performance or like superhuman performance in really complex environments. Then this was alpha go. Again I was like really lucky to just like walk on that and show that scale and planning are really important ingredients and if you just like do that right and you get huge success in an incredibly complex environment.

46:29Alpha fold, another one. This is again by deep mind. So that like this methods are not just like things that you can use to solve games, but they have They they actually will make this world a better place They'll just like ensure that health care is improved that scientific discoveries Are being realized that we'll just like make sure this world is a better place by using AI then Chaturpe tea it kind of like brought AI to to everyone, just like made it accessible to the broad audience. Like everyone knows what AI is now. It has made my life of explaining my job much easier. So, and finally, tip to four.

47:15And I think that, yeah, probably tip to four is like the latest kind of big advancement in AI, because it kind of like showed that, you know, as far as the general intelligence is a matter of years. It's within reach. Yeah, we are getting there. I think that, you know, many, most people now believe that we are like a Few years away from like a GI and you know, that's that's because of like the incredible breakthrough that GP4 was now in terms of like some people I really admire Before I forget so I'd say fast like David Silver. He's And he was my piece, this professor, he was my mentor, DeepMind.

48:00He's an incredibly researcher. He worked, he led to AFCO and Office 0 and he has an early gilding dedication to the field of reinforcement learning and he's probably one of the smallest people, or maybe the smallest person I know went. Amazing guy, amazing reinforcement learning engineer. And the second one I would say is Elya Satzkevich. And he was a co -founder of OpenAI. I had the opportunity to work with him just a little bit in the really early days of AlphaGo. But I think it's like his commitment to scaling IAM efforts and pushing the boundaries of what the systems can achieve is remarkable.

48:42And he got nature that like Jupiter, Jupiter, Jupiter, and Jupiter 4 happen. So, yeah, immense respect towards him. Thank you for sharing that. Let's close out with some rapid fire questions. Maybe first, what do you think will be the next big milestone in AI? I would say the next one, five and ten years. So I think like the next five, ten years, the world would be a different place. I actually really believe that. I think that the next few years, we'll see models becoming powerful in reliable agents that can actually independently execute tasks. And I think that AI agents will be massively adopted across industries, especially in science and healthcare.

49:24So it does send some really excited on what's coming in AI. And what I'm most excited about is AI agents, systems can actually do tasks for you. And this is exactly what we're building that perfection. In what year do you think will pass the 50 % threshold on sweet bench? I think we are one to three years away from the 50 % threshold for sweet agents and 35 years from achieving 90%. So, the reason is, while progress is amazing, I think it's like we still need reliable agent to hit these milestones and it's really, when it comes to research, it's hard to make precise predictions. When do you think we'll hit the data wall for scaling LLMs?

50:10And do you think all the research in RL is mature enough to keep up our slope of progress? Or do you think there will be a bit of a lull as we try to figure out what happens when we hit the wall? So I think like the wall, based on what I've read, I think we have at least one more year for text, just that before we hit the wall. and then we have like this extra model, which might actually buy us maybe a year extra. And I think we are in a really good place to just like start using synthetic data. So the next few years we'll just like figure out the synthetic data problem. So I think that we won't really hit the wall.

50:51Just like we'll hit the wall, but like no one realized it because we have like new methods in place. Do you think LLMs will have their AlphaGo moment? And if so, when? I think it's like LLM's hard draft of a government with the initial release of Chasipiti, where they showed Gaste the power and the progress made over the past decade. I think it's like what they hadn't had yet is their Alpha Zero Mode. And that's the moment where more compute directly translates to increase intelligence without human intervention. And I think it's like this breakthrough is still on the horizon. When do you think that will happen?

51:25I think it's going to happen in the next five years. Wow. Amazing. Yannis, thank you so much for joining us and taking us through the awesome history of Alpha Go, Alpha Zero, Mu Zero, your own journey through DeepMind. And then many of the core research problems that the whole industry is tackling today around data and building for reliability and robustness and planning and in context learning. We're really excited for the future that you're helping us build and that you're pushing forward in the field as well. So thank you so much, Janice. Thank you so much for having me.

From the publisher

Ioannis Antonoglou, founding engineer at DeepMind and co-founder of ReflectionAI, has seen the triumphs of reinforcement learning firsthand. From AlphaGo to AlphaZero and MuZero, Ioannis has built the most powerful agents in the world. Ioannis breaks down key moments in AlphaGo's game against Lee Sodol (Moves 37 and 78), the importance of self-play and the impact of scale, reliability, planning and in-context learning as core factors that will unlock the next level of progress in AI.

Hosted by: Stephanie Zhan and Sonya Huang, Sequoia Capital

Mentioned in this episode:

PPO: Proximal Policy Optimization algorithm developed by DeepMind in game environments. Also used by OpenAI for RLHF in ChatGPT.

MuJoCo: Open source physics engine used to develop PPO

Monte Carlo Tree Search: Heuristic search algorithm used in AlphaGo as well as video compression for YouTube and the self-driving system at Tesla

AlphaZero: The DeepMind model that taught itself from scratch how to master the games of chess, shogi and Go

MuZero: The DeepMind follow up to AlphaZero that mastered games without knowing the rules and able to plan winning strategies in unknown environments

AlphaChem: Chemical Synthesis Planning with Tree Search and Deep Neural Network Policies

DQN: Deep Q-Network, Introduced in 2013 paper, Playing Atari with Deep Reinforcement Learning

AlphaFold: DeepMind model for predicting protein structures for which Demis Hassabis, John Jumper and David Baker won the 2024 Nobel Prize in Chemistry

More from Training Data

All 110 episodes
ReflectionAI Founder Ioannis Antonoglou: From AlphaGo to AGITraining Data · 52 min
Listen in VO