In short
Retrospective on AlphaGo’s 2016 DeepMind Challenge vs Lee Sedol and how its “fast thinking + slow thinking” reinforcement-learning approach enabled breakthroughs beyond human play, influencing later systems like AlphaZero, AlphaFold, and agents such as AlphaTensor/AlphaEvolve.
Guests
Thore Graepel (DeepMind research scientist; key AlphaGo architect; also a Go player who lost to an early AlphaGo “baby version” on his first day at DeepMind). Pushmeet Kohli (leads DeepMind science work; explains how early Go techniques generalize to other search/optimization problems).
Key claims
Go was chosen because simple rules yield an exponential state space. AlphaGo combined policy/value deep nets with search to navigate combinatorial complexity. “Move 37” (a shoulder move on the fifth line) initially seemed wrong but was pivotal; “Move 78” confused AlphaGo and helped Sedol’s comeback. AlphaZero improved without human game data, rediscovering then surpassing human strategies. The match signaled AI could go beyond training distributions, with verification/refutation and verifiable objectives reducing “hallucinations.”
Notable examples
Move 37 probability “1 in 10,000” by humans; final match score 4-1; AlphaTensor framing matrix multiplication as a game; protein folding optimism; AlphaProof verifiable math proofs.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Historic Match of AlphaGo vs. Lee Sedol
0:45 to 2:14
Overview of the historic match between AlphaGo and world champion Lee Sedol.
“Not a single human player would have chosen Move 37.”
Why Go is a Challenge for AI
2:14 to 3:26
Discussion on the complexities of Go and why it was a perfect challenge for AI.
“Tori, I know you're an accomplished Go player yourself.”
Tori's First Experience with AlphaGo
3:26 to 5:58
Tori recounts her first day at DeepMind and playing against AlphaGo.
“And that is because not only because of the breadth of the search space, of the number of moves you can make, but also the depth, how long you have to reason and how long the games are.”
Understanding AlphaGo's Mechanics
5:58 to 9:06
Explains how AlphaGo worked and its unique approach to the game of Go.
“Yeah, so I think if you look at the game of Go, the number of moves that you can make at any given time, there are a finite number of moves.”
The Confidence in AlphaGo's Development
9:06 to 10:41
Discussion on the testing of AlphaGo against a professional player and the team's confidence.
“The slow thinking is not unlike what happened in Deep Blue.”
Preparing for the Match with Lee Sedol
10:41 to 13:00
Insights into the preparations and challenges faced before the match with Lee Sedol.
“But it did give us confidence and gave Demis confidence that we would be able to tackle even harder opponents in the near future.”
The World Watches AlphaGo vs. Lee Sedol
13:00 to 14:00
Describes the global attention and excitement surrounding the match between AlphaGo and Lee Sedol.
“I mean, were you nervous about the performance of AlphaGo?”
The Historic Match Begins
14:00 to 17:08
Explore the intense atmosphere and expectations before AlphaGo's matches against Lee Sedol.
“We also needed to make sure that the system is really stable.”
The Surprising Move 37
17:08 to 19:29
Understand the significance of AlphaGo's unexpected Move 37 and its impact on the game.
“Professional commentators almost unanimously said that not a single human player would have chosen Move 37.”
Lee Sedol's Resilience
19:29 to 24:33
Discuss Lee Sedol's pivotal Move 78 and its implications for human versus AI competition.
“Something that went beyond what a human growth player would normally do, Hushmeek.”
Show all 22 chapters
Reactions from the Go Community
24:33 to 25:40
Learn how the Go community responded to AlphaGo's victory and its effects on the game.
“What was the reaction from the Go community?”
The AI Revolution: AlphaZero
25:40 to 28:06
Discover how AlphaZero advanced AI by learning Go without any human data.
“And to show that you can go beyond that distribution and that insight then can be utilized by the world, I think is an amazing sort of insight that comes out of this whole experience.”
The Learning Process of AlphaGo
28:06 to 29:21
Discover how AlphaGo learns to play by gathering experience from games.
“rules of the game and means of representing and learning these functions that we talked about, the policy net and the value net.”
Rediscovery and Innovation in Gameplay
29:21 to 30:36
Explore how AlphaGo rediscovered human strategies and then surpassed them.
“I'm not going to continue playing in this human way.”
Behind the Scenes of AlphaGo's Success
30:36 to 32:06
Listen to a private conversation that reflects the optimism around AI breakthroughs.
“I'm telling you, we can solve protein holding.”
AI's Role in Understanding Complex Problems
32:06 to 33:34
Learn how AlphaGo's success inspired confidence in solving other complex challenges.
“And this is now the point really where you come on board with the DeepMind Teamfish meet, because when it came to Alpha Fold, I mean, you're an integral part of that story.”
Advances in Search Algorithms Post-AlphaGo
33:34 to 35:34
Understand how AlphaGo's search algorithms influenced scientific applications.
“to come and really lead the charge on how can AI be used for these applications.”
Matrix Multiplication and Its Implications
35:34 to 37:34
Explore how matrix multiplication underlies machine learning systems and AI.
“The issue is that the search space for that problem is even larger than the search space for Go.”
Interpreting AI Decisions and Human Intuition
37:34 to 40:06
Examine the challenges of interpreting AI decisions compared to human intuition.
“or how do you sort of tackle these logistics problems where you are trying to move packets around in a network?”
Hallucinations in AI and Their Verification
40:06 to 42:02
Learn about the risks of AI hallucinations and the importance of verification.
“look, this is a better move than what AlphaGo played.”
Exploring AI Beyond Human Limitations
42:02 to 48:25
Learn about the challenges and strategies in advancing AI beyond human knowledge.
“But then if those large language models are based on human data, is there a danger of you limiting yourselves to what humans have already discovered?”
The Impact of AlphaGo on AI Development
48:25 to 52:12
Understand how AlphaGo marked a pivotal shift towards achieving superhuman intelligence.
“But if we are talking here about advancing scientific knowledge and understanding beyond what humans have done, do you think you've seen examples of Move 37 in science already?”
Transcript
Automatic transcript. May contain errors.0:00Thore Graepel:Welcome back to Google DeepMind, the podcast. I'm Professor Hannah Fry. Picture the scene. It's March 2016. Inside a hotel suite in Seoul, South Korea, two players are playing the ancient game of Go, a game of unimaginable complexity, long thought impossible for a machine to master. On one side is Lee Sedol, a legendary 18-time Go world champion. On the other, AlphaGo, a neural network-based AI system built on a powerful technique called reinforcement learning. Welcome to the DeepMind Challenge live in Seoul, Korea. That's a very surprising move.
0:45Pushmeet Kohli:Not a single human player would have chosen Move 37.
0:50Thore Graepel:After hours of intense gameplay spread over seven days... Yeah, that's an exciting move. Lee Sedol placed two stones on the board to signal his final resignation. And in the blink of an eye, the world changed. The final result of 4-1. Congratulations to AlphaGo and to the entire team.
1:14Thore Graepel:That was exactly one decade ago. And the field of AI has changed unimaginably since then. We have seen the rise of large language models, the growing sophistication of AI agents, and the solving of scientific grand challenges like protein folding. But in many ways, the modern AI revolution arguably began right there on that wooden board in South Korea. So in this episode, we wanted to look backwards and forwards to how a bold experiment in teaching machines to play games became the foundation stone for the AI breakthroughs of today. And with me are the perfect guests to tell that story. Tori Graeful is a distinguished research scientist at Google DeepMind who was right there in Seoul as a key architect of the AlphaGo project.
2:03Thore Graepel:And Pushmeet Kohli, who leads Google DeepMind's science work and is the person to tell us how those early techniques pioneered in Go can tackle crucial problems today. Welcome to the podcast, both of you. Tori, I know you're an accomplished Go player yourself. Just explain to us why Go was seen as a good challenge for AI.
2:24Pushmeet Kohli:Yes, the game of Go seemed like the perfect challenge for AI because the game has such simple rules, yet it leads to such complex gameplay with tactics and strategies and complex patterns. and once the game of chess had been solved as it were or at least you know deep blue had won against the world champion then go was this open challenge it's much more complex than chess by many orders of magnitude and nobody was expecting it to be solved anytime soon yet it it looks so elegant and simple for computer scientists. And so it was the perfect game to tackle at the time.
3:12Thore Graepel:I mean, that idea of nobody thinking it would be solved anytime soon, that sort of hits the nail on the head, right, Vishmi? I know you were working at Microsoft at the time, but just how complex was this problem considered to be? I think it was considered extremely complex. And that is because not only because of the breadth of the search space, of the number of moves you can make, but also the depth, how long you have to reason and how long the games are. In the game of chess, you might think about reasoning about 60 to 70 sort of moves. In the game of Go, it's much, much longer. And that leads to the challenge of the problem.
3:51Thore Graepel:Tori, I know when you first started at DeepMind, being a Go player, didn't you play against Alpha Go on your first day.
3:58Pushmeet Kohli:Yeah, yeah, exactly. So imagine I come first day at work at DeepMind. I know a couple of people, including David Silver, and he asks me, Torre, you're a Go player, right? Couldn't you do us a favor and test our baby version of something that wasn't even called AlphaGo at the time? Of course, you know, it was an internship project and they had just about taken a few thousand games from the internet and had trained a system, or a few hundred thousand games maybe. And I had the opportunity to be one of the first people to play against it. But you can imagine, I was excited, but I was also nervous.
4:42Pushmeet Kohli:It was my first day at work, and there I was being dragged to a centrally located table. On the other side, I think it It was Aja Huang, who would later be known as the Hand of AlphaGo with his poker face. And I got to play against this baby version of AlphaGo.
5:03Thore Graepel:With people watching, presumably. With a lot of people watching all around me.
5:07Pushmeet Kohli:You know, there was no escape. Later, Demis showed up. And of course, David was there the whole time. Yeah. And so what does one do? Play conservatively, right? So I just thought, just don't make a mistake. surely this can't be so hard. But of course, that was exactly what that version of the program was good at. It was trained on human professional games. So it knew exactly what to do against conventional play. And so as this little test match proceeded, my position became worse and worse, and I ended up losing by a small margin. But I took the crown of the first person who officially lost against AlphaGo.
5:48Pushmeet Kohli:So it was quite the experience. And of course, afterwards, everyone knew me. It was a wonderful way of introducing myself. A humbling way. A humbling way, exactly.
6:00Thore Graepel:Absolutely. Prashmi, just remind us, I mean, okay, so I know that the algorithm advanced quite substantially from that early point where it was an internship, but just broadly, explain to us how it worked and this idea about cracking the kind of combinatorial spaces in particular. Yeah, so I think if you look at the game of Go, the number of moves that you can make at any given time, there are a finite number of moves. But if you look and reason about the overall game state, it's exponential. And that exponential growth in the number of states that you have to reason about is what makes the game extremely complicated.
6:39Thore Graepel:So how did they crack it then? Just remind us of the solution that they discovered. The beauty of AlphaGo was there is this element of thinking fast and thinking slow. And AlphaGo, in some sense, was the perfect combination of those thinking fast and thinking slow processes coming together to take on this extremely large search space.
7:02Pushmeet Kohli:And it matches quite well to how humans play the game, I think. You know, if you imagine how a human would play a game of chess or a game of Go, we also have the capacity to look at a position and pretty quickly appreciate if that's good for black or good for white. And we can also look at a position and already see moves that seem promising. We never look at all the possible moves, which would be maybe 20 or 30 in chess or 200 or 300 in Go. we immediately draw on to certain, maybe even aesthetically pleasing moves that seem like just the right ones guided by our intuition. And that element is complemented by planning, where we explicitly reason through the possibilities.
7:45Pushmeet Kohli:If I make this move, my opponent might make that move, and then I have to counter with this move. And these two different ways of thinking come together in how humans play these games, and they also come together in how AlphaGo plays.
7:59Thore Graepel:The intuition and the calculation, as it were. Exactly. So was that the inspiration then? Did you sort of think about how you were playing the game, how other Go players were playing the game and draw that direct inspiration from neuroscience effectively, as it were?
8:13Pushmeet Kohli:Yeah, I think that is definitely one direction because a lot of team members were actually game players who were able to introspect and see how we tackle the game. And then, of course, that comes together with deep learning that at the time, you know, since 2012 had grown as a direction. And now for the first time, gave us the tools to learn these approximate functions. For example, the value function that takes a board and tells us how good it is for either black or white. Or the policy network that takes a board and effectively ranks the available moves according to how likely it would be that a professional player would take them.
8:57Pushmeet Kohli:And so deep learning was just ripe at the time to tackle this problem and gave us the opportunity to implement the fast thinking. The slow thinking is not unlike what happened in Deep Blue. You know, it's the search of the game tree that was already known and that we might now call good old fashioned AI. Okay.
9:19Thore Graepel:Well, I mean, you lost to this thing quite early. But once it had gone through a lot of the people on the team, let's say, I know that you tested it with a professional Go player because you had Fan Hui come into the office.
9:31Pushmeet Kohli:Yeah, exactly.
9:31Thore Graepel:How confident were you at that point that it was going to beat him?
9:35Pushmeet Kohli:Yeah, we had different levels of confidence, which was really interesting. So we had been really lucky to find him. You know, he was the European Go champion at the time. He lived in Bordeaux and came over. We lured him into playing this game with us. And the setup was that he would play 10 test games against the version of AlphaGo at that point. And I personally thought that AlphaGo cannot possibly be at the point already that it beats the European champion, a professional player. And so I had a bet with David Silver. David Silver was confident. He said, I think AlphaGo is going to nail it 10-0.
10:20Pushmeet Kohli:And I said, no, I think AlphaGo will lose at least one game. And the bet was that whoever lost would have to show up at the office dressed as an ancient Japanese go-master and be in the office for one day with that. Well, who showed up like that? It was me. Because it was, in fact, 10-0. But it did give us confidence and gave Demis confidence that we would be able to tackle even harder opponents in the near future.
10:56Thore Graepel:Which you did, of course, on a plane you got in 2016 to Seoul and Korea to play against Lee Sedol. I mean, just tell us, give us a sense of how phenomenal a player he actually is.
11:07Pushmeet Kohli:Yeah, so Isidol was really one of the or maybe the best players at the time with an incredible track record of winning tournaments. He was compared to Roger Federer at the time for his success and intellectual brilliance. And so for us, it was a tremendous honor that he accepted our challenge to play against him. And it was a tremendous challenge because we had to set a date, right? You can't just say, you know, we'll tell you when we're ready. A date was set and we had to work towards that date to actually make AlphaGo strong enough. And what added tension and excitement to it was that Isidol was convinced that he would win.
11:57Pushmeet Kohli:He thought it highly unlikely at the time that AlphaGo would win. And of course, he was basing his assessment on the game records that he had seen against Fan Hui, and he assessed that he was better. But of course, what he wasn't so aware of is that AlphaGo was constantly improving through the training and the algorithmic refinements and so on that we made. And so the entire team basically went to South Korea, and you wouldn't believe the excitement of people there. You know, the truth is, in England, Go is a bit of a niche activity, right? Very few people would be able to play it or even know about it.
12:37Pushmeet Kohli:But in South Korea, people were so excited. The best Go players are celebrities. And, you know, we came there and there were hordes of photographers that took pictures. We had a documentary film crew with us. And so imagine typical computer geeks, as it were, suddenly in the limelight of the world for this match. That was quite the adventure.
13:02Thore Graepel:Yeah. I mean, were you nervous about the performance of AlphaGo?
13:07Pushmeet Kohli:Yes, we were definitely nervous. So, of course, we had a very sophisticated evaluation pipeline. You can test against players that you have access to, like Fonhui. That was super helpful. You can also test against previous versions of the program. And you can calculate what we call the ELO score of the system, which basically takes the outcomes of all the games that you play against other versions, maybe earlier versions of your program, and calculates what the rating of the new version is. And you can calibrate these things quite well. But of course, we didn't know where on that scale E-Sidol would be.
13:48Pushmeet Kohli:And of course, we wanted a cushion as well. You know, it would be nice to be quite a bit better, to have some certainty. Because this is the world stage, right? If you lose this, that's a bit of a hit to the reputation. And so, yeah, we were nervous. We worked up to the last minute. We also needed to make sure that the system is really stable. You know, you don't want to make last minute changes to make it that little bit better. but risked that it now becomes unstable. But in the end, we were quite happy with it. And so we entered that now kind of famous hotel floor where all the action happened, where all the press was waiting and so on and embarked on the match.
14:32Thore Graepel:And people were watching from around the world, including Bushwick. So, I mean, where were you at this point? You were watching on? Yeah, I was in Seattle. I mean, I really started getting into it in the middle of the first game. It became so clear that AlphaGo had reached that specific milestone. And you could even see the reaction from the press and the commentators and LisaDoll himself. Well, it's interesting that you said in the middle of that game, because in the early stages of that game, was it clear who had the upper hand? I think from a person who was just watching it, I felt that in the early stages, everyone felt quite confident that Lisa Doll would win.
15:24Thore Graepel:In fact, only as the game progressed and it became closer to the final outcome that they realized that as you count the territory, AlphaGo had an advantage. In fact, it came as a surprise to people. What did you think?
15:41Pushmeet Kohli:Yeah, so I had this interesting interaction on site with a professional Go player, an American professional Go player, who was sitting next to me while we were watching. And there was some sequence unfolding in a corner. And he kind of approached me and said, you know, I always tell my students not to play that stupid move that AlphaGo just played. So, I mean, it's pretty hopeless. And I was like, I'm not as much as an expert. Let's just wait and see was my reaction. And then after that first game, this gentleman came to me and said, this is the most phenomenal thing I've ever experienced. I'm so grateful that I'm allowed to be here to witness that a machine can play go at this level and there's going to be so much we can learn from it.
16:35Pushmeet Kohli:And he was already embracing this. I mean, you have to imagine these people dedicate their lives to the study of this game. And they've often trained from being young children to their current age just to master this game. And so, of course, it comes as a shock to them that a machine might match or even exceed a human Go player.
16:57Thore Graepel:Because if that was the first game, when AlphaGo won, in the second game, AlphaGo did something that, I mean, really surprised everybody.
17:07Pushmeet Kohli:Wow, that's a very surprising move. Professional commentators almost unanimously said that not a single human player would have chosen Move 37. AlphaGo said there was a 1 in 10 ,000 probability that Move 37 would have been played by a human player.
17:25Thore Graepel:Alphago is actually a probability of improving the machine and losing it. But it's not that you're looking at it. Alphago is quite a creative way. Just explain to us what happened with the now famous move 37.
17:42Pushmeet Kohli:Yeah, so this was a remarkable scene and I was sitting in the international English speaking commentating room and uh and michael redmond our american commentator he had this big demo board on the wall and he would put all the stones up there on the board to show people what was being played and comment on different variations and so he took the stone corresponding to move 37 on the board and then he stepped back and said ah this must be wrong and he took it back and then he looked at the screen again and said, no, no, that is actually what AlphaGo played. And he put it back. He was puzzled. You could see it, that that was such a counterintuitive move for a human player.
Read the full transcript
18:27Pushmeet Kohli:It was a shoulder move on the fifth line. And this is typically something that human Go players avoid. So often in Go, there is some kind of pushing going on along the edges. And one of the players builds territory along the wall of the board, and the other side builds influence towards the center of the board. And if that happens on the third and fourth line, this is considered to be roughly equitable. You know, both sides get something out of it. But what AlphaGo was effectively suggesting is that it's still profitable if you do it on the fifth line, and you give that much more territory to the other party.
19:10Pushmeet Kohli:And that's what was so surprising to people, that there would be situations in which that would be correct. And so not only was it a very special move, but in a way it represented a new way of weighing these two factors of immediate territory versus influence towards the center of the board against each other.
19:30Thore Graepel:Something that went beyond what a human growth player would normally do, Hushmeek. Yeah, absolutely. I mean, there are moments like this where you see the true potential of an AI system. expanding human knowledge where people have regarded in this particular case, the game of Go as a thing to be studied for many, many years. And there comes this particular point where that knowledge is expanded. And people are at first skeptical, which was the case in the game as well. When the move was played, it was considered a hallucination or a mistake, right? for quite a bit of time before its implications became clear.
20:16Thore Graepel:Later on in the game. Exactly. Because it proved to be pivotal to the second win. Yeah. It was not just a moment in that game, but it was also a moment, I think, in the whole sort of history of AI where that particular moment showed us that there will be times when these systems will produce insights which we might not even be able to discern whether they are the right things or amazing breakthroughs. But yet they will have a lot of influence in how we look at whole areas of study in a completely new light. Well, I also want to talk about Move 78. This is a move that was played by Lisa Dahl that confused AlphaGo, causing it to resign the game.
21:08Thore Graepel:What is Lee Sedol up to here? He's just burned like seven or eight minutes just on this move already.
21:22Oh, look at that move.
21:24Pushmeet Kohli:That's an exciting move.
21:25Thore Graepel:Ooh!
21:26Pushmeet Kohli:You know, I'm not actually sure what AlphaGo is trying to do here.
21:32Thore Graepel:What's this?
21:49Thore Graepel:so by this point AlphaGo has won three games in a row and now Lisa Dahl does a move that
21:56Pushmeet Kohli:confuses the system is that fair to say yeah that's absolutely fair to say so move 78 was an unusual wedge move that Isidol played. There had been a very interesting battle, as it were, at the center of the board. And Isidol found this move, and it was also surprising to people, similar to Move 37. and from then on we observed that AlphaGo didn't have a good grasp of the position anymore. We saw that the moves that it made didn't really make sense to us in a bad way. You know, 37 also didn't make sense to us maybe, but these moves even to amateurs like us seemed strange and so it had been confused by the move.
22:50Pushmeet Kohli:And just to zoom out to give you a sense of why this still mattered so much. So you might say, okay, it's a match of five games and AlphaGo has won the first three. What more is there to prove? But then we were thinking, well, if now Isidol was to win the last two, what would you conclude? He's got it figured out, right?
23:14Thore Graepel:He's bound the fragility.
23:16Pushmeet Kohli:Exactly. So it would have been the human triumph. And so that's why that game and the last one were still very exciting to us. But it wasn't entirely the case that we were disappointed. We were certainly disappointed, but also we had so much admiration for Issa Doll to, you know, as a human, to be able to find this move. You just have to imagine this master who has dedicated his life to playing this game in this battle that must have been so hard on him, right? to see this machine play so perfectly and him struggling to find a way. And then in game four, he finds a way. And as he put it in the press conference, I think later, he said that he was so happy and proud that he was able, maybe for the last time on behalf of humanity, to find a way to overcome the machine.
24:14Thore Graepel:Because some people called it the divine move, didn't they?
24:17Pushmeet Kohli:Yeah, yeah. And I think given the tension at that point in time and him really outgrowing himself at that moment and finding that move, I think it's a good name for it.
24:32Thore Graepel:Well, the final score was 4-1 to AlphaGo in total. What was the reaction from the Go community?
24:39Pushmeet Kohli:Yeah, so the Go community followed the match very closely. And of course, the outcome was dramatic and for many people, unexpected. And so people showed very different reactions. You know, some people were absolutely amazed and surprised about the outcome. Some people couldn't believe it. Others, of course, also thought that some era had come to an end because now maybe the strongest Go player was no longer a human, but a machine. but overall what we found amazing is that there was an uptick in interest in the game of Go I think more people play Go now than did before and the Go community really embraced the learning from AlphaGo so there are now many programs that work essentially the same way that AlphaGo does and people use it for teaching purposes they analyze their games through it and overall I I think it has provided a lift to the whole Go community.
25:39Thore Graepel:Let me ask you about the reaction from the AI world to this match. What was the buzz? What was the conversation like? The Lee Sedol match, the AlphaGo Lee Sedol match, was a key pivot point where a lot of people, especially in the machine learning community, who have been sort of working on these models and techniques, as a mathematical and applied project started to see evidence that these systems can self-learn and go beyond human knowledge. And that is a very important sort of point because in machine learning, you train with training data which has been collected and your natural sort of expectation is that the model is going to just be consistent with that distribution.
26:35Thore Graepel:And to show that you can go beyond that distribution and that insight then can be utilized by the world, I think is an amazing sort of insight that comes out of this whole experience. and it really points to what is possible with artificial intelligence in not just the game of Go but in the understanding of the world, in chemistry, in biology, in mathematics, in computer science. What are these amazing analogs of Move 37 that these systems will be able to discover and reveal to us? I think that point that you made there about going beyond human intelligence is just so fascinating. But one of the things that I find most intriguing about the AlphaGo story, even after the victory of 4-1, is that you then built AlphaZero, where you took away all of the human data, all of the games of Go that it had been trained on, and discovered that once you take out the human intelligence, the thing actually improved, which is astonishing to me.
27:48Pushmeet Kohli:Yeah, from a scientific perspective, one could argue that that is an even bigger step than the original AlphaGo. Because as you were saying, the AlphaZero system doesn't have access to any human game records, how humans play, didn't have access to prior knowledge about the game, how the game is played, but really only had access to the rules of the game and means of representing and learning these functions that we talked about, the policy net and the value net. So basically, it starts playing entirely randomly at the beginning because it has no notion of what good or bad moves are, but it gathers experience from playing these games and it It learns what are moves that are more likely to lead to a win, what are moves that are more likely to lead to loss, what are positions that look promising, what are positions that are not promising.
28:43Pushmeet Kohli:And eventually, it starts playing better and better moves. And now, of course, it's not limited by human knowledge. And what it discovered was amazing. amazing. So first of all, it rediscovered ways of how humans play. And that was totally reassuring. You know, there are certain patterns in the corner in Go that we call joseki, or in chess, there are certain opening moves. The system was now more general. It could play chess, Go, and shogi, and could have played any number of other board games if we trained it that way. and so at first it rediscovers human knowledge and we think wow this is so cool it it finds the same openings and so on and then we look at some of these openings and it stops playing them we think what's going on it has found a refutation so it discovered rediscovered human knowledge and then it discards it because it has now gone beyond it and has found there's actually better ways of playing.
29:48Pushmeet Kohli:I'm not going to continue playing in this human way.
29:50Thore Graepel:Stuff that humans hadn't found yet, effectively.
29:53Pushmeet Kohli:Exactly. For AlphaZero, when it played Go, the way it played Go looked alien to me in the end. So this wasn't the kind of Go that I had learned from my Go teacher, you know, which is structured maybe in a way that enables humans to understand it. These moves looked very free and didn't make much sense at the time. But 30 moves later, everything would fall into place and you'd see, oh yeah, oh wow, that makes sense now. And so on, as if it had the foresight in a way, which it did, right? So that discovery from nothing to that level of play was very impressive. Okay.
30:36Thore Graepel:So there's something I want to show you, something that happened actually when you guys were in Seoul because you as you mentioned before you were being filmed for this documentary for um for AlphaGo and there's some footage that didn't make it into the film um but it was it was captured by the cameras as they were sort of packing up but the microphones were still running I don't know if you've um you've heard this little clip let me play it for you hold on this is Demis and David having a sort of private conversation it's just amazing seeing how quickly the problem
31:04Pushmeet Kohli:that is seen as being impossible can change to being practically done.
31:09Thore Graepel:I'm telling you, we can solve protein holding. That's like, I mean, it's just huge. I'm sure we can do that. I thought we could do that before. Yeah. But now we definitely can do it. Beautiful. Isn't that great?
31:25Pushmeet Kohli:Yeah.
31:25Thore Graepel:Tori, do you think that that captured the mood at the time?
31:29Pushmeet Kohli:Yeah, that was the kind of door that AlphaGo opened at the time. If we can do this, then what else could we do? Because this is a game with 10 to the power of 170 different positions. This is super complex. And if we have principled ways of navigating that kind of combinatorial search space, then it seems plausible that we would also be able to handle other large combinatorial search spaces. And at the time, one of the favorites was protein folding.
32:05Thore Graepel:Absolutely. And this is now the point really where you come on board with the DeepMind Teamfish meet, because when it came to Alpha Fold, I mean, you're an integral part of that story. Did Alpha Go, did that project directly influence what you guys went on to do? Or was it sort of like the confidence of a victory that made Demis say things like that? No, I think Demis, from very early on, I think he has a very strong notion of what AI is being developed for. He really sees AI as a tool that will help us understand the world better. In fact, at the time when the AlphaGo matches were happening, I was at Microsoft working on AI for programming.
32:51Thore Graepel:Now AI for coding is everywhere, but at that time, not many people were working on program synthesis and AI for coding. And Demis wanted me to join DeepMind. And my question to him was, I am really interested in having AI systems, machine learning systems for solving the most challenging problems in the world and to make sense of what's happening. And I think his reaction was, if you want to understand the world and if you want to solve the most important problems in the world, then you have to join DeepMind because we will need AI to really understand the world deeply and to tackle these problems.
33:33Thore Graepel:So if you are interested in sort of learning to program, if you are interested in cybersecurity, if you are interested in dealing with climate change, if you are interested in understanding how to deal with impossible to treat sort of diseases, you have to come and really lead the charge on how can AI be used for these applications. I want to ask you about some of the innovations that you guys had made in AlphaGo and how they ended up finding their way into the science projects that you guys were doing. One of the big things that AlphaGo did was to make that gigantic search space more tractable.
34:12Thore Graepel:So how have search algorithms changed since then and how are they being used in science? I mean, search is such an integral part of many problems that you encounter in the real world. We just spoke about protein folding, which could be considered as the search over the space of all possible structures. But just to give a more sort of simpler example, you can think of search as also the search of algorithms for solving a particular problem. So everything around us that computers do has some form of matrix multiplication underlying it. So even the fact that we have these machine learning systems and neural networks that are changing the world today, these neural networks are based on matrix multiplication, essentially taking large matrices of numbers and multiplying them together.
35:07Thore Graepel:And even the very simplest operation of matrix multiplication, which is just taking two matrices and multiplying them, is the simplest thing that you sort of learn in school and college. And yet we don't know, as a whole research community, what is the fastest way of multiplying two matrices. So if you think about that problem, you can reason about it as a search problem. You can say there's a space of possible algorithms and now search over that space of algorithms and try to find me the best algorithm. The issue is that the search space for that problem is even larger than the search space for Go.
35:51Thore Graepel:So one of the first things that we needed to do is we came up with this agent called AlphaTensor, which made matrix multiplication as a search problem, as a game. So instead of, did you win or lose the game of Go, you're saying, did you multiply these two matrices together quickly or not? Yeah. Did you multiply these matrices completely accurately in the smallest number of moves? And that was the game. And there was an algorithm that Strassen in 1969 had come up with. And since then, for 50 years, there was no progress. And then AlphaTensor found a better way of multiplying these two matrices.
36:42Thore Graepel:And that was a key sort of proof point of what is possible with the same sort of techniques. In case there's anyone watching who's sort of, I don't know, maybe not that familiar with the things you're talking about, matrix multiplication, for example. I mean, we need to be really clear on the potential of this thing. I mean, every single large language model in the world is essentially at its heart just a massive matrix multiplication problem, right? Yes. All of the fuss about different chips that are being made is because some of them can multiply matrices faster than others. And what you're describing here is like turning that into a game and even small gains that you might make on how quickly you can do something.
37:21Thore Graepel:Once you scale it up to the size of how much everybody in the world is using AI, we're talking about gigantic differences. Yeah, absolutely. And since then, what we have done is we have said, and let's not just tackle matrix multiplication. Let's tackle all the possible algorithms that you can think of. So our new agents like Alpha Evolve, they search in the space of all possible programs, trying to find the best algorithm that can solve these important problems, whether it's how do you schedule jobs in a data center, which is an extremely important problem and has implications in terms of energy, compute utilization, and so on.
38:02Thore Graepel:or how do you sort of tackle these logistics problems where you are trying to move packets around in a network? So the same basic methodology of tackling these search problems now has been expanded in terms of what you can do with it. Okay, but I'm thinking here about the policy network, the intuition as you described it, where a Go player might look at the board and say, I think this is a fruitful direction in which to search. if you instead of a board instead of a game of go you've got all possible algorithms of everything in the entire world and beyond let's say how on earth do you create intuition in that sort of a situation how do you know how to narrow down the search space yeah so i think this is this is a very interesting sort of research topic that we are now starting to think of when we apply agents like Alpha Evolve to discover these new algorithms, sometimes those algorithms are not very intuitive to us.
39:06Thore Graepel:In fact, they could be counterintuitive. So sometimes you can see the patterns. You can see that there are certain symmetries in the problem that we did not understand. Mathematicians did not understand. Computer scientists did not understand. But somehow there were those symmetries. The agent somehow discovered those symmetries and then it exploited and utilized the symmetries to make the solution much more efficient. In some cases, we just don't understand how it made things faster, but they are faster. And then our challenge is that when you think about collaboration, where humans and these AI agents are working together, then how do we make sure that the systems that are produced and the algorithms that are produced are interpretable by the human computer scientists and engineers.
39:54Pushmeet Kohli:It reminds me a little bit of this situation in AlphaGo where people in the end game were observing AlphaGo and found that it didn't quite play optimally. And they were really surprised to say, look, this is a better move than what AlphaGo played. You know, is it not playing well? Is it making mistakes? And the solution was that AlphaGo was optimizing the objective we had given it, which is to maximize the probability of winning the game. Humans tend to use a heuristic, which is they want to have more territory than the opponent by some margin. And they think the larger the margin is, the better it is for them, which is often true.
40:38Pushmeet Kohli:But AlphaGo doesn't care about the margin. For AlphaGo, it was enough to win by half a point. And so often in the end game, it was almost toying, seen to be toying with the opponent and giving up points just up until the point where it was sure it could win by half a point. And sometimes you get these counterintuitive behaviors, but if you then drill deeper, you can see why they come about.
41:02Thore Graepel:Because the algorithm and the humans are ultimately optimizing for slightly different things. Exactly. Yeah. Okay, but then that does make me wonder. So Move 37 as an example of where it went beyond what humans are able to do. At the same time, when Move 37 first came through, people thought it was a mistake, right? So how can you tell the difference? I mean, if the algorithm comes up with something that is original, can you be sure it's not a hallucination? Yeah, and I think this is a very important point, right? Like with the large language models, especially when they were being developed initially, in the first versions of them, they would hallucinate.
41:41Thore Graepel:They would come up with solutions which were not correct or come up with responses which were completely invalid. And this is where the importance of the agent harness comes into play, where you couple the large language model with a verifier, which is able to sort of prune out when what is being hallucinated and what is actually something that might be remarkable that we need to investigate further. But then if those large language models are based on human data, is there a danger of you limiting yourselves to what humans have already discovered? I'm thinking of what's already in the textbook, as it were.
42:20Thore Graepel:When we build these agents, we deliberately increase the amount of things that they have to explore. So we tell the models that you have to go beyond the distribution that you were trained on, and you should feel free to explore more. And in fact, you might sort of produce new things which might not be appropriate or not be correct. But we have that verifier and evaluation function to prune out those insights.
42:46Pushmeet Kohli:I think this is really how Karl Popper would also characterize the whole scientific process. conjecture and refutation is the famous essay. And, you know, conjecture is maybe hallucination. It's this production capability of producing plausible hypotheses. And then refutation is the step by which you filter out the things that are wrong, that don't work. And I think it also makes clear why the current AI capability landscape looks like it does. Namely, it is very good in verifiable domains. Code is a verifiable domain. You define the objective. You can write down tests for the code. The first test is that it compiles, you know, it's already a good sign.
43:32Pushmeet Kohli:Then you test it on those tests, but you have hard criteria to reject failure, which is super important for these kinds of tasks. If you don't have it, things become much trickier. For example, if you work on open scientific problems, you might not have a verifier who can tell you that this is right or this is wrong. Ultimately, often experiment, physical experiment will be the verification that you need.
43:59Thore Graepel:Right. But that's quite a long, long way down the road, isn't it? I guess the experimental part of it. Because I'm just wondering here about interpretability coming back to the point that you made earlier. Does it matter that you might end up with results that are not easily interpretable here, given that the stakes are so much higher than they are on a board of a go game? Yeah, I think it does matter. Science is also about communication, right? If you can come up with this new insight, but if you are not able to communicate and people are not able to build on top of it, then there are limits to what the impact that will be achieved, right?
44:34Thore Graepel:So interpretability plays a very important role, but it's not the only thing. Take the example of AlphaFold. AlphaFold is able to solve this amazing problem of protein structure prediction. Do we understand completely the conceptual sort of operations that it does? Like at the mechanistic level, yes, but we don't know completely the underlying theory that can be used to recreate a human level reasoning process to make the same predictions. And we will somehow need to convert them to a human digestible form that the bounded rational human mind will be able to comprehend.
45:22Pushmeet Kohli:I think there's a really interesting point there, which is that an explanation not only needs to account for the phenomenon that you're explaining, it also needs to account for the intellectual level of the recipient of the explanation. So sometimes on YouTube, you can see these things, life explained at the level of a six-year-old, an eight-year-old, a 10-year-old, a 12-year-old. I quite like the explanations for 12-year-olds, I have to say. And that reflects this fact, right? An explanation really is a bridge between the phenomenon and our capacity to understand it. So it may very well be the case that future AI systems come up with explanations that might seem simplistic to them, but that are just about right for us to keep up with the AI system, right?
46:13Pushmeet Kohli:Yeah, exactly.
46:14Thore Graepel:I mean, if you look at our agents like AlphaProof, what they are able to do is you give them open maths problems, and they will give you a proof, and that that proof is verifiable. You can tell whether it's correct or not. Exactly. Even if you don't understand it. Yeah, you might not understand it, but you know it's correct, right? The uncertainty about whether the original theorem was correct or not is now resolved. But do we completely understand it? Like, in fact, till now, the results that we have had, we have spent the effort and then converted those results into a form that mathematicians have been able to see and say, yes, it makes sense.
46:56Thore Graepel:I can actually translate it in English and it all works. But there are two key phenomena that come out of it. One is that the importance of framing the problem now rises. Because if you don't, one of the challenges when we are trying to solve these very hard maths problems, when we are giving the agent these hard problems, is to specify the problem accurately so that the agent can now understand what is the reward function that it needs to optimize for. And then once it finds the solution, then there's the challenge of actually converting the solution back to a human readable form. If we do get to a point though, where an algorithm can just come up with its own proof, where's the role for mathematicians in all of this, speaking selfishly?
47:47Thore Graepel:No, I think mathematicians are even more important. today. Because what these agents are able to do is they're able to solve these incredible problems. But what are the problems that need solving? How do you specify that problem? That's where mathematicians and scientists come in. I do like the idea, though, that one day there might be, I don't know, Riemann hypothesis, and it comes back and says, yes, there's a proof. Unfortunately, it's beyond any human's ability to understand it. So, you know, sorry about that. But actually, I mean, I'm joking slightly. But if we are talking here about advancing scientific knowledge and understanding beyond what humans have done, do you think you've seen examples of Move 37 in science already?
48:39Thore Graepel:Yeah, I think absolutely. I think just the example of the matrix multiplication algorithm, it is something that people had studied for many, many years. And yet we have been able to come up with a new algorithm. So that is genuinely a Move 37 moment in algorithmic discovery. And I think we are now seeing the same thing in many other areas of science, in mathematics, in material science, coming up with new structures that we think now are stable. so there are a number of these things but the original move 37 moment is still very relevant because it was in some sense the first and it brought about that concept of going beyond human understanding i am thinking here about alpha zero again and how that really moved away from human data and showed these profound results large language models on the other hand ended up being almost a shortcut to intelligence, I guess, that was based very much on human data.
49:48Thore Graepel:Was that a sort of surprising turn of events for you?
49:52Pushmeet Kohli:Yes, I think that is an interesting thing that we observed. DeepMind was based on this idea that we use games as a microcosm of the real world. And the philosophy of DeepMind had been to place agents within these environments and let them learn how to master them and thereby grow their intelligence. And then what happened with large language models was really this discovery that there's a shortcut, that somehow there's this huge amount of crystallized intelligence, if you like, stored in the form of data on the internet, first text data, maybe images, maybe videos and so on. And that the shortcut is really to first mine all of that data and train systems based on that.
50:43Pushmeet Kohli:And that's basically the first and second generation of large language models that are based on that. But then, of course, you come to the point where, first of all, that doesn't lead you to novelty. You're now within this corpus of existing human knowledge, And we know how competent these models are within that. But it's very difficult to get out of that now. How do we go beyond what we already know? And that's, I think, where now the community, for the past few years, is exploring the methods again that DeepMind pioneered early on. Others, of course, reinforcement learning in environments. Part of the post-training now is routinely forms of reinforcement learning and either on human generated data or also on problems on environments like coding environments and so on.
51:37Pushmeet Kohli:And so now we're in a period where we're going again beyond human knowledge.
51:43Thore Graepel:Pushmi, do you think that we would be here at this moment in the AI revolution if it hadn't been for AlphaGo? I think AlphaGo was that transition point where it became very, very clear that the moment of transition where we go beyond human level intelligence in particular areas is not science fiction or many decades later, it is happening now. and if it could happen in a game of go there was no reason why it couldn't happen in protein structure prediction in fusion in material science and the legacy of that match and move 37 and and that experience is what we are all living in now i think that's a great point to end the episode actually to be honest with you don't push me thank you so much for joining me Amazing.
52:42Pushmeet Kohli:Yeah, pleasure.
52:44Thore Graepel:These big paradigm shifting moments in the story of humans and machines have happened before. But the thing about chess is that it was always just a question of calculation. Can a machine brute force its way to a victory? AlphaGo was different. It was the first time that a machine had demonstrated something deeper, a genuine intelligence that combined intuition with calculation and took us beyond human capability. Now 10 years on from the AlphaGo match, the field has moved at an incredible pace. But many of the questions that preoccupied researchers then are more relevant now than ever. How do you create AI systems that go beyond human knowledge and are capable of new insights?
53:27Thore Graepel:And how do you separate the genuinely new insights from hallucinations? You have been listening to Google DeepMind, the podcast, with me, Hannah Fry. We have got plenty more episodes to come out this year, So please make sure you subscribe to our YouTube channel. We'll see you soon.
From the publisher
Seoul, March 2016. Two players sit hunched over a 19x19 grid covered in a sea of black and white stones. They are playing the ancient game of Go - a game of unimaginable complexity long thought impossible for a machine to master. On one side is Lee Sedol (Sae Dol), a legendary 18-time Go world champion. On the other, AlphaGo, a neural network based AI system built on a powerful technique called reinforcement learning. In the blink of an eye, the world changed. Exactly one decade later, we look back at the match that sparked the modern AI revolution. From algorithmic discovery to the solving of scientific grand challenges like protein folding, the foundation was laid right there on that wooden board.
Join Hannah Fry, Pushmeet Kohli (VP, Science) and Thore Graepel (AlphaGo team & Distinguished Research Scientist) as they unpick the legacy of AlphaGo.
🎥 AlphaGo https://youtu.be/WXuK6gekU1Y
🎥 The Thinking Game: https://youtu.be/d95J8yzvjbQ
Please leave us a review on Spotify or Apple Podcasts if you enjoyed this episode. We always want to hear from our audience whether that's in the form of feedback, new idea or a guest recommendation!
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.




