In short
Summary of Podcast Episode: "Is Human Data Enough? With David Silver"
Podcast Details
- Title: Google DeepMind: The Podcast
- Host: Professor Hannah Fry
- Guest: David Silver, VP of Reinforcement Learning
- Episode Description: David Silver discusses the future of AI, introducing the concept of the "era of experience" and contrasting it with the current reliance on human data. He explores the achievements of AI systems like AlphaGo and AlphaZero, emphasizing the importance of self-generated experience in driving progress towards artificial superintelligence.
---
Key Concepts
- Era of Experience vs. Era of Human Data
- Era of Human Data: Current AI methods largely depend on human data, where knowledge is extracted from extensive human-generated content.
- Era of Experience: David advocates for AI systems that interact with their environment, generating their own data through experience rather than relying solely on human knowledge.
- AlphaGo and AlphaZero
- AlphaGo: The first AI to defeat a world champion Go player, initially trained using human gameplay data.
- AlphaZero: Advanced AI that learned entirely from self-play without initial human input, discovering superior strategies through trial and error.
- Learning from Experience
- Reinforcement learning techniques allow AI to learn from the outcomes of its actions (win/loss) and improve over time.
- David highlights the importance of algorithms that prioritize self-learning over mimicking human behavior, which can limit potential.
- Move 37
- A pivotal moment from AlphaGo's game against Lee Sedol where a non-traditional move demonstrated AI's ability to think beyond human constraints, showcasing creativity in strategy.
- Human Data Limitations
- Relying mainly on human data constrains AI advancements due to the inherent ceiling of human knowledge.
- David asserts that breaking through these limitations requires methods that allow AI to discover and innovate independently.
- AlphaProof
- A new initiative where an AI learns to mathematically prove theorems without human-provided solutions.
- Emphasizes the capability of AI to independently verify and generate mathematical proofs, pushing beyond existing human accomplishments.
- Grounded Feedback
- David discusses the significance of grounded feedback in AI learning processes, advocating for systems that can evaluate their own actions based on real-world outcomes rather than human feedback alone.
- Future of AI and Mathematics
- The podcast explores the potential for AI to transform mathematics, suggesting that with ongoing improvements, AI systems could solve long-standing mathematical problems.
---
Key Takeaways
- Innovation through Self-Generated Experience: Moving towards an AI landscape where systems learn and adapt independently may yield breakthroughs that human-centric approaches cannot achieve.
- Redefining Success Metrics: The podcast prompts a reevaluation of how success is measured in AI, advocating for a diverse range of metrics beyond traditional human performance.
- AI's Role in Society: David envisions a future where AI contributes significantly to various fields by discovering new methodologies and pushing boundaries that human knowledge cannot reach.
---
Conclusion In this episode, David Silver articulates a compelling vision for AI's future, emphasizing the need for systems to transcend human data dependency. He suggests that embracing an "era of experience" could lead to unprecedented advances in artificial intelligence, ultimately enabling machines to achieve superhuman capabilities.
---
Additional Notes
- The conversation between David Silver and Fan Hui at the end of the episode provides personal insights into the impact of AlphaGo on the Go community, showcasing the ripple effect of AI advancements on traditional games and knowledge systems.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:03We're going to need our AIs to actually figure things out for themselves and to discover new things that humans don't know. And I think that's going to be a whole new era of AI that's going to be incredibly exciting and profound for society.
0:18Welcome back to Google DeepMind, the podcast. My guest today is the inimitable David Silver, an original deepminder, one of the key people behind the phenomenal success of AlphaGo, the first program to master the world's most complex board game and achieve superhuman performance. Now, at the end of today's podcast, we have a little extra treat for you, a conversation with David and Fan Hui, the first professional Go player to take on the AI. But now David has a bold idea about the direction that AI should go in next. After all of the current buzz and excitement and achievements of multimodal models, David has a plan for the path towards superhuman intelligence, a new phase which he calls the era of experience.
1:07This is a profound idea and not one without risks. David, welcome to the podcast. Hi, it's great to be here. Real pleasure. Thank you. Okay, so I have spent the weekend with a very enjoyable read of your position paper. And in it, you are talking about the era of experience. Summarise for us, what do you mean by that? Well, what I mean is that if you look at where AI has been for the last few years, it's been in what I call the era of human data, which is that all of these AI methods, they're based on one common idea, which is we extract every piece of knowledge that humans have and you kind of feed it into the machine.
1:48And that's one incredibly powerful way to do things. There's another way to do things, and this is what's going to lead us into the era of experience, which is where the machine actually interacts with the world itself, and it generates its own experience. It tries things out in the world, and it starts to build up its own experience. And if you think of that data as fueling the machine, then that will lead to this next generation of AI that we can think of as the era of experience. I guess this is in a way you're sort of thumping the table saying large language models are not the only AI, right?
2:25Like there are alternatives. There are different ways that we can approach this. That's right. I think we've really got a lot out in the field of AI of building large language models, harnessing the vast quantity of human, like particularly natural language data that's out there and kind of assimilating that all into a machine that knows everything that humans have ever written down. But at some point, we need to get past that. We want to go beyond that. We want to go beyond what humans know. And to do that, we're going to need a different type of method. And that type of method will require our AIs to actually figure things out for themselves and to discover new things that humans don't know.
3:01And I think that's going to be a whole new era of AI that's going to be incredibly exciting and profound for society. Well, OK, then let's talk about some other sort of famous AIs, famous algorithms that have employed different types of methods. most notably AlphaGo and AlphaZero, which of course notoriously beat the world's best Go players about a decade ago, right? Tell us about the techniques that we use in that and how they differ from large language models that we see today. So AlphaZero in particular is very different from the type of human data approaches that have been used recently because it literally uses no human data.
3:41That's the zero in AlphaZero. So there is literally zero human knowledge that's pre-programmed into this system. And so what's the alternative? How do you learn Go knowledge if you're not copying humans and you don't really know in advance the right way to play? Well, the way you go about it is through a form of trial and error learning where AlphaZero basically played itself millions of games of Go or chess or whatever the game is that it was wanting to play. and bit by bit it figured out oh if I play this move this kind of move in this kind of situation then I end up winning more games and then that's a piece of experience that is used to fuel it to become stronger and then it'll play a little bit more like that and the next time what will discover something new and say you know there'll be some new pattern it's like oh when I use this particular pattern I end up winning more games or losing more games and that feeds the next generation and so forth.
4:32And that learning from experience, this learning from the agent's self-generated experience is enough and was enough in AlphaZero to fuel its progress all the way from completely random behavior, all the way up to the strongest chess and Go playing programs that the world has ever known. They didn't start off just as like random empty boxes that kind of found how to play Go from nothing. I mean, initially when you were designing your Go algorithms, you'd worked out a way to encode Go games and then feed them in as a database, right? Yeah, that's right. So the original version of AlphaGo, the version which famously beat Lisa Dahl in 2016, this version of AlphaGo actually did use some human data to start it off.
5:14So we basically fed it a database of human professional moves, and it learned, it ingested those human moves, and that gave it its starting point. And then it learned for itself by experience from that point onwards. However, what we discovered a year later was that the human data wasn't necessary, that you could actually throw out the human moves altogether. And what we showed was that actually the resulting program, not only was it able to recover this level of performance, it actually worked better and was able to learn even faster than the original AlphaGo to achieve a much higher level of performance.
5:54That is such a strange idea. It's such a strange idea that you throw away the human data and you find not only was it not necessary, but it was actively limiting performance in a way. I think one of the hard lessons for people in AI, this is sometimes called the bitter lesson of AI, is that we really want to believe that all of the knowledge that we accumulated as humans is really important. We really want to believe that. And so we feed it into our systems. We build it into our algorithms. and what happens is that actually that makes us design the algorithms in a way which is maybe fitted to the human data and is less good at actually learning for itself and what happens is if you throw out the human data you actually spend more effort on learn how the system can learn for itself and that's the part which can then learn and learn and learn forever the bitter lesson i suppose in a way it's sort of saying accepting that it's possible that something could play Go better than humans can and sort of removing that ceiling in a way.
7:02That's right. That, you know, human data, it's really helpful to get you off the ground. But there is a ceiling to everything that humans have done. And, you know, we see that in Go. There was a maximum level of performance that humans have ever achieved. And we need to break through these ceilings. And in AlphaZero, we were able to break through that ceiling by building a system that learned for itself by self-play and got better and better and better until it blasted through that ceiling and went far beyond. And I think the idea of the year of experience is that we find the methods that allow us to break through that ceiling everywhere.
7:33We build AI systems that become superhuman in all of the capacities that humans seem so amazing, but we find the way to go beyond that. Let me just stick with Go for a second, okay, before we get onto the other ways that you can get rid of human data and thus improve on human ability. because it sort of sounds a little bit, when you say let's just get rid of all of the games of Go that humans have played and start with nothing, it sort of sounds like a magic trick. Just tell me a little bit about the techniques that you're really using there in order to get a machine to, as you say, chain together thousands and thousands of different ideas in order to be amazing at the game of Go.
8:17Well, the main idea is an approach that we call reinforcement learning. And the idea of reinforcement learning is you basically give the outcome of your game a number. And we say, you know, plus one if you win and minus one if you lose. One point. Exactly, exactly. And what we do then with reinforcement learning is we get the system to basically, we give it a reward each time it does something right. And we train the system to basically reinforce, that means do more of the things that gets more reward. And so in terms of, for example, if you've got a neural network, like we do in AlphaGo that's picking the moves, what you want to do is tweak the weights of your neural network a little bit in the direction that gives you more reward.
9:02And that's the main idea of reinforcement learning. But then, OK, I mean, a game of Go is quite long. How do you make it so that you do the right moves in the beginning so that you end up with the right outcome towards the end? How do you work out which bits of the game are important, I suppose? So this is a really important problem. It's called the credit assignment problem. And the idea is that if you've, yeah, you're absolutely right that you could have had, you know, 100 or 200 or 300 different moves. And then at the end, you just get one bit of information saying, you know, win or loss. And you somehow have to work out which of the moves in the game were responsible for winning and which of the moves in the game were responsible for losing.
9:44And there's lots of ways to do that. The simplest way is just to assume that everything that you've done contributes a little bit to that outcome at the end. And it sort of all comes out in the wash. One of the biggest moments of the AlphaGo story was Move 37 that everyone always references. Just tell me about that. So Move 37 was a move that happened in the second game of AlphaGo against Lee Sedol. and AlphaGo played a move that defied everyone's expectations. The traditional idea in the game of Go is that you play your moves typically on the third line or the fourth line of the board because this either gives you territory for the third line or influence on the fourth line and you never go below or above that.
10:29It just wouldn't make sense to humans. AlphaGo played on the fifth line and it somehow played this in a way that just made everything make sense in the board. It kind of connected everything together with this move on the fifth line. And it was so alien to humans that we estimated that only one in 10 ,000 probability that a human would ever think of playing this move. Humans were shocked by this move. And yet it helped win the game. And so it was a moment where humans said, look, here's something creative that happened. Something that a machine came up with that was different from the way humans traditionally thought about the game.
11:06that actually was a big piece of progress and took us beyond the kind of confines of human knowledge. And I guess if we really do want to advance AI, we sort of do want those alien ideas, as you put it. Do you think you've seen an equivalent of Move 37 with large language models? Move 37, in some ways, was special because it was the first moment. It was the first time that people had seen a big breakthrough like this. and because we've been in the era of human data we focused a huge amount on reproducing human capabilities and we've focused much less on going beyond them and I think until we really emphasize systems learning for themselves to go beyond human data we won't see huge breakthroughs the equivalent of Move 37 in the real world.
11:59It just, it seems unlikely to me. Because when you're angered in human data, you're only ever going to have human-like responses. That's right. And I think there are things you can do that allow you to maybe do things in the middle a little bit. So if you push me to say what's the greatest Move 37-like moment, I would probably pick out some work by scientists at MIT who discovered a new antibiotic that no human knew about. And I think that's an incredible discovery of massive importance to humanity. So in that sense, it goes way beyond Move 37. But what I like about Move 37 is that it's not just a single discovery.
12:38It's one of an infinite series of discoveries where the system can just keep on learning and learning and learning. And Move 37 is important to me because it represents just a single point in that infinite sequence of discoveries that can happen once you've got this kind of approach of learning from experience. Rather than actual result in and of its own right. Yeah, that's right. Give me a brief rundown of how AlphaZero worked. So AlphaZero is surprisingly simple. I mean, there's some very complicated algorithms out there in the world, but this one is really straightforward. So all you do is you start with a policy, a way to pick moves, and a value function, which is a way to evaluate moves and say whether they're good or bad.
13:22So you start with that, you run a search. And then what you do is you take the best move according to your search, and you train your policy to do more things like that, to do more of the good moves according to your search. And you train your value function based on how the game actually panned out when you played a game with this search. And that's it. You just iterate that millions of times and out pops a superhuman game player. It's like magic, basically. It does sometimes feel like magic. I remember the first time that really felt like magic to me was when we had just completed AlphaZero on chess.
13:59Someone had the idea of trying it on a different game. So we plugged it into a game that none of us could play, a game called Shogi, which is Japanese chess. And we had no idea how to play this game. What, you didn't even know the rules? So the system knew the rules. The agent, we taught it the rules. But none of us had the first clue of how to really, you know, strategy or tactics. It would have been like blunder after blunder if we'd been playing this game. And we just plugged it in. And it was literally the first ever time we ran AlphaZero on Shogi. We had no idea whether it was good or not.
14:33We couldn't evaluate it. But we sent it off to Demis, who's actually a reasonably strong player, of course. And he said, hmm, this looks quite good. I'm sending it to the world champion. And the world champion said, hmm, I think this is superhuman. And so it literally felt like magic because, you know, we just pressed go on this system and had no idea of the process and how it got there. But somehow out popped a superhuman shogi player. Can AI design its own reinforcement learning algorithms? Well, funnily enough, we have actually done some work in this area. It's work we actually did a few years ago, but it's coming out now.
15:12And what we did was actually to build a system that through trial and error, through reinforcement learning itself, figured out what algorithm was best at reinforcement learning. It literally went one level meta and it learned how to build its own reinforcement learning system. And incredibly, that reinforcement learning system actually outperformed all of the human reinforcement learning algorithms that we'd come up with ourselves over many, many years in the past. I mean, this is the same story over and over again. The more of a human you put into something, the worse it acts, the worse it performs.
15:51Take the human out, does better. Okay, if AlphaGo and AlphaZero then are really exceptional examples of reinforcement learning used to the best it can be, you still find reinforcement learning in the large language models that we have at the moment, right? Tell me about how they're integrated into these systems. So reinforcement learning is used in almost all large language model systems. And the main way it's used is by combining it with human data. So unlike the alpha zero approach, this means that the reinforcement learning is actually trained on human preferences. So the system is basically asked to produce outputs.
16:33And then a human says, this one is better than this other one. and the system becomes more like the one that the human prefers. And this is called reinforcement learning from human feedback, and it's been massively important in LLMs, and it's helped transform them from systems that just blindly mimic any kind of data that you see on the internet into systems that actually usefully produce answers to the kind of questions that people really want to see. And so it's an incredible advance. However, it feels like we've thrown out the baby with the bathwater. These reinforcement learning from human feedback systems or RLHF, they're very powerful, but they do not have the ability to go beyond human knowledge.
17:21Like if a human rater doesn't recognise some new idea and underappreciates that there is some series of actions that would actually end up being far better than some other series of actions, there is no way that the system will ever learn to find that sequence because the raters might not understand that better behaviour. That human feedback element, though, it does seem to give these models some sense of grounding. Like, I know the last time we spoke, grounding was like this really big topic, this idea that you want these algorithms to have a conceptual understanding almost of the world that we're living in.
18:00So if you take away or you remove that human feedback aspect, do you still end up with models that are grounded? I almost want to argue the opposite. I want to say that when we train a system from human feedback, that it is not grounded. And the reason is that we are basically, the way RLHF systems normally work is the system presents its response, its answer to a question, for example. And a rater says that's good or bad before the system actually does anything with that information. So it's like the human is prejudging the output of the system. So, for example, if you're asking for a cake recipe from an LLM, the human rater will look at the recipe that's output by the system and judge whether that recipe is good or bad before anyone has actually made the recipe and eaten the cake.
18:55And in that sense, it's ungrounded. Like a grounded outcome would be someone actually eats the cake and the cake is either delicious or disgusting. and then you've got grounded feedback that says you know this cake really was a good cake or this cake was a bad cake and it's that grounded feedback that allows the system to iterate and discover new things because it can try out new recipes that maybe you know expert chefs presume will be disgusting but actually turn out to be delicious yeah like a yeah monster munch muffin or whatever exactly yeah that's the most delicious food that ever existed okay that's interesting though because I have heard I mean even a conversation with Demis talking about how grounding gets into these models how they kind of have built this conceptual understanding of things and it sounds almost like what you're saying is that the grounding that they have is like a sort of superficial level of grounding maybe?
19:50I think human data is grounded in human experience so it's like the LLMs are sort of inheriting all of that information that humans may be figured out from their own experiments, for example, in science, a human might have tried to walk across water and discovered that they fell in, and then they might have created a boat and discovered that that floated. And all of that information can be inherited somewhat by the LLM. But if we want a system that actually makes discoveries and discovers some completely new form of propulsion across water or some completely new mathematical idea or some completely new way to...
20:28Part the seas. Yeah, new medicine and new approach to biology. The data just isn't there and the system needs to figure out for itself through its own kind of experimentation, its own trial and error and its own grounded feedback, whether that's a good idea or a bad idea. I got to talk to Aurel, Aurel Vinales, who really spoke about how we are running out of human data and that we are going to need to start creating synthetic data in order to fill that gap. I mean, this is related, right, to that idea. It's just rather than using LLMs to create more human dialogue data, you're going about the solution in a different way.
21:08That's right. So synthetic data can mean a lot of things. But, you know, normally it would mean that you've got some process where you kind of take your existing LLM and use it to generate some set of data. And I guess the argument is, similar to the ceiling that we have from human data, that however good that synthetic data is, they will reach a point where that synthetic data is no longer useful to the system becoming stronger. So the beauty of a self-learning system, where the fuel of the system is actually experienced, is that as the system starts to get stronger, it starts to encounter problems that are exactly appropriate to the level it's at.
21:47so it will always be generating experience that allows it to solve the next problem that it's it's encountering and so it can just get stronger and stronger and stronger forever there is no limit and that i think is what differentiates this particular approach of using self-generated experience from other forms of synthetic data just returning to your cake example though i mean if you kind of follow that through somebody eats the cake and says yes this was delicious You're using the human feedback then at the end of the process anyway. Are we talking about that or are we talking about maybe having systems that are completely untethered from humans and are embodied or in the physical world somehow so that they can get their feedback in that way?
22:28Look, I think the ideal is that like AlphaZero, we have systems which are able to generate vast volumes of self-generated data experience that they can then verify for themselves. And in many domains, that's going to be possible. And in many domains, it's not going to be possible. In the ones where it's not possible, we have to acknowledge that humans are a big part of the environment that we're in. We have to acknowledge that they're a part of the world that we want our agents to live in. And so it seems reasonable to think of humans as a part of that environment and to think of the way that they behave as part of the observation that the agent receives.
23:01I think the thing which I'm pushing back against and saying is not grounded is not that. it's the fact that the rewards that the agent learns from is coming from a human's judgment of like whether this sequence of actions is good or bad and the system is not judging for itself based on the consequence of those actions in the actual world and so you know one way to say it is that we shouldn't make you know human data a privileged part of the agent's experience it's just just just observations in the world and we should be able to learn from that like any other data. If we go back to that AlphaGo example earlier of assigning that reward, that one point that it gets at the end, is this almost like the way that we're handling AI at the moment is that the algorithm does its first 10 moves or 15 moves and then we insert a human in who says yes that's a good first 10 moves and doesn't allow the whole process to kind of execute fully before you input that little bit of feedback.
24:00That's exactly right. So imagine that we were training AlphaGo and after every single move, our best Go player comes in and says, oh, that move was amazing. Oh, no, no, that move was totally wrong. And then we get that feedback and we put it in and the system learns to pick the move that the human prefers. It would not end up discovering Move 37 because it would just end up playing like the human thinks is a good game of Go. And it would never discover the new ways to play Go that that human didn't know about. Okay, so I think the environment of Go, what you're saying makes a lot of sense in that environment.
24:37There are other environments too, where I think that this makes a lot of sense. I'm thinking here about the pinnacle of human thought of mathematics. Tell me what's been going on in that space. Like you say, it is an incredible human endeavour that's had millennia of human effort going into it. And so in many ways, it does represent like literally the limits of achievement by the human mind. And so naturally, we turn to it for AI to see, can we achieve those same levels of performance that humans have achieved over all of those years of endeavor? We recently put together what I think is a very exciting piece of work called Alpha Proof.
25:13It is a system that learns through experience how to correctly prove mathematical problems. So if you give it a theorem and you don't tell it anything about how to actually prove that theorem, it will go away and figure out for itself a perfect proof of that theorem and we can actually verify and guarantee that this proof is correct. One thing which is interesting about this is that it's the exact opposite of how LLMs normally work because if you ask LLMs to prove a mathematical problem at the moment, they will normally output some informal mathematics and say, just trust me this is this is correct and it might be correct but it might not be because we know that llms tend to hallucinate a lot they can they can make things up and the nice thing about alpha proof is that it will actually guaranteed produce the truth so let's think of an example here to kind of anchor this in people's minds let's say that uh prime numbers are something that can't be divided by anything but themselves and one and there are infinite number of them off you go prove it?
26:23Yeah. So the way alpha proof works is it's trained on millions of different examples of theorems, not just one. And what happens is it goes off and it trains on them. And to begin with, it can't solve the vast majority of them. 99.999 % of the theorems it just can't do. And these are theorems that humans have already proved. Are you feeding in? We feed into the system something like a million different theorems that humans have come up with themselves, but we don't provide the human proofs. We just provide the questions, but not the answers. So you're giving it stuff that you know is true, but you're just not telling it how to prove it.
27:02And sometimes we don't even know it's true, because what we actually do is we take the human theorem, the human question, and we actually turn it into a formal language. These aren't using language in the sense that language models are using, but they are using a form of language, like a mathematical language. That's right. So in fact, we do use a small large language model. And that large language model allows us to output programming languages. And in particular, we use a programming language that's called Lean that allows all of mathematics to be expressed. And so it's an amazing idea that mathematicians have come up with that you can actually formalize all of these kind of things that we normally talk about in English language or whatever language you happen to be speaking can be transformed into a perfectly clear, verifiable mathematical language that allows all of the ideas of maths to be expressed and also all of the ideas of mathematical proof to be expressed.
28:04So you can say, for example, that if A implies B and B implies C, then there's a way to go from that to A implies C. And that's the kind of thing that you can do in this mathematical programming language. You essentially write a program that takes you from one to the other and you have a proof of this statement. So we take our kind of million human problems and from that we generate 100 million formal problems. and some of those might actually not be possible or they might be incorrectly formulated or they might just be false. And it doesn't matter because all we do is we learn to prove those things and the ones which we can't prove become, we keep trying and keep trying.
Read the full transcript
28:49The ones that we already prove, okay, they're done, they're out the way now. If we disprove them, that's fine, they're out the way. And we're left with the really interesting ones, which are the ones which are really hard to prove. And we keep kind of climbing up from just being able to solve one or two of them to then being able to solve 10 or 20 of them and eventually being able to solve a million of them. Is this the equivalent then, that moment of the proof is correct or incorrect, is that equivalent to AlphaGo, you win the game or you don't? It's exactly equivalent. So if we use the idea that Lean says, well done, you've proved this as a reward, and we give the system, you know, plus one if it solves it and minus one if it doesn't get that correct.
29:29And so this allows us to then train a system by reinforcement learning to get better and better at proving mathematical statements. In fact, we literally used the same alpha zero codes that we used to get better at Go and chess and all of these other games. It's literally the same code, but it's running, if you like, with the game of mathematics. The game, how dare you?
29:54Don't dare trivialise my subject. I'm joking. Okay, how good is it? It's not yet a superhuman mathematician. although that is where we'd like to get to one day. But one thing which AlphaProof did achieve was the most well-known and challenging of mathematical competitions is called the International Mathematics Olympiad. And this is a competition that happens once a year for the most incredible and amazing young mathematicians from all around the world. And the problems, to say the least, are extremely hard. They're spicy. They're very spicy. As a professor of maths, sometimes, I mean, they're spicy.
30:36So you heard it from Hannah. These are hard problems. Very hard. And AlphaProof, amazingly, it actually achieved a silver medal level of performance in this competition. So this is a level of performance that only roughly 10 % of the contestants would actually be able to achieve. In the entire world. In the entire world. This is like the cream of young mathematicians, like the six best from every country. And not only that, but there was one particular question. that less than 1 % of all the contestants were able to solve. And AlphaProof got a perfect proof for this particular problem. So that was nice to see.
31:12What do the proofs look like? I mean, do they follow human-style arguments if you're not inputting any human data into them? I have to say that, to me, the proofs, I don't understand them at all. But Tim Gowers, I mean, the Fields medalist and former IMO, So, I mean, did he get, was he a gold medalist? Tim Gowers was a, yeah, I think multiple gold medalist at the IMO. I mean, mega brain, right? Like extraordinary mathematician. But I mean, he understands these proofs, right? So Tim Gowers actually was the, refereed our solutions to make sure that they were valid solutions and that we hadn't, you know, broken any of the rules.
31:52And he understands the solutions and thought that they were, you know, a huge leap beyond anything that previous AI mathematics could do before. So it's a jump forward, but it's still just the beginning in the sense that, you know, we really want to go beyond human mathematicians. And that's where we'd like to go next. Because at the moment, basically, you've got yourself a very, very, very talented 17-year-old mathematician, basically, right? That's right. And it should be said that the system that entered the IMO did take longer than a human contestant would be allowed to take. So, you know, that's something we're just going to assume will get better over time as machines get faster.
32:30I mean, the IMO is like the perfect test bed because there are correct answers. It can be judged. You can compare it to human performance, all of that kind of thing. But if you are feeding in conjectures, so things that we don't even know are true, you know, I'm thinking of like the ABC conjecture here or the Riemann hypothesis or any of those like really grand unsolved challenges in mathematics. If alpha proof outputs something and says, no, no, no, we've checked this proof. it works. Can you trust it? And maybe even beyond that, is it worth anything if we don't understand it? I think the good news about lean is that mathematicians who are better than myself are always able to take a lean proof and translate it back into something that humans can understand.
33:17And in fact, we've even built an AI system that can do this, which can take any formal proof and what we call informalise it, which means it will turn it back into something which is very understandable to humans. And if we did solve the Riemann hypothesis, and by the way, we're a long way from doing that. But if it was done, there'll be millions of mathematicians who would be very excited to understand whatever new mathematics came out of it and decode it back into things that humans can understand. Okay, but here's my question, right? There's the Clay Maths Institute in the year 2000, offered a million dollar prize for seven different mathematical problems.
33:51And, you know, human mathematicians have had a quarter of a century in order to try and solve them and only one has fallen. Do you think potentially the next one could go to AI? Yes, I do actually. I think that it might take time. I don't think we're there yet. I think there's a long way before AI systems are capable of doing this. But I think AI is on the right track and systems like AlphaProof will become stronger and stronger and stronger. You know, what we saw in the IMO is just the beginning. And you know that once you have a system that can scale and can keep learning and learning and learning, really the sky's the limit.
34:26So what will these systems look like in two years or five years or 20 years? Well, I personally would be amazed if AI mathematicians don't transform the whole of mathematics. I think it's coming. Mathematics is one of the few areas where, in principle, everything can be done completely digitally by a machine interacting with itself and just going and going and going. And so there's really no fundamental barrier to an experience-driven AI system mastering mathematics. OK, I really buy what you're saying about alpha proof, by the way, and the same with alpha zero. I mean, I think they're really excellent examples of how far you can go with reinforcement learning.
35:08But they are also examples where there is a very clear metric of success. You win a game of go or you don't, your proof is correct or it isn't. How did these ideas translate to systems where it's a lot messier and actually these very clear metrics might not necessarily be present? So first, I want to acknowledge that this question is probably the reason why reinforcement learning methods or these kind of experience based methods I'm talking about have not yet broken into the mainstream of absolutely everything that we do in every AI system. so it has to be cracked. If the era of experience is to come about then we have to have an answer to this.
35:46But I think the answer might be right in front of us because actually when you look at it the real world contains innumerable signals. There's just a vast number of signals in the way that the world works. You know if we look at all of the things that we do on the internet for example there's any number of signals like likes or dislikes or profits or losses or pleasure pain signals you might get or yields or properties of materials. There's all these different numbers representing different things about different aspects of experience. And so what we need is really a way to build a system which can adapt and which can say, well, which one of these is really the important thing to optimize in this situation?
36:30And so another way to say that is, wouldn't it be great if we could have systems where, you know, a human maybe specifies what they want. But that gets translated into a set of different numbers that the system can then optimize for itself completely autonomously. So, okay, an example then, let's say I said, okay, I want to be healthier this year. And that's kind of a bit nebulous, a bit fuzzy. But what you're saying here is that that could be translated into a series of metrics like resting heart rate or, you know, BMI or whatever it might be. And a combination of those metrics could then be used as a reward for reinforcement learning.
37:08Have I understood that correctly? Absolutely correctly. Okay. Are we talking about one metric, though? Are we talking about a combination here? So I think the general idea would be that you've got one thing which the human wants, like to optimise for my health. And then the system can learn for itself, like, which rewards help you to be healthier. And so that can be like a combination of numbers that adapts over time. So it could be that it starts off saying, OK, well, right now it's your resting heart rate that really matters. And then later it gets some feedback saying, hang on, I really don't just care about that.
37:44I care about my anxiety level or something. And then it includes that into the mixture. And based on feedback, it could actually adapt. So one way to say this is that a very small amount of human data can allow the system to generate goals for itself that enable a vast amount of learning from experience. Because this is where the real questions of alignment come in, right? I mean, if you said, for instance, let's do a reinforcement learning algorithm that just minimises my resting heart rate, I mean, quite quickly, zero is like a good minimisation strategy there, which would achieve its objective, just not maybe quite in the way that you wanted it to.
38:26I mean, obviously, you really want to avoid that kind of scenario. So how do you, how do you have confidence that the metrics that you're choosing aren't creating additional problems? You know, one way you can do this is to leverage the same answer which has been so effective so far elsewhere in AI, which is at that level, you can make use of some human input. If it's a human goal that we're optimizing, then we probably at that level need to measure, you know, and say, well, you know, a human gives feedback to say, actually, you know, I'm starting to feel uncomfortable. And in fact, while I don't want to claim that we have the answers, and I think there's an enormous amount of research to get this right and make sure that this kind of thing is safe, it could actually help in certain ways in terms of this kind of safety and adaptation.
39:14There's this famous example of paving over the whole world with paperclips when a system's been asked to make as many paperclips as possible. But if you have a system which is really, its overall goal is to support human well-being and it gets that feedback from humans about, and it understands their distress signals and their happiness signals and so forth, the moment it starts to create too many paperclips and starts to cause people distress, it would adapt that combination and it would choose a different combination and start to optimize for something which isn't going to pave over the world with paperclips.
39:49So look, we're not there yet. But I think there are some versions of this, which could actually end up not only addressing some of the alignment issues that have been faced by previous approaches to goal-focused systems, but maybe even be more adaptive and therefore safer than what we have today. Outside of the world of AI, though, I mean, is there a problem with using quantitative metrics as a measure for success at all? I mean, I'm thinking here about exam scores or GDP or the myriad of problems that you can get into when you focus too carefully and end up with a tyranny of metrics. So look, I would be the first to agree that when you mindlessly pursue a metric in the human world, that it often leads to undesired consequences.
40:38At the same time, the whole world of human endeavor is organized around us optimizing for some things. You know, if we didn't have anything that we could optimize for, we wouldn't ever be able to make progress. You know, we have all kinds of signals and metrics and so forth that drive progress. And then people say, oh, OK, maybe that isn't the right metric and they adapt it. Is part of the problem then that at the moment you have an interaction with an AI that is really contained within time? There aren't these sort of longer term learnings or adjustment of what the goals might be. Like once you decide that GDP is the thing that you're going for, it's GDP forever and there's no change.
41:15I think that's absolutely right. That, you know, the kind of AI that we have today doesn't have like a life. You know, it's not something which has its own stream of experience in the way that, you know, an animal or a human might have that kind of goes on for years and years and years and can keep adapting over time. And that needs to change. And one of the reasons it needs to change is so that we can have systems that just keep learning and learning and learning over time and adapting and understanding how to better achieve the kinds of outcome that we really want. Is there something that is quite risky about untethering algorithms with potentially quite a lot of power from human data, really?
41:54There are certainly risks and there are certainly benefits. And I think we absolutely have to take this very seriously and be extraordinarily careful about taking these steps that come next in this journey towards the era of experience. And I should say that, you know, one of my reasons to write this position paper is because I feel that people aren't recognizing that this transition is going to come and that it will have consequences and it will require careful thought about many of these decisions. And the fact that so many people are still thinking only about the human data approach means that not enough people are taking seriously these kinds of questions.
42:34The last time I got to speak to you on this podcast, we talked about a different position paper that you had just written, Reward is Enough, essentially saying that reinforcement learning is all you need to get you towards AGI. Do you still think that that's the case? I think the way I would answer this is by saying that human data might give us a head start. It's a bit like, to borrow a metaphor, it's a bit like the fossil fuels that we discovered in the earth. And, you know, all of this human data just happens to be there. And then we kind of mine it and burn it in our LLMs. And that gives them, you know, a certain level of performance that they have for free.
43:15but then we need in the analogy some kind of sustainable fuel that keeps the world going once all the fossil fuels are gone and i think that's what reinforcement learning is it's the sustainable fuel this experience that it can keep like generating and using and learning from and generating more and learning from it that's really the process that's going to drive progress in ai and i don't want to in any way denigrate what's been done with human data i think it's great i think you know the ais that we've got now are amazing mind-blowing things and you know i love them and enjoy working with them and do research on them myself.
43:48But it's just the beginning. Dave, thank you so much. That was amazing. Thank you.
44:00Of course, there is this monumental amount of progress that's going on at the moment. But when you stop to think about it, there really has been this narrowing in the diversity of ideas around AI. I mean, the success of multimodal models has been so rapid, it's been so profound, so beyond what most people were expecting, that they kind of have sucked a lot of the oxygen out of the broader conversation. And it is noticeable that we're hearing again and again now, these murmurs that we have reached the limit of usable human data. And okay, of course, There are risks involved with this approach of untethering AI from human data, all sorts of areas that need careful thought and attention.
44:42But I can't help but be quite convinced by what David was saying there. If we really want superhuman intelligence, maybe it is now time to step away from the human.
44:57You have been listening to Google Deep Mind, the podcast with me, Professor Hannah Fry. And before you go, we have got an extra special treat for you today in the form of a conversation between David Silva, the man behind AlphaGo, and Fan Hui, the first professional Go player to face it. How are you, Dave? I'm really well. Good to hear from you. It's been a long time. A long time, no see. A decade ago, a little while before the very famous 4-1 victory over Lisa Dahl, Fan Hui became the first professional Go player to test his skills against your algorithm. How long has it been since you spoke to him?
45:35It's been quite a few years. It's so nice to see Fan Hui. It's been absolutely amazing to catch up. Fan Hui played such a huge part in the development of AlphaGo. So it's really just a genuine delight. Thank you so much for joining us, Fan Hui. Oh, thank you. Thank you. For me, it's a very extraordinary experience. OK, so I want to ask you about that match that you had all those years ago. Because I think, I guess now, looking at the full history of it, it almost seems like a foregone conclusion. But at the time, I mean, you must have been pretty nervous, David. And how did you feel about it as well, Fenway?
46:09I remember the first time I saw the Demis email tell me, like, it's an exciting goal project. I still remember when I played with AlphaGo, first game I lost, I feel something strange. I also remember when I lost the second game, I feel fear because I feel maybe I will never win with this program or AI. And when I lost my five game, last game, I feel my old goal world is totally broken. But my new goal world is open. David, I want to ask you as well, though, in advance of that match, how confident were you about the performance of your algorithm? We really weren't confident. It was just so hard to judge where we were because we knew that we'd gone beyond the players that we had at DeepMind and we knew that we'd gone beyond all the programmes that had been written before.
47:04But there's such a huge gap beyond that towards the level of professional players like Fandwe. And we had no idea, you know, are we somewhere in that gap? Are we somewhere beyond that gap? like we just genuinely didn't know and so um this was like the first time we had any opportunity to calibrate our level of performance and i don't think any of us would have been surprised if we'd lost all five games so it was a very pleasant surprise to win all five and yeah we just i genuinely it was like one of those moments where the world could have branched either way and we just didn't know until until the match happened but of course this algorithm then advanced i mean with your help in fact after your match you you came on board and supported the team in in developing it further but that earlier version what did it feel like to play it did it feel fundamentally different to having a human opponent you know i play with another program before alpha go when i play with another program i feel like oh this is a program because they don't play like a human but with alpha go i feel something very strange sometimes i feel like it's really really like human what's the impact been then of alpha go and alpha zero on the go community as there had to be a process of acceptance or was it you know positive from the off first of all when i lost with alpha go so for the all go community nobody really not believe this is true because yeah you know i'm only European champion.
48:29So it's not world champion like Lissedar. But when AlphaGo went with Lissedar and all go community see something different because AlphaGo play really, really well. I remember the second game, the move 37, such beautiful move, really, really beautiful. So creative, it's very creative. For the human, we will never play this move. After that move, everything changed in the Go world. Because for us, everything is possible. Today, even the Go students use AI to learn. So, yes, I think this is really, really good for our Go community. I think it's not just for Go community. It's also for the world, I think.
49:19Fan Hui, thank you so much for joining us. That was such a real treat, especially with the big anniversary coming up. Just great to see you again. and thanks for everything you did on AlphaGo. I don't think it would have been the same without you. I think we would have made some terrible mistakes if we hadn't had your advice to help us along the way. So thank you. Thank you, Dave.
From the publisher
In this episode of Google DeepMind: The Podcast, VP of Reinforcement Learning, David Silver, describes his vision for the future of AI, exploring the concept of the "era of experience" versus the current "era of human data". Using AlphaGo and AlphaZero as examples, he highlights how these systems surpassed human capabilities by engaging in reinforcement learning without prior human knowledge. This approach contrasts with large language models, which depend on human data and feedback. Silver emphasizes the need to explore this path to drive AI progress and achieve artificial superintelligence.
Timestamps
- 00:00 Introduction
- 01:50 Era of experience
- 03:45 AlphaZero
- 10:19 Move 37
- 15:20 Reinforcement learning and human feedback
- 24:30 AlphaProof
- 29:50 Math Olympiads
- 35:00 Experience based methods
- 42:56 Hannah's reflections
- 44:00 Fan Hui joins
___
Thanks to everyone who made this possible, including but not limited to:
- Presenter: Professor Hannah Fry
- Series Producer: Dan Hardoon
- Series Editor: Rami Tzabar
- Commissioner & Producer: Emma Yousif
- Music Composition: Eleni Shaw
- Audio Engineer: Richard Courtice
- Production Manager: Dan Lazard
- Video Director and Editor: Bernardo Resende
- Video Studio Production: Nicholas Duke
- Video Editor: Bilal Merhi
- Audio Engineer: Perry Rogantin
- Camera and Lighting Operator: Robert Messere
- Production Coordination: Zoey Roberts, Sarah Ellen Morton
- Visual Identity and Design: Rob Ashley
- Commissioned by Google DeepMind
Please leave us a review on Spotify or Apple Podcasts if you enjoyed this episode. We always want to hear from our audience whether that's in the form of feedback, new idea or a guest recommendation!
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.



