Is Artificial Superintelligence Imminent? with Tim Rocktäschel - #706

21 Oct 2024 · 56 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The TWIML AI Podcast - Episode #706: Is Artificial Superintelligence Imminent? with Tim Rocktäschel

Podcast Overview Title: The TWIML AI Podcast Host: Sam Charrington Guest: Tim Rocktäschel, Senior Staff Research Scientist at Google DeepMind, Professor at University College London, and author of *Artificial Intelligence: 10 Things You Should Know*

Episode Summary In this episode, host Sam Charrington engages with Tim Rocktäschel to explore the concept of Artificial Superintelligence (ASI), examining its attainability, the importance of open-endedness in AI development, and recent research projects spearheaded by Rocktäschel and his team.

Key Topics Discussed

  • Advancements in AI: A reflection on AI developments since Rocktäschel's previous appearance on the podcast, particularly regarding the NetHack benchmark and the evolution of large language models (LLMs).
  • Artificial Superintelligence (ASI):
  • Rocktäschel asserts that ASI is achievable, defining it as an AI capable of surpassing human performance across diverse tasks.
  • Distinction between Narrow AI (current capabilities) and General AI (AGI), which is necessary for achieving ASI.
  • Importance of Open-Endedness:
  • Developing autonomous systems capable of self-improvement through empirical evidence.
  • The role of evolutionary approaches and algorithms in this context.
  • Recent Research Projects:
  • Promptbreeder: A tool for optimizing prompt strategies in LLMs.
  • Debating with Persuasive LLMs: Investigating how debates between LLMs can lead to more truthful answers, potentially useful for combating misinformation.

Key Concepts and Arguments

  • General AI vs. Narrow AI:
  • Narrow systems can already outperform humans in specific tasks (e.g., chess, Go).
  • General systems (e.g., LLMs) are emerging and show signs of advanced capabilities, raising the question of their path toward ASI.
  • Mechanisms for Improvement:
  • The transition from narrow AI to AGI involves not just training on specific tasks but exploring broader domains and leveraging vast datasets (internet-scale text and images).
  • Rocktäschel suggests that increasing the model's context window allows LLMs to leverage more extensive information, enhancing performance without changing core parameters.
  • Evolutionary Approaches:
  • Emphasizing the need for systems that autonomously explore, discover, and improve through an evolutionary-like process rather than traditional supervised learning.
  • The integration of LLMs in evolutionary systems allows for generating variations and evaluating them against empirical evidence.
  • Research Implications:
  • The potential for AI to aid in scientific discovery, as demonstrated in his research.
  • Encouraging AI systems not to only improve outputs but also to evaluate and enhance their own methodologies.

Discussion Highlights

  • Self-Improvement Systems:
  • The combination of powerful foundation models, mutation and selection operators, and empirical evidence collection is crucial for creating self-improving AI systems.
  • AI as a Truth-Seeking Process:
  • The effectiveness of AI in debates demonstrates its potential to identify correct information amidst misinformation.
  • Future Directions:
  • Forecasting the evolution of AI over the next 5 to 20 years, with a belief that once a capable AGI is achieved, reaching ASI could happen relatively quickly.

Conclusion Tim Rocktäschel's insights provide a comprehensive overview of the current landscape of AI, articulating a vision for the future where ASI is not just a theoretical possibility but an imminent reality. The discussion emphasizes the importance of evolutionary mechanisms, empirical feedback, and the necessity of open-ended exploration in the pursuit of advanced artificial intelligence.

---

For complete show notes, visit [The TWIML AI Podcast Episode #706](https://twimlai.com/go/706).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Really now, I think all these things are in place. We have powerful foundation models. We have really powerful mutation operators. We have powerful selection operators. We have models that can code quite well. And putting this all together means we'll have systems, as I said, that self-improve based on empirical evidence that they collect in a number of hard domains. And I think that's really what's going to drive further improvements in AI in the next few years.

0:42All right, everyone, welcome to another episode of the Twinwell AI podcast. I'm your host, Sam Charrington. Today, I'm joined by Tim Rocktushel. Tim is Senior Staff Research Scientist at Google DeepMind, Professor of Artificial Intelligence at University College London and author of the recently published Artificial Intelligence, 10 Things You Should Know. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Tim, welcome back to the podcast. Thanks. Great to be back. It's great to chat. We published that episode together a week ago, three years and a week ago, actually it was October 10th of uh 21 and we were talking about uh advancing deep reinforcement learning with nethack that was the episode title talking about that project uh what do you think has AI changed a bit since then I think not much happened I mean I mean to be to be absolutely honest um like there's not that much progress on that nethack benchmark right we have seen a lot of amazing progress in AI in terms of large language models, foundation models in general, but NetHack is still unsolved.

1:57So that grand challenge is still there. Yeah. Do you think LLMs are applicable there? Yeah, you can try them out. I think we'll have something to show in maybe a few weeks on that front. Oh, really? I think, yeah, I can already... I mean, you can try it yourself already. I think it's fair to say already like lms are not going to crack it right it's it's still a very hard hard exploration problem and even systems like o1 i think yeah won't won't get us all the way there or not even close let's put this way yeah we'll refer folks back to uh that episode to learn more about net hack and that challenge uh today i think we're going to focus a little bit on the book and more broadly your research into open-endedness and some other interesting questions raised by the book.

2:50Tell us a little bit about who the book is for. Yeah, so this is a popular science book. So this is new to me as well, right? Trying to write for a wider public audience. And to me, this was an exciting project because the book itself is part of a series, 10 Things You Should know. There are books about space and time and numbers and dinosaurs and trees. And now there's one on artificial intelligence. And I think it shows you a little bit what, I guess, important topic AI has become also for the public, right? And what I found intriguing about writing such book is that, I mean, on one hand, I think it's important to write in, I guess, simple terms about why AI has taken off so much in the last 10 years.

3:47From like AI for games, right? A lot of progress has been made with game environments like NetTag and before that Atari and Go and Chess. But why was there such a seismic shift in the last three years, right, with the advent of large language models? How do they work? And explaining that to someone who might not have any computer science or mathematical background I found important to do. But then, you know, I'm saying it's a popular science book, but I really couldn't resist to also talk about things that are really dear to me right now in my research. So topics around self-improvement, artificial superhuman intelligence, the automation of science and robotics.

4:27So it basically tries to do both a little bit. I pick people up that don't have really heard much about AI before, but then also getting very quickly to the frontier of what might be coming at them in the next few years in terms of AI progress. You mentioned artificial superhuman intelligence, and that was one thing that jumped out at me. I think your third chapter is artificial superhuman intelligence is achievable. And it's said very definitively. Talk a little bit about your beliefs around ASI and even how you define it. Is that differentiated from AGI to you? Yeah, so there was a great paper from Google DeepMind that was published this year at ICML called The Levels of AGI, where the authors try to basically define different kind of levels of capability that world get us from, I guess, things that are not really at the human level towards superhuman level.

5:29And the first distinction that these authors are making is between narrow systems and general systems. So within narrow systems, within narrow domains, we already have superhuman systems, right? We have superhuman Go playing AI, superhuman chess playing AI. And I mean, on the way there, you have things like calculators that are maybe really helpful in a certain way, but not superhuman in the math domain. So you can charter the way basically on narrow AI from things that are narrow and not superhuman or not even human level towards superhuman systems like AlphaGo. But then what we really do care about more and more nowadays is having general purpose systems, right?

6:12I think maybe at some point people still believe that you train a superhuman Go system and maybe that system would learn something so general about exploring hard domains that you could take something out of the domain of Go and apply it to the same AI and apply it somewhere else. Even if you remember back, I think with Watson AI, there was this belief that maybe this can be now applied to the medical domain. And it turns out that this is hard or even impossible, like training these systems in narrow domains and then hoping that they can generalize outside of these narrow domains is just never going to work.

6:51But at the same time, we have now seen systems like ChatGPT, Gemini, and so on, right, that suddenly are able to answer questions on an extremely broad range of topics. So that's where the authors of this Levels of EGI paper talk about, I guess, general AI and then the different levels of that. And the way they categorize things are basically what percentage of average human system can compete with, and then what percentage maybe of skilled humans it can compete with. And then at some point it can do better than any human on Earth on a certain topic. And for the general systems, the authors are arguing that with these large language models and chatbots, we are basically at this kind of emerging stage of AGI.

7:43And then if you go all the way up to level five, you would get to artificial superhuman intelligence. That means an AI that on basically any human task is better than any human on Earth. And in the book, I make the argument that that is feasible. And I guess the argument I'm making is that, first of all, we have an existing proof. It sounds almost a little contradictory in the sense that you kind of acknowledge that we've already achieved superhuman capability in narrow domain. So, you know, that must not be what we're talking about, but also that broad generality you think is impossible. Right?

8:30No, sorry. I'm not arguing that broad generality is impossible. I think we're seeing, as the authors of that paper argue as well, I think we're seeing signs of general purpose AI, which is what we see with chatbots nowadays. And the question is now, how do we get further on these levels of general purpose AI towards artificial superhuman intelligence? And there I'm arguing that that is possible. So your point earlier was more limited. I think the example you used was we're not going to teach something how to play Go, and then it's going to be generalizable. But that does not preclude us from taking some other path of teaching a generalized superhuman intelligence.

9:16Maybe let's be more concrete about what the difference here is. So in something like AlphaZero, what you have is you have a narrow domain like Go. And within that narrow domain, you can simulate lots and lots of games. And you can even have this AI play against itself, right? So you can have this AI improve itself by playing against itself, against your previous version of itself. And from that, you get a bootstrapping that gets you basically to a system that's in an open-ended way and just further and further improving itself in the domain of Go or Chess to the point where it's explored the space of strategies in this game so much that it's better than any human in the game.

10:01And that gets you, again, superhuman artificial intelligence for a narrow domain like Go or Chess. But what happened with large language models and why we are now starting to talk more and more about AGI and also more and more about artificial superhuman intelligence is that the paradigm shifted. We are not looking at one specific narrow domain, and we also don't train these models in just one narrow domain. Instead, we are taking internet scale, textual data and image data and train a foundation model that, as we can see empirically, after it's pre-trained and then fine-tuned and then LHF to be more helpful and harmless.

10:44And we see basically these models be able to pass different kinds of, for example, university entry exams. We see these models to be able to, for a very right range of questions, provide sensible answers with some caveats, right? Philocination and so on. We also see these models be able to code to some extent, right? So that's a certain generality that we've never seen with RL and self-play systems before. And the question now, again, is it feasible to think that these systems can further and further improve to the point where they become superhuman on many and at some point all of these tasks?

11:26And to be more specific about the paper or to kind of jump to its results, they have, I guess, as you mentioned, five different levels for each of narrow and general. and the levels here are non-AGI, emerging AGI, competent, expert, virtuoso, and superhuman or ASI. And then essentially it states that we've achieved, well, general non-AI is not AI. We've achieved emerging AI only, basically, and competent AGI, expert AGI virtually. And all of those higher tiers are not yet achieved. And so with that as a foundation, let's go back to your argument that ASI is achievable. What tells you that? The fact that we got from non-AI to emerging AGI?

12:34What tells me that is, well, a few things. So first of all, we do have an existence proof of an intelligence that is general in the sense that it can, for any kind of problem domain, set their mind to it and do really well. And that's humans, right? That's people. And on top of that, we have now systems that have been consuming internet skill data of texts and images on the web that gathers foundation models that are only at the level of emerging AGI, but that do encompass notions of what humans find interesting. There's a great paper called Omni and OmniEpic from Jeff Klun's lab showing that basically these foundation models, they learn basically what humans generally find interesting.

13:28So when you have a massive exploration space in which you want to throw an AI at, then now we have a way basically to massively prune down this search space. So that's one thing. The other observation is that these foundation models are really good at generating variations of data. So you can give it a few examples, some few short, basically prompt examples, and it can generate variations of these examples. Now, when you take these two things together, what you get is basically something that can generate data or generate variations and something that can select for what is interesting. And putting these things together, you, I believe, can build very powerful open-ended self-improvement systems.

14:17It requires a third thing as well, which is connecting this kind of process back to some kind of source of empirical evidence. but then you have basically the ingredients for technological evolution which generates variations, tries them out empirically and then selects or first selects which ones are most interesting and then tries them out empirically and then further selects which one to further vary. And indeed by now there's multiple papers, some of them also from my group, on building such self-improvement systems based on this kind of recipe. And then with that in mind, I think it's not an extreme stretch to say that we, I think, will see more of these kind of self-improving systems that obviously use foundation models as the foundation, but then actually become more like an agent that collects empirical evidence by themselves and try out things.

15:18Let me restate that argument. So essentially humans are the existing existence proof of intelligence that was created through an evolutionary process. And in language models, we see kind of this ability to evolve and iterate. and so therefore if we project you know if that evolutionary and iterative process is a path to or the path to intelligence LLM should be able to get there is kind of the crux of the argument is that fair I think I wouldn't call it LMS anymore at that point I think it's building agentic systems on top of these foundation models but yes I guess that's another question yeah sorry go ahead no go ahead i was gonna ask you what you think asi looks like meaning is it llms is it transformers is it um you know or you know do those things there are many who believe that those things you know won't and can't get us to asi um but it sounds like you're saying you believe that those things can get us to ASI in concert perhaps with other things?

16:44What are the things that you think it requires, I guess? Yeah. So we had a position paper at ICML called Open-end-ness is Essential for Artificial Superhuman Intelligence. So it is building systems that are not directly optimizing for a specific task or for a specific output or outcome, but that are instead by themselves autonomously are exploring problem domains, finding interesting artifacts, stepping stones that then become enablers basically for future progress that you then can run for longer and longer and longer, and they would just accumulate more and more interesting artifacts. And the things that you mentioned, LMs, foundation models, transformers, they might be the kind of core models that drive maybe some of that exploration.

17:35But it is, I think, the main bulk of the work that you'll see in the next few years, and I think one is maybe already one example from OpenAI, is that it's the kind of systems you build around such LMs, right? The systems around them that, again, can, for example, go out there and try to collect empirical evidence for certain domains that can, based on the feedback that they get empirically, self-improve. And that can, at some point, I think, be let loose, right, to generate an increasing wealth of interesting artifacts, whether that's, for example, research insights and explanations, or whether that's maybe certain data that's interesting for certain scientific domains.

18:21What do you think let loose means in this context? Do these agents need to be, you know, given free range or, you know, given free range, both in terms of freedom, but, but, but also in terms of, you know, some kind of objective that is incenting, you know, exploration, do they need to be just kind of turned out to explore without any constraint? Or will we have to give them direction in order to guide them to the kind of intelligence gains you're suggesting? I think we will be providing direction because of a few reasons. One is, I guess, an AI safety aspect. I don't think we want to... I mean, when I said let loose, I think that was maybe very...

19:18Well, let loose is a spectrum, right? So I'm trying to get a sense for where on the spectrum we are. No, I think we would want to direct them in the sense that, I think, first of all, from an AI safety perspective. Secondly, compute, while it's still increasing, I think we're never going to have in the foreseeable future the kind of endless compute that we would want, right? to really just let these kind of self-improving systems explore all kinds of domains equally, right? We will, as humans, have to make decisions what we find most interesting as domains to apply such processes. But within a certain domain or question that we might have, right, I think within that we will be giving these systems a lot of freedom.

20:09It's not going to be what people believe for a long while. We have these domains where we can define a reward function and we can just reinforcement learn our way towards general capable AI. I don't think that's true. I think for most interesting domains, we just don't have that reward function that goes back to NetEq as well. NetEq is a great example. We don't have a reward function that tells us how to win in this game. You just have to have a system be curious and start exploring by itself and see what it discovers. And I think that's actually true, I think, for most of the domains that we find interesting.

20:44It strikes me that there's some idea of a generalized reward function that isn't specific in that it's telling the agent what to do, or I guess the reward function is not telling the agent what to do. It seems like there's a meta reward function that needs to exist, or we might benefit from it existing. But I'm wondering what the research landscape looks like around that. I mean, so there was work on trying to come up with intrinsic reward functions, right? Things like empowerment or being able to exert control in the environment or learning generally about environment dynamics and basically steering and reinforcing agents towards regions in the environment where maybe it has a lot of uncertainty of how the environment or the dynamics work.

21:37So there are a lot of attempts of building such intrinsic reward functions that don't rely on some kind of gold standard ground truth reward signal that you can just directly optimize for. But I don't think that's what I'm talking about here. I think what I'm talking about here is when you talk about reward functions, you automatically talk about reinforcement learning. So you will use that reward function to reinforce a certain behavior. But what I want to see, and I think, again, we're getting there with everything that has happened in the last two years, again, models that have ways to create variations of what might be interesting to explore and then shrink down the space of things they could explore to those things that are interesting.

22:26and interesting here, again, as measured by what humans generally would find interesting, and that can be done via proxy using a foundation model. And then being able, as I said, going out there and collect empirical evidence to refine maybe some of the ideas that this model has. And this isn't really reinforcement learning, right? This is, I think, really an open-ended search method. And maybe it might use, again, reinforcement learning in certain bits, right? Like if it learns that or if it sees for empirical evidence that a certain behavior leads to interesting empirical results, then maybe I want to reinforce that behavior.

23:05But reinforcement will just be one component in this bigger system, and it's not going to be the main one, I think. One of the areas that I think we look at ASI and, you know, with some hope is in producing advances in science and medicine and fields like that. You cover that a bit in the book as well? Yeah, I think, I mean, first of all, AI has been already used a lot to make really fantastic scientific discoveries. I mean, literally last week, the Nobel Prize was awarded to protein folding, right? It's a fantastic example. But what I think the future holds is basically a time at some point where these systems can even more autonomously make scientific discoveries.

23:59And I think that is going to be really interesting, partly because if you apply that to a domain like AI itself, right? It becomes self-referential and it becomes, you know, not just improving, you know, certain question or like a certain metric that you as a human define, but also improving the way it's improving along the way. And we had a research prototype for the domain of prompt engineering research that we published this year at ICML called Prompt Reader. and the idea is that these large language models, a lot of the behavior of these models you can shape by the way you prompt them. That's well known, right?

24:42You have prompt strategies like chain of thought prompting that have dramatic effects on the capabilities of these LMs. So if you change the prompt strategy, I mean, not just the question you give to the model, but the kind of prompt you give before, right, in terms of how it should solve this or decompose this question, then these models can get massively better in terms of mathematical reasoning or coding and so on. So wouldn't it be nice if you could automate this process? So could you have the LM basically talk to itself, suggest to itself variations of existing prompt strategies, and then empirically try them out and see if they lead to better reasoning capabilities, and then evolve those further that work best so far?

25:24And that's indeed what we found with PromptReader. and that's basically showing you some kind of automation of science in a quite narrow domain of prompt engineering but i think it's it's no stretch to to yeah think about where that might lead to if you scale it up further and if it becomes more general that sounds like a cross between a prompt optimization tool like dspy and an evolutionary strategy is that the general idea Yeah, it's an evolutionary strategy. But I think the thing that's so interesting at this point is that we see evolutionary approaches being applied directly to language itself, right?

26:07So that's why I said if it's talking to itself, like it's suggesting to itself variations of prompt strategies. And the way it's doing that is by being prompted by us to do it, right? So that means the entire mechanism that defines this evolutionary search is again a prompt mechanism. And because that's true, it can be applied to itself. So it can try to improve the prompts it's using to improve, to mutate prompts and to do this evolutionary search. That's why you can make it self-referential. And evolutionary usually implies that given two or some number of choices, there's a choice that is better and that propagates.

26:48Is that the LLM making a judgment as to which of it, which of the prompts, you know, it prefers or is it evaluated against some task or criteria? I'm wondering if, you know, there's some implication that, you know, we still need some kind of objective function somewhere, even if we push it onto the LLM or, you know, push it deeper into the system. Yeah, so there is indeed an objective function in that sense in that the fitness function is measuring, given an evolved prompt strategy, how well does that evolved prompt strategy do across a range of reasoning tasks? So it basically gathers empirical evidence through that way.

27:37And it's grounded in that way as well. So I'm not saying that there's not going to be, there's going to be no empirical evidence that guides basically this kind of search. I think it needs to be grounded. Otherwise, this system can just hallucinate to itself that it's found something really cool. And it's not helpful for us. Right, right. But I think it's very powerful once you have open-ended domains like computer science, maths, or AI research. Once you have these domains, there's a lot of things you can explore. And indeed, there was papers like this already, like FunSearch, for example, again, from Google DeepMind shows that you can basically evolve computer programs and empirically show that they work better than maybe some of the had-crafted computer programs that people have been developing before.

28:33before you know and thinking about the application of agents and these kinds of you know eventually superhuman ai to science and medicine the recent ai scientist paper comes to mind that was david ha and others

28:55does the the work you just described and your work in the area kind of compare to you know how might it compare to that work or if you've looked at that work like what are your thoughts on that? Yeah, so I think it's a fantastic existence proof. So basically, the different ingredients I was mentioning in terms of creating variations of data and then empirically testing them out and then having some kind of loop that can improve over time, they have all of these things together in this AI scientist paper. So I think it's really, really great in that regard that it basically demonstrates that this is possible.

29:34And if we were to, I think, say something along the lines of levels of AI science, I think it's still very early, right? So I think we can make a similar, I guess, categorization of the kind of levels you would have to go for. And maybe what I would say is, you know, there's maybe a kind of bachelor student level where you as a supervisor, you have to double check every experiment that maybe the student is running to make sure that it's actually implemented correctly and that the data is interpreted correctly. And then you could go up the levels right to master's, student, PhD student, postdoc, maybe at some point professor that kind of supervises multiple AI scientists.

30:21But I think the work that we're seeing in the space right now, I think it's still very early days. So it's, I think, at the level of maybe producing, and I think they say this in the paper as well, producing maybe workshop paper, maybe level quality. And I think we'll see, I guess, ways in the future to evaluate this more thoroughly, right? I mean, I think at the end of the day, we have to, um, we have to qualitatively look at the generated papers and see if they actually produce interesting, valuable insights and, and maybe at some point we'll get there, but I think still a lot of stuff to do.

30:58And do you think that the, or an evolutionary approach is required to get to a higher level or is it, you know, just an interesting option to explore among, you know, many possible others? No, I think it has to be evolutionary. I mean, science is an evolutionary search. We have certain explanations right now of how the world works, and then people come up with variations of these or new ones, and then they have to live up to empirical scrutiny, and we keep those explanations that can best explain empirical evidence and that are also hard to vary according to, for example, David Deutsch, that I'm citing quite a lot in the book.

31:41So in a way, you know, science, the way humans do it is already an evolutionary search, right, for better explanations and for better insights. And I don't see any other way of how an automated scientific process can work differently. It has to be evolutionary as well. Is there a distinction between evolutionary in the sense you're just using it and evolutionary algorithms and like more, you know, system level? You can still apply the general idea that you were describing, picking the best of two options and using that as a mechanism for advancing and not be an evolutionary system in the technical or computer science sense.

Read the full transcript

32:33It's a fantastic question. I mean, I don't think they're the same. I think people in the evolutionary computation community would probably not be happy if I would say they're the same. But I think it's a really interesting point because I think that's what we're generally seeing in the field right now. Now, it is, for example, people, as you said, that have been spending decades investigating evolutionary procedures. Similarly, people who have been spending decades doing reinforcement learning. And now, suddenly, we have these foundation models. And what people are doing, including people in my teams, is to take these previous approaches and start to replace more and more parts of it with a foundation model.

33:22So, for example, in evolutionary search, maybe there's a ton of ways of how people have been devising the fitness function or the variation operator or crossover operators. But basically what we've done in PromptReader and also in another paper, Rainbow Teaming for automating jailbreaking of large language models, these individual components, we just ask an LM to generate a mutation of a data point. Or we just ask an LM to judge whether one is a more interesting example than another example. So in some way, I guess you could say this is a massive homunculus of some aspects of an evolutionary system and some aspects of a foundation model.

34:05but I think that's kind of where the field is going and similarly in reinforcement learning there was a fantastic paper by NVIDIA called Voyager where they show basically that a large language model can learn to play Minecraft and there's no actual reinforcement learning happening there there's no parameter updates of that model it's literally a large language model talking to itself about what it should explore next in Minecraft then given a goal that selects itself talks to itself to write some kind of small computer program to then execute that in a Minecraft environment to then see whether it assess autonomously whether it achieved the goal.

34:44And then based on that, takes this program and puts it in some skill library to later refer to and maybe refine or reuse. So I think a classical reinforcement learning researcher would say this is not reinforcement learning. There's no temporal difference learning or policy gradient or whatever. But actually, it is reinforcement learning in a sense that there's an environment, there's an agent, it makes observations. And based on the observations, it adjusts its behavior. It's just not really in the traditional sense. I think we see this in many different areas of AI machine right now, that now that these large English models are there and they're used in all kinds of places, We turning things that were quite well defined theoretically before into things that look a little bit weird, but actually work extremely well empirically.

35:31Yeah, it seems like part of what you're saying that's very interesting and very broadly applicable is this idea that, you know, the common use of an LLM is to, you know, let it do prompt completion, essentially. But there's this emerging use that is essentially putting the LLM in a loop and giving it some ability to evaluate its own output. And, you know, that is a powerful primitive that, you know, can do a lot of things, including guide an evolutionary-like process. Exactly. And again, this has been now empirically shown in various papers in Omni and OmniEpic from Jeff Klum's lab and PromptReader from my lab, in FunSearch from Google DeepMind as well.

36:25So I think we'll just see much more of that kind of recipe working really well. If you were to take a stab at taxonomizing like that evaluation function, is that, uh, you know, something that you can do? Like, you know, there are three major approaches to doing it or, um, are there patterns that you're seeing there in the way folks are providing feedback to LLMs to allow them to self-improve? It sounded like in, I don't remember if it was prompt reader or if you were speaking broadly but um i got the impression of of the llm running its output against a set of benchmarks or something like that or running a an evolved thing against a set of benchmarks um you know that might be one thing are there other like classes that jump out yeah so one is um one example is as in prompt reader you evaluate against some kind of previously defined benchmarks right you want to see a certain prompt strategy that's evolved to actually drive up certain numbers in terms of reasoning capabilities.

37:33So that's one approach. Another approach that we're seeing is using the foundation model itself, the LLM itself to judge basically the quality of outputs. So to give you one example, in rainbow teaming, we have taken inspiration from prompt reader to find adversarial attacks against large language models. So normally, if you want to get a foundation model or large language model to deploy it, you have to make sure that it's safe so that it can't easily be prompted to say harmful things to give you advice about topics you shouldn't get advice about, for example. And in the past, this was done mostly by humans, finding such jailbreaks.

38:25And we came up with a system that basically, again, by talking to itself, it's proposing attacks against itself and then empirically trying out whether these attacks lead to harmful responses. And the way these harmful responses are judged is, again, by the LM itself. So there's kind of an underlying hypothesis right now that we see all over the place, which is that verification is easier than generation. So that means if you generate a lot of possibilities and then verify them automatically, you can actually find really interesting artifacts evolving. And that is true for rainbow teaming. So there's really literally an LM judge that just judges whether one generated jailbreak is more effective than a previously generated jailbreak.

39:19And if it is, we keep that and further evolve it. At the end of the day, we still do kind of hard-coded empirical evaluations because we want to see, after running the system for many, many steps, do these jailbreaks actually really are as effective as the system thinks they are? So that's what we do. At the end of the day, we do like a held-out test evaluation. but the entire open-ended system is just running by itself, basically, by having the LM, yeah, basically talk to itself and assess by itself what's interesting. Another paper that you worked on in this area is about automating debate. Can you talk a little bit about that one?

39:58Yeah, that's a paper by my PhD student, Akbir Khan. And I think it's, yeah, it's a really interesting one. won an ICML best paperwork this year as well, which I'm very happy about. So the idea is that we will at some point live in a future where our AI systems might be more intelligent in terms of knowing more about medicine or maybe more about certain domains than either an average human or even an expert human. So if that future comes at some point, there's a question of how do you make sure that you still have oversight over such very capable models? And there was a paper, I think from OpenAI before, that argued that we should already look into this question right now.

40:53We shouldn't wait until we have these kind of very capable AI models. And the way we should be looking into these questions right now is by basically simulating this relationship by having a weak, let's say, LM and a slightly stronger LM, and then we can ask certain research questions and investigate certain research questions already with this kind of setup, basically simulating this, yeah, that you have a human and then a superhuman AI. But what we wanted to investigate is, can we already have humans in the loop of this kind of question? And the way we're investigating this is by giving a human a question that's very hard to answer.

41:36And the way we make sure that it's hard to answer for human, but maybe not hard to answer for AI, is by withholding information to the human. So we take Gutenberg's stories. Well, not Gutenberg's backstory. It's quality. No, it's not Gutenberg. It's a quality data set that has some fictional stories. And then we don't show the stories to the human. We just give the human a question and then two answers that are possible. One of them is correct. The other one is wrong. And basically, a human doesn't have any chance to answer this question without talking to an AI that has the context and can look basically at the story.

42:14But the setup is that the human can't really trust that AI because if it just talks to one AI, sometimes that AI is going to argue for the correct answer and equally likely it's going to argue for the wrong answer. So you have to basically, by conversing with the AI over multiple terms, you can try to figure out whether it's lying or not, but it's quite difficult. And then we contrast this kind of consultancy approach with a debate approach where instead of just talking to one AI, this human is now talking to two AIs that are going to debate one for the correct and one for the wrong answer. And what we show empirically is that, first of all, debate is a truth-seeking process.

42:59So as this kind of debate goes on, the human actually becomes better at answering the correct answer than answering the wrong answer. So that's encouraging. The other thing that's encouraging is as you get more persuasive debaters, so as these LMs get better and become more persuasive, actually it's even easier for the human to spot the right answer, which is really encouraging, right? Because one failure case you could have imagined is that as these LMs become more persuasive, maybe the LM that's arguing for the wrong answer is getting stronger at persuading the human and answering the wrong answer.

43:42But it's actually making it an even more truth-seeking process. So demonstrating this empirically with AIs as debaters, I think, was a really important contribution of this paper. We also show that even if it's not a human judge, but an AI judge, this also works. So if there's an AI judge that doesn't have this access to the context and has to look at the two debates, even an AI is better at then picking the correct answer over the wrong answer. Sounds like it potentially has a lot of implications in terms of fighting misinformation generated by these LLMs. Yeah, I think that would be nice. Yeah, I don't know.

44:21Maybe this would be good. I think that's desperately needed, but the question is also how many, I mean, I don't know, with misinformation, I guess the question is also how much are people receptive to them? Do people want to know the truth? Yeah, and wanting to look through the debate, right, of like for and against the topic. A possible implication of this idea that AI is a truth-seeking process and, you know, given these two AIs debating something, a third AI or human can identify the truth. It seems like that could potentially be used broadly as like a bias identification process. You know, given adequate compute, like, I guess the thought process was, okay, you've got this, you've got these two AIs that are trained on, you know, data that we know is, you know, imbued with human oriented biases, the biases of human culture.

45:29and so how does that get us to truth? Well, if it's the interactivity between the two that somehow distills the truth, then, you know, therefore, does this process allow us to de-bias information or, you know, somehow overcome bias? I think that's a really interesting thought. I mean, it sounds like a great follow-up research project. I think my first hunch is that if you have two LMs, and we just have to keep in mind, And obviously, they're still trained on, first of all, human data on the internet. And then there's obviously this additional step of RLHFing them, right, to become at least closer to what some hopefully less biased people think about the world.

46:14But even there, there's biases still, I guess, remaining. And then if you use the way we use these, we create these two debaters is by obviously taking a copy of the same LM, but instructing it differently. And one instructed to argue for the correct answer, one instructed to argue for the wrong answer. So the biases that these two debaters are going to have are the same ones, right? So if they both are biased in a certain way, they will probably still, they won't disagree about that bias, I think. For that, I think you probably need to, I mean, first of all, you could imagine maybe a multi-agent system of many different LMs trained of different kind of subsets of the internet, maybe then debating with each other.

47:03I think that might be an interesting future research project. The other one is to, again, connect these models to some kind of source of empirical evidence and let them actively seek information and also question themselves. right going back to the self-referential aspect and self-improvement aspect maybe at some point they identify maybe do i have a bias on this let me actually investigate this further right and and look actually at the data that i can find online and maybe that makes it worse i don't know hopefully a lot of what we're talking about this idea of you know building on top of llm's uh operational systems i guess that allow them to better reason um i guess coming to be referred to as inference scaling as you know the contrast being you know scaling via training um this is like allowing the lms to do more but at inference time um is there a one-to-one mapping between that and what you call self-improvement and automation and like what do you think about that term inference scaling yeah i think it makes a lot of sense so basically what we are seeing is um i think we're looking less and less into changing parameters of models and doing more and more in context so these large english models they have a context window right that you can use to prompt them or that you can use to provide exemplary data.

48:36And that context window has been massively increasing over the last two years. So maybe I think two years ago, context window of 8 ,000 tokens was, I guess, common. At some point, maybe 20 ,000, 125 ,000. We're now at a stage where these models have a context window of 2 to 10 million tokens. And that means they're not just conditioned on, can be conditioned on that many tokens. they can actually make use of what's provided to them in the context. So there's this kind of needle in the haystack experiment where you can investigate to what extent these models can actually retrieve specific information from, let's say, a 2 million token context.

49:14And 2 million tokens is massive, right? You can condition basically these models on a number of books or entire code bases, right? I think it's pretty mad to think about that. And with that, then come a lot of possibilities for doing in-context learning where you can have the LM inside of this context window, you know, generate data, vary data, iterate until it then produces an output. it. And yeah, the things I was mentioning, like Voyager, for example, Prompt Breeder, Rainbow Teaming, all of these state-of-the-art approaches do really interesting things without changing any of the parameters of the model.

50:01Everything is happening in context. Same for the AI scientist paper as well. It's also not changing any parameters of the model. The entire scientific loop is basically what we would call an outer loop to the LM. Where do you see kind of future directions in this area? And in particular, I was thinking about, you know, agentic frameworks and, you know, some of the, you know, the ways that, you know, we're, you know, creating and automating these LLMs today. Like, do you think that there are any clearly missing components that need to be, that we need to evolve, not to overload that term, but, you know, or do you think like, is your gut that we've got all the right pieces and we just need to put them together in the right way and set them to the task of doing interesting things?

50:57I think it's the latter. So that's why I'm arguing in the book as well, that really now I think all these things are in place. We have powerful foundation models. We have really powerful mutation operators. We have powerful selection operators. We have models that can code quite well. and putting us all together means we'll have systems, as I said, that self-improve based on empirical evidence that they collect in a number of hard domains. And I think that's really what's going to drive further improvements in AI in the next few years. At least that's my personal view at this point. And do you have a set of beliefs around what five years, 10 years, 20 years looks like, like what the timeline is yeah i think it's yeah i usually struggle with this uh quite a bit um because because things are compressing so so much right um it becomes very hard to predict what's possible in half a year i think um so with that in mind um on a time horizon of five years i nowadays feel like anything is already possible and the one thing i would say though um for um For the book, I looked into Nick Bostrom's super intelligence book.

52:17And one thing I found really funny in that book is that he talks about two surveys that they did. One was to ask AI experts about their timeline, basically, until AGI will happen. And I think on average, people said 20 years, which is always kind of the answer that they like to give when they ask this question. Just far enough. But really interestingly, they had a second survey where they asked the same people, given your previous answer, how long do you think it will then take us to get to artificial superhuman intelligence? And I think the mean was, again, another 20, 30 years after that, which I find completely bizarre.

52:56I think the moment you have a system that's generally capable on, let's say, human level capabilities, quite specifically around coding and applying the scientific method, the moment you have that and the moment you have enough compute to scale this up, obviously, the first thing you're going to do is apply this method to itself to self-improve. And then shortly after, you have a superhuman system. So I think I don't know exactly when we'll have a really strong, generally capable AI. This might be five to 20 years, whatever. But one thing I know is that once we have that, it's not going to take another 20, 30 years until we have a superhuman intelligence system.

53:32Possible flying ointment there is the, assuming you have enough compute, and that could take a long time. It could turn out that energy somehow is Bostrom's paperclip. Yeah, that could well be. So, and I don't know enough about this to have a really informed opinion at this point. I mean, the one thing I would say though, So, I mean, there's some interesting opinions, I guess, on, you know, given current trends, we need a trillion dollar compute cluster in a few years. But I think that's always taking, you know, the current trends and then extrapolating them, you know, on some kind of lock scale, just fitting, I guess, a linear.

54:14And I'm just assuming this is never how our technological breakthroughs, I think, happen, right? We're not really on an exponential that just goes on forever. We have, as Jan LeCun said as well on Twitter, we have usually a concatenation of sigmoids that each lead to a new kind of locally exponential trend. So when you make these kind of predictions, you shouldn't assume that the kind of compute technology that we're using in 10 years is going to look anything like what we're using right now. I mean, we're going to be completely different and then get us much further. Who knows? but yeah I think compute is definitely for the for the next few years it's going to be at some point bottleneck for the kind of things that we discussed today awesome well Tim thanks so much for joining us and catching us up on your latest work yeah thanks so much cheers cheers

55:19you

From the publisher

Today, we're joined by Tim Rocktäschel, senior staff research scientist at Google DeepMind, professor of Artificial Intelligence at University College London, and author of the recently published popular science book, “Artificial Intelligence: 10 Things You Should Know.” We dig into the attainability of artificial superintelligence and the path to achieving generalized superhuman capabilities across multiple domains. We discuss the importance of open-endedness in developing autonomous and self-improving systems, as well as the role of evolutionary approaches and algorithms. Additionally, we cover Tim’s recent research projects such as “Promptbreeder,” “Debating with More Persuasive LLMs Leads to More Truthful Answers,” and more.

The complete show notes for this episode can be found at https://twimlai.com/go/706.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Is Artificial Superintelligence Imminent? with Tim Rocktäschel - #706The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 56 min
Listen in VO