Are we ready for human-level AI by 2030? Anthropic's co-founder answers

1 Apr 2025 · 52 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Azeem Azhar's Exponential View - Episode with Jared Kaplan

Episode Overview In this episode titled "Are we ready for human-level AI by 2030? Anthropic's co-founder answers," Azeem Azhar speaks with Jared Kaplan, the co-founder and chief scientist of Anthropic. They discuss the rapid evolution of AI, the newly predicted timeline for achieving human-level AI, and the potential implications of advanced AI systems like Claude.

Key Themes and Discussions

  1. Predictions on Human-Level AI
  2. Timeline Acceleration: Kaplan predicts that human-level AI may arrive in the next 2-3 years rather than by 2030, emphasizing the rapid advancements in AI capabilities.
  3. Defining Human-Level AI: There is no clear objective measure. Kaplan suggests that the increasing complexity and capability of AI systems are more important indicators.
  1. Advancements in AI Capabilities
  2. Complex Task Handling: AI models are evolving to handle increasingly complex tasks that would traditionally take humans significant time to complete, such as analyzing lengthy documents.
  3. Test-Time Scaling: Kaplan discusses the importance of allowing AI models to "think" longer during inference to improve their performance on hard tasks.
  1. Constitutional AI and Interpretability
  2. Need for Safety Mechanisms: The conversation highlights the importance of research into constitutional AI and interpretable systems to ensure powerful AI operates safely and aligns with human values.
  3. AI Monitoring: Kaplan shares insights on a framework to monitor AI behavior using advanced systems that can reflect on and evaluate other AI systems.
  1. Economic and Societal Impacts
  2. Comparative Impact of AI: AI's influence on productivity and the labor market could be realized much faster than previous technological advancements like electricity or the iPhone.
  3. Capitalism and AI: The episode touches on concerns about the misalignment of AI incentives with public interests, stressing the need for a governance framework around AI development.
  1. Future Outlook and Excitement
  2. Expectations for the Coming Year: Kaplan expresses excitement about advancements in scalable supervision of AI to ensure beneficial outcomes as capabilities improve.
  3. Innovation Cycles: Discussion on the expected frequency of new AI model releases, suggesting continuous improvements and iterations will occur, potentially every few months.

Timestamps for Key Moments

  • 01:27 - Jared's updated prediction for reaching human-level intelligence
  • 08:12 - What will limit scaling laws?
  • 11:13 - How long will we wait between model generations?
  • 16:27 - Why test-time scaling is a big deal
  • 21:59 - DeepSeek's competitiveness algorithmically
  • 30:08 - Managing the paradoxes of AI progress
  • 32:21 - Interpretability and AI safety
  • 39:43 - Misalignment of model incentives with public interests
  • 42:36 - Preparing for electricity-level impact
  • 51:15 - Jared’s excitement for the next 12 months

Conclusion This episode of "Azeem Azhar's Exponential View" delves into critical discussions surrounding the future of AI, particularly focusing on the rapid developments that could lead to human-level intelligence much sooner than previously thought. The insights from Jared Kaplan provide a nuanced understanding of both the excitement and the caution required as AI technology continues to evolve.

For follow-up, listeners are encouraged to explore additional resources and updates through Azeem Azhar's Substack and Anthropic’s website.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You last year put forward the prospect of human level artificial intelligence by 2030. If anything, I expect it probably sooner than 2030, probably more like in the next two to three years. What would need to be true in terms of DeepSeek for it to propel itself beyond the capabilities of U.S. frontier models? There's so much low-hanging fruit to collect that it's unpredictable who's going to sort of find which advances first. There's no reason why they can't be very competitive algorithmic. What does it mean to be interpretable if the machines are operating in spaces that make us look not so much like silverback gorillas, but like hamsters?

0:36One of the directions we're moving is where you could have AI systems that think about what another version of plot is doing in order to sort of monitor it and steer it in a good direction. So by the time you're at a point where AI is sort of as smart as people or beyond, you're able to leverage those smarter than human AIs. The way in which these models impact economic productivity and the labor market could be much faster than the canal, electricity or the iPhone. What is the debate that ought to be happening that would most help us prepare for that kind of fast deployment scenario? Is it really safe to have AI that is smarter than you?

1:16And I think that is a real question. Like, should we be having these super intelligent AI aliens kind of invading the Earth or should we decide not to? I'm delighted that I've got Jared Kaplan, who's the co-founder and chief scientist of Anthropic. Jared, it's great to have you here. Thanks so much for having me. It's great to be here. You, last year, put forward the prospect of human-level artificial intelligence by 2030. Given everything that's happened since then, what's your current assessment? I mean, I think, if anything, and I mean, maybe I'm drinking too much of my own Kool-Aid, but I mean, if anything, I expect it probably sooner than 2030, probably more like in the next two to three years.

2:01But what is human-level AI exactly? I mean, it's not something like an objective measure that you either cross the line or you don't. I think AI is just going to keep getting better in a lot of different ways. So you've raised a really important question there, which is, what is it? You know, it's not like landing two astronauts on the moon and bringing them back safely. That was very, very clear and well understood. What is the purpose of having a test that says human-level AI? Yeah, well, we don't really have any tests. I mean, at Anthropic, I mean, we have probably tens, hundreds of different tests and evaluations that we're running on Claude.

2:37And as time goes on, I think, honestly, the experience of working with Claude, collaborating with Claude, and like kind of getting productivity benefits from that is in some ways a better measure for how useful Claude is, I think, than any given test. I guess the way that I think about sort of how capable AI is, is sort of on two axes. There's what environments can an AI actually go out and act in? So I go back to sort of AlphaGo, which was this superhuman Go playing program, better than any human, smarter than any person. But it was restricted to be on a Go board, a very, very restricted little grid that you can act on a very specific game.

3:22And as we sort of developed large language models, large multimodal models, the different environments that AI could interact in has grown. So I mean, it grew a lot where you could just talk to chatbots like Claude. I think it's grown further where AI can understand images. It goes further when AI can use computers. and eventually, obviously, the thing that we all imagine, the sort of sci-fi thing, is AI being embodied in a robot that can go out into the world. So I guess that's one of the directions I think of. The other is just like how complex are the things that AI can do? Like can it do something that would take me a minute or 10 minutes or an hour or a day?

4:04And I think we're just going to keep moving in that direction and that's where sort of AI has to sort of actually take action to the world and learn things the way that we do to be useful in that way. Yeah, I think that's a really helpful way of thinking about it. So if I play that back, one is, what is the range and breadth of domains in which it can operate? And, you know, of course, we've gone from text to multimodal images. And of course, that next boundary being the physical world. Although I always think that we do a lot of our most useful work in our heads anyway. And then the second one, I think, is so interesting.

4:39It's this idea of what is that unit of human time that the machine is able to operate on? Because the very early large language models, if you go back to, I guess it was BERT, right? They did tasks that were a second. Look at a sentence and find a noun. And then when you got the first versions of, I guess, GPT-3, you had tasks that could last maybe 10 seconds. Look at a paragraph and pull out a sentence. And of course, if you go to Sonnet 3.7, I mean, I can give Sonnet 3.7 tasks that might take me hours. So here is 20 ,000 words, distill out eight or nine of the key arguments, identify where they are coherent with each other and where they don't agree with each other.

5:22And that's a job that would take a graduate student half a day. And it's quite interesting to see the rapid progression of that sort of duration of task from these models. So I suppose one question I would have for you is, is that something that you track and you can forecast in your headspace, 3.7, which is the latest Claude, how long can it operate for? It's a great question. Yeah, no, this is something that I track and it's definitely something, certainly it's something that I think about very actively. And a lot of our research is oriented around this. I think there's a sort of, we talked about it as sort of the horizon that Claude can operate on.

5:56I think the way that at least if you're a developer, and obviously not everyone is, that maybe this is most visceral with something like Claude code, or you can ask Claude to sort of search through a repository of code, make changes across all sorts of different features, and maybe iterate and test the code itself. So I think those kinds of capabilities feel like they're the most complex. As you said, a lot of what we do happens in our heads, and that's true for me, too. But I think that the way that we really get a purchase on the world is by trying things and see what works and see what doesn't.

6:30And so I think that's what really allows you to extend this horizon. It's definitely something I track. I mean, it's something that also I remember people years ago who were AI enthusiasts talking about, well, maybe AI won't be able to do things that take longer and longer. And I think we are seeing this horizon expand. And so the utility of AI goes up. Why does the horizon expand? Is it more memory in the GPUs? Is it some bit of magic that you're tweaking? It's a great question. So I think it's a few things. I think one aspect is just sort of the model intelligence kind of in a general sense is going up.

7:06So the model is able to attend to more different issues to track more things. Another is sort of the context length. So the context length of our models keeps going up. And we find that we can extrapolate it much further than anything we've shipped yet. And so it should be possible for AI to understand more and more, like to go from understanding like a paragraph to a chapter to a book to something much, much longer. That's helping it. And then finally, I think that we're training AI using reinforcement learning to do more complex tasks in a sort of useful way. So we're training AI to do more complex coding tasks, to study longer documents, exactly like the example you gave, and distill out more information.

7:49We're always sort of trying to find the tasks that push the envelope on what Claude can do and train it to get better, better there, just as like in our own educations as people, like we're always trying to solve harder and harder problems as we get older and we progress from elementary school to high school to university. So I think it's all of those things together that are sort of pushing this envelope. Right. But, you know, when we think about large language models, a lot of the emphasis is on that first L, the large. And we lived in this regime of scaling laws where, to some degree, there was a predictability that if you 10x'd the size of a model, which would mean 10 times as much data, 10 times as much compute, 10 times more complexity in the end, you got this sort of, you know, predictable linear improvement in how well the model worked.

8:36And the argument has been that that kind of scaling is getting harder and harder. Either it's that we're running out of data, or it's really expensive, or it's actually just really complex. I mean, you've been at the front line of that. When you look at that pre-training scaling, what has been the bit that has started to put the brakes on the rate at which we're seeing results from it? Maybe zooming out for a second. I think they're sort of scaling laws as a quite precise, and that was what was so surprising about them, very precise empirical finding just from studying AI and AI training. And that was, you put it beautifully, that if you increase the size of neural networks, the number of parameters they have, if you increase the amount of data, if you increase the amount of compute you use to train, then you get these stunningly predictive curves for how well the AI can model its data, can make its quote-unquote loss go down.

9:41What this really means is how well can large language models predict the next word in a sentence, paragraph, document, etc. And that just very, very precisely improves as you scale up. And we haven't seen any limits to that. I mean, we are seeing that, I think, as you make models bigger, as long as you have all of these ingredients that you mentioned, model size, compute, and data, you still get improvements. I think probably the limiting factor that people talk about the most for good reason is data. Eventually, one is going to run out of data. I actually don't know that we have reached that point yet.

10:19I guess we will see. But I do think eventually in the next couple of years, we'll reach that point. Certainly cost also matters, but there are all these different other ingredients, I think, that are driving cost down. I think we're finding algorithmic improvements that make model training much more efficient. And we're also seeing that hardware is improving very quickly as well. And so costs are going down for those reasons. So I think the sort of scaling is going to continue. Now, there's a separate question. There's this very nice empirical statement that the AI can model its data better, but that doesn't necessarily mean it's more useful for you.

10:54That doesn't necessarily mean it's more useful as Claude. I think generally, it has that implication, but it's much less precise. And so it's possible that sort of the gains that we get in what AI can do for you will come more from training it to do useful tasks after pre-training rather than pure scale. Right, right. Okay, I want to talk about the after pre-training in a second. But it's been, I think, one year and 10 days since Claude 3 was released. So belated happy birthday to Claude 3, which I think became everyone's most personable LLM a year ago. One of the things that we've seen happen in the full AI stack is that the generation time has got shorter and shorter.

11:38So semiconductors used to be on a three-year cycle. Jensen Huang has put them on a one-year cycle. AMD is responding in a similar sort of way. And there has been an acceleration, but we're over a year since Claw 3 came out. So what is the right time between generations for these large language models? And what should we expect as consumers on the other end? I think that the generation time for models has been really, really fast. At least to me, it feels fast. And I think that's basically going to continue. So I think that we should expect a new generation of cloud models in not too long, certainly in the next six months or so.

12:18And I think that basically that's going to continue. And it's both because we're improving sort of post-training or reinforcement learning, training cloud on more tests, and because I think we're able to improve the efficiency and intelligence from pre-training. So I think that's not slowing down anytime soon. I think in some ways the model cycle is even faster than the hardware cycle. We'll see if the hardware cycle is really one year, but it's definitely moving quickly and we're getting new chips sort of as we speak. There was that very, very fascinating and challenging paper written by a young man called Leopold Aschenbrenner last year, which I'm sure you would have read.

12:58And in that, he had a two-year generation time between these sort of order of magnitude improvements in models. And I remember reading that and I was thinking, I don't think it can be two years because honestly, it takes time to build a data center to get the chips from NVIDIA and to find the power and to generate the data, assuming you needed synthetic data in some cases. When you look back at that now with a bit of distance, you know, how would you communicate that sense of practical generations in terms of how we should expect as consumers, as members of society, these models to improve? Is it a three-year clock cycle?

13:36Or is it going to be a two-year one? I guess I think of it as being smaller than two years, but that's maybe because we're iterating very quickly. So I would say that there's some pre-training life cycle, but that's usually measured in months rather than years. Now, there's a question of research, like how quickly can researchers come up with new innovations that are sort of worth shipping? But I think of that as being quite a bit shorter than a year. And then with reinforcement learning, I think that historically that's been much less compute, much less resource intensive training that may be changing now.

14:14And for that, I think we can iterate much more quickly. So I think that there's sort of a desire, at least as we develop Claude, to every time we think that there's like a significant improvement that we can deliver in Claude. And that might be for any of those reasons because of pre-training improvements or because we've just realized that we can train Claude on some new task, like what goes into Claude code and that will be useful to people. I think we're expecting to ship it. So I think it's really, it's more of a continuum. It's more like, you could ask how quickly does Moore's Law develop?

14:48And it's really more of a continuum where it's like, I don't know, it's like maybe it doubles every 18 months or something like that. I think with AI, it's faster than that. I don't know how viscerally that feels when you're playing with each generation of models. But I mean, you can tell me, I mean, you've been playing with Cloud 3.7 Sonnet, you played with Cloud 3 Opus. Like, I don't know what you feel is the biggest. I would know. I mean, it's completely wild. the rate with which we have to update our behaviors as somebody who uses these these tools is really really astonishing and i find myself coming up with something you guys introduced something called called clawed projects um a while ago and so i built lots of projects to help in my research work then as the models got better and of course i use i use every family of the models.

15:39I use the OpenAI ones and the Perplexity and U.com and lots of other ones all at the same time. And I also often use them adversarially because they all have slightly different flavors. I used to think of Claude as being that really super smart history grad student when you're an undergrad, kind of charming, knew a lot of stuff and was never cleverer than thou. And I would think of another model being a bit like that nerdy mathematician who always had to get it right. You know which one I'm referring to there. And so there was this different impersonalities, but I have found that the rate with which they feel like they're getting better means that I almost don't document my changes in behavior.

16:20I'm just literally living it through practice. And yes, you're right. So I think it feels very fast. So I want to come to this other question, which is this phrase that Satya Nadella and of course Jensen has now started to use quite a lot from the middle of last year, which was test time scaling or inference time scaling. What is it and how big a deal is it? I think it's a big deal. So the claim here is that as you let an AI model think for longer, then you can get predictable improvements in the accuracy when it's doing a hard task where just, say, pure thought improved performance. So I think the classic example is like solving a really hard math contest problem or a competition coding problem.

17:12What we see is that as you let literally like, say, Claude's movement, seven son, think for, say, a thousand words or two thousand words or four thousand words or eight thousand words or sixteen thousand words, you get sort of predictable improvements in which each doubling of the amount of time Claude can think, you get sort of a constant increase in performance. And you can actually also see that extended in other ways. Like you can train a separate AI model, or you can even just ask plot itself to analyze the solution, but you can train it to decide sort of which possible solution to a problem is best.

17:52And then in parallel, you could generate one solution, two solutions, four, eight, etc. And you can ask it to choose which is the best of the ones that have been generated in parallel after the fact. And again, I think you tend to see pretty clean scaling where you can get better and better performance that way. And so I think that for very difficult tasks, this means that you can either choose to have a smarter model, sort of solve it in a shorter amount of time, or you could ask a smaller model to work longer and maybe get the same performance. And so I think this is exciting because for the very, very hard tasks that you might want AI to solve, maybe if you just kind of throw enough test time compute at it, you can solve them.

18:35I mean, I think the things that we imagine in the future are things like helping to cure diseases or making new breakthroughs in theoretical physics, things like that. Right. So the additional thinking, this doubling of time spent thinking gives you this predictable improvement. I noticed in the new version of Claude 3.7 Sonnet, you've got this thinking time. And it's just a little feature that flag that I tick. Is there some way of looking at my query and then you make a judgment of how much thinking should go into that type of a query to give me a sufficiently improved response rather than have me sit around waiting for forever?

19:16Yeah, indeed. So, I mean, Claude 3.7 Sonnet is this first hybrid sort of reasoning model where it can act just like Claude 3.6 Sonnet or 3.5 Sonnet new. I mean, we're great at naming. Your naming is, yes. Are you still using AI to help with your naming? A very, very primitive or like, yeah, we're using like GPT-2 to name our models. But yeah, Claude 3.7 Sonnet can behave very similarly to prior generations where it doesn't think at all, or you can ask it to think, and it tries to sort of decide from its training how much thinking to do. And sometimes you wish it would think a little bit more, sometimes a little bit less, but But basically, based on sort of the difficulty of the task you assign to it, it will think the amount that it expects is best.

20:07And that is something that we're working on and we think will get better over time is that it should become possible. Kind of let it think as much as it wants, but kind of get a response in whatever the most reasonable time frame is. And I guess the question is just like, I don't know, if you start a new job and your boss gives you something hard to do, you might really want to spend a lot of time thinking because you really want to get the right answer. you don't want to get fired. But on the other hand, in some situations, maybe once you're comfortable at your new job, you might feel like, oh, I'm just going to give a quick answer.

20:38Like, we have a good relationship. So I think Claude is sort of in the same space where it doesn't know, like, am I expected to try really, really hard to get the best answer? Or should I just, am I wasting someone's time? And so that's something that we want Claude to sort of learn from context over time. Try to make a little bit of that in, but hopefully that will improve with with future generations. But architecturally, is it a pre-processing step where the query comes in, you do some kind of a process, you make some kind of judgment, you send a parameter to the underlying model to say, think for X thousand tokens or think for 2X thousand tokens, or is it all in the single model?

21:17It's all in the single model. If you're a developer, you can specify precisely sort of what budget Claude gets. Right. Sort of 99 plus percent of the time, it will sort of stay within that budget. And a lot of the time, it will actually undershoot that budget quite a bit. So you'll say, you can think for 16 ,000 words, but it'll only think for 4 ,000 or something. And it's using its own judgment in deciding, should I use all the space I have or not? As a user on Cloud.ai, I think there are sort of fewer parameters that you can set just to keep things simple. But it's all one model that's sort of deciding based on what you've asked, how much to think.

21:58For a lot of people, this thinking time experience exploded into their understanding. Inauguration Day this year, January the 20th, because of DeepSeek R1. How big a deal was that in Anthropic? I had been following DeepSeek's progress for at least sort of a year, year and a half, because they've been writing papers and improving their models. So it wasn't actually very surprising to me or too anthropic. It was interesting to see, though, like the reaction sort of globally of, wow, China has this great model. And there are people that I talked to in the US who have thought historically maybe China's many years behind.

22:39And seeing Deep Sea's progress in the papers they were writing, I kind of thought, well, they're like, I don't know, maybe they're six months behind, but they're not very far behind. Yeah, yeah. I mean, it was apparent with what was called V3, which was, I think, October 2024, where they were achieving GPT-4 level capability at 30, 50 times cheaper. And then I think the earlier paper was maybe November 2023 with version 2, and you were starting to see some pretty good results. From the outside, I suppose, partly we're often looking at that chart that shows the gap between the top ELO-rated models.

23:18So an ELO rating is like a chess-like rating for how good models are in chat. I know you know this, Jared, just for people listening. And, you know, you see the frontier is a bunch of American firms, and the best Chinese model is quite far behind. And over the course of the months, the delta has got shorter and shorter. And so you're looking at that and you're saying, well, is it just going to be a delta where the Chinese are slightly behind the best U.S. models? Or is there a real momentum within those Chinese firms that will take them past the frontier U.S. models? What would need to be true in terms of the kind of research breakthroughs that DeepSeek might need for it to propel itself beyond the capabilities of U.S.

24:03frontier models? Yeah. I mean, research breakthroughs are happening very, very quickly. I mean, maybe in a certain sense, they're not even breakthroughs. I mean, one thing that I always say is that when you see really rapid progress in science, it's not because the scientists suddenly got much smarter. they're not superhuman. It's because people have found an area where there's a lot of like very low hanging fruit, a lot of iteration that can be used to improve. And so I think that's what's been happening in AI. I mean, maybe for 10, 15 years, certainly the last five years. And I think there's so much low hanging fruit to collect that it's unpredictable who's going to sort of find which advances first.

24:45My expectation, though, I don't know for sure, is that going forward, I think that there are these sort of export controls that are in place that I think mean that sort of Western firms will probably have an advantage in terms of the amount of compute available. And I think that will probably make it more difficult for DeepSeq and others to be competitive. But I think that in terms of the basic algorithms, all of the sort of leading AI companies, I think, are finding ways to do very simple things that work well and scale well. And there's no reason why, I mean, I think DeepSeek, based on their papers, has also found a lot of these ideas and techniques.

25:26And there's no reason why they can't be very competitive algorithmically. But it does seem that there's been a tone shift in the discussion of AI and AI development. And I know that your co-founder, Dario Amadai, has been speaking quite a lot about increasingly the importance of some sort of export control or licensing regime around chips as regards getting them out to China. And I even detected, and please correct me if I've got this wrong, a shift in Anthropik's approach to how quickly development needs to happen. And I got a subtle sense that maybe Anthropic, which is often argued very strongly in favor of a sort of a safety first approach, had slightly changed its gate and said, we just need to go faster than we have.

26:17Have I misread that? The way that Anthropic thinks about sort of speed of development and its interplay with safety is primarily through our responsible scaling policy. So it's true that in those sort of very early days when we founded Anthropic, I think there was a general sense among us, although I think this was not something that, I mean, AI was kind of not a big deal in the wider world. Our sense was that AI was going to make very, very rapid progress, and that that was primarily going to be very beneficial to the world, but there were also a lot of risks associated with it. And we did have some sense that this powerful technology being developed slightly more slowly could actually be better in terms of sort of getting things right.

27:01Now, the way that we've sort of, we kind of figured out, and I think this was kind of a breakthrough of its own, was that sort of we created this responsible scaling policy as kind of a way to help coordinate with other labs to make sure that AI development is beneficial and isn't creating harms. So the idea there was that we would think carefully about what kinds of real risks from AI exist that we want to actually take seriously. And we would measure the capability of our systems to sort of actually do harm, to be risky. And then if we crossed certain thresholds, we would basically commit to have mitigations in place to avoid those problems.

Read the full transcript

27:48And so in a certain sense, this meant that when we were thinking in these terms, we were more free to move quickly because the idea would be we have this framework in place where we can move as fast as we can, both in terms of AI capabilities and also safety research and risk mitigations. and we would be sort of bound by, to sort of not cross certain lines until we were ready. And so I think that's something that we've sort of alluded to with even the release of Claude 307 Sonnet and the way that we've sort of discussed it in its system card, that we think that we are approaching some of these thresholds and therefore future models may need more protections.

28:31But at the same time, in terms of our research, we are getting those protections in place. So we had this both research release that we called constitutional classifiers and an associated sort of jailbreaking demo where we asked people, sort of anyone on the Internet, to try to sort of jailbreak this new system in order to sort of test out this method. And so basically what we think is that we want to move as quickly as we can for a variety of reasons, but we want to make sure we have these systems in place. And so that's kind of how we're kind of coordinating to, we hope, scale responsibly. Yeah, I mean, this gets to the heart, I think, of some of the paradoxical ideas that you have to hold in your head, you and your colleagues.

29:18You know, on the one hand, the way in which we're building AI today has many critics. And the critics say, this system is uncontrollable, right? You don't have any verifiable safety and we don't have the science for how we build safety around this method of making AI. And you have a great comment. I'm just going to read it here. Maybe supervising a thing that's smarter than us is hard. Maybe not. But once you make a thing that's broadly much smarter than you, and given that it'd be easy to run millions of copies of that thing once you have one, you're going to lose and be disempowered if there's a conflict.

30:00given the stakes, being 90 % sure it'll work out is very far from okay. I think that was you. Was that you? That sounds like me, yeah. It sounds like you, right. Okay, good, good. So I guess one thing I'm curious about is how do you personally manage that paradoxical piece of the sort of risks that you've yourself articulated and the work that you are doing and the speed with which it's moving, generation times that are less than 18 months. I mean, is there some internal cognitive mechanism that you've built for yourself? Yeah. So, I mean, I guess I can talk about all of the different kinds of research we're doing to try to meet the moment.

30:44I mean, I think very broadly, I mean, you mentioned DeepSeek. There are all these AI labs in China that are near the frontier or almost at the frontier, at the frontier. There are a lot of different AI researchers and labs in the U.S., scattered across the world. Everyone's sort of pushing forward this technology. And obviously, I mean, you've mentioned the potential for risk and critics that say that this technology is dangerous. On the other side, obviously, there are a lot of people saying that's really silly. That's totally unnecessary. What's most important is delivering the benefits of this technology.

31:24because other things that I've said, that Dariou have said, maybe this technology can help us to cure cancer 10 times faster than we would otherwise. And so people say things like, how could we possibly slow down when we're going to be able to deliver those kinds of broad improvements for human welfare? So there are a lot of competitors. There's a broad capitalist ecosystem moving very quickly. There are a lot of benefits to this technology, but also if it's really as powerful as millions of geniuses in a data center, etc., then there are also risks. And so it's a difficult thing to juggle. I think that as AI becomes more capable, I think the stakes go up.

32:05And I think the confidence level you would like to have before you develop or deploy such systems also goes up. The way that we're thinking about this is through a lot of different research directions. Interpretability is one we talk about a lot. And it's something that even I was unsure about, I would say, I don't know, three or four years ago. I wasn't really sure whether we'd be able to get any benefits. But I do think that we're starting to see... To say, what's the benefit of interpretability? The benefit of interpretability is basically, if you have a really advanced AI, and you're not sure if you trust it, it sure would be useful if you could read its mind.

32:46And it sure would be useful if you could not even just read its mind, but understand how it puts its thoughts together to take actions in the world. So interpretability, I think, first and foremost, at least, I think could be a very, very powerful way of checking whether AI really is doing what you want it to do. Right. But as AI gets better on this continuum, the way in which it's going to make its decisions are going to become less and less interpretable to us. You're a theoretical physicist. You are somebody who understands tensors, you know, multi-hundred dimensional spaces, and you have formalisms that allow you to navigate and make sense of those.

33:27I can say the words in English, but all of the work you do as a theoretical physicist is not interpretable to me. And so at some point, maybe it's 18 months, maybe it's 36, maybe it's 54. On your trajectory, We will have systems that will be smarter than everyone. We have the Astronomer Royal in the UK, Lord Martin Rees, watching this right now. They'll be smarter even than him. And the number of people who could interpret that then drops. It gets to a world where only Terence Tao, the mathematician, can interpret it. And then he can't. So what does it mean to be interpretable if the machines are operating in spaces that make us look not so much like silverback gorillas, but like hamsters?

34:13Yeah, it's a good question. So there are a couple of things I would say. So one is that if you want to understand everything that the AI is doing, I agree that's going to be very difficult. But you might be able to study a bunch of specific examples where you can understand, for example, what is the goal that the AI is seeking? Maybe you don't fully understand all of the steps in the process, maybe you can break it down and study it intensively. And you can. Maybe you can use AI to help you to analyze what's going on inside. Maybe you can use a dumber AI to help you to understand a smarter AI.

34:48But generally, there may be components of what the AI is doing, like specifically its goals or the way that it might respond in particularly dicey situations that you can set up and analyze that might give you insight into whether the AI is sort of ultimately aligned or not. The other thing I would say is that interpretability is really just one tool. So I think that it hopefully will give us insight and some way of auditing AI and checking what's going on. There's a few other things that we're doing simultaneously. One is trying to figure out better and better ways to use AI to help supervise and monitor AI.

35:28I think both of those. So constitutional AI, which sort of Anthropic developed a few years ago, I think was like the very first example of this, where we used AI systems to check whether other AI systems were obeying a constitution and to guide them towards behaviors that were in accord with a list of principles that we call the constitution. So what we're trying to develop now, one of the directions we're moving is kind of like constitutional AI souped up, where you could have AI systems that think using the reasoning we were talking about earlier about what another version of Claude is doing in order to sort of monitor it and steer it in a good direction.

36:09And so the goal there is that you want your monitoring and the supervision you have of AI to improve with the intelligence of AI. So by the time you're at a point where AI is sort of as smartest people or beyond, you're able to leverage those smarter than human AIs for alignment. So can I jump in with a historical parallel? So that's a little bit like an enlightenment idea, right? It's an enlightenment idea in the sense that you layer on top of what's gone before. And I have a sense of, when I use these large language models, that they do embed that idea already. It's really, really difficult to imagine a large language model that will believe the earth is flat.

36:56Because in order for it to believe the earth is flat, it can't then be helpful in other ways, right? Because so many of the internal relationships have to be broken, that it's not going to help. So there is some sense, it almost feels that this is like a teleological argument, which is that if each subsequent generation is roughly aligned, there is an arrow of progress that suggests that you build one on top of another. And I mean, is it that there's a kind of a selection pressure that goes on that models that end up diverging from that path become less useful and therefore they don't benefit from economic incentives?

37:35Or is that, am I just drawing too strong a parallel with history? Well, I think there's, no, I think it's a good parallel. I would say there's two things. So one is we got extremely lucky And I remember talking to folks concerned about safety who were excited about this back, I don't know, in 2017, 2018. There was this vision that AI was going to look like AlphaGo, where it starts off as a completely blank slate, and it trains just to optimize to do a particular task. we got, I think, lucky, or maybe we made good choices in developing large language models, because fundamentally large language models, the very first thing they learn is to understand human writing, human ideas, the way humans use words to conceptualize the world.

38:27And obviously that has encoded a lot of our biases. It has a lot of the intuitions we have about ethics and morality, common sense, human history, the things that have gone well, the things that have gone wrong. And furthermore, it means that the very first thing we got these models to do was to chat with us. I mean, I think it's not a coincidence that dialogue, chat, was one of the very first breakout applications because these models are trained on language. And so I think that is sort of exactly in sync with what you were saying, where there's a very, very strong bias that these models will at least understand and be able to communicate, and their basis for thinking will be very much a mirror to our own.

39:05Now, it might be a strange mirror in a lot of ways. It might be a funhouse mirror, but nevertheless, they really need to sort of understand all of those ideas. And so I think that does give us a pretty strong foundation. And it also means that at least some informal senses of interpretability, like being able to look at the chain of thought thinking these models are using and understand it, I think is sort of baked into the model. So I think that is a major advantage of the way that AI has developed. It didn't necessarily have to be that way. a lot of people, no one really predicted it five or six years ago.

39:38But I do think that it makes our task of alignment a little bit easier. Right. But you know, you have within this idea, the constitutional AI, you have character training. These are all decisions that you make about how the model is going to behave and what its sense of, you know, goodness is, right? I mean, goodness giving us things in ways that are honest and helpful and harmless. But those are choices of technologies that could end up being quite infrastructural. I think many people now believe that if the 19th century infrastructure was canals and railways, and then sewage systems, 21st century infrastructure will be a layer of AI.

40:24And the nature of infrastructure is that it is a social product in of itself. But the way that we currently guide the success of these models is actually through the interactions in the market. And it's through the interactions in the market with lots of different SKUs. So if you're building a based model like Grok, you're not competing on a fair basis with Claude because you have 200 million people a month on X that you can pour at that model. Or if you're at open AI, you've developed some forward momentum. So these models are not going to compete in a way that's necessarily aligned with public interest.

41:08Is that a problem? It's definitely, I think for any advanced technology, it's a problem. I think it's very much a social problem that's fundamental to sort of capitalism, though, in the sense that with any new technology, there are certain capitalist incentives. There's a sense in which those often are not completely anti-aligned. I mean, people are buying products. They put their dollars where their opinions are. But I do think that there are externalities. There can be problems introduced with new technologies. Capitalism is very, very, very far from perfect in terms of making sure human welfare is improved, is broadly available.

41:46But I think a lot of the problems are similar to problems we already have in the world. They're just potentially, if the technology moves very quickly, then maybe we'll have to face them more quickly than we do with other technologies. But I think they're broadly similar. Also, I mean, we do try to make in some flexibility. Like, you could ask Claude to roleplay in a lot of different ways. Instead of sort of sounding just like Claude, it's happy to sort of sound different. And I think we try to set some guardrails of sort of basic harmlessness on some of this. But you can change Claude. And I think part of that is that we do think that it shouldn't be up to us to completely dictate the value.

42:24I think AI is fundamentally very empowering and should be sort of broadly accessible. And we want people to sort of be able to benefit in whatever way they see. I mean, I think freedom to use the technology in the way that you want is really important. I think one discontinuity is between what's happening in the Bay Area, then what's happening within the frontier labs and Main Street, for sake of argument, and a sense of how quickly or not things are moving. And I get a sense from you that, you know, you really feel things are moving very, very quickly, that the generations between advances is quick, quicker than perhaps we feel outside.

43:02There's a contrary perspective, which is that, you know, short-term equilibration is often constrained by social organizational inertia, by the fact that you have to build power stations or just get people to change their behaviors. I mean, I've been using large language models now for, you know, regularly since November the 30th, 2022. I'm still doing searches on Google. I mean, not so many, but I'm still doing a few. So there is a sort of contrary view that says paradigmatic change still takes quite a bit of time. But that still leaves the possibility that the way in which these models impact economic productivity and the labor market could be much faster than the canal or electricity or the iPhone.

43:56What is the debate that ought to be happening that would most help us prepare for that kind of fast deployment scenario? That's a great question. I mean, so I was a theoretical physicist until pretty recently. And often when I talk to physicists, I sort of say, if you believe me, and I mean, obviously, people should be skeptical about AI. People should think about it carefully. I, if you'd asked me six or seven years ago, didn't in any way expect this. I sort of slowly became more and more convinced that AI was going to make very rapid progress. Now I'm pretty convinced, but I might still be wrong.

44:36So people should have appropriate levels of skepticism. I certainly question, is AI really going to be on this trajectory? But if you accept that, I mean, the thing that I say to my fellow physicists is, look, we really need people with all different kinds of intellectual training and experience to work on this, because if this is really happening, it's pretty revolutionary. And we want sort of our best and brightest to be paying attention. want everyone to be paying attention and thinking about what it means. In terms of what should we be thinking about, I think there's a lot of interesting debates about what the economic impacts of AI really will be.

45:12I think it's different from a lot of other technology. I think something that's very interesting, and I don't know what's going to happen with it, is that in terms of where is AI most useful, most productivity enhancing, it's really sort of more educated, sort of white-collar jobs that might be impacted. And so I think that's a difference, I think, from the way we often think about automation. So I think parsing the consequences of that is pretty interesting. We ourselves, I guess, we're trying to study empirically, we're very focused on empirical approaches for AI and otherwise, how AI is getting used.

45:45So we have this tool, Clio, that allows us to, in a privacy-preserving way, sort of aggregate how Claude is being used. And we're studying things like, is that complementary? Is it productivity enhancing? To what extent is it potentially replacing tasks that people would otherwise do? And we're sort of opening that data set up to economists to try to study. So I think that's an example of a place where I think there's a lot of interesting work to do to try to understand, like, what is the progress that it's going to be? Because right now we see a huge uptick in usage of AI for software engineering.

46:21And I think that's sort of the perfect area because software engineers love to adopt new technology. It's very exciting. And also software is verifiable, right? So if Claude produces something that doesn't work, it won't execute, it won't pass its unit tests. So you can go back and say, listen, this didn't do the thing I needed to do, do it again. Exactly, exactly. So I think there's a lot of reasons why software engineering is a natural place for AI to be adopted. But I think it's a great question. Is something like what we see in software engineering with so many folks using AI, is that going to happen all across knowledge work or is it going to be much slower?

46:57And then how is it going to kind of permeate our day-to-day lives? I think it's interesting to think about that. I think it's interesting to think about, do we really want AI that's at or beyond human level? You were asking me all these questions about, is it really safe to have AI that is smarter than you? And I think that is a real question. Like, I think it's a society level question. Should we be sort of having these super intelligent AI aliens kind of invading the earth or should we decide not to? And how, from an international perspective, do we decide how we roll out this technology? You know, I frame it slightly differently to that because I see you and your colleagues are evolving, right?

47:37Each subsequent AI is really impressive. And if you had shown me Claude 3.7 sonnet thinking five years ago, I would have 100 % said we have AGI. We have it. And what you're doing in a sense is you're sort of progressively disappointing us because we're getting so used to it. It's like, oh, wait a second, 3.7 sonnet's two weeks old, right? I'm not interested in that anymore. But one of the reasons I slightly disagree with the framing of a superintelligence that that appears out of nowhere is because it's not appearing out of nowhere. It's appearing in an evolving way in our hands. We're also starting to recognize that as with any system, there's an efficient frontier.

48:23No model is better than every other model in every way, right? There are trade offs. And the way it will get embedded across our economy will be, to be honest, the model that's going to run the OCR on a scanner in a warehouse is just not going to be as smart as Claude 5 Sonnet. Why would you do that? Why would you run up the cost? Why would you have all the latency? And so, you know, I think for me, part of the challenge actually is really about what is that evolving governance of this machine intelligence systems where there will be thousands or millions of these models. I think some of the AI risk debate emerged from Nick Bostrom's excellent book a decade ago, where the view was you'd have a singleton, right?

49:04You'd have a solitary, all-powerful AI, rather than many, many thousands or millions or tens of millions of them, all interacting in slightly different ways, which I think creates a different control problem, to be honest. I completely agree. I think being able to continue to iteratively deploy improvements to Cloud, I think, is valuable because we can sort of see where it's headed. We can sort of identify problems if the next model has an issue. People can complain to us. We can discuss it. To the extent that it's that important, we can kind of discuss it as a society and then we can remedy it.

49:41So I agree. Superintelligence is not going to be one specific moment where we hit superintelligence. It is very much a continuum. AI is evolving in that way. And yeah, I agree. We're going to be using AI systems that are not very sophisticated basically as tools forever. I mean, yeah, there's no reason to use an expensive model for OCR. So I agree. We're going to have this ecosystem. It's a good point that this ecosystem itself, especially if AI is interacting with AI to do a lot of the work we do now, like maybe your AI doctor interacts with your AI pharmacist and negotiates with like the AI health insurance.

50:17To deny you your cover, no doubt. But yeah, something's going to change, right? Yeah. Yeah. I think that like that ecosystem could have a lot of problems that maybe Claude is safe then it's aligned, but the ecosystem has problems. And that's something that is so new that no one's really studied it, but it's something to worry about. I think there's this general worry that the way that AI development will cause harm eventually is that things kind of go off the rails. In the modern world, I don't understand how most things work. Like, do I understand how my car works? Do I understand how my iPad works?

50:50So when you get more and more systems that no one really understands, things can kind of go off the rails in ways that are really hard to predict or kind of due to the interaction of the ecosystem. So that's definitely a risk. That is another conversation, though. I mean, and you've tickled my brain now and there's so many places to explore that one. You've been super generous with your time. I wanted to ask just one last question, really, which is about the things that we should be excited about. So if you look out over this next 12 months and you look at Claude, you look at Anthropik, what are you most excited about?

51:22I'll say something that has nothing to do with Claude. I think I'm most excited about getting what I described in terms of sort of scalable supervision of AI helping us to sort of make sure that AI is useful, is doing good things for us. I'm very excited about kind of getting that really working, going beyond constitutional AI, because I think that's really the lever for becoming more confident that we can continue to improve the capabilities of AI and make it more useful and make it beneficial for all of us. Well, I look forward to that as well. I will be stopping this call and going straight into Claude and asking about all the questions I should have asked you.

52:00Jared Kaplan, thanks so much this morning for chatting to us. Thanks so much for having me. It was a lot of fun.

From the publisher

Anthropic's co-founder and chief scientist Jared Kaplan discusses AI's rapid evolution, the shorter-than-expected timeline to human-level AI, and how Claude's "thinking time" feature represents a new frontier in AI reasoning capabilities.

In this episode you'll hear:

  • Why Jared believes human-level AI is now likely to arrive in 2-3 years instead of by 2030
  • How AI models are developing the ability to handle increasingly complex tasks that would take humans hours or days
  • The importance of constitutional AI and interpretability research as essential guardrails for increasingly powerful systems

Our new show 

This was originally recorded for "Friday with Azeem Azhar", a new show that takes place every Friday at 9am PT and 12pm ET on Exponential View. You can tune in through my Substack linked below. The format is experimental and we'd love your feedback, so feel free to comment or email your thoughts to our team at live@exponentialview.co.

Timestamps:

(00:00) Episode trailer

(01:27) Jared's updated prediction for reaching human-level intelligence

(08:12) What will limit scaling laws?

(11:13) How long will we wait between model generations?

(16:27) Why test-time scaling is a big deal

(21:59) There’s no reason why DeepSeek can’t be competitive algorithmically

(25:31) Has Anthropic changed their approach to safety vs speed?

(30:08) Managing the paradoxes of AI progress

(32:21) Can interpretability and monitoring really keep AI safe?

(39:43) Are model incentives misaligned with public interests?

(42:36) How should we prepare for electricity-level impact?

(51:15) What Jared is most excited about in the next 12 months

Jared's links:

Azeem's links: 

Produced by supermix.io


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from Azeem Azhar's Exponential View

All 44 episodes
Are we ready for human-level AI by 2030? Anthropic's co-founder answersAzeem Azhar's Exponential View · 52 min
Listen in VO