Why I don’t think AGI is right around the corner

3 Jul 2025 · 14 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Dwarkesh Podcast Episode Notes: Why I Don’t Think AGI is Right Around the Corner

Episode Overview In this episode, the speaker shares his thoughts on the timelines for achieving Artificial General Intelligence (AGI), reflecting on discussions with various guests on the podcast. The speaker argues that while some believe AGI is just a few years away, he holds a more skeptical view as of June 2025, highlighting the complexities and limitations of current AI systems.

Key Concepts & Arguments

  1. AGI Timelines
  2. Discussions around AGI timelines vary widely, with some guests suggesting 2 years and others suggesting 20 years.
  3. The speaker presents his perspective as of June 2025, implying a longer timeline for meaningful AGI development.
  1. Limitations of Current AI Systems
  2. Continual Learning:
  3. Current Large Language Models (LLMs) do not possess ongoing learning capabilities, limiting their effectiveness compared to human workers.
  4. The speaker emphasizes that LLMs do not improve from experience as humans do, affecting their ability to perform complex tasks.
  5. Human Value:
  6. The speaker argues that human workers excel due to their capacity for contextual understanding and incremental improvement which AI systems currently lack.
  1. Challenges in AI Improvement
  2. The current AI’s learning methods are compared to ineffective teaching methods for learning complex skills (e.g., playing the saxophone).
  3. The speaker shares personal experiences attempting to use LLMs for tasks, noting their inability to adapt and learn preferences over time.
  1. Skepticism Towards Predictions
  2. The speaker expresses skepticism regarding predictions from AI researchers about the imminent development of reliable AI agents capable of performing complex tasks autonomously.
  3. He believes that significant hurdles remain, including:
  4. Increased Complexity: The complexity of tasks increases the difficulty of developing effective AI agents.
  5. Data Limitations: The lack of sufficient multimodal data for training is a significant barrier.
  6. Algorithmic Development Timeline: Historical evidence suggests that algorithmic innovations take considerable time to develop.
  1. Future Predictions
  2. Short-term (2028):
  3. An AI capable of managing end-to-end business tasks (e.g., tax preparation) is projected but the speaker remains cautious about its feasibility.
  4. Long-term (2032):
  5. Expectations for AI to learn organically and function like a human in any white-collar task, potentially leading to a breakthrough in AGI.
  6. The speaker anticipates a "broadly deployed intelligence explosion" once continual learning capabilities are solved.
  1. Final Reflections
  2. The speaker acknowledges the potential for rapid advancements but underscores the unpredictability of timelines in AGI development.
  3. He maintains a probability distribution outlook, where both optimistic and cautious scenarios are possible.

Key Takeaways

  • The timeline for AGI is uncertain, with significant dependencies on overcoming current technical limitations.
  • Current AI systems lack the ability to learn and adapt like humans, which is a critical barrier to their integration into complex workflows.
  • Predictions about the future of AI must be tempered with an understanding of past developments and existing limitations.

Call to Action Listeners are encouraged to engage with the speaker’s blog and subscribe to the newsletter for future insights and developments in the AI landscape.

For more information, visit

[Dwarkesh Podcast](https://www.dwarkesh.com)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Okay, this is a narration of a blog post I wrote on June 3rd, 2025, titled, Why I Don't Think AGI is Right Around the Corner. Quote, Things take longer to happen than you think they will, and then they happen faster than you thought they could. Routager, Dornbush. I've had a lot of discussions on my podcast where we hack a lot our timeline to AGI. Some guests think it's 20 years away, but there's two years. Here's where my thoughts lie as of June 2025. Continual learning. Sometimes people say that even if all AI progress totally stopped, the systems of today would still be far more economically transformative than the internet.

0:36I disagree. I think that the LLMs of today are magical, but the reason that the Fortune 500 aren't using them to transform their workflows isn't because management is too stodgy. Rather, I think it's genuinely hard to get normal human -like labor out of LLMs. And this has to do with some fundamental capabilities that these models lack. I like to think that I'm AI forward here at the Thorek -Ech podcast, and I've probably spent on the order of 100 hours trying to build these LLMs tools for my post -production setup. The experience of trying to get these LLMs to be useful has extended my timelines.

1:10I'll try to get them to rewrite origin in a transcripts or readability, the way a human would, or I'll get them to identify clips from the transcript to tweet out. Sometimes I'll get them to co -write an essay with me, passage by passage. Now, these are simple, self -contained, short horizon, language in, language out tasks, the kinds of assignments that should be dead center in the LLMs repertoire, and these models are 5 out of 10 at these tasks. Don't get me wrong, that is impressive. But the fundamental problem is that LLMs don't get better over time the way a human would. This lack of continual learning is a huge, huge problem.

1:45The LLM baseline at many tasks might be higher than the average humans, but there's no way to give a model high -level feedback. You're stuck with the abilities you get out of the box. You can keep messing around with the system prompt, but in practice, this does not produce anywhere close to the kind of learning and improvement the human employees actually experience on the job. The reason the humans are so valuable and useful is not mainly their raw intelligence. It's their ability to build up context, interrogate their own failures, and pick up small improvements and efficiencies as they practice a task.

2:17How do you teach a kid to play a saxophone? Well, you have a try to blow into one and listen to how it sounds and then adjust. Now, imagine if teaching saxophone worked this way instead. A student takes one attempt, and the moment they make a mistake, you send them away and you write detailed instructions about what went wrong. Now the next student reads your notes and tries to play a Charlie Parker call. When they fail, you refine your instructions for the next student. This just wouldn't work. No matter how well honed or prompt is, no kid is just going to learn how to play saxophone from reading your instructions.

2:49But this is the only modality that we, as users, have to teach Ella Lums anything. Yes, there's RL Fine tuning, but it's just not a deliberate adaptive process the way human learning is. My editors have gotten extremely good, and they wouldn't have gone that way if you had to build bespoke RL environments for every different subtask involved in their work. They've just noticed a lot of small things themselves and thought hard about what resonates with the audience, what kind of content excites me, and how they can improve their day -to -day workflows. Now it's possible to imagine some ways in which a smarter model could build a dedicated RL loop for itself, which just feels super organic from the outside.

3:24I give some high level feedback, and the model comes up with a bunch of verifiable practice problems to RL on. Maybe even a whole environment in which to rehearse the skills that things it's lacking. But this just sounds really hard. Now I don't know how well these techniques will generalize your different kinds of tasks in feedback. Eventually the models will be able to learn on the job in this other organic way the humans can. However, it's just hard for me to see how that could happen within the next few years. Given that there's no obvious way to slot in online continuous learning into the kinds of models these LLMs are.

3:52Now LLMs actually do get kind of smart in the middle of a session. For example, sometimes I'll cool right in essay with an LLM. I'll give it an outline, and I'll ask it to draft an essay, passage by passage. All that suggestions up to paragraph four will be bad. And so I'll just rewrite the whole paragraph from scratch until it, hey, you should cite, this is what I wrote instead. And at that point, it can actually start giving good suggestions for the next paragraph. But this whole subtle understanding of my preferences and style is lost by the end of the session. Maybe the easy solution to this looks like a long rolling context window like Claude code has, which compacts the session memory into a summary every 30 minutes.

4:26I just think that titrating all this rich task that experience into a text summary will be brittle in domains outside of software engineering, which is very text -based. Again, think about the example of trying to teach somebody how to play the saxophone using a long text summary of your learnings. Even Claude code will often reverse a hard or not optimization that we engineered together before I hit slash compact. Because the explanation for why it was made didn't make it into the summary. This is why I disagree with something that Shoto and Trenton said on my podcast. And this quote is from Trenton.

4:57Even if AI progress totally stalls, you think that the models are really spiky and they don't have general intelligence. It's so economically valuable and sufficiently easy to collect data on all of these different jobs, these white color job tasks, such that to Shoto's point, we should expect to see them automated within the next five years. If AI progress totally stalls today, I think less than 25 % of white color employment goes away. Sure, many tasks will get automated. Claude for Opus can technically rewrite audited transcripts for me. But since it's not possible for me to have it improve over time and learn my preferences, I still hire a human for this.

5:35Even if we get more data without progress and continue learning, I think that we will be in a substantial use in a position with all of white color work. Yes, technically, AI might be able to perform a lot of sub tasks somewhat satisfactorily, but their inability to build up contacts will make it impossible to have them operate as actual employees at your firm. Now, while this makes you bearish on Transformer to AI in the next few years, it makes me especially bullish on AI over the next two decades. When we do solve continuous learning, we'll see a huge discontinuity in the value of these models.

6:05Even if there isn't a software only singularity with models rapidly building smarter and smarter successors systems, we might still see something that looks like a broadly deployed intelligence explosion. AI is what be getting broadly deployed through the economy during different jobs and learning while doing them in the way that humans can. But unlike humans, these models can amalgamate their learnings across all their copies. So one AI is basically learning how to do every single job in the world and AI that is capable of online learning might functionally become a super intelligence quite rapidly without any further algorithmic progress.

6:37However, I'm not expecting to see some open AI livestream where the announced that continue learning has totally been solved because labs are incentivized to release any innovations quickly. We'll see a somewhat broken early version of continue learning or test time training, whatever you want to call it before we see something which truly learns like a human. I expect to get lots of heads up before we see this big bottleneck totally solved. Computer use. When I interviewed and theropic researchers, Shultzor Douglas and Trenton Brickett on my podcast, they said that they expect reliable computer use agents by the end of next year.

7:09Now we already have computer use agents right now, but they're pretty bad. They're imagining something quite different. Their forecast is that by the end of next year, you should be able to tell an AI, go do my taxes. Go through your email, Amazon orders and Slack messages. And it emails back and forth where everybody need invoices from a compiled lawyer receipts. It decides which things are business expenses, ask for your approval on the edge cases, and then submits form 1040 to the IRS. I'm skeptical. I'm not an AI researcher, so far be it to contradict them on the technical details. But from what little I do know, here are three reasons I bet against this capability being unlocked within the next year.

7:43One, as horizon links increase, rollouts have to become long. The AI needs to do two hours worth of agentic computer use tasks before we even see if it did a right. Not to mention the computer use requires processing images and videos, which is already more compute intensive. Even if you don't factor in the longer rollouts, this seems like it just slowed down progress. Two, we don't have a large pre -training corpus of multimodal computer use data. I like this quote for mechanizes post on automated software engineering. Quote, for the past decade of scaling, we've been spoiled by the enormous amount of internet data that was freely available to us.

8:16This was enough to crack natural language processing, but not for getting models to become reliable, competent agents. Imagine trying to train GPT -4 on all the text data available in 1980. The data would be nowhere near enough, even if you had the necessary compute. End quote. Again, I'm not of the lab, so maybe text -only training already gives you a great prior on how different you eyes work and what the relationships are between different components. Maybe RL fine tuning is so sample efficient that you don't need that much data, but I haven't seen any public evidence, which makes me think that these models have certainly become less data hungry, especially in domains where they're substantially less practice.

8:52Alternatively, maybe these models are such good front end coders that they can generate millions of toy UIs for themselves to practice on. For my reaction to this, see the bullet point below. Three. Even algorithmic innovations were seen quite simple in retrospect, seem to have taken a long time to iron out. The RL procedure, which DeepSeek explained that their R1 paper seems simple at a high level, and it took two years from the launch of GPT -4 to the launch of a 01. Now, of course, I know that it's hilarious to arrogant to say that R1 or 01 were easy. I'm sure a ton of engineering, debugging, and pruning about alternative ideas was required to revive at the solution.

9:26But that's precisely my point. Seeing how long it took to implement the idea, hey, let's train our model to solve verifiable math and coding problems, makes me think that we're underestimating the difficulty of solving a much gnarlier problem of computer use, where you're operating on a totally different mortality with much less data. Reasoning. Okay, enough cold water. I'm not gonna be like one of these spoiled children on hacker news who could be handed a golden egg laying goose and still spend all their time complaining about how loud its quacks are. Have you read the reasoning traces of 03 or Gemini 2 .5?

9:55It's actually reasoning. It's bridging the problem is thinking about what the user wants. It's reacting to its own internal monologue and correcting itself when it notices that it's pursuing an unproductive direction. However, we just like, oh yeah, of course, machines are gonna go think a bunch, come up with a bunch of ideas and come back with smart answer. That's what machines do. Part of the reason some people are too pessimistic is that they haven't played around with the smartest models operating in the domains that they're most competent. Giving Claude code a vague spec and then sitting around for 10 minutes until a zero shots of working application is a wild experience.

10:27How did I do that? You could talk about circuits and training distributions and RRL and whatever, but the most proximal, concise and accurate explanation is simply that it's powered by baby artificial intelligence. At this point, part of you has to be thinking. It's actually working. We're making machines that are intelligent. Okay, so what are my predictions? My probability distribution is super wide. And I want to emphasize that I do believe in probability distributions, which means that work to prepare for a missile line 2028 ASI still makes a ton of sense. I think that's a totally possible outcome.

10:58But here are the timelines at which I'd make a 5050 bet. An AI that can do taxes end to end for my small business as well as a competent general manager could in a week, including chasing down all the receipts on different websites and finding all the missing pieces and emailing back and forth with anyone we need to hassle for invoices, filling out the form and sending it to the IRS, 2028. I think we're in the GPT -2 era for computer use, but we have no pre -training corporates and the models are optimizing for a much sparser reward of a much longer time horizon using action primitives that they're unfamiliar with.

11:28That being said, the base model is decently smart and might have a good prior over computer use tasks. Plus, there's a lot more compute and AI researchers in the world, so it might even out. Preparing taxes for a small business fuel is like for computer use, what GPT -4 was for language. And it took four years to get from GPT -2 to GPT -4. Just to clarify, I'm not saying that we won't have really cool computer use demos in 2026 and 2027. GPT -3 was super cool, but not that practically useful. I'm saying that these models won't be capable of end -to -end handling a week long and quite involved project, which involves computer use.

11:59Okay, and the other prediction is this. An AI that learns on the job is easily organically seamlessly and quickly as a human for any white collar work. For example, if I hire an AI video editor, after six months, it has as much action -able, deep understanding wave preferences or channel and what works for the audience as a human would. This I would say 2032. Now while I don't see an obvious way to slot in continuous online learning into current models, seven years is a really long time. GPT -1 had just come out this time seven years ago. It doesn't seem implausible to me that over the next seven years we'll find some way for these models to learn on the job.

12:34Okay, at this point, you might be reacting. Look, you made this huge fuss about how continual learning is such a big candy cap. But the newer timeline is that we're seven years away from what at a minimum is a broadly deployed intelligence explosion. And yeah, you're right. I'm forecasting a pretty wild world within a relatively short amount of time. AGI timelines are very log normal. It's either this decade or bust. Not really bust, more like lower marginal probability per year, but that's less catchy. AI progress over the last decade has been driven by scaling training compute for frontier systems over four X a year.

13:04This cannot continue beyond this decade, whether you look at chips, power, even the raw fraction of GDP that's used on training after 2030. AI progress has to mostly come from algorithm progress. But even there, the low heating fruits will be plugged at least under the deep learning paradigm. So the yearly probability of AGI craters after 2030. This means that if we end up on the longer side of my 50 50 beds, we might well be looking at a relatively normal world up to the 2030s or even the 2040s. But in all the other worlds, even if we stay sober about the current limitations of AI, we have to expect some truly crazy outcomes.

13:40Many of you might not be aware, but I also have a blog and I wanted to bring content from there to all of you who are mainly podcast subscribers. If you wanna read future blog posts, you should sign up for my newsletter at thewarcache .com. Otherwise, thanks for tuning in and I'll see you on the next episode.

From the publisher

I’ve had a lot of discussions on my podcast where we haggle out timelines to AGI. Some guests think it’s 20 years away - others 2 years.

Here’s an audio version of where my thoughts stand as of June 2025. If you want to read the original post, you can check it out here.



Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe

More from Dwarkesh Podcast

All 94 episodes
Why I don’t think AGI is right around the cornerDwarkesh Podcast · 14 min
Listen in VO