Just How Fast is AI Evolving?

26 Jan 2025 · 15 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The AI Daily Brief: Episode Summary

Episode Title

Just How Fast is AI Evolving?

Podcast Description A daily news analysis show exploring the multifaceted impact of artificial intelligence (AI) on creativity, industries, and the ethical implications of advanced AI technologies.

Episode Overview This episode discusses the rapid evolution of AI, focusing on advancements in reasoning models and the implications for society. The discussion is rooted in a piece by Professor Ethan Mollick, titled "Prophecies of the Flood," which contemplates the approaching timeline of artificial general intelligence (AGI) and the transformative effects it could have.

---

Key Concepts and Discussions

  1. Rising Tide of AI Capabilities
  2. AGI Definition: Machines capable of outperforming expert humans in most intellectual tasks are becoming more imminent.
  3. Urgency from AI Researchers: There is a growing narrative among researchers emphasizing the quick arrival of super-smart AI systems.
  1. Skepticism Towards Predictions
  2. Motivations of AI Labs: Researchers may exaggerate timelines to attract funding or boost their significance.
  3. Inconsistency of Current AI: Large language models exhibit both strong performance and significant limitations, complicating predictions of their future capabilities.
  1. Adoption and Integration Challenges
  2. Speed of Human Adaptation: Even if AGI were achieved, the slow pace of organizational and societal adaptation to new technologies could delay its impacts.
  3. Theoretical vs. Practical Implementation: Current technologies must find relevant applications in the real world, which is often a slow process.
  1. Significant AI Benchmarks
  2. Three major benchmarks highlight the advancements:
  3. Graduate-Level Google Proof Q&A Test (GPQA): OpenAI's O3 model scored 87%, surpassing expert humans.
  4. Frontier Math Problems: O3 achieved a score of 25%, a significant improvement.
  5. ARC-AGI Test: O3 scored 87.5%, outperforming previous AIs and human averages.
  1. Implications of Narrow AI Agents
  2. Emerging Agentic Systems: There are early examples of AI systems capable of autonomous action, like Google's Gemini, which can conduct complex research.
  3. Limitations of Current Agents: While capable, these systems often lack depth and nuance compared to human-produced content.
  1. Transformation of Knowledge-Based Tasks
  2. The rapid advancement of AI capabilities poses the possibility of transforming many knowledge-based professions. However, the effective integration and ethical deployment of these technologies remain crucial considerations.

---

Key Takeaways

  • AI Evolution is Accelerating: Current advancements may indicate a more rapid approach to AGI than previously thought, but skepticism about timelines and practical implementation remains.
  • Preparation is Essential: Organizations must begin discussions about the implications of AI now, rather than waiting for widespread adoption to address potential challenges.
  • The Role of Society: The transition to an AI-integrated society necessitates collective input from various stakeholders, not just developers.

---

Conclusion The podcast emphasizes the urgency of preparing for the societal changes that AI will bring, advocating for proactive discussions and adaptations to embrace the evolving landscape of artificial intelligence. This is particularly critical as AI capabilities continue to advance and gain accessibility in the coming years.

---

Additional Information

  • Sponsors: KPMG, Vanta, and Superintelligent provide insights into the practical applications of AI and the importance of compliance and readiness.
  • Engagement: Listeners are encouraged to join the Discord community for ongoing discussions about AI developments.

---

For further details, subscribe to [The AI Daily Brief](https://pod.link/1680633614) and explore the newsletter and Discord community links provided in the episode description.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today on the AI Daily Brief, just how fast is AI evolving? The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. To join the conversation, follow the Discord link in our show notes.

0:18Hello, friends. Back with another Long Reads episode of the AI Daily Brief. Today, we turn once again to Professor Ethan Mollick's One Useful Thing blog, reading a piece from now a couple of weeks ago called Prophecies of the Flood, What to Make of the Statements of the AI Labs. The setup and context of the piece is something that we've been talking about a lot on this show for the last month or so, which is the sense that the tide is rising to continue the water analogy, that AGI is getting closer and closer, the capabilities are increasing, and that something big is on the horizon. As usual, what we're going to do today is turn it over to an Eleven Labs version of myself to read the piece, and then I will come back and give you some thoughts of my own to close it out.

0:59Prophecies of the Flood Recently, something shifted in the AI industry. Researchers began speaking urgently about the arrival of super-smart AI systems, a flood of intelligence, not in some distant future, but imminently. They often refer to AGI, artificial general intelligence, defined, albeit imprecisely, as machines that can outperform expert humans across most intellectual tasks. This availability of intelligence on demand will, they argue, change society deeply and will change it soon. There are plenty of reasons to not believe insiders as they have clear incentives to make bold predictions.

1:31They're raising capital, boosting stock valuations, and perhaps convincing themselves of their own historical importance. They're technologists, not profits, and the track record of technological predictions is littered with confident declarations that turned out to be decades premature. Even setting aside these human biases, the underlying technology itself gives us reason for doubt. Today's large language models, despite their impressive capabilities, remain fundamentally inconsistent tools. Brilliant at some tasks while stumbling over seemingly simpler ones. This jagged frontier is a core characteristic of current AI systems, one that won't be easily smoothed away.

2:05Plus, even assuming researchers are right about reaching AGI in the next year or two, they are likely overestimating the speed at which humans can adopt and adjust to a technology. Changes to organizations take a long time. Changes to systems of work, life, and education are slower still. And technologies need to find specific uses that matter in the world, which is itself a slow process. We could have AGI right now and most people wouldn't notice. Indeed, some observers have suggested that has already happened, arguing that the latest AI models like CLOD 3.5 are effectively AGI 1. Yet dismissing these predictions as mere hype may not be helpful.

2:40Whatever their incentives, the researchers and engineers inside AI labs appear genuinely convinced they're witnessing the emergence of something unprecedented. demanded. Their certainty alone wouldn't matter, except that increasingly public benchmarks and demonstrations are beginning to hint at why they might believe we're approaching a fundamental shift in AI capabilities. The water, as it were, seems to be rising faster than expected. The event that kicked off the most speculation was the reveal of a new model by OpenAI called O3 in late December. No one outside of OpenAI has really used this system yet, but it is the successor to O1, which is already very impressive too.

3:14The O3 model is one of the new generation of reasoners, AI models that take extra time to think before answering questions, which greatly improves their ability to solve hard problems. OpenAI provided a number of startling benchmarks for O3 that suggest a large advance over O1 and indeed over where we thought the state-of-the-art in AI was. Three benchmarks in particular deserve a little attention. The first is the called the Graduate-Level Google Proof Qanda Test, GPQA, and it is supposed to test high-level knowledge with a series of multiple-choice problems that even Google can't help you with.

3:47PhDs with access to the internet got 34 % of the questions right on this test outside their specialty, and 81 % right inside their specialty. When tested, O3 achieved 87 % beating human experts for the first time. The second is Frontier Math, a set of private math problems created by mathematicians to be incredibly hard to solve. And indeed, no AI ever scored higher than 2 % until O3, which got 25 % right. The final benchmark is ARC-AGI, a rather famous test of fluid intelligence that was designed to be relatively easy for humans, but hard for AIs. Again, 03 beat all previous AIs as well as the baseline human level on the test, scoring 87.5%.

4:25All of these tests come with significant caveats, but they suggest that what we previously considered unpassable barriers to AI performance may actually be beaten quite quickly. As AIs get smarter, they become more effective agents, another ill-defined term, see a pattern, that generally means an AI given the ability to act autonomously towards achieving a set of goals. I have demonstrated some of the early agentic systems in previous posts, but I think the past few weeks have also shown us that practical agents, at least for narrow but economically important areas, are now viable. A nice example of that is Google's Gemini with deep research, accessible to everyone who subscribes to Gemini, which is really a specialized research agent.

5:03I gave it a topic like, research a comparison of ways of funding startup companies from the perspective of founders for high-growth ventures. And the Agentic system came up with a plan, read through 173 websites, and compiled a report for me with the answer a few minutes later. The result was a 17-page paper with 118 references. But is it any good? I've taught the introductory entrepreneurship class at Wharton for over a decade, published on the topic, started companies myself, and even wrote a book on entrepreneurship, and I think this is pretty solid. I didn't spot any obvious errors, but you can read it yourself if you would like here.

5:36The biggest issue is not accuracy, but that the agent is limited to public non-paywalled websites and not scholarly or premium publications. It also is a bit shallow and does not make strong arguments in the face of conflicting evidence. So not as good as the best humans, but better than a lot of reports that I see. Still, this is a genuinely disruptive example of an agent with real value. Researching and report writing is a major task of many jobs. What deep research accomplished in three minutes would have taken a human many hours, though they might have added more nuanced analysis. Anyone writing a research report should probably try deep research and see how it works as a starting place, even though a good final report will still require a human touch.

6:15I had a chance to speak with the leader of the Deep Research Project, where I learned that it is just a pilot project from a small team. I thus suspect that other groups and companies that were highly incentivized to create narrow but effective agents would be able to do so. Narrow agents are now a real product rather than a future possibility. There are already many coding agents, and you can use experimental open-source agents that do scientific and financial research. Narrow agents are specialized for a particular task, which means they are somewhat limited. That raises the question of whether we soon see generalist agents where you can just ask the AI anything, and it will use a computer and the internet to do it.

6:50Simon Willison thinks not, despite what Sam Altman has argued. We will learn more as the year progresses, but if general agentic systems work reliably and safely, that really will change things, as it allows smart AIs to take action in the world. Agents and very smart models are the core elements needed for transformative AI, but there are many other pieces as well that seem to be making rapid progress. This includes advances in how much AIs can remember, context windows, and multimodal capabilities that allow them to see and speak. It can be helpful to look back a little to get a sense of progress.

7:21For example, I have been testing the prompt Otter on a Plane Using Wi-Fi for image and video models since before ChatGPT came out. In October 2023, that prompt got you this terrifying monstrosity. Less than 18 months later, multiple image creation tools nail the prompt. The result is that I have had to figure out something more challenging. This is an example of benchmark saturation, where old benchmarks get beaten by the AI. I decided to take a few minutes and see how far I could get with Google's VO2 video model in producing a movie of the otter's journey. The video you see below took less than 15 minutes of active work, although I had to wait a bit for the videos to be created.

7:57Take a look at the quality of the shadows and light. I especially appreciate how the otter opens the computer at the end. And to up the ante even further, I decided to turn the saga of the otter into a 1980s-style science fiction anime featuring otters in space and a period-appropriate theme song, thanks to Suno. Again, very little human work was involved. Given all of this, how seriously should we take the claims of the AI labs that a flood of intelligence is coming, even if we only consider what we've already seen, the O3 benchmarks shattering previous barriers, narrow agents conducting complex research, and multimodal systems creating increasingly sophisticated content.

8:34We're looking at capabilities that could transform many knowledge-based tasks. And yet the labs insist this is merely the start, that far more capable systems and general agents are imminent. What concerns me most isn't whether the labs are right about this timeline. it's that we're not adequately preparing for what even current levels of AI can do, let alone the chance that they might be correct. While AI researchers are focused on alignment, ensuring AI systems act ethically and responsibly, far fewer voices are trying to envision and articulate what a world awash in artificial intelligence might actually look like.

9:06This isn't just about the technology itself. It's about how we choose to shape and deploy it. These aren't questions that AI developers alone can or should answer. They're questions that demand attention from organizational leaders who will need to navigate this transition, from employees whose work lives may transform, and from stakeholders whose futures may depend on these decisions. The flood of intelligence that may be coming isn't inherently good or bad. But how we prepare for it, how we adapt to it, and most importantly, how we choose to use it, will determine whether it becomes a force for progress or disruption.

9:36The time to start having these conversations isn't after the water starts rising. It's now. Today's episode is brought to you by Vanta. Trust isn't just earned, it's demanded. Whether you're a startup founder navigating your first audit or a seasoned security professional scaling your GRC program, proving your commitment to security has never been more critical or more complex. That's where Vanta comes in. Businesses use Vanta to establish trust by automating compliance needs across over 35 frameworks like SOC 2 and ISO 27001. Centralized security workflows complete questionnaires up to 5x faster and proactively manage vendor risk.

10:14Vanta can help you start or scale up your security program by connecting you with auditors and experts to conduct your audit and set up your security program quickly. Plus, with automation and AI throughout the platform, Vanta gives you time back so you can focus on building your company. Join over 9 ,000 global companies like Atlassian, Quora, and Factory who use Vanta to manage risk and prove security in real time. For a limited time, this audience gets$1 ,000 off Vanta at vanta.com slash nlw. That's v-a-n-t-a dot com slash nlw for$1 ,000 off. If there is one thing that's clear about AI in 2025, it's that the agents are coming.

10:54Vertical agents by industry, horizontal agent platforms, agents per function. If you are running a large enterprise, you will be experimenting with agents next year. And given how new this is, all of us are going to be back in pilot mode. That's why Superintelligent is offering a new product for the beginning of this year. It's an agent readiness and opportunity audit. Over the course of a couple quick weeks, we dig in with your team to understand what type of agents make sense for you to test, what type of infrastructure support you need to be ready, and to ultimately come away with a set of actionable recommendations that get you prepared to figure out how agents can transform your business.

11:31If you are interested in the agent readiness and opportunity audit, reach out directly to me, nlw at bsuper.ai, put the word agent in the subject line so I know what you're talking about, and let's have you be a leader in the most dynamic part of the AI market. Hello, AI Daily Brief listeners. Taking a quick break to share some very interesting findings from KPMG's latest AI quarterly pulse survey. Did you know that 67 % of business leaders expect AI to fundamentally transform their businesses within the next two years? And yet it's not all smooth sailing. The biggest challenges that they face include things like data quality, risk management, and employee adoption.

12:08KPMG is at the forefront of helping organizations navigate these hurdles. They're not just talking about AI, they're leading the charge with practical solutions and real world applications. For instance, over half of the organizations surveyed are exploring AI agents to handle tasks like administrative duties and call center operations. So if you're looking to stay ahead in the AI game, keep an eye on KPMG. They're not just a part of the conversation, they're helping shape it. Learn more about how KPMG is driving AI innovation at kpmg.com slash US. All right, back to the real non-AI NLW here. As usual, Ethan does a great job, I think, of summing up a lot of what's going on as well as a lot of the sentiment out there.

12:46It has definitely been the case that a vibe has shifted. The labs seem more and more comfortable and even eager to talk about how quickly AGI is coming. Reasoning models are the watchword of the moment. And there have also been some advances that make it feel like not only are these things coming, but they're likely to be very widely accessible. The Chinese model DeepSeq, which has everyone in such a tizzy here because of how close it performs to open AI models at a tiny fraction of the cost, has everyone thinking even more about what the implications of incredibly cheap and abundant intelligence really are.

13:18Also, since this piece was released, we got the release of OpenAI's Operator, which, while still limited in what it can do, is the sort of generalist agent that Ethan is talking about. There was an interesting interview with venture capitalist Chris Saka earlier this week with Tim Ferriss, where Saka became the latest person to articulate just how disruptive this wave of new intelligence and cheap and abundant intelligence could really be when it comes to people's jobs and livelihoods. As I've said before, I think that this transition, while hugely full of potential, will require nothing less than a total re-evaluation of the social contract.

13:51A new way of thinking about work, a new way of thinking about expectations, a new way of thinking about how we judge our own value, and so much more. I don't know if things are moving faster or if it just feels like it. I do think that things that have been theoretical for some number of years are now moving into production. I think that this year we're going to see more and more people actually deploying agents in a way that makes the assistant era of AI look quaint. And I agree with Ethan wholeheartedly that the time to be having these conversations about what we want out of a society that has AI embedded is now.

14:20I don't think we're turning back the tide, but that doesn't mean that we have no agency in the world that's being created. Big ponderous thoughts for your weekend. And with that, we will close the AI Daily Brief. Appreciate you listening as always. And until next time, peace. Thank you.

From the publisher

Between reasoning models and agents, things feel like they're heating up. But just how fast are things really going? A reading and discussion inspired by https://www.oneusefulthing.org/p/prophecies-of-the-flood


Brought to you by:

KPMG – Go to ⁠⁠⁠⁠⁠www.kpmg.us/ai⁠⁠⁠⁠⁠ to learn more about how KPMG can help you drive value with our AI solutions.

Vanta - Simplify compliance - ⁠⁠⁠⁠⁠⁠⁠https://vanta.com/nlw

The Agent Readiness Audit from Superintelligent - Go to https://besuper.ai/ to request your company's agent readiness score.

The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614 Subscribe to the newsletter: https://aidailybrief.beehiiv.com/ Join our Discord: https://bit.ly/aibreakdown

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
Just How Fast is AI Evolving? The AI Daily Brief: Artificial Intelligence News and Analysis · 15 min
Listen in VO