In short
Podcast Notes: The AI Daily Brief - Episode: Claude Sonnet 4.5 Can Code Autonomously for 30 Hours 🤯
Episode Summary In this episode, the host discusses the impressive capabilities of Anthropics’ Claude Sonnet 4.5, which can autonomously code for up to 30 hours, significantly surpassing previous benchmarks. The episode also covers recent news from OpenAI regarding their upcoming model Sora 2 and its associated AI-only video app, alongside discussions on AI job impacts and regulatory news from California.
---
Key Topics & Discussions
- Claude Sonnet 4.5 Autonomy
- Capability: Claude Sonnet 4.5 can autonomously code for 30 hours without interruption, creating a chat app similar to Slack.
- Significance: This advancement marks a major leap in AI autonomy, indicating the potential for AI to handle complex, long-term tasks effectively.
- Key Innovations:
- Enforced modular artifacts for managing longer segments of code.
- Persistent memory surfaces and planning loops that help maintain context and continuity during coding tasks.
- Runtime constraints that enhance operational efficiency.
- OpenAI's Upcoming Releases
- Sora 2 Model: Expected to launch soon, with claims that it produces video indistinguishable from real life.
- AI Video App: A TikTok-style application for AI-generated content, restricting uploads to only AI-generated videos.
- Concerns & Reactions:
- Discussions around the implications of AI-generated content versus traditional media.
- Debate over copyright arrangements for content generated using the app, highlighting the complexities of intellectual property rights in the AI realm.
- Job Market Impact
- Lufthansa Layoffs: Announcement of 4,000 job cuts by 2030, primarily in administrative roles, attributed to digitalization and AI-driven efficiency improvements.
- Industry Trends: Acknowledgment that AI is reshaping workforce structures and roles, prompting companies to reassess staffing needs.
- Regulatory Updates
- California AI Safety Bill SB 53: Recently signed into law, requiring AI companies to disclose safety protocols and risks associated with their models.
- Industry Reactions:
- Mixed responses from tech companies: support from Anthropic, caution from Google and OpenAI, and a neutral stance from Meta.
---
Key Takeaways
- The advancements in Claude Sonnet 4.5 highlight the escalating capabilities of AI in coding, showcasing a significant step toward greater autonomy in AI systems.
- OpenAI's upcoming products could substantially influence public engagement with AI-generated media, raising important questions about content ownership and quality.
- Companies are increasingly recognizing the necessity of integrating AI into their operations, leading to strategic workforce adjustments.
- Regulatory measures like California's SB 53 may shape how AI technologies are developed and deployed, balancing innovation with safety concerns.
---
Conclusion The episode underscores the rapid advancements in AI technology, exemplified by Claude Sonnet 4.5’s coding capabilities. As companies adapt to these changes, both opportunities and challenges arise in areas such as employment and regulatory frameworks. The continuous evolution of AI tools promises to reshape the tech landscape, prompting ongoing discussions around ethics, safety, and the future of work.
---
*For more insights, subscribe to The AI Daily Brief on your preferred podcast platform.*
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the AI Daily Brief, Anthropics new Sonnet 4.5 model can apparently code independently for up to 30 hours. We're going to talk about what that means for the state of AI autonomy. And before that, in the headlines, OpenAI is apparently not only about to launch Sora 2, but an AI-only TikTok-style video app as well. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
0:29All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Robots & Pencils, Notion and super intelligent. To get an ad-free version of the show, go to patreon.com slash AI Daily Brief. And if you are interested in sponsoring the show, set us a note at sponsors at AI Daily Brief dot AI to find out about all the opportunities. This truly is the smartest, most engaged, and most high power AI audience in the world. So if you are interested in accessing that, please do reach out. And with that, let's dive in. Welcome back to the AI Daily Brief Headlines Edition, all the daily AI news you need in around five minutes.
1:03We kick off today with the latest rumors out of OpenAI, where that company is expected to launch not only their next generation video model, Sora 2, but also a social app for AI-generated video to go alongside it. There is actually a lot to unpack here. This is more than just a model release, so let's dig in. Sources speaking with the Wall Street Journal said that the model and its companion app would be arriving in the coming days. In fact, you might have noticed this set of new commercials that OpenAI dropped yesterday, seemingly as a marketing campaign, some think that they are actually doing double duty, not just advertising chat GPT, but sneakily showing off the video generation capabilities of the new Sora.
1:43Right, smart Kretschmann? It certainly doesn't look generated, but some are arguing that Sora 2 could be indistinguishable from real video. Some broke down the video frame by frame and thought they found evidence it was AI generated. And what people are excited about is that each video appears to show camera motion that would be difficult bordering on impossible even if you were using a drone. In the second ad showing a man cooking Italian food for his date, we zoom out from an extreme close-up through a cluttered kitchen, out the window, and across the street. Summing up the feelings of many, software engineer Jolson Ribello wrote, if this really is Sora 2, nothing will be the same anymore.
2:18Now, the scoop from Wired is that alongside the new Sora, OpenAI is poised to release a short-form video app powered by the new model. The app, which features a vertical video feed with swipe-to-scroll navigation, appears to closely resemble TikTok, except all of the content is AI-generated. There's a For You-style page powered by a recommendation algorithm. On the right side of the feed, a menu bar gives users the option to like, comment, or remix a video. The app reportedly does not allow users to upload photos or videos, with the only source of content being Sora 2. However, users will be able to verify their likeness and have themselves appear in generated clips.
2:52Other users can also generate clips featuring verified likenesses, but users will receive a notification when clips featuring them are generated. Wired continued, OpenAI appears to be betting that the Sora 2 app will let people interact with AI-generated video in a way that fundamentally changes their experience of the technology, similar to how ChatGPT helped users realize the potential of AI-generated text. And apparently, this is more than just showing off a technology, but also a business recognition of the opportunity of the moment. Wired continues, Internally, sources say, there's also a feeling that President Trump's on-again, off-again deal to sell TikTok's U.S.
3:25operations has given OpenAI a unique opportunity to launch a short-form video app, particularly one without close ties to China. Now, not everyone is thrilled about this. In fact, there has been an explosive conversation around the lamentability of short-form brain rot ever since Meta announced its Vibes feed last week. You'll remember I did a whole segment of an episode that was all about how much people did not like the idea of Meta having an AI video-only feed. Now, as OpenAI apparently gets ready to release something like this, we'll get to see how much of that was about the format versus how much of that was people just not liking meta.
3:59Interestingly, in the wake of all of that controversy, OpenAI insider Rune had posted, there is a moral panic around short form video content, in my opinion. He added that he understands the concern, but that he's just not certain that hours on TikTok are meaningfully different to hours in front of the TV. He said, I basically agree with Postman on the nature of video and its corrupting influence on running a civilization well as opposed to text-based media. I'm just not sure that it's so much worse than being glued to your TV, and I'm definitely not sure that AI slop is worse than human slop.
4:26Ahman Osman noted that Rune appeared to be breaking the narrative on X a few days ahead of this key announcement from OpenAI. With the report that we're getting this app from OpenAI, Ahman said, bro was running narrative ops. Now, what other interesting dimension of the story has to do with the copyright arrangements that OpenAI will be putting in place? Sources said that the company has begun notifying talent agencies and studios about the product over the past week. The communication notified rights holders that they will need to explicitly opt out. Otherwise, their intellectual property will be included in the generated videos.
4:56With the small nuance being that recognizable public figures won't appear without explicit permission, but fictional characters will require an op-out. The debate around that could be an episode all on its own. But look, I think that we are very close to actually getting this app, so I'm going to pause it here. We will come back and talk about all the implications when we see what the thing actually is and we get people's actual first reactions to it. Moving on to our next story, more layoff news seemingly related to AI. German airline Lufthansa said they would be eliminating the equivalent of 4 ,000 full-time roles by 2030.
5:26That's around 4 % of their 102 ,000-strong workforce. However, this is a highly targeted downsizing, with Lufthansa aiming to make the cuts primarily from their 10 ,000 administrative roles. The layoffs are also a sharp change in direction, as Lufthansa stated that they would be adding 10 ,000 new hires over the course of the year back in January. In a press release, the company said, The Lufthansa Group is reviewing which activities will no longer be necessary in the future, for example, due to duplication of work. In particular, the profound changes brought about by digitalization and the increased use of artificial intelligence will lead to greater efficiency in many areas and processes.
5:58Like Accenture's downsizing announcement last week, this doesn't appear to be a case of a struggling company using AI to mask around a belt tightening. After a troubled year in 2024 where operating margins dropped to 4.4%, Lufthansa has guided that they expect margins to reach 10 % by 2028, up from their strategic target of 8%. They also expect to see 2.5 billion euros of free cash flow by that date. The stock was up 0.9 % on the news to bolster a year-to-date gain of 25. Nat Lufthansa claims that this is purely a restructuring effort to get ahead of reduced workforce needs as they accelerate AI adoption.
6:31And we are certainly going to be keeping an eye to see whether this type of announcement, i.e. forward telegraphing of AI shifts, becomes a trend. Lastly today, a bit of regulatory news. California Governor Gavin Newsom has signed AI Safety Bill SB 53. The bill is a watered-down version of last year's SB 1047, which was vetoed by Newsom in September. SB 53 requires leading AI companies to report the safety protocols they use in producing models and disclose the highest-degree risks posed by the models. The law is largely concerned with catastrophic risks like aiding in bioweapons production or facilitating mass casualty events.
7:04In addition, the law strengthens whistleblower protections for employees of AI labs. California State Senator Scott Wiener, the chief sponsor of the bill, said, this is a groundbreaking law that promotes both innovation and safety. The two are not mutually exclusive, even though they are often pitted against each other to be. Now, last year's SB 1047 featured a huge amount of very public pushback, while this process has been quite a bit quieter by comparison. Anthropic came out in favor of the bill while Google and OpenAI opposed it. Meta was on the fence, not endorsing the bill, but giving Newsom a soft green light to sign it.
7:34And there are still some concerns. Colin McKeown, the head of government affairs at Andreessen Horowitz posted, we're fighting for a national AI strategy that gives little tech a fair shot and keeps the U.S. in the lead. California's AI Bill SB 53 includes some thoughtful provisions that account for the distinct needs of startups, but it misses an important mark by regulating how the technology is developed, a move that risks squeezing out startups, slowing innovation, and entrenching the biggest players. As well as railing against the idea of state-by-state regulation, Colin argued that the rules should govern how AI models are used rather than how they are trained.
8:03Still, the bill was drafted explicitly as a compromise. Senator Weiner worked with California's Joint California Policy Working Group on AI Frontier Models, which was set up last year following the veto. That group was chaired by Dr. Fei-Fei Li and includes numerous industry stakeholders. Overall, Newsom said that in passing the bill, quote, California has proven that we can establish protections to protect our communities while also ensuring that the growing AI industry continues to thrive. This legislation strikes that balance. A last note before we move over to our main episode, yesterday was one of those days where we had two very distinct big stories that could easily be a main all on their own.
8:38The first, which is what I went with, is all about Claude 4.5 and the expansion of the autonomy frontier for agents. But there is a ton about agentic commerce and OpenAI's new checkout feature in ChatGPT that really deserves its own space as well. I decided that rather than crowding that into the headlines today, the plan is currently for it to be the main episode for tomorrow. Although if we get that Sora 2 app or something new and big, who knows? Suffice it to say that sometime this week we will get into all of that. For now though, that's going to do it for today's actual headlines. Next up, the main episode.
9:10What if AI wasn't just a buzzword, but a business imperative? On You Can With AI, we take you inside the boardrooms and strategy sessions of the world's most forward-thinking enterprises. Hosted by me, Nathaniel Whittemore, and powered by KPMG, this seven-part series delivers real-world insights from leaders who are scaling AI with purpose. From aligning culture and leadership to building trust, data readiness, and deploying AI agents. Whether you're a C-suite executive, strategist, or innovator, this podcast is your front row seat to the future of enterprise AI. So go check it out at www.kpmg.us slash AI podcasts, or search You Can With AI on Spotify, Apple Podcasts, or wherever you get your podcasts.
9:53AI changes fast. You need a partner built for the long game. Robots and pencils work side by side with organizations to turn AI ambition into real human impact. As an AWS-certified partner, they modernize infrastructure, design cloud-native systems, and apply AI to create business value. And their partnerships don't end at launch. As AI changes, Robots & Pencils stays by your side, so you keep pace. The difference is close partnership that builds value and compounds over time. Plus, with delivery centers across the U.S., Canada, Europe, and Latin America, clients get local expertise and global scale.
10:25For AI that delivers progress, not promises, visit robotsandpencils.com slash AI Daily Brief.
10:56It can now build documents from your entire company's knowledge base, organize scattered information into organized reports, basically do tasks that used to take days, and get them complete in minutes. These agents don't just help with work, they finish it. Getting started with building on Notion is easier than ever. Notion agents are now your very own super user to help you onboard in minutes. Your AI teammates are ready to work. Try Notion AI for free at the link in our show notes. Today's episode is brought to you by Superintelligent. Now, one thing that we are having a lot of conversations with folks about is the fact that for some of you, your fiscal year is coming to an end.
11:29And that means two things. One, it means planning and thinking about what you're going to do in the next year. And two, it means using up those last of budgets so you don't lose them. If you are an enterprise that happens to find yourself in that situation, Superintelligent would love to help on both fronts. We are moving increasingly towards an annual AI planning model where we map out how you can create an action map of your organization's agent opportunities that represents an executable backlog of AI and agent use cases that you can deliver on over the course of the next year. Additionally, for those end-of-year budgets, we have worked out deals with a number of partners where we can pre-lock in general implementation packages even before you've figured out exactly what use cases are going to require them.
12:10If you'd like to learn more about superintelligence agent readiness audits and this new end-of-fiscal-year plan, visit us at bsuper.ai, click Get Started, and make sure to use the word Fiscal somewhere in the description. Welcome back to the AI Daily Brief. Today we are talking about a much-anticipated model release in the form of Claude Sonnet 4.5. Now, on the one hand, people have been excited about Anthropic releasing their latest Claude 4.5 model in general, but really when push comes to shove, the coding implications of Sonnet 4.5 are what people have been most focused on. Today we're going to talk about the response to that model, an interesting new user experience that came with it, and about how our sense of the autonomy frontier might be fundamentally off as this thing apparently has coded for up to 30 hours completely autonomously.
12:57First up though, let's talk about what was announced. No surprise, Anthropic decided to focus the announcement on the coding implications. In fact, in their opening tweet, they call it the best coding model in the world. Now, if you are a regular listener, you'll know that all the way back since 3.5, Claude really has been for most of that time the preferred set of models when it comes to coding use cases. The only exception to that really has been in the last month or so, where GPT-5 and OpenAI's codecs have started to win back market share from the Claude models, both because of the gains of GPT-5, but also because of some issues during August with model performance on Anthropics' side.
13:31Sonnet 4.5 is very much Anthropics' attempt to reclaim that crown. They write, it's the strongest model for building complex agents, it's the best model at using computers, and it shows substantial gains on testing of reasoning and math. Benchmarks, as you know, are one of my least favorite ways to understand a new model, but the published benchmarks do show some big jumps, especially when it comes to these coding use cases. For example, on SweetBench Verified, they're up to 77.2 % raw, as opposed to GPT-5 Codex's 74.5%, and all the way up to 82 % with what they call parallel test time compute.
14:03On the TerminalBench benchmark for agentic terminal coding, they claim 50 % as opposed to GPT-5's 43.8%. And basically all of the other benchmarks put them in and alongside the Opus 4.1 and GPT-5 class models of the world. The company did also announce a number of upgrades to Cloud Code itself. The first is the Cloud Agent SDK, which basically gives users access to the tools, context management systems, and permissions frameworks that are embedded in Cloud Code. They've also got an updated terminal interface and a new VS Code extension, so people can work with Claude code in their IDE instead.
14:36They also added this little checkpoints feature, which is getting punted aside based on all the other news, but as they put it, lets you instantly undo Claude's latest changes, which seems like a super valuable feature for any sort of agentic coding use case. Still, the big show is of course this new model, and that's what everyone was focused on. And as tends to happen with a new model, there is some variety in the first impressions. While I didn't see anyone that had an outright bad experience with it, there were certainly some meh type of shoulder shrug experiences. Jeremy Mack writes, Early results for Sonnet 4.5.
15:08Code quality not markedly different than 4. CSS is improved. Outputting markdown when not asked. Same price in TPS as always. Gosu Coder writes, First impression of 4.5, keep in mind this is after three hours of head-down coding, so still early. One, I don't think I can see a difference versus 4.0. In fact, if you told me this was actually 4.0, I'd believe you. 3.5 and 3.7 were noticeably different. 2 still had to go back to GPT-5 for a few things that Sonnet couldn't figure out. We have definitely hit a wall in coding progress. Now, a lot of people responded that they hadn't had that same experience.
15:43Ming, for example, said that he had found that it was better at following instructions and better at parallel tool calling. And others just generally said that they were more impressed. On the other end of the spectrum, you had a lot of posts like this one from Leo Synthwave who wrote, My verdict on 4.5 Sonnet, very good vibes, very fast. Although at the same time, he also said, Thinking, which is a particular mode of this model, often doesn't seem to yield a significant improvement in output, and I still prefer codecs with GPT-5 codecs for agentic use. Tool use seemed to be a thing that Anthropic was focused on.
16:13Kim Anismus called out this section of the announcement post as related to tool usage. The model more effectively uses parallel tool calls, firing off multiple speculative searches simultaneously during research and reading several files at once to build context faster. Improved coordination across multiple tools and information sources enables the model to effectively leverage a wide range of capabilities in agentic search and coding workflows. Simon Willison did a deep dive, headlined by the statement, I think it may live up to Anthropik's claims of being the best coding model in the world for the next few weeks at least.
16:42And in his post, he definitely talked about this enhanced tool usage as one of the big upgrades. Dan Shipper and the team at Every summed up by saying that it was faster than GPT-5 Codex and smarter and more steerable than Opus 4.1. And the big thing that they noted was the speed and the performance for the cost. They said that the new Sonnet 4.5 felt about 50 % faster than previous versions of Claude. They also said that it was smarter than Opus and more than anything else, it was 5x cheaper. Dan writes, it's still the same pricing as the old Sonnet 4, so there's basically no reason to use Opus in the API anymore.
17:14Sonnet all day. Some other folks noticed benefits in areas other than coding. For example, Bindu Reddy wrote, so far, definite improvement on coding, math, and data analysis over Sonnet 4. Ethan Mollick wrote, it's a really good model. I saw especially big jumps in doing finance and statistics, which tend to get overlooked in the focus on coding. And in fact, if you go to Anthropics announcement post, the focus on finance was one of their big notes. For example, in their published benchmark for financial analysis, Sonnet 4.5 got a 55.3 % compared to, for example, GPT-5's 46.9%. Peter Wilderford wrote, everyone talking about 4.5 being great at coding, but I'm taking way more notice of that huge increase in computer use score.
17:55The jump he's noting is from 44.4 % in Opus 4.1 to 61.4 % with Sonnet 4.5 on the OS World test. Peter writes, that's a huge increase over the state of the art, and I don't think we've seen anything similarly good at OS World from others. Claude agents coming soon? Now, speaking of agents and just production use cases of these models, some of the big agentic coding companies instantly started to put this model into production. The factory team, which focuses on agentic coding for enterprises, wrote, after testing with Anthropic, we find the strengths of Sonnet 4.5 to be significantly more reliable and accurate file editing, high environmental awareness, snappier than previous models on quick questions, not overthinking simple tasks.
18:37Walden from Cognition wrote, when our team tried Sonnet 4.5, we realized it was worth building a whole new version of Devon around it. This model behaves very differently. They actually published an entire blog post about what they changed. They wrote, Because Devin is an agent that plans, executes, and iterates rather than just auto-completing code, we get an unusual window into model capabilities. Each improvement compounds across our feedback loops, giving us a perspective on what's genuinely changed. With Sonnet 4.5, we're seeing the biggest leap since Sonnet 3.6. Planning performance is up 18%, and to end eval scores up 12%, and multi-hour sessions are dramatically faster and more reliable.
19:14A couple of other notes that they shared. They write, Sonnet 4.5 is the first model we've seen that is aware of its own context window and this shapes how it behaves. As it approaches context limits, we've observed it proactively summarizing its progress and becoming more decisive about implementing fixes to close out tasks. Interestingly, they said that this context anxiety, which is their term for it, can actually hurt performance, where they've observed the model taking shortcuts or leaving tasks incomplete because it believed it was near the end of its window even if it had plenty of room left.
19:42Moritz Steffen from Cognition also noted that the model tracks all modified features and doesn't stop until they work. He writes, one particularly impressive moment was when I asked it to build a Datadog clone and it ran a log emission script in the background while using Devin's browser to test the live event ingestion UI. Now with all that, so far I haven't seen people who had switched over to GPT-5 codecs rushing to get back into the anthropic sphere. Peter Gostev writes, Hmm, definitely better than Sonnet 4, but not obviously better than GPT-5 thinking high in codex models just now. Victor Talen writes, I really like Cloud 4.5 for coding.
20:16It's fast, reliable, surgical, high quality in a good way. I think I will use it a lot, especially for style refactors and things like that. But it is nowhere near as smart as GPT-5. I wouldn't leave it alone making large changes on HVM. Yes, it sucks to wait 30 minutes for a codex refactor, but debugging AI introduced errors takes way more time than that. Peak intelligence is very important. GPT-5 is not nearly as smart as I need, and Sonnet is less smart than that. Eric Provincher had a really interesting way of putting it. He writes, I'm starting to see anthropic models as light reasoning models while open AI models are deep reasoning models.
20:49With only light reasoning, Sonnet 4.5 excels at efficient context usage to pinpoint information. Codex tool calls are bulky, and they're interspersed with reasoning tokens to test hypotheses. It craves context to understand more of the problem. Gap between GPT-5 and Sonnet 4.5 becomes apparent when you have a hot context window where no new tool calls are needed. GPT-5 can think for a few minutes on end to find a detailed complete solution, while Sonnet 4.5 is satisfied with a few seconds for a serviceable one. Deep reasoning only works with sufficient context, but allows the model to really evaluate problems so exhaustively that it appears almost superhuman.
21:23By contrast, light reasoning stays closer to the surface, but serves as breathing room for models to collect their thoughts. It is in many ways much more human. Anthropic is far and away ahead on light reasoning. Which is super interesting. I think this is a much more useful diagnostic than a simple better or worse. And once again, comes back to the idea that we live in a world where at least for the moment, the best strategy if you truly want optimal performance is going to be model switching based on different contexts and needs. Now, there are two more things that I think are really worth noting about this launch.
21:53The first is Imagine with Claude. In their announcement, Postanthropic called this a bonus research preview. They write, In this experiment, Claude generates software on the fly. No functionality is predetermined, no code is pre-written. What you see is Claude creating in real time, responding and adapting to your request as you interact. It's a fun demonstration of what Claude Sonnet 4.5 can do, a way to see what's possible when you combine a capable model with the right infrastructure. Sean Strong from Anthropic wrote a little bit more about Imagine. He said, It pioneers the concept of model as backend, using a model to not only generate interfaces on the fly, but also power all the functionality behind it.
Read the full transcript
22:30An example he gave was a choose-your-own-adventure version of his founder journey. He writes, For the prompt, I asked Claude to generate an interactive choose-your-own-adventure game based on my startup experience. It accurately retold our pivot from VR games to management, even making an interactive management dashboard and app launcher to showcase key functionality. It then had us go through our fundraise, massive growth, and ultimate shutdown due to COVID. Peter Yang asked it to, quote, show me the desktop of a bad PM on the left versus a great PM on the right. Swix from Latent Space and now Cognition validated that 4.5 is a very good coding model in general, but chose to focus on Imagine in his post about it as well.
23:06He writes, most generative UI today is no more than glorified tool calling of pre-made components. Imagine with Claude is the first mainstream adoption of the web sim paradigm that went viral last year, generating entire UIs on the fly that you can immediately use. 4.5 Sonnet enables vibe coding to be so fast and so good that you can conjure up ephemeral apps to explore the latent space of what's possible, just in time as you explore it. Now he caveats, it isn't perfect yet. Buttons and dense UIs, like simulated email clients, often don't work, or are slow enough that the illusion is gone. But it's a generation away from replacing the tyranny of designs made for the media in person, and ushering the age of truly personalized, malleable software.
23:45Josh Bickett picks up on that and writes, CloudImagine could become a new form factor for how we interact with AI. It's completely different than chat. It's like a generative computer that we talk to in a natural language. I'd guess that vision is that everyone gets their own persistent generative computer instance with a Cloud Code generating the UI, processing data and files under the hood. I'd guess that what's happening is a frontend is passing the prompt directly to a Cloud Code terminal agent, which writes back to the frontend. It looks like a beautiful feedback loop. I'm going to put in some reps with Imagine this week before the preview goes away.
24:16and will certainly share what I discover. Now, the other big thing that people were really jumping onto was the immense time that Sonnet 4.5 is apparently able to work autonomously for. Hayden Field from The Verge wrote about this in her piece about the announcement. She sums up, Anthropic's latest AI model spent 30 hours running by itself to code a chat app akin to Slack or Teams. It spat out about 11 ,000 lines of code, and it only stopped running when it had completed the task. Now, some tried to figure out how this was possible. Carlos Perez Rez writes, how is it possible that Sonnet 4.5 is able to work for 30 hours to build an app like Slack?
24:51The system prompts have been leaked and Sonnet 4.5 reveals its secret sauce. Some of the ways it accomplishes this, it forces quote unquote big code into durable artifacts. Anything over about 20 lines is required to be emitted as an artifact and only one artifact per response. He writes that gives the model a persistent append-only surface to build large apps module by module without truncation. He also points out things like it enforcing runtime constraints, governs tool loops, supports long horizon autonomy via planning and feedback loops, and ultimately, whatever the combination of things, if this is really true, it is a total game changer when it comes to the autonomy horizon that we've been working on.
25:27When Replit announced Agent 3, they shared that it had reached autonomous agent runs of 200 minutes. And a few days later, OpenAI announced their coding-optimized GPT-5 codex model, where that company said, quote, during testing, we've seen GPT-5 codecs work independently for more than seven hours at a time on large complex tasks, iterating on its implementation, fixing test failures, and ultimately delivering a successful implementation. At the time, which was literally just two weeks ago, people were saying that even that was insane. But now we've got this claim for 30 hours that just obviously blows that out of the water.
26:00And one example that Anthropic gave to really sum up and dramatize the progress that has been made in the AI coding space over just the last couple of years, is that they asked every previous version of Claude to make a clone of Claude.ai. It wasn't until 3.6 that you even had something that you could try to log into, and it wasn't until Sonnet 4 that there was even a functional clone. Now it was able to build something that actually worked, working autonomously for over five hours to do so. Nick Dobos takes a step back and points out, it's honestly insane how fast these are improving. Sweebench from 33 % to 82 % in just around a year.
26:34Part of the reason that we spend so much time on the coding use cases on this show, even though many in this audience, in fact, most of this audience are not software engineers by training, is not only that thanks to these new tools, all of us get to be software developers to some extent or another. It's that coding is so clearly the frontier where we are seeing the biggest changes take place when it comes to model capabilities. Agenda coding improvement is not just a bellwether of where models are. It's also the mechanism by which they get better at everything else as well. I'm going to be keeping a close eye to see if anyone outside of the lab setting gets anywhere close to that 30 hours of performance.
27:08But if it's true, it really is a game changer. Rohan Paul went back to a recent Axios interview with Dario Amadei, where Dario said, The vast majority of code that is used to support Claude and to design the next Claude is now written by Claude. It's just the vast majority of it within Anthropic, and other fast-moving companies the same is true. Rohan adds, Now it all makes sense. Claude Sonnet 4.5 can keep its coding focus for non-stop 30 hours. The shift has started in all of tech. Now, things move fast in this space. From the first reads, it's not even clear that Sonnet 4.5 is definitively the best coding model compared to GPT-5 Codex.
27:44And even the people who think it is are still kind of waiting to see what comes with Gemini 3. But it is yet another moment that shows the relentless pace of change in this space, and I'm excited to see what new opportunities it unlocks. For now, that's going to do it for today's AI Daily Brief. Appreciate you listening or watching as always. And until next time, peace.
From the publisher
Anthropic's Claude Sonnet 4.5 reportedly demonstrates groundbreaking autonomy by coding for up to 30 hours non-stop, significantly outpacing prior benchmarks like GPT-5 Codex’s seven-hour runs. This leap is enabled by innovations such as enforced modular artifacts, persistent memory surfaces, planning loops, and runtime constraints—transforming the way AI tackles complex, long-horizon tasks. The broader implication is that AI is now not only capable of building sophisticated applications autonomously but is also recursively engineering its own future iterations, rapidly accelerating progress across the tech landscape.
Brought to you by:
Is your enterprise ready for the future of agentic AI?
Visit AGNTCY.org
Visit Outshift Internet of Agents
Try Notion AI today with Notion 3.0 https://ntn.so/nlw
KPMG – Discover how AI is transforming possibility into reality. Tune into the new KPMG 'You Can with AI' podcast and unlock insights that will inform smarter decisions inside your enterprise. Listen now and start shaping your future with every episode. https://www.kpmg.us/AIpodcasts
Blitzy.com - Go to https://blitzy.com/ to build enterprise software in days, not months
Robots & Pencils - Cloud-native AI solutions that power results https://robotsandpencils.com/
Vanta - Simplify compliance - https://vanta.com/nlw
The Agent Readiness Audit from Superintelligent - Go to https://besuper.ai/ to request your company's agent readiness score.
The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614
Interested in sponsoring the show? nlw@aidailybrief.ai
