The good, the bad, and the future of AI agents

2 Oct 2025 · 47 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Decoder Podcast Episode Summary

Episode Title

The Good, The Bad, and The Future of AI Agents

Host

Hayden Field

Guest

David Hershey, Applied AI Team Lead at Anthropic Podcast Description Decoder is a show from The Verge about big ideas in business and technology, hosted by Nilay Patel. This episode features Hayden Field discussing the advancements and challenges of AI agents with David Hershey.

---

Key Highlights

Introduction to AI Agents and Claude Sonnet 4.5

  • Claude Sonnet 4.5: Anthropic's latest AI model, introduced as a significant breakthrough in autonomous AI agents, particularly for coding.
  • Claims that the design allows agents to operate tasks for extended periods—up to 30 hours—without human intervention.
  • The episode discusses the potential of AI agents to reshape productivity and augment human labor.

Current State of AI Agents

  • David Hershey shares his perspective on where AI agents currently stand:
  • Some areas show impressive progress, especially in coding.
  • However, there are still significant limitations in other sectors.
  • Examples:
  • AI agents excel at programming tasks but struggle with tasks involving complex user interfaces and data handling (e.g., financial spreadsheet manipulation).

Progress and Challenges in AI Agents

  • Progress:
  • Coding is currently the domain where AI agents are achieving faster advancements.
  • Consumer Applications: There’s potential for AI agents to assist in various consumer tasks, though they are not yet reliable for long, complex operations.
  • Challenges:
  • Agents remain ineffective in interpreting and interacting with complex user interfaces.
  • Difficulty in executing simple tasks (e.g., completing financial spreadsheets), highlighting the need for further development.

Surprising Areas of Adoption

  • Legal Sector: Notably, AI applications have seen rapid adoption in the legal field, surprising many due to its complexity.
  • David Hershey notes that the legal domain offers substantial work volume, justifying swift integration of AI tools.

Claude Sonnet 4.5 Features

  • Exciting Capabilities:
  • The model can autonomously develop software over sustained periods, as demonstrated by a project that involved recreating Anthropic’s own chat application, Claude.ai.
  • It can handle complex tasks like implementing features and testing applications, illustrating a leap in capability.
  • Operational Insights:
  • The model's ability to break down tasks into manageable chunks resembles effective human collaboration styles, enhancing productivity.

Limitations and Future Prospects

  • Despite advancements, Hershey notes that models need to overcome various "hills" to achieve full efficacy across various tasks.
  • Current limitations include:
  • Spatial awareness problems in games (e.g., chess, Pokémon).
  • Continued struggles with niche domains where expert feedback is essential.

AI Coding Market Insights

  • The coding sector is identified as crucial for AI development due to its potential to improve productivity significantly.
  • Hershey emphasizes that Anthropic aims to make Claude Sonnet 4.5 the leading coding model, with positive feedback from users and measurable improvements observed through testing.

Consumer Engagement and Future Directions

  • Anthropic is focused on both enterprise and consumer markets, emphasizing the importance of creating versatile models.
  • The discussion touches on how AI models should be integrated into consumer-facing applications, balancing first-party offerings with partnerships in the tech ecosystem.

Conclusion

  • David expresses optimism about the ongoing advancements in AI technology, highlighting the transformational potential of AI agents while acknowledging the work still needed to enhance their capabilities across various fields.

---

Key Takeaways

  • AI agents, particularly Claude Sonnet 4.5, represent a significant leap in autonomous capabilities, especially in coding.
  • There remains a gap between AI's potential and current performance, with many sectors still requiring improvements.
  • The legal industry shows unexpected growth in AI adoption, while coding remains a prime focus for future advancements.
  • Continuous user feedback and enhancement of models are crucial for the development of more sophisticated AI agents.

---

Further Reading

  • [Anthropic releases Claude Sonnet 4.5 in latest bid for AI agents | The Verge](#)
  • [AI agents are science fiction not yet ready for primetime | The Verge](#)

---

Credits

  • Episode produced by Kate Cox and Nick Statt.
  • Edited by Ursa Wright.
  • Music by Breakmaster Cylinder.

For feedback, reach out to decoder@theverge.com or follow Hayden Field on social media platforms. Check out the Decoder social media for more updates.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Hey there, and welcome to Decoder. I wanted to have David on because earlier this week, Anthropic released a brand new AI model called Claude Sonnet 4.5 that's been making waves. For reference, Claude is to Anthropic what ChatGPT is to open AI. This new model, Sonnet 4.5, is being built as a big breakthrough in autonomous agentic AI, especially for coding purposes, which is a big battleground in the AI market right now. All these companies want to get a slice. These types of AI products can, in theory, be given complex tasks and then go off and complete them over the course of many hours, or in some cases, even multiple days.

1:07And Anthropik says this particular model, Sonnet 4.5, can run for up to 30 hours straight without any human intervention, all while working on a singular task like building a software application from scratch. For the last year or so, companies like Anthropic, Microsoft, OpenAI, and others have been promising that this agentic technology would be the next phase of AI, the next big hype-filled thing that comes after general-purpose chatbots. They say it could really unlock generative AI's potential, and it's true they've made some strides. But as we've seen so far, agents aren't quite there yet, and they have a ways to go.

1:46Most of us are not, in fact, sending agents off on the internet to do our bidding, and we're certainly not giving them tasks that might take 12 or 24 or even 30 plus hours of autonomous work without human handholding, at least not yet. At the same time, many companies are looking at agents as the breakthrough that's supposed to unlock huge productivity gains from AI models, including the opportunity to use them to replace or augment human labor. So I wanted to sit down with David, who spends a lot of time testing out what models like Claude Sonnet 4.5 can and can't do to ask him where we are on this promise of AI agents.

2:22I wanted to talk about what these types of products are good at from a consumer standpoint beyond just programming purposes and also what the path forward looks like as AI agents progress. Okay, here's Anthropics David Hershey on the state of AI agents. Here we go.

2:52I wanted to ask you about your view of the current state of play for AI agents. We hear all the time that agents are the next big thing for generative AI, but are we still in the prototype stage, the testing phase, or what? How would you characterize what AI agents do today, right now, relative to what AI companies actually want to offer in the end, which I hear from a lot of execs is basically Jarvis from the Marvel movies. I am less confident in our ability to project the end, but I'm happy to talk about it now. I've seen agents come a long ways in the last year working with customers. And my view is there are places where we're starting to see what it looks like when agents work really well.

3:36and there are still a lot of places where they don't work really well. And I think that kind of makes it confusing. I think for some people, code is a great example. When you're writing code, and especially with Sonnet 4.5, when you watch the model spend a lot of time as an agent developing software itself, it's incredible. It can do a ton. It's gotten much better literally this week. As we release models, it's really visible and obvious how much better they're getting sort of if you're plugged into that at being an agent doing really long-running complex tasks. When you look at other sections of the economy or other like jobs that people want agents to do large pieces of, sometimes there's stuff they're still not good at.

4:16In some cases, they're not good enough at deciphering what's on a computer screen yet to be able to navigate complex UIs or whatever it is. And so they fall over and stumble over themselves on something sort of silly. And it's easy to point at that and say, what are agents all about? Like, is this kind of fluff or hype or whatever it is? And I think the way that I see this sort of generally happening is we've like slowly been ironing out the kinks of the stuff that models fall over themselves on. And we've done that the best so far in coding. I think the industry has done that the best so far in coding where we're making really fast progress on how much an agent can accomplish when writing code.

4:57And we are, I think, starting to make that progress in a lot of other domains. For example, like I talked about clicking on UIs in a computer. This model, Sine 4.5, is way better at that. I don't know if it's exactly the point we're going to tip into people trusting it to automate a whole bunch of stuff they do when they're clicking around a browser yet, but we're getting there. And so I think my general view is we're making really fast progress. It's just not necessarily visible in every part of the economy yet. And then every job and every person and every individual, I have a feeling sort of each model that comes out will get one bit closer to being something that everybody can sort of interact with and see.

5:34So it's great at developing software. It's great at coding. It kind of reminds me of the robot hand moment, you know, how robots can do things that are really, really complex and hard for humans, but actually grasping something has always been a real headache. What about consumer facing stuff? You talked a little bit about UIs, but what do you think agents right now are the absolute worst at? What's the simplest thing that they truly just cannot do? I honestly sometimes struggle to put my finger on this because I think it's surprising little funky things in a lot of different agents. So for example, if you're trying to do something related to finance, maybe there's a bit of manipulating a spreadsheet that is really hard and it falls over.

6:18And so it's like 99 % of the stuff it can do, it can do the math and it kind of understands how a finance model works, but it will stumble over a spreadsheet. This is the thing you talked about. I don't know if there's like, I think it's hard to boil down all of the jobs that we want to help people do. And one tiny little thing that the models aren't there. If it was that easy, I guess we would probably be on top of it already. And I think like maybe a better model is for each of these different things that we want models to help us with. There's like a million little components it breaks up into.

6:47You need to be able to see the right cell on a spreadsheet and know how a formula works and know the macroeconomic model. And thing by thing, you can find in each different task the little thing that's not quite there yet. And so as we think about it, coming from Anthropic, when we think about it, it's like we just have to sort of be able to iron out, like fix each one of the little gaps in each of these to help everybody use it. But I honestly, this is probably an unsatisfying answer, but I can't quite put my finger on like there's just like this one thing. I think this is why this field is hard.

7:17We're just sort of like constantly working on the whole universe of the stuff people do in a remote computer and trying to help them out. And that just means it's a really wide scope. And there's just a lot, it's like from my side, it's a lot, it's fun. There's a lot of hills to climb to try to make the models better at all of the things that we wish they were good at. So we're working on all of them at once. Being someone that works with a lot of different startups and a lot of different industries on how they're actually applying AI, what are the industries that have surprised you the most? What sectors are clients in that you just wouldn't really expect?

7:50Or maybe the ones that you've seen the most growth in in the past six months, the past year. What are the trends you're seeing in terms of which sectors are really adopting agents at scale? I have two versions of this. but I'll give you the fun one first. I think like one of the domains that surprised me the most in my personal customer work is the legal domain, which is at some value, at face value, it's apparent why that can be really useful. There's all this information you need to know about case law and studies. And it's just fusing information is something that's like pretty obvious models are good at.

8:26But actually, there's so much depth and complexity to the legal field, which I didn't appreciate. I'm not a lawyer. I didn't appreciate it until I started working with people in the legal domain. I used to write a lot about legal AI and how the law sector was kind of the last to adopt a lot of AI tools because they were super old fashioned. I did a couple of trend pieces on that. And then when they started to adopt it, I was really surprised at how quickly it came. Yeah, that's exactly what caught me off guard. Honestly, there's so many things that make it hard. Writing a really good legal system, you typically need a lawyer to tell you if it's good or not.

9:00Having to have lawyers in the feedback loop of how to build is challenging. and so I've been really impressed with a lot of the companies I've worked with and their ability to sort of like work around that where they like have this interesting build of companies that have like lawyers on staff to help with product building like like like big amounts of lawyers on staff to help them build products they build agents and cool things like be able to like comb over case law and look for the right details and and pieces and then obviously use the stuff that AI is obviously good at of synthesizing answers at the end but how quickly that field I think it's just like a it's probably like the scale of the upside there of how much just the pure volume of work that needs to be done and how much that it can help that has driven the speed that people go but that domain has surprised me i promise you a two-part answer my second part of this answer is i think the funny thing about this space is it's really hard to guess where the next agent is going to take off because of that thing we talked about before or i talked about before where it's like sometimes there's just this like one little tiny thing that's not super obvious to any of us that's blocking an agent from working so like if you wanted a model to do your taxes for you i think that'd be very nice i don't look forward to doing my taxes every year and you just find out that like actually what it's really bad at is you upload your w4 and it can't quite see the difference between the two boxes on your w4 and that's why the whole thing doesn't work you know this is a toy example i don't think that's actually where models are but yeah that It would be pretty complex.

10:32I remember when I worked in two states or I had a job in one state and I lived in another, it was also really hard for me to figure that out. So not surprising that AI can't either. Yeah, it is complicated. But I think we can get there. It feels in the domain of what the models can do. I've seen them do much more complicated. I don't know. They're better than me at math by a decent amount. So I would expect them to be able to do this. But sometimes it's just like this one little thing. And so I'm kind of constantly surprised. And it mostly is around when we release new models. Like, it turns out there's some set of things that the model can do now.

11:04And then you see new agents crop up that we're doing things we couldn't do before. And so there's like nice little micro surprises that come out. I don't spend a lot of time thinking about accounting normally in my day job, but then you run into an accounting startup that can suddenly do something new and interesting. And that's pretty cool. Do you think that data annotation is going to be a big part of that? Do these models get tripped up because there's some data that they don't have enough of, especially with a niche industry? Is that kind of where you see some of the obstacles come into play?

11:32I think we need to learn from specialists, and we need great ways to learn from specialists. And I don't know if it's necessarily data in the classical sense that there's some pile of data that we're going to learn from that makes it better, or there are just better ways we need to incorporate the intelligence of accountants and lawyers and other people into our models. Yeah, I think that could come from data. Some of it certainly will. I think that can come from talking to and interfacing and working with people in those domains. And I think a lot about learning directly from our customers. I think there's a future where more companies can contribute more directly to making the models do the stuff that they care about.

12:14It would be really nice if someone saw this big, important thing that they'd like to achieve instead of having to sit around and wait for a lab to hopefully build a model that helps them do it. They can sort of like directly work with the lab. And I think that's probably somewhere that, well, I know that's probably somewhere that we'll head where we can have more direct mechanisms of working with sort of experts. And yeah, it's partially a data thing, but I also think it's just like, I don't think it's surprising that one of the things that models are great at today is software engineering. When the building that I'm in, in Anthropic headquarters is filled with software engineers who know how to make models great at software engineering because they write software all day, you know?

12:52Yeah, that's true. And yeah, exactly. Like it's, it's deeply unsurprising that that's true. And I think as we grow up as a company that's only a few years old and has spent most of our time hiring software engineers and figure out how to consult and work with more people in more diverse places and doing more diverse jobs, we'll get better at building models that are great at all of the other things that people wish models were great at. We need to take a quick break. We'll be right back.

13:27Support for this show comes from LinkedIn. Imagine if any of the movies that included the line, I need the right person for the job, settled for, I'll just take about anyone. How many heists would have failed? How many deals would have fallen through? How many secret spy missions would have ended in disaster? So why would you accept just anyone when hiring for your business? When you need the right person for the job, you can turn to LinkedIn Jobs. And now LinkedIn Jobs is stepping things up with their new AI assistant. so you can feel confident you're finding top talent that you can't find anywhere else.

14:01With LinkedIn Jobs' AI Assistant, you can skip the confusing steps and recruiting jargon. It filters through applicants based on criteria you've set for your role and surfaces only the best matches so you're not stuck sorting through a mountain of resumes. Hire right the first time. Post your job for free at linkedin.com slash partner. Then promote it to use LinkedIn Jobs' new AI system, making it easier and faster to find top candidates. That's LinkedIn.com slash partner to post your job for free. Terms and conditions apply.

14:43We're back with Anthropics' David Hershey discussing the landscape for AI agents. Before the break, you heard David explaining some of the trends he's seeing, both from his work internally at Anthropic testing new models and from clients who are now using this tech in their industries. But now I want to ask David about Anthropic's big announcement this week, Claude Sonnet 4.5, and why it's being billed as such a step forward for AI agents. Just released Claude Sonnet 4.5. And to me, that was a big deal. I'm really in the weeds on this stuff. So I was really into all the specs. But for the average person, that's just a word and a number with a decimal and another number.

15:26So why is this a big deal? And how does it differ from your latest models before that? What's the meaningful difference and the meaningful step change here? Yeah, I'm very excited about SONI 4.52. And I'm also very cognizant that when my mom texts me asking about it, it just is a model with a decimal and a number after until I talked to her too. So there are a lot of things that I think are exciting. One thing that I want to say up front before I get into a lot of details that like have jumped off the page to me about the model is it's sometimes hard to predict like what the big impact is going to be on everyone.

16:04These models get generally smarter. They get generally more capable. and not to keep harping on the same point, but sometimes we just don't know the blind spots that they had until we get over them. And so when we released a model last year, Sonnet 3.5, and suddenly all of these vibe coding startups happened, I don't think we actually knew that that was true. We didn't know that there was going to be, the moment we released that model, I don't think we could have predicted that so many companies would crop up and be able to help people start writing code with agents in this amazing way that we've seen happen.

16:34in. So part of what is exciting about Sonic 4.5 is the unknown to me. I'm really confident this model is the smartest model we ever created. I'm really, I've seen, and I'll get into some of the things I've seen it do that I've never seen another model do before. And that's really cool. And part of it's concrete stuff. We'll talk about, I'll tell you, I'm excited to talk about some of the software engineering stuff I've seen it do, but some of it's stuff that we really don't know until our customers and the people we work with try to build cool new things that they couldn't make work before and then they suddenly make a big work before and so when i if i had to guess how my mom might see sonnet 4.5 and it might make a difference to her is that there's some product that she couldn't use before or didn't exist before in the way that perplexity has happened in the past for consumers or or cursor exists now for developers that it's going to happen that's going to impact your life because this model is capable of a new thing that we didn't know before And I have a hard time predicting it.

17:28It's part of the fun part and the hard part about being someone who works for customers now is it's really hard to predict this stuff the day we launch a model. But it tends to happen. And that's cool. Have you seen any trends start, even though it's been like one day? Anyone that's starting to use it, you know, in a new way, one of your clients or even in beta testing? Too early for customers, I think, to see anything brand new. It typically is like, it's a very fast field, but I have like a give it a month rule to find out what the new companies are going to be. My testing, I have certainly seen some new stuff.

18:02And with the testing some of the team has done, that is really exciting. One of my favorites to talk about is my team has been working on seeing how much of a software engineering task a model can take on at a time. And so I think we're all aware that models can be used to write code. like that's sort of obvious to a lot of people at least in this industry and probably people listening to this podcast at this point but it's often in pair with a human really directly like going back and forth using quad code or an ide like write one thing at a time check and review it and the longer you let a model go trying to implement something big and complex the more likely it is to like do a whole bunch of stuff you didn't want to happen and so we were like really curious, what is the most we could stretch that?

18:45If I give Quad a really huge task, what can I accomplish? And one thing we've seen with this model is in a way that is really not true of models I've tested and seen in the past, if you give a sufficiently good overview of what you want a model to accomplish, I haven't really seen a ceiling on how much it can keep Costa consistently working on making something better. And so my favorite actual example of this, We released a video of this yesterday, but I'm going to talk about this until people get bored with me. We asked Claude to recreate Claude.ai, our consumer chat application, from scratch.

19:24Sonic 4.5 just worked overnight. We woke up and it just did it. This beautiful clone of Claude.ai that works incredibly well. in my favorite moment from it there's a feature that people like of ours called artifacts where when you ask quad to make a document or make a web page it will make it and render it next to the chat so you can like play with an app that you built live or whatever it is and i was saying to the person on my team who was working on this demo like work on this thing it'd be really like cool if quad could like build artifacts itself that would be amazing it's like a complex feature it's pretty hard to figure out.

20:02It's like, we should, we should try and see if that happens. And then he messaged me like two hours later and he didn't do anything. He didn't intervene. And he's just like, Hey, quad just built artifacts on its own. And it's like currently testing it live. It's just like trying it out itself. And I really, it's just like, we released artifacts a year ago. It was like artisanally created by a whole bunch of like wonderful engineers around me at Anthropic who were doing all of the software engineering to build this really complicated things themselves. And then a year later, we released a model and we asked it to build Quad.ai and let it go for 12 hours and just did it itself.

20:41Instead of this giant team of people at Anthropic that work so hard on this thing, the model is just really capable of biting off really meaty, complicated tasks. This is something that would take me months to do if I did not have quad. And overnight, we sort of looked at it and watched it happen. And that progress of just like, like, going from a point where it was neat that a model could write a snippet of code for me a year ago to like, oh, it can do like a big chunk of the stuff of the complex developer work that I need to do. It's just like, I don't know, it blows my mind. It was like a pretty big, Wow.

21:21So that took about 12 hours, you said? That specific one was like a 12-hour thing. We've seen up to 30 hours of continuous dev work. Okay, this is what I was going to ask you about, because something that unexpectedly went viral from my own article about Sonnet 4.5 was the 30-hour bit, the fact that Sonnet 4.5 could code autonomously for up to 30 hours with no interruption. So I heard one engineer at Anthropic used it to code a chat app that the company compared when they talked to me to Slack or Teams. Obviously, it was only 11 ,000 lines of code, so much smaller than Slack or Teams. And it seemed like just an example project.

21:58But people online are really excited about that detail and calling for Anthropic to release it, and they want to know more about it. So can you give me any details? Was that someone on your team that was testing it? And how impressive was it really? Or was it just kind of rudimentary? Give us the deeds. Yeah, yeah, yeah. This is the thing that also, I don't know, like all of my excitement and the reason I'm years because I've been working on this exact thing. So it was someone on my team. His name's Justin. Shout out Justin. You should give praise to him. He's amazing. Recently joined the team on a side note.

22:26He's pretty good. Nice. This was born out of this thing of a lot of people have built demos or proofs of concept. And there's this sort of vibe coding trope that you can, yeah, you can build it, use it to quickly mock something up, but can it really like build a real application? Justin really wanted to like test that and so he was experimenting using our quad agents SDK which is just sort of like a more programmatic version of quad code to some extent he was testing like can I give a full spec of like a complex thing like an app similar to Slack to the model and just like watch it build and we did some takering and experimentation around it to get it right but But yeah, you asked, was it impressive?

23:09It's impressive. It has DMs and threads and channels and a slick search functionality. And you can upload images and GIFs and render them. And multi-user authentication. And Claude, we didn't ask it to, but implemented a whole bunch of AI users for testing. So if you log in, you can send a message. And there's Alice, the PM, is in there that you can send a message to who will respond to back stuff about PM work. it is it is remarkable it is like it is not by any means slack but like if you didn't spend a lot of time thinking about it you'd like look at it and think that was a pretty reasonable productivity app that you would use to chat with your co-workers wow what did you guys name it i don't think we have a name i need to ask justin to give it a name we need i actually no i lied he has a work chat very boring we're engineers we're clearly not the product folks in the org so So I love that.

24:06Yeah, someone who worked at Slack, I think, messaged me and said, you know, our code base is a thousand times that. And I was like, yeah, it's just an example. But I mean, it happened in 30 hours. So it's a big deal. I was also going to ask you what surprised you the most during your own testing of it, of Sonnet 4.5. I'm just going to continue double clicking on this thing. The way that it built the app was surprising and really interesting. I'm not surprised that the model is getting better at building complex apps. the thing that was really interesting for the way it does this is it just has a tendency to like bite off little pieces that it can handle really well and do that continuously so a lot of times like models in the past would get really eager and ambitious they would say like i'm going to build this whole thing and i have like these grand ambitions and it would kind of just like meander everywhere trying to do this miraculous piece of work and the thing that was really cool about outside of 405 is just like pragmatic kind of like it it's like okay right now I'm going to test like does image upload work and then it's going to do that it's going to spend a little while doing that but it just like bites off one little chunk at a time and that feels a lot more like what I want a co-worker to work like or a collaborator work like it's like if I ask you like hey I need you to go build workshop I don't want you to go off on this like crazy escapade trying to make everything magical I want you to just like bite off a piece at a time show it to me, like committed to get like that kind of thing.

25:32And it's just a little bit more natural and collaborating with feels with, it feels more natural. And funny enough, like, I think this is unrelated, but we've been chatting in our internal company Slack with Claude a lot lately and just like chatting with it. And it's been really natural. Like it feels just a little bit more like working with a coworker does. In terms of the tone? Yeah. Like the, the tone, how it, how it responds, like how it acts in Slack, how it, how it participates in a conversation, what it tries to do. It's a little less like over the top and eager that has jumped out a little bit, which was surprising.

26:07Like that's not something that I normally expect, but yeah, it's also been funny, I guess, which is a funny thing. Like it like cracks better jokes. Like it's a little bit more witty, that kind of thing. Do you think it also dialed back the sycophancy a bit? You know, I know that's been a big word in this space lately in terms of, you know, being over eager, over validating? Did you feel like it was more real and that it was less like that a little bit? I have felt that. And I think it's certainly something we're pressing on. I, it's like one of my least favorite thing about every model is when there's a phantic like that.

26:37I think it just gets in the way of doing good work. And it's been a big focus as far as like, I think nobody ananthropic likes that. And, and also just like, I think it's, it's bad for the, it's just bad. So I do think we have made meaningful progress and this model seems to be a little bit more willing to push back. And that's part of that thing. Like being a natural coworker is someone who actually can tell you when you're wrong. When it comes to rebuilding claw.ai, that's pretty big. Did that worry you at all? Because, you know, it's kind of doing, like you said, you know, organic work that you worked on for months.

Read the full transcript

27:11Does that worry you about job replacement for engineers, anything like that? What were your thoughts? We have a great team, and I'm really not worried about, I'm going to frame this two ways. I'm not currently worried. Right now, Quad is a collaborator. It works really well with me. It accelerates me. I think it makes our whole team better and faster writing software. I am general. This is not a now thing. To be really honest, watching Quad go for 30 hours, it does trigger a little bit of like, oh my God. like that's like a pretty different thing it is a meaningful step change and i think it does like this technology anthropic is founded on the principle that this technology is going to be hugely impactful in the world and part of that is it's going to change how we do jobs and something like doing a whole week of work for me like that's just like meaningfully going to change the industry of software engineering and so yeah there's like a little bit of and i would be like It'd be goofy for me to say there's not any amount of, how are we going to work next?

28:15How do I incorporate this? What is it? My net-net, though, is I just think there's a ton of room to make us better, to make better software for users, to make the world a better place with this technology. And I'm really confident in that still. But there's a smidgy of, we're going to have to figure out some new ways that we build with Quad and that we operate as software engineers to work with a thing. If it's really going to build the whole app itself, we have probably a different role that we need to play here. We need to take another quick break. We'll be right back.

28:54Support for the show comes from Rippling. If you're a business owner, here's the truth. SaaS promised to make work easier. But now the average company is buried by hundreds of apps that silo your teams, slow you down, and simply don't work together. That's not SaaS. That's SAD. Software as a disservice. That's why you need Rippling. Rippling is the unified platform for global HR, payroll, IT, and finance. They've helped millions replace their mess of cobbled-together tools with one system designed to give leaders clarity, speed, and control. By uniting your employees, teams, and departments in one system, Rippling removes the bottlenecks, busywork, and silos your software created.

29:37Automated, perfectly in sync, and seriously simple to use, Rippling gives your company one source of truth for your people, their data, and everything they touch. With Rippling, you can run your entire HR, IT, and finance operations as one. Or pick and choose the products that best fill the gaps in your software stack. And right now you can get six months free when you go to rippling.com slash decoder. Learn more at r-i-p-p-l-i-n-g dot com slash decoder. That's rippling.com slash decoder for six months free. Terms and conditions apply.

30:18Refresh your good vibes with Tic Tac. Mit den richtigen Vibes wird... Hey! Oh man! zu

30:51We're back with Anthropics Applied AI Lead, David Hershey. Before the break, we were talking about the capabilities of Claude Sonnet 4.5 and whether he sees a future where this technology might even automate parts of his job. But now I want to ask David where he sees the model falling short and what might be next for AI agents. Well, what are the primary limitations to Sonnet 4.5 that you wish you guys could have offered with this release that you couldn't? And what was something that you tried to make it do during testing that it couldn't do? So basically, yeah, the things you, the features or, you know, the context window or anything else that you wish you could offer that you couldn't.

31:29And also what was dumb about it in testing. I have a pet thing that I always test on, which is I really like to make quad play games. I'm accidentally famous for creating quad plays Pokemon in the past. Oh, that was you. I've heard a lot about that. Yes, that is my project. And I have also tried a lot of other games. Like I had quad playing chess this time. And one of my favorite comics people really wants me to make quad play Catan effectively. And it's really bad at spatial awareness still. And they say it annoys me to no end. Like it just like basically doesn't know the difference between left and right and up and down.

32:08And this is just like, it's just one of those examples that exists in the world of it can do PhD level math and I can't. And just for it to not really understand that it can't walk straight through a building, it hurts my brain. And models are kind of weird and funny that way. So that one was the one that probably jumped out the most, is I really wanted Quad to be this great chess player. It's so logical and interesting. And then it doesn't know where the pieces are on the board and where other pieces are. And it's like, ah, you're almost there, Quad, one day. And features that I think are still out there, I have been so in on this model that I have to think about the next level.

32:48Honestly, like the real answer, I don't know. I can't think of some like feature that I wish was there. There's like a lot that I want quad to get better at. There's like so much that I want quad to get better at. I want it to be able to beat Pokemon one day on its own. And I want to see it like how it does it. I want to like, again, there's all of these fields where I'm like, know that quad still needs to get better. Like I, we talked about legal, like I think quad is becoming a better lawyer, but I don't think it's like as good as a lawyer as it is a software engineer yet. And like, I still wish we worked on that.

33:18And, and we're making progress. Like I know the teams that are working on all of these little things and focusing and thinking about all these little things. And there's always this, like so many hills in my mind that we have to climb still. So I could probably like, there's probably this intersection of, there's a million things I would call it was better at. Not one that I am personally making except for Pokemon, which is what one day the research team will listen to me that it matters. They haven't quite listened to me yet. Well, when I went on a reporting trip to London and visited the office, it was all I heard about.

33:46So it's making waves somewhere. That's good. It makes people happy here. It's a fun experiment. I don't think it has crested the peak of what we need to train Claw to be good at yet, unfortunately, one day. Let's talk a little bit about what's next. So talk to me about why you think Anthropic is pursuing such gains on the AI coding front and how it stacks up to competition. Obviously, in testing, you had to take into account what the market looks like right now and what other models can do. How did you think Sonnet 4.5 stacked up during testing, and what stuck out to you there? Coding market is really important for a lot of reasons.

34:25It's really well-positioned for people to build with our models. It's a place where you can have a huge impact. People have figured out great ways to integrate models into how they do a job, writing code. and I think more so than essentially another industry. And so if there's like a right now where you can make a huge difference by continuing to make models better, I think coding is probably like the best use case. And so it's a huge focus for us. We want to keep helping people who are relying on our models to write code be able to get more out of them. And our belief, we've talked to a lot of customers, done a lot of testing.

34:56Like I'm pretty confident that this is the best coding model in the world. We have a lot of benchmarks and other things to prove that out. And it certainly has like for all of the people here in our early testing just been a huge step function improvement in what it feels like within quad code or other surface areas to develop and write code. And from my perspective, this is like a really big step function. And personally, like this is just talking for myself, this is the most noticeable change and improvement I've seen since we released Sonnet 3.5 last year, which I think is a funny coincidence.

35:32I don't think we necessarily knew that these 0.5 sonnets were going to be such special models for us. But just mechanically, this is some interface of vibes and numbers and things I've seen. This feels like a really, really meaningful step change improvement. One of my favorites, just to maybe call it out, I love the team at Cognition. I spent some time working with them and they put out a post yesterday about how much their product, Devin, improved with the model and some of the work they did to make that product better that I loved. It's a cool blog. And that like really, I don't know, I think that just that validation from the customer, like it's a huge jump in a benchmark that I really haven't seen have a benchmark, like a jump for them in a while.

36:18I just think that's really cool. Like for the right product, built the right way around this model, I have a feeling this is a huge step function change in ability to write coding. And I'm quite confident it's going to be the best model in the world for people like that. Yeah, that's really interesting because I usually don't mention benchmarks much in my own reporting because sometimes they can be subjective and sometimes they're, you know, created by the companies that are testing their own models in very specific areas with very specific sets of questions. But I think usually what I hear from engineers is that they go based on the feeling and based on what things it can do that it couldn't do before in their own anecdotal testing.

36:57So that's why it's interesting to see, you know, you have your own pet projects that you like to test on, and you've seen changes in that regard instead of just in the benchmarks specifically. I also wanted to ask you about anthropics chasing consumers versus enterprise versus governments. So, you know, all AI companies right now, it seems like, are kind of like working with those three tiers. And I think it's probably because, you know, those are two of them maybe are more concrete areas for potential profit. So do you think with Sonnet 4.5 and just all the stuff you're working on in general, I know you work mostly with the enterprise side of things, but is there a specific slice that Anthropic is chasing more right now and why?

37:41I think they're basically all really important. And I think one of the things that we have luckily seen, and I think is true for all of the labs, is that when we make our models generally smarter. They service all of those segments. They service enterprises who are building with us. They serve as consumers who want to use our chat app or write code or get into that. And they're useful for the public sector too. Just to tie a little bit of a link here, I've worked with customers and startups that have become big consumer hits. Like Lovable is a great example where it's like that's an enterprise customer from my perspective that I spend a bunch of time working with and helping them be successful and our team does.

38:25But then they go turn around this thing that everybody can use to build just-in-time apps to service every part of your life. And so I think this kind of just bends back on itself where in reality, our focus is building really great models that are safe. And we have seen time and again that that results in sort of like success in all of these places. and there's just like a lot of different ways from our perspective. This is talking a little bit from my personal perspective instead of just Anthropics, but I think there's just a lot of ways that you can make progress building great models. And whether it's helping enterprises or helping consumers or helping the government, like these all have a way of warping back in on themselves where the thing that we really do that makes a difference is make great models.

39:10And there are a lot of different ways you can impact different segments with that. And when it comes to consumer use cases, OpenAI seems to be pushing pretty hard into that right now, at least this week. They launched Pulse recently, last week, and then yesterday they debuted their instant buy button. You know, I wanted to ask where you think Anthropic is planning to meet consumers. Do you think it's more likely you'll be reaching customers with Claude directly in the future, or is it more likely customers will use Claude through something like Cursor? You know, and that can also apply to your startup clients too, right?

39:42Where are you seeing people really find Claude? I definitely would think we are growing. It's just like the presence of the applications we build to interface with consumers. And Claude Code is a great example of this. I know it's not like the traditional consumer, but there's like a very prosumer-y thing that we've captured with Claude Code. And I think we've demonstrated that we do have some of the muscle to capture, like for some sets of clients, a thing that they love that, that's really like sparks the excitement of a consumer product that people love using. And I guess like, obviously we would love to do that 50 times over.

40:21We would love to be able to invent a bajillion beautiful uses of quad that interact or that, that consumers love and love to build. And we're investing in trying to find and build great experiences that help people do important things with quad. We have a sort of constant tinkering we we we had this imagine demo that came out it's just a limited time thing i don't think that's the product we're gonna launch that is consumer facing but like we just like it's part of a portfolio i guess of like how we think about we need to keep building and trying and innovating and inventing products seeing if there's something there and we have a good opinion of what models are capable of that helps us have an interesting lens on what we build so it's a big focus like we We need to get there.

41:09It's important for us to build direct relationships and help people have lots and ways of interfacing a quad. That said, I'm biased as a customer guy a little bit, but I also just think I wouldn't ever, and I don't think we should, and I don't think Anthropic is betting against the ecosystem of people who are trying to build amazing products. And I think it'd be really silly to claim that that's all our market to grab. I just don't think it is. If we build great models, then there's an incredible ecosystem here in Silicon Valley and abroad building amazing products that I think is probably a bigger upside than our first-party applications will ever be.

41:47So I think that it will always be a huge focus of ours. Awesome. And then last question for you. Back to the AI coding market and how important it is and how much most AI labs are chasing that right now. So we talked a little bit about Vibe coding earlier. And here at The Verge, a lot of us have tried it. And to no avail, we had some success with really simple stuff, but not a lot of success with building large-scale applications. A couple of us did, but not as much as we actually expected. You know, did you see a big change in terms of testing out, like, Vibe coding with Sonnet 4.5? Or was it the same as before?

42:27I have noticed, like, my own personal coding. like every model we release i am trained as a software engineer so it's just like cheating i'm like not the best guinea pig for vibe coding but i still do it on the side sometimes and you notice meaningful improvements on like what how much it can turn out before it goes off the rails and how much you can trust it i actually i think this is actually an interface problem and i think there's something funny about ai which is that it has this tendency to outgrow interfaces really fast So if you look at the history of coding with AI, there was probably like Copilot with Ghost Text.

43:02So Copilot would automatically complete your code. For a while, there was, you would go to us or ChatGPT or wherever it was, and you would ask it to write some code for you in a browser window if you were a developer, and then you'd copy and paste that into your editor. Cursor figured out how to sort of like bridge those two things where it could be side by side. A lot of people have started building this sort of agent that was alongside your ID that lets you build things. I don't think any of that is quite the thing that we need for everybody to build production applications though. Like there's some interface where I actually do think Sonnet 4.5 is the first model that could be that thing where anybody could build a sort of like production ready application.

43:40I've seen enough evidence of when left to its own devices, how it can build complex applications. I've seen it be able to deploy a complex application to AWS, like fully autonomously and do like a security audit on it. Those are like both really incredible things that like make me think this is the model that could sort of like cross that chasm and get to the point where like anybody can make something production ready. I have a feeling though we need like one more interface that isn't quad code and isn't cursor. But like the next step past that, I think needs to happen to get so that like it's more obvious to everybody instead of having to like try to figure out if you're on the right path by coding.

44:20Awesome. Awesome. Well, thanks so much. I'm so glad we got to talk and appreciate you coming on with so last minute. I really appreciate it. Yeah, it was fun. I appreciate it. Of course, it was really nice to be on. I'd like to thank David for taking the time to speak with me and thank you for tuning in. I hope you enjoyed this episode. If you'd like to let us know what you thought about this show or what else you'd like us to cover, drop us a line. You can email us at decoder at the verge we really do read every email or hit me up directly on x blue sky or threads i'm at hayden field on all platforms decoder also has a tiktok and an instagram and now also a youtube channel check those out at decoder pod they're a blast if you like decoder please share it with your friends and subscribe wherever you get your podcasts decoder is a production of the verge and is part of the vox media podcast network our producers are kate cox and Nick Stat.

45:14Our editor is Ursa Wright. The Decoder music is by Breakmaster Cylinder. See you next time.

From the publisher

This is Hayden Field, senior AI reporter at The Verge and your Thursday episode guest host. Today, I’m talking with David Hershey, who leads the applied AI team at Anthropic. I wanted to have David on because earlier this week, Anthropic released a brand-new AI model called Claude Sonnet 4.5 that’s been making waves.

So I wanted to sit down with David, who spends a lot of time testing out what modes like Claude Sonnet 4.5 can and can’t do, to ask him where we are on this promise of AI agents, and also what the path forward looks like as agentic technology progresses.

Links: 

Anthropic releases Claude Sonnet 4.5 in latest bid for AI agents | The Verge

ChatGPT’s built-in Buy Now button has arrived | The Verge

OpenAI really wants you to start your day with ChatGPT Pulse | The Verge

Anthropic’s Claude AI is playing Pokémon | The Verge 

AI agents are science fiction not yet ready for primetime | The Verge

Agents are the future AI companies promise and need | The Verge 

Amazon is betting on agents to win the AI race | Decoder

Credits:

Decoder is a production of The Verge and part of the Vox Media Podcast Network.

Our producers are Kate Cox and Nick Statt. Our editor is Ursa Wright. 

The Decoder music is by Breakmaster Cylinder.
Learn more about your ad choices. Visit podcastchoices.com/adchoices

More from Decoder with Nilay Patel

All 153 episodes
The good, the bad, and the future of AI agentsDecoder with Nilay Patel · 47 min
Listen in VO