Claude Sonnet 4.5 Review, AWS Director on AI Agents, Vercel’s $9B Valuation | Sep 30, 2025

30 Sep 2025 · 40 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: The Information's TITV - Episode on Claude Sonnet 4.5 and More

Episode Overview

  • Title: Claude Sonnet 4.5 Review, AWS Director on AI Agents, Vercel’s $9B Valuation
  • Date: September 30, 2025
  • Host: Akash Pasricha
  • Main Guests:
  • Jeanne DeWitt Grosser, COO of Vercel
  • Zach Lloyd, CEO of Warp
  • Andrew Filev, CEO of Zencoder
  • Theo Wayt, reporter from The Information
  • Shaown Nandi, Director of Technology at AWS

Key Highlights

AI Agenda Live Conference

  • The episode coincides with The Information's AI Agenda Live conference held in Times Square.
  • The show promises to provide highlights and insights from the conference.

Vercel's Series F Funding

  • Funding: Vercel raised $300 million at a valuation of $9.3 billion.
  • Business Focus: Vercel aims to become the "AWS of AI," enhancing their AI cloud capabilities for developers.
  • Product: Vercel's V0 is a tool for building AI-native applications, with aspirations to support various applications like conversational agents.

Claude Sonnet 4.5 Model Review

  • Early Impressions:
  • Both Zach Lloyd and Andrew Filev praise Claude Sonnet 4.5 for its improved performance over previous versions.
  • Enhancements include better context management and reasoning capabilities.
  • The model is seen as user-friendly, lowering the barrier for non-technical users to effectively leverage AI coding tools.

Discussions on AI Models

  • Model Effectiveness: The conversation touches on the potential plateauing of AI model improvements, with emphasis on the need for context and understanding of the intent behind user queries.
  • Performance Benchmarks:
  • Speed and accuracy are critical metrics for assessing AI model performance.
  • Claude Sonnet 4.5 is noted to resolve issues faster than its predecessor and competes well with more expensive models.

Insights on XAI's Organizational Changes

  • Theo Wayt discusses significant changes in XAI, Elon Musk's AI company, particularly around org chart dynamics.
  • Observations include:
  • Less hierarchical structure compared to traditional tech firms.
  • The importance of key figures like Tony Wu and Guedong Zhang in the company’s leadership.

AWS Perspective on AI Agents

  • Shaown Nandi's Definition:
  • An agent is described as an autonomous or semi-autonomous system capable of reasoning, planning, and executing tasks, differentiating from traditional automation.
  • Application Examples:
  • Discussed enhanced seller agents for Amazon's marketplace that optimize catalog management and pricing.
  • Case study of Formula One using intelligent root cause analysis for faster network outage resolutions.

Key Takeaways

  • Investment Trends: Vercel's significant funding highlights the booming intersection of AI and cloud services.
  • Model Advancements: Claude Sonnet 4.5 represents an important step in making AI tools more accessible and effective for a wider audience.
  • Organizational Flexibility: XAI's flat structure facilitates rapid decision-making and adaptation compared to more traditional corporate frameworks.
  • Future of AI Agents: Continued evolution in agent technology is anticipated, with a focus on improving communication and collaboration between systems.

Conclusion The episode provides a comprehensive analysis of the latest advancements in AI, particularly through the lens of new models and funding ventures. It underscores the dynamic nature of the tech landscape, driven by innovation and the strategic positioning of companies like Vercel, AWS, and XAI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:13Welcome, everyone, to the Information's TI TV. My name is Akash Pasricha. It is Tuesday, September 30th. We are all excited here at The Information because today is our AI Agenda Live conference, which is happening today in Times Square. Be sure to follow along our website for more coverage on that. We are also going to be bringing you the highlights from that event over the next few days. So stay tuned to this stream. Today on the show, we've got some exciting guests coming on. We are talking about the latest release of Claude Sonnet 4.5. We're going to be talking to some founders who have been testing out the model.

0:47And by the way, Vercel is coming on the show first, actually. We're coming on first to talk about their latest fundraise. We also have some of the latest reporting on XAI's org chart. And separately, we've got a great conversation planned for you with our friends at AWS. But before we get going, I want to highlight for you a big story that we published last night at The Information. The Information is first to report that OpenAI generated more than$4 billion in revenue in the first six months of 2025. And to put that number in context, that figure is more than the total revenue that the company generated in all of last year.

1:26Now, one thing to note here is that we know the company was recently generating$1 billion a month in revenue, so it is likely making good progress towards that$13 billion annual revenue target for the year. The story also has a lot of great details about where the company's cash burn is at, and so I highly recommend that you check that out. Okay, let's get to our first guest. Developer tools have changed significantly as a result of AI, and Vercel is one company that has gained a lot of traction as a result of that trend. The company raised$300 million today in a Series F funding round at a$9.3 billion valuation.

2:06And I want to bring on the company's chief operating officer, Jean Grosser, to talk about the news. Jean, welcome to TITV. It's great to have you. Thanks for having me, Akash. So I want to talk about the business more broadly later in, but very quickly, what does Vercel do? Well, we do a lot of things. So historically, we've been known for providing a front-end cloud. So basically, the best place to go and build full-stack front-end applications. And then about a year ago, or a little more, maybe 18 months, we released V0, which is a Vibe coding application that many are using for prototyping and app building.

2:43And then underneath the hood of V0, we had to build a bunch of infrastructure to support an AI-native application. You can't just get that stuff easily off the shelf these days. So in June, we released the AI cloud, which is now the best place to go and build full-stack agentic applications, whether that's for conversational agents or background agents. You can build them all on the AI cloud. Right. So AI coding, essentially, you're helping people do that painful task in much more automated ways with your technology. What are you going to do with the money? Well, we'll invest in all three of those areas.

3:19So a lot to do in building out the AI cloud. You can sort of think of our aspiration is to go become the AWS of AI, giving everybody off-the-shelf primitives that they need to build AI applications. We've recently seen a lot of traction in enterprise. So we'll be building out both enterprise product and the go-to-market org to support that. And then I think one of the things that's been incredible we'll see about AI is just how instantly global it is. So continuing to expand our global reach. Right. Now, one of the things I want to talk about is there are so many AI coding tools out there. We've got actually a couple other founders who are working on their own AI coding tools coming on the show in a minute to talk about Claude Sonnet 4.5, which I want to get your thoughts on in a second because I know you guys have been using it.

4:06But if we look at the AI coding landscape more broadly, I mean, how do you see this shaking out? Are you looking to buy other companies right now? Do you see consolidation happening? That's what I'm forecasting here very quickly. Yeah. I mean, certainly it's a fragmented market. And I think as both the makers of vZero, but also the infrastructure that goes underneath that, primarily we really want to support all of these applications getting developed. So Cursor runs a bunch of their agentic infrastructure on us. Many of the other Vibe coding applications out there similarly are using our AI cloud.

4:43And in fact, we open sourced V0. So you can go, we actually enable the fragmentation. You can go and use what's under the hood of V0 to create your own. Got it. Neuron, one of the things you'll probably see is actually the fragmentation is beneficial because people can go deep on specific types of use cases where you really want to fine tune the AI. That's one of the things with V0, we're deeply focused on React and Next.js. So, you know, which is what? it. So Next.js is the top framework for building front-end applications. So if you're a developer writing code to build a front-end application, there's a decent chance you're going to be writing it in Next.js.

5:26And what we can do with v0 is actually we have a composite model. So we'll take a model like Sonnet, but we have a fixer model that runs on top of it that catches all the places where Sonnet got it wrong and adds all of our Next.js expertise to make it more likely that you can one-shot a prompt to an app or build a really robust application that actually works. Right. You've been using Claude's new 4.5 Sonnet? Yes. What's the review? We're going to have a panel on it in a second, but we might as well ask you, how are you finding it? Yeah. I mean, we often are working with a lot of the model providers before they actually publicly drop their model updates.

6:06So we've been working on with that one for a bit and still think very highly of it. it's it's powering b0 right and what what's your take on on you know the the sort of discussion around the effectiveness of these models sort of plateauing you know the idea that it's harder to make substantial improvements when the things are already quite good in some cases do you see that plateauing of the models um so far we're definitely seeing improvements with with each uh version um i think one of the things you're seeing a lot of the last is it as good is it as good all right Like, I mean, the margins must be shrinking here a bit.

6:44Yeah, that's a good question. I probably have to get our VZero engineering team over to give you a more detailed answer on that one. We have like a pretty robust eval process that we go through with each one of these model drops and ability to compare them to other models, et cetera. Right. But I do think one of the things you see a lot of the model providers do is now, you know, focusing in specific areas. So obviously coding being one of those, but it seems like there probably are still returns for more area-specific models. Right. You guys are one of the bigger coding companies. Are you looking to do acquisitions?

7:17Do you see them coming in the next 6 to 12 months? Vercel has been highly acquisitive, typically of smaller startups doing really interesting things. Most recently, we actually made a pretty significant acquisition in buying Nuxt, which is another open source front-end framework. So we are always out there looking at things, whether it's a one-person open source developer or folks with interesting functionality that we think we could add to the platform. Right, right. And that is sort of the reason we talk about consolidation here is because you talk about the infrastructure that enables these coding tools.

7:56That's the business that you're on. you talk about the coding tools themselves, you talk about AI code review, right? I mean, we've had companies from all three buckets. I kind of see this as, you know, why not have a one vertically integrated entity that does it all? I mean, is that what Vercel wants to become? I mean, that's how we describe ourselves, right? So you've got the framework in Next.js, which we build and maintain. And you've got, you know, the platform in all of Vercel's infrastructure that's specifically built for front-end and now agentic applications. And then you have built into our platform a bunch of agentic capabilities.

8:38So exactly what you just said, code reviews were releasing the Vercel agent. And so it will do code reviews, and it does them quite well because it's doing them in a specific context of XJS and code built for Vercel. So you can imagine that we can be quite accurate and effective at that relative to more generic-based code reviews. Right, right. So I think in general, our view is being vertically integrated is real strength here and a differentiator of our platform. And that, you know, the goal is basically to take agents across the entire platform. And so we're providing folks solutions rather than problems.

9:17Right. Gene, before I let you go, I want to just take you back in time a little bit to your previous role at Stripe. You were at Stripe for more than nine years. You were most recently chief business officer. And the reason I want to ask you about some of your insights in Stripe land is because we saw OpenAI yesterday release this agentic checkout tool, I guess, you know, you can now check out with ChatQBT. And we've been really fascinated by the intersection of AI and commerce, which I imagine is something you were looking at very closely in your previous role. You know, indulge me here for a minute.

9:48Where do you think this is going? Because, you know, I saw that last night, the feature, and it felt a little bit, a little bit under, you know, I don't know that it was quite the agent that everyone imagined doing the purchases for, you know, sorting through different options, deciding what the best option is. I mean, do you see this technology being a threat to something like Amazon? Or how do you see that playing out? That's a good question. I mean, it could, right? Like, I think right now you're probably going to want a human in the loop where you get to say, okay, I agree, agent, you did go find the thing I was looking for.

10:22But there are a lot of more basic, you know, purchasing needs where sort of you as the human don't have a strong opinion. You just need the thing. And, you know, arguably that's what Amazon's awesome for among many things. But I could see that's probably where agentic commerce goes first is things where you're highly unlikely to disagree with the agent. Right. Versus in other cases, it teeing up. Like toothpaste. Yeah, exactly. You're welcome to go get Crest for me. Right, right, right. But the more discretionary purchases, you need a human in the loop. What did you think of the release yesterday?

11:00I have not gone into it in detail, although Rafael actually is also working with Stripe on a bunch of agentic commerce runs. So I will be spending Stripe tours today. So I think they went through it in depth at that event. And I will be watching the keynote later this evening. And so what, I mean, you know, is Stripe, you know, you know what? I got so many questions for you on this. I'll tell you what, why don't you come back on the show, you know, a week today? Let's keep this conversation going because I would love to ask you more about the Agenda Commerce Front and also what Bracel is doing with Stripe.

11:34Thank you so much, Gene, for coming on the show. That is Gene Grosser, the Chief Operating Officer at Purcell. Okay, well, we are marching on ahead to our next guest, folks. As we just mentioned, Anthropic released Claude Sonnet 4.5, and developers have quickly been running to try it out and see just how good it is. I want to bring on two more people who have been using the new model already to give us their early reactions. Zach Lloyd is the CEO of Warp, and Andrew Filev is the CEO at Zencoder. Both those companies are AI coding software companies. Welcome to you both. It's great to have you.

12:10Thanks for having me. This is awesome to be here. So I'll tell you what, I got a really open-ended question here, and it's what are our initial reactions? So who wants to go first? Zach? Sure. Initial reactions, very good. We march on as far as the quality of these coding models improving. It's a significant improvement over the last generation of Sona models. It keeps what's good about them. It's fast, but it can also do sort of longer horizon reasoning tasks. It does a bunch of things really well with multiple parallel tool calls. So overall, I'm impressed. Very good. Andrew, what do you think?

12:47It's an awesome model. And it does what models, what great models have been doing recently is they're taking in something that works in the industry and they're training the model to do that. So if you think about the previous big splash, it was reasoning models. So they basically took their chain of thought technique that was popular in industry and like kind of lusted it out, right? That they're LLM scale. And then what this model is doing, one of the things that I don't know if people picked up on this, it's very good at context management, which is one of the most important things to get these agents to do sizable tasks.

13:25So if you worked with any kind of professional AI-first engineer, they're getting really good at breaking down the big tasks into smaller ones, right, to manage context. And basically my read from how this model behaves, how it operates, is that Anthropic took that paradigm and trained the model to do the same automatically. So this means that a lot of people who are not yet good at that skill, who did not go through this kind of personal journey of discovering how to best use the models, how to do context management. Now they can get it out of the box. They can just throw an idea at the model and the model will automatically create the requirements document, the plan for them and whatever.

14:05So it lowers the barrier for regular folks to get the most out of these models. And as a byproduct, another cool thing is when Frontier Labs train models to solve those things, because it's their general intelligence, they automatically pick up a lot of other things. So another big improvement is that Sonic right now is becoming, in my opinion, Sonic 4.5 is becoming as good as writing those plans as the previous generation of Opus, Opus 4.1, right? And Opus is five times more expensive than Sonic. And I will take that five times cheaper price any day. And it's also faster, right? Because Sony is a smaller model.

14:47Yeah. So it operates faster. And that's awesome. Zach, talk about the pricing here. Yeah. So it's the same price as the last Sony model, as Andrew is saying. But it's better quality. And on a bunch of the benchmarks, it's as good as Opus, if not better. And Opus is five times more expensive. so from the perspective of offering value to people who are using it um this is like a huge step up and you know we hear a lot about these benchmarks you know a lot of the uh readers of the information are highly technical people but a lot of them also are people who may not have as much technical acumen as people coding when they talk about the benchmarks what exactly in layman's terms are you assessing any on?

15:32I mean, one thing is speed, I guess. Another is sort of accuracy. What else are we looking at? Oh, sorry. Go ahead, Andrew. Yeah. So for the benchmarks that are published, when labs assess it, they typically just assess their resolution rate. And there are different methodologies. We won't go into details, but they assess how frequently does the model solve the issue now when we run benchmarks internally we also look at cost and we also look at the speed and so for example one of our early insight is that their sonnet 4.5 resolve the issues about 20 25 faster um than sonnet 4 which is again awesome for the real life and you don't see it in the public benchmarks because they're all about just just the high mark right uh but internally when you're evaluating those agents uh whether it's enterprise or whether you're building products, you do need to look at Takash, as you said, as like the whole picture, right?

16:28It's not just the resolution rate, but it's the speed, it's the cost and everything else. Zach, jump in here. Anything else you guys are looking at in terms of benchmarks? The interesting thing about the latency is there's two ways of looking at it. So there's like, if you're sitting there using Sonnet, how long does it take for the token to stream back to you? And the Sonnet models are very, very fast. There's a second way of looking at it, which is like, How long does it take for the model to actually complete the task? And it's better on both of these. But one of the interesting things is like we find that some of the the reasoning models, like for instance, GPT-5 is slow to return the first token.

17:08It like spends a ton of time thinking and you can see it thinking, but it actually completes the task faster. Got it. Two pieces of latency are not exactly the same. Other than that, I agree with Andrew. It's like they're reporting resolution rate. We ran all the benchmarks on Sonnet 4.5 internally. It scores higher on SweetBench out of the box for us. It is definitely a state-of-the-art experience. Zach, I want to stick with you for a second. Andrew, I'm going to come back to you. You know, as we saw this in our last segment, we were talking to Vercel about sort of this discussion around model effectiveness sort of plateauing and the idea that we made so many big jumps, right?

17:46It's now hard to make everything that much better every single time are you seeing sort of a plateauing in the efficacy of the models that you're using zach i think if you look at the jump from like three five to three seven to four to four five that i think it's a slightly smaller delta i don't know what andrew would would think here but when we were using say like sauna three seven and it went to four it was like whoa things that were not possible before are now the model can do. I haven't spent quite enough time with 4.5 to know if it feels that same way. But it feels like we're getting a little bit closer to the edge.

18:29And I would say the limiting thing is starting to become more, does the model have the right context? And can it reason over your big code base? Does it understand your intent? Does it have all the right tools? So all of that. Yeah, contacts basically for people, you know, who might not be coded, basically, do they have access to the right data, essentially? The right data, does it like understand how your organization works? There's all these other things that go into, say, great software engineering that is not just a question of like pure inference. And so that's a place where I think we're starting to, that's a little bit more of the limiting factor at the moment, in my mind.

19:06Right. And so, so Andrew, I mean, it sounds like Zach is saying, Hey, I mean, yeah, the, the efficacy of the model is one thing, but it's, it's not actually the model now that is sort of the rate limiting step. It's like, how else do you build the other parts around the model to sort of get the most out of it? Is that, is that the way you understand it? To me, it's a multiplier. There's the model itself and how good is the model. And then there's what we call harness, like how good is their, all the tooling and applications that use the model and, and both their successes and their failures kind of multiply uh so if you got bad model and bad harness it's terrible and and then vice versa you've got also model and awesome harness it's great right and then what happens is they both also move upstream so the models start to do more of the job that harnesses did before um and they're getting trained on that um but then right harnesses are starting to get higher levels so for example we started working on orchestration of multiple agents and whatever and how do they pass over information like how do you manage that whole fleet, right?

20:02So it's all moving up. And from that perspective, I think it continues to level up, just like, you know, human brain was probably pretty much the same as it was 20 ,000 years ago, but we as a civilization now are much more capable. And I think the same thing happens with the models and harnesses and kind of all the ecosystem. Well, I'll tell you what, you know, I think we, I know for myself, I got a little more Googling to do about some of these termers, but I want to thank you both for coming on. These models are coming fast and furious, and it's always great to get people on the ground using them to tell us how it is they are liking them.

20:38That is Zach from Warp and Andrew. Thank you so much. Really appreciate your time. Let us get on to our next segment, folks. XAI has been a fascinating company to follow, especially with the latest acquisition of X, Elon's social media platform, of course that is, And all of that has meant big changes for the company's org charts. Now, if you follow the information's coverage, you know that we love to geek out over company's org charts over here. And this week, we published an exclusive look at what XAI's company structure currently looks like. I want to bring on Theo Waite, who covers all things Elon Musk to tell us more about what he's found.

21:17Theo, welcome back to the show. It's great to have you. Thanks for having me. You're dressed up today, man. You got the AI Agenda Live Summit. Look at that. See you there. I'm wearing the same blazer, man. I wear the same five blazers over and over again, Monday, Tuesday, Wednesday, Thursday, Friday. Okay, let's talk about XAI. You used to cover Amazon, okay? And look, Amazon, it's one of the biggest companies in the world, right? I'm sure they have a lot of org chart infrastructure in place. It's a conglomerate. You're now covering XAI, which I take, I mean, the org chart must be, it looks so different.

21:53So what I want to know is, what were your observations from XAI and how do they differ from like a traditional tech company that we think of? It's pretty refreshing covering XAI and working on the org chart compared to Amazon, I have to say, because at Amazon, there's a lot of, you know, agonizing over director versus VP versus L7 versus L5. Like there's so much hierarchy and so much terminology that just doesn't exist at XAI. A lot of times people don't even really have a formal title besides a member of technical staff. And it's just not a very hierarchical company in the same way, just largely by virtue of its size, but also because of Elon's management style.

22:40There are some people that report to Elon that have, you know, hundreds of direct reports or not direct reports, but hundreds of people under them. Other people that report directly to Elon that have zero, but are just, you know, an engineer working on a project that he really cares about. And so he, you know, wants to speak with them directly. It's all very impromptu. So, you know, impromptu means people are getting promoted and people are getting demoted, I guess, very quickly. Let's talk about sort of who is in his inner circle right now. Who are the people that we should have on our radar for the XAI inner circle?

23:14Sure. So the two most important research scientists that are, you know, they call them engineers, but they're research scientists at XAI right now are Tony Wu and Guedong Zhang, who are former research scientists from Google who were promoted by Elon this year, kind of at the expense of some other co-founders of the company. um but when Elon you know pays a lot of close attention to to one of his products had to start to roll and in this case they they came out on top and and so so these are people you know I was looking at the terminology you know these are people who are technically called members of the technical staff right yeah that's right but that you know in reality they're they're managers now right and what about people who you know may have recently been you wrote about this this uh guy, Jimmy Ba.

24:05What do we need to know about him? So he used to be in charge of the majority of the engineers at XAI, or at least, you know, a huge number of them. He kind of got demoted, it seems, over the summer when Elon took a lot of his responsibilities away, although he is still, you know, in Elon direct report, and he's in charge of the enterprise business. And then there's another guy, Igor Babushkin, that left the company entirely a couple months ago after Elon also took some responsibilities away from him. So with all of these changes happening, what is, do we have any reporting on the culture of XAI?

24:44I imagine these are people that, you know, are really burning the midnight oil trying to make Elon Musk happy. Yeah, it's incredibly intense based on everyone I've talked to. I mean, you know, on the one hand, I don't think it's an industry you get into for work-life balance but i i do think there are also you know some demands from elon that are especially uh galling or intense compared to open ai or meta um and when you know a company like open ai or meta tries to poach these people and offers them you know much larger paychecks like uh you know i i think the money is probably a bigger consideration for them at that point especially if they're like 10xing your your pay or something like that yeah what about x i mean linda yaccarino left i we haven't heard anything about who's running x right now did you find anything about you know interim leaders or sort of piecemeal solutions to leadership right now we have a kind of tbd space for the ceo of x because uh there is not one and it's unclear uh what the long-term plan there is.

25:53But, you know, at the moment, there's this trio of executives, Monique Pintarelli, John Nitti, and Angela Zepeda, that are kind of traditional media and marketing types that joined under Yaccarino. And they're, you know, running X for the most part at this point. But it really does feel like, you know, a pretty secondary part of the company at this point compared to XAI. I I just don't really get the sense that it's what Elon is spending his time obsessing over this year. Right. Great. Well, Theo, look, thank you so much for coming on and talking to XAI. I know that you're going to be moderating our robotics panel at AI Agenda Live this afternoon.

26:33I'm excited for that one. And yeah, we'll have you back on the show maybe to talk about that later on with our favorite robotics correspondent, the information Rocket Drew. I mean, he's also an encyclopedia. Thank you, Theo, for coming on. We appreciate it. See you this afternoon. Okay, our next segment is with our presenting partner, Amazon Web Services.

Read the full transcript

27:00There has been a lot of talk about AI agents lately, and one of the biggest moving targets is really what the right definition of an agent is, what things agents are good at, and more importantly, what they're not good for. To talk about that, I want to bring on Shao Nandi, a director at AWS. He is also the former CIO of Dow Jones News Corp. I'm really excited to have him on. Sean, welcome back to the show. It's great to have you. It's good to see you again, Nakash. So we talked last time about the chief AI officer. Today, I want to talk about agents. And one of the things I want to get your help on is, look, we've talked a lot about agents on the show, and I've been struggling to sort of understand really what the difference is between sort of old school automation that we've, you know, read so much about over the past decade, it seems.

27:47And now we have this new paradigm of agents. I mean, talk to me about how these are different and how you think about what an agent actually is. Yeah, no, it's a really fun topic, actually. Let me start with sort of a formal definition from when AWS thinks about agents. We look at them as autonomous or semi-autonomous systems that can reason, plan, and act to accomplished goals. They're not just answering prompts, but they're executing tasks, adapting to context, and driving outcomes. And what helps is thinking about how they differentiate from all of the generative AI action we've been seeing for so long.

28:23And generative AI, of course, generates amount of text, codes, images. Agents are really a step change. They don't just create, they do things. Jet draft a business proposal, update a CRM system, schedule the next meeting, kick out the workflow without you doing all the clicking. And you asked like sort of the difference between what we've seen in the past, robotic process automation, all those pieces. What I'll say is if you think about the old follow concept, which was RPA, really you had a rules-based engine and you had to sort of follow a pre-built path. It was very brittle. Of course, we've been seeing the assist for the last couple of years, co-pilot type activities, sort of predetermined goals.

29:04but now we're getting to collaboration. That's the goal defined agent. So we have that sort of definition of an agent. What are some ways that you've been helping customers work with agents or some examples about how you've been using them in your own business? Yeah, I mean, look, so many. I mean, you said own business. There's productivity. There's new product innovation with Agentic. There's helping solve problems faster. That's not just about productivity. Amazon, you know, we are doing a lot with agents and we just announced enhanced seller agents. Now, seller agents are really in our marketplace.

29:38And this is not AWS. This is big Amazon retail. And you think about all the third-party sellers selling products on Amazon. We wanted to help them automate catalog management, do better pricing, have product descriptions. And we initially released a Gen AI-powered seller support tool that would create images or identify images and reformat text. Now it's actually helping them do tasks, do research, reason for them. And that's a huge productivity boom for all those people selling on Amazon. But it also, even more importantly, can improve quality, which turns to a better customer experience. Right.

30:14But you do have to, I just want to make sure I understand it, you do have to prompt them, right? So, I mean, how does it work? Yeah, it's not just magical, right? So it might prompt you to say, give us some initial information, give us a sell sheet, give us some images, give us a link, give us some information. But it's going from simply sort of generating suggested copy to actually going and helping you research and interacting with you. It's a really much better experience. And, you know, love to have people go out and see that. We've been writing blog posts about it. But of course, those who actually sell on Amazon, okay, most of us aren't selling on Amazon.

30:47They're actually using this. And this has to work at scale, Akash. That's the difference. We have millions, tens of millions of sellers and third-party sellers. You can't just do this for one or two people has to be bulletproof. I mean, that's sort of our case. I could tell you about a fun customer case. Sure, yeah. I'm always down for a good story. So I'll tell you what. So XCIO here, you mentioned that. Dan Jones, yeah. Yep. Network outages, connectivity outages, all is a pain point. You know who it really matters for? One of my favorite sports, Formula One. So for those who follow Formula One, you know, high-speed racing worldwide, happens everywhere.

31:24And you go from race to race. There's a race in Mexico. The next week, there's a race in Austin. A few weeks later, a race in Singapore. And they broadcast all these on channels like F1 TV. So when they have a network problem, it's disastrous, right? Like, you got to be able to broadcast all those cameras, all those angles. Right, right. The problem resolution during setup, when they're setting these sites up, it would sometimes take 15 engineering days to investigate why the setup was coming. So it had to start way in advance, three weeks in advance. F1 built a really intelligent root cause analysis agent for networking outages at race sites.

31:59They connected multiple systems. They do natural language troubleshooting. It drove an 86 % reduction in resolution time of issues. So triage time went down from one day to 20 minutes. And that's just incredible. And so this is sort of the agent sort of finding the problems for you. I want to talk a little bit about how F1, decided to focus agents on that particular problem? Because it's really hard. I mean, you could apply agents to anything and everything, seemingly, but there appear to be better use cases, better applications than others. How did they land on that particular problem? So I'll give you these sort of pressures from both sides.

32:41You know, like how do you select? How do you select what use case, what situation? It's when you need intelligent decision-making. You need better automation and you have complex dynamic situations. And network outages, while not always the most fun-sounding thing, there are so many possible areas that could be coming from. Physical fiber breaks, bad configuration. It's really complicated. It's really hard to process all that information. That's low-hanging fruit for an agent. However, you're not probably going to use agents in the most high-stakes, zero-error-tolerance situations yet. You're not going to use them, for example, to give real-time feedback to a driver on the cockpit.

33:15You're going to use sort of your classic sensors for that, right? But the network outage, this was research and resolve. You already have a problem. You want to make it solved quicker. There's no risk being added by this. There's only risk reduction. And of course, there's still a human in the loop. The other element, and this was less true for F1, but this is for everyone listening, you have to sort of make sure you're organizationally ready, that you have the right partner to work with on building these, that you have a scalable infrastructure. And once you have that in place, you can experiment more broadly.

33:46You can go to production more broadly. But F1 had selected us. They're working with us. Great for us. They already understood how to apply this. That was some of the mental model. I can go deeper on some of those situations. No, I mean, I think that, you know, I hear you on sort of the how much risk there is to implementing it and whether or not that makes sense. I hear you on the leadership piece. The other piece that we've been talking about a lot in the show is the ROI and what the business case is for agents. How do you sort of think about that? Is this is it all sort of cost savings and quantified in terms of time?

34:20Do you sort of say that, hey, this is doing the work of 10 full time employees? And, you know, that's that. How do you think about that? I'll give you one example that's not cost savings based. And then we'll talk about cost. Because most of your listeners are really worried about ROI and cost, right? All of them. I think all of them. The not cost one is Alexa Plus. You're going to see a ton of devices announcements this week from Amazon. Really exciting things we use. And Alexa Plus is the core of them. And that's really a new product for us, right? It's customer experience. It's Alexa being able to do things that it's never been done before.

34:52It's been completely rebuilt. We're connecting LLMs with thousands of services. is it's just new product creation. But for most of us, it is what you mentioned. It's efficiency, it's speed, it's customer experience. And for the ROI, I'll tell you the basic advice I give leaders all the time. Think big, start small. So don't get yourself in terms of what's possible. When you start with a scenario, pick something where you have a measurable outcome. So for a call center, a contact center, we have a great product set, Amazon Connect. And one of the things we look at when we're deploying Connect or any sort of call center, contact center experience is really around what are we looking to accomplish?

35:36Are we going to reduce supervisor escalations? You know, customers hate supervisor escalations. It sounds like that experience, but it's actually expensive. Supervisors are hard. They cost a lot. So you have a defined metric. I want to bring escalations down by 35%, and it will result in X. This cost savings. Then you can look at the cost of the agentic solution. And I'll give you a hint. build cost isn't a big deal lately. Like you can build these agents very quickly and easily. So what ends up being the most expensive part then? Running it. And by the way, people aren't stressed about running it yet because they deploy an agent.

36:09They're running it for like 20 agents, 10 agents, five people, right? Fine. But if they're successful, consumption is going to go through the roof. So we're spending lots of time saying to customers, build quickly, experiment quickly with those measurable metrics. But then let's work together on how to really make sure you've picked the right model. We just launched a whole stack of OpenWay, aka open source models, over the last several weeks. We have the latest Anthropic models announcing yesterday, running in Bedrock, Anthropic 4 or 5 model. And picking across the different models for the one that makes sense for your use case, when you've built a really narrow agent, and I could go on for an hour about this, Akasha, I won't, I promise, you might need a really simple model to answer that question fast and quickly.

36:51And that costs an incredibly different amount of money than a full reasoning model. That's why it's not like one model to solve them all. Right, right. Last question before we let you go. You know, as it relates to agents, what are some of the technical challenges or the breakthroughs that still need to happen? You know, I'm talking about, you know, things that are high up on your list that if we could really just figure this one thing out, not just at Amazon as an industry, right? You know, this would really propel agents forward in terms of functionality. Explain what those are to us and try to keep it in layman's terms for us so we understand that.

37:25Yeah, of course. Look, I'll tell you what's amazing right now. We're all seeing the movie trailer. Like we're all seeing the art of the possible and getting super hyped up. And much of it is possible. It just needs attention and a way to run it resiliently. And I think one of the things that's really mattering is as you're building agents, how are they going to communicate with each other? How are they going to have the right authentication? Like we announced in Preview a series of services. We call it AgentCore. There's a bunch of great open source solutions out there like Crew AI and Langchain that let you run agents at scale.

37:54But customers being able to do that easily, it sounds less hot to talk about evolution. But you really want to drop an agent and know how it's going to talk to your neighboring agent using open source protocols like A to A, what authentication looks like. Having that all be set up and easy, that's pretty critical. And we put the groundwork in place, but getting it to executed scale. And the second part, I think you're going to laugh, but it's about us. It's getting the organizations educated and ready. And getting them on board. Yeah. Just like you said, getting leaders to actually buy into the idea that, okay, this is how the business is going to run.

38:30The tooling is so capable, Akash. Even now, it's moving faster than we can keep up. So we haven't taken advantage of what's available today yet in most cases, much less all the great things that are coming over the next three months that I can't wait to tell you about. Great. Well, Sean, I'm excited to have you on again to talk more about those things. It was a fascinating discussion about agents. Thank you again for the time. That is Sean Nandy, the director at Amazon Web Services. And with that, folks, that does it for today's show. A reminder that we are live on this stream Monday through Friday at 10 a.m.

39:04Pacific, 1 p.m. Eastern. I want to thank Amazon Web Services, who is our presenting sponsor for this production. And I want to thank you for tuning in. We really do appreciate your viewership. I am already excited for our next show tomorrow. Again, we're going to talk about our AI Agenda Live Summit happening this afternoon. Stay tuned on our website for more coverage on that. And with that, have a good night. I'll see you tomorrow.

39:30Thank you.

From the publisher

Vercel's COO Jeanne DeWitt Grosser, talks with TITV Host Akash Pasricha about the company's $300M series F funding round and its aspirations to become the "AWS of AI." We also talk with Warp's Zach Lloyd and Zencoder's Andrew Filev about their first reactions to the new Claude Sonnet 4.5 model. The Information's Theo Wayt breaks down the latest xAI org chart shake-ups, and we also get into AI agents with AWS's Director of Technology, Shaown Nandi.

Articles discussed on this episode:

https://www.theinformation.com/articles/people-running-elon-musks-xai


TITV airs on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.


Subscribe to: 

- The Information on YouTube: https://www.youtube.com/@theinformation4080/?sub_confirmation=1

- The Information: https://www.theinformation.com/subscribe_h


Sign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agenda

More from The Information's TITV

All 304 episodes
Claude Sonnet 4.5 Review, AWS Director on AI Agents, Vercel’s $9B ValuationThe Information's TITV · 40 min
Listen in VO