In short
ROI and pricing for “agentic AI” in enterprises—whether AI agents can reliably resolve customer issues cost-effectively, and how companies should budget and pay for outcomes.
Guest backgrounds
Clay Bavor is co-founder of Sierra, an agentic AI platform focused on customer service/support, sales, and recommendations.
Key claims
Agents are software that reason, decide, and take actions via tools without hardcoded if-then-else logic. Hype is real but warranted: coding agents saw a step change (e.g., Cloud 4.5/Codex 5.2). Scaled deployment requires guardrails: supervisory agents, escalation to humans, and post-conversation monitoring. Sierra uses “constellation” orchestration (multiple models, sometimes fine-tuned) and “Ghostwriter” to build agents, then runs thousands of simulated conversations (including messy voice conditions). ROI improves via fully autonomous resolution and outcomes-based pricing (pay only when the agent resolves the job). Token “maxing” is driving CFO scrutiny; token budgets will be allocated like other role budgets.
Notable examples
Rocket Mortgage agents gathering borrower/asset/credit info and originating mortgages; Cigna deployed behind a toll-free number in ~58 days.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOFrustrations with AI Customer Service
0:00 to 0:31
Discuss the common frustrations experienced with AI in customer service.
“Have you ever had that really frustrating experience where you're on with a customer service agent, which is basically an AI, either on the phone or via text, and it can't answer your question.”
Understanding Agentic AI
0:45 to 1:59
Exploration of what agentic AI means and its relevance in the business context.
“And effectively, it is a term that encompasses more advanced AI that is able to do more complex tasks, longer form tasks on your behalf.”
Introduction of Clay Bavor
1:59 to 2:21
Host introduces guest Clay Bavor, co-founder of Sierra, and their background.
“We're going to talk about that and much more with the co-founder of Sierra, Clay Bavor.”
Defining AI Agents at Sierra
2:21 to 3:16
Clay explains what AI agents are and their functionalities at Sierra.
“Clay, thanks so much for joining me here in London on the Tech Download.”
The Reality of AI Agents
3:16 to 4:56
Discussion on the current capabilities and limitations of AI agents.
“So first of all, what we do is help companies build customer-facing AIs, as you said, for all parts of their customer experience.”
Technological Shifts in AI
4:56 to 7:08
Exploration of recent technological advances that have impacted AI agent functionality.
“5.2, where all of a sudden you had agents that could build complex software that could perform at the level of like a staff level software engineer.”
Sierra's Approach to AI Models
7:08 to 8:31
Discussion of Sierra's strategy in using various AI models for different tasks.
“And so I think it's not been any one thing.”
Implementing Agents with Clients
8:31 to 11:23
Detailed overview of the process Sierra follows to implement AI agents for clients.
“My other fun word at the moment is orchestration.”
Monitoring and Improving AI Agents
11:23 to 13:36
Insights into how Sierra monitors performance and improves AI agents post-deployment.
“And the process of testing is really interesting because unlike deterministic software where you put an input in, you get an input out, conversation is messy.”
The Future of AI in Customer Service
13:36 to 14:00
Discussion on the balance between human augmentation and full automation with AI agents.
“And anytime you identify an opportunity for improvement in one place, you make that change.”
Show all 21 chapters
The Impact of AI Agents on Customer Inquiries
14:00 to 15:00
Learn how AI agents can significantly resolve customer inquiries, enhancing efficiency and satisfaction.
“I think many companies think, oh, well, we'll dip a toe in the water with augmentation and we'll help our team be a little bit more productive.”
Outcomes-Based Pricing in AI
15:00 to 17:00
Discover how outcomes-based pricing models align costs with successful customer outcomes in AI services.
“What does that look like for some of your customers?”
Trends in AI Spending and Adoption
17:00 to 19:00
Explore recent trends in AI spending and the challenges companies face in measuring ROI.
“Seat-based, consumption-based, and so on.”
Token Maxing and Corporate AI Usage
19:00 to 23:00
Understand the implications of 'token maxing' as companies increase their AI usage amid budget concerns.
“Companies urging employees to use AI, employees racing to use AI, whether it's productive or not.”
Shifts in AI Pricing Models
23:56 to 28:00
Analyze how the shift from seat-based to outcome-based pricing affects AI service implementation.
“And I want to just get into one of the terms that we've spoken about.”
Balancing Model Performance and Cost
28:00 to 29:56
Explore the trade-offs between advanced AI models and practical applications in businesses.
“And so one of the things I think our customers appreciate about working with us is they don't have to think about token usage and runaway token.”
Understanding AI Model Selection
29:56 to 32:13
Learn how companies choose the best AI models for different tasks to optimize efficiency and costs.
“And there'll be something of a dispersion of applications and what models are used.”
AI's Impact on Traditional Software Business Models
32:13 to 34:29
Discuss the implications of AI on traditional software models and the challenges companies face.
“And so what these companies do like Sierra is they focus on a layer known as orchestration, where their systems effectively are choosing the best model for a specific task.”
Valuation of Frontier AI Companies
34:29 to 36:55
Examine the unprecedented valuations of AI companies and the potential for future growth.
“Systems that store data, project management systems, right, where, okay, a project may last for a week, a month, I think somewhat more challenging.”
Future Developments in AI Technology
36:55 to 39:54
Evaluate the next steps in AI development, including voice models and self-improvement.
“And so what these, to me, what the valuations of these companies are predicated on is that intelligence, the ability to think, solve problems, invent, discover is immensely valuable.”
Exploring Recursive Self-Improvement in AI
42:00 to 44:10
Learn about the concept of recursive self-improvement in AI agents and its implications for the future of AGI.
“A couple of things, the voice improvements.”
Transcript
Automatic transcript. May contain errors.0:00Have you ever had that really frustrating experience where you're on with a customer service agent, which is basically an AI, either on the phone or via text, and it can't answer your question. And you're just constantly asking loads of questions. You're not getting right the answer. And eventually, I just give up and I just type in human, get me a human. Well, this is one of the areas where more advanced AI could have a big impact.
0:30Hello, everyone. I'm Arjun Karpal, senior technology correspondent at CNBC, and welcome to another episode of The Tech Download. Now, I'm going to throw another word at you, which you've been hearing over the last few episodes, agents, agentic AI. I'm sorry for keep bringing up, but we've got to start here because it is one of the biggest topics right now in AI. And effectively, it is a term that encompasses more advanced AI that is able to do more complex tasks, longer form tasks on your behalf. It's something that all companies are talking about right now. The question is, can these agents, these AI agents, deliver value?
1:08And that really is the big questions. Can they give you good business outcomes at a good price and, of course, help companies make money? And really, that's a big focus and conversation right now in the world of enterprises as they consider whether and how to adopt AI. The huge focus right now is on cost because we've seen these stories of companies that have blown through their budgets dedicated to AI because people are just using it so much. But the problem is they're not necessarily seeing that return on investment. The stakes are pretty big here if you think about it. If AI can't be delivered and given in a cost-effective manner to the point where businesses are seeing the return on investment, then it won't get adopted.
1:55And that could have implications across the entire tech ecosystem as well. We're going to talk about that and much more with the co-founder of Sierra, Clay Bavor. We're going to talk about pricing, what agents can do right now, where they're going in the future, customer services as a battleground and so much more. and I started by asking what AI means in Clay's view.
2:21Clay, thanks so much for joining me here in London on the Tech Download. It's so great to be here. I'm back in one of my favorite cities in the world, lived here in my mid-20s, and so it's nice to be back. Yeah, we first met in Singapore. Now we're here in London. No waterfall in the airport. Yeah, just for some context for our listeners, CBC hosted an event in Singapore called Converge Live, And it was in the Jewel, which is this massive fountain in the airport. And that was the backdrop for our interview, which was pretty amazing. Not as snazzy, but certainly very fun. This is all brand new, by the way.
2:51This is great. This is a great setup. So, Clay, let's get into the conversation. Because Sierra, for people who don't know, is an agentic AI platform that is focused on customer service. And one of the agents in agentic is such a buzzword right now. And with buzzwords, we get in technology. There's lots of different definitions around it. When someone says to you, AI agents or agentic AI, what does that mean at Sierra? So first of all, what we do is help companies build customer-facing AIs, as you said, for all parts of their customer experience. So customer service and support, definitely, but also sales and product recommendations.
3:27We originate mortgages. We can send satellite signals from space to help you get set up on a new vehicle, satellite radio. So as I think about agents, they're a new type of software. To your point, it's the buzzword du jour. Under the hood, as I see it, agents are software that can reason, that can make decisions, that can take action using tools, and that can understand and generate language fluently. And the whole point is you can define what you want an agent to do, and then it can go do that autonomously or semi-autonomously without kind of the need for a pile of if-then-else statements where you've deterministically programmed everything.
4:11How much are agents being overestimated right now? Because I think that, you know, the view I've had is, you know, we can just let these agents go off and do all these things for us. And which sounds amazing. You know, if there was an agent who could answer all my emails, wow, that'd be incredible. And do it successfully? Fantastic. But, you know, that's what we're sort of hearing as the sort of promise of this technology. But right now, what is the reality? Yeah, well, I think it is funny seeing people with their open claw setups and their Mac mini arrays and I'll let my agents do anything. And yeah, I probably wouldn't recommend that.
4:45But I do think some of the hype around agents is warranted. There was, for instance, a real step change in what coding agents could do in November and December of last year with Cloud 4.5, Codex 5.2, where all of a sudden you had agents that could build complex software that could perform at the level of like a staff level software engineer. That was a real aha. And I think wake up call for folks who are working in agents. I think in other domains like our own, the reality is agents are making an enormous difference already. And so we work with companies like Rocket Mortgage, one of the largest mortgage originators in the U.S., their agents are gathering income, property, asset, credit information, originating mortgage folders, making outbound calls, helping people find better properties.
5:35And the key in deploying these agents at scale is ensuring that you can do so safely. And so it's putting guardrails around them. It's having what we call supervisory agents looking over the shoulder of the primary agent to make sure that things are happening correctly. And if you have the agent architecture right and then these guardrails around them correctly, you can have agents doing a lot out in production. We'll definitely get into, I think, some of those guardrails because those are really important as well. But just before we do, what are some of the big technology shifts that have allowed us to get to this point that you're describing?
6:09Oh, it's such an interesting question. As I see it, there've been a few inflection points. Of course, there was GPT-3.5 and ChatGPT. And I think that showed the world, okay, language models can understand and generate language. GPT-4 was the first model that was capable of calling tools. And so think about giving a language model or an agent access to APIs and systems that it can use to look up information and get stuff done. And then one of the, I think, less appreciated breakthroughs over the last couple of years was in 2024, OpenAI released reasoning models, the O1 model. And these were models that were taught to basically think out loud to themselves.
6:50And it turns out when you do that, the model performance can go up almost arbitrarily. If you give it more time to think, the thinking gets better. And then finally, I think just the step change in coding agents enabled a new level of reasoning, decision-making, and so on in these models. And so I think it's not been any one thing. It's models that can reason and decide. I think GPT-4 was kind of a proto version of that. Gemini and the Claude family have gone on to similar advances. Tool calling, reasoning, and those all add up to just more and more sophisticated agents. And I think important to note for this conversation, Sierra, you work with all of these various different models that underline your product.
7:34One of the approaches we've taken, rather than kind of wrap a single model, we built a thick platform on top of what we call a constellation of models. And what that lets us do is choose the right model for the right task. And so for complex reasoning and language generation understanding, we'll use the frontier models, hosted frontier models like the GPT family, like the Gemini family. And then for other tasks where we have a particular viewpoint on how things should work, we'll fine-tune our own models. So, for instance, a subtle thing in voice agents is understanding interruptions. Are they trying to get a point in or did they just cough?
8:15Distinguishing between those two is very important in determining whether the agent should kind of keep going and explaining something or back off and listen. So we'll fine tune our own voice activity detection models and other smaller models to support our agents. I like that. Constellation. That's a nice word. Yeah. My other fun word at the moment is orchestration. Sure. That's another because it's quite visual. You can imagine a conductor. Yeah. Cue the violins. Yeah. And cue this model. Cue that model. Yeah. Yeah. That's pretty fun. So in the enterprise, when you go to a customer or a customer comes to you, OK, we want to bring on some of your tools.
8:53We need an agent for this. How does the process work from that discussion through to actually implementing this successfully? One of the things we're super proud of, first of all, is that we've been able to get through that process very quickly. And so, for instance, with Cigna, multinational health care insurer and health care company, we were able to take them from first conversation with them to live behind their toll free number in something like 58 days. So what happens in that? So again, when we're building these customer facing agents with our customers and partners, you really want to understand and define what excellent looks like in every touch point they have with their customers.
9:35So typically we'll sit with a technology team and their operations team to map out, okay, here are the key customer journeys that you want the agent to be able to support and perform. And then working closely with those teams, basically express those in our platform to capture, okay, here's what the agent should do. Here are the guardrails. Here are the goals it's pursuing, what it's allowed to do and not do. Here are the tools it should integrate with. And then one of the more interesting breakthroughs we've had in the past six months is a tool we call Ghostwriter. Ghostwriter is an agent for building agents.
10:13And so what we can use Ghostwriter for is actually providing a lot of the scaffolding for a company's agent. You give it some example transcripts. Hey, this is what great looks like. You can even give it a napkin sketch, like a literal sketch on a literal napkin of kind of a workflow you want it to implement and let Ghostwriter cook, and it will build out this first version of the agent. The next step is actually running the agent through a gauntlet of simulated conversations. So before the first live conversation, our agents may have hundreds or thousands of simulated conversations with basically virtual customers.
10:50That lets us test and harden the agent and improve its capabilities. And then we'll roll into live production, where we'll usually do a percent, 2%, 5%, and then roll it out fully. So that's kind of the process. It's about understanding what the agent should do, modeling that in our platform, and then incremental testing and refinement. So once you are satisfied in the test that this is working fine and this is okay to interact with real customers, that's at that point, then you'll start those sort of incremental steps to roll that out. That's right. And the process of testing is really interesting because unlike deterministic software where you put an input in, you get an input out, conversation is messy.
11:33In particular, voice conversations, different accents, different languages, different background noise. And so our simulations are able to capture all of that, including introducing dogs barking and street noise into the background so that the simulations accurately reflect what the agent is likely to encounter in the real world. What happens in a real life situation where the agent messes up? What are the steps taken to rectify that and then use that for improvement? Yeah, a couple things. So first of all, all of our agents as they're running, I mentioned this, have these supervisor agents. It's kind of like a Jiminy Cricket, like looking over at the agent's shoulder.
12:15Are you dispensing financial advice? That's not okay in the financial services setting. Are you dispensing medical advice? Not okay in the medical setting. And so the supervisor agents can kind of catch errors and send a response back to the primary agent and have it try again. If an agent realizes that, OK, I've kind of gone outside the bounds of what I'm able to do and I need to phone a friend, so to speak, the agent can escalate to a person and hand off seamlessly to a person who can pick up. And in that process, we're able to hand the person kind of a warm readout. Here's all that's happened.
12:54It's kind of a cheat sheet on the conversation so far, and they can pick up from there. So are those potential errors picked up before they happen or as they happen, and then that sort of rectification process happens? Yeah. In the vast majority of cases, as they happen or before they happen so that we can handle them properly. And then after the fact, if we see opportunities to improve the agent, we have a whole system of kind of post-conversation monitors. Think of this as like the supervisor doing a detailed readover of every conversation and saying, oh, this could have been better. I would have said this a little bit differently and things like that.
13:31Those can then be incorporated into the agent's behavior forever into the future. One of the neat things about agents, in particular those built on Sierra's platform, is you can define them once and they can show up on every channel, meaning WhatsApp, text, phone calls. And anytime you identify an opportunity for improvement in one place, you make that change. It appears forever into the future and across every one of these channels. How much are companies looking for augmentation, human in the loop versus, you know, complete replacement? There's a spectrum we see. I think many companies think, oh, well, we'll dip a toe in the water with augmentation and we'll help our team be a little bit more productive.
14:15I think the more forward-thinking companies really are going straight to these autonomous agents. The impact of an AI agent that can help fully resolve 50%, 60%, 70%, 80 % of all incoming customer inquiries, whether it's, hey, how should I think about what mortgage might be right for me? Or, hey, I'm having a bit of trouble with my speaker system. Help me troubleshoot it. The difference between that and helping someone be 10 or 15 or 20 percent more productive, it's very significant. And so the ROI I think companies are seeing from deploying fully autonomous agents to fully resolve customer cases, not only is the ROI financially very significant, the experience for customers is so much better as well.
15:03You never have to wait on hold, right? You're never transferred and so on. And with the ROI question, are companies sort of having to figure out new ways of calculating what is a good return on investment in this age where, you know, there may be different pricing models that are not traditional SaaS models and sort of what success looks like? What does that look like for some of your customers? At least in the AI space, we have really pioneered what we call outcomes-based pricing. And our view is agents represent a new type of software, software that can actually fully get a job done. And you should only pay when it gets that job done and gets that job done well.
15:44And so in the service and support setting, as an example, a phone call here in the UK might cost 10 or 20 pounds for a company to handle. It's extremely expensive. And what we do with outcome-based pricing is we say, OK, if a Sierra agent fully resolves a customer conversation, fully resolves a problem for you that would have cost 10, 20 pounds, we'll charge some fraction of that. You save money. That's the only time Sierra is paid for its services. And what's neat about the model is it deeply aligns incentives. It's like we only win when our customers win. Our success is aligned with our customers' success.
16:25And one of the things we've seen is in our proofs of concept, over 90 % of all of them convert to long-term partnerships because the ROI case is so clear. You don't need to build a fancy model. It's, okay, I only pay when I'm saving money or making money through an upsell or a retained customer. And so I think with outcomes-based pricing, actually the ROI math has gotten simpler, not more complex. So this is very obviously different to sort of the recurring revenue subscription-based models we've seen on software businesses traditionally, right? Seat-based, consumption-based, and so on. I think there's an analogy in online advertising, actually.
17:06We started with kind of dating myself here a little bit, but taking over the Yahoo homepage, then it was CPM, then it was CPC, cost per click, and then cost per conversion. And I think as you get closer to the actual source of value for a company, the more aligned you are with the company that you're serving. We really like that. When you think about the sort of metrics, is it, let's say that the phone call you were saying sort of cost 10 or 20 pounds, let's say Sierra does it for a pound, but the company still loses money in that case. Would that, would they still pay you for that? Or is that sort of not a situation you see?
17:42But even if the outcome was resolved, so for example, you know, a customer wants to return something, obviously, you know, they're returning it, getting a refund. Oh, I understand. You know, something along those lines. I mean, it's quite a specific example, but I'm just trying to work through some of the examples. Maybe to zoom out a little bit, there are all sorts of different flavors of outcomes. And so with one of our media and subscriptions customers, the outcome there is a retained customer. And if we're able to retain a customer who is calling in and maybe thinking about canceling at a rate that's higher than their baseline, we're paid a fraction of lifetime value of that customer.
18:20And so we save them when they otherwise wouldn't have been. In the support setting, it's really, hey, there would have been a much higher cost to serve that. In the sales setting, it's if we're able to attach a premium product or upsell to an initial purchase. Again, we share in that value creation. And so I think we try to align what we charge with the value being created and the outcome being delivered. Yeah, really interesting. So just on this ROI and spending, if we just zoom out a little bit, because there's been a couple of interesting trends, I think, over the past year and then certainly the past few months.
18:59Token maxing. Yeah. One of those. I guess this is the idea of companies. There's a lot of that. Yeah, a lot of that. Companies urging employees to use AI, employees racing to use AI, whether it's productive or not. Yeah. But just to say, hey, we're using AI lots, right? And that's obviously contributing to the spending that's happening on certain AI companies and certain AI model companies as well. And there are some comments quite recently from Sam Altman, the CEO of OpenAIR, who was saying that actually this spending on AI has become a huge issue for enterprises more recently, certainly towards the start of this year.
19:39And so what's happening right now is companies look at AI spend and ROI, because I think we saw this wave of everyone getting some form of chatbot, right? Whether it's Microsoft Copilot, OpenAI, Cloud, whatever it might be. And everyone sort of using it and saying, this is fun. And companies rolling this out as part of enterprise packages, whatever it might be. And now there seems to be a sort of, right, everyone's had a go. But actually, where's the value? What are we getting out of this? from people using all these chatbots. And that feels like a harder thing to define. And that feels like there's a bit of a more of a scrutiny on that spend and companies saying, right, we need to really define if we're going to adopt AI, what it is we need to adopt.
20:25So what's happening with these various trends going on, token maxing, spending, and are we facing a reckoning of sorts from enterprise on spending at this point? Well, I think the inflection point in spend really was these coding agents. And I think what you're seeing there is a reflection of a couple things. One, there is immense value in these coding agents, just unequivocal value in these coding agents. And speaking from my own experience with Sierra, we have engineers who estimate they're between three and 10 times as productive, some as high as 20 times as productive in terms of pull requests, lines of code shipped, not a perfect proxy for productivity, but roughly maps to new feature development.
21:13That's an astonishing increase. I was like, that has never happened before. And so I think you see engineers, you see companies recognizing, my goodness, with Codex, with Cloud Code, I can fire up half a dozen agents in the morning with ideas for new features before my morning commute, come in, make a coffee and review the work done and ship six features before it's 9.30 in the morning. The other thing we're seeing is pressure from executive teams like use AI. And so, okay, use more AI, more tokens. And as a proxy for being an AI native company, I think token use or being an employee leaned into using AI, using a bunch of tokens is kind of the, okay, it's some, you know, marker of intensity and leaned in this to using AI.
22:07So I think the value is real. And also as kind of a blunt instrument for measuring AI adoption, you know, token usage is probably not the best. I think what we're going to see, in part because for software engineers in particular, the cost is so high. I think some engineers spending more than$100 ,000 per year on tokens. You're going to have companies start to budget token usage and allocate tokens to employees basically as a tool for getting their job done. Similar to how, okay, you've got this travel budget or this budget for this part of your role, I think companies are still catching up to how to budget and allocate resources in this new world where it is, on the one hand, extremely high value, these tools, also very costly.
23:00And I think you want people being discerning in where they're using their tokens. With a feature, just because you could build it doesn't mean you should build it. And so I think some editing, some discretion around where to apply these amazing but expensive tools is warranted. And I think we'll see that reflected in how companies allocate resources and budget.
23:25I'm Steve Sedgwick from CNBC. And in my new podcast, Executive Decisions, I ask powerful leaders about their decisions that changed everything. I'm not frightened of making tough decisions. And I think leadership can be very lonely. Business leaders should not stay quiet in a world which is super complex. It is absolutely fine to also change your mind. Tough calls, personal crossroads. These are the stories that we can all learn from. It's Executive Decisions from CNBC. Get it wherever you get your podcasts. Welcome to this first breakout session. And I want to just get into one of the terms that we've spoken about.
24:01And that is token maxing. Now, we've gone through this period where companies have been encouraging their workforce to use AI. for everything. And so people within organizations are using AI and effectively each instance they use AI, make a query, whatever, costs money. And they're using it so much as companies push their workers to use it that they're blowing through the entire AI budget of a company. And that's now led companies, CFOs at companies to really focus on what is the return on investment of our spending on AI. We had this sort of time, this fear of missing out where companies are like, we need to use AI, we need to adopt AI.
24:42And so they're all using it at whatever cost. Now there's this stepping back and this focus on return on investment. And that's interesting because if you think about enterprises, large enterprises, they often buy loads of different pieces of software or subscribe to lots of different pieces of software. And that's created this entire business model known as SaaS or software as a service. And effectively, that's run on what's called seat-based pricing. So something along the lines of how many users are using this, this is how we're going to price it. What's happening now, there's this shift, and we'll hear a little bit more about it from Clay, is there's a focus on outcome-based pricing, i.e.
25:24you only pay if there's a desired outcome. And so that's fascinating because that is really hitting at the heart of measurable return on investment on AI. Now it's not all in on AI necessarily, but more a refined focus on it. And as I think through this a little bit, the implications are quite interesting because if companies start thinking about the cost of AI, do they look at alternatives beyond the frontier labs? Now, when we talk about frontier labs, we're basically talking about OpenAI, we're talking about Anthropic, talking about companies like Google's right at the forefront of the latest AI development.
26:04But those are expensive models, AI models that these companies use. And there actually are a plethora of cheaper models as well. And we're going to get a little bit more into that. But ultimately, as the focus on cost happens, what does this mean for frontier models?
Read the full transcript
26:22I guess just thinking about this a bit further, the coding part of the equation also is it's a specific use case. Yeah. And it's a specific part of a company or certain companies, technology forward companies with lots of engineers are using these tools. But when you think about the broader economy, it feels like the first wave of agents has certainly hit coding. And that's without a doubt. And those early adopters being used, being used quickly. And that's where a lot of this usage has come from. But actually, when you think about the broader economy and agents across businesses, across so many different industries, it's not going to be the same, is it?
27:02In terms of you may have agents or agentic products like yours, which are relying on various different types of models, orchestrating those and figuring out what the best model is for a specific task. And that may not necessarily be the most expensive frontier model, right? That could be something that's cheaper and open source because it can do that lift. And that then I think changes the equation around cost, doesn't it? Yeah. In terms of the way you're pricing it is very different. It's outcome-based, but also the cost on your end in terms of calling up these models will be very different as well.
27:36And that's how you can serve perhaps more cost-effective solutions. I think that's right. I think the coding agents have converged around a consumption-based model because there's a really direct mapping between amount of reasoning, amount of token spent, lines of code written, and so a very direct connection there. I think in other domains that have reached the mainstream, I think AI for customer experience and sales service support and so on, AI for legal with companies like Harvey, the pricing models there will differ and may not be directly about spend on tokens. And so one of the things I think our customers appreciate about working with us is they don't have to think about token usage and runaway token.
28:22Like we take on the risk there and we're then able to optimize under the hood to use best model for performance, latency and so on. And again, we're focused on delivering the outcome, the made sale, the solved problem and so on. Are there implications and read through for the growth of these kind of frontier companies in a world where, you know, maybe a lot of enterprise tasks don't necessarily need the best models right now and the most expensive models? Coding, you know, certainly has been on that frontier side, but maybe other things inside of an organization won't need that kind of front. Is there a read through into what that means for the growth of the frontier models?
29:00In the near and medium term, we are likely to see just unbounded demand for higher and higher levels of intelligence. I think if any company could hire a software engineer that was 10x better, they would. And so as the models improve, I think there will be huge demand for not just coding agents, but agents in science and drug discovery and material science, in legal, in areas where there's kind of no ceiling to how much intelligence you can use. I think at the same time for applied companies like ourselves, there are tasks where you don't need to drive the Ferrari to the grocery store. You don't need a 10 trillion parameter model.
29:45You don't need mythos to help parse a warranty policy or determine what sweater would go best with that pair of trousers. And so I do think you'll see applied companies like our own being discerning in what model you use for what task. And there'll be something of a dispersion of applications and what models are used. I think for the labs, the rate limiter will be how much compute can they get access to, because I think we're likely to see almost unbounded demand for kind of the very highest quality tokens. Yeah, compute still remains the bottleneck, right? Very much so, you know, until we get, I don't know, data centers into space and fusion and so on.
30:33This is sort of a devil's advocate question. Given the advanced capabilities of these labs, why wouldn't they just dominate the agentic enterprise space? So one of the things we've seen with the labs is frontier models and developing kind of the core intelligence pretty different from the last mile in specific domains, whether it's legal, whether it's serving customers at scale and regulated industries. and what we found is that the last mile is more like the last 80 miles in a lot of this. And so you think about all of the things that you need to get right to have an agent pick up in 56 languages, in six different dialects of English, fluently with low latency in a regulated industry like insurance to handle a first notice of loss claim.
31:25There is a lot, a lot of things that need to come together there, well beyond just kind of the raw intelligence of the language model and even the agent framework. And so while I do expect the foundation models to get better and better and better at following and executing kind of open-ended tasks and agent frameworks generically to improve, I think it's the vertical specific domain specific knowledge and solving the 200 things you need to get right in that last mile where there's real space and need for applied companies that are very, very deep in those spaces.
32:07I just want to tackle the topic of AI models as well, because there's so much focus on AI models with the view that they are the be all and end all, but that's not really the case as we're seeing it now enterprises are not necessarily choosing ai based on their models they just want to see what ai systems make their processes more efficient allow them to perhaps see revenue growth allow them to to make cost cuts all sorts of different things and that's where we've seen the rise of companies like sierra for example like perplexity where their systems are based on multiple different models or a multimodal approach to AI.
32:44And so what these companies do like Sierra is they focus on a layer known as orchestration, where their systems effectively are choosing the best model for a specific task. If there's a model that actually doesn't require, or if there's a task even that doesn't require the very frontier models, then their systems will choose a cheaper but capable model to carry that out. And that's a trend we're seeing more of as we continue that discussion around the cost and the return on investment around some of these enterprises adopting AI. Now, the last part of this conversation really is focusing on the future.
33:22And you're going to hear a term I bring up and Clay brings up called recursive self-improvement or RSI, this idea that actually AI will be able to improve itself. Now, it's something that's been theorized about. It's something that's been spoken about. by the Frontier Labs. But the question is, how close are we?
33:43We've spoken about ROI and pricing, et cetera. The other part of this conversation is how traditional software businesses adapt in this world, because there's been a big concern in the market. Yeah. Saspocalypse. Saspocalypse. Yeah. Choose your funny word for it. Yeah. Yeah. And some of our listeners and viewers might've heard these words, the idea that AI is going to wipe out a lot of these companies that are built on traditional software models, these subscription-based services. What is your read? I think there's some truth to it, and I think parts of it are overblown. First of all, I think in general, companies that store data with a longer half-life, so data that is more durable, Systems of record that store data that's more durable.
34:29I think they're in a good spot. Systems that store data, project management systems, right, where, okay, a project may last for a week, a month, I think somewhat more challenging. The other area where I think we're likely to see more defensible kind of software business models is the closer a company is to modeling something in the physical world, the kind of safer it is from some of this compression. And so, for example, like ERP, you're modeling supply chains and this cascade of interconnected parts and components that all need to come together. That's not something that someone is going to vibe code over a weekend, even with the latest coding model.
35:13So I think some of it is overblown in that these systems of record, I think, are not going away. And yet, I think AI in particular on business models. I think some of these companies that have grown up with seat-based business models, I don't know that in some of these application areas, seat-based licenses make sense anymore. So how do you transition from a seat-based model to an outcomes-based model? That kind of feathering up and feathering down just from a business model transition standpoint, I think is pretty challenging. Clay, with the growth we've seen in the AI industry over the past few years, we've seen some very big companies form, right?
35:55OpenAI, Anthropic right at the top. And as we're recording this, you know, OpenAI has just filed confidentially for an IPO. Anthropic is gearing up for an IPO. Meanwhile, SpaceX is also getting ready to trade as well. We've seen these huge valuations on these companies, these mega IPOs coming to market. How do you assess the valuation of these almost unprecedented valuations of these frontier labs right now? What are they predicated on? We use the word unprecedented. I think we are in an unprecedented time and modern AI is an unprecedented technology. I mean, I've been reading the history of Unix.
36:34And remember, we're coming from punch cards, mouse and keyboards, typing computer code into a terminal. And all of a sudden in the last four years, we've had AI models that can think, reason, speak, generate images of anything, all emerge. You have intelligence that you have access to via an API. That's crazy. That's unbelievable. And so what these, to me, what the valuations of these companies are predicated on is that intelligence, the ability to think, solve problems, invent, discover is immensely valuable. I think that's true. It's the whole foundation of our economy, right? Smart people discovering new ways to apply information, to manipulate technology in service of humanity.
37:29And so these are companies, each of which is going to have a major role in kind of the full realization of artificial intelligence, making intelligence available at effectively unbounded scale. And downstream from that is better products and services, better drugs, better materials, better engineering, better science. and so what is the value creation in that? I mean, it's hard to imagine something with a higher ceiling than the fundamental ingredient in invention, discovery, and so much of human progress. I guess in the near term, there's going to be a sort of tension between these big companies in an unprecedented way, their valuations, the fact that some of them are still loss-making versus traditional public markets as well, which I think is going to be an interesting dynamic.
38:26how the public market investors react to these type of companies, how they value them, what they're willing to pay in the near term, right? That feels to me something that's going to happen. It is. I think in more ways than one, both the size of the IPOs and the nature of the business, where some of these companies are loss-making at IPO. I think the public markets are also able to value things correctly, right? It's a forward-looking bet on, okay, they can see right now we are in the largest capital buildout, I think, in history. I think we've now exceeded the railroads. And it's not only OpenAI, Anthropics, SpaceX making these giant capital investments, it's all of the hyperscalers.
39:12And so you see Google issuing new equity, you see Meta and others issuing debt to pay for all of this. And so I think what the markets are likely correctly seeing is, OK, you've got to make these huge investments in order to provide the compute for training and inference to deliver this unbelievably valuable service at scale. OK, there's some up from an investment, but with then potentially unbounded return down the line. Claire, let's focus on technology as we close out this conversation. We've seen some huge developments in technology over the past few years around AI. Over the next 12 months, what do you think are going to be the next steps on the sort of AI development journey?
39:53Gosh, first of all, things are changing as fast as they ever have. And whatever crystal ball I think anyone in technology felt they had three years ago, it's gotten cloudier. The headlights extend not as far as they used to. A few things that I'm especially excited about, One is the development of multimodal models, in particular voice-to-voice models. And so today, most voice agents take speech, transcribe it to text, reason about it, and then generate speech out the other side from synthesis. With a voice-to-voice model, you have voice in, voice out with some reasoning in between. And in our case, where more than 60 % of all of our conversations are voice-based, you have much lower latency, greater fluency, and other benefits, like not losing the tone.
40:49Like the difference between that's a great idea and that's a great idea is very different, right? But lost in today's model. So voice-to-voice models, I think, are going to be quite interesting. I think as these coding models become more capable by the month, I'm very excited to see how organizations transform software development. We've already made major changes to how we run Sierra. It's interesting to see how the bottlenecks have shifted around from writing code to now reviewing code. I think the next bottleneck is, well, how do you decide what features are worth building? So I think we're going to continue to see step change improvements in the coding models as well.
41:32And then finally, I think as agents get built out, agent frameworks, frameworks for agents to, in particular, self-improve. So evaluate themselves, kind of give themselves notes and improve in a kind of semi-supervised way. I think that's really, really interesting. software that can understand what a job well done looks like, evaluate itself, kind of coach itself and improve the next time, be better tomorrow than it was today. A couple of things, the voice improvements. I hope that still means we can record another podcast. Oh, definitely. Yeah. I won't send an agent. I won't send an agent. You'll be here in the flesh.
42:15Okay, perfect. And then the second point, self-improvement. So there's been this term recursive self-improvement RSI. Oh, yeah. Yeah. That's been, you know, capital R, capital S, capital R. That's it. And the idea here is that these are models that can almost autonomously improve themselves and effectively models training the next generation of models. Yeah. That's what's been been spoken about. Where are we in that journey right now, as you see it? So what I spoke about was a much more limited version of that, which is in, For instance, in our own platform, we have a research agent that can research how the agents perform better and then ghostwriter an agent that can build agents.
42:54And those two can kind of talk to one another and in a semi-autonomous way improve the agent over time. Recursive self-improvement in the labs is something entirely different. This is models acting as AI researchers, agents acting as AI researchers, and improving the model architecture itself, doing experiments, running training runs, and so on, and not being inside one of the labs. I don't feel I am an expert to speak to that. What I can say, though, is I think when very smart people from each of the major labs are talking about this thing, it is almost certainly going to be a thing. And so I do think it is not just hype.
43:41I think we are to use, I think, an expression Demis Hissabas used. We are in the foothills of AGI. and I think this recursive self-improvement, coding agents and then research agents that can kind of pathfind to more and more capable models, it's going to be an important part of the path from here to AGI and it does seem like we're approaching it. Clay, it's been a pleasure speaking with you. Thank you for joining us here in London. Arjun, great to be here. Thanks so much. Really hope you enjoyed that conversation with Clay Bevor of Sierra. You can get in touch with us at thetechdownloadatcnbc.com.
44:17You can get in touch with me directly. I'm on X, TikTok, Instagram, LinkedIn, wherever on the internet somewhere. You can find me at Arjun Karpal. Thanks for listening and watching and we'll catch you next time.
From the publisher
AI agents are designed to do more than answer questions. They are meant to complete tasks.
Sierra co-founder Clay Bavor joins CNBC’s Arjun Kharpal to discuss how AI agents are moving from demos into real business workflows, especially in customer service, sales and support.
Bavor explains how Sierra builds and tests customer-facing AI agents before they go live, why companies want clearer ways to measure AI’s return on investment and how outcome-based pricing could challenge the way software companies get paid.
The conversation also covers coding agents, rising AI token costs and why the hardest part of enterprise AI may be the “last mile” of deployment.
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.




