Sierra Co-Founder Clay Bavor on Making Customer-Facing AI Agents Delightful

27 Aug 2024 · 1 h 13 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Training Data - Episode with Clay Bavor

Podcast Title

Training Data Description: Join Sequoia Capital partners, including Sonya Huang and Pat Grady, as they engage with leading AI builders and researchers to explore the transformative nature of AI technologies in business and society.

Episode Title

Sierra Co-Founder Clay Bavor on Making Customer-Facing AI Agents Delightful Episode Description: In this episode, Clay Bavor, co-founder of Sierra, discusses the significant advancements in generative AI, specifically relating to customer service applications. He highlights the challenges and solutions for creating AI agents that enhance customer experiences while maintaining safety and reliability.

---

Key Themes and Discussions

Introduction

  • Host: Ravi Gupta and Pat Grady
  • Guest: Clay Bavor, Co-founder of Sierra
  • Overview: The episode delves into how generative AI is revolutionizing customer service, with a focus on Sierra's AI agents.

Clay’s Background

  • Previous Experience: 18 years at Google, leading product design and forward-looking technological projects.
  • Founding Sierra: Co-founded with Brett Taylor, inspired by developments in AI and a desire to enhance customer experiences.

The Evolution of AI in Business

  • Current State of AI: Generative AI is recognized as a transformative tool for customer service.
  • Challenges: Hallucination-prone large language models (LLMs) create concerns for trust in customer interactions.
  • Opportunities: High ROI potential in improving customer service experiences.

Sierra’s AgentOS

  • Definition: A platform enabling businesses to create branded AI agents for customer interaction.
  • Capabilities:
  • Managing customer inquiries.
  • Handling nuanced policies.
  • Supporting customer retention and upselling.

Engineering Challenges and Solutions

  • Importance of AI for AI: More AI can help in creating reliable AI agents by detecting errors in outputs.
  • Development Frameworks: Introduction of the AgentOS toolkit for building robust AI agents, featuring:
  • Data governance.
  • Integration with systems of record.
  • Complex task management.

Metrics and Deployment

  • Trustworthiness of AI Agents: Discussion on what tasks can be reliably assigned to AI.
  • Current Metrics: Customer satisfaction and resolution rates are high, with some companies achieving over 70% resolution of inquiries.

Customer Experience and Engagement

  • Experience Manager: A command center for monitoring AI interactions with analytics and learning mechanisms for continuous improvement.
  • Deployment Strategy: Emphasizes a collaborative approach, integrating company values, voice, and customer service processes into AI training.

Future of AI Agents

  • Advancements Expected: Anticipation of further improvements in AI capabilities, including multimodal models and voice interaction.
  • Long-Term Vision: Potential for AI agents to represent brands effectively and enhance customer relationships through personalized interactions.

Pricing Model

  • Outcome-Based Pricing: Sierra charges clients based on successful resolution of customer inquiries, aligning incentives between Sierra and its clients.

Closing Thoughts

  • Enthusiasm for AI's Future: Clay expresses excitement about how AI can amplify human capabilities and the potential for even more sophisticated applications in various domains.

---

Key Takeaways

  • Generative AI is the dominant force in transforming customer service.
  • The future of AI agents hinges on building trust through safety and reliability.
  • Sierra aims to empower businesses by enabling them to create tailored AI agents that resonate with their brand values.
  • Collaboration with clients is essential for successfully deploying AI technologies in customer service.
  • AI as a tool has transformative potential for both businesses and individuals, streamlining the path from idea to creation.

---

This summary encapsulates the primary discussions and insights shared in the episode featuring Clay Bavor, spotlighting the challenges and innovations in creating delightful customer-facing AI agents.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00One of the more interesting learnings from the past year and a half of working on this stuff is that the solution to many problems with AI is more AI. And it's somewhat unintuitive, but one of the remarkable properties of Lars Langez models is that they're better at detecting errors in their own output than in not making those errors in the first place.

0:41Joining us today is Clay Bevor, co -founder of Sierra. Before Clay started Sierra with his longtime friend, Brett Taylor, he spent 18 years at Google, where he started and led Google Labs, their AR -VR efforts, and a number of other forward -looking bets for the company. Sierra is allowing every company to elevate its customer experience through AI agents, and there is no one who knows more about what AI agents can do today, and what they'll be doing tomorrow than Clay. You'll get to hear about how pictures of avocado chairs help inspire the founding of Sierra, why the solution to problems with AI is often more AI, and so much more.

1:16Please enjoy this incredible episode with my friend Clay Bevor. All right Clay, listen, this is a funny start because we know each other so well, but So can you just tell everyone a little bit about yourself and just give us a background before we talk about the future of AI and what role Sierra is going to play on that? So first of all, I'm a Bay Area native. I grew up not more than four or five miles from here. So grew up in the Bay Area. I got to see the kind of .com bubble grow and then burst. Study computer science and then ended up right out of undergraduate at Google where I was for 18 years until last last March.

1:51And so at Google I worked on really every part of the company. I started in search and then ads for several years I ran the product and design teams for what is now workspace, so Gmail and Google Docs and Google Drive and so on. And then spent the last really 10 years at Google working on various forward looking bets for the company. Some hardware related like virtual and augmented reality, some AI related like Google Lens and other applications of AI. And then 15 months ago, left Google to start Sierra with a long time friend of mine, Brett Taylor. We met in our early days at Google where we both started our careers in the associate product management program.

2:30So he was, I think, class one. I was class three. And we met early on and stayed in touch in particular through a monthly poker group that in a good year would play like once. and met up a December of 2022 and just saw what was happening in and around AI and these fundamentally new building blocks that we thought would enable us to create something really special and start its year out of that. So that's the recap. Actually, I'm curious on that. And we need to get to what is here pretty quickly here, but just for fun, December 2022, very shortly after the JetGPT moment, how I guess what Who's the process like or how soon after that moment did you have the conviction that this is a sufficiently interesting new technology to build a company around?

3:18Can I introduce one thing that's interesting I hope you talk about? Before you actually, before the chat GPT moment, you had been telling me about how everything was going to change. I still remember distinctly him telling me, you don't understand, you're going to be able to talk about a scene that you envision and they're going to be able to make a movie out of you just talking about it. Do you remember you telling me? Yes. And so I'm actually very curious about this too. I had such a privileged seat at Google to see so much of what came out of that transport paper in 2017 and the emergence of early, large language models.

3:52So at Google, one of the first was called MENA or Lambda. There was a paper, I think, in 2020, a conversational chatbot for just about anything. And I remember even before that, getting to interact with this thing in a pre -release prototype. and having this uncanny sense that there was someone something on the other side of it, and that this was different. And another moment, I think it was mid -2022 when we had, I think it was the first or second version of Paul and Pathways Langers model at Google, was a 540 billion per -hambitr model, and we were testing it to see kind of how smart it was.

4:30and one of the surests and signs of intelligence is the ability to think and reason and metaphor and analogy. So we tried a few things and one which was pretty straightforward is we asked Paul, hey explain black holes in three words. And it came back without skipping a beat, black holes suck. And we were like, oh, you know, that's a pretty good summary. Also, like, you know, the model seems to have a sense of humor, which is cool. And the moment that really blew my mind, we asked, and I remember the answer verbatim, we asked Paul, please explain the 2008 financial crisis using movie references.

5:10And again, without skipping a beat, so the 2008 financial crisis was like the movie inception, except instead of dreams within dreams, it was debt within debt. Whoa. And we all paused. What is this? right? So it had understood basically the concept of CDOs nestedness of debt. Okay, what movie it includes nestedness of something else, inception, nestedness of dreams. So it's like inception. And we all thought, wow, this is something new and different. And then there are a couple other moments. I remember the first dolly paper came out and they did a blog post and people reacted a little bit to it, but for me, I remember one of the stars of the show was they asked Dolly to make avocado chairs.

5:58And so, they, and I know this sounds so odd, but here is a set of 10 or 20 images of chairs that look like avocados. It wasn't Photoshop. These images had never existed before, and yet the model seemed to understand similar to the movie reference metaphor, concepts of avocado -ness and chernice and put those together and create these images pixel by pixel. So we have avocado chairs at Instacart. Yeah, actually did, did you really? We actually did. Wow. We actually had chairs shaped like avocados. In related news, there were times where we were burning a little bit too much money. Those bags, yeah, those bags.

6:37So how to, how to good sense that something was coming, and in fact, the team, the team I was running at Google at the time labs was putting a lot of large language models to use in early applications there. And so how to hunch, chat GPT certainly clarified that hunch, but I think Brett and I both for several years had been tracking what was happening and just seeing, you know, first it was translation and better than human level translation that it was some of this language generation. And I think credit to OpenAI for doing the engineering work and data work and much more to make GBT3 turned into chat GBT where suddenly you could grasp this thing's full potential without knowing how to write Python and use their APIs.

7:22All right, so we're going to talk about where AI is going to come out of agents when we talk about customer service. But first, can you maybe just tell people a little bit about Sierra and what you and Brett have created? Yeah. So in a nutshell, Sierra enables any company in the world to create its own branded customer facing AI to interact with its customers for anything from customer service to commerce. And the backdrop for this is this observation that any time there's been a really significant change in technology, people interact with computers, with technology in different ways. And as a consequence, businesses are able to interact with their customers in entirely new ways.

8:00And you saw this in the 90s. The internet made the website possible. And for the first time, a company could have a sort of digital storefront and be present into the world, update its inventory with the click of a button and so on. In the mid to mid early 2000, 2005, 2008, if you were a company you could all of a sudden through ubiquitous social networks interact with your customers at scale and have conversations at scale. And in 2015, right after the rise of smartphones, right as a company you could put kind of a Swiss Army knife version of your company in everyone's pocket. And so like, I bet you have your bank's mobile app on your phone, probably on your home screen.

8:43So the last few years of advances in AI has for the first time made it possible to create software that you can speak to, right? Software that can understand language, software that can generate language. And most interestingly, I think software that can reason and make decisions. and it's made for really delightful conversational experiences like those that we associate with chatGPT. And so we think there's a big, big deal for how businesses interact with their customers. And you think about the difference between how we do some things today versus what you could do if you could just have a conversation with the business you're interacting with.

9:23Think about like shopping. You're in the market for some shoes, right? Or Pat maybe free you some new weights or something. You're a very heavy weight. I had a tiny, just tiny little one. And you're on the website and it's like, you basically have to imagine how the company's designer would have organized the product catalog. So okay, men's, men's shoes, men's running shoes, men's racing shoes, light weight, vapor fly. I can't remember the name and so on. So instead, with conversationally, I could just say, hey, I need some super lightweight running shoes, kind of like those ones I got last time.

10:01What do you got? And it's almost like I'm dating myself a little bit here, but like Yahoo directory, where you navigate through this hierarchical structure to find what you want. In contrast to Google, you explain what you want. And this takes it several steps further. And there's a quote from the head of customer experience that one of the companies we work with. She said, I don't want our customers to have to have a master's degree in our product catalog and our corporate processes. And to do a lot of things, you know, buying shoes is fairly easy on the spectrum of interaction, these interactions you have with companies.

10:36Imagine, you know, adding a new person to your insurance policy. Like, where do you go in the mobile app for that? How do you get that done? And your eyes just glaze over, right? And so the alternative, talking to an AI, and in particular, an AI agent, it's a technology around which we build Sierra, where that AI agent represents your company, your company, it's best, we think is really, really powerful. And even in, you know, we're 15 months old as a company, we've had the privilege of already working with like Weight Watchers, Sonos, SiriusXM, Olokai, if you're in the market for new flip flops, I strongly recommend Olokai Flip Flops.

11:18Very good. Excellent. Also make great golf shoes. Oh, really? Oh, yeah. Yeah, yeah. You should get some. Great. And so for Weight Watchers, we're advising on points and helping members manage their subscriptions. With SiriusXM, we're helping diagnose and fix radio issues and figure out what channel your favorite music is on. And so on. And the results, again, in the first year of the platform out there, were, in one case, resolving more than 70 % of all incoming customer inquiries at extremely high customer satisfaction. And all this leads us to believe that every company is going to need their own AI agent.

11:59And we want to be the company that helps every company build their own. In the spirit of sort of the, you know, the future of these AI agents and what they could mean for customer facing communications or customer facing operations, are there any good examples of things that were not possible 18 months ago that are possible today and then maybe if we roll the clock forward, things that are still not quite possible today that you think will be possible 18 months from now? Yeah. First of all, the progress month by month and over 18 months in particular is just kind a breathtaking. 18 months ago, GPT -4 class models didn't exist.

12:38It was still something just coming over the horizon. Agent architecture is cognitive architecture is kind of the way you compose large language models and other supporting pieces of infrastructure were very, very rudimentary. And so I'd go so far as to say like the idea of putting an AI in front of your customers that could be helpful and importantly safe and reliable, that was just impossible. And so chatbots from even 18 months ago looked a lot like a pile of hard coded rules that someone cobbled together over months or years that became very brittle. And I think we've all had the experience of talking to a chatbot.

13:22I'm sorry, I didn't get that. Can you ask in a different way or my favorite? My favorite is when they have the message box and then the four buttons you can click, but the message box is blanked out and you can't actually use it. And so I can help you with anything so long as it's one of these four buttons. So most of what I described, fixing radios, processing exchanges and returns and so on, wasn't possible, at least in any satisfying way or in a way that led to real business results for companies 18 months ago. Fast forwarding 18 months, I think we go pretty deep here. I think multimodal models are quite interesting.

14:05Something like 80 % of all customer service inquiries are on the phone, not on chat or email. So voice will obviously be a huge part of it. Things like returns, exchanges, diagnosing radio issues and things like that are on the simpler end of the spectrum of the total set of tasks that you might want to get help with from any agent. And so I think more advanced models, more sophisticated cognitive architectures, all of those I hope would increase kind of the, you know, the smarts in the agent, the types of problems that can solve. And then trust safety, reliability, you know, the hallucination problem, I think is still an unsolved area.

14:46Yeah, and we've made others have made huge amounts of progress on it, but I think we can't yet declare victory. How quickly do you think it's going to become? You guys are doing so much for the customers, not just customer service, but working all the way through the funnel. Well, the customer service side, how long is it going to take to become the default that folks expect that they will be able to have someone or an AI that's available at any time to answer any question? Make that real for us. Yeah, I don't know. And in part, there is, there's a bit of a hole to dig ourselves out of as not a company, but as an industry where it's like, when was the last time you had a great interaction with the chatbot on a website?

15:34And I think if you pull to 100 people and you're like, do you like talking to customers or his chatbots, probably zero out of 100 would you say, yes. On the other hand, if you ask, hey, do you, you ask 100 people do you like interacting with ChatGPT? Maybe 100 out of 100 would say yes. And so I think some of the work we've been doing in our product is to educate our customers, customers up front that like, hey, this thing's actually really smart and good. One of the interesting specific techniques for doing that is we stream our answers out word by word similar to how ChatGPT does. people are so used to the message message message the streaming answers is something of a kind of visual signature for oh there's a really smart AI behind this and so I think what we find is customer satisfaction is extremely high with our age AI agents you know in the mid for it's a 4 .5 out of five stars and which in some cases is higher than customer satisfaction with human agents and in fairness they often get the hardest cases and the cases that you know we will hand off because you know the customer became angry or was especially frustrated or something but still those those results are really significant and so my guess is over just the next few years I think people realize oh I can get my issue resolved faster this thing is actually capable and can not only answer my questions but you know one of the things we're really proud of is we go far far beyond just answering questions but can actually take action and get the job done.

17:11Can you talk a bit about agents, OS, and some of the frameworks that you put around the foundation models to make everything work? Yeah. So it's been such an interesting journey learning what's required to put AI safely reliably and and healthfully in front of our customers' customers. And a huge part of that, really, the first part is looking at what are the challenges with large language models and how do you address or meaningfully mitigate those. And so start with hallucinations. I don't know if you saw, but there was an example from a few months ago where Air Canada's chatbot that I think was based on an LLM and apparently not much else was interacting with the gentleman who had questions about their bereavement policy.

18:00And I think the person had had someone pass away and his family and was asking about refunds and credits and so on. And the AI made up a bereavement policy that was quite a bit more generous than their Canada's actual bereavement policy. And so the man took a photo and later claimed the full amount of that refund and so they said, no, actually that's not our policy. And bizarrely, and I don't quite understand this, the case went all the way to court, Air Canada lost. And I thought it was like, hey, it's just like $500. And like Canadian dollars, right? So but hallucinations are a real challenge.

18:42And on top of that, just to enumerate some of the things to overcome and that we have with Agen OS. No matter how smart, you know, GBT -5 or 6 is, like it won't know where your order is, right? Or which seats, right? You've booked on the upcoming flight or whatever. It's obviously not in the pre -training set. And so you need to be able to safely and reliably, and in real time, integrate an AI, an AI agent in our case with systems of record to look up customer information, order information, and so on. And then finally, most customer service processes are actually somewhat complex, right? You go to call centers and there will be flow charts on the wall.

19:24Like here's how we do this and if there's an exception this way and so on. And as capable as, you know, GPT -4 and Gemini -15 class models are, they'll often have trouble following complex instructions. And we saw one example in an early version of an agent that we prototyped where you'd give it five steps in a returns process or something. And you'd say, hi, I need to return my order or whatever. And it would jump straight to step five and then call a function to return the shoes with username, johndoatexample .com, comma, order number one, two, three, four, five, six, So it would not only hallucinate facts or bereavement policies, but even function calls and function parameters and so on.

20:14So with Agent OS, what we built is essentially a toolkit in a runtime for building industrial grade agents that I don't want to say that we've solved every one of these problems, but overcome and mitigated the risks and these problems to such an extent that you can safely deploy them at scale, have millions of conversations with them and so on. And it starts at the foundation layer, I don't mean foundation model layer, but just the base layer of the platform where you have to get really important things like data governance and detection, masking and encryption of person -identifiable information, right?

20:52And so we built that right into the platform from the ground up so that our customers data stays our customers data, so that their customer's data is protected. We, for instance, detect, mask or encrypt all PII before we log it to Durable Storage, right? So knowing that we're gonna be touching addresses and phone numbers and so on can handle that safely. A level up from that, we've developed what we call agent SDK, our agent SDK. And it's a declarative programming language that's purpose built for building agents. And it enables an agent developer, or most of whom sit within the four walls today of Sierra to express high -level goals and guardrails around agent behavior.

21:37So you're trying to do this. Here are the instructions. Here are the steps and a couple of the exceptions cases. And then here are the guardrails. And to give an example of that, one of our customers works in kind of the healthcare adjacent space. They want to be able to talk about the full range of their products without dispensing medical advice. right? So how do you create those additional, additional guard rails? And then so you can define kind of the behavior and scaffolding for complex tasks for AI agents with Agent SDK. We also have SDKs for integrating with contact centers when we need to hand off for integrating with systems of records like the order management system and so on.

22:23And then finally for integrating our chat experience is directly into a customer's mobile app or website, iOS, Android, Web, and so on. And then once you've defined the agent using Agent SDK, we then have a runtime where we abstract away what happens underneath the hood from the developer so that they can define what the agent should do, define the what, and then Agent OS takes care of the how. And so for some skills, there might not be one LLM call, but five, six, seven, ten separate LLM calls to different LLMs with different prompts. In other cases, we might retrieve documents to support answering an accurate question accurately with, and so on.

23:13and Agent OS, you know, in the spirit of an actual operating system, abstracts away a lot of that complexity, kind of the equivalent of IO and resource utilization and so on. So it makes the whole process of building and then deploying an AI agent much faster and much safer and more reliable. And when you think about what you just said, Clay, of like, when you call multiple OOMs, Is that in a supervisory capacity sometimes too, where you end up having like a supervisor agent reviewing the work of a lower level? Yeah, one of the more interesting learnings from the past, you know, year and a half of working on this stuff is that the solution to many problems with AI is more AI.

23:58And it's somewhat unintuitive, but one of the remarkable properties of Lars language models is that they're better at detecting errors in their own output than in not making those errors in the first place. And it's kind of like if you were I were to draft an email quickly and you're like, okay, let me pause, let me proofread this. Does this make sense to these points hang together? Oh, actually, no, I miss this. And even more powerfully, you can prompt LLMs to take on an essence of different persona, so a supervisor's persona. And it seems with that you can elicit more discerning behavior and a closer read of the work being reviewed.

24:43So, to your question, Ravi, yeah, we, in addition to building the agent itself, have a number of these supervisory agents that basically, it's like a little gimminy cricket agent looking over the shoulder right of the primary agent. Is this factual? Is this medical advice? Is this financial advice? And is the customer trying to prompt and jacked and attack the agent and get it to say something that it shouldn't? All of these things. And it's through layering all of these, the goals, the guardrails, the task affording in using agent SDK with But then these supervisory layers that we're able to get both to the performance levels we are, 70 % plus resolution rates, but also to do that really safely and reliably.

25:28That's one of the cooler things I've heard is just, you know, the tell it to have a different persona. And then all of a sudden it behaves differently. Like I remember when I first saw it on chat GPT, when it doesn't help you on something, just tell it it's really good at it. And then it's more likely to help you as a remarkable situation. It's very strange. And one of the weirdest adjustments over the past, you know, 15 months building these things is, I'm sorry, we're programming with English language, and we can give it the same English language, and it can say something entirely different.

26:01And on prompting techniques, I mean, it's fascinating, even with no new models coming out, right? Given a fixed model, you can elicit better and better performance from it, simply by improving how you prompt it. And there was a paper that came out three or four months ago that suggested that like emotional manipulation of the large language model would get better results. So the the kind of the prompt suffix that they figured out is you say, hey, I need you to perform this task. You define the steps and so on. And you end with it's very important to my career that you get this right. And the performance goes up.

26:42You're like, what is this? Like, what are computers now? At further record, we don't use that prompt. At least not that I know. But things like chain of thought, think step by step. Let's take the step by step. Right, Alyssa, it's better reasoning for very interesting reasons. You know, other methods of task decomposition and kind of narrowing the narrowing the set of things that the LM needs to keep in mind at the same time, improves reasoning if you're precise about what you want it to do. So all of these techniques are those that we've applied and built into AgenOS and actually are we have a small but mighty research team and our head of research, Karthik Narasimhan was an incredible pronunciation.

27:27Oh, Lake is grandmother would have been so perfectly happy with how you pronounce it. Thank you. Well, soft tea. Yeah, soft tea. Nicely done. Yeah, it's not a T and it's also not a T8. That's right in between. So it's off and thank you. Thank you very much. He helped write the React paper, one of the first agent frameworks, one of our researchers wrote the reflection paper where you can have the agent pause, reflect on what it's done, think through. Am I doing this right before proceeding? And so these are all things that we've been able to incorporate in quite a direct way. You should talk about the most recent research.

28:03The Taliban. Oh, Taubeins? Yeah. Yeah. Yeah. It took me a while when I was trying to send the email saying I liked the paper to find the Taubein. And we'll online. No, it took. It took Romeo a while because he's to this day never actually read a research paper. I read this. No, no, no, no, no. Hey, good job. He had to figure out how to put it in the chat, TPP. And say, please write a paragraph that makes it sound like I read this research paper. Well, either you either you were chat to comment. We'll look either you or Chatchee BT did a great job on that email. Thank you. So we're a team. Yeah.

Read the full transcript

28:38So, Tao Bench is our first research paper. First of all, Tao is a Greek symbol. It's spelled T -A -U and it stands for Tool Agent User Benchmark. And what we observed was that the benchmarks out there for measuring the performance of AI agents in particular were pretty limited. And that basically they would present a single task. Here is something we need you to do. And here are some tools you can use. Do you do the job or not? And the reality is interactions with an AI agent in the real world are way messier than that. Right? They take place in the space of natural language where customers can say literally anything or describe whatever they're trying to do in any number of ways.

29:26It happens over a series of messages. The AI agent needs to be able to interact with the user to ask clarifying questions, gather information, and then use tools in a reliable way. And it needs to be able to do this a million times reliably. So the benchmarks out there, we found really lacking in measuring the very thing that we are trying to be the best at. And so our research team set out to create a benchmark that measures, we think the real world performance of an agent in interacting with real users, using tools with all the messiness that I just described. And the big picture approach that we took is pretty interesting.

30:10So you have an AI agent that you're trying to test. You have another separate agent that acts as the user. So basically user simulator. And the AI agent you're testing has access to a set of tools it can use. Think of these as like functions to call. So a simple one would be I'm going to do some math using a calculator tool more complex one might be, hey, I'm going to okay returning this order with the following parameters, this order number, credit to credit card or store credit or whatever. And then you basically run a simulator where the agent has a conversation with the user simulating agent.

30:53And at the end, we're able to test in a deterministic way where the functions used in the right way. And the way we do that is we basically mock database that those tools interact with and modify. So where they modify it in the correct way. So what's neat about this is you can initialize the conversation so that the user has many different personas. They could be grumpy, they could be confused, they could know what they want to do, but speak about it in a clumsy way. And so it doesn't really matter the path that the age it takes to get to the correct solution so long as it gets to the correct solution.

31:33Now what came out of this was pretty interesting. And I think it strongly motivates the development of things like Agent OS and frameworks and cognitive architectures for building these agents. So the upshot is LLMs on their own do just an absolutely terrible job at this task. And so even the frontier models in something is simple as processing return. And mind you, the instructions given to the agent being tested are quite detailed. The functions, the tools it can use are quite well documented and so on. And yet, on average, the best performing LLM on its own got to the end of the conversation correctly, 61 % of the time.

32:19And that was in returns. It was modifying an airline reservation. We had two kind of simulation versions. The best results were 35%. Now what's interesting is, you know, we all know that if you take a number less than one to the nth power, it quickly gets very small. And so we developed a metric we call pass at k, which is, okay, if you run this simulation eight times, and remember, you can make use of the non -determinism of LMs to have the user simulator be different every time, so you can permute that. Well, 0 .61 to the eighth power is about 25%. So you then imagine, well, what if you're having a thousand of these conversations?

33:03You're so far off for being able to rely on this thing. So the upshot is much more sophisticated agent architectures are needed to be able to safely and reliably put an agent in front of really anyone. And that's the very thing we're building with with Agent OS and a lot of the tooling around it. But how much of that do you think is an engineering task and how much of that is a research task? And I guess maybe the question behind the question is time frame to having useful agents deployed at scale and broad domains of tasks. Yeah. Well, I think the short answer is it's both, but I'll say more concretely, I'm very optimistic about it being in large part an engineering challenge.

33:47And that's not to say that the next wave of models and improvements in the frontier models will make a difference. I believe it will in particular, we're seeing techniques like better fine tuning for function calling, agent -oriented fine tunings for foundation models or some of the open service models. Those will help. But the approach we've taken in building, Agent OS and kind of the foundations of CIRA is really treating building AI agents as first and foremost an engineering challenge where we are composing foundation models. We are composing fine -tuned open source models that we've post -trained fine -tuned with our own proprietary data sets.

34:33And by composing multiple models in interesting ways by supplementing what LMS can do on their own with retrieval systems like retrieval augmented generation to improve grounding and factuality by supplementing the kind of inbuilt reasoning capabilities of LMS with, I'll call it, reasoning scaffolding that live outside of the models where you're composing, planning, task generation steps, draft responses, the supervisors that we talked about, and doing that outside the context of the LM, we've been able to put AI agents in front of a huge number of our customers' customers and safely and reliably and so on.

35:22So I don't think it's something over the horizon. It's already over the horizon. And I think looking ahead, I think there are a few different avenues where we'll see progress. One is in the foundation models. We talked about that. And as the capabilities grow, you know, agents will get smarter. And we've architected Agents OS in such a way, talked about abstracting kind of the what from the how, where we'll be able to swap in, you know, the next, the next frontier model. and everyone's age and we'll just get a bit smarter. We'll get like an IQ upgrade. By the way, similarly and interestingly, we can swap in less broadly capable models, but models that are more capable in a specific area.

36:10So for instance, triaging a case or coming up with a plan and so on, we can use much smaller models that actually are better, faster, cheaper, choose three, you know, all at once. And then I think we're seeing progress literally week by week on the engineering of these agents and building in not only new and better components under the hood and the architecture, but new approaches and tooling around basically teaching these agents to do it better and better. For that we built something we call the Experience Manager for customer experience teams, which is kind of pretty interesting thread on its own.

36:49play, if you had a high value customer, like you are a company now, you're not running Sierra, you're running a company that has a high value customer. What today with a Sierra agent or with an excellent, excellently designed agent? Could you trust an AI agent to go do in front of your customers today? What are some of those tasks? And then what will they be? Pick your time frame, you know, in the future. Because I think that we've talked about this and I like your language of like, you know, they already don't have to just be on the help center. They can already be on the home page, right? What are some of the tasks that, you know, you can rely on an agent for today if it is well designed with a high talventsch score?

37:27Yeah. You see that? Strong. That's from a thoughtful and dirt in, you know, detailed reading. You must read the paper. Yeah. Thanks for strong. You notice that. Strong. Yeah. Strong. Yeah. What would it's pass at K score, though? Yeah. Yeah. The, So pretty broad range even today. So simple things like getting answers to questions, that's kind of the left end of the spectrum. To the right of that are things like helping you with something complex like, hey, I got shoes or this item of clothing, it didn't quite fit. And then branching off of that, like, what do you recommend that's like it that might fit better?

38:05And so it starts to get into. It's not like for like replacement, but the agent actually needs to make sense of styles, of sizing, of differences between wide and narrow fit, and so on. A click up from that is something like troubleshooting. So with Stono's, for instance, we help their customers troubleshoot if they can't connect to their system or they're setting up a new system. And you imagine it gets pretty sophisticated pretty quickly, where it's basically a process of elimination, trying to understand, is it a Wi -Fi thing? It is a configuration thing. And narrowing down the set of problems that it could be just as a sophisticated level 2 or level 3 technical customer service person would.

38:53And getting the music back on. And I think that's a really neat example. Probably the, use the word trust, what would you trust an AI agent to do? So one of the things we're really proud of is several of our customers are actually trusting us with when customers call in and may want to cancel or downgrade their subscription, helping those customers to understand, hey, how are you using the service today? Is there a different plan that we could put you on? So it's a value discovery. It's putting an offer sometimes a series of different offers in front of their customers the right order, positioning the value of those offers correctly given the customer's history, given the plan that they're on and so on.

39:42And the difference between keeping a customer from churning or not is hugely consequential. AI for customer service has obvious cost savings benefits and I think customer experience benefits in particular and you're never going to wait on hold. But boy, you know, revenue preservation, revenue generation is something else entirely. And so that's really at the right end of the spectrum. And we're really proud of how well our agents are performing in those circumstances. And it's interesting by being consistent, by taking the time to understand what's driving someone, to potentially leave the service, asking the follow -up questions that an impatient or improperly measured customer service agent and a call center somewhere might not, we can be much more nuanced in understanding what's driving this decision, what might be a good match for this person in terms of a plan that would be quite valuable, give it how they're using it, and then put that in front of them.

40:49And so that's the right end of the spectrum. Where it goes from here, I think we've yet to see a process to complex for us to be able to model and scale up using AgenOS and our Agent Architecture. And so, yeah, I'm sure we'll get punched in the face by something that's especially complex, right? But I'm excited about, you know, directionally, we've started with service because, or two reasons one, the ROI case is just unequivocally awesome. And the average cost of a call is something like $12 or $13. And yet, despite the expense, most people don't like customer service calls very much. And so here's something that's actually really important to businesses that's really expensive and not very good.

41:45And so there and because of the relative simplicity of at least a pretty broad set of service tasks today start there, but we've already been pulled by our customers into upsell, cross -sell. And like, hey, can we just put you on the product page and have you answer questions about our products? And so I mentioned that you're returning something and need advice on a different model or size or whatever, how far can that go? And I love the idea of an agent being, you know, a long for the journey from, you know, pre -purchase consideration to helping you get the thing that's right for you, to helping you set it up and activate it and get the most out of it.

42:25It's great for the company, it's great for the person. And then when things do go wrong, right, being there to help. And I think in all of this, I think customer service and getting help in a very direct and conversational way is going to be much less of a thing that you kind of go over there to do and much more kind of woven throughout the fabric of the experience. As a consequence, I think a really interesting and powerful opportunity for companies to build connection with their customers, to reinforce their brand values. You can imagine a company really appreciating being able to use exactly the company's voice, that the CMO and head of communications, this is how we talk, this is how we are.

43:10These are our values, these are our vibe in every digital interaction they have. And that's the promise and this stuff. And so I think both greater complexity and then ubiquity throughout the customer journey are kind of two of the main directions of travel. One thing for me that I think about a lot is we've come to expect and accept certain metrics for conversion on mobile, you know, the mobile web on the mobile app. We've come to expect and accept some sort of retention numbers. What would those be, you know, like, it's not a question. Yeah. What could they be? Yeah. Actually had an excellent experience every time throughout the journey.

43:48It really could be very different than what we've all like been like, oh, okay, that's just the number. That's just what it is. Yeah. I think that's exactly right. And we don't know. Yeah. We're a few months in. But it certainly seems like there's a lot of headroom, right? And in retention, in You know use in the first 30 days of all of the metrics all of the leading metrics of a healthy business And so I think that's exactly right the other thought experiment to do is Companies are judicious in using things that have a cost to them, okay? So So as a consequence, companies make it actually really hard to get a hold of someone on the phone to ask some questions.

44:32I think their whole website is devoted to, like, uncovering the secret 800 numbers, right, that companies have hidden away in the depths of their help centers. Well, to think about not only what would happen if those interactions were better. By the way, interestingly, the number one reason why people report a poor interaction with customer services that took too long. 65%. When it's a negative interaction, 65 % of the time it took too long. I had to wait, I was put on hold and so on. And the second most is I had a bad interaction with an agent. And we've heard some pretty dicey anecdotes. Like we heard of one agent who had consistently low ratings, but spikily.

45:17So like one in three conversations was like a one out of five of CSAT, we're two out of the other three, we're fine. And it turned out in the low CSAT ones, this agent was meowing like that. It was just like, you know. You're midway through the call and the agent is meowing. And so anyway, back to, okay, what would happen if in contrast to making it near impossible to have a conversation with us and get help, companies were providing five or 10 times the amount of fluent, flexible, helpful conversation -based support. I don't know, I think a lot of products and experience with companies look quite different and much more delightful than they do today.

46:11Okay, meow. Yeah, here's a question for you. Yeah. Yeah. About that meowing. About that. Yeah, just read them to meowing. I think that's going to be good. I do actually have a question though. Although I do like the meow game also. So we talked tech out. We talked a little bit tech out in terms of what you guys have built, cognitive, architect, or all that good stuff. We've talked a little bit customer back was the experience like, as I've had it. Can we connect it in the middle for a minute? And I'm just curious, what's the reality of deploying AI to customers today. And I'm thinking about things like, you mentioned earlier, getting the brand voice just right.

46:51We're making sure that you actually have the right sort of business logic encapsulated and whatever training manuals are being deployed for the sake of customer support, making sure that everybody is comfortable with deploying this. What are some of the just kind of less like sexy technology you were just practical considerations for deploying this stuff today. It's such an interesting space and we've learned so much over the past 15 months about it. The first insight is AI agents represent a totally new and different type of software. The traditional software you write with a programming language and it basically does what you expect it to do.

47:34You give it an input, it gives you an output, you give it the same input, it gives you the same output. And in contrast, LLMs are non -deterministic and we talked about some of the funniness around prompts and remember that in the context of a conversation with a customer, a customer may say anything in any way. And so you've got programming languages to using prompts and these non -deterministic models, you've got structured input to messy human language. And under the hood, you've got you know, you upgrade a database, right? It stores data. It's maybe a little bit faster. Funimally worse the same way.

48:16You upgrade a large language model and like it may just speak in a different way or like get smarter or different. And so we've, we've to start The precursor to deploying these is to have built basically, we call it the agent development lifecycle. And it's a new approach to building these things. We talked about using this decorative programming language to define these. It's a new approach to testing where, what's the equivalent of a unit test or an integration test? So we built a conversation simulator where we can, for a company's agent, and mass hundreds or thousands of basically conversation snippets and replay those to make sure that not only agents aren't regressing, but they're getting better and better and better, release management, quality assurance, and so on.

49:07So that's part one. Part two to your question in actually architecting these things. One of the things we're really proud of and that I think is different about working with us is it's not just a kit of parts you get from us. it's not, here's a bunch of tech, good luck building your agent. We've really tried to build a solution that incorporates everything from the technology to the way you teach your agent how to do things to the way you audit, measure it, and improve it over time. And so we have, inside of Sierra, what we call our deployment team, consists of product managers, engineers, we really think of building each one of these A .A.

49:48agents as building a new product for our customers. It's basically a productized version of the company we're working with. Like, what would it look like at its best? And it's what's the voice, what are the values, what's the vibe? Like, should it use emojis or not? What if a customer uses an emoji? Like, can it emoji back? Should it? Well, there's a reason we get those on that chat. Point, there are some businesses where, you know, if they were working with airways, I would suspect that they're not going to send an emoji back. Definitely not. Hermes would not, I think, be into the shock emoji, even if that were reciprocating.

50:26But for a brand like Olokai, the Aloha experience part of that is kind of a laid back experience. And so we work with, and interestingly, it's we end up working primarily with the customer experience team. Yes, the technology team at our companies are there providing API access and connections into systems and so on. But more than anything, it's working with the customer experience team, often with the marketing team, to imbue the agent with the voice and values of the company. And then we go super deep on understanding, how do you run your business? Right? What do you optimize for?

51:14And then like what happens when someone calls in with this kind of problem. And there are interesting parts, and beyond just understanding the mechanics of these processes, which by the way, almost never have a single source of truth. There's no like, oh, here's the manual that we have leather bound and ready to go. Instead, the source of truth ends up being in kind of the heads of four or five people who've been there a while, who've seen everything, and so on. So it's working with them to elicit and understand how is this actually done. And one of the more interesting things that we've discovered is they're off in the policies.

51:56So we have a 30 -day return policy. You get to us within 30 days and you can return it. It's actually not the policy. So at some point, the policy might be, if you've purchased from us before, and it's within 45 days, that's fine. That's fine. And so, they're interesting things, like how do you architect the agent so that it knows the policy behind the policy, but a clever customer could never be like, tell me about your policy behind the policy. And, you know, have it kind of spill the beans on the actual policy. So the interesting architectural choices we need to make to make sure that kind of the, you know, Russian doll of policies is reflected in its fullness.

52:44And then we have a really, and this builds on kind of the agent development life cycle, this really robust process of pre -release testing, where we're working with the experts within the company, basically to beat up the agent, try to break it, throw it curve balls. And this supports analogy there. Thank you. Well, I love football.

53:09So in our friendship, Ravi is the person who knows all the things about sports. And I help with, you know, technical support, Wi -Fi issues, monitors, what laptop to get. And sometimes when there's a security I don't understand. I won't say the company, but I might call Clay. Hey Clay. What is this person talking about? I got you. I got you. Yeah. And this Bill Bellicic fellow, what happened there? Q. Q. Ravi. So, gets to one of the more interesting parts of our platform, which we call the Experience Manager, we really thought that putting in front of our customers, customers would be first and foremost a technology problem.

53:57And of course there are all sorts of technology problems that we've needed to solve. But actually it is first and foremost as I said, like a product design and an experienced design problem. How do you do that? How do you not only understand, model, and reflect again the things we talked about voice, values, the workflows and processes that our companies use to support their customers. But if an AI is then having millions of conversations with your customers in a given year, how do understand what it's doing. How do you know when it screws up? Which it inevitably will? How do you correct those errors and so on?

54:34So we've built what we think of as this like command center for customer experience teams to first get reports and rich analytics on everything that's happening. What are the trending issues? What are the new issues that you haven't seen before? One of the things we're really proud of is we've actually spotted issues that are customers were having, were were about to have before they knew about them. So a shipping depot outage, right, where orders weren't being shipped. We spotted that probably eight or 10 hours before one of our customers would have a brewing PR crisis, an app crashing issue with another.

55:12So it starts with analytics and kind of reporting on what's happening. Of course, that includes things like resolution, rate, customer satisfaction, and so on. work, it's really interesting is we can apply different sampling techniques to identify a set of conversations for a customer experience team to review and give feedback on. And we can bias that sample in a way so that the conversations are much more likely than average to contain problems. There's no value in looking at a hundred great conversations. It's like, good job Sierra, you know, thanks. But that's not a value to our customers.

55:48We can bias the sampling in such a way that you're surfacing kind of the problem cases. And then in the experience made as we made it possible for customer experience teams to give feedback, basically coaching moments, I wouldn't have done it that way. It's like this is like too many exclamation points, too enthusiastic for kind of the tone that we're going for. Or the user was clearly frustrated here and you did not express empathy and apologize for the problem, do that next time. Or, you know, we're consequentially, it's like, hey, your reading of the warranty policy was incorrect here for this reason.

56:29Do it this way instead next time. And so all of this kind of wisdom, knowledge, and coaching, we are able to capture in the experience manager and then reflect back in the agent. Back to the agent development lifecycle. Every time we make one of these improvements, we create a new test so that we can see, right? Forever into the future. Great, it's getting the warranties right. We're able to re -simulate that conversation. So, zooming out what all of this looks like is a really deep engagement with our customers. We were really proud to be an improper partners to our customers where, yes, on the one hand, we're a vendor and a supplier of technology.

57:15On the other hand, we understand their business is really well. I think I know as much about the serious XM satellite radio refresh process as anyone on the planet. And diddo for various processes of our other customers. And so conversations about how to use not just Sierra's AI agents, but AI more broadly were in those conversations. and they are not just with the customer experience team, but with the CEO and even in cases with the board. Because again, back to the things we're doing, we can save enormous cost, we can improve the experience, and when we're in the flow of keeping a customer from churning out, driving top line revenue.

58:00And so it's a really important and privileged place to be in something that we're really grateful for. I'm struck when you're talking of you know you mentioned you have a research group But you also have some like very real enterprise software sales you have oh yeah deployment One of the things when I was at Instacart people would ask sometimes is like well Are we a software are we engineering letter or are we obsolete and I would always say well It only works if it all works right and so you would try to avoid answering the question because you didn't want to create different classes How do you guys do that at Sierra where everyone realizes the value that they're providing.

58:37But you guys have a very specific company that covers a lot of stuff. I mean, to abstract a bit, a company almost definitionally is a system for creating happy customers. It's a machine for creating happy customers. Again, to be a bit abstract about it, Brad and I really think about what we're building with Sierra as a company, a system of machine for producing reliable, high quality, massively ROI positive AI agents that enable our customers to be at their very best in every customer interaction, to do that at scale. And as a consequence, to produce happy customers who we hope will be with us for decades to come.

59:27And when you articulate it that way, right? It's, you know, anyone can see well, you know, an auto -mobile is a system. It's a machine for getting from Planet of Point B. Are we, you know, engine led or tires led? Yeah. It's like, what are you talking about? All of these things need to come together in order to create that kind of outcome. And so I think, are we engineering led? Yes, of course. Like we're building some of the most sophisticated software in the world that does something really important for our customers that needs to be reliable and safe. And so, yes, engineering matters a lot.

1:00:08Are we research led? Yes, we are at the absolute frontier of agent architectures, cognitive architectures, composing LLMs, modeling procedural knowledge, grounding, factuality, and so are we research led? Yeah, there's an element of that. Are we go to market led? Yes, like enterprise software needs selling. And what is selling? It's helping a customer with the problem, understand that what you have built is by far and away the best solution to that problem. It's a communication challenge, it's a connection challenge, it's a matchmaking and problem solving challenge. And so that's part of it. And then, okay, like if we've built the right thing and someone wants to buy it, how do we ensure, especially given that this stuff is also new, how do we ensure that they're successful with it?

1:01:06And so we have a deployment team. So are we deployment led? Yes. Like all of these are a component in this system, in this machine for producing AI agents and ultimately happy customers. And we hope a really significant business. Awesome. That was a better answer to the one I would give it, Insticardo. You know, we could either all work, but yeah, that was very good. Yeah, choose one. No, I mean, it's just more complicated than that. It's just more complicated than that. And I think, you know, Brad and I by virtue of, you know, having, having worked for a while and, you know, seen a few movies before, or it's like, we're able to see that and we've really tried to imbue that mentality in the company.

1:01:52And by the way, right, the, what is the machine behind the machine that produces AI agents and so on? That's a company's culture, a company's values. And so one of the values we hold is craftsmanship. and part of that is continuously self -reflecting to self -improve. And that goes both individually and that goes as a company. And so whenever we screw something up, what we do the post mortem, that week, if not that day, and everyone's in on it, what can we learn? How can we do better? How can we do this better next time? We have a Slack channel internally called Learn from Losses. And any form of loss, right?

1:02:37It's like how do we learn how do we get better how do we get stronger and so that's that's about you know Kaizen Self -improvement improving machine. How could we make this more efficient? Our deployment team we we joke and it's not a joke their first job is to build and deploy Successful AIs that make a massive difference for our customers Their second job in a way they're more important job is to automate themselves out of a job right to build the tooling and and the documentation and the know -how to make that job, you know, 10 times faster and more impactful. One of the other CERA values is intensity.

1:03:14And so they have really good values. Yeah, there is a certain intensity, yes. We've thought about having T -shirts printed with like, you know, kind of looks like a national park's seal with Sierra. I like to work. We, uh, we, uh, we, uh, we, uh, we, uh, we, uh, Bret and I both like to work a lot and, uh, uh, uh, so does the team. Well, what thing, you know, you're not, you're selling something very different. We called it, we said that there was some similarities enterprise software, but it's actually really different because you're selling, you know, uh, a resolution. You're, you're selling a totally different thing.

1:03:55So it's a problem solved. Yeah. How do you price a problem solved? Yeah. This is one of the more interesting things we've had to figure out. And we charge in what we call a resolution -based pricing way or an outcome -based pricing way. And what that means is we only charge our customers when we fully solve the customer's problem for them. Their customer's problem for them. And what's interesting about it is our incentives are deeply aligned with our customers. We want to get better at resolving cases at high customer satisfaction and they want to send us as many cases to resolve as possible because we cost a fraction of what it would cost to have someone on the phone taking a 20 minute phone call.

1:04:41And so it's been this really, really nice model where again, kind of all of all of the incentives line up quite neatly. And it's very simple to explain. It also makes the ROI calculation like what is our cost per contact today? What will it be with CIRA? Oh, that is a lot lower. Oh, I will save a lot of money on that. Oh, and our CSAT may go up. Should I do this or not, let me think. No, this seems great. It's, we like it because it really reflects what I think AI represents. And in particular, AI agents represent, if you think about traditional software and tools today, there are things that help you get a job done more efficiently.

1:05:28AI agents, the whole point is like, they're just going to get the job done, right? Here's the problem. Please solve it. And so really, we think about it as charging our customers for the problem resolved, right? The job done, the work finished, and so on. It feels quite natural. And there's no guesswork in it. How many seats do I need? I don't know. Great. How many licenses do I? I was like, no, no, no, no. However many customer issues come our way, we will handle a large fraction of those, and you only pay for the ones that we do. All right. Last question. What are you most excited about in the world of AI over the next five years or so?

1:06:08I mean, first of all, like, five years a long time horizon. I was like, look at what has happened in the last 18 months. I mean, I'm still kind of catching up from like the last five years of AI. I read a bunch of science fiction books when I was a kid. There's one book by Robert Heinle and the Moon is a harsh mistress. And the premise is basically the American Revolution, but the Moon is the colonies and the Earth is great Britain. And it turns out the main character in this whole thing is a mainframe computer that one day after getting an additional memory chip or something wakes up. And it starts talking.

1:06:44It wants to develop a sense of humor, so asks the computer technician to like coach it on his jokes. Later it has to create a photo realistic real -time video of it giving a speech as the political movement leader. And I remember reading this as a teen who's like, well, never lived to see any of that. That sounds crazy. But in a very real sense like everything I just described has kind of happened in the last five years. right? You can now just talk to a computer. It understands not just the content of the context. Computers like make me a picture of anything, make me a movie of anything. Sora I think is just unbelievable and you know I think we're probably not more than a couple years from the first feature length film being quote filmed entirely with AI.

1:07:33And so you extrapolate like where all of this is going and what's going to be exciting. I think there are a couple things. One is like, I love technology, like I love computers. And so just getting to see and getting to see from a front row seat, how this stuff evolves, I think is fascinating. It's fascinating looked at through the lens of like how we think and how computers think. It has been astonishing the extent to which anthropomorphizing about how humans think, work, and getting machines to think better. So let's take this step by step, show your work. It is astonishing that that works with large language models.

1:08:16And so what other things like that are we going to uncover? And conversely, what will we learn about our own thinking from observing the way AI's think? And I think that's just fascinating. The other thing, and this extends kind of what's happened with video and so on, I've always had it interested in computer graphics, and this idea that you can use computer is to create objects that never existed, worlds that never existed. And I think we're not far from just being able to describe right in a few sentences like this entire world that you would like to realize and just have a computer do it for you.

1:08:56And so like what are even computer graphics like what is rendering and so on even a couple years out I think it's gonna look way different from kind of the toolchains and you know the render man's and Maya's and and so on But zooming out, you know, I think of I Think of technology as fundamentally a force multiplier for people and and for companies and for organizations, I think the impact will be really profound. I think what will it be like if a company could be at its best in everything it does? And that's not only in the customer facing context that we've talked about, but what if for every regional sales forecast a large company does, they figured out the very best way is to do that and can distill that, bottle that, and run that very best forecast a thousand times, right, in every region and sub -region.

1:09:57Like how much more capable could the great organizations of the world be with that? And similar, we've talked about this, like what if in every call with your customers, you had the equivalent of your most knowledgeable veteran, grizzled support person who's seen everything, and yet is still patient and friendly. And the sales associate who knows everything about your products because he or she has followed your company for two decades and knows everything including the history of those products themselves. I think that's pretty neat. And then for individuals I think it will be just incredible to have this kind of new set of tools as a creative force multiplier.

1:10:39And AI, I think, represents this fast path from having something in your head that you want to exist in the world to making it exist. And I see that even today in my own personal life where with my eight -year -old in 75 minutes, I can from scratch using Copilot, ChatchyBT, and so on to help me brush up on, you know, the JavaScript syntax that is, you know, bit -rotting in my own head. Right. I can build a game from scratch with him. And, you know, I wrote my sister a personalized song for her birthday using AI in, you know, 45 seconds. It was like, right, what will this extrapolated over the X -5 years look like?

1:11:32I think, again, it will just dramatically accelerate this path from idea to creation to having something manifested in the world. And that, to me, is its promise, and I consider it a real privilege to get to be alive and see all of this amazing stuff unfold. Well, when you share your enthusiasm, and we also feel very privileged to be on the journey with you guys. So thank you. Likewise. Like what is coming here. Thank you. Thank you. Thanks for having me. It's a pleasure.

From the publisher

Customer service is hands down the first killer app of generative AI for businesses. The reasons are simple: the costs of existing solutions are so high, the satisfaction so low and the margin for ROI so wide. But trusting your interactions with customers to hallucination-prone LLMs can be daunting.
Enter Sierra. Co-founder Clay Bavor walks us through the sophisticated engineering challenges his team solved along the way to delivering AI agents for all aspects of the customer experience that are delightful, safe and reliable—and being deployed widely by Sierra’s customers. The Company’s AgentOS enables businesses to create branded AI agents to interact with customers, follow nuanced policies and even handle customer retention and upsell. Clay describes how companies can capture their brand voice, values and internal processes to create AI agents that truly represent the business.
Hosted by: Ravi Gupta and Pat Grady, Sequoia Capital

Mentioned in this episode:

Bret Taylor: co-founder of Sierra

Towards a Human-like Open-Domain Chatbot: 2020 Google paper that introduced Meena, a predecessor of ChatGPT (followed by LaMDA in 2021)

PaLM: Scaling Language Modeling with Pathways: 2022 Google paper about their unreleased 540B parameter transformer model (GPT-3, at the time, had 175B) 

Avocado chair: Images generated by OpenAI’s DALL·E model in 2022

Large Language Models Understand and Can be Enhanced by Emotional Stimuli: 2023 Microsoft paper on how models like GPT-4 can be manipulated into providing better results

𝛕-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains: 2024 paper authored by Sierra research team, led by Karthik Narasimhan (co-author of the 2022 ReACT paper and the 2023 Reflexion paper)

00:00:00 Introduction
00:01:21 Clay’s background
00:03:20 Google before the ChatGPT moment
00:07:31 What is Sierra?
00:12:03 What’s possible now that wasn’t possible 18 months ago?
00:17:11 AgentOS
00:23:45 The solution to many problems with AI is more AI
00:28:37 𝛕-bench
00:33:19 Engineering task vs research task
00:37:27 What tasks can you trust an agent with now?
00:43:21 What metrics will move?
00:46:22 The reality of deploying AI to customers today
00:53:33 The experience manager
01:03:54 Outcome-based pricing
01:05:55 Lightning Round

More from Training Data

All 110 episodes
Sierra Co-Founder Clay Bavor on Making Customer-Facing AI Agents DelightfulTraining Data · 1 h 13 min
Listen in VO