In short
Sierra’s approach to building “agentic commerce” customer-experience agents that are simpler than typical agent harnesses: low-latency voice architecture, parallelized reasoning/listening/speaking, multi-model orchestration, and PCI-compliant payments separated from LLM infrastructure.
Guest
Zack Reneau-Wedeen, head of product at Sierra. Background: product leader focused on agent building for enterprise customer experience; discusses Sierra’s platform used by major brands and its voice, payments, and governance capabilities.
Key claims
- Agentic commerce will surpass e-commerce; Sierra agents can drive sales with commission via outcome-based pricing.
- Sierra conversations differ from a standard LLM call: 10–15 models may run per turn, sometimes classifying and responding simultaneously.
- Best performance comes from parallelism and context engineering (“when you think the model’s being dumb, it’s probably you”).
- Payments must be isolated: Sierra is PCI DSS Level 1 certified; payment info is not sent to external LLMs because providers aren’t PCI-certified.
Notable examples
- Voice transcription ensembling: two models run in parallel; if one flags “silent,” its output is trusted.
- Redfin AI search: a Sierra agent returns listings; the same agent is available as a ChatGPT app via MCP-style integrations.
- Explorer: “ChatGPT deep research” over customer conversations/data; can proactively ask questions and feed fixes to Ghostwriter.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOSierra's Unique Agent Model
0:45 to 2:52
Discussion on how Sierra builds agents for different customer interactions.
“You know, 10 or 15 different models might be invoked for a given conversation turn.”
Platform Architecture and Use Cases
2:52 to 4:48
Exploration of how Sierra's platform is structured and its extensibility.
“more to all of our existing and future customers.”
Building and Analyzing Agents
4:48 to 6:44
Overview of the process of building, analyzing, and releasing agents.
“processes and change management that have just pushed us to develop from the very beginning, you know, very buttoned up procedures and collaboration and review and all of this stuff.”
No-Code Agent Development Experience
6:44 to 10:34
Insights into the no-code experience for building agents on Sierra.
“Who is doing that analyzing and that iterative improvement?”
Agent SDK and Ghostwriter Functionality
10:34 to 14:00
Discussion on the role of the Agent SDK and the Ghostwriter in agent creation.
“So you come in and you just say, hey, you know, I want to orchestrate order returns or I want to do flight booking or I want to do car rental or referral from primary care provider to a specialist.”
Evolution of Agent SDK
14:00 to 15:00
Learn how the Agent SDK has evolved to improve user experience.
“Like you're saying agent SDK is somewhere in the middle and that might actually confuse it.”
Integrating Code and No-Code Solutions
15:00 to 17:30
Discover the balance between code and no-code solutions in agent development.
“So the core agent SDK is part of the Sierra platform, but building agents in code is totally something you can do.”
Agent Communication and Latency
17:30 to 19:00
Understand how agents communicate and the importance of low latency.
“And so I think it's kind of equal parts model improvements and platform improvements.”
Agentic Commerce vs. E-Commerce
19:00 to 21:40
Explore the potential of agentic commerce and its future implications.
“And because so many of our customers have their own technology teams and a robust array of different AI projects internally, they might be experts in a particular area.”
Payments and Security in Agent Platforms
21:40 to 23:00
Learn about the security standards for payments in agent platforms.
“I know that we were certified by a QSA, which is a qualified security assessor.”
Show all 40 chapters
The Future of Consumer Interaction
23:00 to 25:00
Discuss how consumers may interact with agents for various services.
“I do think that the majority of this will be from personal agents.”
Parallel Processing in AI Agents
25:00 to 28:00
Understand how parallel processing boosts efficiency in AI agents.
“commercial, that all still feels relevant to me.”
Modular Architecture for Voice Models
28:00 to 29:28
Learn about the advantages of a modular architecture in voice AI.
“We've learned so many things from being a modular architecture on voice.”
In-House Model Development
29:28 to 30:54
Understand when and why to build AI models in-house versus using existing solutions.
“You mentioned you have some in-house models as well.”
Agent Data Platform Insights
30:54 to 33:16
Discover how the agent data platform enhances customer interactions using AI.
“You mentioned something like every run of the agent would have 10 to 15 different model calls.”
Explorer Agent Functionality
33:16 to 34:36
Explore the features of the Explorer agent in enhancing customer service.
“is it can either integrate with your customer data platform, all your internal systems, or you can have it just on Sierra or you can do sort of a zero copy integration.”
Future of Agent Architectures
34:36 to 37:08
Discuss the potential evolution of agent architectures in the coming years.
“And so what it allows you to do is basically instead of having to go spelunking for the specific insight in reports or in monitors, you can just ask the question.”
Model Transitioning Strategies
37:08 to 39:24
Learn how to effectively transition between different AI models.
“They'll say, oh, well, you'll just ask, you know, GPT-12 to make that model, so it's actually not a big opportunity.”
Context Engineering in AI
39:24 to 42:00
Understand the principles of context engineering and its importance in AI.
“Do you also change out some of the tools themselves?”
Understanding Prompt Caching and Its Impacts
42:00 to 43:15
Learn about the role of prompt caching in AI performance and its trade-offs.
“basically like, yeah, prompt caching is great.”
Exploring Reinforcement Learning in AI
43:15 to 44:28
Discover the benefits and challenges of using reinforcement learning in AI systems.
“There's two topics that we were discussing earlier, and I'm curious.”
Capacity and Performance During High Demand
44:28 to 45:50
Understand the importance of capacity planning for high-demand scenarios like Black Friday.
“But if we're not pushing the state of the art, we really want to be thinking about what's going to be true three months from now, six months from now.”
Evaluating Multi-Agent Systems
45:50 to 47:49
Learn when multi-agent systems are beneficial and when they can hinder performance.
“using whoever has the capacity to serve us.”
Voice Technology in AI Agents
47:49 to 49:27
Explore the challenges and innovations in implementing voice technology in AI.
“Is there a right time to build a multi-agent system?”
The Future of Voice-to-Voice Models
49:27 to 54:28
Discuss the current state and future potential of voice-to-voice AI models.
“and they have a ton of volume over voice, even more than they have over chat.”
Predictions on Voice Technology Adoption
54:28 to 56:00
Get insights into when voice-to-voice models might dominate AI interactions.
“and I would say early on the APIs missed the mark on the ergonomics.”
Guesstimating AI Adoption Timelines
56:00 to 56:40
Explore predictions about the future adoption of voice-to-voice AI models.
“which I think, not to get too philosophical, but that's kind of the direction of the industry in general.”
The Importance of Memory in AI
56:40 to 59:10
Learn how memory plays a critical role in AI agent interactions.
“I vividly remember the first time staying up till 3 a.m., just trying to jailbreak the prompt.”
Implementing and Managing Memory
59:10 to 1:04:00
Understand how to save and extract memories within AI systems.
“So if you call, for example, my wife lived in Hawaii for a year.”
Evaluating AI Agents: Best Practices
1:04:00 to 1:06:00
Discover evaluation processes and tools for AI agents and simulations.
“at least or identification at least from you.”
Continuous Learning and Improvement in AI
1:06:00 to 1:10:06
Discuss continual learning in AI and how it ties into memory and evaluation.
“and collaboration and review and making sure that, you know, you have workspaces so you can let Ghostwriter run free, but still review it before you make any changes.”
Benchmarking AI Performance
1:10:06 to 1:12:04
Learn about the benchmarks for AI agents and their development process.
“I don't think anyone knows more than we do.”
User Experience in AI
1:12:04 to 1:14:13
Explore the importance of user experience in developing AI agents.
“So it's just too customer specific for us to rely on something as general as a benchmark.”
Outcome-Based Pricing Models
1:14:13 to 1:16:25
Understand how outcome-based pricing aligns incentives in AI services.
“We've seen, and I think one of the reasons vertical companies have been pretty successful lately is that understanding the contours of each industry and each company really makes a difference.”
Differentiating Pricing Strategies
1:16:25 to 1:18:38
Examine how different customer needs affect pricing in AI services.
“I think that's particularly interesting.”
The People Behind Sierra
1:18:38 to 1:20:44
Discover the traits and skills that help individuals thrive at Sierra.
“So we make sure that incentives are deeply aligned.”
The Evolving Role of Product Management
1:20:44 to 1:23:57
Learn how product management is changing with the rise of coding agents.
“I love being like, oh, I could imagine, you know, my friends using this or my parents using this.”
The Ideal Agent Builder Profile
1:24:08 to 1:25:11
Learn about the key traits and skills that define successful AI agents.
“who ends up fitting this agent builder profile the best?”
Interviewing for Agency
1:25:11 to 1:26:02
Discover methods for evaluating agency in potential candidates.
“My own personal rubric, which is like very much in beta, is kind of this customer intuition, agency, product judgment, technical depth, communication, intensity.”
Insights from the AI Native Interview
1:26:02 to 1:27:07
Understand how building a product in an interview reveals candidate agency.
“So the most concrete way that we've changed our interviewing process is we have this AI native interview.”
Transcript
Automatic transcript. May contain errors.0:00Sierra:Agentic commerce will be bigger than e-commerce.
0:03Zack Reneau-Wedeen:There are cases where the Sierra agent is actually getting paid a commission on a sale. Today, I'm talking to Zack Reneau-Wedeen, head of product at Sierra, the platform-powering customer experience agents for most of the Fortune 20.
0:16Sierra:Coding agents are really good at file systems. They're really good at Git. They're really good at grep. Let's materialize everything into those structures so that coding agents can just, you know, cook.
0:26Zack Reneau-Wedeen:He breaks down how Sierra builds for voice and why the architecture looks nothing like a standard agent harness.
0:32Sierra:One of the big unlocks for Sierra agents was how to parallelize thinking, listening, and talking. And if this model says it's silent, you trust it. If this model does not say it's silent, you trust this one.
0:42Zack Reneau-Wedeen:We get into why a Sierra conversation is unlike a typical LLM call.
0:46Sierra:You know, 10 or 15 different models might be invoked for a given conversation turn. So sometimes you're classifying and you're responding at the same time. And Zach explains how and why Sierra built an entirely separate infrastructure layer for payments. We have isolated infrastructure where payment info doesn't go to an external large language model because none of the LLM providers are PCI certified in that way.
1:08Zack Reneau-Wedeen:Welcome to Max Agency, the podcast that goes deep into how the best agents are being built by builders like you. Most people are probably familiar with Sierra as a customer support platform. But from what I understand, recently you guys are going broader than that. Could you talk a little bit about the types of agents that you help people build?
1:28Sierra:Yeah, so this has been the vision since the beginning. I think because many companies have an RFP process where they're very specific about, hey, we want to solve customer service. That is often where we start. But we think of Sierra as the full engagement platform across all of the moments that matter for your customers. So if you're an airline, that might be browsing for a flight, might be booking the flight, might be choosing your seat. Might be, in my case, I have a small dog, so adding a pet in cabin, then the flight might get rescheduled or delayed or canceled, et cetera, et cetera. You need to get your bags there.
2:05Sierra:There are just so many different things across that process. Some of them are sort of sales. Some of them are more service. Some of them are more loyalty. But they all kind of ladder up to the relationship between a business and its customers. And Sierra agents are present at all of these different parts of the customer lifecycle. So as an example, there are cases where the Sierra agent is actually, because of our outcome-based pricing model, getting paid a commission on a sale, which I think is quite different from how most people would imagine service. And so we get really excited about those more exotic opportunities because they also give us an opportunity to push the platform forward and, you know, continue adding to what the platform can do and kind of turn that into a product, package it up nicely and bring more to all of our existing and future customers.
2:55Zack Reneau-Wedeen:How similar is the platform between these different use cases? I would say that it's very extensible.
3:02Sierra:And so you can kind of take it in different directions. we like to say i think it's originally attributed to the um one of the creators of the programming language pearl but we try to make the easy things easy and the hard things possible so out of the box pretty similar the starting point but you can kind of take it in any direction that you want so it's not that oh there's like a separate product for you know one company versus another but the agents that you can build on it and we'd like to think of these agents as products unto themselves can be arbitrarily customized. What does it look like to build on the Sierra platform?
3:40Sierra:So we have, it's basically a web app. There's three main sections. There's analyze, build, and then there's release. Within the analyze section, you have things like our Explorer agent, which is kind of the long running ChatGPT deep research for all of your customer conversations and data. You have reports, you have monitors, which are kind of always on evaluators of conversation data as well. And then within build, you have ghostwriter, which is the agent similar to codex or cloud code for building agents. You also have journeys, kind of the underlying source code layer, although it's not really code, it's more like natural language or standard operating procedures, as well as kind of different variables and everything like that.
4:27Sierra:On the release side, you have all of the collaboration and change management and governance procedures. And so Sierra, I think at this point, we're working with most of the Fortune 20, something like 40 or 50 percent of the Fortune 50 or Fortune 100. So very much with a lot of the largest companies in the world. And they have needs around governance and release processes and change management that have just pushed us to develop from the very beginning, you know, very buttoned up procedures and collaboration and review and all of this stuff. That's basically what it's like, I think, on the surface, probably similar to a lot of other, you know, places that you go to build things, whether that's Figma or Cloud Code or these different places, but just very much optimized around no-code agent building and giving you
5:19Zack Reneau-Wedeen:all those capabilities. And are those different steps intended to be done in that order, like analyze, build, release? Can you analyze basically human transcripts before you build the AI agent? Or does analyze really come after you build and release the first version of the agent and now you're iterating on it? It's both.
5:35Sierra:So typically, you'll come in with some sort of resource of how you want the agent to be structured and architected, how you want it to behave. Maybe that's transcripts. Maybe that's standard operating procedure. Maybe that's a conversation that you have with Ghostwriter. And that will typically be how you build the agent. So I'd say most people will start with build. But then once your agent is live and production conversations are happening, your daily routine probably starts more with analysis. You're probably thinking, how can I optimize the metric that I care about, whether that's customer satisfaction or resolution rate or, in the case of the customer I mentioned, like sales converted.
6:18Sierra:And so you get those insights and then you want to make improvements to the agent, whether it's, you know, fixing an issue or finding a new opportunity to hill climb on a metric or please customers in one more way. And so that typically involves, you know, working with Ghostwriter. Often Ghostwriter will actually proactively suggest an improvement on the insights to kind of close that loop and build that flywheel. But I would say the day-to-day is more analyze, build, release.
6:47Zack Reneau-Wedeen:Who is doing that analyzing and that iterative improvement? Is this engineers? Is this product folks?
6:54Sierra:It's primarily people that have the most depth and insight about the ideal customer experience, which tends to be operations people. So customer experience managers, folks in that department at our customer companies. it's also a number of engineering teams will build either other agents that interface with Sierra agent or they can extend the platform via basically tools and packages that you can then kind of see and introspect on Sierra so it's very much kind of the same way that you have the person that knows everything about your knowledge base we want them to be able to come in and self-serve on day one and just, you know, make the perfect instantiation of knowledge.
7:40Sierra:The person who knows everything about the standard operating procedures should be able to just do that in the product. And so we're constantly kind of trying to sand down all of the barriers between the people with the most context and their ability to contribute directly to the platform.
7:54Zack Reneau-Wedeen:You've said no code a few times. So what does this agent building experience look like? Is it truly no code? And I'm assuming it's maybe something like quad code where you talk to it and it generates something under the hood. Is it generating code? Is it generating a custom DSL? Yeah, good question.
8:11Sierra:So the layers of the stack, you have kind of what we call agent OS, which has our constellation of models. So translating the tasks that need to be done on the platform into prompts, into a data injection across 10 or 15 different models that might be invoked for a given conversation turn. Some of those might be frontier models that need to do top tier reasoning. Some of them might be in-house models that are very good at a specific task. And some of them might be just classifier models that run really well on a model that's a little bit cheaper and more performant. And so that's kind of the base layer.
8:50Sierra:On top of that, you have the agent SDK, which is the code-based layer of agent orchestration and context management. That's kind of where Sierra started. But over the last 18 months, most of the agent development, pretty much all of the agent development, has shifted to our no-code layer that we call journeys. It compiles down to agent SDK code deterministically and isomorphically, which is a fancy word for you can turn it one way and then turn it back and it's the same. And so you can have code that you transition over to no code. You can have no code that you transition over to code. But the language of specifying it is very much declarative.
9:35Sierra:Here's how I want the agent behavior to be. when customers ask about this, we want to unlock these conditions and kind of flow in this direction. And we find that that's pretty intuitive because it maps the type of document that we would write for someone joining the team in a customer experience role or a sales role. You would explain to them how to do the job. And that's kind of what you're doing here as well.
9:56Zack Reneau-Wedeen:But there is some DSL for journeys. It's not pure raw kind of like text. Correct.
10:03Sierra:And it's very hard. I'm not sure we could get into a discussion about it. If you're just doing text, you have to choose between this is non-deterministically compiled, which all of the experiments we've done in that direction, you end up, I think, with more harm than good. Or this is a prompt engineering task, which then puts you in the realm of engineering teams. And we are very proud to be more in the realm of operations teams, where a lot of that domain-specific knowledge resides. The other big piece of it is that Ghostwriter has totally changed the learning curve for building agents. So you come in and you just say, hey, you know, I want to orchestrate order returns or I want to do flight booking or I want to do car rental or referral from primary care provider to a specialist.
10:50Sierra:And Ghostwriter just kind of already knows those concepts and is an expert in journeys. But Ghostwriter is using the journeys product. So it's not writing code. It's writing journeys directly so that you can go inspect that after the fact as well.
11:04Zack Reneau-Wedeen:I imagine there's some format that these journeys have to adhere to. And I imagine that's not in the model's training data at all. Was it hard to teach it that format or was it pretty easy?
11:15Sierra:It's a really good question because at every point there's this conflict between here are the perfect abstractions for me and here are the abstractions that the models are most familiar with. And similar to in math how you're often taking one problem and reframing it in another problem to do a proof or something like that, you have to decide if you want to reframe this problem into something the models understand or build a skill and inject context in the right way so that the models can understand your way of thinking. The truth is that we do both. So there are cases where we'll say, coding agents are really good at file systems.
11:50Sierra:They're really good at Git. They're really good at grep. Let's materialize everything into those structures so that coding agents can just cook. Then there are other cases where it's like, no, no, no, our way of thinking about this is the correct way of thinking about this. And there's not really a way to shoehorn it into what models are already good at. So let's do the investment to make the models good at this. My personal perspective is that 80 % of the time you want to do the first thing and just meet the models of where they are on their turf. And you should reserve the second one for that really special case.
12:27Sierra:I'm curious if that's been your experience as well.
12:30Zack Reneau-Wedeen:I think recently it's probably gotten to be, we see a lot of people using the file system as an abstraction. And so I think recently there's been a lot of talk, especially as the labs talk about how they're RLing the models to be really good for their harness to try to fit everything into a file system or this particular like edit file tool or things like that. I also think that the models are really good at writing certain packages. Like if you're in the training data, I think a lot of lane graph is in the training data. So I think at least anthropic models recommend lane graph for a lot of use cases.
12:59Zack Reneau-Wedeen:And that's great. But for newer things like deep agents, which is a new package we have, it's not in the training data at all. We spend a little bit of time, maybe not as much as we should, but we spend a little bit of time thinking about what makes these models good at writing certain things. We really have no clue how to know how to affect what goes in the training data, but it's a really interesting thing. And so I think there's definitely been cases where we see that people choose technology because the models are really good at writing it. And so one question I was also going to ask for the agent SDK, I imagine that's your own custom kind of framework built in house.
13:30Zack Reneau-Wedeen:I don't know if you experimented with having it sounds like you didn't you're not having ghostwriter write that directly. But that's obviously much more closer to code. And so I was curious if you experimented with ghostwriter or any model like writing agent SDK versus just writing like raw code. So, yes.
13:47Sierra:One of the things, too, is if you almost do the abstraction that the models are really good at, it can be overconfident or it can be familiar and successful. And so you have to be very thoughtful about going either all the way there or just not going there at all.
14:02Zack Reneau-Wedeen:Like you're saying agent SDK is somewhere in the middle and that might actually confuse it.
14:05Sierra:Exactly. Exactly. We have kind of reinvented the agent SDK two or three times as models improve. So it used to be you had to have more deterministic guardrails in order to get the behavior that you want. now there's more room for reasoning at each individual step and you can kind of push out the frontier of that reliability versus reasoning trade-off. So that's been very interesting. The reason for Ghostwriter primarily or entirely editing no code is just that that's where the vast majority of activity is on the platform today. So that's what our customers know. And so making Ghostwriter good at it is really where all the payoff is.
14:46Sierra:I think if we tried to do it For code, it would be a similarly scoped task, but it would be hard to get it to be really good at both at the same time. There's always going to be some trade-off.
14:56Zack Reneau-Wedeen:Do you still let users edit the agent SDK code if they want, or is that now completely abstracted away from them in terms of just its journeys? So the core agent SDK is part of the Sierra platform, but building agents in code is totally something you can do.
15:12Sierra:One example is a number of our customers have CICD, continuous integration pipelines, that they want to make sure their agent is released on. And so they need a Git repository, which is where their agent lives. Another example is sometimes you have a particularly complex tool that interacts with a streaming API or something in a way that is just easier to model in code than in no code. And so the way that these work is because no code compiles down to code, you can kind of import or under the hood it will import code files and compiled no code files kind of all as though they're the same thing because they are so i think this is a benefit of starting out as a code-based platform is that we still support it we have a number of customers that have dozens or in a few cases 100 plus developers building on the platform and sometimes for that you know they work in git they release in git and so being part of their enterprise change management protocol just means supporting Git.
Read the full transcript
16:10Zack Reneau-Wedeen:You mentioned that the agent SDK has changed over the past few years, as everything in the space has. What does it look like now and how has it changed? What does that evolution look like?
16:20Sierra:So it started out, we call this now flow-based, very much like, you know, do this. And it wasn't just like, do these things. I think I'm a big fan of your not another workflow builder blog post. So it wasn't that rigid, but it would be, hey, you know, make sure you collect their email before you say that you're going to send them a confirmation email, right? Very clear to us, but you would want to do those things in that order. Now, I think if you think about just the way that agents can reason through tool calling, instead of having to specify that in the actual structure of your agent, you might just give it the context that in order to call this tool, you know, as a prerequisite, you should have their email and it will know how to ask for it.
17:05Zack Reneau-Wedeen:So it's really like you'll just say that in the prompt.
17:07Sierra:Yes. Or, you know, eventually all things end up in prompts, but it would be in the journey. And so as you're kind of writing it out, you would specify that that's one of the rules or policies of the journey. And then the agent can take care of the rest. I think it's a mix of the models getting better and our orchestration platform becoming even more sophisticated and robust. So when I talk about that constellation of models, there's a lot of, we talked a little bit before about Linnaeus and Darwin, there's post-training that goes into that, there's model selection and eval and prompt engineering as well.
17:43Sierra:And so I think it's kind of equal parts model improvements and platform improvements.
17:47Zack Reneau-Wedeen:I want to talk about this model in a second, but I want to stay on the harness for a little bit. How similar does it look like in its current form to a coding agent harness? Does it have access to skills and subagents in the same way that someone using CloudCode would have?
18:00Sierra:So the one constraint we have that Cloud Code doesn't have is latency. Majority of Sierra conversations are voice. And if you're not responding in one or two seconds, then people wonder where you went. And so we are highly optimized for these low latency use cases. There's a ton of parallelism. That being said, at a high level, it is using a lot of the same models. It has access to tools. So there's a lot of similarities. There are also, you know, you can invoke other agents from the Sierra agent. So the core, what's best for the core conversation loop isn't typically what's best for software development, but you might want to say, hey, let me actually give you a call back in 20 minutes after I figure this out.
18:46Sierra:And then you would have a type of loop that runs, you know, more like Cloud Code.
18:49Zack Reneau-Wedeen:And for those longer loops that might happen in the background, are those also built on the Sierra platform and are just a separate type of agent that remove the latency constraints? You can do it either way.
18:59Sierra:So you could have a Sierra agent calling out to another Sierra agent, or you could also have a Sierra agent calling out to an in-house platform. And because so many of our customers have their own technology teams and a robust array of different AI projects internally, they might be experts in a particular area. like they might have document generation handled themselves and that might be a long running agent and then a Sierra agent can call out to it and wait for a response. So it's kind of up to you to choose and we find that enterprises are varied enough that they appreciate kind of having choice.
19:35Zack Reneau-Wedeen:When you do these agent to agent communications are you using A to A or one of the protocols specifically for that or MCP or just an in-house I don't know REST API call?
19:46Sierra:The most common is an API call. When you know who you're talking to in advance, oftentimes you can save a lot of tokens and make sure that you're 100 % accurate that way. That being said, CR agents support the MCP and agent-to-agent protocols. You can kind of install that integration and then your agent can be an MCP client. You can also set up your agent to be an MCP server. So this is how we support ChatGPT apps, which rely on MCP servers. Basically, the tools of the agent can be made available to ChatGPT, and then you can at-reference a Sierra agent. The example that would be most familiar is Redfin.
20:25Sierra:If you go to redfin.com and do their AI search, under the hood, it is a Sierra agent that is returning the home listings and having the conversation with you. And that agent is also, I believe, available in ChatGPT. Interesting.
20:41Zack Reneau-Wedeen:I didn't realize that CR agents could be ChatGPT apps. Is that the right terminology for them?
20:46Sierra:It is. Yeah, exactly.
20:48Zack Reneau-Wedeen:Do you have an opinion or hot take on in the future, do you think people will be interacting with the agents that represent brands on dedicated chatbot websites or in ChatGPT or central chat engine?
21:03Sierra:I think that agentic commerce will be bigger than e-commerce. So if I think about how I get things done today, it used to be that I went to websites and clicked around. Now I ask Codex or Claude to do things for me. And I don't see why I won't do that to manage my subscriptions, to order supplies to my home, to make dinner reservations. It just feels like that's where we're headed. And so if that's happening, I think brands will want to be ready on the other side of that. So we are very much planning for that world. we were investing in payments before it made sense i think and it's a long process but a few months ago we announced you know we're uh fully pci dss level one certified i have no clue what that means what does that mean payment card industry okay um oh man you stumped me on dss uh i have uh some of the other acronyms um in my head but we'll put it in post okay okay thanks Oh, man, it must be like digital.
22:07Sierra:I don't know. I know that we were certified by a QSA, which is a qualified security assessor. And what that means is we're able to do the only voice payments platform. Certainly at launch, I think still this is the case where you don't need to transfer to another platform. So it's a cohesive experience throughout checkout. out. And all the work that went into that, we have isolated infrastructure where the payment info doesn't go to a large language model, doesn't go to an external large language model because we have none of the LLM providers are PCI certified in that way. And so putting that all together is like spinning up a separate cluster, you know, getting certified, making sure all of our operational rituals, you know, conform to what the security assessor is looking for.
22:52Sierra:and we put in that work because we believe in this future where agentic commerce is actually bigger than e-commerce and i think e-commerce is in the hundreds of billions of dollars at this point a couple percentage points of gdp or something like that just in the u.s and so if
23:08Zack Reneau-Wedeen:you think about that space um it's pretty big and by agentic commerce do you mean like chat gpt talking to a sierra agent that represents redfin or do you mean someone going to redfin's agent no matter where it is and talking with it there? Both. Both.
23:25Sierra:I do think that the majority of this will be from personal agents. Just looking at user behavior, we spend so much time in Cloud and ChatGPT and Codex that you have to think
23:39Zack Reneau-Wedeen:that's where a lot of that behavior will accrue. And do you think those agents will interact with another agent? Why not just the raw APIs themselves? I do.
23:49Sierra:I think that as you think about being ready for that world, the same way that you might want to use Shopify or you might want to use certain software on your website to do product recommendations, to do checkout, you might want to use Stripe. Similarly, you'll want to use a platform that can make sure that you're presenting your products in the right way, that you're making checkout as easy as possible, and you're showing up at your best, whether it's for a customer that's browsing or for an agent that's browsing. The one thing that I think is pretty different is the attention isn't necessarily valuable in the same way.
24:30Sierra:Like our eyeballs are more valuable than an agent just spewing out tokens, assuming that no one's ever going to look at it. At least that's true for now. At some point, it ends up in some future training run and maybe has value. But I think that's de minimis relative to getting us to look at things. And so I do feel like maybe that's a bit different, but the presenting yourself in the right way, making it easy to check out, making it easy to understand what products are available, to express the preferences of whoever is responsible for that agent going off and doing something
25:02Zack Reneau-Wedeen:commercial, that all still feels relevant to me. I've seen some dev tools provider, I think Sentry doing a similar thing where they have a bunch of APIs, obviously for the underlying platform, but they also have an endpoint to just ask questions of the agent directly. And I think you could make a counter argument that like, great, brands should absolutely care about how the platform is being used and how it's being presented. But you could do that with skills or some other mechanism to expose that to the agent. And I honestly don't know which one's right, but it has been interesting to see the whole space is so new, but increasingly so companies exposing agents as endpoints to interact with rather than the endpoints themselves.
25:36Sierra:I agree. I think that all of this stuff, you could try to do it yourself. It might be that certain companies, that's the best option. What we've seen is that because there's often tens or hundreds of millions of dollars on the line, in some cases, billions of dollars on the line, you really want to make sure you're getting the best solution. And so if you're going to be 90 % as good at it as you could be partnering with a company like Sierra, it still makes sense to partner and you know get that extra few billion dollars one last question on this fun side tangent payments how early are we i think we're really early i personally still don't order paper towels with codex i don't know if you do no and that's why i asked i'm glad that
26:22Zack Reneau-Wedeen:you said that because i don't i'm not close to doing that and so i was wondering how far behind I was.
26:25Sierra:I mean, like, I also didn't do it with Alexa. You know, I think for some of that really easy stuff, you could probably have done it already. The one that I think will definitely become a thing is there are a lot of apps that claim they can, you know, go through all of your subscriptions and cancel the ones that you're not using. That feels like as a consumer, that's a useful service. I'm definitely closer to doing that with Codex than I am with an app. You know, know, it would be so much work to tell it all the things and try to tell it which ones to cancel. It's a very manual process. If I gave Codex or Cloud Cowork or something just access to my browser and said, hey, you know, go to all of the streaming apps and like the ones that I'm not logging into, just, you know, cancel those and let me know if you need my password.
27:15Sierra:And obviously you have to figure out how to make that secure and everything. But I feel like that I would have demand for that product.
27:21Zack Reneau-Wedeen:Going back to the harness for a little bit, you said something earlier about things running in parallel. Is that like guardrails that you're running or retrieval steps or what's running in parallel in this process? So many things.
27:33Sierra:One example, knowledge. We will often look up answers before we know if we want them. So you'll, you know, before you decide whether this question needs an answer, you'll at least have the answer ready or in parallel with deciding. So sometimes you're classifying and you're responding at the same time, basically speculative execution. Another example would be transcription, what we call ensembling. I think we might have published a blog post on this today, which is great. Go read it. We've learned so many things from being a modular architecture on voice. This was an early decision we made that I think has totally played out to our advantage, where we have the ability for any language, for any customer, for any use case, to multi-home providers across transcription, across synthesis, and across native voice-to-voice models.
28:30Sierra:And so on the transcription side, for example, it just turns out when you have a thick UK accent from northern UK, or at least parts of northern UK, I don't know exactly the region, there is one model that has the highest quality transcription, but it hallucinates during silence more than other models. So we run two models in parallel. And if this model says it's silent, you trust it. If this model does not say it's silent, you trust this one. And so that's just an example where we're running those in parallel. We have logic for when you take the right one, and it's very specific. And if you had all your chips in with one provider or one system, or you weren't doing things in parallel, you would inevitably hit the limits of what that provider can do.
29:17Sierra:So the same way we use Cloud and Gemini and the GPT class models, we're also able to use all of the leading players on the transcription and synthesis and speech-to-speech side as well.
29:28Zack Reneau-Wedeen:You mentioned you have some in-house models as well. What do those models do and why did you guys decide to build those in-house?
29:35Sierra:So knowledge is a great example. I think whenever we are pushing the limits of what's possible, we always consider whether we should build this in-house. whenever it's limiting our ability to deliver more for our customers. So an example where we're probably not the company is these, you know, many millions of dollar training runs that produce the GPT-55 class models. And, you know, Mythos, I'm sure, is a many, many millions or tens or hundreds of millions of dollars training run to get that produced. And that's stuff that OpenAI and Anthropic are just the best in the world at. I think what we're the best in the world at is going really deep with customers, understanding all of the process knowledge specific to their industry, specific to their company, specific to their customer base, and then having the products that can allow them to serve those customers as best as possible.
30:29Sierra:And so an example like knowledge where we were hitting the limits of the retrieval and re-ranking that we could do with out-of-the-box models, we asked the question of, you know, should we create our own models here and eval them? And we have a research team that's pretty sizable and tightly integrated with our product teams. And so we can flex that muscle when we need to, but we try not to be doing it just for the sake of doing it.
30:54Zack Reneau-Wedeen:You mentioned something like every run of the agent would have 10 to 15 different model calls. If you had to guesstimate, like how many of those are frontier model calls versus like in-house fine-tuned versus not frontier model, but third party?
31:06Sierra:So for a typical turn of a conversation, I would guess that, and this is just ballpark, but, you know, being precise rather than being accurate. I think maybe a couple frontier model, a handful of classifiers that probably don't require that, a handful of speculative execution in the case of voice in particular, to make sure that it's low latency. Sometimes there will be an interim response that's generated to, you know, the same way you'd say, hold on a minute, I'm just pulling up your account, like that kind of thing. Roughly like a third, a third, a third or a quarter, a quarter, a quarter.
31:43Sierra:But I would say the frontier models, just because they might be slower or more expensive, would probably be, you know, more doing the bulk of the reasoning, but in one or two inferences for a given conversation turn.
31:57Zack Reneau-Wedeen:Do you ever end up training models specific to a customer?
32:00Sierra:It's not something that would be out of the question, but I can't think of a specific example. The reason I pause is because we do have our agent data platform and there are machine learning models that power strategies that are specific to customers. But in terms of like a, you know, language model or generative model, we don't have cases of that.
32:23Zack Reneau-Wedeen:What's the agent data platform and what does it power?
32:26Sierra:Basically, one thing we realized pretty early on is large language models are really good at in-the-moment empathy, oftentimes better than we are, of understanding, okay, you know, I understand you're having a hard time. I'm really sorry about that. And it's the same way when you walk into a restaurant that has amazing service or a hotel that has amazing service. They recognize the moment you walk in, okay, this person just got off a really long flight. or this person's 10 minutes late to their reservation and they were stuck in traffic and I'm just gonna let them know that that is not a problem, their table is ready.
32:59Sierra:And large language models have that, especially on a platform like Sierra, but they don't necessarily know what you care about at a level deeper than that. And oftentimes the previous generation of AI or recommender systems have a better understanding of some of those things. And so what agent data platform does is it can either integrate with your customer data platform, all your internal systems, or you can have it just on Sierra or you can do sort of a zero copy integration. And it can take that structured data that knows what to recommend along with the unstructured data of the here and now in the conversation and then use those to generate better conversations, better orchestrations around how you want customers to feel and what you want to do for them.
33:47Sierra:So that all sounds maybe a little bit abstract. One example would be during sales. Oftentimes, there's structured data that knows the right offer to present, but doing it just with that structured data with the previous generation of AI and before Sierra feels very stilted or it feels like, you know, I don't know why you're doing this. And so large language models can really understand how to present an offer, how to attribute it, weigh two different offers based on conversation context and pick the right one for the moment and that kind of thing. So we see it a lot with sales, with loyalty and retention, those types of conversations.
34:23Zack Reneau-Wedeen:One of the last agents we haven't talked about too much is Explorer. Yes. What is Explorer? What does it look like under the hood?
34:29Sierra:I've described it as ChatGPT deep research for all of your customer context and conversations and all of the data on Sierra. And so what it allows you to do is basically instead of having to go spelunking for the specific insight in reports or in monitors, you can just ask the question. You can say, hey, I noticed my resolution rate dipped. You know, why was that? Or how can I generate more sales? Or I wish that more people were converting from trial to full-time paid plan. How come that's not happening? And then more than that, you can set up automations so that on a daily basis, for example, Explorer can ask the same questions proactively and then partner with Ghostwriter.
35:14Sierra:We currently think of these as kind of two separate agents, the analysis agent and the authoring agent to say, oh, here are some fixes that are suggested to improve. and you can chat with Ghostwriter and kind of pick it up from there. Under the hood where this is converging, I think, is a shared harness that is an expert at using Agent Studio, Sierra's platform. And so that's kind of what we've been setting up in terms of what we talked about at the beginning, like, you know, figuring out the file system architecture that maps to the product. And so as we've exposed more and more tools, you know, like building knowledge bases to these agents, they get more and more powerful.
35:53Sierra:and we see a lot of emergent behavior between both.
35:56Zack Reneau-Wedeen:Does this harness end up looking more similar to like a coding agent harness than the harness that's part of AgentOS or Agent SDK?
36:03Sierra:Yes, so this is less of a quick-turn conversational agent and more of a longer-turn deep analysis agent. And so it ends up looking a lot more like a cloud code or a codex.
36:16Zack Reneau-Wedeen:One of the things I've been thinking about, I'm curious if you have a take here. In a year, two years, three years, Will there be this split just in terms of harness ones that's optimized for kind of like, yeah, lower latency, external facing customer experience type things. Voices may be heavily involved and another that's really focused on these deep research, maybe coding, like you run in a sandbox, things like that. Or will it just end up converging into one harness that, you know, depending on how you prompt it or, you know, has these async sub agents in the background that can maybe run for longer periods of time?
36:47Sierra:I think there will always be latency, performance, cost trade-offs, and different architectures that emerge because of that. I've actually been surprised by how many different types of model companies there still are. and when I talk to people who are particularly AGI-pilled about it and I say, hey, what's like a cool model opportunity that's not flying like so close to the sun that the labs will do it? They'll say, oh, well, you'll just ask, you know, GPT-12 to make that model, so it's actually not a big opportunity. But I think in reality, at least up until now, I'll probably look dumb when AGI comes out.
37:26Sierra:you do see a lot of success in areas like voice models from transcription to synthesis. You don't always see leadership from the model labs. You see the model labs actually trying to focus more on specific problems for Anthropic. I think it's coding for OpenAI. It's been consumer now maybe shifting a little bit more to enterprise for Google, definitely consumer as well. and it really does feel like there are still trade-offs to me. So I expect there will still be multiple architectures up until that event horizon of AGIs here, so all bets are off.
38:00Zack Reneau-Wedeen:As you guys build your core agent harnesses, and I'm assuming want to build them in a model agnostic way, what do you need to change to go from an open AI model to an anthropic model?
38:10Sierra:Usually you have the evals that are designed to work across both. So if you have really good evals and a really good harness or really good architecture, then you should be able to kind of hill climb toward eval performance without too much effort. Often what will happen is you'll learn the first time you're switching one task from one to another or making it possible to run on multiple systems that your eval wasn't quite as good as you thought
38:33Zack Reneau-Wedeen:and so then you make your eval better and you continue to improve.
38:36Sierra:But the short answer is that it's pretty simple for a given intelligence level of a model to run a task on one or the other. And so again, it's that basically latency, quality, cost trade-off, but not more than that. And because we have customers that have very specific requirements around what clouds they can run in, what models they can use, and our company approach is to meet them on their terms, you don't serve most of the Fortune 20 without that approach. It's not really a choice. And so because that's our approach and we've built a lot of products around that, we also have made sure that we can kind of move between models of comparable intelligence without too much heartburn.
39:21Zack Reneau-Wedeen:And what do you end up changing when you're hill climbing? Is it just the prompts? Do you also change out some of the tools themselves? It depends on the case.
39:28Sierra:And I might not be the expert on the exact history of each. I think that if you change the tools, it's pretty hard not to have downstream effects of that. And there might be certain tasks that can only run on certain models. and other tasks that can run on other models. And so there's always kind of a set of eligible models for specific tasks. I don't know exactly how tools change, but I know the eval getting more robust and the prompt conforming to the quirks of each model is definitely part of the development.
40:00Zack Reneau-Wedeen:You guys recently wrote a blog around context engineering, and I think you said it was the key to great agent building or something like that. How do you guys think about context engineering and what tips or tricks would you have for others?
40:13Sierra:I think it's showing agents everything they need to do the right thing, but nothing more. And as models get smarter, you can be a little bit less precise with everything they need and certainly less precise with nothing more. So early on, the agent SDK was really about only giving the model exactly what it needed and kind of spoon feeding the context. Now to extend the meal analogy, it's probably more like, you know, putting out the right dish. And maybe in the future, it might be something that is even less structured. One concept that I think is in that blog post is kind of progressive disclosure.
40:51Sierra:You'll probably know more about this than I do. But when you bring something into the prompt, you don't want to do it before it's relevant. And then you also risk incoherence if you then yank it out of the prompt. So when you do things like prompt compaction, you just want to be really thoughtful about not making it lossy. Because if you keep something in the history that is incoherent with the rest of the system prompt, it's not going to end well. And so I think to the degree, you know, when we're fixing issues or when we've seen hallucinations, it's often because one part of the prompt was this and the other part was this.
41:30Sierra:And actually, one of my main learnings from Sierra is anytime you think the model's being dumb, it's probably you.
41:38Zack Reneau-Wedeen:I like that. I think that I think a lot of people have learned similar lessons. Yeah.
41:43Sierra:Whenever you think the model's too dumb, the model's actually too smart.
41:46Zack Reneau-Wedeen:How much do you guys care about prompt caching and maintaining that cache? I've heard I've heard kind of like two mindsets on it. One is like, yeah, do everything you can to maintain the cache, like don't invalidate it until you like absolutely need to. And then I've heard another theory that's basically like, yeah, prompt caching is great. But like what matters most is like performance. And sometimes you need to just like break the cache in order to insert the right context or give it a system reminder or something like that. How how strictly do you guys try to adhere to prompt caching?
42:14Sierra:Is the purpose for those who are prompt caching loyalists for speed or cost or quality?
42:23Zack Reneau-Wedeen:I think the first two mostly speed and cost. Speed and cost. Yeah.
42:26Sierra:I haven't I haven't heard anyone argue that it's for quality, but maybe maybe it works better. It's a nice to have. we definitely don't want to invalidate a cash for no good reason, but quality comes first. So we aren't zealots about it at all. I would also say that when the outcomes that your agents are delivering are very valuable, you have the luxury of not being extremely focused on cost in particular. And so that probably is part of the reason for that is that, you know, So a conversation with a customer could sell a$100 product or a$1 ,000 lifetime value plan. And so those are valuable enough that quality almost always comes first.
43:14Zack Reneau-Wedeen:We've talked a bunch about the agent itself. There's two topics that we were discussing earlier, and I'm curious. We've talked a little bit about them, but I'm curious if you have any more thoughts. First being RL. When is RL good? When is it bad? How much have you guys explored it?
43:26Sierra:We've explored it a lot, in part because it has two great promises, you know, increasing the ceiling of the quality of models and then also making it so that you can do a similar task on more models. I'm curious for your take, but in practice, I've seen a little bit more of the second one when it comes to like enterprise RL. It's taking an open model or open weights model and saying, how can we get similar performance to a frontier model? The two things that make it hard are, number one, the way that that gets delivered is non-deterministic and might include, you can't include any data that you don't want the model to regurgitate.
44:04Sierra:So we basically would never fine-tune a model on something when it could lead to regurgitation risk. That's just a non-starter. And then also, in general, just the way that you would train the model, you have to think about preparing all of that data. The other big one is that the frontier models are improving so fast that you want to remain as agile as possible. And so in many cases, doing something like RL makes a ton of sense for something like knowledge, where we feel like we are pushing the state of the art. But if we're not pushing the state of the art, we really want to be thinking about what's going to be true three months from now, six months from now.
44:43Sierra:And oftentimes RL is a rounding error against that.
44:46Zack Reneau-Wedeen:Yeah, I feel like to your point earlier, we've started to hear it a little bit more recently, I think because of cost. So I think like most people are interested in it when they're using these frontier models and the performance is good, but now whether it's coding or other things, their cost is just going through the roof. I think the, and we're starting to investigate this more, but I think the places we're hearing it most are basically in those where performance is good, cost too high, how can I bring it down? Let's see if I can train a model to get similar costs and a fraction of the cost or similar performance fraction of the cost.
45:16Sierra:Yeah. Interestingly enough, a lot of our progress here has been driven by capacity, not cost, where we have a lot of customers that are in the retail space. And when we go into Black Friday, Cyber Monday, for example, you need a lot of capacity to deal with the spikes that they face. We've also done load tests that are on the order of, if you were to have that rate of conversation over a year, it would be billions of conversations. And so that level of concurrency and those spikes just mean that we need to be resilient to downtime with a particular provider and ready for, you know, using whoever has the capacity to serve us.
45:59Sierra:And so it's funny because it's useful in so many ways, but a lot of the reason why we have such good support for multiple providers is specifically preparing for Black Friday, Cyber Monday, and running load tests for really large customers.
46:13Zack Reneau-Wedeen:One other harness agent engineering topic, multi-agent systems. Where do you think they're useful? Where are they not useful?
46:20Sierra:I think they are often not as useful as people think.
46:25Zack Reneau-Wedeen:My thoughts on this would be people should be really thoughtful about why they want a multi-agent system.
46:31Sierra:If you want a multi-agent system so that one team can work on one agent and one team can work on another agent, then you're shipping your org chart. If you want a multi-agent system because it just makes you more comfortable to think about this problem over here and this other problem over here, then you're also not optimizing around impact. If, for example, you had an agent that does triage and another agent that does a task, by building it as a multi-agent system, you're often depriving of the agent doing the task of the information from the triage and depriving the agent doing the triage of all the procedural information from the task.
47:09Sierra:And that's typically destructive of value. And so we are often just want to make sure that we're doing multi-agent systems for the right reason. If you're kicking off even a sub-agent, you want to make sure that it has everything it needs to do that task and that there's no reason why it shouldn't just be part of the main agent. And so I think I've seen a lot of cases where people are reaching for multi-agent systems the same way you might reach for microservices before you're necessarily ready for that level of optimization and also for reasons that might not be just about building the best possible agent.
47:48Sierra:And so Sierra agents tend to be kind of one agent representing the brand. You certainly can have multiple agents and build a multi-agent system, but if you're managing context correctly, if you're doing really, really good context engineering, then typically it's just not a problem because you're not exposing the wrong context to the wrong agent. Is there a right time to build a multi-agent system? I think if you have truly separable jobs, right, where there's not any purpose of the first context being part of the second context, I will say that in my personal opinion with, you know, May 18th, 2026, there are not a lot of great times for it.
48:30Sierra:There might be times where it actually, the organizational difficulties are worth the quality drop. But if you're doing it specifically for quality, I think it's pretty rare that you can't just solve it with better context engineering. And I'm kind of a monolith loyalist on that.
48:45Zack Reneau-Wedeen:I feel like voice is one of the things that is getting more and more popular, but there still aren't a ton of people doing a lot of, but you guys are. Can you give me a voice 101 or 201? What should I and other agent builders know about voice compared to just building, you know, simple chat agents?
49:02Sierra:Voice has been maybe the most fun project that I've worked on in my whole career. So, and I, for context, I joined Sierra as an agent PM working on building agents, specifically with customers in a forward deployed role. And one of the first customers I worked on is SiriusXM, the in-car streaming radio service. And so I'm a big SiriusXM fan before that and as a result. and they have a ton of volume over voice, even more than they have over chat. And so many of their touch points with customers are over the phone. And so early on, it was very obvious that voice was going to be impactful for the business.
49:41Sierra:And we got to think from first principles, basically from the ground up, what makes a voice experience great? How is that similar to chat? How is that different from chat? And so latency is probably the most obvious one. you need to be really thoughtful about parallelism you need to be really thoughtful about what we call progress indicators which is where you say you know hang on a second while i look up your account that's number one number two is naturalism this is a combination of a number of different things so oftentimes when something sounds a little bit robotic i'll i'll read what the agent said and i'm like wow i sound robotic too so it's a combination of what the agent is reading and then and also the quality of the voice itself.
50:25Sierra:There's multilingualism. It's very easy to speak different languages over chat using large language models. It's a lot harder to be fluent in, I think it's about 60 languages on Sierra platform than it is on chat. And each of those languages, sometimes the very best transcription provider might have a 20 % word error rate. I think that's true for a language like Hungarian, for example. And so it's like, how can we ensemble multiple transcription providers in order to get that
50:57Zack Reneau-Wedeen:down and kind of be better than a single model is on its own?
51:00Sierra:The other big factor is I think we all believe that a few years from now, most voice agents will be running voice native models. So, you know, real time, I think they might be up to 2.5 at this point. They've had like three big real time launches at OpenAI this year already. There was the really cool demo from Thinking Machines Labs as well. So there's been a lot of increased momentum here. And as of a few months ago, we now have production agents live with the voice-to-voice models. And so it's fully end-to-end doing that. You still need the transcript in order to make API calls and that sort of thing.
51:38But the agent is responding with audio as the input.
51:43Sierra:The other big piece of it, I think the Thinking Machines demo was a really good example. up until now we basically had like 50 lines of python i think silero is the most popular voice activity detection library deciding when to speak and then a trillion parameters deciding what to say and that balance feels very off to me if you think about the conversation we're having right now i'm actually using a lot of my brain power to decide when to speak in addition to decide what deciding what to say. And it's probably more like 50-50. And so one of the big unlocks for Sierra agents was deciding to think about not only how to parallelize a task, but how to parallelize thinking, listening, and talking.
52:29Sierra:So that when I'm listening, I'm already thinking about what I might say next. When I'm talking, I'm listening for interruptions. And so that was a big unlock in terms of the product design. The other one I would say is just modularity. Like I said earlier, no one is the best at everything in this space. And when there are, you know, 100 plus languages worldwide that, you know, really deliver meaningful results, when many of our customers are global brands, global companies, you need that flexibility to use one provider here and another provider there and to ensemble them together in a specific place as well.
53:03Zack Reneau-Wedeen:How much of that modularity and that parallelism and thinking about different things goes away when it's like a native voice-to-voice model?
53:13Sierra:In one specific conversation, it goes away. But if you think about the businesses we serve, the voice-to-voice models today are just reaching a level of reliability where you would trust them for English. And so if you still want to support all the different languages, you need that modularity for the foreseeable future. The other thing is they're still almost an order of magnitude more expensive. They aren't quite as good at reasoning yet. And so the cases where they are live in production, they're not quite as reliable with tool calling and instruction following. The cases where they're live in production, it's cases where we know in advance that the journey is a little bit simpler and where the naturalism matters even more than usual.
53:58Sierra:And the procedure is not as complex as some other cases. And so it's still, I would say, a fraction of our market that we can use voice-to-voice models for.
54:09Zack Reneau-Wedeen:My perception also, and I've never built a voice agent, so I know truly nothing here, but my perception here is for the voice-to-voice models, you probably have less control over what goes on inside of the loop, basically, of tool calling and reasoning. Is that correct, or are there pretty good controls for what happens inside?
54:27Sierra:You may not have built a voice model, but you're an expert in developer ergonomics. and I would say early on the APIs missed the mark on the ergonomics. And so they got the integration points wrong. And it was exactly what you said. It was, hey, if you want our model, you need our voice activity detection and you need the whole thing. There was still an underlying model that was available. So I'm dating myself in AI, but the GPT-40 audio model was extremely exciting. It did things that no model before it could do. I think maybe like people that are real AI OGs would say this about like GPT-2 or something.
55:06Sierra:And so you could see that this was coming. And I think we all would have said five years from now, this is where we're going to be. But the way that we wired that up in our system was basically using the entire Sierra pipeline and then holding on to the input audio and piping that in with all of the prompt context into the audio model to do the last mile. So we were basically still doing everything ourselves and using it for the last mile. I think you're right that over time, there's more and more that you can do with the audio model, the same way there's more that you can do with the text models.
55:43Sierra:The fallacy would be that, okay, so then you don't need the harness or you don't need all of the orchestration and simulations and everything because you can make that choice. You can either do the same thing a little bit more easily, or you can set your sights on new and more impressive things, which I think, not to get too philosophical,
56:02Zack Reneau-Wedeen:but that's kind of the direction of the industry in general.
56:05Sierra:It's like, are we all obsolete? Or are we going to find new things to do that raise our horizons even farther?
56:11Zack Reneau-Wedeen:If you had to guesstimate a time, we're big into guesstimating on the podcast, apparently. When do you think more than 50 % of your either traffic or customers will be served by a voice-to-voice model as opposed to this speech-to-text, text-to-speech pipeline?
56:28Sierra:I will be surprised if it happens in the next 18 months. I've been surprised before. I was surprised by Opus 4.5 late last year. I was certainly surprised by ChatGPT. I vividly remember the first time staying up till 3 a.m., just trying to jailbreak the prompt. And so I know you're a sports fan too, So if we're doing over-unders, it would be like 24 months in one day or something like that. You know, like over-under 24 months would be probably my personal guess, Demation.
57:03Zack Reneau-Wedeen:How, if at all, do you guys think about memory, specifically long-term memory? It sounds like you've got users potentially interacting with multiple different agents that a single brand can be building. How do you think about the memory that's shared across them?
57:19Sierra:Memory is very important to the platform. So I mentioned the agent data platform earlier, which kind of brings together machine learning data or, you know, big data, as you might say, about customers, and then marries that with in-the-moment context. That can only happen if you have a sense of identity and can also bring in memory from the past. So in every Sierra conversation, there's the possibility of identifying the customer, saving memories, either implicitly, automatically or explicitly, and then extracting those memories at a future date for use in the agent. So it's very much first class primitive on the platform.
58:02Sierra:I think you'll see that happen more and more over time as well. Just as these journeys get more complex, as we see more and more wins from the personal touch, we already have a number of cases where resolution rate has gone up meaningfully from memory, whether it's just greeting you by name, remembering what you called about last time, knowing that yesterday you were on the phone for an hour and it was really frustrating. And so early on, we had that memory through customer systems only, but we found just from customers asking over and over, hey, can you just have this first class on the platform that it's helpful to have both seamless integrations with the CRM as well as on platform memory that really understands AI better than most CRM software does.
58:47Zack Reneau-Wedeen:How do you guys think about memory? I feel like you've got agents, multiple agents, interacting with customers throughout various stages of their buying experience lifecycle. So I imagine memory must be important. How do you guys think about it?
59:01Sierra:So memory is extremely important to the platform. And since the agent data platform introduction, which we launched back in early November, it's been a first class primitive on Sierra. So if you call, for example, my wife lived in Hawaii for a year. And so I was flying Hawaiian Airlines back and forth quite a bit. And on a couple occasions, for anyone who's brought a dog to Hawaii, there's a lot of paperwork involved. I'm excited for the Sierra agent that can help with that. But I would often add a pet in cabin, not that often, a couple times. And if I call back, you know, it's nice for them to remember why I'm calling, to know about me, to know I prefer aisle seats.
59:40Sierra:I'm a big user of the in-flight internet. Hawaiian has Starlink back and forth from Hawaii. And so these things, just what we've seen in practice is that if you know who someone is, you greet them by name, you remember what's important to them, and you show empathy in the moment, it increases all of the metrics that are most important to businesses, from resolution rate to conversion rate, et cetera. And so we've made memory first class on the Sierra platform, where during a conversation, implicitly or explicitly. You can basically store memories. And then the agent, if the same person calls back, can extract those memories.
1:00:19Sierra:The one thing to be aware of is you have to be really thoughtful about authentication. Because oftentimes if someone calls over the phone, you don't necessarily know 100 % from their phone number that it is this person. You know, some office networks all have the same phone number, maybe it's a family line, etc. And so every business has to think about what the policy is for allowing the extraction of memories and which memories are sensitive versus not so sensitive. Saying, hey, Harrison, thanks for calling again. That's probably fine. But if it's like, hey, Harrison, are you calling about your social security number?
1:00:55Sierra:That's definitely a different standard. And so we try to be very thoughtful about that with our customers as well.
1:01:01Zack Reneau-Wedeen:When you say you can implicitly or explicitly save memories. What exactly does that mean?
1:01:06Sierra:So there's kind of three layers of it. Number one is on a given conversation turn, you could say, I want to save this to memory. Number two would be at the beginning of a conversation, you could say, these are the things that are important to remember. Remember their birthday. That's always nice. I remember Cold Stone Creamery growing up, they would give you a free scoop on your birthday. You know, it's a great opportunity for brand
1:01:30Zack Reneau-Wedeen:Can you say you could say, so would the brand say, would the customer say this in the system prompt or would this be the end customer talking to the agent saying, hey, remember for future things that my birthday is on XYZ?
1:01:42Sierra:So the first one you said would be an example of journey building, an example of what I just said, and you would do it in journey building. You'd say, I care about birthdays as an agent developer or an agent builder at any one of our customer companies. The second thing you said would be the third category of memory, which would be just sort of remember important things. And it would be an important thing if the customer said, hey, I want you to remember this when I call back in the future. So whether you're deciding something in the moment, this is important. As an agent builder, I care about these things.
1:02:10Sierra:Or, you know, let the agent decide. Those are kind of three ways to structure memory storage.
1:02:17Zack Reneau-Wedeen:And when you think about the structure of memory itself, do you guys think about it as a knowledge graph, a vector store, a file system, TBD?
1:02:24Sierra:It's not super important. I guess I would say you want to optimize around retrieval. But the reason why I said it's not super important is that typically your knowledge base is three orders of magnitude larger than the memories for an individual customer. And so the retrieval and ranking problem is pretty simple. And I don't think it matters what structure you use, at least in our system today.
1:02:48Zack Reneau-Wedeen:I feel like memory is this really hot topic and everyone loves to talk about it. And there have been memory startups now for like two years, but I don't see any of them being massive breakout success. Why is that? Is memory not that important in the grand scheme of things? Is it so bespoke? Is it just too early on? Is it really hard? Like why isn't there a more established memory company or memory pattern? Do you have it turned on with Claude or ChatGPT? Not on purpose, although I think it is accidentally. And do you find it useful with your accidental turning it on?
1:03:22Sierra:I don't really there.
1:03:23Zack Reneau-Wedeen:Okay.
1:03:23Sierra:I ask because I would say that those are useful to me. I think what I said earlier about how when you're trusting us with memory, you're trusting us with authentication. That's part of it is that actually in order to pull off memory, you need to be trusted with something that has higher risk, you know, as well. And so the reason I mentioned ChatGPT and Claude is those are products that you are already trusting. And so I think they have more freedom than a B2B player would have, where it's like, hey, if I want to buy memory from you, I also need to buy authentication or verification at least or identification at least from you.
1:04:04Sierra:And I don't know exactly what the startups are in the space, but I would imagine like you're biting off more than you think when you sell memory.
1:04:11Zack Reneau-Wedeen:Talking about observability and evals for a bit, you guys have an interesting problem, I presume, where you have evals for your internal agents and for the maybe like general purpose agent SDK. But then I'm assuming your customers want to do evals themselves as well. Are those the same? Do you use the same tools for both? Or if they're different, why and how are they different? Typically not exactly the same.
1:04:33Sierra:So internally, the agent OS, you know, if you think about it as a series of tasks, and some of those tasks might be very complex, and some of those tasks might be more simple. The eval problem is more similar to the eval problem that any applied AI company has. When you think about our customers, I think the eval problem is much more complicated and involves things like what happens when there's background noise in voice and what if I have an adversarial user and I want to save these 20 personas and run all of my simulations against all 20 of the personas and make sure that it works. And so you end up with just a more complicated topography because a conversation by nature is very complicated and can go in so many different directions.
1:05:22Sierra:And so we built a product specifically for our customers to eval agents called simulations, and it supports all of these different things. I think it is probably you can tell when someone's building an agent, if they have good simulations, it's such a great unlock because you can make changes in a way that is constantly improving the agent and being sure that you're not regressing, especially as you get into big teams with complex agents that are doing so many things. I mean, I know you see this at Langchain as well. Like having really good evals is such a great unlock. And so we pride ourselves in addition to, you know, government governance and collaboration and review and making sure that, you know, you have workspaces so you can let Ghostwriter run free, but still review it before you make any changes.
1:06:11Sierra:We also have that simulation layer so that every change you make is tested against all the assumptions of the platform across voice and chat and many languages and many personas and this high dimensional space that you're going to experience in production.
1:06:24Zack Reneau-Wedeen:Going out from evals for just a second, because you said something around continually improving the agent. I want to talk about continual learning. That also ties into memory, I guess, a little bit. How do you think about continual learning in general? Does the Sierra platform support it in a fully, I'm assuming not like completely automated way, but like how far along are you guys and what do you think the future in continual learning holds?
1:06:48Sierra:Where we are today is you can automatically detect an issue with a monitor. Ghostwriter can automatically suggest a fix to an issue and you can review that issue and push it to your agent. And so you're still in the loop or people are still in the loop in all of the cases, but it's as automated as it can be with still giving you authority over that. I think in the near future, you will start to see the first cases of Sierra agents improving themselves where they have a confidence level to the fix. For example, if there's an error in a knowledge article and it can tell that there's a contradiction and it can go check the website and, you know, for whatever reason, it's very clear what the true answer is.
1:07:33Sierra:It could give you an FYI instead of needing approval. Same way, I do some work, I ask for approval, I do some other work, I give FYI. And so all of the primitives are there. It's just around the confidence that people have and the level of control that they want to have. And so we also don't want to get ahead of our skis there. Most of our customers, they want to review every change that goes into the agent. This is a really important part of their business. We don't want to pull the future forward too quickly. We want to move at the pace our customers are excited about.
1:08:03Zack Reneau-Wedeen:One of the things you mentioned, going back to evals, is monitors. What are monitors? And then you guys also wrote a blog called Monitoring the Monitors or something like that. I would be curious to hear about that.
1:08:12Sierra:Yeah, we have a saying in the company that the solution to all problems with AI is more AI. And so oftentimes you have something that's 90 % accurate and you figure out how to verify it 90 % of the time. You figure out how to verify that 90 % of the time. and so on and so on. And you have something that's, you know, three or four nines of reliability. And I think with non-deterministic systems, that's just quite a bit about how it works. And so similarly with a conversation platform, you can set up monitors that run on every conversation and look out for the things that you want to flag either for review or to create issues from, et cetera.
1:08:51Sierra:And it just basically gives you peace of mind, narrows the set of, hey, I don't have to wake up every morning and try to read 10 ,000 conversations. I can read five and I can say, OK, these five look good. I feel comfortable going on with my day. And so that frees up a lot of our customers to think about how do I actually improve customer satisfaction or resolution rate or some of these more strategic levers, as opposed to feeling like they need to review everything. So I think that's why it's one of our more popular features.
1:09:21Zack Reneau-Wedeen:You guys released TauBench, which is an eval for a few different agentic use cases. And I think you released a few other benches as well. Why do you guys invest in these and why should people check them out?
1:09:34Sierra:So I mentioned we have a research team and it's very exciting when you're building something to also think about how other people could use it. I mean, the distance that the AI space has come and how we've benefited just from all of the contributions to open source, you know, our knowledge engine, as I mentioned, runs on open models that we've fine tuned. It has felt like one of the areas where we can contribute because we actually know a lot about what it takes to build a good voice agent. I don't think anyone knows more than we do. We know a lot about knowledge retrieval. We know a lot about tool calling and following process.
1:10:14Sierra:And we know a lot about transcription. And so we've released, I think those are the four areas. There might be another one where we've released benchmarks in the sort of Tau cinematic universe. There's Tau voice, Tau knowledge, Tau bench, and Mu bench, which is the multilingual transcription benchmark. And so it really just started because the first Tao Bench was a lot more popular than we expected. And we were like, oh, people trust us to kind of say what good looks like in this space. And so we've continued to do more and more. And our research team has grown. And there's appetite. I think it also has this ancillary benefit of causing us to think about these problems in a very principled way.
1:10:52Sierra:And, you know, from kind of the first principles of what good looks like. And then we can evaluate our agents that way as well. So I think it has that benefit, but it is very path dependent on TauBench being a hit and, you know, TauSquared being the SQL being a hit as well. And then us just deciding, OK, let's do more of this. People seem to like it.
1:11:11Zack Reneau-Wedeen:How much does the core agent team use these to guide their harness choices?
1:11:17Sierra:Most of the benchmarks we use to evaluate providers more than to evaluate agents. And so, for example, we had, there's a really exciting new transcription model that came by the office and presented it to us. And so we were able to say, this looks really exciting, but we'd like you to run it against MuBench, and then it'll be really exciting. And so it really helps in the modular approach that we've taken. Like the reason we discover that this model works really well when there's silence in northern United Kingdom, but this other model works really well when there's speech is because of things like MuBench in particular for that one.
1:11:58Sierra:Internally, simulations is the main way that we eval the actual agents that are going out to production. So it's just too customer specific for us to rely on something as general as a benchmark.
1:12:10Zack Reneau-Wedeen:How do you create these benchmarks? Are they synthetically generated? Do you do a lot of data labeling internally, outsource it? I think it's a mix of all three.
1:12:18Sierra:I don't know all of the details for all of the benchmarks, but I know that we do a lot of stuff internally, just in terms of, especially when you're kind of in the zero to one phase, just figuring out what the right shape of the data is. Even when you work with external companies, they often want to see some number of examples from you. And then I think also being able to synthesize data when scale matters a lot, especially if you can do it in a reliable way, is very helpful too. It's harder for something like transcription where audio synthesis might be, you know, already in the training set of the transcription and that kind of thing.
1:12:57Sierra:But for things like text, I think it's easier.
1:12:59Zack Reneau-Wedeen:One of the things I think is pretty underrated in building agents is UX. So we've already talked about voice as a modality. We've talked about actually showing up as a chat GPT app. How else do you guys think about modalities or UXs? Have you experimented with generative UI in any form?
1:13:16Sierra:We have quite a bit. I think it's pretty vertical dependent as well. To give you an example, when you're checking in for a flight, if you have a hypothetically 12-letter last name with a hyphen in the middle of it and a first name that's hard to spell as well, hypothetically, then it might be helpful to type that in while you're on the phone. And so we see in industries like airlines, appetite for multimodal experiences, especially when there's a lot of reservation retrieval or input. For something like retail, we see exactly what you described, where really polished UI around product discovery
1:13:54Zack Reneau-Wedeen:and around recommendation moves the needle and makes a difference.
1:13:58Sierra:I think where Sierra is particularly differentiated is going really deep with customers, especially in specific areas and learning, what does it mean to build an amazing retail discovery experience? And then just from first principles, what's the agent that could help drive that versus what does it mean to build a great airline check-in experience or flight disruption experience, to the degree that can be great, it can be not terrible, I guess, then, you know, what's the right form factor for that? We've seen, and I think one of the reasons vertical companies have been pretty successful lately is that understanding the contours of each industry and each company really makes a difference.
1:14:43Zack Reneau-Wedeen:One of the things that I think you guys are actually best known for is your revenue model and charging for outcome-based pricing. How do you actually do that? How do you estimate the value that an interaction has? And is it specific to each customer?
1:14:56Sierra:This, I think, is maybe the number one operational reason or business reason why Sierra has been successful. It aligns the incentives between our company and our customers. And I think the phrase I like to use, which is a little bit cheeky probably, is if you don't understand the value of outcome-based pricing, your outcomes are probably not that valuable. Because when you're delivering, you know,$100 outcomes, and you get to keep a portion of it, everyone wants to row in the same direction. And it cuts through all of the prioritization and decision making that often will cloud and resource allocation that often will cloud enterprise partnerships.
1:15:39Sierra:So it's extremely valuable. And I think it's a big reason why we've been successful. I think it will just become the norm for companies that are doing differentiated high value activities. If your product really like feels a little bit more like a commodity, you'll start to see more usage based and seat based pricing because it's just simpler. An area, for example, knowledge based lookups are a little bit more that way, just question answering. And so in the case of question answering, that's not an area where you would have a high premium for an outcome of any particular sort. But if it's making a sale on a membership or selling someone a car, that's a really big outcome.
1:16:21Sierra:And so companies will be more than happy to pay for that. I think where we're seeing things going is intra-conversation outcomes to also thinking about more, as I mentioned, kind of the moments that matter across the customer lifecycle and driving outcomes on top of our agent data platform that kind of span that whole lifecycle.
1:16:43Zack Reneau-Wedeen:I think that's particularly interesting. You guys support multiple different outcomes. So you've got customer support and you've got sales. How different is the pricing between those and how many different of these like categories or templates do you guys end up having?
1:16:56Sierra:It really depends on the value. So you asked if it was customer specific. The answer ends up being that it sort of has to be. In certain cases, you are troubleshooting very complex setup to a device or something. and you have to try 15 different things to get it to work. And the average conversation might take 20 turns and the amount of context engineering to make that work might be very high. In other cases, you might have something where, you know, you're just resetting the signal on your TV and it's very quick and easy, or you're checking your balance with the bank and that's very easy. And so, you know, one outcome is very valuable and drives a lot of loyalty and one outcome is somewhat commoditized.
1:17:40Sierra:You might have some cases where there's, you know, an outcome that's tens of dollars and in terms of the, you know, money that the agent would earn. And then you might have some cases where it's, you know, much, much lower than that.
1:17:54Zack Reneau-Wedeen:And does that ever differ within a customer? So like in your example, I could imagine you could have an agent doing a really simple task of, oh, tell them to unplug the computer and plug it back in or something like that. And there's another one where like, oh, my God, who knows what's going wrong? And it like is a miracle that it solves that at all. If it's the same customer, will it be charged the same amount or do you differentiate even within those different types of requests?
1:18:18Sierra:There are cases where we differentiate. We're not dogmatic about it. What we found is that oftentimes the benefits of having our incentives aligned are so high that it's not worth negotiating every detail of what counts for what. And it kind of evens out over time and you do right by your customers over time and you build trust and contracts aren't infinite and you want to have a really high renewal rate and have them trust you with more use cases and these kinds of things. So we make sure that incentives are deeply aligned. And then on top of that, I think you can get really pedantic about the engineering of specific outcomes.
1:19:01Sierra:And maybe over time, the market will move in that direction. But I think you're missing the forest for the trees in that case, because of just how powerful the concept is. And so most of our customers are eager to find something simple that we all understand that feels fair, as opposed to trying to engineer like the perfect value for the for each outcome.
1:19:20Zack Reneau-Wedeen:Why don't you think there's more outcome-based pricing right now? Is it because there's not enough agents doing valuable things or because it's so operationally intensive for now because it's early on that you guys have just a built-up muscle of doing it and that's what allows you guys to do it so effectively?
1:19:35Sierra:I think it's probably a bit of both. I think that there are a lot of products that probably, as models have improved, find themselves in a position of being more similar to what you could just buy tokens and create. And then also there's just we're very early here. If I had to say, though, I would guess that the second one is more important and there will be a lot more of this. The same way someone doesn't care how many hours I work as long as I produce, you know, new products that are good. and I think that that will become true of agents as well. There will be a mix of building agents in-house on platforms like Landgraf and then there will be also, you know, buying products like Sierra to build agents on.
1:20:27Zack Reneau-Wedeen:Maybe switching to the last topic, which is just the type of people that thrive at Sierra. I think you guys are also pretty famously known for your forward deployed engineering or agent builder approach. Could you talk a little bit about that, both in terms of what those people do as well as the right persona to grow into that role?
1:20:43Sierra:I joined Sierra about two and a half years ago, and it was my first B2B job ever. I'd only worked in consumer products. And I love building consumer products. I love being like, oh, I could imagine, you know, my friends using this or my parents using this. But I never really loved growth and the idea of figuring out how to drive a couple percentage points of attention or a couple percentage points of usage. And what I learned when I joined Sierra is I love enterprise sales. I got a tattoo. So basically, the process of caring about each customer individually, saying one customer is upset, I'm going to call them right now and find out why and see how I can help, just felt very empowering as a builder in a way where building for a billion users on Google Search, for example, it was exciting in other ways, but it didn't feel like you could listen to each user and help them.
1:21:43Sierra:And in many cases, we have customers of Sierra that have gotten promoted in their organizations. They're building careers because of the agents that they built on Sierra. And so it just feels very deep in terms of those relationships. What I love as well, though, is that the end user of a Sierra agent is still a consumer in the vast majority of cases. And I think it's pretty rare to have a product that needs to be consumer grade where the product that you're building, it's a platform, but then the end user is really a consumer and you have to have them in your mind the whole time. But where you have kind of the enterprise sales process of building trust, of solving problems, of discovering value and then delivering that value for people.
1:22:25Sierra:And so I think the people that really appreciate those two things, the customer obsession and the craftsmanship, do very well. I think we've also discovered just with the rise of coding agents, certain things are more important than they used to be. Deep customer intuition. GPT 5.5 doesn't really have that. Agency, the ability to say, why can't I do this? One of our engineers that has really high degree of agency, her status message is just like, why not today? And so having that mindset, I think, is really important. And then the other thing, just as someone with a product background, is I think we kind of have a faster car than you need more pit stops kind of thing.
1:23:11Sierra:So like a Formula One car needs to get its tires changed more often than my Hyundai Kona. And the reason for that is, you know, it's driving faster, it's burning more rubber, etc. And I think we have a similar thing building products as well now where coding agents have allowed us to write code a lot faster and even to review it faster now. But certain things like product judgment and customer intuition are therefore actually needed more often, not less often. And so people that can bring that to the table themselves are in this amazing loop of moving fast. But people where it's one person's job to bring that and another person's job to do engineering, they need even tighter collaboration and, you know, more daily standups and that kind of thing to be successful.
1:23:57Zack Reneau-Wedeen:I really like that car analogy. I hadn't heard that before and totally resonates with what I'm seeing where product is becoming the bottleneck because it's so easy to code and you can make so much of things, but that doesn't mean you should. who ends up fitting this agent builder profile the best? Is this product people then? Is this engineers with good product intuition? Like what does it look like practically?
1:24:17Sierra:We're still figuring it out. I will say that people that have done both roles are often successful in the company. Our head of engineering, Aria, has been a product manager in the past. We have a number of engineers that have been product managers. I think those skills, knowing how to talk to customers, not just like what to say when you're in front of a customer, but how to find your way into the right conversations, having a high degree of agency, being really strong with communication so that you're getting, you know, product isn't the bottleneck anymore. Those are really important skills. I still think kind of knowing the right questions to ask and the right things to tell coding agents is really important.
1:24:58Sierra:So the systems thinking and the architecture design are really important. And so if you have not been an engineer before, it can be difficult. And so I think that the multidisciplinary approach is more important than ever. My own personal rubric, which is like very much in beta, is kind of this customer intuition, agency, product judgment, technical depth, communication, intensity. Because when the car, you know, you need to be really locked in when you're driving a Formula One car. And then one which is a little harder to pin down, but it's just kind of leadership where when there's more activity going on, the ability to draw it into the correct direction is really important as well.
1:25:42Sierra:So this is kind of the working framework in my head, but I'm sure there are lots of other things, too.
1:25:48Zack Reneau-Wedeen:How do you interview for agency? And I ask this because I think the guest we had on in the previous episode said the exact same word agency for one of the traits that they look at. And I asked him the same question. So now I'm going to ask you the same questions.
1:26:00Sierra:How do you interview for agency? So the most concrete way that we've changed our interviewing process is we have this AI native interview.
1:26:08Zack Reneau-Wedeen:And you wrote a great blog on it the other week.
1:26:10Sierra:Yes. And so you, by the way, Vijay and Aria and our engineering leaders, but I've seen it done and participated in the interview panels. And basically it involves building a product end to end over the course of a few hours and then reviewing it with a team. I think in that environment, you can see what people think is off limits or what's their job and what's not their job and how far they extend sort of what they're allowed to do. And if they're able to find opportunities that you would have thought, oh, maybe that they would think that's out of scope, bring them into scope and build great products on top of it.
1:26:48Sierra:You kind of see agency. You see that they have a sense that a lot is in their control instead of feeling like certain things are not in their control. And if you think about coding agents, they bring so much more into, I think, the locus of control, right? And so you can do more things. And if you appreciate that, I think it comes through in that AI Native interview.
1:27:11Zack Reneau-Wedeen:Thanks for listening to Max Agency. If you liked this episode, leave a review and subscribe. Send feedback or questions to maxagency at langchain.dev. We want to hear from you.
From the publisher
Zack Reneau-Wedeen is the Head of Product at Sierra, the conversational AI platform behind customer-facing agents for most of the Fortune 20. Before Sierra, he spent seven years at Google as the founding PM for Google Lens and Google Podcasts, then led product at Robinhood and CoinTracker. Sierra is mostly known for customer support, but Zack reveals how and why the company is building agents that span the entire customer lifecycle, from browsing and booking to sales and loyalty. In this conversation, he argues agentic commerce will be bigger than e-commerce, explains why he's a "monolith loyalist", and unpacks why, when a model looks dumb, the problem is usually you.
–
We also discuss:
- How Sierra's no-code layer compiles down to agent code, and back again
- Why most multi-agent systems just ship your org chart
- Inside Sierra's modular voice architecture: thinking, listening, and talking in parallel
- Why Sierra built a PCI-certified stack for voice payments
- How outcome-based pricing aligns incentives
- Why there's no breakout memory company
–
Timestamps:
(00:00) Introduction
(03:39) Analyze, build, release: how you build on Sierra
(07:54) Inside Ghostwriter
(11:04) Meeting models on their turf “80% of the time
(17:47) The one constraint Claude Code doesn't have
(19:35) Agent-to-agent: when an API call still beats MCP
(21:02) Why agentic commerce will be bigger than e-commerce
(27:31) Running models in parallel and ensembling transcription
(32:22) Inside the Agent Data Platform
(40:00) Context engineering: everything it needs, nothing more
(41:38) "Whenever you think the model's too dumb, the model's actually too smart"
(46:13) Why multi-agent systems are a trap
(48:44) Voice 101: latency, naturalism, and 60 languages
(56:11) When voice-to-voice passes 50%: the over/under
(57:03) Making memory a first-class primitive
(1:02:47) Why there's no breakout memory company
(1:08:02) Why the solution to all AI problems "is more AI"
(1:09:20) Why Sierra open-sources the tau-bench universe
(1:14:42) How outcome-based pricing aligns incentives
(1:20:26) Who thrives as a forward-deployed agent builder
(1:22:16) The Formula One analogy: why product is the bottleneck
(1:25:47) How Sierra interviews for agency
–
References:
- Agent2Agent (A2A) Protocol
- Anthropic
- ChatGPT
- Claude
- Claude Code
- Claude Mythos
- Claude Opus 4.5
- Codex
- Deep Agents
- Gemini
- Hawaiian Airlines
- LangGraph
- Model Context Protocol (MCP)
- Not Another Workflow Builder
- Redfin
- Sentry
- Shopify
- Silero
- SiriusXM
- Stripe
- Tau-bench
- Thinking Machines Lab
–
Where to find Zack:
–
Where to find Harrison:
–
Where to find LangChain:
–
Send feedback or questions to maxagency@langchain.dev




