How Capital One Delivers Multi-Agent Systems with Rashmi Shetty - #765

16 Apr 2026 · 54 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Capital One’s deployed multi-agentic AI approach and the enterprise platform that governs, evaluates, and scales it, using “Chat Concierge” as the main example.

Guest

Rashmi Shetty, Senior Director of Enterprise Generative AI Platform at Capital One. Background journey: thesis work on pervasive/real-time decisioning with actuators; then enterprise scaling via ML pipelines/AutoML; now focused on “closed-loop” agentic systems that plan, act, and observe.

Key claims

Multi-agentic is needed when a complex, goal-oriented task must be broken into steps handled by specialized agents (planner, intent disambiguation, tool execution, governance/evaluation/refinement). Capital One embeds model risk/compliance via platform guardrails, policy-based tool permissions, and runtime governance. Safety comes from policy-bound execution and human-in-the-loop checkpoints. Observability and end-to-end latency monitoring are critical across agent layers; evals must be end-to-end (not just per-agent). Closed-loop learning uses production telemetry signals feeding experimentation.

Notable examples

Chat Concierge for auto dealers/customers—scheduling test drives, financing approval, and trade-in estimates—chosen as a “high surface area, low risk” observability target.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to Chat Concierge

0:00 to 0:30

Discover how Capital One's Chat Concierge simplifies car shopping through multi-agent AI.

“Capital One's tech team isn't just talking about multi-agentic AI.”

Transitioning to Multi-Agent AI

0:30 to 1:08

Learn about the shift from classic ML to multi-agent AI systems in complex scenarios.

“To learn more about AI at Capital One, visit CapitalOne.com slash tech slash AI.”

Rashmi's Journey to AI

1:33 to 3:38

Explore Rashmi's journey in AI and how her thesis laid the groundwork for her career.

“Yeah, so I think my personal journey is something many, many folks might relate to.”

Capital One's Generative AI Role

3:38 to 6:11

Understand Capital One's approach to generative AI and its implementation timeline.

“And here we are now in the agentic world where we are moving beyond that decisioning systems and actually taking actions, which feels like a closed closure of the loop for me.”

The Philosophy of Multi-Agentic Systems

6:11 to 10:01

Learn about the motivation behind multi-agent systems and their application in complex scenarios.

“Let's maybe make this more concrete by talking about a specific system that you developed and deployed.”

Regulatory Aspects of AI Agents

10:01 to 11:30

Delve into how regulatory frameworks influence the design and operation of AI agents at Capital One.

“So multi-agent fits in very, very seamlessly in this scenario.”

The Role of Developer Experience

11:30 to 13:40

Examine how Capital One focuses on developer tooling and experience for deploying agentic solutions.

“One is that, you know, whatever platform you're deploying these agents on, you know, has, you know, regulatory aspects to it.”

Ensuring AI Safety in Business

13:40 to 14:01

Discuss the importance of safety and governance in deploying AI agents in business contexts.

“The first thing that you think of is can these agents execute safely in an environment?”

Speed of Development in AI

14:01 to 14:29

Learn how rapid development is integrated into AI systems.

“And then comes the speed of development.”

The Risks of Chatbots

14:29 to 15:04

Discover the potential pitfalls of AI chatbots in customer interactions.

“I forget the very specific scenario, but a business had a chatbot.”
Show all 31 chapters

Policy and AI Safety

15:04 to 15:49

Understand the importance of policy in ensuring safe AI operations.

“Yeah, I mean, this is where policy comes in, right?”

Requirements for Developer Systems

15:49 to 16:36

Explore what developers need when transitioning to agentic AI systems.

“And then you can, customers bring in their own policies and add layer on top of that.”

Operationalizing Agentic AI

16:36 to 18:13

Delve into the complexities of operationalizing agentic AI systems.

“Now we are talking about agentic systems.”

The Developer Experience

18:13 to 19:44

Learn how Capital One enhances the developer experience in AI.

“to the developer experience to keep in mind when you're implementing your systems.”

Observability in Agentic Systems

19:44 to 20:39

Understand the significance of observability in developing agentic systems.

“that the platform brings to the fore that helps them bridge that gap.”

Challenges in Multi-Agent Systems

20:39 to 22:45

Examine the unique observability challenges in multi-agent workflows.

“So I can go back to our WeChat example, Chat Concierge.”

Observing Agentic Behavior

22:45 to 24:19

Explore how agentic behavior is observed in complex systems.

“So they do want the service availability in productive systems as well, as much as it is needed in offline design time systems.”

Evaluating Agentic Systems

24:19 to 26:11

Learn how evaluation frameworks adapt for agentic systems.

“system to observe an end-to-end latency profile of how the different layers and systems are functioning together.”

Capital One's Approach to Model Evolution

26:11 to 28:00

Discover Capital One's strategies for evolving AI models.

“Talk a little bit about how the teams there approach evals for these types of systems.”

Capital One's Model Evaluation Strategy

28:00 to 28:30

Learn how Capital One accelerates model evaluation and deployment through specialized mechanisms.

“makes the evaluation of new models quicker and easier, easily provisioned within the platform infrastructure.”

Reasoning and Specialization in AI

28:30 to 29:20

Discover the importance of reasoning capabilities and specialized models in AI applications.

“This is another core underlying philosophy of Capital One, right?”

Architectural Decisions for Agentic Systems

29:20 to 30:40

Explore key architectural choices when building complex agentic systems.

“So this can be achieved primarily using specialized models and fine-tuning.”

Managing Complexity in Multi-Agent Systems

30:40 to 32:50

Understand how to optimize latency and manage complexity in multi-agent architectures.

“first application, the agentic application that we were putting in production, we already had a platform mindset from the word get go.”

Human Involvement in Agentic Workflows

32:50 to 33:50

Learn about integrating human oversight in automated systems for enhanced decision-making.

“What is that juncture at which human handoffs need to be done?”

User Experience in Agentic Systems

33:50 to 35:40

Explore how to design user experiences that effectively incorporate human feedback in agentic systems.

“So these are some things we absolutely keep in mind when we are building our platforms.”

Navigating the Multi-Agentic Journey

35:40 to 39:50

Find out how to support clients during their transition to multi-agent systems with various tools.

“Most of the times our lines of businesses already have a strategy for how this user experience is delivered to their end customer.”

Frameworks for Internal vs. External Agents

39:50 to 42:00

Understand the differences in governance and risk profiles between internal and external agent use cases.

“in the least constrained manner as possible while we make sure that the execution of this is safe and scalable.”

Implementing Closed Loop Agentic Systems

42:00 to 43:58

Learn about how Capital One integrates closed loop systems in their AI platforms.

“So the underlying infrastructure, the way we look at it is like we are all teams delivering unified GenEI SDKs for our agentic experience for our customers.”

Agentic AI Best Practices and Lessons Learned

43:58 to 46:36

Explore best practices for implementing agentic AI in enterprise systems.

“Like, you know, certainly they've been on this Gen AI journey with you, so they've gotten comfortable with LLMs versus traditional models.”

Scaling Agentic Systems at Capital One

46:36 to 50:56

Understand the challenges and strategies of scaling agentic systems.

“You have to look at latency as something that needs to be optimized end-to-end.”

Developers and the Learning Curve in AI

50:56 to 53:54

Discuss the evolving dynamics of developer experience in AI integration.

“And a lot of these elasticities baked into our underlying platforms.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:01Rashmi Shetty:Capital One's tech team isn't just talking about multi-agentic AI. They already deployed one. It's called Chat Concierge and it's simplifying car shopping. Using self-reflection and layered reasoning with live API checks, it doesn't just help buyers find the car they love. It helps schedule a test drive, get approved for financing, and estimate trade-in value. Advanced, intuitive, and deployed. That's how they stack. That's technology at Capital One. To learn more about AI at Capital One, visit CapitalOne.com slash tech slash AI. We moved from a classic ML world to a world where we have LLMs generating responses.

0:45And now we want to move on to a world where actions need to be taken. And when the problem that we are working on is a complex one, that's where multi-agentic comes into play.

1:08Rashmi Shetty:All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Rashmi Shetty. Rashmi is Senior Director of Enterprise Generative AI Platform at Capital One. Rashmi, welcome to the podcast. Thank you, Sam. Thanks for having me here. Thanks so much for joining us. We're going to be digging into the why and how of multi-agentic AI at Capital One. To get us started, I'd love to have you tell us a little bit about your personal journey. How'd you get to where you are? Yeah, so I think my personal journey is something many, many folks might relate to.

1:48It's funny how you ask me about my personal journey in a day and age where agentic AI is in its explosive form. so we have stopped pausing to think about it. How did we get here? So, yeah, I mean, yeah, it's been more or less organic and a natural evolution in my specific case, but if you ask me how can I pin it to a specific phase of my life, I would go back to my thesis, my thesis work. The topic of my thesis was around pervasive computing. Very interestingly, it was around perceiving your environment, getting signals from your environment, taking context-aware decisions in real time and having actuators trigger actions to change that environment based on your specific goals.

2:41It was a physical world in which this was manifesting. But in a sense, if I think about it, the principles were very much the same. It's about distributed intelligence. It's about intelligence being embedded in the system that you're living and operating in. And if you think about the governing principles where we were working on a design at that point in time, which my thesis topic was around, was specifically around decisioning systems, real-time intelligence. And that's where we are today. I think moving from academia to the industry just kind of taught me how to scale that in an enterprise world.

3:27So the next phase was explosion of data, building out AI ML pipelines, the AutoML platform journey where you need to operationalize pipelines at scale. And here we are now in the agentic world where we are moving beyond that decisioning systems and actually taking actions, which feels like a closed closure of the loop for me.

3:52Rashmi Shetty:How long has Capital One had a generative AI title or role? So generative, we have been, for those who don't know much about Capital One, we've always been at the forefront of modernization. So we've been in the AI ML journey for the past decade, even governed data. Even before that, we've been one of the first banks to go on modern cloud platforms. Similar to that, the Gen AI journey began a few years ago, like early 2023 was when we embarked on this mission. We had our first pilot, agent assist pilot, which went live in 2023. And then now we've had our GenEI presence and now an agentic presence in the past two years.

4:43Rashmi Shetty:Talk a little bit about kind of the motivation for multi-agentic systems. Where does the need for multi come in? So here is the philosophy behind it, right? We moved from a classic ML world to a world where we have LLMs generating responses. And now we want to move on to a world where actions need to be taken, specific goal-oriented actions need to be taken. And when the problem that we are working on is a complex one with multifaceted aspects associated with it, that's where multi-agentic comes into place. So basically, we have a large complex goal, which we have to break down into specific steps.

5:33And each step is basically narrowed to a specific agent. And that agent is tasked with the goal of executing that specific task. And then you move on to the next agent. So a multi-agentic system or architected orchestrated system comes into play when you have a large complex goal, which can be broken into steps. And you can take this entire system to fruition of that one complex goal that you have. So that has been the founding principle behind our backing of multi-agentic architectures.

6:11Rashmi Shetty:Let's maybe make this more concrete by talking about a specific system that you developed and deployed. It's called Chat Concierge. Tell us about Chat Concierge. What is it and why does it exist? So Chat Concierge for us was our beachhead initiative around deploying a multi-agentic solution. So Chat Concierge is essentially an auto dealership project or application that was deployed out to our auto dealers to basically bridge that experience between dealers and their customers and make it very seamless. So this is an auto buying experience that we wanted to make sure that we deliver the right solutions or cars to the right customer needs.

7:03So it was a multi-agentic chat experience that was brought to the fore with a human in the loop to our dealer customers, car buying customers to get the right match between the car that they need to buy. And we need to understand we are moving to a world where the car buying experience doesn't start at the dealership. It starts before. It starts when they go to their website and try to figure out, okay, what's the inventory? What does this dealership look like? And that's where the customer experience begins. It's at this juncture that we wanted to make this experience very seamless. that you know, okay, you figure out based on your family needs, your demographic needs, your specific personalized needs, what is that car that fits your, matches your specific intent.

7:58And then the subsequent actions can be generated, scheduling a test drive, etc.

8:03Rashmi Shetty:And did it start with the, did you know that you would end up doing multi-agentic at the beginning of this project, or did that approach evolve? Yes. So that's a very good question. So most of the times, I mean, when we are trying to think about the right tool for the right problem, we need to understand the complexity of the scenario and the use case that we need to determine. So we go in with a multi-agentic scenario where there is a certain amount of autonomous execution that is necessary in terms of decision-making, in terms of complex reasoning that needs to be done. When there is a multitude of user intents that needs to be serviced by a single application, this is where we need to discern if this can be solved by a single deterministic model or do we need to go with a stochastic reasoning system, which is the multi-agentic.

9:02In this specific scenario, there were a multitude of intents, So there had to be one agent that understands specifically this intent and tries to disambiguate by asking clarifying questions back to the customer. That is that narrow job. And then from there on, we had multiple tools that can get executed in the form of different actions that need to be taken based on the intent that comes in. So we have a planner agent that does this discernment. And then we have all other agents in terms of making sure that what we are delivering is well-governed, well-vetted, evaluated, validated against risk standards, validated against response accuracy, validated against the latency needs that are provisioned.

9:48And then finally refining the response to your customers in a format in which they humanly understand that. So this entire task goal had to be broken down. So multi-agent fits in very, very seamlessly in this scenario.

10:04Rashmi Shetty:One of the topics that comes up frequently when I'm speaking to your colleagues at Capital One is the idea of, you know, operating within this highly regulated environment. To what degree, you talked a little bit about validation as, you know, one of the roles of these different agents. To what degree does regulation play into that? And do these agents have a role in conforming to various regulations? So this is one of the biggest strengths and why we have been so forward in our agentic journey is because this is the DNA of Capital One. You can say that we are a tech organization that do banking, which means that we have all of the modern stacks that is necessary to build robust, scalable enterprise systems.

10:55At the same time, we are deeply embedded within the model risk framework that the banking industry gives you. So this means that we have a very, very robust model risk office that we work very, very closely with. We have all of the risk and compliance frameworks embedded within the platform, which appear as policies, as guardrails, as security enforcement, cyber enforcements across our different layers of the platform that get implemented across different threat boundaries of the agents.

11:29Rashmi Shetty:And so that speaks to a couple of things for me. One is that, you know, whatever platform you're deploying these agents on, you know, has, you know, regulatory aspects to it. Are the agents themselves like enforcing regulation relative to the platform or as distinct from the platform? So there are two aspects to what you're saying, right? So one is the building of the agents and the one is the runtime execution of the agents. These platforms come into the fore when you are governing agents in runtime. And that's where the massive, huge benefit of platforms comes into the fore. This gives the architects of that specific agentic framework the flexibility of focusing deeply on the design, whereas the platform brings in all of the governance and risk compliance that needs to be bounded to make these agents execute safely in any environment.

12:38The second aspect, the first aspect that you spoke about is the agent development in and of itself. Can the agents themselves bring in governance? Yes. For the specific domains, absolutely. Right. So you can definitely bring in that kind of governance layered on top of the base risk and governance that we offer from an enterprise platform standpoint.

13:04Rashmi Shetty:Is your team primarily focused on either the kind of developer tooling and developer experience for the Capital One developers that are building systems like this or the runtime platform that the agents operate in or both of the above? The purpose of this enterprise platform is to bridge the gap for our customers to rapidly deploy agentic solutions, develop and deploy agentic solutions at scale safely. That's the North Star. So whatever it takes to get our customers to that North Star is what this platform is responsible for. Absolutely. The first thing that you think of is can these agents execute safely in an environment?

13:51That's the first order of priority that comes in. Can they scale? Can we hit all of our customer metrics while bounding these specific agents? And then comes the speed of development. I mean, the speed of development is something that we are seeing all around us. All of these are datafully embedded within our platform to bring that experience, that developer experience, the tooling experience, the SDKs, the frameworks to the fore, so that customers have all of that at their disposal to build at lightning speed, deploy safely and securely.

14:28Rashmi Shetty:Yeah, I'm thinking about an example that we've all heard about. I forget the very specific scenario, but a business had a chatbot. I think it happened in Canada, had a chatbot on their website and the customer asked for a discount and the chatbot basically gave them a discount. You know, this would clearly be disastrous in a car dealer type of scenario. Like, how do you make, you know, generative AI and agents safe, you know, for, you know, these dealers that have a lot at stake? Yeah, I mean, this is where policy comes in, right? Policy bound agentic operations. Policy can be enforced in multiple different ways across multiple boundaries, guardrails, tool-based permissions, what an agent's goal is, what an agent's manifest is, and what is this agent allowed to act on.

15:25These are all bounded policy-based decisions that need to be made at specific egress and ingress boundaries within the platforms. And that is what the platform brings to the fore. These are the tools, capabilities, platform brings to the table. And at the platform level, you have some mandatory enterprise cyber level policies. And then you can, customers bring in their own policies and add layer on top of that. I mean, this is where evaluates come in, right? evaluation is all the more important in your current platform strategy more than ever.

16:08Rashmi Shetty:Digging into the users of the platform, the developers building systems, you know, talk a little bit about, you know, what you've observed about their requirements. What do developers need when coming from either traditional software or traditional ML AI? What's unique about this agentic world that requires new tooling? So I think one of the key differences in how we have considered classic ML as operationalizing classic ML versus the agentic AI systems is truly we have a world where we are deploying models, We are evaluating individual models, making sure it's precise, accurate, secure. Now we are talking about agentic systems.

17:01It's basically systems that span across multiple levels. There are several foundational things that need to come together to make this agentic AI system, the build and deployment of it at scale, a reasonable ask of the platforms. in the sense that agentic systems span across multiple things. There is data lineage and context lineage that needs to be passed across multiple agents without blowing up your context windows. There's memory service integrations, tool service integrations. There is easy location of data, bringing that governance, data governance, along with your agentic AI systems, scaling, latency mechanisms.

17:54So there are multiple layers. So what in the past, what used to be thought of as non-functional requirements such as latency today is product feature. It is baked into the experience of a developer. So these are some things that, you know, we are seeing a paradigm shift in terms of what we need to bring to the fore, to the developer experience to keep in mind when you're implementing your systems. For example, Capital One has been fabulous in this space because we have always considered, even in the classic ML world, that your AI advantage is actually, your data advantage is actually your AI advantage.

18:31So we have had a legacy where our fingerprint, Our DNA is to actually figure out all of these enterprise data pipelines, make sure it's locally available across the enterprise, build on the modern tech stacks. And then finally, when we are ready to go on this agentic journey, we had all of these things in place. So from that perspective, providing those reusable services and enterprise hooks, governance hooks, cyber hooks, became very, very easy because we had that in place. So which means that the developer experience is now highly focused on building out their layer of the specific agentic architecture.

19:23And we have the SDKs, the frameworks, the tools, all of that provision for them to accelerate that. So going back to your question, what is it that they need? they need a need for speed while staying safe. It's that governed, safe deployment with the necessary tools that the platform brings to the fore that helps them bridge that gap. And it's all the more important today than ever that the cycle times are squishing. So earlier model development takes months. You don't really have to start worrying about how do I take this to production yet. But today, you have to have that vision in mind. Pilot to production.

20:09What's that path? Is that path clear for me once I'm done with my experimentation? Because these days, experimentation are in days, not weeks or months.

20:17Rashmi Shetty:In terms of things developers need, one thing you didn't mention specifically, but you spoke to the requirements, for example, really understanding what your latency profile is. is observability. Talk a little bit about how the platform and the developers like approach observability for agentic systems. Yeah, yeah. This is also a paradigm shift, right? So I can go back to our WeChat example, Chat Concierge. So one of the reasons why we picked Chat Concierge was that this was a high surface area, low risk scenario to pick. This is a brilliant way to observe patterns of failure modes in production as well.

21:05So basically, essentially what you end up doing is you have your own mechanisms for experimentation. The golden data set approach is the most widely known. So you know what you know, you build a golden data set, you do offline batch of ALS, observe that and design your system for your North Star metrics and you deploy. But what we actually intend to do in agentic systems is to form a closed loop system. So as you know, agentic systems has got specific things, right? An agentic system classifies as an agentic system if it is having the ability to plan, to reason, to plan, to think, act, and then observe and learn back.

21:51So we're really looking at a closed-loop system. So the way developers are approaching observability today is that heightened awareness of that closed-loop system. So you are designing your system as best as you can in design time, doing offline evals, observing it. And then in production, you are observing all of the specific metrics that span across layers. Observability comes in different flavors depending on what you're trying to observe. It can come as how are the metrics showing up for a business user? How are the metrics showing up for an ops engineer? How are the metrics showing up for scientists, a research scientist?

22:30How is my model drift going away? So this is something that is observed in real time in your production systems. And that feedback is what comes in for continuous learning and continuous iteration. So, yeah, so I think developers have heightened understanding of this. So they do want the service availability in productive systems as well, as much as it is needed in offline design time systems.

22:58Rashmi Shetty:It strikes me that in a multi-agentic world, you know, this problem is the challenge of observability is compounded. Now you've got, you know, these highly, these fundamentally probabilistic systems that interact with one another. What are the unique observability challenges that you've seen with multi-agentic workflows? All the more important that observability comes to the forefront in stochastic systems like a multi-agentic application, right? All the more important for us to be able to replay agentic actions and try to understand how it functions. So, as I mentioned, there are several different layers that you observe, several different junctures that you observe.

23:43agent behavior needs observability along many different dimensions in terms of what are the tools invoked? What was the reasoning mechanism that led to that tool invocation? And overall, what was this context that passed across systems? Is there any potential for monitoring for latency optimizations across these layers? So observability across agentic systems, there is a standardization for sure in terms of observing agentic behavior, but there is also a standardization across the system to observe an end-to-end latency profile of how the different layers and systems are functioning together. So latency today is really a cross-functional effort, right?

24:32It has to be observed from an end-to-end mechanism. It cannot be a siloed approach today where just the models are evaluated for latency. So from an agentic standpoint, it's very much like any other system except for that agentic systems have specific nuances that needs to be observed. As I mentioned, the behavior, whether it's hitting its goal or not, whether it is invoking the right tool or not, what was the reasoning process that resulted in that tool invocation.

25:07Rashmi Shetty:And did you find that the combination of your existing observability tooling and infrastructure plus, you know, traces, you know, which are, you know, popular for without, you know, observing LLMs, like did those two things provide you what you needed or did you need to, you know, build new to support these types of agentic workflows? Yeah, as I mentioned, I think one of the key strengths of Capital One has always been that we have been designing for the future in the sense that we have had a robust observability stack. And whenever anything new comes in, all you need is these additional SDKs that cater to that specific nuanced layer or component that is being built out.

25:58You can think of it like plugins. So this is something that we, yeah, to answer your question, I think we've had the foundation always to go into this world today.

26:11Rashmi Shetty:And then you mentioned evals earlier. Talk a little bit about how the teams there approach evals for these types of systems. Yeah, evals is basically, you know, I mean, you can, there is a lot of literature around agentic eval systems versus classic ML eval systems. But the fundamental principle is very much the same, right? You have your golden data sets, you have your specific matrices that you want to kind of adhere to or hit, and your experimentation evolves around that. The only difference is that now eval frameworks are basically tuned to end-to-end evals rather than individual agent evals because individual agent evals gives you nothing unless it works upstream and downstream for the whole system.

26:59So that's the nuanced difference. So you have to kind of come up with your own evaluation framework for your use case. But as a platform, we offer different eval tools that is needed for you to build your eval pipelines, deploy them, provision that a golden dataset offer you, offer those sandboxed environments where this can be executed.

Read the full transcript

27:30Rashmi Shetty:How do you approach the speed at which models evolve in this space? Yes, this is not unique to Capital One, right? But there is a nuance there as well in terms of Capital One's risk-first platform strategy coming to help us, benefit us in this specific space, right? As new models come into the fore, there are specific layers that has been built in in terms of provisioning, inference layers, inference optimization layers, and benchmarkings that makes the evaluation of new models quicker and easier, easily provisioned within the platform infrastructure. So whenever there's a need and there's a new kid in the block, we have mechanisms to rapidly experiment on it, deploy and experiment on it.

28:28Rashmi Shetty:And does the presence of inference and inference optimization imply that you tend to favor self-hosted models? This is another core underlying philosophy of Capital One, right? When you are the most successful, when you can offer two things, reasoning and specialization. So reasoning capabilities with our agentic platforms, agentic frameworks in the platform, we are bringing that to the fore. Specialization is something that is very, very crucial, right? So for going back to the chat concierge use case, for us to provide that personalized experience to our customers. And we are privileged to be in the position of catering to many such lines of businesses who have similar needs of distinct personalization needs for their customers.

29:22So this can be achieved primarily using specialized models and fine-tuning. Student distillation gives you that control you need to have on providing personalized experiences as well as having some control over your latency metrics. So, yes, that's been our strategy.

29:49Rashmi Shetty:Which goes back to the point you made earlier about the value being the data, right? That's a unique asset that you have and can build around and doing so requires customizing the models. Right. Exactly. Exactly. I think that's one of the easier questions for us to answer, right? Specialization is not an easy task, but I think for us it's easier because we have all of the pipelines in place. Yeah. When you think about the chat concierge effort and other efforts, like what, and thinking broadly from the perspective of, you know, things that listeners might encounter in their own environments, like what are some of the key architectural decisions or compromises that come into play when you're building a system like this?

30:39Yeah, so I think one of the few things that I did already mention to you about while we are approaching this, we already kind of thought through in terms of a risk first platform approach, although it was our Beachhead customer, it was the first application, the agentic application that we were putting in production, we already had a platform mindset from the word get go. So that is one of the architecture decisions that we kind of made, if you may say. So the other thing that we were thinking of is in terms of a full stack approach, thinking through a full stack approach of, okay, where does our data reside?

31:20Where does our model optimization strategy stay? How does our across layer monitoring look like? How does it need to be optimized for latency? Our strategy around specialization using open source foundation models. Our approach for multiple agents to fulfill this complex goal that needs to be met in terms of understanding which of those different agents are going to be provisions. How does governance look like at this agentic layer? How does execution of tools be provisioned? So essentially, I mean, and how do we improve any part of the stack based on production learnings? So this is famously known as the orchestra problem, right?

32:12How do you orchestrate the entire lifecycle from pilot to production back to experimentation? So these are some of the key decisions that we had to make quickly to understand, you know, how do we optimize for latency for smaller specialized models? How do we optimize for cross model conversations? How do you optimize the context? Whether some specific techniques had to be taken in terms of choosing the right model for the right task. Yeah, and definitely the other decision is human handoff. What is that juncture at which human handoffs need to be done? What are those governed or policy-based actions that determine human handoff?

33:01Rashmi Shetty:To what degree do you rely on agentic standards like MCP or A to A or A to I? Like there are a lot of these emerging approaches to agentic standards. Is that something that you, you know, you're thinking about and worried about or are you, you know, mostly trying to leverage what you already have and not necessarily trying to apply these external focused approaches? No, that's the purpose of the platform. The purpose of the platform is to abstract away any complex underlying technical decisions that have been taken. For example, as you mentioned, tools, whether we go in for MCB or native tools, local tool calling.

33:45These are all things that are abstracted within the platform and the customers are open to, you know, just leveraging what we are exposing to them in terms of accelerating their journey. So these are some things we absolutely keep in mind when we are building our platforms. So we have definitely a standpoint and viewpoint on how we approach MCB, A2A, and all of that decisioning is baked into our platforms.

34:11Rashmi Shetty:You mentioned human in the loop a few times. Talk a little bit about those considerations and how you make it easy for developers to incorporate humans into their agentic systems. Right. These are a part of the agent orchestration, right? And along the multi-agentic architecture, these are junctures where we need to decide the bounded actions that agents need to specifically take. Where is that juncture? Where is that node where human needs to be brought in for validation? So this is a classic multi-turn, multi-pass agentic architecture where one of the decision makers could be a human. And these are things that we bake into via the orchestration mechanism and the frameworks that we expose for these orchestration mechanism in terms of, okay, where is that loop interrupted?

35:05And by a specific human after which the loop can continue.

35:10Rashmi Shetty:To the folks on the Capital One side who are monitoring the systems and those humans that are pulled into the loop, like, how do you think about that from a user experience perspective? Are you rolling out like agent inboxes to folks to kind of centralize all the decisions that they need to make or approval steps around these agents? Or are you integrating into, you know, existing channels, whether it's Teams, Slack, email, that kind of thing? Yeah. So I think one thing to understand is that human loops are typically deployed in high risk use cases, high risk scenarios. Most of the times our lines of businesses already have a strategy for how this user experience is delivered to their end customer.

35:56So meaning that there is a specific strategy in which these human loops, there could be a 24 by 7 support system, a support mechanism where the response times, it's all baked into the end to end SLA to your final customer, right? So the delivery mechanism in which the human is involved in the loop, every customer has a unique nuanced way in which they want to get it done. In some cases, there's a human gatekeeper who's always there. And in some cases, it's an offline notification. So as a platform, we need to be agnostic to how this notification mechanism needs to be. And we have to have an easy way to embed it within whatever is the operating environment within our line of business.

36:39Rashmi Shetty:And then when you think about this from the platform's perspective, like, are there, what are the major components that a user's interacting with? So when you say what are the major components a user is interacting with, technically the user is interacting with every single layer of the component where a service exists. So these services are brought very, very seamlessly to the user where there is a need for configuration. For example, guardrails. There are many types of guardrails. If the user has an input to be provisioned, experience is delivered to the user in a homogenized manner, but they are actually interacting with the actual service, underlying service, which needs to be configured.

37:31This goes for pretty much all of the layers, right? guardrails, memory, any kind of retrieval systems that you have in play. Yeah. So, yeah.

37:41Rashmi Shetty:I think I was imagining the platform there and trying to get a sense for how many boxes are in that architecture and thinking that it was probably a lot and wondering what they were. Yes. Yes. I think we hide the complexity of it. Okay. Okay. So, yes, there are a lot, but there's abstractions and you try to hide the complexity. What does that look like from a developer perspective? Are they, is it, I think that question, you know, what I'm thinking about in that question in particular is, like, are there, you know, libraries that the, and the answer to this is yes, but right. Are there libraries and they have to understand like this, you know, the model that's exposed by these libraries or are there like systems that they're clicking through to configure things?

38:35Rashmi Shetty:Probably a little bit of all of the above. Multiple layers of everything. So it's always a journey, right? What you're talking towards is, you know, a self-service journey for customers. There are multiple layers to it in terms of, you know, how they interact with the system. There are CLIs, there are SDKs, there are UXs, UIs. We have all the ranges. depends on which persona you're trying to serve. So, yeah, we do have all of those. And maybe the thing that I'm trying to move towards here is from the perspective of a person that's providing this platform, how do you manage the complexity for teams that are new on this multi-agentic journey?

39:19So I think customers that are new on the multi-agentic journey are one of our more naive users, right? So we have to offer them pre-built blueprints that they take and run for their use cases. There are savvy customers who know what they want to do and you offer them those developer kits. So we have all of these different ranges that we offer to our customers depending on where they come from, what their use case is, and how can they get to the fastest path to production in the least constrained manner as possible while we make sure that the execution of this is safe and scalable.

40:03Rashmi Shetty:Can you give us examples of the kinds of things that have been kitted so we get a sense for like the level of abstraction that you're operating at? Is it like customer-facing agent versus, you know, workflow agent, that kind of thing, or is it higher or lower? Yeah, well, internal and external agents definitely have a completely different risk profile associated with it. But that is all governance. That's all policy and governance. In terms of the developer experience, there are different kinds of use cases that can come into the fore. And I think it's pretty known in the industry what are the different levels and different bands of use cases that exist.

40:46you have simple summarization scenarios or complex multi-agentic scenarios like you just saw which is the chat concierge so we have this range how does your team interface with like the lower level

41:01Rashmi Shetty:infrastructure i've spoken to teams there that have operated like the kubernetes environments and things like that is that all hidden behind apis that your team has built to manage that lower level infrastructure? So as I mentioned, this is a cross-functional effort to kind of get all of this experience out the door for our customers. So we do have, as I mentioned, an inference optimization layer, which we partner with. They provide all of the hard work with regards to making sure that the right models are hosted in the right environments and optimized for inference are provisioned through specific SDKs and APIs.

41:44And these are SDKs and APIs that get embedded within our own systems along with the other services that we offer. For our customers, it doesn't matter, right? These are all different services, multiple of which you need to kind of go through with your agentic development journey. So the underlying infrastructure, the way we look at it is like we are all teams delivering unified GenEI SDKs for our agentic experience for our customers.

42:12Rashmi Shetty:You talked about on a couple of occasions the closed loop nature of agentic systems. What have you created to facilitate that? Traditionally, you have your golden data set and the way you close the loop is you you know, take runs that are either outliers or, you know, issues and you put them back into your golden data set for the future. In the, you know, Gen AI world, now you can, you know, change your prompts, you can, you know, change, you know, other aspects of your context to facilitate incorporating in like this closed loop, you know, signal. How do you think about that approach? As a platform, the platform's role is to make sure that there are hooks and tools provisioned along the closed loop journey and for the customers to come up with their own mechanisms in which how they want to specifically do it.

43:17So you have observability, which observes specific failure modes. Signals for that failure modes are captured. Those go back into the experimentation environment, which can be used by customers. And they can tune whatever they wish to tune to, you know, kind of realign their agentic systems to the newly changing environment in which it's actually operating in. So multiple hooks, some of which you already mentioned, you know, pure prompt tuning, model tuning, context management, retrieval, grounding, any of these. There are, these are all available in the experimentation environment. Got it.

43:56Rashmi Shetty:So then it sounds like from the platform perspective, you know, the key for you is to ensure that the data pipeline is solid and that they can access the data at various places where it exposes the signal that they're trying to take advantage of. Correct. Correct. Exactly. you. You know, talk a little bit about the, you mentioned the model risk officer or that model risk office that you have there, you know, there, anything you can, you know, share in terms of, you know, how agentic types of, you know, approaches has changed their relationship with that group? Like, you know, certainly they've been on this Gen AI journey with you, so they've gotten comfortable with LLMs versus traditional models.

44:47Rashmi Shetty:Does agentic change that at all? No, no, no. We are partners in this journey, right? I think that's one of our strengths, biggest strengths in terms of banks. I think the world is waking up to agentic governance now and trying to build frameworks around it. We have a robust framework. So these are our strong partners that help us understand how to bake into any new systems modern stacks, tech components that we develop, and how do we bake into what are the current regulatory compliance policies into this framework? So we partner very closely with this organization. So I wouldn't say how this relationship has changed.

45:31We are partners delivering towards the same goal. So I think that's one of our biggest strengths as a tech organization operating to run the bank.

45:41Rashmi Shetty:Taking a step back, you know, if you were talking to folks that have, you know, similar enterprise requirements, but are maybe further behind in the journey, you know, how would you talk to them about, you know, lessons learned, best practices, things that have made your successes possible? I kind of answered a little bit along our discussion topics today. So one of the biggest lessons is that treat - You've dropped a lot of breadcrumbs for sure. True, true. It's time to reassemble that bread now. So I think the core is that you really need to treat agentic AI as a system. It's truly a system.

46:27You have to start with governed data. You have to kind of put in that risk controls baked into multiple layers of your application or your system. You have to look at latency as something that needs to be optimized end-to-end. Take into consideration handoffs as well in this end-to-end experience for your customers in terms of latency expectations. And understanding that your biggest gains do come from post-production telemetry is also critical. I think, yeah, basically in a nutshell that, you know, as you move from the AI ML to the Gen AI world and now to the agentic problems, I think bridging that, bridging that world is important to understand.

47:12You don't throw away, at least for Capital One, we had a very, very good legacy of tech stacks that we had built, which we are leveraging to quickly jump onto the agentic bandwagon in the last year. So for us, especially in our case, reasoning became explicit. Specializing became modular, became very easy. Yeah, so data quality, governance, integration, keeping an eye on end-to-end latency, post-broad improvements. I think these are some of the key things that I would say that key lessons for anybody else wanting to go along this journey.

47:55Rashmi Shetty:For folks that are in similar platform types of roles, is there a different slant or take on that or more nuance or detail? For folks that are thinking about expanding existing platforms or building out new platforms for that matter to accommodate agents? Yeah, I mean, some of the nuances that I mentioned, you need to understand what services you are lacking to build this agent capabilities. What are the policies that you're lacking to build these agentic possibilities? Look at observability from a completely new lens. Instruments sooner. Define escalation thresholds. I think the problem of building agents has been solved for.

48:43We are moving very rapidly into a world where agent execution and governing of agents in execution environments is becoming more critical.

48:53Rashmi Shetty:Talk a little bit about kind of next steps for chat concierge, your team, agents at Capital One broadly. I think basically we have the necessary components. It's now scaling. We are at a lifecycle phase within the agentic lifecycle, development lifecycle, where we are now figuring out a scaling problem. How do we get across to millions of customers at the same levels of speed and governance and scale it out to multiple users, both from a developer standpoint as well as from a consumer standpoint. And when you break that down, like what are the, like how do you get there? Like, you know, certainly part of that is like technical scaling.

49:46Rashmi Shetty:Like, what does it mean to run agents and, you know, many, many agents in production serving many, many customers? But I imagine that there's organizational scale as well, onboarding scale. Like, how do you think about the various components of that? Yes, it's very nuanced, as you rightly mentioned. Like, there's definitely a technical component associated with it in terms of there's also a strategic component associated with it in terms of how do you manage your fleet of agents? across your company, across different lines of businesses. All of these things, you know, it's like kind of keeping in line with what is happening in the governance world around agents, keeping a close eye on it while you keep your scaling standards to these evolving standards out there in the industry and everything that's new and out there.

50:41So there's definitely a strategic nuance to that. But as with all software deliveries, you have to kind of figure out how your tech stack scales. And I mean, we, again, I'll go back to how we said that we are cloud-enabled, cloud-native. And a lot of these elasticities baked into our underlying platforms. It's just about, you know, kind of observing as we scale what we need to adapt or readapt based on changing standards.

51:12Rashmi Shetty:Are there particular friction points that you've observed in the relationship between platform and developer? Are there kind of next steps that you're pursuing in that direction? I don't believe this is, friction is the right word again. We're all partnering in terms of making sure we as a platform provider are providing to our developers and consumers. Developers, on the other hand, we have the benefit of a very fast feedback loop from all our user bases because we are all in-house. So I think it's more than friction. I think there's no time for friction. The world is moving at a rapid pace. So I think we have mechanisms that we have built in in terms of getting this rapid feedback loop with our internal and external users so that we can kind of get that back into the platform.

52:11Contributions from our partners is a huge part of it.

52:15Rashmi Shetty:Yeah, and I didn't mean friction in a negative way. It's just like it's a lot for developers. And, you know, when I'm out talking to folks, I encounter folks from, you know, a perspective of like, you know, I just still don't really believe in this stuff. It's broken. Like it hallucinates. Why are we doing this to, you know, let's go, go, go. And, you know, not to mention if you, you know, once you're in go, go, go, like there's just a lot to know and learn. I think one of the key things to remember here is that we don't take away agency from our developers. they need to have the agency to know to you know be able to configure and play around the system in uh based on their line of visions that is key so what you do is you offer all the toolkits developer kits tools framework services observability is key right to kind of give an agency as they observe the system they know what they need to do so yeah i mean learning curve yes like it's a compounded learning curve that we're all in.

53:15It's massive. Every small evolution in the past is compounding. So for example, we were in a world where data was commodified, like big data the past decade. That became storage, became cheap. And then compute became cheap. And with each of this, we are in a compounded learning curve right now. And our developers are experiencing that as well. So how do we bring the seamless experience to them in a way where it is beneficial for them so that they can jump onto the bandwagon very quickly without having to worry, should I, should I not? I think that's the goal. That's the goal.

53:53Rashmi Shetty:Awesome. Well, Rashmi, thanks so much for jumping on and sharing a bit about how you have been approaching Agentex Systems at Capital One. Super interesting stuff. Thank you. Thank you, Sam. Likewise, it was a pleasure being here.

54:10Thank you.

From the publisher

In this episode, Rashmi Shetty, senior director of enterprise generative AI platform at Capital One, joins us to explore how the company is designing, deploying, and scaling multi-agent systems in a highly regulated environment. Rashmi walks us through Chat Concierge, a multi-agent chat experience for auto dealerships that handles intent disambiguation, tool invocation, and human handoffs to deliver safer, more personalized customer journeys. We discuss Capital One’s platform-centric approach to AI agents and how it separates design from runtime governance, embedding policies, guardrails, and cyber controls across agent threat boundaries. Rashmi shares how the team approaches the developer experience for agent builders, observability, and evals for stochastic, multi-agent workflows; and strategies for model specialization, including fine-tuning and distillation. We also cover standards and abstraction, closed-loop learning from production telemetry, and key lessons for enterprises building agentic systems.

The complete show notes for this episode can be found at https://twimlai.com/go/765.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
How Capital One Delivers Multi-Agent Systems with Rashmi Shetty - #765The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 54 min
Listen in VO