In short
Podcast Notes: Evolving MLOps Platforms for Generative AI and Agents with Abhijit Bose - #714
Episode Overview In this episode of The TWIML AI Podcast, host Sam Charrington interviews Abhijit Bose, Head of Enterprise AI and ML Platforms at Capital One. The discussion focuses on Capital One's evolution in MLOps and the adoption of Generative AI (Gen AI) within their platforms. Key insights include the integration of cloud-based infrastructure, the use of open-source models, and the future of AI workflows in the enterprise.
Key Topics Discussed
- Abhijit Bose's Background
- Experience: Head of AI & ML Platforms at Capital One, previously led engineering teams at Facebook AI Research.
- Role: Oversees the development and deployment of AI infrastructure, including both traditional ML and Gen AI applications.
- Capital One's Platform-Centric Approach
- Platform Company: Capital One is described as a tech company that does banking; this reflects their commitment to a strong data culture and centralized platforms.
- Investment in Platforms: Aimed at optimizing resources like GPU for high-leverage use cases and applying governance frameworks centrally across the organization.
- MLOps Evolution and Gen AI
- Challenges of Gen AI: Focus on the complexities of deploying Gen AI, including the need for observability, monitoring, and managing model performance.
- Observability Tools: Discussion on the need for enhanced monitoring to accommodate new challenges posed by Gen AI, such as LLM hallucinations.
- Use of Cloud Infrastructure
- AWS Utilization: The platform is built on AWS, utilizing a robust control plane based on Kubernetes that allows integration of various tools and services.
- Flexibility with Models: The team combines open-source models like Llama with proprietary training and governance mechanisms.
- Fine-Tuning and Training
- Model Fine-Tuning: Emphasis on fine-tuning open-source models with proprietary data to meet regulatory standards and improve performance.
- Infrastructure for Training: High-Performance Computing (HPC) clusters are employed for training LLMs, addressing the demand for sophisticated resource management.
- The Role of Agents and Agentic Workflows
- Future of Workflows: Exploration of multi-agent workflows as a new frontier in AI, emphasizing collaboration between agents and human operators.
- Research and Development: Efforts to integrate existing frameworks (like Langchain) while also innovating internally to address scalability and regulatory needs.
- Talent and Skills for the Future
- New Roles: Introduction of AI Engineer and Applied AI Researcher roles focused on Gen AI capabilities.
- Skill Adaptability: The need for talent that combines technical expertise with an understanding of business processes and AI governance.
Key Takeaways
- Investment in MLOps: Continuous development in MLOps is crucial for effective deployment of Gen AI, with a focus on user-centered design and operational efficiency.
- Complexity of Observability: Monitoring Gen AI applications requires advanced capabilities to handle new challenges such as hallucinations.
- Role of Open Source: Leveraging open-source models allows Capital One to customize and adapt models while maintaining control over sensitive data.
- Agentic Workflows: The potential for agentic systems to streamline back-office operations and enhance software development is significant, with ongoing research into optimizations.
Conclusion Abhijit Bose shares an optimistic view on the transformative potential of Generative AI at Capital One. With ongoing investments in platform capabilities, the integration of advanced AI workflows, and a commitment to developing specialized talent, Capital One is positioned to lead in the evolving landscape of AI and machine learning.
For complete show notes and additional resources, visit [TWIML AI Podcast Episode 714](https://twimlai.com/go/714).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00To create that feedback loop, you know, from all the way to the IVR system, to all the way data coming back into our model for refinement. we realize that there is that part of deploying Gen.AI that's not always talked about, but it can get very, very complex, you know, chaining all those events.
0:36All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Abhijit Bose. Abhijit is head of Enterprise AI and ML Platforms at Capital One. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Abhijit, welcome to the podcast. Thank you, Sam. It's a pleasure to be here today. I'm super excited to have you on the show. We'll be digging into all the things you're working on and thinking about from a platform's perspective in support of Gen.AI at Capital One. But before we dig into that, I'd love to have you share a little bit about your background.
1:15As you said, I lead the enterprise AI ML platform engineering teams at Capital One. What it means is that we build and deploy all of the core AI ML infrastructure for the company, and that includes classic machine learning models like GBMs, neural nets, but more recently, Gen AI applications, you know, that involves LLMs, or agentic workflows. I have been at the company for four years. In that time, you know, my team and I, we have pretty much rebuilt the machine learning stack from the ground up in AWS to the point now where almost all of the companies, data scientists and engineers, use our platforms to build and deploy ML at the company.
2:06It's one thing to experiment with machine learning, but when your most critical, mission critical credit approval and fraud detection models run on your platform, it's a whole different kind of a ballgame. So it comes with a lot of operational responsibility. just building that MLOps muscle over the last four years has been a really amazing experience for me personally and for my team. And with Gen.AI, it's the scope and the possibilities of what we can do is just exponential from here on. And before Capital One, I was the head of research engineering for Facebook AI Research in the East Coast.
3:04So New York, Montreal, and Pittsburgh labs, the engineering. I founded a lot of those engineering teams. That was, again, a great experience. Yeah. That sounds like an entirely different level of scale than what I'd expect at Capital One. How do they compare? It's, you know, in terms of complexity and the kind of the sophistication, I would say what we are doing now at Capital One is probably similar. But FAIR was much more academic style research, you know, so we built, we did a lot of academic style, you know, research in computer vision, NLP, you know, robotics. We build prototype systems. And then for some of those systems, we actually scale them to the entire company.
3:56So kind of similar things we are doing here too in our science team, you know, especially in AI. So we have the research to production path here as well. So yeah, I've been very lucky to be able to do this for a job. That's awesome. That's awesome. Now, I've interviewed many of your Capital One colleagues over the years. Many of those conversations have talked about platforms and MLOps. We had speakers from Capital One at our conference talking about some of the data platforms you were building at that time. Platforms is something that Capital One has been building out and investing in for quite a long time.
4:44You know, talk a little bit about the role of platform at Capital One and kind of how you see, you know, what you're trying to accomplish in investing so much in building out platforms. I will say that, you know, we are a platform company. And having come from Facebook, I have seen the power of centralized platforms like enterprise, you know, scale platforms. Very similar what I see here. There is a very strong... Meaning a platform company as opposed to a bank? That is true, actually. We say Capital One is a tech company that does banking. And it is really true. And part of being a tech company is to have platform mindset and having a very strong data culture.
5:32I see those same kind of ethos in our DNA as well. So, as we advance in our AI journey, we are taking a very similar platform-centric approach. So for Gen AI, for example, we are building all of those AI capabilities in a way that the rest of the company can build on and leverage. And as you know, if you are investing a lot in GPUs and training clusters and inferencing services, those are expensive investments. So being a central platform actually helps to make sure we are using those GPU resources for the highest leverage use cases. We can optimize continuously. The other thing that I see is that when you have central platforms, you can apply a lot of the governance and controls in one place rather than having a federated set of controls everywhere in the company.
6:38And I have seen that personally for machine learning. Like we have all of the kind of the data, cyber, all of the kind of the governance that you have to do as being a bank. We apply that in our machine learning platform in one place for all of our users rather than every organization trying to figure them out. So that I see a big advantage of being a platform company. The other thing I'll tell you in terms of our platforms, ML platform happens to be the platform with the highest NPS score of all platforms in the company. You know, this is a personal kind of pride, you know, for me and our teams is that we are so focused on our users.
7:29We have thousands of users across the company that use our platform. We really listen to them, you know, solved a lot of their pain points. And that's kind of the power of the platform. When you do this right, it solves many, many challenges in one place. I'm curious who your users are for the platform. Sonoma Enterprises offer a variety of levels just like the vendors offer, where a user can interact with the platform at a fairly low level, at a notebook or developer-oriented level. and maybe you have applications running on the platform that they consume directly, forecasting, for example. Are you doing similar things there?
8:20We have different user personas on our platform. There are what we call power users. They want absolute control over their machine learning workflow. So they use our SDKs to actually code the model and the features, but they can deploy it on our platform without our engineers, like platform engineers, helping them. So it's a very self-service environment, even for very sophisticated users. and even over the years, what we have done is to really make it easy for them to distribute their compute over a bunch of CPUs, and we are doing the same thing with LLM on GPUs, for example. So we do a bunch of things to make it easy for our users, but there are also users who like to use more graphically.
9:23They can connect to data and add a few steps. So it's all like low code, no code kind of environment. But behind the scenes, those are automatically compiled and into code and run on our platform. So we address the pain points and needs for a range of users on our platform. And you mentioned that over the past several years, you've been migrating the platform into AWS. Does that mean that you're consuming the AWS services and that is your platform like SageMaker? I have the impression that you're actually doing more kind of building your own on top of whatever infrastructure you're using. So, I mean, this is a great question, right?
10:15A lot of companies are grappling with that, right? How do we build a machine learning platform on the cloud? We spent a lot of time early on, like in my first year, for example, on designing a robust platform control plane. And that's based on Kubernetes. and we enhanced that control plane in a bunch of ways so that it's easy for us to take something from SageMaker, something from open source or wherever we feel like we can add a lot of IP. For example, when it comes to automating a lot of the governance steps, which no other platform or publicly available platform can do it because it's very specific to our processes and model risk framework.
11:13So we have that flexibility of choosing different things from different places, including contributing our own in our platform today. and yeah the other thing i'll tell you that you know actually helped us on the jnai journey uh because now we are in a position to you know kind of extend our original machine learning platform to more like llm hosting or llm training you know fine tuning uh you know agentic workflows So we took the basic kind of the architecture and just extended that to Gen AI use cases, and we were up and running very quickly. I think if we just went for a vendor solution in the beginning, it would probably have blocked us from making so much progress in AI and Gen AI.
12:17Before we dig into Gen.ai and how you're tackling that from a platform perspective, I'm curious how the relationship between traditional ML workloads and Gen.ai rolls out. are you have you kind of shifted or are you seeing your users kind of shift their focus to dive into Gen.AI or are all the same traditional MLAI work streams still kind of continuing when I talk to some folks the sense is that Gen.AI kind of took all the air out of the room of traditional ML, AI, you know, workloads. And then other folks describe it more as it kind of piqued the interest and it kind of raises it to an executive awareness level.
13:15But a lot of the problems that they end up wanting to solve are best solved with traditional ML and AI techniques. And so it actually, Gen AI kind of pulls those along. I'm curious what you see there at Capital One. We are actually investing in both. There are solid examples of where machine learning, as we have practiced traditionally, would still be very, very relevant, especially when it comes to create modeling and a few other things, where it's not easy to just switch to another technique without a lot of research and development. So we continue to invest, you know, in traditional machine learning techniques.
14:03And then in their cases like customer servicing or, you know, other like fraud detection, where we have invested a lot in AI and Gen AI to be able to harness the power of, you know, those kinds of approaches. So, yeah, we are doing both. Taking a step back and thinking about Gen.AI from a platform perspective, how do you think about the extension of an existing MLOps platform and set of workflows into Gen.AI? There are similarities, but at the same time, I would say there are quite a few things that are very different for Gen.AI. I will start with the observability as the area where it really starts to make a difference between traditional machine learning and Gen AI.
15:02In traditional machine learning, you deploy a model in production. We have many hundreds of models every day making all kinds of decisions. you have to have solid monitoring in place in terms of drift of data, input features, the model scores themselves. And so there are alerts, alerting that needs to happen, and then you have to be able to collect all the logs to figure out what's going on with the model. Is it a feature? Is it something changed in the model execution environment? so all of that is you know all of that equally transferred to the JNI side but you also have LLM hallucinations you know that's net new in the observability world you have to have guardrails and then you have to log everything in those guardrails both input and the response in case of agentic workflows you also have to see which tools they are planning to call because tool execution is so much part of the agentic workflow.
16:16You need to make sure that the execution of those tools is being properly logged. They're being properly governed. They're called in a properly governed way. So observability becomes very, very not just important, but it also becomes very complex. in the LLM world. We had to build a lot of infrastructure on top of our regular anomaly detection and model monitoring in case of Gen AI. It just, you know, and we are still building a lot of infrastructure on that. And you said on top of your existing monomonitoring and anomaly detection, how does the observability for Gen.ai interact with or take advantage of the services that you had in place?
17:12So this is an active area of work for us. But what I would say is that, you know, the way a lot of our anomaly detection algorithms work, you know, that's also a platform that we have. So platform is very much part of our DNA, as you can see. So all of our algorithms for detecting deviations, you know, and then alerting, that's a part of our machine learning platform that also operates as a separate platform. So you can actually, you know, basically point it to any data source and it will automatically start generating these alerts and, you know, with some little bit of configuration that you have to do.
17:59So we are using very similar algorithms, but it's just pointing to text and other types of data, not always like structured data. So we are leveraging what we have built, but extending them to new types of monitoring. From a model perspective, how are you thinking about model use and model selection? Are you using open models? Are you using models delivered as a service via Bedrock, for example, or another service? So we generally use open source models like Lama as the base. And then, you know, we take those models and then we fine tune them with our own data and customize them to meet our regulatory and governance thresholds.
19:03You know, and then we also test them very extensively, you know, with respect to hallucinations, accuracy, depending on the task that, you know, you are deploying them for. And the other thing I will say is that we also host these models completely within our premises. like within our own AWS environment. So no data ever leaves our, you know, perimeter. Was there a transition point at which, or what was the transition point at which the LAMA models, you were able to commit to the LAMA models in that way and didn't feel like you were leaving too much on the table by not choosing a best-of-breed, you know, vendor model?
19:50So, I mean, we constantly look at all our options, including third-party hosted models as well. So this is like, LAMA is just one example of the open source models that we have been using. But our strategy is that we will constantly benchmark these models and pick the right one for the right use case. But we wanted to look at models that we can host ourselves and, you know, we can properly safeguard our data. And but at the same time, not, you know, sacrificing the accuracy that we want for some of our use cases. Are you to what degree are you exploring fine tuning as a way to get the most benefit out of those models?
20:43We do a lot of fine tuning in-house. We fine-tune to specific tasks and specific domains. For example, as you know, summarization is one of the harder tasks. Even if you look at MMLU benchmarks, often even the frontier models, you'll see the summarization accuracy is a little bit lower than, let's say, search, for example. so you know you will see much better accuracy and relevant scores if you take one of these open source models and then fine tune and customize it you know with your own data in case like in our case it will be capital one data you know some corpus of text you know for that use case that we have identified yeah so that's pretty much has been like our approach and it has served us really well.
21:46You know, if we want to get into like specific use cases, I'm happy to talk about one use case where, you know, we took Lama and then we fine-tuned it and then we saw much higher accuracy for both search and summarization. Oh yeah, I'd love to dig into that. So this is a system, you know, it's kind of like a rag pattern in terms of like, you know, the workflow pattern. When you call our call center, you know, let's say you have a question like, if I have a declined card transaction, you know, would it count towards my, you know, like my limit, you know, credit limit? or, you know, if you're going to, you know, Canada and you want to put the rental car, you know, on the card, you know, would it come with insurance, for example?
22:49Like common questions, you know, sometimes, you know, it's also about policy. Sometimes it's about a procedure, you know, like early payment on something. So traditionally, our agents would do a very similar to TF-IDF kind of search, keyword search, and try to find the link to the document that has that particular answer. And it would take them quite a bit of exploration of whatever the system returned, going to different links, and then try to find the answer and then answer the question. You would take the agents a lot of manual. Looking around manually. Yeah. Okay. And the human agents we're talking about.
23:39The human agents. AI agents. That's right. Sometimes agent is actually overloaded, especially when you talk about human-in-the-loop systems, like a human agent and an AI agent. And customer service use cases, it comes up all the time. Exactly. Yeah, that's right. So with our new system, basically it runs on our platform. We took all of those knowledge bases that the agents look up, we indexed them and put them in a vector database. And we also fine-tuned our LLM with a lot of the knowledge and the type of question and answer tasks these agents produce. perform. And then we noticed that the accuracy went up a lot.
24:34So now they only have to look at the first or the second link. They don't have to hunt. And the quality of the summaries that they can actually relay back to the customer on the line, those are also getting much better. And it's also like from an AI safeguard and putting together the appropriate guardrails, It's still human in the loop because our human agents are still in charge of the conversation. And we can also collect their feedback. They can indicate on the system what the LLM and the RAC system actually returned as the response. And we get a quick feedback from them. And yeah, the one thing I would say is that it was really interesting working on that project because to create that feedback loop, you know, from all the way to the IVR system to all the way data coming back into our model for refinement.
25:43that, you know, we realize that there is that part of deploying Gen AI. That's not always talked about, but it can get very, very complex, you know, chaining all those events. What's a specific example of where you see complexity there? You know, chaining all the events together, like, you know, exactly all the logs from different systems. if you don't put automation in there, it'll be very difficult to kind of like create that data set, that annotated data set that you want to, you know, retrain the model, for example. Are you trying to understand from the IVR logs whether the user gets to where they're trying to get the first time or whether they have to, you know, spend a lot of time navigating the system just to get to their responses?
26:34Is that the kind of feedback you're pulling into the loop? No, we analyzed those, but that was not the intent of this particular system, you know, because this was really our one of our forays into Gen AI in customer servicing. But we were mainly focused on the question answering part of the process. Got it. Got it. That system is now being used by 20 ,000 agents, human agents at the company. So it's very widely deployed, you know, pretty much a large portion of our, you know, agent community. They have been using this system now. Oh, wow. From a platform perspective, what did you have to put in place to enable the fine tuning?
27:18A lot of work went into this collection of the data, you know, properly annotations of the data. We also set up HPC clusters, you know, like the difference between traditional machine learning, you know, where you are training a GVM or, you know, like decision trees. in almost all of those cases, you can train them on one or two GPUs and they may run for like a few hours to a day. But for LLMs, you need HPC clusters, very tightly interconnected file system, high-speed networking and the GPUs to work together for weeks at a time. And for sure, you know, there'll be a few GPU failures, you know, every day.
28:09And that can bring the whole, you know, training or fine tuning job down. And then you don't want to waste a lot of that expensive hardware, you know, time to go wasted. So we had to do a bunch of things on AWS to create our own HPC environment for training and fine tuning. So that was all in-house. We basically got the GPUs from AWS, but then we built our own fine-tuning stack on top of that. We have checkpoint restart and all of those things that you need to have a production system. And is that also Kubernetes-based? That's also Kubernetes-based, yeah. And do you see, it's been a while since I dug into Kubeflow.
28:58do you see Kubernetes evolving or are there specific projects on the Kubernetes side that are evolving to support these kinds of fine-tuning HPC workflows? Or is that something that you're expecting to continue to have to build yourself? Yeah, so that's a great question, Sam. In terms of Kubeflow, we do use it. We are actually part of the, like some of our engineers are part of the open source Kubeflow project. They commit in the GitHub. And a large part of the company actually use Kubeflow on our platform. We offer that as a service, a hosted service, basically, on our traditional machine learning part of the platform.
29:45In terms of, I think, Kubernetes-based training, whether it's HPC or traditional machine learning, I think that will continue for some time. If you think of some of these large clusters that the labs have set up, some of where the frontier models are built, a lot of them are based on Kubernetes clusters, Kubernetes as the orchestration for containers. But I think the community needs to develop more tooling for scheduling jobs and workload management. I think those would become more common as time goes so that we don't have to build a lot of those things ourselves. Interestingly, I'll tell you, when I first joined four years ago, we have lots of great Kubernetes engineers, but I was like, why do we have to build our own Kubernetes stack?
30:49Why can't we just go and just get that as a service? So over the years, we have worked with AWS to leverage EKS, for example, Elastic Kubernetes Service. So that has made our life a little bit easier. but there are lots of other things that you have to build on top of the Kubernetes stack to really make it easier for our users, scientists, engineers, analysts to be able to use the infrastructure. So that's where we focus these days. Got it. So the Kubernetes community has some work to do to build the HPC semantic layer on top of the underlying container management. Yeah, yeah. You mentioned data annotation was a big part of the work that went into fine-tuning.
31:49Can you talk a little bit about how that differs between the traditional MLAI and Gen AI and what are the things that you had to build to support it? So in traditional machine learning, you take a, like, let's say you want to build a model for fraud detection. So typically you will get a, you know, a set of transactions, you know, you'll divide the data into two parts, like your, you know, training data set, maybe three parts like validation and test, you know, data sets and all that. But then you will have a labeler to label them, you know, as fraud or not, right? Like positive or negative samples.
32:32And then in some cases, you can use some automated knowledge that if nobody complained or the customer did not come back with a fraud or there was no fraud alert, etc., those are maybe not a fraud transaction. So all you need is basically a labeling, like a methodology, and there are systems that can automatically label with a model and with some human levelers validating them. In case of LLM, whether it's, you know, fine tuning or pre-training, it's order of magnitude more complex because you are really creating, for example, an input prompt, you know, like let's say it's a, you know, it's post-training input prompt, you know, you want the answer.
33:29you need to extract that answer, you know, maybe from your corpus of knowledge, you know, so you need much more refined, you know, setup capabilities, not just for humans to review and evaluate, but you also need things like LLM as a judge, you know, to actually, you know, score some of these responses. So it becomes much more complex, this whole annotations and data creation portion of the pipeline. Is the solution to that to deliver platform services or is it like analogous to annotation tools? It's a combination of both. for us. We think there is an opportunity for leveraging platform even in that context.
34:24Plus, there are lots of tools today in the market as well as in open source that we are looking at to be able to do those annotations as well. And what are some of the ways that the platform can support that workflow or pipeline? I would say automated scoring automated evaluation, you know, from production systems all the way to the training system that we can automate, you know, and the tools can play a role, but they will be embedded in our platforms. I guess when thinking about deploying all these fine-tuned models, the next question is like, how are you doing inference? You mentioned the supercomputer for training or supercomputer cluster for training or HPC cluster, I guess is what you said specifically.
35:17I imagine you've got inference clusters as well. That's a different level of scale than in the traditional world. Talk about how you approach that. Yeah, we use GPUs in our inferencing platform. But the interesting thing is that there is so much science happening in the inferencing space. As you know, reasoning during inferencing is actually becoming more and more like a more popular way of being able to answer complex questions, for example. So those things are not going to be possible without GPUs, for example. So we built an inference platform. There are a bunch of GPUs behind the scenes where we do a lot of optimization around caching.
Read the full transcript
36:15And there are a lot of known techniques like speculative decoding, those going to our platform as well. And then this is one area where I would say having great scientists in your team really helps. So this is an area where science and engineering, they have to work together from the point the model is trained all the way to deployed in your inferencing system. On quantization, a lot of these optimization like caching, speculative decoding, you can't just build a system without considering all of the scientific aspects of the model itself. So that's something we do. We are very fortunate to have a science team working very closely with the engineering team to make these things happen.
37:14And how early are you thinking about optimizations from an inference perspective when you're kind of building out a project? Oh, from the get-go. Like we look at, you know, you know my my manager once told me that you know like inferencing if we are not careful you know can kill us right in in this business so so you have to like one of our targets you know for my team is to continuously lower the cost per token and the latency you know those those are two KPIs that we have to lower on a continuing basis. And your objective is to kind of keep them flat while you are ramping up your use cases and number of users you are scaling.
38:07But they cannot be linearly increasing with the number of users because that will be very expensive. So, you know, that's another area where we are very focused on, like, you know, having, we have a few scientists and engineers whose, like, focus is to lower the cost per token and latency from our inferencing system. Would you say that all of your models end up using some type of optimization technique, whether it's quantization or similar, or is that only a portion where it's most critical? How broadly are those techniques applied to the various models that you have in place? It depends on the use case.
39:00And, you know, we have, you know, we have, you know, like playgrounds where people kind of like experiment and all that. You know, those are not production systems, you know, so, but in production, you know, we try to squeeze as much as possible, like effective utilization of the GPUs, you know, memory, GPU like processing, you know, throughput. maximizing throughput through the system. So that's where you need those techniques. And what thoughts do you have on GPU alternatives, whether these are, in the case of AWS, Tranium and Inferentia, but there are many others coming online. Those things that you see in your future, are you excited about them?
39:51Do you think the GPU is going to get you there the whole way? How do you think about that whole world? I would say most enterprises or labs today use GPUs. I have seen some teams actually effectively use GPUs, for example. In some cases, I've seen that. We are also looking at alternate architectures, but Tranium, as you said, but I would say the workhorse for us is still the NVIDIA GPUs. But we are keeping an open mind about this. So we are going to explore what will work for us. And then we'll add up some architecture, alternate architecture. But one thing I would say, Sam, as you know, we are 100 % on the cloud on AWS.
40:52That actually has a huge benefit. benefit. It has been for a while. It has been for a while, actually. I think we started our journey in 2017, three years before I joined. So I was lucky. I already inherited a cloud environment for machine learning. So being in the cloud actually makes it much easier to be able to to test and learn from different architectures. If you are a traditional on-prem enterprise, you have to literally figure out how to get a Tranium or TPU. It's not even possible, right? I mean, those are offered in a cloud environment. So I think being in the cloud in this area really helps in AI ML.
41:43You mentioned kind of inference scaling, the O1 style reasoning oriented models. Is that something that you started experimenting with building around LAMA or are there already? I think I actually saw something recently where someone created a reasoning style model around LAMA or is it something that you're waiting to mature? How do you think about that? And even more broadly as a follow-on, I think it's an interesting question. How do you think about waiting for things versus building them given the pace of innovation in this space? I would say we are not quite ready to talk about it. But what I will say, though, that we now have the ability to pick and choose whether to use an open source model like Lama.
42:43or build our own. We have that ability now. And then we will continuously experiment what's available outside versus what we can build ourselves and make those trade-offs decisions. You talked earlier about agents and agentic workflows. Can you talk a little bit about about where you're seeing opportunities to apply that pattern? We are super excited about agent-take and multi-agent-take workflows. You know, I truly believe, you know, it is probably the next frontier where you have like, you know, one or more agents, you know, responsible for understanding intent or planning or, you know, executing a tool, you know, taking some action.
43:35And it's an active area of research for us. We are also building our platform infrastructure to be able to orchestrate and coordinate these agents. So as you know, Langchain, Langgraph, we have incorporated them. We have brought them into our platform, but we are not stopping there. We are also researching the agentic framework itself to see a lot of the gaps in the current agentic frameworks. We are going to address them in our platform. But there is huge potential if you look at how the back office works in most banking, how we execute a lot of contract document review or document understanding tasks, how we develop software.
44:42You can think of agentic frameworks actually doing different parts of the software lifecycle. So I think of all the LLM patterns, like RAG has been very popular. I think the next frontier for us is truly going the agentic way and then try to see how agents can solve different tasks within the enterprise, like routine tasks, mundane tasks. And there are hundreds and thousands of them. So we are very excited about it. You mentioned LangChain and LangGraph. Can you talk a little bit about why in that area you use existing off-the-shelf tech versus build-your-own agentic framework? We try to align with open source frameworks as much as possible.
45:37If you think of Kubeflow, Kubernetes, open telemetry for the base for our monitoring systems. Because as you know, with open source, you have the power of the community to harness. and we also put our engineers in those open source projects. So they contribute back to the community as well. So we felt like LangChain, LangGraph are projects where we can contribute as well as we can also harness the power of the community. But we are not stopping there. We have extended some of these things into our own kind of framework within the platform wherever we see issues with scalability or being able to orchestrate many agents, for example.
46:32So we do modify them to suit our scale and complexity. And not to mention a lot of regulatory controls that you need to have as you build into the platform. Do you have any early results with these agentic-style workflows that you can talk about? still very experimental for us. You know, if we talk in six months, I'll probably have, you know, like a solid something to tell you. But, you know, I would say maybe in the software side, we have seen double digit, you know, efficiency gains, for example. With CodeGen, you know, the goal for us is not to automate everything, but really the mundane task, where you are figuring out configuration files, routine code maintenance and development.
47:28A lot of those with Kubernetes. Yeah. And as you know, those YAML files tend to be very large and somebody has to look through. As you know, those things can be automated even with agentic code generation agents, for example. But then leave the really interesting part of software engineering, like system design, those are like end-to-end test harnesses for a project. Things that require more experience and expertise, those we think would be the creative part of software development and that our engineers will enjoy doing more. creating code in stuff files and just filling up the stuff code. Even basic unit test generation, a lot of engineers, I'm sure they would welcome an agent doing those kinds of things.
48:36We see that in our research with our software engineers. With your current approach to CodeGen, are Are developers choosing their own tools or do you have a standard stack for code gen for developers there? We have a standard stack. So, you know, one thing we have invested a lot in is like the governance of how we are going to deploy AI and gen AI in the company. so there is a governance process where we look at every aspect of any kind of third party tool that we are going to use including how do you have guardrails do you send data to a server what kind of evaluation framework we are going to have in place before we bring in that tool is the code generated, will be deployed in production, and what's the CICD process for that?
49:45So there are a lot of things that we consider before we decide what tool to bring in. So we have standardized around a couple of frameworks. And are they homegrown or off-the-shelf IDE plugins or services? or? It's a combination of both. I think, you know, most large companies will be probably like ours. Like, you know, they will have one or two, you know, off the shelf. And then for certain specific tasks, you know, you may fine tune your model because with your own data, you know, so that gives you much better flexibility and control, you know, what the accuracy and relevancy for that particular task.
50:30And just to be clear, I'm assuming that if you were able to say what you're using, you would have said it.
50:39Is that the case? It is. Again, in six months, maybe I'll be able to come back and tell you more about these specific tools. But frankly, we are also learning at this stage. you know it's not like you know it's deployed to a large portion of our software engineers but a lot of these tools are you know we it's still you know like the the it's not clear who is going to be the clear winner in this space you know so we are constantly you know evaluating different of you know off the shelf tools to see which tasks can be better done with the tool or is it a general purpose solution. So that's why I'm a little bit hesitant to declare a clear winner and say we are using that tool.
51:28JOSHUA MACHTENBAUM - We've already talked about a couple of things. Like in six months, we'll be able to talk about X and Y. But more broadly, how do you see Gen AI evolving at Capital One? KUNAL KUMARANI - I would maybe talk about both the areas that we are exploring and where we think we'll see some breakthroughs. Plus, I think we should also talk a little bit about how important talent is for us, again, going into next year and beyond. So I think agentic workflows, as I said, would be very, very important for us. As I said, so much of back office and software engineering, et cetera, document contract verification, document verification could be automated.
52:21through that. You know, agent tech workflows is also a place where business process and domain expertise actually meets technology, right? So it's an interesting intersection where the traditional definition of talent, you know, is no longer kind of the same definition here. Like you need somebody who understands Gen AI, who understands the, you know, the business process, and they also know probably enough of prompt engineering. They know enough of guardrails to actually build that solution together, right? So it's not the standard software engineering skill that you need. You need a few other things to be able to really build those agentic workflows.
53:11And what does that look like? Product management, analysts, some combination of all of the above? Do you see a new role emerging, I guess? So we have created two roles actually at the company in this year in anticipation for AI and kind of being at the forefront of Gen AI. We have created an AI engineer role in addition to our machine learning engineer role that we started about three years ago. And then we also created an AI researcher role, applied AI researcher, I think we call it, who are experienced in building LLMs and retraining, fine-tuning, agentic workflows, et cetera. But then we are not looking for just pure engineering experience.
54:11We are looking for people who can adapt themselves with different problems that they will be solving. So that agility and adaptability, I think, is going to be important for all of us as we work on this AI and Gen AI, because a lot of things will be transformed. A lot of our processes, systems can be reimagined. So we need people who are going to reimagine the future instead of just executing a task. So that's one area. As I said, we are also seeing a lot of reduction in inference costs. As you know, one of the reports, I think from A16Z, like Anderson Horowitz, that showed the 100x drop in inference costs, for example, from two years ago.
55:09So, you know, and we see that also in our own work, that inference cost, you know, cost per token, etc. Those are dropping, you know, rapidly. I think that will encourage a lot more, you know, companies and even a lot more open source developers to be more creative. Right now, it's still kind of expensive to be a freelance innovator and kind of doing research on your own. So those things will also help in the coming years. And then I think we will also see more open source models that are incorporating O1-style reasoning. you know I'm already seeing a couple you know but I think that will become more and more you know common and so that will help us to solve you know much more complex problems you know in even in multimodal you know kind of problems for example in the enterprise Any other thoughts looking forward?
56:16I'm very excited and I'm lucky to work for a company you know where I get to play with all of this new technology. I'm happy. I feel really lucky. It's a fun time. It's a fun time. I have been doing this for the last 20 years probably, but never ever in my career, I have seen a phase where every week there is something new coming up. It's actually a challenge if you think, Sam, like it's so hard to stay, you know, current with last week's news, right? So, you know, Twitter and, you know, some curated emails are my friend these days. Like, you know, it's so much news coming every week. Well, Abhijit, thanks so much for jumping on and taking the time to share a bit about what you're working on and how you're thinking about Gen AI platforms at Capital One.
57:14Thank you very much. I really enjoyed our conversation.
From the publisher
Today, we're joined by Abhijit Bose, head of enterprise AI and ML platforms at Capital One to discuss the evolution of the company’s approach and insights on Generative AI and platform best practices. In this episode, we dig into the company’s platform-centric approach to AI, and how they’ve been evolving their existing MLOps and data platforms to support the new challenges and opportunities presented by generative AI workloads and AI agents. We explore their use of cloud-based infrastructure—in this case on AWS—to provide a foundation upon which they then layer open-source and proprietary services and tools. We cover their use of Llama 3 and open-weight models, their approach to fine-tuning, their observability tooling for Gen AI applications, their use of inference optimization techniques like quantization, and more. Finally, Abhijit shares the future of agentic workflows in the enterprise, the application of OpenAI o1-style reasoning in models, and the new roles and skillsets required in the evolving GenAI landscape.
The complete show notes for this episode can be found at https://twimlai.com/go/714.




