AI Access and Inclusivity as a Technical Challenge with Prem Natarajan - #658

4 Dec 2023 · 42 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: AI Access and Inclusivity as a Technical Challenge with Prem Natarajan - #658

Overview In this episode of The TWIML AI Podcast, host Sam Charrington speaks with Prem Natarajan, Chief Scientist and Head of Enterprise AI at Capital One. The discussion revolves around the technical challenges of AI access and inclusivity, as well as the multidisciplinary approaches employed at Capital One to tackle these complexities. Key topics include bias, class imbalance in data, the integration of research initiatives, and foundation models for financial data curation.

Key Themes

  1. Background of Prem Natarajan
  2. Early Interest in Language: Prem shares his upbringing in Pune, India, and how his exposure to multiple languages ignited his interest in language processing and eventually led him to AI.
  3. Academic Journey: He discusses his education in India and the US, where he worked on significant projects related to speech recognition and natural language understanding.
  1. Research Priorities at Capital One
  2. Graph-Based Machine Learning: Exploring how structured data (like graphs) can be integrated with advanced AI techniques.
  3. Understanding Model Inference: Investigating how deeper insights into model inference processes can enhance performance.
  4. Inclusivity and Access: Focusing on designing AI systems that cater to a diverse range of users.
  1. Balancing Research and Practical Applications
  2. Curiosity and Real-World Needs: Emphasizing the importance of aligning research with practical applications to provide tangible benefits to customers.
  3. Partnerships with Academia: Maintaining collaborations with academic institutions to tackle complex AI challenges that require diverse expertise.
  1. Tackling Bias and Class Imbalance
  2. Computational Implications of Inclusivity: Exploring how biases in training data affect AI outcomes and the importance of using representative datasets.
  3. Class Imbalance Solutions: Discussing various techniques, such as data augmentation and label smoothing, to address class imbalances in training data.
  1. Foundation Models and Data Quality
  2. Role of Data Curation: Highlighting the critical nature of data quality in training foundation models and the importance of privacy-preserving techniques like federated learning.
  1. Future Directions in AI
  2. Building a Strong AI Team: Capital One is focused on attracting top talent in AI and creating an environment conducive to innovative research and applications.
  3. Leveraging Generative AI: Exploring the potential of generative AI and LLMs (Large Language Models) across various use cases beyond just generative applications.
  1. Investments in AI Research
  2. Mission-Inspired Research: Prem describes how research at Capital One is driven by mission-oriented goals that aim to solve problems for customers, which in turn leads to broader applications and innovations.
  3. Creating Broadly Applicable Solutions: Solutions developed for specific problems can often have wider applicability across different contexts.

Key Takeaways

  • The intersection of AI, access, and inclusivity is a major focus in developing AI systems that serve diverse user groups.
  • Tackling biases and ensuring data quality are essential to improve AI performance in practical applications, such as fraud detection.
  • Collaborative partnerships with academic institutions enhance the capability to address complex AI challenges.
  • Investing in AI research aligns with broader business objectives and drives innovation that can benefit the entire financial service landscape.

Conclusion The discussion with Prem Natarajan provides valuable insights into the ongoing efforts and challenges in making AI more accessible and inclusive while maintaining a strong focus on practical applications in the banking sector. Such initiatives not only aim to enhance customer experience but also strive to advance the field of AI as a whole.

For more detailed show notes, visit [TWIML AI Podcast Episode #658](https://twimlai.com/go/658).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:09All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Prem Natarajan. Prem is the Chief Scientist and Head of Enterprise AI at Capital One. Prem, welcome to the podcast. Thanks, Sam. Excited to be here. I'm looking forward to our conversation. We've got a ton of interesting topics to cover, including how you've crafted an AI research program at Capital One in collaboration with academic partners, Some of your recent work on foundation models, graph machine learning, and much more. But before we dive into those topics, I'd love to have you share a little bit about your background and how you came to work in AI.

0:53So what's a kid who grew up in Pune a few decades ago playing cricket, doing with AI, right? It's a good question to ask. A good place to start. I grew up, you know, my family is ethnically Tamil. So at home, we spoke Tamil. I grew up in the state of Maharashtra, where the local language is Marathi. We had a school system where we had something called the three-language formula. The medium of instruction was English. Hindi was one of the languages that we were kind of required to learn and which I ended up loving. And Marathi was the local language, and Tamil, of course, was the lingua franca at home.

1:32As it turns out, while English and Hindi and Marathi might roughly fall in the category of Indo-European languages, the same language family. Tamil does not. So growing up, many of us who were kind of with similar background as me, we kind of make observations about language structure just as kids. You know, I wonder why the order of these verbs and nouns is different. We may not even have known what verbs and nouns are, but we kind of knew these words signified these things. And somehow, I think deep in there, There was kind of a native interest in, or a deep interest in kind of language developed, especially from just the comparative use of language in daily life.

2:11And that eventually led to interest in, I guess, language processing technologies, which very naturally then led to a broader interest in AI and machine learning, et cetera. That led to, anyway, undergraduate education in India, graduate school here in the US, was blessed to work on a series of DARPA and other similar DOD-sponsored efforts, which were all breaking new ground in speech recognition, natural language understanding at the time. Just a phenomenal time to learn from some of the champions of the trade at the time. Community was small, but very fertile. So many research, so many unanswered questions to dive into, and so many people had thought about it for so long to talk with and learn from.

2:58And then it was great because I had my period of education in that world. And all of a sudden, it became of mass market interest. And so by the time it became of mass market interest, I was ready for it. So that's what brought me here. Awesome. It's always great to be in the right place at the right time. Indeed. Yeah. I mentioned looking forward to our conversation and being excited about it. That goes pretty far back. I mentioned this to you previously before Capital One, we're at Amazon working on Alexa and the speech aspect of that, of Alexa is obviously a key element of it. And I tried for quite a while to get an interview with you and was never able to make all of the pieces fit together.

3:45And then we met not too long ago at Capital One dinner at KDD that I emceed and really enjoyed that conversation. I remember some really interesting topics that came up there from my perspective. We had a really interesting exchange about reasoning, about agent-based systems, about achieving fairness. Great group of people that you and the team there assembled. But it just led me to look forward to kind of digging in a little bit deeper with you. So thanks once again. I think a really interesting place to maybe start is to just have you riff on what's top of mind for you. What are some of the research priorities around AI that are most interesting for you nowadays at Capital One?

4:30Yeah, you touched on some of those already in your opening, Sam, but I'll just recount a few. There's just tremendous progress as we were talking about at the dinner that you just mentioned, which, you know, thank you. I greatly enjoyed the perspective you brought to, the curiosity and the question that led to like kind of folks talking about it in a different way than they might otherwise. Part of it, though, is that a lot of progress in just the general area of deep learning and transformer-based learning, etc. But some of the data that we're interested in has structure to it, like graph kind of structures.

5:02And you mentioned this in your opening about graph-based ML, etc. So one of the areas that we are very interested in is how do we combine the structure that is inherent in much of our data or some of our data with the power of these new techniques. And I think chatted with one of the people in my organization, Bayan, a while ago, and he and his team have been working on this area and some others on this topic. So that's one. A second area is how do we kind of understand the inference processes of these larger models? As always, people started looking at the smallest possible unit of these models, the neurons, as kind of the basis to try to understand inference processes.

5:50But there's been some interesting new work where groups of neurons are discovered that might kind of be doing things together. So those are all some interesting directions for us. From a more philosophical perspective, but one that makes it into our computational work, things like how do we make all of this more inclusive? How do we increase access to it across the full spectrum of people who come to us for our services? And there, we are looking into the intersection of core AI ML technology with UX affordances, like what kind of design approaches help amplify access or democratize availability of these things to folks.

6:32So our interest is very broad. It tends to be very specific in the near term as we're trying to build things and launch them into production. But all of those near-term specificity is grounded in kind of this longer arc of very broad set of interests that we have. And so given that level of breadth and your academic background as a kind of pure researcher, how do you balance that with, I would think, more tactical needs of a bank that's trying to make customers happy and kind of keep the lights on and avoid fraud and that kind of thing? How do you think about prioritization? So the good news here is from a curiosity perspective, all of us want to work on things that we're solving new problems, we're advancing the state of the art in some ways.

7:22And the very happy state that I find ourselves in, I don't mean just at Capital One, but in general, this community that's applying these new AI ML techniques, is that there is so much potential, yet there is so much distance to be traveled in taking these things and applying them to specific use cases. So these are all invention rich areas. So there is very little right now that has simply turned the crank, right? Pretty much anything we do, especially if you come at it from the perspective of, I want to take a responsible, thoughtful approach to applying AI to these problems. So it's not just, let me throw it at something and see how it works.

8:01And oh, if it works well enough based on things I can measure, I use it. We're actually saying, how do I do this in the context of how I want to do this? And in framing the problem that way, we actually discover new invention challenges. So the daily course of our lives, even though there might be a capability that we want to launch in six months or eight months or nine months, a set of capabilities, getting to those capabilities is not merely putting together available things. There's a fair amount of invention. So in some sense, the happy place that I was talking about is that the field itself, there's so much invention to be done in every step of the way, that is just a joyful journey of discovery and curiosity right now.

8:46On the academic side, I'll say there are aspects of it that I carry with me all the time. And one of the ways I kind of blend it with my work, whether it was at my previous place, you mentioned Amazon and Alexa, and now at Capital One, is that some of the hardest problems in AI, the ones that are the longer arc problems, they're beyond, I believe, the capacity of any single organization to solve, right? So therefore, multidisciplinary, multi-stakeholder, multi-sector partnerships are key. And in the kind of society we live in, these sectors are academia, industry, nonprofits, government. And so we have taken an approach of engaging with academia broadly.

9:31We have partnerships with NYU, CMU, South Maryland, many, many more. I'm just mentioning a few. And we are continuously broadening that. For example, even on the data side, beyond AI, we are involved in the MIT Future of Data Consortium, both contributing to the advancement of thinking around data and learning from others. So those partnerships help us shine a light on some of the hardest problems and then work with academic and other partners to solve them. I'd love to dig into the way you tackle some of these longer arc problems from a research and technical perspective. You mentioned that inclusivity and accessibility is a thread that kind of weaves across the various research efforts.

10:21How do you think about that as a set of technical challenges? And how do you kind of dig into that from a research perspective? Yeah. So most of these problems, when we talk about inclusivity or access as red room, we kind of think about what are the computational implications of those things, right? And maybe I'll illustrate it with a very old example from speech recognition. One of the first things when I went beyond research to actually deploying stuff in the field was actually speech recognition for call centers, right? I talk about it more openly now, But there was a time when I tried to hide the fact that I was involved in IVR systems that was annoying my friends, right?

11:02And so now I feel fine talking about it. So with speech recognition and my team was building some of these things that were going to some of the biggest call centers and some of the biggest telephone companies in North America. And as we were testing it, you know, I'd find me and so many others with similar accents like mine were part of building these technologies and, you know, deploying them. I found that they didn't work uniformly well for me. And so like certain words, I would have to repeat them multiple times to get it right. And it was kind of hard for me to remember exactly what I said, which way that made it work the last time, but that shouldn't be the burden on the user anyway, right?

11:42The purpose of AI in my mind should be to transfer the cognitive burden from the user to the system. To me, that's the noblest of missions, right? Like you transfer the cognitive burden. All of us have so many cognitive burdens every day. But if we can transfer a little bit of that cognitive burden over to our system, we're doing good. So that started some thinking. And then we said, you know what, if you simply collect data that is representative of whatever set of users you want to serve, then that might introduce maybe a majoritarian bias in the data, where on average, it performs well, because if your test sets are similarly this thing.

12:18So one of the old truths in ML or what was acknowledged as a truth was in order to truly test the ML system, test data has to be similar in distribution to the training data. Well, it's not until you start really deploying them in the field that you start doing. No, no, no. The test data has to be representative of all the different types of people who will end up using the system. And so that's really what we want to test for. So it comes down to things like that. How do you take experiences that each one of us has, translate that into computational considerations, and introduce that thinking at every stage of the process of designing these models, of testing these systems, and then taking those test results back to the algorithms and saying, what do I need to do?

13:03Oh, the math here is such that it embodies the distributions. Is there a way for me to break that mathematical learning into factors that then focus on different parts of the distribution. So it's just every part of the thing has to be involved. All of it, though, starts from having the right set of equities represented at the table. So another aspect of our academic partnerships is that because we do believe that we increase the diversity of thought around these problems by engaging more broadly. And so how does that translate into some examples of specific research projects? Yeah, so I'll give you a very specific name that sounds super technical, right?

13:44But it's actually, like I said, like, you know, all the challenges in translating these into computational things. So if you think about what I talked on the data side, right, it's really about the fact that from a mathematical perspective, you say this data is in many different categories, right? And just by happenstance, there is imbalance in the amount of data that falls into these different categories. And how do I make sure that doesn't happen? So one of our recent publications, actually upcoming ones, is about how to simplify neural network training in the face of class imbalances on data, right?

14:17This is actually the name of the paper that some of our folks are working on. And so it's things like that that then say, oh, I know how to deal with class imbalance in training data while still optimizing performance for everyone. And then those training techniques are then used in production, in training of production models, and then the performance is improved across the board. Now dealing with class imbalance is kind of an age-old problem in machine learning. Can you give us a flavor of how this paper approaches that problem? Yeah, so many different techniques. So a lot of this, Sam, what happens is as some of these new techniques come in, we kind of think about what were some of the older learnings that we had and how do we adapt some of those approaches mathematically to newer ones.

15:05So here, broadly speaking, we'll talk about data augmentation or label smoothing or things like that. And so now we're getting even deeper, right? So we start with the end thing that we see, hey, different users are seeing different performance. We want to make that the same. Oh, we trace it back to class imbalance and we say, let's address that problem. And then we say, okay, what are the specific ways? What is the way in which I smooth the distribution? Oh, one of the ways is if all of these different bins have different amounts of data in them, Can I do some interesting math to make it look like they're all the same amount of data?

15:36Oh, wait, data augmentation, is there a way in which I can do it in a meaningful way to augment data that is in the original training set? Are there ways in which I can smooth the distribution of labels? So things like that. So many such techniques. One of the things I find, by the way, exciting about the new deep learning stuff is that you can design objective functions that are very directly focused on addressing either defects that you might think about or optimizing for outcomes. So one of my PhD students at USC did some work on what we'll call adversarial invariance, right? And then later extended it to unsupervised adversarial invariance in training of these models, where you can actually design the objective functions that you're using for learning to make sure that your predictions on different classes are equally accurate or not disproportionately different in different ones.

16:33And so there, the kind of data balancing happens as part of the learning itself, where you don't have to explicitly do it. So whereas historically, we used to have to take this very explicit way of data augmentation, et cetera, and that still helps. But we can now do more sophisticated things like actually engineer the math a little bit in the optimization itself to achieve the objectives that we want. Yeah, I love that example as a very clean articulation of kind of the pushback on, well, bias in machine learning is a data problem and you just have to collect better data and it has nothing to do with machine learning as a technical field or a set of techniques.

17:12Well, we have these things called objective functions and we can architect those objective functions to address some of the biases that are inherent in the data, if we're thinking about that. That's right. Actually, Sam, if I can just digress a little bit, but on that specific topic, you referenced the fact that we met at this dinner at KDD. One of the faculty members we had at our table, you might remember, was Kai Wei Chang from UCLA. I've known Kai Wei for several years now. And in fact, one of my PhD, current PhD students, Yi Hong-Soo, is kind of collaborating with him as well as part of his PhD studies.

17:49But Kaiwe's pioneering contributions in studying bias in natural language generation was that while for the longest time people said we're only as good as our data, he kind of broke the myth and said, you know what, we're not even as good as the data and showed that many of the learning algorithms were amplifying the biases that were in data. So they were not just replicating the biases, but they were amplifying the biases. And I remember when I read that paper and I said, my God, what a brilliantly simple demonstration. The experiments that he did were just so elegant, like in showing how they amplify the biases in visual ways, like he was generating labels for collections of images.

18:32And even though the image might have somebody of a particular gender in a particular context, because the model has a bias of in that context, in that environmental context, I expect the other gender or appearance of that gender. So it would generate labels and that the distribution of that labels in the generated set was very much more skewed than it was in the training data. It was something that everybody could relate to. I remember as I was reading that paper, I said, gee, if they can amplify the biases, maybe they can reduce it too, right? That's what you latched onto with this whole objective function thing.

19:07But yeah, Kai Wei at our table was actually talking about this stuff. Interesting. And so this particular research effort, the simplifying neural network training under class imbalance, does that align specifically to a core Capital One use case, like fraud, for example, or something else? Or I think I'm just pulling on a curiosity, Like how closely aligned are the individual research projects with what the bank's doing? It's a fair question. I'll give you a couple of thoughts on that. One is, in the case of fraud, we definitely want to reduce fraud across the board. And there's huge skews from a class imbalance perspective.

19:50There's a huge skews in this capital. But without tying it to any specific use case, I'll just say almost everything that we want to use it for, right? We want to address this challenge, right? And it's not just us. almost everything that everyone in the world wants to use AI ML for, these are fundamental challenges. But what happens in our case is we work back from specific things we're building, and we use it for a very broad set of applications. Fraud is the one that we highlight the most because it's just so broadly beneficial to attack that as a phenomenon. And it's kind of interesting too, Sam, because in addressing fraud, it's not super useful to find fraud super accurately four days after it happened.

20:34Kind of want to detect it in the moment. There's a time window. There's a time window. So we're almost satisfied if it detects it in the moment and the transaction is paused in some way, right? And so in those cases, the same approach of class imbalances can also come with other considerations. How do I also make sure that I can do it in real time and balance it and stuff like that. But without tying it necessarily to just that use case, I'll say that this is a fundamental problem and this won't be our last publication on that topic. I'd like to think of this as just one more investigation of how to improve things.

21:10But that's kind of one of the fundamental journeys, I think, in machine learning. There's just no way to get around class imbalance. Like whatever I do, I'm still going to speak the way I speak and I can't convince other people to speak like me. And so that kind of thing, yeah. Yeah. Pulling on the fraud thread a little bit more, one of the things that classically characterizes academic approaches to problem is that it's very, you simplify it quite a lot and you say, hey, this is the specific thing that I'm looking at. I'm not even sure that this is relevant to what I'm trying to get at. What I'm trying to get at is I'm thinking about all of the conversations that I've had with Capital One and others about solving this important fraud problem.

21:58And there's some set of research that is going at it as, hey, let's think about fraud as this graph problem because it has inherently graphical qualities. And there's, you've got this class imbalance, you know, let's approach it from a class imbalance problem and try to address that. I'm really curious, like when you have many disparate research efforts trying to tackle this problem, like how do you integrate all of those results into tangible system progress? Is that an interesting challenge in and of itself? Yeah, I mean, I think there are a couple of interesting things in there, which I think I understand what you're getting at.

22:36Like the problem may be the same, but there are so many different methodologies or approaches at getting at that problem. how do you bring them together in a way that they're at least additive, if not super additive? Because ideally, what you'd like to do is if I'm able to reduce fraud with one technique by 5 % and by another technique by 3%, at a minimum, I would like the combination of the two to at least be greater than 5%. But ideally, I would like it to be greater than even 8%. They come together and they do magic, right? So there are two kind of ways in which one can think about it. One is that you have an overall framework of how you are addressing fraud, some machine learning framework, et cetera.

23:15And there are specific subcomponents of that framework. And here we can tie back to our recent discussion. One might be, what do I do explicitly with the data that I'm using to train the models to improve that performance? Another is, what do I do with the objective functions to improve the performance? Yet another might be, what do I do to network configurations to improve that performance? A third might be, what kind of external knowledge sources can I bring in to improve the performance? Now, notice in the way I crafted this, maybe there's a little bit of a slate of mind, if not slate of hand in it, where I explicitly listed things that are kind of complementary, right?

23:56So if there were like four different ways of improving the data, the possibility that all four of them will come together and give an even greater lift is less than if I have an effort in improving the data stuff and then at the objective function. So partly, it's in how we construct the research agenda. So at the core of what you're asking, if I take the more constructive or proactive approach to that thing, you're saying, how do you make sure you're investing in approaches that can actually be brought together to give you a greater lift, right? So I don't have to wait at the other end of the road for people to do certain things and then say, oh, how do I bring these together?

24:36One can actually set up the problem to say, hey, there are these different things we could explore. Can we ask different folks to explore these different things? If we look at it from this framing, Sam, then you can also see, gee, maybe the research into how to bring in external knowledge sources into the prediction or inference process is a bigger, longer-term research agenda, right? Oh, maybe that's where I should do partnership with academia because there are people who can focus on this longer term, you know, and we can bring in the constraints of how it needs to work in the real world. Updating the objective functions, et cetera, maybe that is so tied in the specific use that maybe we are best positioned to work on it internally.

25:21I would say I approach that same question in advance of having to answer that question, right? So, you know, I was going to say, I know what I want to do at the end is I want to bring together the outputs of these four or five different research efforts. How do I set it up in advance? So that's one part of it. The other part of it, of course, is that in order for you to say, I want to choose these different techniques, you have to be able to compare them. There's only one real way we know how to compare things. You measure their performance on the same task on the same test set. And so it's also important for us that we define the research task or the specific task to be accomplished clearly so that everybody's working on the same objective overall, and that we also have the same test set, which is, I guess, the colloquial way of saying is it's a level playing field.

26:13Same tasks, same test set. What are these different Thanksgivingly. And do you find that the prevailing kind of academic test sets generally correlate strongly with your requirements or have you invested significantly in coming up with your own test sets and benchmarks that better measure the real world use cases that you're trying to model? Yeah. Correlate well is a contextual question, right? Is that do they correlate it well enough for you to be able to use some of these open source test sets. And I wouldn't say necessarily just academic, but just open source benchmark tests. Let's say there are 20 people working on it.

26:54Maybe they give me enough insight to say these are the four or five different techniques that seem most promising in terms of how they might apply to my research challenges. And so then you can look at those techniques to bring them in and then either optimize them or test them on your task, your data set, et cetera. But I think what you're really getting at, Sam, and I agree with you if this is what you're really getting at, is one of the ways in which we can accelerate the impact of much of the work happening across is by working with academia to define test sets and training sets that allow students and others who are working on these problems to focus on exactly the right problems or exactly the right dimensions of those problems.

27:35And one of the things we do want to do with our academic partnerships is promote a culture of where we can help define those kinds of datasets. And there too, actually, maybe if you look at generative AI, we kind of think of it in the context of generating responses or useful predictions, et cetera. But you can also look at it from the perspective of being able to generate useful datasets that are synthetic datasets that are not representative of real world phenomena, but are not actually grounded in any specific real world entities. So they make them that much more possible to be released openly or things like that.

Read the full transcript

28:14Awesome. Interesting. I did want to touch on some of the work that you're doing with MIT around foundation models for financial data. Can you talk a little bit about that work, its focus, how it's evolved? The thing I'll say is on the data side, a lot of the work, the focus is on data curation, the partnership with MIT. One of the things that has come about in a lot of the research around foundation models more and more is the recognition that it is not just the volume of data, but the quality of the data that goes into training these models that is a determinant of the performance of those models.

28:52So from that perspective, data curation, data quality, controlling, or assuring a certain level of data quality, those are all important things. And it turns out creative people can figure out how to use AI for the purposes of curating data that is then later used to improve the AI, things like that. So kind of almost lifting yourself up by your own shoelaces kind of sense, like bootstrapping. But the other aspect of it also in our partnership there is also what you might think of as privacy preserving learning or federated learning. How can we actually federate learning? Like if certain data sets are in different places, right?

29:32Instead of moving the data sets around, can we actually do the learning where the data sets are, bring back the learnings, update the model, then do the next iteration of learning? So, you know, this is really federated learning. So that's another area that we're working on. When you talk about data curation and even more broadly, the training data for this foundation model that you're working towards, what is that training data? What is the model modeling? That depends. So remember we were talking about like, how do you bring together different techniques? And I said, there's a research task.

30:05And so we have many different tasks and these tasks are tied to use cases. So you mentioned one, which is fraud. And so what we're trying to detect might be, is this particular transaction fraudulent, right? And so now you have to see what kind of information globally is useful for that. What kind of network configuration, when I say network, I mean neural network configuration might best give you those predictions. And if the scores come out, how do I normalize these scores so that they're comparable across different events or different transactions, et cetera. So all of these are important. And for all of these, the quality of data is critical because if, for example, for certain merchants or certain classes of entities, systematically certain attributes of data are missing, then your performance on those things is less reliable than in other contexts where the data is more complete.

31:04Or if the data is noisy, if the time registrations, For example, I'm just making up an attribute. Time registrations are off simply because of some noise somewhere in the processing. All of those might affect the way in which the system learns and introduce variability into the scores that it produces. So I'm kind of making up a generic notion of quality without tying it to anything specific in the data we deal with, just to stay away from talking about specific aspects of the data, like grounded in financial stuff. But these are the things. So the data has attributes, there are events, there are attributes around those events, there are entities involved in these events.

31:44And if any of this data is either missing or even things like my name is Prem or full name is Prem Kumar, and if different people spell that differently, instead of all of that transaction data being able to be tied to me or the one entity, maybe it gets fragmented across different things and that reduces the effectiveness of the data. So data quality issues come in a variety of ways. And so we find different ways to address those. I'd say entity resolution, like saying that these three different entities are either the same entity or different entities, is an important problem regardless of the use case.

32:24And one of the things that we want to improve performance on by curating the data is entity resolution. And I was giving you a very simple-minded example of like, oh, you know, my name is misspelled or spelled in a few different ways. And then I want to find out that it's the same entity. But even if my name is misspelled slightly differently, if my address for both of the names is the same, that gives me some more confidence that this is the same person. And then I might be able to normalize the two things and improve the quality of the data. Now, instead of having two separate people, it's always the same entity, right?

33:00Whether it's used for fraud, whether it's used for something else. And those kinds of spelling mistakes can happen with merchants also, especially if you have a lot of different merchants. Different people may spell them differently, or different databases have them spelled differently, but you don't want to treat them as different entities if they're the same entity. So I'd say more like, rather than take a particular problem and try to chain together a very point solution, part of our academic collaborations focus on what are some cross-cutting problems like entity resolution or the quality of data that goes into improving our entity resolution or anomaly detection and how do we improve the data to improve performance on those problems because when you improve performance on those things they can lift a lot of different boats so you were kind of putting those together in a particular configuration maybe maybe not i actually don't even know if anybody's using it simply because that's a level of detail or granularity that I'm not usually engaged in.

34:02But the fundamental problems that we are trying to solve through these things should help many, many use cases. Yeah. And am I correct that there's an LLM angle to that data curation approach? I have to actually check. I don't know, but I'll tell you this. I thought I remember reading that somewhere. Yeah, yeah. I'll put it more generally. There's an LLM angle to almost everything nowadays. because the thing is, for example, you can easily see how you can do it. Like if you can use an LLM to generate reasonable variations of entities, you can use that to train a system to say, if I see these variations, bring them together in one.

34:37So it's just like in some sense, the use of the LLM is just so natural nowadays in many creative ways, yeah. Great, so we've talked through several examples of your work. I think you mentioned that you can kind of sprinkle LLMs on everything. I'm assuming that that's part of direction that you're headed. I'd love to have you riff a little bit on where you see this all going. What, you know, with the popularity of LLMs and generative AI and the partnerships that you have in place, like how do you approach the future? Yeah, a couple of different thoughts on that, Sam. One, while LLMs and generative AI in general are kind of the headline terms under which we discuss a lot of these things, the underlying technology of deep learning, especially transformer based deep learning, is something that we think has a lot of applicability even beyond generative applications.

35:32But if we believe that, then we have to develop our own expertise in terms of how do we take that technology and apply it to problems of interest to us, sometimes maybe uniquely of interest to us, but at a minimum, kind of specifically of interest in the financial space. And so one of the things we're doing and super excited about is we're building a world-class AI team, scientists, machine learning engineers, product managers, technical program managers, UX designers. So we're kind of going back to this notion of we need multiple perspectives in order to build the best systems that we can. But we also need top talent to kind of solve the hardest problems and the hardest invention problems in these areas.

36:19So by way of partnership with these universities, for example, I think we become a place where top talent wants to come and do research, addressing problems of interest to us because they can collaborate with other top talent in the world, often resident in academia, and also bring in fresh new PhD students who are working on some of the more challenging problems for the future, et cetera. So a big part of it is building the talent base. Second part of it is building kind of the infrastructure elements so people can do this work efficiently. What are the platforms that we have for experimentation?

36:53What are the platforms that we have for modeling? What are the platforms that we have for testing? Because all of these accelerate the work of lots of different folks who are joining us. And then the diversity of use cases, of course, will attract creative, smart product managers, technical program managers across the board to come and work with these scientists and engineers in helping to advance the application of these techniques to newer use cases, to new customer facing applications, etc. I had an interesting recent interview with Miriam from the team there at Capital One. And one of the topics that came up in that conversation was justifying investments in the work that her team was doing, machine learning engineering in that case.

37:36And I'm curious to ask you the same question, like how the bank or how do you think about the investments in research for a bank? What is the rationale and how do you think about the importance of it? I think historically, if you look at research, even in the context of NSF vis-a-vis DARPA, one of the paradigms for research is curiosity-driven research, where you're discovering new things for the sake of that discovery. That's very properly an academic mission and kind of the mission that NSF will take on itself, the National Science Foundation. For other modality of research, DARPA and maybe more of industry adopts, is what I'll say mission-inspired research, right?

38:18What is the mission I'm on? What are the services I'm trying to provide to certain sets of people, certain sets of customers? Where is today's technology and what is the gap between today's technology and what we need in order to serve these customers in specific ways that we want to serve them, right? And that turns out to be a very powerful framework for research. that is designed to deliver impact. So in that sense, I would say we kind of pick the problems that solving which would drive impact, drive benefits to our customers, advance our business, et cetera. And then we focus our attention on that set of problems.

39:00Now, it turns out if those sets of problems are hard enough, then the solutions you come up with actually end up being generally applicable to lots of different things, right? This is why we see DARPA's impact is far beyond what it has, the specific programs that it has done, like it's transformed society, the world, et cetera, in many ways. And I think each one of us in our own ways is doing that, in that when we focus on the mission at hand and a set of users within that mission, and then look at the gaps in technology and use that to frame the specific invention problems we need to solve, then not only are you guaranteed that if you solve that problem, you will have impact because that's the whole framing, right?

39:44It also turns out, I mean, and this is not necessarily intuitively true, but just experience shows us, it also turns out that solutions to those problems are very broadly useful to other things. And I think there is something about, you know, I think, I don't know, maybe it was Mahatma Gandhi who said something like, The solutions to the problems of humanity are to be found by living amongst people themselves and not in the remote peaks of some mountains. And so I think by rooting ourselves in like here and now problems that are big, they're actually examples of other problems, related problems.

40:20And so we're sampling from that set. And therefore, we end up creating solutions that have very broad impact. Awesome. Well, Prem, once again, it's been a pleasure chatting with you about some of the ways you're thinking about research at Capital One. Thanks so much for taking the time to share a bit with us. Yeah, Sam, thanks again. I think you have a way of making conversations so interesting and free-flowing. I really appreciate the opportunity to talk with you. Thank you. Yeah, take care. All right, everyone, that's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimla.ai.com.

40:59Of course, if you like what you hear on the podcasts, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening and catch you next time.

From the publisher

Today we’re joined by Prem Natarajan, chief scientist and head of enterprise AI at Capital One. In our conversation, we discuss AI access and inclusivity as technical challenges and explore some of Prem and his team’s multidisciplinary approaches to tackling these complexities. We dive into the issues of bias, dealing with class imbalances, and the integration of various research initiatives to achieve additive results. Prem also shares his team’s work on foundation models for financial data curation, highlighting the importance of data quality and the use of federated learning, and emphasizing the impact these factors have on the model performance and reliability in critical applications like fraud detection. Lastly, Prem shares his overall approach to tackling AI research in the context of a banking enterprise, including prioritizing mission-inspired research aiming to deliver tangible benefits to customers and the broader community, investing in diverse talent and the best infrastructure, and forging strategic partnerships with a variety of academic labs.

The complete show notes for this episode can be found at twimlai.com/go/658.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
AI Access and Inclusivity as a Technical Challenge with Prem Natarajan - #658The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 42 min
Listen in VO