In short
AWS Podcast Episode #724: Accelerated Computing - From Fraud Detection to AI Innovation
Episode Overview
- Release Date: June 9, 2025
- Host: Shruthi Koperkar
- Guests: Ray (Container Specialist) and Sudhir Kaldindi (Principal Solution Architect in Financial Services)
- Main Topics:
- GPU-accelerated computing on AWS
- Use cases in various sectors including autonomous vehicles and fraud detection
- Architectural patterns for AI workloads
---
Key Concepts
- Accelerated Computing with AWS
- GPU Acceleration:
- NVIDIA GPUs and AWS AI chips (Trainium, Inferentia) enable AI and machine learning workloads.
- Facilitates success across various customer applications.
- Use Cases Discussed
- Rivian's AI for Autonomous Vehicles:
- Optimizing GPU use with Amazon EKS for object detection in simulated driving scenarios.
- Focus on maximizing GPU utilization due to high costs and the need for efficiency.
- Real-Time Fraud Detection in Financial Services:
- Processing over 100 billion events annually with NVIDIA technologies.
- Importance of real-time data processing and accurate fraud detection systems.
---
Insights from Guests
Ray's Insights on GPU Workloads
- Kubernetes and EKS:
- Many organizations are moving to Kubernetes for managing GPU workloads.
- Challenges include ensuring maximum GPU utilization to avoid costly inefficiencies.
- Architectural Considerations:
- Must consider networking, storage, and resource management to optimize performance.
- Use of Elastic Fabric Adapter (EFA) for low-latency communication between GPUs.
Sudhir's Insights on Financial Services
- Architecture Patterns for Fraud Detection:
- Efficient data storage and retrieval are key for analyzing vast historical and real-time transaction data.
- Leveraging AWS services (S3 Data Lakes, EMR) for scalable fraud detection systems.
- Customer Example - FeatureSpace:
- Achieving high fraud detection accuracy by combining AWS scalability with NVIDIA GPUs for real-time processing.
---
Key Takeaways
Architectural Patterns
- Maximizing GPU Utilization:
- Importance of efficient design to handle high computational costs associated with GPU workloads.
- Combining Technologies:
- Graph Neural Networks, Large Language Models, and traditional ML models work together for robust fraud detection.
Tools and Frameworks
- NVIDIA RAPIDS:
- Speeds up data processing and machine learning pipelines, reducing costs.
- NVIDIA Inference Microservices (NIMS):
- Simplifies deployment of AI models by providing optimized containers.
Cost Management
- Carpenter for Kubernetes:
- Optimizes resource allocation by determining the best EC2 instance type based on workload needs.
- Dynamic Scaling:
- Systems should be designed to adapt to changing workloads without over-provisioning resources.
---
Conclusion This episode highlighted the transformative role of accelerated computing on AWS across various industries, particularly in AI innovation and fraud detection. With strategic architectural decisions and the effective use of advanced tools, organizations can achieve significant improvements in performance, cost-efficiency, and scalability in their AI applications.
---
Additional Resources
- [AWS News Blog: New Amazon EC2 P6-B200 Instances](https://aws.amazon.com/blogs/aws/new-amazon-ec2-p6-b200-instances-powered-by-nvidia-blackwell-gpus-to-accelerate-ai-innovations/)
- [Accelerating Fraud Detection in Financial Services with NVIDIA RAPIDS on AWS](https://github.com/aws-samples/ai-credit-fraud-workflow)
---
Feedback For questions and feedback, connect with Shruthi Koperkar on LinkedIn or send an email to awspodcast@amazon.com.
---
Keep on building!
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00This is episode 724 of the AWS podcast, released on June 9, 2025.
0:09Hello, everyone, and welcome to another episode of the AWS podcast. My name is Shruti Koperkar, and today we are going to do something fun. We are going to dive into a couple of different use cases for X-rated computing. This is essentially the compute on AWS that is accelerated using NVIDIA GPUs, as well as the AWS AI chips, AWS Trainium and Inferentia. And it's purpose-built for AI and machine learning workloads. And so we are going to discuss how accelerated computing is enabling success for several different customers across different verticals. And we are going to do that by chatting with a few different guests on this episode, starting with Ray.
1:02Ray, welcome to the show. And can you please introduce yourself? Thanks for having me, Shruti. So I'm a container specialist solutions architect. I've been with AWS for about eight years. And for most of my journey, I've worked with customers trying to architect applications in the cloud. and yeah so my primary role is to now work with customers that are trying to build these massive systems on Kubernetes and a lot of them is accelerated computing since everyone is currently now trying to run some kind of machine learning solution or some kind of generative AI solution so a large part of my work now is working with customers that are trying to build ML applications within their organization.
1:55So they are trying to build these platforms that enable hundreds of machine learning experts to build models and do things like distributed training or to model serving or serving large language models, etc. So that's going to be what I did. Right, that's really cool. So, you know, just to elaborate that on a little bit, There could be teams who are directly using Amazon Bedrock, for example, and just building generative AI applications with that. But then we also have customers who have internal teams who are building their own AI ML models or sometimes fine tuning, sometimes pre-training and then deploying.
2:40and you work with the teams that help these other teams at a customer do that. So the teams that you work with typically are the ones that own the AIM and infrastructure piece where they are responsible for securing the GPU capacity or the, you know, X-rated instances powered by Tranium or Infantia capacity and then figuring out how to manage that, the cluster management, sort of tracking utilization, all of that. Is that sort of what you were trying to describe? Yeah, yeah, exactly. So there's a vast array of customers that are using AWS ML solutions or building machine learning applications on AWS.
3:31And on one hand, you do have customers that really just don't want to invest anything on building net new applications on AWS. but would still like to integrate some kind of machine learning solutions. And for that, a turnkey solution like Bedrock is perfect, where you don't have to understand infrastructure. You can just say, hey, I just want to be able to use a large language model. For that, Bedrock is amazing. And then you also have customers that want more control, that are building their own models, that really would like to tune the underlying hardware for the specific use case that they may have.
4:10So for example, we have many, many customers that do object detection when they may be simulating like a car going through a highway. And what happens when all of a sudden it sees a pedestrian crossing? So that may not be a real world situation. That entire scenario, driving scenario, may be a generated scenario. So for that, you may need really, really a lot of compute. Sometimes it's a distributed compute jobs. We are talking about tens of GPUs. This is very similar to how customers also train large language models or vision models. So these customers would benefit from having access to the underlying hardware because their needs are very unique and not what most customers would use.
5:08these models for. You know, they're building their custom models. Yeah, yeah, that makes sense. So, you know, as organizations scale their AIML workloads, as you said, many of them are turning to Kubernetes and specifically Amazon EKS. So what are some of the key challenges you're seeing as customers try to orchestrate these GPU-exilated workloads at scale? Yeah, so I think the biggest challenge that the entire compute community faces with these new workloads is that the GPUs that power these workloads are very expensive. That means anytime you're using the GPU, you as a user need to ensure that you're maximizing its usage.
6:03and that's where we see a lot of use cases suffer where a GPU may only be utilized 30 % of the time. Right. And that may not sound quite a lot because, you know, if you come from the old infrastructure world where I came from, we go for a CPU percentage and CPUs are okay. You know, they're not that expensive. If you run 30%, that's fine. GPUs are expensive. You know, we're talking about thousands of dollars that get wasted because our current applications are unable to maximize the value out of GPUs. Okay, so that's the current challenge. That's faced across the board. Anyone that's using GPU for machine learning, whether they are training models, whether they're hosting models, they are struggling with maximizing the utilization of the GPU.
6:58Okay, so, and Kubernetes community is not unique. We're still running the same application. So that's where the challenge is, is to how to maximize the GPU. How do you make sure that the GPU cycles are fully utilized so you're getting the value out of your money? And the benefit with Kubernetes is also that there are many, many frameworks that are tuned and designed to optimize GPU utilization. like their whole concept is how do we make sure we serve a workload for example let's say you may say i want to be able to build a service that translates this text from english to french right and you may build this and thousands of people are using this across the board but each person may be translating different lengths of text you know you may be translating a sentence versus somebody else maybe translating a whole page.
8:01Now, if you stack them up and want to serve them together, there's no really good way, easy way to optimize it because you don't know how much a user may be trying to translate. So you now need to have some kind of intelligent way to batch them together so that you maximize the amount of GPUs you may have. And so these are complex algorithms. And there are frameworks out there like VLLM that are designed to maximize the GPU utilization. And these products are increasingly easier to deploy on Kubernetes. So where we see a lot of customers that already have significant expertise in Kubernetes, they have teams that already manage tens of Kubernetes clusters in the environment.
8:52For them, moving to Kubernetes moving to deploying EKS for the ML solution becomes just a natural progression of their architecture. Makes sense. Now, what are some of the architectural considerations then when they are using Kubernetes for these GPU-experated workloads? What are some of the design decisions that the team need to be thinking about regarding, you know, around networking or storage or resource management as they use EKS? Yeah. Yeah. So if you're going into the self-managed world, which is where I live, it's you're architecting an end-to-end solution. And so whenever you're building a solution like that, you're not thinking in isolation.
9:41You're also keeping things like cost in mind. So firstly, it's very difficult to give a generic answer because A, we're still learning. This is still new for many of us in the community. But it also depends on the type of workload that you may be running and the amount of traffic you have. I think the biggest way you look at it is what is the need for my use case? Am I building a chatbot? Am I building a document processing solution? What is the solution? and how would I consider my solution to be success? Is it going to be speed? Is it going to be accuracy? And so that is really going to define what is going to be the metric by which you're going to gauge success.
10:34And once you're going to get that metric, you would find out what is the actual, then you would figure out how much money am I going to make from building this solution? Because GPUs are hard, you'd have to have a solid plan on how are you going to recuperate this cost. Based on the amount of benefit that you have downstream from building this application, you would then define the architecture. So there's a wide... AWS gives you a huge variety of options from computes to storage to, you know, the type of storage to type of GPUs. So you do have a bunch of options. We have customers that are building the next generation LLMs on AWS.
11:20And then we also have customers that are not trying to do the most advancing, but would still like to benefit from having the option of having different types of GPUs being available. So I have the option of using not the latest generation NVIDIA GPU, but I can use maybe a few generation old because it's fine for my workload. so first you define what that metric is how much what kind of compute what kind of performance you would need and then you work backwards from that so if you really need fast storage fast distributed storage you would then start working through like what are the storage options in AWS and I would go through each one of those options as an as an architect you would go to each one of those options let's say you know this is the most appropriate option but hey this is also the most expensive option.
12:12Where do I find that balance? So it's the kind of thing that you do, but in terms of distributed compute, you're looking at the fastest, whenever you're talking about multi-GPU solutions, you're looking at having the fastest link network-wise between the two GPUs available. So on AWS, that's EFA, the Elastic Fabric Adapter. So you definitely want to make sure that there's the lowest amount of latency between your GPUs. You definitely want to make sure that they're sitting as close as possible. So in terms of 8-APS, it's the same AZ. You want to reduce any kind of storage latency. So GPUs often have to copy data from disk into memory.
13:01and if your GPU and this, but the way this operation occurs is there's also CPU in between. So if your CPU is a bottleneck, if your disk is a bottleneck, then your GPU is just waiting, just trying to extract data from the disk. That's not a good situation to be in. You want a GPU to be like, always be processing. You know, it never waits for data. It never waits to exchange data between other GPUs when you're doing distributed training, and it never has to wait to load data from the disk because you're wasting dollars there. Right. Is that when something like Amazon FSX for Lustre comes in is where it can really feed data very fast to the GPUs?
13:51Absolutely, yes. So FSX just recently added the support for GPUs to directly read data from a distributed storage. So that is really, really helpful because you bypass a CPU. And FSX for Lustre is a great service for anyone that's looking to do distributed training because in that situation, what you're looking to do is you have a vast corpus of data, basically a bunch of files. and these could be terabytes to petabytes of files. And you may have hundreds of GPUs trying to create data from each one of these files. And so that's where you need distributed storage. That's where you need the ability to scale to that kind of hundreds of instances reading at the same time.
14:43So FSX for Lustre is great for that kind of use case, yes. Awesome. Okay, so let's dive into a customer example. And I know that you've worked closely with Rivian, and they are obviously using a lot of GPUs and together with EKS. Can you walk us through that use case, perhaps, and like sort of to the extent that you can share what were some of those sort of criteria in terms of performance or budget or time to market that they were trying to hit? And then what did that mean for their solution architecture that AWS and them designed together? Essentially, what they have built is they have these massive jobs that may be doing things like object detection.
15:35Let's take that as a generic workload. You may be running a car and simulating that through a scenario where you may say that a car goes from spot A to spot B. And during that spot, there are a bunch of objects that it has to detect. Maybe it has to stop for pedestrians or stop at a stop sign or slow down at a school crossing, things like that. So these tasks are typically multi-GPU tasks, and sometimes these may take tens of instances, P5 instances. And P5s are pretty expensive instances. They have eight NVIDIA H100 GPUs, if I'm correct. but these are massive instances and they are also very, very difficult to get in terms of availability.
16:34So when you do get them, you want to make sure that you're maximizing the utilization. So what Rubian has built is that they have built a stack on top of this project that AWS started, which is data on EKS. And as part of this project, we have an ML reference architecture implementation called the JARC stack, which is Jupyter Argo Ray on Kubernetes. And so Rivian has extended that stack and in the GTC talk, they've gone through the architecture in detail. in detail. But what that stack does is it runs this scenario across a distributed array job. And that distributed array job is orchestrated by Argo.
17:40So it runs on a bunch of EC2 nodes that are Kubernetes cluster worker nodes. And the way it does it is that they have a bunch of EC2 instances and they have a bunch of jobs running and their job as the infrastructure team is to ensure that A, the number of jobs remain low and there are not that many pending jobs in the queue because at that time, a developer is waiting, wasting their time. And at the same time, try to not make sure, try to ensure that there's not a P5 instance per developer. So they want to ensure that whatever the number of instances they have, they're utilizing it fully and efficiently.
18:29So the way they use this, they use Argo workflows, which is a workflow engine that runs on top of Kubernetes. And they use Argo and EKS to then make sure that, hey, I have, let's say, 200 P5 instances. These are the number of jobs that I have, and Kubernetes then makes sure that they have the loads, the jobs get scheduled as needed. And so a lot of this is also bad scheduling. So they also implement bad scheduling to make sure that the jobs are co-located and if they need to intercommunicate, the latency is low, et cetera. Awesome, awesome. I think you mentioned Jupyter. Argo for the Argo workflows, which is an open-source container-native workflow engine.
19:19And then you also mentioned Ray, which is a distributed computing framework that is designed to build and scale AI ML applications. Yes. Is that right? Yeah, and I think Ray by any scale, any scale is one of AWS partners. Yes, yes, yes. Awesome, awesome. So one other question I had, since we are talking about containers, containers is NVIDIA has launched NIMS, which is the NVIDIA Inference Microservices, which is essentially sort of a container that comes packaged with the LLMs and all the optimizations and libraries that might be needed. How do those work with EKS? Have you seen customers sort of use NIMS with EKS and across what verticals have we seen those?
20:13Yeah, so what is NIMS? First of all, NIMS is also addressing the same problem that we started with, which is how do I ensure maximum efficiency of a GPU? So that's one challenge. The other challenge also is that how do I ensure that I can run models on the limited hardware that I have? Okay, so not everyone has the luxury to get a P5 instance. I can't afford it, but I still want to be able to run, let's say, an open source model. And so a lot of times what you have to do is that you have to do a lot of model customization to be able to run it on a commodity, EC2 instance. You know, Bedrock or like chat tools really make it really super easy to start chatting.
21:11but if you were to deploy your own large language model it's not that easy because you may not have the latest hardware you may not have the best hardware you may not have multiple EC2 instances so a lot of times you may have to do a series of optimizations to be able to get it to run on your own hardware which is okay you can do it but not everyone knows the best techniques and not everyone knows what to trust you know at this point So NVIDIA has these containers that are optimized that you can say, hey, these are optimized for this kind of hardware, and it will maximize the utilization. And these are tuned to run on most hardware.
21:56So depending on the type of GPU you're running, you may be getting a different setting. So NVIDIA has made sure that you don't have to be running those per model optimization. Like Lama may have a different set of optimization and then a different model may have a completely different set of optimization. How do you know which one is the best way to run? And so there's a lot of research that's needed. And so NVIDIA NIM basically eliminates that research. You can take the NIM container and then run it in your environment. And the beauty is that it plugs in with a lot of other backends. So you can use it with the tool of your choice.
22:39Awesome. Since we are talking about containers There's one more thing that comes to mind Which I've seen a lot of customers use And that is Carpenter And can you maybe chat a little bit about What Carpenter is and how it helps With sort of optimal resource utilization Okay I would have preferred a longer conversation Longer time but let me give you a very, very quick summary of what Carpenter does. So a lot of times, the people that are running applications in the cloud versus the people that are building applications in the cloud are different people. And so I may be a developer and I say I'm writing my code and I write my code, I build my application and say, okay, operations guy, run it.
23:34Now the operations people are like, okay, fine, I'll run it. but how much capacity should I allocate to your application? And I would say, you know what, it's not my job to find out what the application capacity is, so give it the maximum, like 50 ,000 GPUs. But that's not practical, right? So that may be too expensive. So what ends up happening is you have a lot of wasted capacity that's not really fully utilized. So a lot of times we have to do, as infrastructure architects, you may have to do right-sizing. You may have to say, for this application, this is the best site of EC2 instance. And that may be good for one application, but how do you do it at like this 500 application level?
24:20Impossible. So the right way to do it is, hey, let the environment decide when to scale. So what Carpenter does is in Carpenter or with EKS Auto Mode, you just go and create containers. And in creating containers, you'll say, this is the amount of CPU and GPU or memory I want. And Carpenter will then decide, and EKS auto mode will then decide what type of EC2 instance you should create. So I don't have to decide whether this should run on C5 instance, because as an operations person, why should I decide what this runs on? And developers, most of the times, they don't know what type of instance it should run on.
25:04They're like, Give me an instance with a disk, you know? So Carpenter does, optimizes that kind of compute scaling, and it does two things. One is it makes sure that when you scale up, you scale up reacting to the workload that you need to deploy. So if you just want two vCPUs and two gigs of RAM for a small workload, Carpenter will deploy an EC2 instance that just matches just that amount. Now, over a period of time, if you deploy many, many instances and Carpenter thinks that replacing them with a larger instance will be cheaper for you, Carpenter can also do that. So you can let Carpenter decide that and say, hey, it's okay if you disrupt my workload between this time and that time just to save costs.
25:53So Carpenter is basically a Kubernetes autoscaler, which now is open source, by the way, and also available across other clouds. And so it's designed to make scaling really seamless on EKS and Kubernetes in general. Awesome. This was really fun, Ray. Thank you so much. I know that personally, I would like much more time with you to dive deep and understand some of these technologies, but you really provided a great overview of what folks should be thinking about when they are running GPU workloads or X-rated workloads on EKS and the various different frameworks and tools like Ray and Carpenter and how they can help them.
26:42So thanks a lot for joining us. Yep. Thanks for having me. Okay. So for the next segment in this episode, we are going to look at X-rated computing use cases across the financial services sector. And joining us to talk about this is Sudhir. Sudhir, can you please introduce yourself? Sure, absolutely. Thank you for having me, Sudhir. Hello, everyone. It's great to be here. My name is Sudhir Kaldindi. I work as a principal solution architect and worldwide financial services at AWS, specialized in payments. Awesome. Welcome again. Can you maybe talk a little bit about how is X-rated computing, which is, you know, are compute instances powered by the NVIDIA GPUs or are purpose-built AWS, AI chips, train them, and Infantia?
27:32How is this portfolio of compute being used by customers in financial services? What are some of the challenges that they are trying to solve using these instances? When we're trying to build real-time fraudulent solutions, what we see is there are several key architectural considerations especially will come into play especially when leveraging nvidia technologies on aws so financial institutions right they face the challenge of processing vast amounts of transaction data especially in real time which requires a robust infrastructure and also efficient data processing uh what we have seen is using nvidia snippet speed which is integrated with aw services like amazon emr which accelerates data processing and also using machine learning pipelines, which enables faster fraud detection.
28:20And the use of graphical neural networks in NVIDIA's AI workflow, which enhances fraud detection accuracy by analyzing complex relationships within transaction data. And we also see, right, customers today using NVIDIA's Morpheus framework and also the Trident Inference Server, which are very crucial for real-time transaction analysis and model deployment, which ensures swift decision-making and also AWS services like Amazon SageMaker and Amazon EC2 support model training and deployment, which allows for scalable and efficient fraud detection systems. So what we encourage customers is by again combining NVIDIA's AI solutions with AWS Cloud Infrastructure to the financial service customers, they can achieve up to 14 times faster data processing and model inference alongside side significant cost detection.
29:10So this integration allows for real-time fraud detection and also reduces false positives, improving customer trust in building the fraud detection solutions on AWS. I see. So, you know, a lot of customers in financial services are using these GPU-experated instances to combat fraud. And this requires, as you said, a lot of real-time data processing, instant decision-making, and that's where the, you know, obviously the NVIDIA GPUs, but also some of this other software that NVIDIA has built comes into play. Could you break down maybe some of the architectural considerations when building these solutions?
29:56Like, especially given that, you know, there is sort of this real-time processing required. What are those architectural considerations? What does that solution architecture look like beyond, you know, using the GPUs and some of the software services you used? Because AWS provides a wide array of services that kind of go around this X-rated compute. Right. Yeah. So that's a very valid question what we hear from the customers, right? When we're trying to work with financial service customers to build a product system, the several key challenges, what they said is, especially the huge amount of transaction data, which they handle daily.
30:33Again, this definitely requires efficient processing and systems to process it quickly. And again, it's not about today's data, right? We're also talking about years of historical information, as this history is very valuable to identify the patterns, and which also had presence its own set of challenges. So one of the first challenges what we see is, how do you try to store and retrieve this vast volume of data, information which is very efficiently? and the second thing is how do you try to correlate the older data with emerging trends so coming to the architecture patterns here right the first thing is where you need to move your data again if you're trying to run from your own premises we do have many services where you can move the data into AWS and the second thing is if you have your own data residing in the silos we can also store efficiently meaning process the data effectively and then push it into again S3 Data Lakes.
31:29And through S3 Data Lakes, what you can do is you can able to design where is the cross storage. Coming to the volumes, what we have seen is using the Trident entrance server, we could be able to process close to 350 ,000 transactions per second, which is a huge scale in identifying the fraud. Awesome. That's really great. And that's a helpful sort of guidance on what type of architecture patterns you've seen and what helps. It sounds like data is, I mean, we've all known that data is central to AI, but this is where it really becomes apparent where you have not just like live new data, but also a lot of historical data.
32:12And how do you process all that data at scale becomes important. That is correct. And we also have the workshop related to this. again we encourage customers to get into this public git repository so that they can start building the architecture what we have built on top of yeah no that would be great we can actually include a link for that on our show notes and so listeners please check out if you're interested we will include one link um let's dive into a sort of a customer example maybe like specific customer story you've worked with FeatureSpace and their FeatureSpace's ARIC risk hub is something that they are building on AWS they're processing over 100 billion events annually with really impressive fraud detection rates can you walk us through maybe how they are leveraging AWS as well as NVIDIA technologies and achieving these results And maybe some of it is things that you've already mentioned, but it will really be helpful to understand from a very specific use case perspective.
Read the full transcript
33:25Sure, sure, Sruti. So, FeatureSpace, Eric, Risk Hub, right, which is definitely a great example of how cutting-edge technology can tackle financial fraud at scale. Especially today, FeatureSpace, right, they are passing close to, as I said, right, over 100 billion events annually, which they have built a robust system with remarkable fraud detection rates. So the big part of this success which comes with the usage of both AWS and NVIDIA technology. So coming to the AWS side, right, the feature space leverages, again, the scalability and reliability which are definitely neat to handle vast volumes of transactions in real time.
34:03And again, especially the feature space, they use our elastic computing power and storage which means Eric Riscop can seamlessly scale up to, again, thousands of transactions per second, which during peak times, without even compromising on performance. And also, it ensures uptime and low latency, which is essential for fraud detection, where every real, every second matters, right? Because whenever a transaction is happening at a point of sale or at an e-commerce transactions, customers do not really want it to wait at the transaction part, right? So, immediately, the transaction has to make sure that it is not fraudulent.
34:39and the transaction has to happen seamlessly. So that is where AWS is really helping the feature space to build the live transactions on AWS. And coming to the NVIDIA part, so they've been leveraging NVIDIA GPUs and here they use NVIDIA GPUs to train their deep learning models which power their adaptive behavior analytics. And again, GPUs are incredibly efficient at handling complex mathematical calculations, right? which is required for AI training. And again, this deserves in faster model development and also more accurate fraud detection. So by combining NVIDIA GPUs acceleration and also with AWS scalable code and format, FeatureSpace has created a dynamic fraud detection platform that not only reacts to fraud attempts, but also learns and evolves to predict new threats.
35:30Awesome. I think you may have mentioned NVIDIA Rapids and Rapids Accurator for Apache Spark on Amazon EMR earlier. Can you maybe talk a little bit about that particular solution and how it's helping institutions, you know, tackle some of these more challenges while keeping costs down, which is always important? Yeah, sure. Especially, right, the costs which are tied to the feature engineering and model machine learning models, right? They can really add up. Again, if you're not really trying to optimize the cost is going to be significantly high. So that is why finding ways to cut those costs is very super important, I would say.
36:13So one of the game changers in this space is NVIDIA Rapids. Again, it's a tool that speeds up data processing and machine learning pipelines, which means they can handle huge amounts of data much quicker than with traditional CPU-based systems. Again, this speed is absolutely critical, right? When it comes to real-time fraud detection because it lets institutions react to emerging threats right away. again but it's not about the speed it's also about saving money so here the traditional cpu systems often need more resources and take more time to process data which again drives up cost so that is where the gpu accelerator solutions like rapids they can process data much more efficiently cutting down on the need for extra hardware and also reducing oral expenses things get even better with rapid accelerator for apache spark on amazon emr which boosts the speed of model training this is crucial because fraud tactics which are always evolving and institutions right they need to adapt fast so by using these technologies together right especially the financial institutions they can create fraud detection systems that are faster more scalable and can handle new threats while keeping costs under control.
37:26In short, right, what I would say is the solutions are reshaping fraud detection, making it faster, more affordable, and also more scalable. And that's exactly what financial institutions need to stay ahead of, potential fraud and protect themselves from big financial losses. Yeah, so true. I mean, it sounds like these institutions have to constantly innovate because, you know, the fraud, you know, keeps evolving. That is true. And being able to take that is really important. One last question. You know, as you mentioned, like, they have to constantly evaluate. And we are seeing this convergence of different AI technologies.
38:07So there's graph neural networks. I think you mentioned those earlier. There's large language models. There's large transaction models. Correct. How are all of these working together to kind of create sort of that innovation or that, you know, more sophisticated detection systems? Yeah. So when we look at how AI technologies like graph neural networks, which we also call as GNNs, and also large language models, LLMs, and large transaction models, right? So they do work together, again, for fraud detection. What we see, it's very exciting. so these technologies are combining to create very advanced systems so if you look at graph neural network what they do is they do analyze large complex transaction methods again through networks which i would say to find patterns that might show fraud so the example i can give you is again the linkedin right again which we are connected to each other so that is how hcnns work uh behind the scenes and they are good at catching unusual things that other models might fail.
39:11So when it comes to large language models, they are good at processing lots of unstructured data like invoices or emails to find anything suspicious. So again, this helps business or even financial institutions, right, to catch problems faster and also work more efficiently. Coming to large transaction models, they also play a big role by processing a lot of data to find patterns that indicate fraud. But the best part is when all three technologies work together, they give a complete picture of transactions, reducing false alarms and also catch fraud in real time. I would say, obviously, by combining graphical neural networks with traditional machine learning models, right, which could be XGBoost, which makes them more accurate and also easier to understand.
39:55And overall, this combination of AI technologies, right, is really changing fraud detection. And again, one fraud model really may not really serve the purpose. So that is where we encourage customers to start building like different several models fit for the purpose and which makes systems more accurate, efficient, and also better prepared for the future. Awesome. That's a really, really great overview. And you really help break down what each of these different types of models or algorithms are good for. And of course, all of them can be accelerated or rather are accelerated using the GPU accelerated EC2 instances, as well as several different sort of software libraries available from NVIDIA.
40:39Thank you so much, Sudhir, for joining us, for talking to us about how customers are using, you know, accelerated EC2 instances in the financial services sector. Thank you. Thank you for having me, Sruti. Thank you all. Okay, so that's it for this episode, everyone. It was great chatting with Ray and Sudhir on how X-rated computing is helping customers achieve great business outcomes across different industries. If you have any questions, you can find me as Shruti Koperkar on LinkedIn or X, or you can send us feedback to awspodcast at amazon.com. And until next time, keep on building.
From the publisher
Join host Shruthi to discover how organizations use GPU-accelerated computing on AWS. Container Specialist Re Alvarez Parmar shows how Rivian optimizes GPU usage for autonomous vehicles with Amazon EKS. AWS Financial Services expert Sudhir Kalidindi explains real-time fraud detection processing 100B+ events annually. Learn architectural patterns and tools to maximize performance while controlling costs for AI workloads and next-gen applications.
Learn More:
AWS News Blog: New Amazon EC2 P6-B200 instances powered by NVIDIA Blackwell GPUs to accelerate AI innovations: https://aws.amazon.com/blogs/aws/new-amazon-ec2-p6-b200-instances-powered-by-nvidia-blackwell-gpus-to-accelerate-ai-innovations/
Accelerating Fraud Detection in Financial Services with NVIDIA RAPIDS on AWS: https://github.com/aws-samples/ai-credit-fraud-workflow
