Going Serverless in Financial Services with Brian McNamara

7 Jan 2025 · 38 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Going Serverless in Financial Services with Brian McNamara

Podcast Details

  • Title: Software Engineering Daily
  • Episode: Going Serverless in Financial Services with Brian McNamara
  • Description: The episode discusses serverless computing, its adoption in financial services, challenges faced, and insights from Brian McNamara, a Distinguished Engineer at Capital One.

Key Concepts What is Serverless Computing?

  • Definition: A cloud-native model allowing developers to build and run applications without managing server infrastructure.
  • Benefits:
  • Scalability
  • Reduced operational overhead

Challenges in Financial Services

  • Unique complexities in adopting serverless models due to regulatory and compliance issues.
  • Importance of governance controls and cost management.

Brian McNamara's Role at Capital One

  • Distinguished Engineer focusing on serverless integration at a tactical and strategic level.
  • Engages with teams to adopt serverless technology and collaborates across different business lines.

Capital One's Shift to Serverless

  • Cloud Journey: Initiated in 2014 with a focus on improving developer experience and reducing developer burden.
  • Serverless First Approach: Established in 2021, prioritizing serverless technologies for new projects.
  • Cost Considerations:
  • Three components driving total cost: engineering cost, cloud infrastructure cost, and maintenance cost.
  • Serverless may seem more expensive at the cloud compute level but offers significant savings in maintenance and operational efficiency.

Advantages of Serverless

  • Minimal Management Costs: Reduced focus on server management allows developers to innovate.
  • High Availability and Resiliency: Serverless services like AWS Lambda provide built-in resilience and scale automatically.

Overcoming Challenges of Serverless Adoption

  • FUD (Fear, Uncertainty, Doubt): Addressing misconceptions about serverless capabilities and workload suitability.
  • Best Practices: Importance of rethinking application architecture; avoiding monolithic designs in favor of microservices.

Current Status and Future of Serverless at Capital One

  • A significant number of applications are running on serverless compute, leading to enhanced focus on delivering business value.
  • Continuous evaluation of new AWS services for compliance and security.

Monitoring and Observability

  • The need for robust monitoring solutions even in serverless environments.
  • AWS provides metrics, but teams must adopt best practices for observability to ensure application performance.

Governance and Compliance

  • Security as a Priority: Rigorous processes for evaluating new services and maintaining compliance with regulatory standards.
  • Tools and Automation: Use of Open Policy Agent to ensure deployments conform with internal policies.

Developer Experience and Productivity

  • Balancing speed and governance to minimize friction for developers while ensuring security and compliance.
  • Continuous feedback loop with developers to improve tools and processes.

Challenges and Lessons Learned from Migration

  • Initial resistance and concerns regarding serverless capabilities.
  • Importance of understanding AWS constraints and the right use cases for different compute types (e.g., Fargate vs. Lambda).

Conclusion and Future Directions

  • Anticipated growth in serverless technologies, focusing on durable workflows, better visibility into the software supply chain, and improved operational experiences.
  • Emphasis on the importance of observing and understanding applications in serverless environments.

Key Takeaways

  • Serverless computing presents unique opportunities and challenges in financial services.
  • A robust governance framework is essential for successful serverless adoption.
  • Developer experience and feedback are crucial for fostering a productive environment in a serverless-first organization.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Serverless computing is a cloud native model where developers build and run applications without managing server infrastructure. It has largely become the standard approach to achieve scalability, often with reduced operational overhead. However, in banking and financial services, adopting a serverless model can present unique challenges. Brian McNamara is a distinguished engineer at Capital One, where he works in serverless integration and development. Brian joins the show with Sean Falconer to talk about why Capital One shifted to a serverless approach, how to think about cloud costs, establishing governance controls, tools to stay well-managed, and much more.

0:38This episode is hosted by Sean Falconer. Check the show notes for more information on Sean's work and where to find him.

0:57Brian, welcome to the show. Hey, Sean. Thanks so much for having me. Yeah. So to start our conversation, can you explain a little bit about your role at Capital One? Sure. So I'm a distinguished engineer at Capital One. Essentially, that's a senior IC role. And my role really revolves around considering how we do serverless at scale, both at a tactical level and a strategic level. So at a tactical level, I have the opportunity to engage with different teams who were looking to adopt serverless compute, help them run through any outstanding questions that they might have around the technologies, help them decide which compute may be better for their given workloads.

1:37But then we also have the opportunity on my team to actually work with others across the enterprise. So in terms of where I sit at Capital One, I'm in the retail bank, but I have the opportunity to really work with partners in other lines of business, like our card business, enterprise, cyber, ML. So it's really nice to operate really both at a very high level and really a very low level too. When you reach sort of the level of distinguished engineer within an organization, is that kind of your primary thing that you're doing day-to-day? Instead of necessarily sort of hand-on-keyboard coding, you're a little bit more like a thought leader within the company and working with some of these other teams to help them invest in certain technologies where you have deep expertise in?

2:23I'll say yes and. So one of the nice things about the distinguished engineer role at Capital One specifically is, you know, you do have the ability to influence really how a large engineering organization works or considers problems. So in that sense, we do get to work on the big boulders. In my role, though, it's actually a really good blend in that I do get to get my hands on the keyboard. So while I may not support feature teams and actually write production code, I am working on ways, whether it's in sample applications or templates, looking for ways to help streamline the overall adoption of serverless like Cap1.

3:01Okay. And what would you say is kind of like the primary difference between operating at this level as an engineer versus, say, being at maybe like a senior staff engineer level? Yeah. So I think, and I mean, obviously, this is my experience. It's the ability to influence strategy, I think, is one of the main differentiators. So while I do get to work with those future teams who are looking to adopt serverless compute, it's working with our partners across the enterprise, not just within a single line of business. And that really, I think, allows us to take a more holistic view in how our engineers approach solving business problems with serverless technology.

3:45I see. And why did Capital One decide to make this transition to service technology? Great question. So, you know, if you look back at Capital One's history, we really started our cloud journey in 2014. And, you know, if we look at how we've adopted cloud technologies over the years, I think it followed more or less, I mean, I hate to say a traditional pattern, but, you know, you can look at initial efforts revolved around, you know, like literally lift and shift. So, you know, moving existing stacks from on-prem to the cloud. But really, as the years progressed, you know, our leadership was looking for ways to improve the overall developer experience and quite honestly, reduce developer burden, you know, letting developers focus on what they do well.

4:31So really in 2021, we made a declaration that we were going to be a serverless first company. And what that's meant is for newer projects, newer applications, not that you have to go serverless first, but serverless technologies like Lambda and like AWS Fargate should be considered first. And if we look at the reasons why, one of the canonical examples that people provide for why any organization goes serverless is lower cost. I think it makes sense to unpack that a little bit more. If we look at what's driving that cost, there are really three components. There's the engineering cost. There's the cloud infrastructure cost.

5:15And then there's the maintenance cost. If we look at the cloud compute cost, you may actually find that serverless, when compared with an equivalent or an analog for an instance-based compute, may actually be more expensive. So people often hold up cloud cost as a reason to not consider serverless. I think our leadership has looked beyond that and said, well, really, if we look at what's really driving overall cost or the total cost of ownership, it's not necessarily only cloud cost. We have to consider maintenance costs and the ability to innovate. And I think really that's what serverless provides.

5:50You know, if we look at the dimensions of what constitutes serverless compute, you know, really it's all about minimizing management costs. So, yes, there are servers and serverless, but you don't have to manage them. You don't have to patch them. You can have compute that scales to really, really high levels if needed, but can also scale down to zero if needed, too. You also have compute that really runs when it's needed to. So in response to business events, and you have built-in resiliency. So looking at, let's say, a service like AWS Lambda in particular, you get high availability out of the box.

6:25Whereas if you look at instance-based compute, you may need to consider like, okay, how resilient do I want this application to be to different types of failures? Well, with a managed service like AWS Lambda, well, we as customers don't have to worry about that. And we can let developers focus on the interesting problems. Yeah, I mean, I think the point you raised there about sort of speed of innovation and being able to sort of potentially plug and play different parts of the stack as essentially the public cloud introduces new technologies. it's a little bit easier to be adaptive versus if you are sort of running this stuff yourself, then in the same way that like a monolithic application, you end up with like this tight coupling of different services and it makes it hard to be essentially as adaptive because every team kind of needs to be in sync and it's going to really slow down your price.

7:14If you're running this yourself, you're sort of almost like you're tightly coupling your software to the sort of on-prem system that you're running versus having a little bit more of a decoupling between the service that you're creating and essentially taking advantage of something like public cloud. Does that make sense? Yeah, exactly. Like we're working with wonderful cloud partners like Amazon. You know, this is a great time of year where we get to see what Amazon's been working on for the past year. You know, it's always exciting knowing that we have a partner that's innovating on our behalf.

7:44Yeah. And then some of the problems I've seen companies run into when it comes to actually like cloud costs being expensive comes from not really going through the process of like rethinking or re-architecting the way that they are building and running their software to be best suited for the cloud, right? Legitimately just sort of lifting and shifting something into the cloud is not going to help you save money. You need to go through a process of sort of, you know, rethinking how those services are going to take advantage of the fact that not everything has to run in the cloud, like, you know, 100 % of the time.

8:19And that's how you can actually help reduce your cost if you're optimizing for the things that the cloud is actually good at. Yeah, I would agree with that. And with like, you know, let's say Lambda applications in particular, even when we see, you know, teams that are interested in moving from other compute to Lambda, there's this initial thought to build essentially like a monolithic Lambda or Lambda-lif. Yeah. And without even, you know, if they don't initially go through the process of decomposing their application, breaking down that monolith, they may find that Lambda may not be the right choice for any number of reasons.

8:52But I think for teams that take the time to consider what their applications do, how they're invoked, and really break things apart, they find that Lambda can really be a suitable landing place for them. So I believe you said that this sort of project kick-started in 2021. So what is the status of things today in terms of Capital One's journey to serverless? Yeah. So we have a large number of applications that are in fact running on serverless compute. I can't share the exact numbers, but it's a non-trivial number. And what we've seen for the teams that have made the transition is that they are able to spend time or more time on adding business value, on delighting customers and much less time on patching infrastructure or managing availability in the event.

9:40Like as a part of our normal process of launching applications, we ensure that applications are resilient to different types of failures. Well, with serverless compute, that honestly becomes a lot easier. You know, we've talked about a couple of the advantages of serverless in terms of, you know, the team being able to work faster, be a little bit more adaptive to changes in technology. Is there also an advantage in terms of, you know, bringing engineers into the organization where they might have experience with a lot of these services already? And it's kind of like the way that they're used to working.

10:12Yeah. I mean, I would say for us, you know, as a large enterprise, we have about 14 ,000 engineers. So when we do bring new engineers on, it's often easy, you know, yes, there, you know, as with any, any new hire, there is a ramp up period, but, you know, we find that people are happy to embrace serverless technologies, you know, without necessarily being beholden to legacy architectures, you know, figuring out what the right server is to jump to and, you know, all those activities that don't necessarily add value. So it's, you know, we find that the more that we lean into managed services, the more productive our developers can be.

10:51And then, you know, you don't have to share specifics, but can you help maybe for the audience, like shape the amount of like data or traffic that you're dealing with and actually running through AWS? Like how big is this essentially? So we are, you know, a very, very large AWS customer. Without going into specifics, I can say, you know, you can think of CAP1 as a very, very large user of AWS, a strategic partner with AWS, you know, akin to a Netflix. So, you know, we are all in on the AWS public cloud. In fact, we shut our last data center back in 2020. So really, whether it's, you know, serverless compute, instance-based compute, storage.

11:32We are very, very large users of AWS. And what are some of the high-use services that you use within AWS? What were some of the original workloads that you wanted to move over? So we are large users of serverless compute, as you might imagine. But we do also have a large footprint in more traditional compute. So whether it's instance-based like EC2 or ECS on EC2, if we look at other services that are more serverless in nature, S3, You know, we are very, very large users of S3 that powers, you know, a lot of our AI and data analysis workloads. You know, then there are the connective services like Amazon Simple Node Notification Service, Amazon Simple Queue Service.

12:16So really, we don't use all AWS services, but we do use a lot. And what we do use, we're pretty heavy users of. as you might imagine with, you know, despite being, you know, yes, we are a very, very innovative tech company, but we are a financial services company first. So we do have to be mindful in how services are approved and in how we govern usage, you know, within the developer community. Yeah, I'd love to get into details on that. But one thing before we jump there is like with many industries, like banking's gone through a lot of transformations, I think over the last certainly 10 to 15 years.

12:54And I think like one area that has changed banking significantly is around the fact that now if I'm a customer of Capital One, a lot of my access to my accounts is going to be coming from like a mobile device. Whereas before, you know, maybe I actually had to go into a bank or reach an ATM machine or something like that to make something happen. And I would think that that would have a pretty massive effect on sort of the traffic that you end up seeing, because I can essentially just be hitting that banking app multiple times a day, doing transfers, doing refreshes, all kinds of stuff that you never would have seen before.

13:27Yeah. So for us, we have recognized the industry changing from being exclusively to a physical experience to the virtual experience as well. And I think Capital One places a really high value on delighting customers wherever they are. So whether it's in a retail bank itself, whether you're a customer of our retail bank who uses our mobile app or our website or similarly a credit card holder, we want to meet you where you are. Yeah, and I guess some of the elasticity of the cloud really helps with potentially spiky traffic. Yes, yeah. And you bring up a great point there. I think when people were initially migrating to the cloud, like when you saw the public cloud become more of a thing, like in the mid, we'll say 2010s, where you saw larger businesses adopting cloud technologies.

14:21I think the initial value that people held up was that you'll save money. And I don't know if that's necessarily true. I think what you're really buying is elasticity, the ability to be wrong. Right. You know, like if you think about what it takes, like the process of actually racking and stacking a physical server, like there's a whole lot that goes into that. And you better be right. You know, otherwise, you know, you're living with that decision for a non-trivial amount of time. I think what the cloud offers you is, you know, one, the ability to be wrong. So if you don't get like if you're looking at instance based compute, you don't have to get it just right, right out of the gate.

14:56You know, it's just an API call to destroy that instance. It's an API call to create a new instance. But I think the point that you bring up about adjusting to customer demand and having that elasticity built in, like even if you get that instance sizing right or the Lambda function configured with just the right amount of RAM, well, you may need to scale horizontally. And I think that's really the power of the cloud. We don't have to have capacity just waiting. And with each type of compute that is offered by cloud providers like Amazon, it's interesting to look at what the unit of scale becomes and how fast that unit of scale can be applied.

15:34With services like Lambda, you can see the unit of concurrency, which is really the unit of scale. You get a large number in your account and in your region out of the box. But depending on your needs, you can work with Amazon to increase those limits as needed. And we've certainly seen use cases where we have had to go well beyond those account default limits. But beyond that, you're able to get more and more capacity and scale to really high levels pretty fast, too. Whereas even if you compare lambda scaling with, let's say, instance-based auto scaling, you're talking about potentially milliseconds or hundreds of milliseconds versus minutes.

16:19So it lets you more closely associate scale with the business value. You don't have to over-provision or worry that you're under-provisioned. Right. What you need as you need it. And does working in sort of these elastic service environments change the way that you have to think about monitoring and observability? I'll say yes and no. So, yes, you know, it doesn't change the need, right? You know, there is a need to understand what's happening in your application. That need doesn't go away. Just because AWS is handling, you know, we'll say the underlying infrastructure, you know, it doesn't absolve you of making sure your applications are monitored appropriately.

17:02It does change how you do it or the mechanism that you use. So, like, in many regards, I think working in serverless environments forces you to be more disciplined in how you approach observability. You know, there's no instance SSH into or, you know, where you can run, you know, HTOP and see which processes are consuming CPU, what you're, you know, running IO stat or, you know, VM stat, you don't have an instance, right? So like, I think you do need to be more disciplined. The good news is, you know, AWS does provide, you know, a number of metrics out of the box, whether it's to deal with concurrency performance, you know, like there are a number of metrics that are there.

17:44but you also have the ability to add your own instrumentation as well. So if you want to write custom metrics that are associated with, let's say business value, user signup, user abandonment, or deposits made, you have that ability. You can build that logic in itself. And the nice thing is there's a really good, vibrant community that has, I think, shown how to do this well. So whether it's using utilities provided by AWS. So I'll give a shout out to a capability that AWS offers called AWS Lambda Power Tuning. Initially, power tuning was built to help with observability. So if you look at what the module provides, so there's power tuning for Python, Node, and for Java.

18:29The nice thing is logging, metrics, and traces are all first class citizens in those modules. So you don't have to do a whole lot to get a lot of visibility, which is really nice. But if you wanted to embrace industry standards like open telemetry, like you can do that too. And there is a good story for OTEL with Lambda in particular. So if you want to instrument code, you can. Otherwise, you can lean into the SDKs to auto instrument code. Obviously, there are tradeoffs there. If you are using a non-compiled language, auto instrumentation, you'll see a penalty at cold start. But you can do it.

19:06You don't have to be an OTEL expert to get that visibility, which is great. So I wanted to talk a little bit about essentially governance and security compliance. A fun stuff. Banking is a very regulated, sensitive industry. You have to take care with anything that you're building. Obviously, you're sort of dealing in customer trust. How does that kind of change the way that you have to think about building products when you work at a bank? Yeah. So, I mean, at Capital One, security is job one. As you point out, we are working in a regulated industry. We need to make sure that we're doing things the right way from a security standpoint, from a governance standpoint.

19:45What that means is that as new services are introduced, we may not be early adopters. We need to do evaluations to determine whether or not services have the necessary controls that we feel they need. So, you know, there's that part of it, making sure, you know, services are imbued with the necessary controls that we need internally. But beyond that, we need to make sure then that our process of building and deploying code also is very rigorous and stands up to compliance. So making sure that all artifacts are versioned, making sure, you know, like some of the practices that we adhere to, we use tools like Open Policy Agent or OPA to make sure that anything that we deploy conforms with our policies, that we're not deploying services that we're not supposed to deploy.

20:37Making sure that even for the resources that we can deploy, that they are compliant, that we're not using certain properties, or if we are using certain properties, that they're configured a particular way. So you can get really, really granular. And the nice thing is you can offload burden from your developers in figuring it out, right? So we use OPA so that developers can deploy compliant applications. You know, it's an important part of how we deploy code. And then in terms of, you know, the same care that you have to take when potentially adopting new services within AWS, like what is that process like when you look at potentially bringing in, you know, a new library within the actual source code that isn't something that was developed at Capital One Accru?

21:25Obviously, I'm sure you must have some checks in place to make sure that there's not some sort of malicious supply chain attack that is hidden within the library or a reference library or something like that. So, you know, as a part of our security process, we do vet new libraries that come in and we do have internal processes that continually check for vulnerabilities and notify teams when they're using modules that are no longer compliant because they may have a critical CVE. So yeah, we really try to secure the entire supply chain from build to deploy to really your running applications as well.

22:03Like these governance controls that meet these standards without sort of compromising some of the speed and efficiency that you get from this serverless development environment. Yeah, I mean, it is certainly an interesting question. There is this natural tension, I think, between developers who want to iterate fast and want to deliver and, you know, want to execute on the, you know, the latest, greatest thing. But I think many of our engineers also recognize the importance of the work that we do. And they're willing to accept a certain tradeoff there because we're dealing with people's money and their finances.

22:39You know, we need to partner closely, you know, as in like in engineering organizations or more developer focused organizations. We have to partner with our developer experience teams, teams that are actually, you know, managing the CI and CD processes. We have to partner with our cyber teams. We have to partner with our open source management teams. There is a lot of coordination. And, you know, I think if we consider, you know, all that's involved there, you know, really there is this effort to try to shift as much of that as possible away from our developers. So in many respects, you know, we try to both shift left and shift right.

23:19So like when we talk about centralized controls, you know, we want to make sure that our CICD pipeline is the choke choke point, you know, the ultimate decider to determine whether or not something can be deployed. Is it compliant? Is it secure? But we also want to minimize friction for developers. So their day to day should not be spent wondering, am I doing this right? Am I doing this securely? And we do things internally. So like one of the nice things we've built, you know, we certainly lean into open source tooling, like I mentioned, open policy agent, we have the ability to determine, like let our developers determine, am I doing things in a secure and compliant manner?

24:00or do I have to wait for something to go to a pipeline and see it fail? Ideally, we want to shift that left. So we use some internal tooling, but we also do lean into tooling like AWS, BFN, Lint. So you can determine like for, let's say, cloud formation oriented deployments, are the templates that you're deploying, are they syntactically correct? You know, is it valid YAML? Is it valid JSON? All the way down to like, you know, are there rules or like, you know, resources that you have to find are the properties that you're specifying valid. You can also lean in and write your own rules to determine like whether or not a template is compliant and, you know, add that to your CFN Lint run.

24:41So really, it's a really delicate balancing act, you know, making sure that we do, you know, provide our developers with the means to be agile, iterate quickly, while balancing the need to be secure, compliant, and well governed. What role does a strong notification messaging system play in this? For messaging systems, you think it's important that developers understand when things change, making sure that they have visibility into what has changed and why it's changed. Ideally, we would, you know, allow those developers to see notifications. Like when we move things, like when we shift things right, you know, when we have that central CICD process, if things, you know, aren't going according to plan, if let's say builds fails, if deploys fail, like making sure our developers understand where the failures occur.

25:31And similarly, you know, actually having that visibility tied back to shift left efforts as well. So like if resource that you're trying to deploy is not compliant, making sure they understand why and making sure that we have supporting documentation to help them understand how to get that non-compliant resource compliant. Do those notifications ever get too noisy? Not going to lie. It can be noisy for our developers. But the great thing is we have a really, really strong developer productivity group internally. and they're constantly looking for ways to improve that experience. Because it's one thing to see that a build failed and you see this huge dump and you wonder, how the heck am I going to troubleshoot this?

26:14They've been spending a lot of time in narrowing down where failures occur, making sure that messages are only as verbose as they need to be and providing that supporting documentation. And also meeting your developers where they are. So whether you're looking at a build log or whether you receive Slack node notification or an email node notification. You know, just making sure that developers understand at that necessary point why things didn't necessarily go according to plan. What role does the Servoist Center of Excellence play in all this? Yeah, so for an organization the size of Capital One, I think it'd be arrogant to say, I know, you know, like what our developer community needs.

Read the full transcript

26:56And, you know, like I'm the only one who can speak authoritatively because I am not. like I am not at all. Different lines of business have different priorities and they have different needs. So really what the center of excellence allows us to do is group people, you know, like people who have an interest in improving the serverless developer experience, letting them come together and share like what's working, what's not, what are the pains, what can be done, you know, like how do we appropriately leverage the knowledge and experience that we may have as a group of domain experts to improve the lives of all developers, whether or not you have that domain expertise.

27:34And also really, I think, work with other teams outside of, let's say, the serverless center of excellence, work with cyber partners, work with enterprise partners, work across multiple lines of business. So really, it's a good platform to receive feedback from the developing community, but also influence how other teams consider the work that we do as developers. And when you talk about the developer community, you're talking about the internal developers at Capital One. Yeah. How big is that roughly? There are about 14 ,000 developers at Capital One. Okay. So a good size community. Yeah. Going back to some of the things that we were talking about in the beginning there in terms of some of the reasons for moving to a serverless for Capital One?

28:26What were some of the challenges with actually putting that migration in place? Yeah, I think there are a few. One, I would say, is plain old FUD, fear, uncertainty, and doubt. I think people for a long time have associated Lambda with, I mean, I hate to say toy applications, but we'll say operations-oriented things that may not be business critical. Lambda doesn't scale like I need it to scale. You know, there's no way I can run my application in Lambda, you know, like initially, you know, five minutes isn't enough, 15 minutes isn't enough. You know, I need more resources. I think developers will be surprised at, you know, the work workloads that can be handled by serverless compute like Lambda.

29:08The other important thing to note too, is that we're really big on saying we're serverless first, but we're not serverless only. There are going to be certain workloads that are not suited for Lambda that are not, you know, like if they're not well suited for Lambda, we would ask teams to look at, you know, serverless container services like Fargate. But even beyond that, like if a service is not right, like we have the means to support other compute and other services as well. But I think the biggest issue is helping overcome a lot of FUD, even helping teams understand what good looks like. So what I mean by that is, you know, Earlier, you mentioned how does observability change?

29:44Well, we want to make sure that teams are empowered to know how their serverless applications are running. When it comes to things like splitting apart the monolithic application, what's the right way to do it? When is the right time to do it? How does your Lambda function scale? How does your application traffic change over time? Are serverless services like Lambda able to keep up with what you need? So really overcoming FUD, but then ultimately just doing the hard analysis to say, like, is Lambda right for you? If it is, great. If it's not, that's okay, too. For those teams that did decide to make the plunge, you know, I think it was initially a struggle to help them understand, like, what the different metrics really meant.

30:27So, you know, for non-serverless compute, a really common metric, let's say for APIs in particular, is, you know, requests per second or transactions per second. And in Lambda, if you look at the unit of scale, it's concurrency, which is really a product of the number of requests that come in. So that RPS, that TPS, but also duration. How long does that Lambda function run for? And the goal is to minimize the amount of concurrency that your function is consuming. With that, it's a matter of helping teams understand, like, what's the right way to impact performance? And in Lambda, really, there's only one knob to turn, and that's memory.

31:06So with Lambda functions, you'll find that both memory and CPU scale together linearly. So like a 256 meg Lambda function has twice the compute and RAM as a 128 meg Lambda function. So for some teams, we would see them come in and allocate 10 gigs of RAM. Like I need all the RAM in the world. When you look at what they're actually consuming, it's like drop in the bucket. Maybe we ratchet that down. But we also saw a lot of teams under allocate the amount of RAM. So, you know, like knowing what the calculation is for Lambda functions, you have invocations. There's a number of invocations. That's a component of the price.

31:48But then there's the amount of RAM consumed and the duration that that RAM is consumed for. and we would see teams allocate like 128 megs. That's the minimum RAM, you know, like that we can allocate. So I'll do that and I'll save all this money. Well, what we would see is teams would actually starve their Lambda functions. So functions would actually run longer because they didn't have the necessary resources. So the great thing is, you know, there are open source tools. AWS Lambda Power Tuner is a great open source tool. If you're not using it and you're a Lambda shop, please use it. Alex Castleboney was a solutions architect at AWS at the time, wrote it.

32:23It is awesome. And it really helps you determine, like, what's that right number based on either cost or performance needs? So, you know, you actually can dial that in a lot more. You're not wasting resources or money. When it comes down to, you know, teams having to make decisions about is serverless the right thing to use versus, you know, using something like, you know, Fargate. What's sort of the framework for making that decision? Yeah, like I would say first consider the AWS constraints. Like, do you have a function, like do you have a workload that needs to run consistently for more than 15 minutes?

32:59If so, you know, Lambda's not right for you. Do you have the need to consume more than 10 gigs of RAM? If so, Lambda's not right for you, right? Like there's some easy ones to look at, but it's interesting when you consider then too, like, you know, beyond those obvious constraints, looking at things like, well, what is your Lambda function being triggered by? And what I mean there is like, you know, Lambda is event-driven compute. It'll be, you know, your code is, you know, is triggered in response to an event. And AWS has, you know, over a hundred event sources that can trigger Lambda functions.

33:33But let's say you're using, you know, you want to write a Lambda-backed API and you're using, let's say, ALB, Amazon's application load balancer service. Well, that's a synchronous indication of a Lambda function. Now, Lambda can handle a 6 meg payload size at this time, but ALB can only pass in 1 meg, and 1 meg in, 1 meg out. So if you're writing an API and you need to handle more than 1 meg in or 1 meg out, Lambda may not be the right choice if you're using ALB. with a service like Amazon's API Gateway REST service, that number jumps up to the full six megs. So API Gateway can handle 10, 10 meg, but you're going to be constrained by the six meg lambda limit.

34:19So like, look at the obvious things like, hey, you know, like what are the physical constraints? Consider what you're looking to integrate with. Beyond that though, it really gets interesting. One common thing that comes up that would drive someone away is, you know, I have a Java application, You know, Java will never run well in Lambda. Well, I would ask you to revisit your assumptions. You know, there are ways to mitigate things like cold start pains. Consider how often things like that, you know, those cold starts happen. Like, is it something that's happening a lot or not? If it's, you know, if it's a synchronous invocation where you have a user on the other end of that request, that cold start may really matter a lot more than an asynchronous invocation.

34:58Like someone uploads an object to S3, well, you may not have somebody waiting on the other end of that request. So if you have a cold start that takes a little while, so what? It may not matter as much. The other thing, too, is the AWS Lambda team has worked really, really hard over the years to improve that developer experience for, let's say, job in particular. So you can use services or capabilities like provision concurrency, where you essentially have an AWS management capability that will keep a certain number of Lambda execution environments warm. So you won't have those cold cold starts for the number of provision concurrent units that you've set.

35:37Last year, Amazon also introduced a capability called Snapstart for Java. And I'll be honest, I need to revisit the exact Java version. but I want to say Java 17 and higher, I need to double check. But this year they actually introduced that for Python as well. It was a pre-invent announcement. So, you know, AWS is looking for ways to improve that developer experience and help developers, you know, force them into making a tough choice. Like, is Lander right? Where do you think re-invent's going on this week? There's lots and tons of announcements. Like, where do you think, you know, serverless is going in the next couple of years?

36:14Interesting places. So like one of the more interesting announcements I heard this week was on DSQL Aurora. We have like that, I think is going to be massive in ways that we can't yet comprehend. That's one. I think durable workflows will be another area. So like if you consider how like how important it is for certain workloads to run to completion, you know, when you have ephemeral compute services like Lambda, state becomes really important. So like, how do we, how do we do that for a really important look? That's going to be another one. I think continuing to manage the, the soft software supply chain is going to be really important too, and providing visibility into that.

36:55And I think the like last, I would say improving the, the operator experience. One thing that I, that makes me shudder is when people say serverless is no, no ops. It absolutely is like, it doesn't absolve you from your operational responsibilities. The good news is, as your cloud provider, whether it's AWS, Google, Azure, or anybody else, they're assuming more of the responsibility. But it doesn't mean that you don't have any responsibilities. So, like I would say, I would look for ways to improve the understanding of what's happening in applications. So, the ability to observe what's happening is going to become really important, even more often.

37:33Yeah. Well, I know we're coming up on time here, Brian. I want to thank you so much for coming on the show. I really enjoyed this. Yeah, Sean, thank you so much for the invitation. I really enjoyed the conversation. Cheers.

From the publisher

Serverless computing is a cloud-native model where developers build and run applications without managing server infrastructure. It has largely become the standard approach to achieve scalability, often with reduced operational overhead. However, in banking and financial services, adopting a serverless model can present unique challenges. Brian McNamara is a Distinguished Engineer at Capital One where

The post Going Serverless in Financial Services with Brian McNamara appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Going Serverless in Financial Services with Brian McNamaraSoftware Engineering Daily · 38 min
Listen in VO