In short
AWS Podcast Episode Notes
Episode Title
#726: Single region, zero excuses: Mastering AWS resilience Release Date: June 23, 2025 Hosts: Jillian Ford Guests: John Formento and Tarik Makota
Episode Overview This episode delves into the intricacies of single-region resilience using AWS services. Experts John Formento and Tarik Makota address common misconceptions, outline strategies for robust application building, and provide actionable tips for identifying and mitigating failure modes in cloud applications.
Key Themes
Understanding AWS Terminology
- Regions vs. Availability Zones (AZs):
- Availability Zone: A physical data center or multiple data centers grouped closely together.
- Region: A collection of multiple AZs, designed to protect against geographic events.
Fault Isolation Layers
- Multiple Layers of Fault Isolation:
- Partition Level: Separation for GovCloud and commercial AWS regions.
- Regional Isolation: Each region can fail without impacting others, similar to a bulkhead in shipbuilding.
- Availability Zones: Further divides regions, providing additional fault isolation.
Misconceptions About Resilience
- Single Region vs. Multi-Region Architecture:
- Single-region setups can achieve high availability if designed correctly.
- Multi-region setups may introduce complexity and dependency issues if not managed properly.
Designing for Resilience
- Key strategies include:
- Understanding Control and Data Planes: Knowing the difference helps in building resilient architectures.
- Failure Mode Analysis: Identifying potential failure points is crucial for resilience planning.
- Scalability and Load Management: Implementing strategies such as load shedding and throttling to maintain service during high demand.
Actionable Strategies for Resilience
- Recovery Procedures:
- Regularly test and exercise recovery processes to ensure effectiveness.
- Follow the AWS Resilience Analysis Framework for identifying common failure modes.
- Critical User Journeys:
- Focus on critical paths in the user experience; prioritize resilience in areas that directly impact user satisfaction (e.g., e-commerce checkout processes).
- Infrastructure Choices:
- Evaluate whether to implement high availability (quick recovery) or disaster recovery (longer recovery time), each requiring different architectural approaches.
- Service Utilization:
- Leverage AWS services designed for resilience, such as:
- Amazon Application Recovery Controller: Helps manage application traffic during failures.
- Fault Injection Service: Allows testing of application resilience by simulating failures.
Advice for Customers
- Rethink Resilience:
- Shift focus from aiming for a specific number of availability 'nines' to understanding and managing potential failure modes.
- Continuous Improvement:
- Treat resilience planning as an ongoing process, not a one-time task. Regularly reassess architecture and operational strategies as business needs evolve.
Conclusion The conversation emphasizes the importance of understanding AWS's architecture and services in building resilient applications. Both guests encourage listeners to approach resilience planning with a critical mindset, focusing on failure modes and the impact on user experience.
Further Reading
- [AWS Resilience Analysis Framework](https://docs.aws.amazon.com/prescriptive-guidance/latest/resilience-analysis-framework/introduction.html)
Key Takeaways
- Multi-AZ architecture is foundational but not sufficient alone for critical applications.
- Businesses must actively question and assess their resilience strategies and infrastructures.
- Continuous testing and improvement are essential for maintaining resilience in cloud architectures.
---
This markdown entry provides a comprehensive overview of the podcast episode, summarizing key discussions and advice on AWS resilience strategies.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00This is episode 726 of the AWS podcast, released on June 23rd, 2025. Welcome, everyone, to the AWS Podcast. I am your host for today, Jillian Ford, and I always love episodes that apply to every person, every single listener, and this episode is one of those. We are going to be talking about single region resilience, so stick around because there's going to be a takeaway for you wherever you are at with your single region resilience journey. And I've got two experts on the topic, which I can't wait to pepper them with questions for selfish reasons, of course. So let's do some introductions. John, why don't you introduce yourself?
0:50Hi, everybody. I'm John Fermento. I am a principal product manager in our Resilience Infrastructure and Solutions organization. I focus a lot working with customers who are wanting to operate critical workloads on AWS. and then I also helped build Amazon Application Recovery Controller.
1:10Hey, I'm Tariq Makoda. I'm a Senior Principal Solution Architect, which means I spend most of my time working with customers on the resilience aspect. I've been at AWS about seven and a half years and I'm also an extended member of the team we call internally Customer Resilience Engineering. Super cool. I can't wait to ask you all these questions. So let's start really simple because we've got people who are brand new to AWS. So Tariq, can you explain the difference between a region and an availability zone? How much time do we have? All right. 30 seconds. Yeah. Both are basically logical constructs under which are the actual physical implementation.
1:58What I mean by that is availability zone is one or more data centers that are closely located. And then multiple availability zones form AWS region. AWS region generally is give you protection against the events such as electrical grid feeders, such as nature events, things like that. regions are generally spread around geographically or about 60 miles and these two constructs are heavily used to create resilient and reliable applications. Love it. And I know a lot of people listening are like, okay give me the stuff to be able to implement right now but I think it helps to really have a better understanding of AWS's fault isolation boundaries before you can start already implement some actual tactical things into your architecture.
3:01So, Tariq, like, when you're talking to customers, can you help them explain, can you explain how they think about default isolation? Yeah, so there's multiple layers of default isolation, so I'll stop from the top and kind of navigate down. I presume we have listeners who are both in public sector that would be utilizing our GovCloud, and then we have the customers who are in enterprise companies basically utilizing what we call commercial cloud. So the first separation when it comes to the fault boundary is what we refer to as a partition. So in this case it would be partition for the GovCloud and it would be partitioned for the commercial AWS regions.
3:44So that's the first layer of isolation. Then under that layer of isolation there's a regional isolation. So we basically every region itself can fail and have an impairment without impairing other regions. Meaning, this is basically, if you think about it, very similar in shipbuilding to the bulkhead pattern, where each region is a part of the bulkhead, and flooding one of the bulkheads does not basically take the whole ship down, aka other regions are not impaired. After the regions, the next step, the next fold boundaries, the availability zones. Availability zones themselves are just basically extension of that bulkhead pattern.
4:30So if you think about it, region being a part of the bulkhead, then separating that bulkhead additional layers is basically what the availability zones are. So those are the basically basic construct. Diving deeper, when it comes to the protection, there are services that are at the regional level, there are services that are at the AZ level. So for example, EBS would be availability zone called boundary type of a service, EC2, similar. In addition to that, some services implement what we refer to as cell-based architecture, which is another layer. Well, before the cell-based architecture, a number of services implement concepts of control planes and data planes.
5:22Control planes very simply are used to persist and make the changes to the service. The data planes actually carry the workload of the service itself. An example of this would be spinning up the EC2 instance, basically starting up EC2 instance, whether it is through the autoscaling group or a start instance or watching through the console would be the control plane function, meaning I'm instantiating something brand new two EC2 instances communicating to each other would be the data plane operation of the EC2 data plane. After those two constructs, a number of services implement cell-based architectures, which are probably out of the scope of this, but it's another layer of isolation.
6:08We're the largest set of fleets. If you think about the data plane, data plane might be composed of the fleet of servers that might be in tens of thousands, and they logically get separated in the smaller groups called cells and that contains the blast radius as well. So you just, as a customer, right, like you explained a lot to me about, you know, different boundaries AWS has and then got into some of the plumbing about how services, you just think about data plane and control plane, like why does it matter? What can I do with that? The reason it matters is knowing how to use these contracts is important in your own application.
6:49So to give you an example, if I am to take my application from my on-prem data center, my fault isolation boundary at that point is my application, a.k.a. my data center. I take that application and transpose it to AWS, and I stretch that application across the three availability zones. Each availability zone is the fault isolation boundary. However, since I'm crossing those isolation boundaries, I'm basically poking the holes into the bulkhead pattern. So I need to either modify my application to be availability zone aware, have availability zone isolation, or availability zone affinity. There's a number of ways to do it and to take the advantage of these things.
7:38Also, in terms of when it comes to certain failure scenarios, some services. If I am relying on a control plane functions to basically repair my application after the impairment, those control plane functions may or may not be available. So understanding what is available, what is not available, what are the constraints basically help me construct my application in such a way where I can have the highest availability possible. Is that what you were looking for? You wanted specific details. I think that's something. And the reason why I poked a little bit is because when I'm talking with customers, there's usually like a so-what moment that happens of like, great, John, you're telling me about control plans, you're telling me about data plans, and you're telling me about this Zono contract.
8:41And what's worth, the way the AWS infrastructure is built is a very unique property of AWS, right? Where we have this regional isolation and something happening in one region shouldn't have an impact in another region. And then the isolation between availability zones gives customers very powerful mechanisms to build resilient architectures. And then so, like, again, so what is, as a customer, when you're building for, you know, let's say, you know, a highly resilient application in a single region and using a multi-AZ approach, understanding the different, like you have to think about it from like an API perspective, not just a service, right?
9:21So if you think about EC2, there are some actions which are going to be control plane. And what you don't necessarily want to do is if something is going wrong, dependent on launching instances, right? For your most critical apps, that's a very complicated process. And that's part of EC2 control plane, where it's attaching ENIs and doing a bunch of things within the process of spinning up an EC2 instance. But interfacing with already running EC2 instances, such as that HTTPS connection or whatever port or protocol you're using, is a data plane action. And that's kind of the nuts and bolts of regular operating of that service.
10:03And so that's just one example, but even thinking about a database, spinning up another replica as a mechanism to recover, that's a control plane operation of spinning up a replica. Let's have that replica there already and do a SQL promotion, which is more of a data plane operation on an existing data store. so kind of you need to understand how those relate to the services you're using to make informed choices on a resilient software architecture yeah you bring up some really interesting points because i when i talk to customers i work with startups and i think a lot of them they see exactly what you were saying about all the what aws does to really have make it easy for customers to build these resilient architectures.
10:53And it gives people, unfortunately, a false sense of security that they don't really have to do much in order to be resilient. Like, oh, I can just use like a managed service, for example, and then I'm good. So I think it's great that you were kind of unpacking what are some additional things that people should think about. And I'd love to understand more on that thread of like, what are some other misconceptions that maybe you see customers have when they're thinking about designing for single region resilience? I'll go first. I think the biggest misconception is that single region architectures cannot achieve high availability or resilience or otherwise said, multi-region architectures can achieve the higher resilience.
11:47there's a lot of ifs in that statement but it's technically if not done appropriately multi-region architecture could have lesser availability. Example of that being if I am synchronously replicating from region A to region B and if I have dependency on that synchronous replication That means if any one of the two regions is unavailable, my whole application is unavailable. Statistically speaking, two regions may have impairment more often than single region, just as a statistical thing. So I think one of the things, that's number one. I think the other part of that misconception is usage of availability zones.
12:42So availability zones are physically isolated with independent power, independent cooling, networking infrastructure, right? But in order to use them correctly, application needs to be distributed and architected correctly. I spoke a little bit about this earlier. if I spray my traffic across all three availability zones at any point, meaning my traffic increases in availability zone 1 for some reason I send it to availability zone 2 because that's where the service is and then that goes back to availability zone 1 and then to availability zone 3 and then to the database I'm basically crossing those zonal boundaries multiple times, which may be fine if I have ability to detect that any one of those availability zones is impaired and have either circuit breakers or some sort of other mechanism that basically would isolate that availability zone from my application standpoint.
13:48So just running application on multiple availability zones does not mean that that application is highly available and can use all of the benefits of availability zones, an application needs to basically prepurposely ensure to take advantage of those. So that, to me, the biggest misconception, I believe, is that going multi-region by default gives me more availability and resilience. I think how the application is architected, how the failovers work is just as important. I think what you were just saying there, even within a region of how to architect for multi-availability zones, I think that's going to cause a lot of people to really think more critically about how they're designing it.
14:45So I'd love to understand, John, how should customers start to think then about resilience for their critical applications? Yeah, that's a good question. And maybe I'll tie back to the example I was given around the control plane, data plane. And if you would just hear that snippet, right, of like, hey, you should rely on running instances only, it's the wrong message, right? And so when you think about critical applications, you need to think about, there's normal operations you're going to do, right? Like naturally as load increases, you want to have scaling to be able to absorb that load. And you're going to naturally rely on various control planes to do that.
15:24And that's a good thing, right? As more customers are using my service or product, I want to be able to scale to absorb that load and then scale back down to, you know, a predefined kind of baseline to save costs, right? Now, with a very critical workload, what you want to think about maybe is like that baseline of like, what is the minimum we need at any given time to service our clients? And so then what in my mind, it now becomes, you need to really separate like a conversation around critical workloads around like what are specific recovery operations we have. A lot of times customers will think about these in terms of like run books or, you know, SOPs for recovering from various failures and things like that from like the normal state of the application.
16:08And so for like recovery procedures, you really need to double click on for these critical workloads. And that's where you want to make sure like things we're doing are not introducing like a new modality. This would be one of Tariq's favorite of like bimodal operations. And what we mean by that is what you want to avoid in a critical system is as a part of recovery, going down a path you never test or normally don't go down, for lack of a better word. And what we've seen, you know, more times than not is there will be this kind of fork in the road, so to speak, when something's going wrong and you go down this new path and you try to do things to recover that you have never tested.
16:45it. And eventually, typically that backfires because you need these recovery processes to be tested and you want them to be, you know, you know, just regularly exercised. So when you use them, it's going to work as you expect. It's, you know, I think that's a big thing to think about. And it also goes into like the one thing, you know, we talk to customers a lot about is understanding failure modes and maybe doing like a failure mode analysis. Like for folks out there who are familiar with like FMEA, failure mode effects analysis, it's kind of like our version of that, so to speak, where we just kind of think about it through a lens of cloud-based systems.
17:25And the idea is, if you look at any application, there are common failures that just occur. And we've categorized these and just maybe a plug to the document that's available for customers. If you Google or search wherever, like the AWS resilience analysis framework, you'll find this. But we basically categorize through internal learnings, working with customers, common failure categories or modes. And you think about shared fate, excessive load, excessive latency, misconfiguration, bugs, and single points of failure. And so if you look at your application, and it's not just components. Try to break it down to, say, a user journey or story.
18:07And an example of this may be for an e-commerce site. Path to purchase, you'll hear that a lot. Like what's the path my customers go to actually complete a purchase is a critical path. More so than maybe me being able to update, you know, my, you know, address or phone number, right? Like it's okay if I can't do that right this second, but most likely I want to purchase something and I want that to be available. And for like a banking application, one of the most common critical paths is like being able to check my balance, right? Like I want that level of trust with that financial provider that can go and I can see how much money I have.
18:42Like I worked hard for that, right? And so if you think about it as a frame of there's these critical user journeys, and there's interactions between components, in this case, maybe AWS services of where, you know, say like an API gateway interacts with a Lambda that interacts with the DynamoDB table. You can then look at, well, what are common failure modes that you see? Maybe excessive load is one of those, right? What happens if a client all of a sudden is misconfigured and is bombarding my service and overloading it to the point where I can't handle all these requests? What can I do? And maybe I'll send this to Tariq.
19:17What would you do if your application was being bombarded by requests? What's a mitigation you could do? Well, there's a couple of them. they're not popular but um you know the the load shedding would be the one of the things um you know the load shedding meaning if i'm getting overwhelmed with requests um i have a decision to make all right decision is am i going to load shed some percentage of the customer or the request um and am i going to impair let's say 20 of those or am i going to allow this to go on until the resources of application are exhausted and I potentially have full-on failure of my application aka 100 % of the customers are affected So in that case it may be better to give one client a bad day or a bad 20 minutes and preserve other customers or clients in my system.
20:17Yeah the analogy that I like that all of us can relate that Mike Haken actually wrote the paper of this in a builder's library it's like when we all go to the restaurant sometimes we go in there and we are told hey 15 minutes wait 20 minutes wait we are basically being throttled right so they could take all of us in the restaurant and then i could be sitting in john's lap with his family eating the dinner at the same time experience is bad for both john myself or they can allow john to finish his dinner open up the table and then i go have my dinner so um that's kind of like an everyday example of these things um so definitely you know load shedding throttling as well um throttling based on number of the customers and i could treat customers differently um of my application if i have a customer that needs the high level utilization and is an important customer i might give that customer more tokens and more space to use in my application versus other customers, or I can distribute it evenly.
21:26So traveling the customer requests. The last part of it is, you know, the load shedding also allows me to sometimes, I think in the e-commerce applications, it's about four seconds for the user abandonment, meaning the user will wait for about four seconds before they abandon their cart, purchase, whatever. so for example if i'm having some sort of a slowdown or impairment um where the record responses was being returned to the customer you know after four more seconds i have a decision to make i already know that there's going to be abandonment by the customer that's likely um i have projected that the request might take more time that that customer is willing to wait and if i proceed with that work that's going to be wasted work to that point So I have a decision to make, should I do that work or not, given the time limit.
22:21So those are some of the techniques to use. Exponential backoff is another one. Or I can just stop certain things from working using the circuit breakers. Depending on the impairment that I'm trying to handle, there are a number of techniques that could be used. A lot of these things need to be implemented in an application. Some of AWS services can help with these things. Safety at Gateway can help with rate limiting and things like that. ALB can help with making the load shedding easier. It's not out of the box, but I can put the path for the load shedding, kind of quote-unquote black hole, the traffic.
23:04Long story short, knowing how these services behave, what the capabilities are, and implementing the functionality in my application basically give me the best chance of having a high level of availability. One theme that I'm hearing from both of you is that it sounds like customers need to really rethink how they are thinking about resiliency. A lot of people will work backwards from, oh, I need a certain number of nines, where it actually sounds like, based on what you're both saying, is that they should really be thinking about it in terms of failure modes, like all the scenarios you were talking about of different points of failure and how you want to be able to respond from that failure.
23:49So, John, I'd love to hear from you of walk us through how customers can start to make that shift as part of their resiliency strategy. Yeah, sure. And, yeah, I think, like, you know, number of nines, you know, like I want 99 or measuring like against service level objectives, SLOs. It's really common. And I'm not saying, you know, change the way you measure, like that makes sense. But like what we like to see or what is typically helpful when we talk to customers is thinking about how to like, what's the means to get to that end. And it's typically looking at your application and figuring out like, hey, yes, these three failure modes are pretty plausible.
24:34and everything's a trade-off with resilience and to your point earlier like like maybe like it's not uncommon for customers to say like yeah i'm using these you know man services and they're multi-az so i should be good and quite frankly that's a really good baseline to start you can build some very resilient applications that way um and so like i guess to for like the audience and for like this conversation like tarik and i are really coming through the lens of like the most critical apps like you know if flights may not take off transactions may not clear like things that are really going to be impactful um and and i'm sure every business has some of these and typically and my point is typically it's not every application within a business um and so that the you know when you think about these things you have to make trade-offs um around how much you invest in the resilience um because i'm sure for every application team there's a counterpart in the business that's saying we need these features.
25:29And then right, like, how far do you take this thing? And so, yeah, with that said, like, with every kind of failure mode, there's a trade-off. And we like to think about it in terms of, like, what's the plausibility that it could happen? How realistic is it? And then if it did happen, what would be the impact? And so when you think about the different categories and how they can manifest into failures for your application, you know assess it against those two things and if you kind of think of and like obviously like the high highs like high probability high impact you probably want to take care of those but now if you get something to like well hey the risk is super low but the impact may be really high that's probably something you want like that's a harder trade-off to make and then right as those ranges or inputs kind of change you can make that assessment but you really want to make that assessment against you know the engineering effort and the cost of implementing because like Tariq's talking about throttling and circuit breaker technology and, you know, creating cell-based architectures, which is something we do.
26:29Like when we're building services at AWS, you will not find a conversation that doesn't incorporate concerns or conversations around blast radius reduction. But, you know, that's an engineering tradeoff to make. And so as a customer, you really want to think about the level of complexity you're building into the system and does that, you know, get you what you're after. And then right now it's back to the end of the day of like what availability do your customers expect? And I think most customers nowadays expect everything to be available whenever it is the heck they want to use it. So it's a challenging environment, right?
27:08Totally. Yeah, it really is. And I'm sure there are other customers as they're like listening to this and it's getting them thinking about what the failure modes are. There's probably people who are like, wait a minute, but I don't know what I don't know. So is there some sort of list that AWS has curated that can help me see what are some of the common failure modes that maybe I should think about? I almost want to just tell everyone your alias really quick, but that's probably not going to do you any good. Maybe you tell them where they can find it.
27:43I've actually been playing with AI a lot. To be honest with you, I was playing with QCLI the other day and I asked about the failure modes for API gateway. but it gave me 14 failure modes. And then I asked what are some of the mitigation factors for it and gave me those and some of the links to the blog. It was fairly accurate. So when it comes to, I would say that each, I shouldn't say each, most of the services have probably resilience section in the documentation where they're going to talk about some of these things.
28:29The Gen AI, the AI could definitely help because it adjusted so much public data, so it's like a quick way to kind of get like an 80-20 type of a rule. But we've been talking this before. While we've been focusing here how AWS services can fail, my own application has failure points that these things will not tell me about it. So for example, most of the application have dependency on some sort of a single sign-on system, things like Okta, Ping, and so on and so forth. I as a customer could implement these things myself in my own infrastructure. I can self-manage them and things like that. At that point, I'm taking responsibility to making these highly available.
29:20but my application might be dependent on single sign-on to be there. Or, as John was mentioning before, thinking about critical paths. If I have a travel application and that travel application is giving me turn-by-turn directions, and in addition to that it's going to tell me what the weather is in the city or town that I'm going to arrive to, which one is more critical? Is it critical that I know the weather in the city that I'm going to arrive? I would say not as much. It's nice to have. So if that's a dependency on weather.com or weather service of some sort, I may want to basically make a choice that when that dependency is not available, that basically my application still operates.
30:06So I have to take ownership. A lot of times when it comes to how I build my application, So for example, if I have built my application in such a way where the application has availability zone isolation, when I do deployments, I might have ability to deploy in a 1AZ out of 3, or I might have ability to use the feature flags to deploy my function only to the subset of things. And I can test that, and if there's issues, I basically roll back. It is more common that deployments cause, or the config changes through the deployments, cause an outage than anything else. That's a very common thing, even with all of the testing.
30:53So there is no one single way to get all of these, all of the potential failure modes. But one of the things that I would suggest to the team and some of our internal teams do this on a weekly basis as the team goes through the backlog, they might come together and basically say, what happens if the power goes out? This can be done even if my system runs on on-premises. What happens if the power to the data center fails? And then I would reason about all of the things that actually happen or what happens if the mainframe system fails. And then I reason basically to all the potential impacts that I have.
Read the full transcript
31:37Once I have those impacts, it will tell me the significance to my business. So, for example, if I'm running the re-scrolling service that is not super important and the mainframe fails, I may say, okay, I'll be down for a few hours if I'm just going to recover, versus if I'm running some sort of a critical payment system, etc., or i may have to have a high level recovery so what we what we haven't really talked about is like very basic choices like the first choice that i need to make it is you know whether my system requires disaster recovery or high availability and it's a huge difference as far as of how i'm going to architect how i'm going to recover um analogy to this would be almost kind of like, am I going to build an airplane, aka high availability, or am I going to build a car, a car with a disaster car?
32:38Example being, the way an airplane moves is by having the jet engines. Most of the airplanes, commercial ones, at least have two jet engines, and they have everything else pretty much redundant. So they can operate on single engine, meaning during the impairment and the failure, there's a high availability because there's a redundancy. In terms of the car, if my tire actually gets deflated or impaired, I basically have to pull over. I have a spare tire in the back, so I'll take the 30, 40 minutes to replace the tire. That would be, quote-unquote, disaster recovery. So making those choices and knowing what's important is critical.
33:23Also, a lot of applications have recovery time objectives, recovery point objectives, which technically should be defined by the business. Basically, it's continuity and disaster recovery documents. What I find quite often is that we as a technologist have a much more aggressive recovery time objectives, recovery point objectives, than the businesses necessarily need. Sometimes this is just due to the lack of miscommunication. Sometimes it's just for the other reasons.
34:05But deciding disaster recovery or high availability would be the first step in a lot of these things. I would say disaster recovery is probably a little bit less complex to handle. having high availability we're talking about being able to recover within minutes not necessarily hours that is much more difficult than it may require different things from my application simple example if my recovery time objective is 12 hours I may just decide not to do anything because statistically any impairment that actually happens would probably be done at 12 hours. I can wait for six hours before deciding that I want to do disaster recovery in the secondary location, whether it's an AWS region or something else, whereas I may not have that ability if I have a critical application.
35:06In a critical application, the concept of airplane is then questioned around high availability versus the cost. So cloud being provision on demand and matched consumption, we match the demand with a number of the resources. The question then becomes, if I use an airplane scenario and if I'm running in two availability zones, if I lose one availability zone, assuming that both availability zones are utilized, almost, let's say, 80 % each, if i lose one that means that traffic from the one that i have um what the quote lost once i actually move to the other i'm going to have 160 of the traffic going to the single az and i may not have the resources so what john was talking about earlier is the concept of static stability i should have excess capacity in my availability zone an excess capacity meaning that means extra compute that might be idle.
36:17And in this term, if you think about the airplane analogy, it's definitely cheaper to build up airplanes with one engine, but then the airplanes are so critical. If I lose one engine, basically everyone on board perishes. So if you use the same concept with the critical application, if I lose one availability zone, it's critical application, I might have reputational risk. I might have a financial impact on things like that. And if a financial impact is, I'd say if a financial impact for 30 minutes of outage is$20 million, and having static stability, having the excess capacity, is going to cost me$10 million a year, I will choose that probably every time.
37:00So it's a trade-off. So I think those basics also come in play. and then after the things, once I have decided those things, all of the things about the failure mode analysis go. But I would actually suggest everyone plays with the AI. It's actually pretty good at identifying, I would say, 70 to 80 % of these failure modes. Wow, that's a really good call out on QDeveloper. So what are some other services that can help customers in AWS build resilient applications? shameless plug here Amazon application recovery controller can help so we have capabilities called routing control to help ship traffic between regions and then for more applicable to this conversation with single region architectures with multi-AZ we have two really awesome capabilities one called zonal shift both of these we're about to talk to zonal shift and zonal autoshift integrate with EKS clusters application load balancers, network load balancers, and EC2 autoscaling groups, but they allow you to essentially say, hey, something's going on in this availability zone, shift my work out of that AZ.
38:09And that's a pretty unique capability to AWS. Then there's also the fall injection service, which we should have talked more about how critical testing is. I remember I mentioned you need to test your recovery procedures, but testing's paramount to success when it comes to these types of things. And so fall injection service allows you to, they call them actions, but allow you to inject certain kind of symptoms of various failure modes and test how your application handles those types of things. And so, yeah, I mean, there's a couple, I'll call it just a doc again, like Tariq talked about using Q.
38:47if you rather just read something, you can also read that doc I mentioned, the resilience analysis framework that will guide you through that process and give you some ideas on different failure modes. Yeah, I would just add to John's. John kind of mentioned the services that are specifically intended to help me as a customer to either test or validate or move my application when there's an impairment. I think the other aspect of this is we talk about the shared responsibility model in terms of resiliency. And different services have more responsibility on AWS services versus others. To give you an example, if I am using Connect, Connect is pretty much I just have to kind of define and configure the Connect and then everything else, compute, connectivity, etc., is managed by the Connect team.
39:46So the Connect team is taking more responsibility from the resilience perspective than I am. So this makes my resilience easier. If I'm running containers, I have multiple choices. I can run containers on my own. I can use any one of the container services and manage my own fleets, or I can have the container services manage the fleets and node workers. What this does in that analogy of static stability, if I use, for example, ECS Fargate or EKS Auto Mode, the new service that was released, then I'm basically shifting the responsibility on that capacity management and having enough capacity for me to those services.
40:30Lambda would be a similar example. I can run the code on EC2 instance, or if that's some sort of request-response application that basically can be moved to Lambda, Lambda is going to actually manage provisioning of those instances, the capacity of those things, and I just have to provide the code. So what I would say is most of the services give the customer some level of resilience boost. In some cases, as a customer, may need to utilize those services and know how they actually operate and basically build my own resilience and application. In other cases, I might be able to remove that heavy lifting to that service like Amazon Connect and basically not really, you know, have a minimal responsibility when it comes to the resilience of that replication.
41:30So it sounds like there's really, I would say, a lot of things that customers should think about. not just on like the framework of how they're thinking about the resiliency strategy, but also the services that can really be able to help them. I would say that's a huge takeaway for me from this conversation. So one last question for both of you is just like some parting advice. So, John, do you have any parting advice for customers that are thinking about their multi-AZ resiliency strategy? Yeah, I think multi-AZ is a great place to start. Just out of the box, you get so many benefits. And so I hope that didn't get missed in today's conversation.
42:15Out of the box, multi-AZ is the best practice and you can build a really reliable, resilient application with that. But for critical workloads, which, again, if you look at a customer, not every workload is usually a small percentage. There are things you can do that we talked about here to really get that next level of resilience, again, still within a single region. And that kind of saves you, like Tariq mentioned, some of the complexity of going multi-region, which I think we may talk about at another time of what that all entails. But, yeah, that would be just my kind of parting thoughts. Let me see.
42:50I would say question everything. So if I was a customer, I would question recovery time objectives, recovery point objectives. Like what I am being told is that accurate? Like does my business continuity plan, does it match? Do my regulators actually asking for the level of the resilience that I'm basically trying to implement? Like is that adding additional complexity? The reason I said question everything from the beginning was because the more cogs in a wheel, the more likely are the failures. So I want to simplify my architecture as much as possible. To that same thing, having 30 services in my architecture might be cool, but those are the 30 failure points, potential failure points.
43:44statistically speaking we're talking about the regions of this like if i have 10 this actually you know that will do less so um i would say question all the choices along the way and then that's what then leads into the failure points the failure points that we talked about in resilience analysis framework is basically questioning myself how can my application fail and what other things that support my application would fail and how that actually so um i guess it requires There's a lot of scully and questioning everything along the way to optimal results. It may be obvious. One last thing, right?
44:25It's not a one-time deal, right? You don't do it once and you're done or you spend a month doing it and you're good forever. Think about this as a continuous process.
44:38Totally. And especially for some customers that maybe they're on a single region right now and they're thinking, oh, maybe I actually do need to use multiple regions, which, yes, as John was hinting, is going to be another episode. So you'll definitely want to make sure that you are subscribed to the AWS podcast so you can be notified when that is available. but thank you so much John and Tariq for being here on the 8 of his podcast this was super interesting I learned a lot I'm sure the listeners did as well thank you appreciate it hopefully we didn't scare anyone well hopefully scared them enough that they're actually going to start looking at the resilience so that that was my secret motive there you go hopefully mission accomplished yes awesome thank you so much everyone.
45:30Awesome. Take care everybody.
From the publisher
Dive into single-region resilience with AWS experts John Formento and Tarik Makota as they debunk common misconceptions and share practical strategies for building robust applications. Learn why multi-AZ isn't enough on its own, discover key AWS services for resilience, and get actionable tips for identifying and mitigating failure modes.
Learn More: https://docs.aws.amazon.com/prescriptive-guidance/latest/resilience-analysis-framework/introduction.html
