#739: Applying Multi-region Strategies on AWS

29 Sep 2025 · 41 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AWS Podcast Episode #739: Applying Multi-region Strategies on AWS

Release Date: September 29, 2025 Host: Jillian Ford Guests: Jeff Ferris, Scott Wainstock Episode Link: [Listen Here](https://docs.aws.amazon.com/r53recovery/latest/dg/region-switch.html)

Episode Summary In this episode, the AWS Podcast focuses on the essential topic of multi-region architectures, which is critical for businesses as they expand their operations and seek resilience in their infrastructure. Hosts Jillian Ford, along with experts Jeff Ferris and Scott Wainstock, share insights on best practices and strategies for implementing multi-region frameworks on AWS.

Key Concepts and Discussions

  1. Introduction to Multi-region Strategies
  2. What is Multi-region Architecture?
  3. A strategy where applications are deployed across multiple AWS regions for redundancy and low latency.
  • When to Consider Multi-region?
  • After mastering multi-AZ (Availability Zone) setups within a single region.
  • Driven by business needs such as compliance or geographical expansion.
  1. Prerequisites for Multi-region Implementation
  2. Understanding AWS Regions:
  3. Physical and logical separations between regions that require careful architectural planning.
  • Isolation Requirements:
  • Applications must be architected to maintain operational integrity even if one region experiences failure.
  • Operational Planning:
  • Importance of documenting processes and understanding application dependencies.
  • Coordination among engineering and business teams is crucial.
  1. Challenges and Solutions in Multi-region Operations
  2. Common Challenges:
  3. Difficulty in coordinating across multiple teams and accounts during operational events.
  4. The complexity of managing dependencies among resources.
  • Strategies for Success:
  • Establish a solid communication plan.
  • Implement regular operational reviews and practice failover testing.
  • Foster a culture of resilience within teams.
  1. New Tools and Services
  2. Amazon Application Recovery Controller (ARC)
  3. A service designed to facilitate recovery from outages in both single and multi-region environments.
  • Region Switch:
  • A new service that orchestrates recovery for multi-region applications, simplifying the recovery process with recovery plans and workflows.
  1. Best Practices for Testing Multi-region Strategies
  2. Testing Approaches:
  3. Regular testing of active-active architectures is easier than active-passive configurations due to ongoing activity in both regions.
  • Investment in Testing:
  • Emphasize the importance of testing frequency in line with business requirements.

Conclusion and Recommendations

  • Actionable Takeaways:
  • Understand your application architecture and build through a resilience lens.
  • Consider the trade-offs between cost and availability when designing multi-region architectures.
  • Utilize AWS resources such as the Well-Architected Framework for continual improvement.
  • Further Learning:
  • Listeners are encouraged to explore additional resources and reach out to AWS experts for tailored advice on multi-region strategies.

Contact and Resources

  • AWS Application Recovery Controller Website: [Visit Here](https://aws.amazon.com)
  • Well-Architected Framework: [Learn More](https://aws.amazon.com/architecture/well-architected/)
  • Follow AWS on Social Media: Stay updated with AWS announcements and best practices.

Episode Notes

  • The podcast emphasizes that building a culture of resilience and operational muscle is crucial for managing complex multi-region infrastructures.
  • Both Jeff and Scott highlight the importance of continuous improvement and learning from operational events to enhance application reliability.

Final Thoughts This episode serves as a comprehensive guide for developers and IT professionals looking to enhance their understanding and implementation of multi-region architectures on AWS. By focusing on resilience and leveraging new AWS tools and practices, businesses can ensure their applications are robust and responsive to the challenges of modern cloud environments.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00This is episode 739 of the AWS podcast released on September 29th, 2025. Welcome, everyone, to the AWS podcast. I am your host for today, Jillian Ford. And today's topic is one of my favorites because it's a topic that the vast majority of businesses are going to think about at some point in their business journey. So there is definitely something here for everyone. And we're talking about multi-region architectures. We've got two experts that I'm really pumped to learn from on this topic. So Jeff and Scott are here today. So let's do some quick intros. Jeff, introduce yourself. Yeah, I'm Jeff Ferris.

0:49So I am the technical lead of our resilience technical field community. That's a community of essays and other interested parties that focus on resilience and resilience related topics, resilient products, resilience technologies, just across all industries. And we share our experiences with the communities and with the rest of our customers and other ISAs in the field to ensure that everyone thinks about resilience in the right ways. Right ways. I definitely like the way that you just said that. And I hope people here are definitely going to learn the right ways to think about resiliency. Scott, please introduce yourself.

1:30Hi, Jillian. Thanks for having me. My name is Scott Wainstock. I'm a senior software development manager working with a team called Amazon Application Recovery Controller. And what we do is we build recovery mechanisms for customers that they can use in a multi-region or single-region fashion. Super cool. So there's definitely a lot that we're going to dive into this topic from Scott and Jeff. So let's first start off just with the basics. Jeff, when should someone consider multi-region? I think before considering multi-region, it's really important for a customer to be operating very well in a multi-AZ model in a single region.

2:19Once they have a lot of comfort in building and running their applications in single region multi-AZ, then that helps you identify where you have candidate applications to potentially run in multi-region. And we often say just don't run multi-region just to be running multi-region, right? Sometimes it's driven by compliance requirements or latency needs to make sure that your application is distributed closer to your actual users. So there are other characteristics that will help drive that decision. But the decision to go multi-region should never be because it's there, right? It really should be a piece of your business requirements.

2:58I think, Jillian, you had my friends John and Tariq on a couple episodes ago talking about this very topic, right? We did, yes. And I'm so glad that you made a plug for that episode. So, folks, if you missed it, definitely listen to that one as well if you really want to just master multi-AZ before getting into the multi-region topic. so i love that so okay so they've looked at multi-az with and they feel like they've got that down within a single region and i know a lot of customers they start to really think about it maybe from a business requirement which totally makes sense maybe they've got they're expanding into a different part of the world or maybe there's compliance requirements that require for them to think about multi-region.

3:48So now that they've decided, okay, check the boxes, it makes sense to actually go multi-region. What are some additional prerequisites that customers should think about in order to be able to approach a multi-region architecture? I think one of the really important things is to understand the logical and physical separations between AWS regions, right? It's going to require architecting your application in a specific way to not lose the benefits of that logical and physical isolation. You don't want to build a multi-region app that still experiences downtime if there's an impairment in either one of the regions that you're running in.

4:33So maintaining that strict isolation can be challenging. It really requires thinking of your application architecture in a different way. And multi-region give you that kind of bounded recovery context. So you don't want to architect your way out of the benefits of that separation that we've created between our regions. Yeah. Can you elaborate on that of how customers can think about that those tight isolation boundaries? Because I could totally see that being a common mistake, especially maybe if you're someone that is new to the cloud. Yeah, you know, I think, you know, if you think about failing an application between regions, having that separation between your application stacks, they really need to be able to stand alone, you know, to survive an impairment in either space.

5:29Any of those dependencies, you need to fail, when you fail from one region to another, you need to make sure that everything is failing over altogether, right? That requires just a lot of operational muscle, a lot of planning and coordination. It'll be coordination between your engineering teams, the business. It's more than just a technical challenge when you're designing these types of systems. We often will say that the culture of resilience is an important part of being able to succeed with these larger, more complex multi-region infrastructures. You know, kind of having that built into your engineering practices, having that operational muscle to regularly practice failover between regions or, you know, simulating outages of various components, being able to make sure that you understand how your system is going to behave in those scenarios.

6:27And it also requires thinking a little bit differently about how you're managing your data, right? If you're looking at things like consistency versus availability, right? The standard cap theorem trade-offs. Partitioning is part of the multi-region infrastructure. So now you're left with the decision between consistency and availability. So really understanding where your priorities lie and engineering your architecture to support the requirements of your application. Yeah, definitely. I think you bring up a lot of interesting points that I think then causes for people to realize, OK, there's actually a lot that needs to go into planning this out.

7:12So for customers where this is their first time going from one to two AWS regions, walk me through some suggestions how you would think about that customer should go about that process. Yeah, I think first, you know, a full understanding of your application architecture is just a table stakes requirement, right? and then you know building a a run book you know don't just rely on everybody getting things right in the moment you've really got to have everything kind of documented and understood you've got to know when your components um you know come up in a different region you've got to understand uh what are all the dependencies because you can't some you know if you're going with like an active passive model and you're doing an actual failover then coming up in another region requires that all of your components have been brought online before you are able to fully start running your application in that secondary space, right?

8:14And with data replication, you've got to decide if you're going to use an asynchronous or a synchronous approach. There are differences in the time that it takes, the data latency, because there's a need to copy that data between regions. So making sure that you aren't out of sync. if you're using a synchronous replication across regions, then you've got to avoid issues with your data being out of sync when you've started your application in that new region. But with the latency characteristics, your write commits need to be accounted for in your cross-region design. Yeah, it sounds like there's really definitely a lot that I think customers need to think about.

8:54And so I can imagine that even like once you've actually done the design and have it actually in production, there's still a lot that needs to be managed in terms of just operating it. So, Scott, I'm curious from your perspective, what are some of the challenges that you've seen customers face when they're operating multi-region applications? I think maybe take a step back. The customers who are thinking about moving to multi-region or starting from multi-region at the beginning, the first thing that we tend to talk about is calculating how much downtime would actually cost them. So if you were in a single region and your application was to be unavailable, how much does this cost you per minute of outage?

9:48And then when you consider the move to multi-region or the build in multi-region, how much is it going to cost to have this redundancy that Jeff's talking about in multiple regions? And what we find is that usually the redundant cost is kind of outweighed by the cost of being unavailable entirely. So it just helps frame the question a little bit for me. Can you afford to not be available? What challenges are there once you are in multi-region? Coordination resonates with me a lot. You see multiple engineering teams engaged on a single operational event. it's hard or it can be difficult for them to communicate state between the teams to get the right approvals of online when they need to be cross-account operating cross-account can be challenging it just increases the number of eyes that you need on the problem and makes it difficult to understand the status of your recovery yeah and sounds like there's just a lot that really needs to go into to actually be able to operate like with these multi-region applications.

11:01So I'm curious, Jeff, what are some strategies that you've seen customers do up until this point? And so I think, you know, having a solid communication plan is important. I like to talk about some of the things that we do at Amazon because they were very unique to me. And I've been here for 11 years, so I've seen this all in motion here at Amazon for years. And I still remember the very first operations call that I jumped on, you know, way back in the beginning. Everyone at Amazon is invited to listen in on these operations calls. And I thought that was very strange. I was in professional services at the time.

11:41It made no sense for me to be participating in a call about, you know, things that were going on in the operational business of AWS. But then I really started to understand the reason why everyone's invited, right? It makes sure that as we're building things, we understand not only how our own stacks are built, but we understand the issues that people have run into that came before us. The other part of that call that was really fascinating to me was we talked about the operational wins. This wasn't just a call to talk about issues or concerns, but what went right? What were some good things that happened?

12:20And are there practices or policies that could come out of that that would be valuable to other engineering teams? So that really introduced to me the idea of resilience as a core part of a company's culture. It was a blameless call, and they did talk about things that hadn't gone quite right. No one ever asked whose fault it was. They just wanted to understand how do we prevent these things from happening again? There was another very interesting part of that that we call the wheel spin. Just never seen anything like it where just at random, two different groups were asked to explain their current operational metrics.

13:01Now, they weren't having any issues with their services, but the idea is by always needing to be ready to look at your operations dashboard and explain what's going on, you kind of build that muscle so that you're never letting things just kind of run on their own, right? People were always aware of the latest details for their service. If they were called on during the wheel spin, they had to talk about any deviations in performance, whether that was positive or negative. And did they understand why those things had happened? And that just really reinforced how deeply ingrained in the culture that operational element was.

13:42And I think that's a great way to make sure that you are constantly working on those communications channels, constantly understanding the operational status of your service, whether it's good or bad, understanding where you have visibility or where you don't have visibility, where you have gaps in your observability. And those sorts of things come out through those sort of regular practices. It was just, it's still, I join those calls anytime I can. You know, we have, it's about a two-hour call every week. And, you know, I probably make one a month at least. It's just really informative and really interesting to see how that communication flows and how other teams are doing their observability and monitoring.

14:28I've been on the receiving end of that wheel spin, Jeff, a number of times. And I'll tell you, it is an interesting experience. And observability, definitely a requirement in order to answer the sort of questions that you would get on a call like that. And it actually makes me think of one more thing that customers sometimes have challenges with during their multi-region journeys, which is providing evidence that they're in compliance with regulatory restrictions, whether it's data sovereignty or similar items. The fog of war during these operational calls can make it difficult to piece together the event after the fact, especially if you're lacking that observability.

15:14This is a challenge we hear from folks all the time. It's such an interesting process that Amazon does, And I'm sure a lot of customers who are listening right now are like, wow, that sounds really cool. How could I do something similar? Like, I mean, I'm not Amazon, but I definitely want to make sure that my application is resilient. So do you have any suggestions of how to apply that type of meeting structure to a smaller IT team? Yeah, I'll start by saying it's not going to happen overnight, right? It is a significant difference from the way many companies run their ops calls. And if you've already got a good, strong operations culture, you're a step ahead.

16:00But it really has to be driven from the top down, right? This is an executive-level decision. It has to become just an executive-sponsored component that has intent behind it. it's not going to happen by accident, right? It has to be driven from the top and it has to be a priority of all the teams involved from the engineering teams to the business teams to your finance, right? Some of these things, there's cost in taking two hours out of every engineer's day to jump on these calls, but there's value in jumping on those calls. And I think it's critical to understand that that value is there. It is worth the cost to invest that time in building an operational muscle.

16:43Scott, I'm curious from your perspective, having been on the receiving end of those calls, like, would you have any advice for smaller IT teams of how they can implement like a similar process? I think Jeff's summary is pretty good. It's an investment. But the way I think about it is you're either spending that time preparing for an event or you're taking that additional time during an event when your customers aren't able to access your application. So I do think it's a worthwhile investment. Top down sounds right. I also feel like a couple motivated engineers or managers can start this work on their individual teams, start building up best practices there.

17:28And, you know, results tend to speak for themselves. This sort of thing catches on with other teams, build culture from the bottom up. Yeah, I love that. It sounds like even though it's a big time, I say that with air quotes, even though people can't see me, time investment, it actually ends up simplifying your operations long term because you're doing so much planning ahead of time. So I like that from the people culture operations part. I'm curious, Scott, is there anything else that AWS is doing that can actually help customers with their entire infrastructure to be able to manage this regional switch?

18:16I would say, well, just a couple mechanisms that I really like, and then I'll talk about a new service that we've launched that we're pretty excited about. But the mechanisms that I think are pretty good, one of them is called ORR. And what this is, is kind of a premortem for your service where you can look at an application and say, here is 50 things that we've seen gone wrong in the past. How have you addressed those 50 things? Like, are you prepared to actually take on traffic from a standby region if there was to be an issue? Can you provide that evidence for preparation? So think of it as a premortem.

19:00And then the other mechanism that I'm a big fan of is a kind of a postmortem called COE, which is correction of error. And what this is, is it's an extremely deep dive into an operational event where we say, why did this happen? Okay, why did that happen? And we keep asking ourselves why until we get to the root cause of what the problem was. It really shines a light on events or root causes of events after they occur and helps you prevent those in the future. So I'd say that pre-mortem and post-mortem in combination are pretty easy to roll out and pretty effective. So you can think of it as a list of best practices for building and operating applications that are built up over time and are informed by actual issues that happened in production.

19:51and a team would take this list and while they're building their application they would check it say oh we have a S3 bucket over here do we have the right level of encryption on that bucket have we made sure that we're scaling appropriately for compute in a redundant region that we operate in things like that most teams would just sort of do over time but it's a checklist to make sure that no one forgot something. So I think of it as a pre-mortem, making sure that what you're building will stand up to issues from the past. The post-mortem mechanism we call COE, which is after a live event has occurred, we really spend a lot of time thinking about how this happened.

20:40Not just what the first event or what the first issue was, but what caused that first issue and then what caused that issue? And we drill down until we find the root cause of the problem. Then we talk about all the ways that we can remediate this and make sure it doesn't happen again in the future. Once we find enough of these frequent issues, those become premortem questions. So it's this little flywheel of success for operations that it's really just a couple meetings and some deep dives on events. Wow. Yeah, it's super interesting. I think a lot of people I hope at least are going to expire to want to implement a form of that just like pre-mortem as part of their own operations process.

21:27But Scott, you were telling us earlier that there's this new service that's come out? There is. I'm very excited about this one, Jillian. So like I mentioned, I work with a team called Amazon Application Recovery Controller, and we have a number of services that help with multi-region and in-region recovery. But you can think of these previous solutions as kind of triggers for recovery. You flip an on-off switch to help move traffic away from an impacted area. The problem is with a multi-region application, like Jeff was saying, it's very complicated. There can be a lot of dependencies. There's a lot of people on the call who are trying to recover the application.

22:12Coordination can be hard. So what we've built is a new service that orchestrates recovery for multi-region applications and takes some of the concepts that we've already built, takes concepts from other teams within Amazon, within AWS, and centralizes all that. so teams can just go to one place and work through their recovery journey. Yeah, I was just going to say, if you can elaborate on that more about being able to coordinate recovery, especially if the combination of having that orchestrated so it's not solely relying on just like human intervention versus maybe you might want some human intervention at some point.

23:02So maybe like how does application recovery control or region switch like help in both scenarios? Yeah, let me I'll talk your ear off about this for a minute. I'll just put it all out on the table. So the service is called region switch. That's what we launched. It supports active, passive, and active-active applications. So either one of those slots right in. We built it to address some of the common issues that Jeff had mentioned, like customers traditionally had written custom scripts that span multiple accounts, and those are hard to coordinate during a live event. It's hard to orchestrate recovery of multiple resource types in a single spot.

23:47Like you might need to fail over your compute resources and your databases, but they tend to be two separate swim lanes. That's extra coordination. It's difficult to actually test that your recovery mechanisms are working correctly. You tend to find that they're broken during a live event, which is no good. So we've built a service that we think will help customers address this. So how you can think about it is we have a concept called recovery plans. And a recovery plan, you could tie one-to-one with an application. And the recovery plan is the record of all the steps that need to be taken to recover this application during a regional impairment.

24:33So within that plan, we have a concept called workflows. And you can think of workflows as a flowchart. In fact, you can create a workflow using a drag-and-drop GUI from the console to create that flowchart, or you can put it in infrastructure as code. It's pretty flexible there. Once you define that workflow, what you're doing is you're setting execution blocks, is what we call them. And you can think of these as building blocks for recovery. So at launch, we had nine, I believe, nine execution blocks. And these help you scale your application, trigger traffic shifting between regions for your application.

25:20We have manual approval steps. So this gets back to your question about when a person would be involved. So you might say you'd recover your database layer and then wait for a senior engineer to make sure everything works and manually approve it. we also support custom recovery actions through lambda so chances are folks already have some set of recovery steps that they take today we allow you to bundle that up in a lambda and execute that way one of the more powerful things that we do as well as we have region switch plan orchestration which you can think of as a plan of plans so as you're recovering if you have dependent applications those become little forks in the road of the top level application recovery plan.

26:14Yeah, I have a lot more to say about this, but do you have any questions about that so far? I do. Yes. So I work with startups and I'm sure this question could apply to maybe companies that are really new to the cloud, where this is their first time actually planning an active-passive, active-active architecture. So they might not necessarily know what the sequence should be. So are there is is there like maybe any. Like pre-made like templates, default recommendations for customers that are like, well, well, what does AWS recommend what the sequence of steps to be if I'm not really sure what they should be?

27:01Yeah, I think Jeff is going to have a really good set of answers for that, I think. But I'll tell you programmatically, at launch, we don't offer drag and drop template selection. But that is something that we're looking to add in the future is to say, I'm a startup. I have a standard three-tier architecture. Can you just tell me which plan I should use? And we're actually looking for ways to make that selection more dynamic in the future. But I'll say from my perspective, we have a well-architected framework is something that we offer through Amazon, and that gives you a general overview of best practices for building in the cloud.

27:44But Jeff, you probably worked with more startups than me. Yeah, in addition to the well-architected practices, we have some prescriptive guidance out there on our website where we actually walk people through specific infrastructures for, you know, different types of industries or different types of workloads. And with these new features, you can very much expect that we'll be updating a lot of the technical content that's out there for Application Recovery Controller to give some more guidance around what high -level scenarios might look like. It is a relatively new release. So depending on when you're listening to this podcast, it might not all be out there just yet.

28:28But that is something that we're very interested in getting updated. And that is heavily driven by the resilience community. Yeah, I'm really excited for customers to start using this. And it sounds like there's even just like a layer deeper of, okay, you can have active, active, active, passive, but even within those, there's different, it sounds like levels of granularity of what that even means. So maybe Scott, Scott, could you elaborate? Like, for example, you were telling me that there's a concept of passive auto scaling. Um, so yeah, so we, we have blocks that handle different phases of the recovery journal or journey.

29:13Sorry, I'm always going on a recovery journal. I need to start going on a recovery journey. And one of the first things that we would do is scaling. We have an EC2 auto scaling block, ECS scaling and EKS at launch. And what this is is customers can set a percentage of scale that they would like to have in their target region, which is something I should mention real quick. we build these recovery tools and it's always interesting to build them because they have to be more available than the applications that run on top of them. You know, if the recovery tool isn't available during a live event, then, you know, what have we done?

29:56We've built something that broke. So we go to extra steps to make sure these are available. And what we did with Region Switch is we built a regional data plane, meaning that our data plane is in all commercial regions at launch. And the dependency posture that we've taken is very similar to a customer's recovery strategy, where we say, if you've built in the West and you recover in the East, you've taken a dependency on the East. That's where our Region Switch data plane would live. So we would live in both sides of your recovery journey. Scaling is the beginning of this. And through these blocks, customers can preemptively say, I need to be scaled above this percentage in my target region in order to operate successfully.

30:53These blocks do have a concept of passive evaluation, which runs every 30 minutes. and make sure that the resource configuration across your active region and your target region and IAM permissions, things like that are sane. They check out, they didn't break between runs. So that's one of the things that we do proactively with our scaling. We also have recovery blocks, like actions that will actually shift DNS, shift traffic away. We offer that through Arc Routing Controls, which is another application recovery service that helps you move traffic between regions. And we have a new execution block or a new concept here called Route 53 Health Checks.

Read the full transcript

31:44And what this is is a managed function that updates a Route 53 Health Check and redirects traffic based on your DNS configuration. So I think a lot of customers are using Health Checks along with DNS config to move traffic. We just kind of bundle that experience in an execution block. That's probably more than you wanted, Jillian. You just wanted to hear about passive. No, I think this is exactly the kind of stuff that gets people super pumped about multi-region. And there was something that you were saying earlier. No, I think I forgot what it was, but you said something and it made me think, oh, I wanted to go there.

32:26Oh, okay. So, Scott, question for you. So Jeff was making a very great recommendation earlier about being able to test often their multi-region strategy. And I'm curious if this new service, Region Switch, how that influences their ability to test their multi-region strategy. Yeah, that's a good question. And it's, I mean, to be clear, it's always going to be challenging to test multi-region recovery. It's a big action to take. And if you're not able to confidently do it, it's something that people are afraid to do. And there's a lot of finger crossing that things work during a live event. So we've tried to make that easier by consolidating all your recovery actions in a single service.

33:23and we've added a considerable amount of monitoring, both console-facing and back-end monitoring and logging so that you can see exactly how your recovery is going at each step of the process and then at the end get a consolidated report of how that recovery went, which we're making available as evidence for even regulatory requirements. So I guess what I'm trying to say is we've taken a process that involves dozens of steps across dozens of teams, and we've put that on a single place. It still can be challenging and scary to test this out, but now there's a lot fewer threads that you need to chase down during that testing.

34:14I wanted to jump in real quick there. You made a great point about regulatory compliance changes. In the past, simply stating that you had designed a system to be resilient and redundant was sufficient for many, many regulatory areas. But we've seen a lot of shifts in modern framework to actually want evidence, right, to need to see the details and to see that these things have been practiced. And having that output from a region switch helps right away with hitting some of those regulatory requirements. That's the hope. Wow. So I've already learned a lot. I'm sure all of the listeners as well have really learned a lot.

34:56And I'm sure they're super pumped to start using this. So Jeff, I'm curious. You've been working with probably hundreds by now of customers on their just like resiliency strategy. Any last piece of advice that you've had from based on the hundreds of conversations you've had that can help folks today? Yeah, I think it's important to understand, and this is a little bit of a deeper topic, so I'm going to brush it very quickly. It's important to understand where the different data plane and control plane actions are in your infrastructure and to build applications that can fail over without dependencies on control plane operations.

35:40You want things to be running in the data plane. That means that it's already something that's standing up and functioning. You want your recovery efforts as much as possible to be things that can be executed through data plane changes without that heavy dependency on control plane. And understanding where those boundaries are, it takes a deep dive into some of the documentation to really understand what that impact can be. Wow. It sounds like based on that suggestion that that's potentially a mistake that you've seen some customers make. And it shows then the importance of really testing because I could imagine if you, even if it's a smaller application, you would really have no idea what would happen unless you actually test.

36:34Yeah, absolutely.

37:04You're not going to solve every problem all at once, but as they come up, you know, capturing that information somewhere, prioritizing and moving forward until you've got a clean and repeatable process. I love it. Scott, so I'm curious for you if there's any, like, advice that you have. I mean, you know, it's a balance between, and I think Jeff actually said this earlier, the balance between cost and availability here. I would say for testing specifically, it's going to be a lot easier to test an active-active application than active-passive, just because by definition you're still active in one region and maybe you're balancing traffic there.

37:50And it does help you test the mechanisms of your shift, if not the full end-to-end active-passive story. So I think it's really down to the customer to try to see where that balance is. Are you happy just operating active-passive and maybe testing once a quarter, or do you want to spend a little more money with more infrastructure and testing more frequently? It's all going to depend on the individual needs of your customers. I love it. Really exciting. And last question for you, Scott, where can customers go if they want to learn more about Region Switch? Well, I would start at the Amazon Application Recovery Controller website.

38:42That's the landing page for all of our services. Region Switch should pop right up from there. Awesome. And is there any place that you suggest customers go to if they want to maybe learn more, reach out to you on the Internet? um i the learn more you can't go wrong with uh well architected framework i mean spend spend time on that that's um constantly evolving constantly growing um though that's a pretty good record of our best practices um but beyond that uh contact information's on the website you're always welcome to reach out we'd love we love hearing from people especially with a new service like which parts are resonating which parts aren't um we've got a pretty full roadmap for building onto this service.

39:27We'd like to have it informed by your needs. So please don't hesitate to reach out. Perfect. And Jeff, anything with you? Any place that you'd suggest if people wanted to reach out to you with any questions? Yeah, yeah. I mean, Scott mentioned the Well-Architected, the reliability pillar there is all about resilience practices. and there's a lot of information, a lot of guidance that is not specific to AWS, right? Well-architected is about how would you do this? You know, what are the things you need to consider regardless of where or how you're building? And of course, we do have some suggestions that lean heavily on our tools, but well-architected is really meant to be agnostic best practices regardless of where you're built.

40:17if you have an account team your account team can always reach out to specialists through many internal channels we have ways to coordinate with our RSA community if you want to have an executive briefing those can be coordinated we have resilience topics within all of those different things and of course we're just publishing content all the time mostly under architecture You'll also find us at any local summit, and we have a huge presence at reInvent. So stop by the Resilience booth or the Ask the Experts table at any of those events, and we'd be happy to talk about anything that's on your mind.

41:03Great call-outs. Well, Scott, Jeff, thank you so much for being here today on the AWS podcast. Thanks, Jillian. Yeah, thanks for having us. Thank you.

From the publisher

Amazon Application Recovery Controller (ARC) Region Switch - https://docs.aws.amazon.com/r53recovery/latest/dg/region-switch.html

More from AWS Podcast

All 45 episodes
#739: Applying Multi-region Strategies on AWSAWS Podcast · 41 min
Listen in VO