The 12 Scenarios of Failure: Applying Chaos Engineering to SAP at AWS

6 Feb 2024 · 30 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Dev Interrupted: Episode Summary

Episode Title

The 12 Scenarios of Failure

Applying Chaos Engineering to SAP at AWS

Episode Description In this episode, host Conor Bronsdon talks with Guilherme Sesterheim, an SAP DevOps Site Reliability Engineer (SRE) at AWS. They explore the application of Chaos Engineering and DevOps principles to SAP, a domain traditionally regarded as risk-averse. Guilherme discusses how these modern practices can enhance the resilience of SAP systems, particularly focusing on HANA, SAP’s in-memory database.

Key Highlights

Introduction to Chaos Engineering

  • Definition: Chaos Engineering involves intentionally injecting failures into systems to test their resilience and prepare for worst-case scenarios.
  • Purpose: The aim is to understand how systems respond to failures and develop strategies to mitigate real-world problems.

Applying Chaos Engineering to SAP

  • Traditionally Risk-Averse: SAP is characterized as a mature and conservative technology, often resistant to rapid innovation.
  • Modern Practices Implementation: Guilherme is integrating open-source technologies and modern practices into the SAP ecosystem, focusing on enhancing operations and automation.

Challenges in SAP Systems

  • Installation and Migration Complexity: Migrating and installing SAP systems is often complex, making it difficult to implement modern practices.
  • Limited DevOps Adoption: SAP's environment is not conducive to many DevOps practices, with limited automation capabilities within the SAP framework itself.

The 12 Scenarios of Failure

  • Key Scenarios: Guilherme outlines 12 specific scenarios in which failures can be injected into SAP systems to test their resilience.
  • Example Scenario: One scenario involves stopping a primary server in a high-availability setup to see how the system reacts and whether it properly switches to the secondary server.

Strategies for Practicing Chaos Engineering

  • Mindset Shift: Encouraging teams to challenge the traditional views on SAP practices and consider the benefits of chaos engineering.
  • Proactive Testing: Implementing more frequent resilience tests beyond the conventional six-month intervals.

Future of DevOps and SAP

  • Cultural Shifts: Guilherme emphasizes the importance of a supportive culture where teams feel empowered to explore new methodologies.
  • Potential for Modernization: SAP's future developments, such as the Business Technology Platform (BTP), may gradually facilitate the adoption of microservices and more modern architectural practices.

Key Takeaways

  • Integration of New Practices: There is a potential for integrating chaos engineering and DevOps practices in SAP to improve system resilience and operational efficiency.
  • Continuous Learning: Encouraging a culture of experimentation and learning can enhance teams’ ability to handle failures effectively.
  • Customer-Centric Approach: AWS focuses on understanding customer needs and fast-tracking their journeys towards adopting modern practices in SAP environments.

Conclusion This episode offers valuable insights into the intersection of Chaos Engineering and traditional SAP practices, highlighting the potential for innovation in a traditionally slow-moving domain. Guilherme's experience and strategies provide a roadmap for other organizations looking to modernize their operations and enhance system resilience.

Additional Resources

  • Book Recommendation: [Wiring the Winning Organization - IT Revolution](https://itrevolution.com/product/wiring-the-winning-organization/)
  • Trial Offers:
  • [Start Free Trial: LinearB's AI Productivity Platform](https://linearb.io/start-free-trial?utm_source=podcast&utm_medium=referral&utm_campaign=devint-shownotes&utm_content=shownotes)
  • [Book a Demo: LinearB](https://linearb.io/book-a-demo?utm_source=podcast&utm_medium=referral&utm_campaign=devint-shownotes&utm_content=shownotes)

Closing Thoughts The discussion with Guilherme Sesterheim not only sheds light on the challenges faced by SAP environments but also points towards a brighter future with the integration of modern practices such as Chaos Engineering and a shift in organizational culture towards embracing change.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Let's break the application. Let's kill some process. Let's see what happens. You can inject failures on the network. So let's add some latency. And wow, latency in the SAP world is one of the hardest things to spot. Let's inject a system crash. Let's inject some DNS failure, some resolution failure, and let's see what happens. Is your engineering team focused on efficiency, but struggling with inaccessible or costly Dora metrics? Insights into the health of your engineering team don't have to be complicated or expensive. That's why LinearBee is introducing free Dora metrics for all. Say goodbye to spreadsheets and manual tracking or paying for your Dora metrics.

0:37Linear B is giving away a free, comprehensive Dora dashboard packed with essential insights, including all four key Dora metrics tailored to your team's data, industry standard benchmarks for gauging performance and setting data-driven goals, plus additional leading metrics including merge frequency and pull request signs. Empower your team with the metrics they deserve. Sign up for your free Dora dashboard today at linearb.io slash Dora or follow the link in the show notes. Hey, everyone. Welcome back to day two of Dev Interrupted at DevOps Enterprise Summit. I'm your co-host, Conor Bronson, and I'm joined by Gil Herme Sesterheim, SAP DevOps SRE Engineer at Amazon Web Services.

1:14Gil, welcome to the podcast. Very nice to meet you guys. Thank you for the introduction. It's a great pleasure to have you with us. I know you're a speaker here at the summit. I know you had a really interesting talk on applying chaos engineering techniques to testing SAP installation. Can you dive into that? Chaos engineering is a fascinating topic. I come from a very strong open source and DevOps culture. So I don't know, I'm pretty comfortable with regular languages like, I don't know, Terraform, Jenkins, Ansible. So all of these open source technologies I'm very comfortable with. Around three or four years ago, I joined AWS.

1:48I worked for the professional services and I joined an SAP team. So I do have some SAP background as well from, I don't know, the very beginning of my profession. But then I bring in all this load of knowledge and information around the open source part into the SAP world. Which is atypical for SAP. Very much. I don't know if everyone is familiar with, so it's a huge company. Yesterday I brought one very nice number just to show how big they are on the presentation. So basically, if you look for that on their website, they say that 77 % of world's transaction revenue pass through at least one SAP system.

2:26Wow. 77 % of everything, all the money that flows through this world touches at least one. That's incredible. That's amazing. I don't know any company, any other company that has a metric huge like that, right? So SAP is this very mature and also a risk adverse company. Yes. So yes, they are slower than other companies into modernizing their applications. So basically what we approached yesterday was then how do we bring some modern technologies, some modern practices actually, like the chaos engineering into the SAP world. So basically there are some stuff that we can do right now, nowadays, testing how resilient your servers on SAP are.

3:11Yesterday we focused 100 % on HANA, which is the database from SAP, and in-memory database. So basically approaching some concepts of how do we inject failures? How do we check of how things went? So did the database behave exactly like I was expecting before the failure injection? Or I don't know, did something go wrong while I was doing that? So there are some techniques, some things that we can do using open source technology into the SAP world without going inside the SAP. So that's when things get more strict when we talk about SAP. But when we are still talking about the server, there's a lot we can do around the operations already.

3:49Interesting. So it's more about using open source techniques and chaos engineering to extend SAP's capabilities. Exactly, exactly. So when we talk about DevOps for SAP, the most common question that I have, questions that I get, I can separate them into two. So DevOps inside SAP, and yeah, that's very complex. We don't have a lot to do there. and frankly it's hard. Of course SAP has some offer because it's a buzzword so it makes sense to talk about DevOps inside SAP but honestly there's not much, in my opinion, not much value that is generated by doing DevOps inside SAP because of the limitations.

4:29You're just so limited on what you can do. Yes, so it's hard for automating tests. So a very small number of companies is already using Git for SAP. So everybody's using Git except SAP. except the developers that go inside SAP, right? And the second bucket, so DevOps inside SAP, it's very shy, it's hard, there's not much value added to the companies that want to use it. Of course, there are some things that we can do, but again, not much value added. And on the second bucket, we get then the DevOps around SAP. And basically, that's how I call it, around SAP, because you're not going inside, of course.

5:05Then you can get the servers, you can interact with the servers. So I work for AWS, my examples go into the, EC2s and how we do things inside AWS. So basically when you install SAP in a server, you have there your EC2s standing and the regular procedure for, I don't know, any major application inside big organizations is that, okay, you install something, it's working, you leave it there, you don't touch it. So that's the regular, how things go, right? Besides, I don't know, your patching windows, your security maintenances, you don't touch that very often. So all the suggestion presented yesterday was around, So let's try to shift that mindset a bit more from that famous analogy from pets versus cattle, which is pretty famous for DevOps, but it's extremely new when we go to SAP.

5:51So let's try to shift that focus a bit more. So there are companies that used to do their high availability testings every six months. We can, yeah, so that's a lot. And the majority of them... If you're not watching on YouTube here, my eyes just got very big. But that's the typical. So the majority of companies, at least that I interact with, big companies here in the States, they do like every six months. Some of them, they do three months. And I heard about one customer that want to do that every two weeks. So that was the most aggressive one I've ever heard. So yes, we can combine some of the DevOps practices, some of the DevOps non-technologies.

6:32So we can speed that up. we can do those testings more often and we can be better prepared for, I don't know, failures and things that typically happen. And we are just waiting for them, right? So instead of just waiting here, seated at my chair, let's do something more proactive to be there. Let's try to go ahead and solve the problem in advance. It makes total sense. Like you've talked about a couple of examples here of how you're leveraging DevOps practices to extend SAP's capabilities and enhance the approaches you can take around automated testing and other things. Are there different categories of DevOps practices that you're applying?

7:05Is it mostly around testing? What's your approach? Yeah, so what I've seen and the opportunities, so I come from a very strong SAP, again, professional services inside AWS. So we interact with biggest companies here in the States, like huge shops of SAP. So we try to focus more on the opportunities and the struggles we see on the market. And this is very nice because we interact and we learn a lot every single day whenever we talk to a new customer or an existing customer about the struggles that they are having. So what I guess that all of the automations and all of these improvements we've been working for the last two or three years, I guess they kind of categorize in something around the operations part.

7:46Because again, we cannot go inside SAP. We do, but it's not very much nice. But then we go into the operations. So again, automating this chaos engineering part, automating the auto-scaling for SAP. So again, auto-scaling is something very usual, I can say, but that's pretty tricky when you go to SAP and you cannot do auto-scaling in the majority of the servers. So you do auto-scaling in just one or two of them. So even that, the automation is not like a streamline, it's not a simple flow like we have with other applications. So you have to have a strategy of how do you want to auto-scale, and also the limit of your auto-scaling.

8:29So I can summarize that into the operations part. Yeah, I'd love to understand the strategy piece. What's maybe a typical or example strategy that you would take to implementing this? So a typical SAP landscape, you have the database. So let's use HANA here as the example. So the in-memory database from SAP. You have an ASCS and ERS, and I'm sorry, I don't remember all the, what does that mean? We're going to use a lot of acronyms, that's okay. But basically that's an enqueue server. So that guy centralizes all the requests between the platform internally. So it's pretty important. And here we already start seeing some stuff, so different names that SAP gives to common things that the open source world knows.

9:12So the ASCS is always the primary one and ASCS is always working. If that guy goes down, ERS takes up, which is basically the same software, the same software installed, but once ASES is primary and ERS is secondary. That's the rule. So it's not the same application that you're deploying to different servers. So they are different applications, but they work in a high availability scenario. Now getting to your question, the third layer of this, so database, queue server, and then the third layer is actually where your users connect to. So that is the PAS, primary application server, something like that.

9:48That one I remembered. And when you want to auto-scale SAP, so you don't talk about database, you don't talk about the NQ server you just talk about where your users connect and if you have just one server right now and the database is working fine with that one server that same database will have to handle well if you have 10 servers so that's a struggle in planning from the very beginning so that is the only guy, the PAS that you can always scale I don't know, 1, 2, 3, 6, 10, 20 I don't know, 30 servers for the biggest shops out there then your users will connect there and again, your database has to handle that So you have to have a very powerful, big instance.

10:26Because again, just one of them, even though you have two databases, let's say, for high availability, two queue servers, only one of them is always the primary. And the other one is just idle, waiting for something to happen. Interesting. So, I mean, this is intriguing. Obviously, there's a massive scale list, as you mentioned. 77 % of worldwide monetary transactions have at least one SAP connection in them. There's clearly some distinct challenges that you're facing in trying to extend the capabilities of DevOps to SAP because of, to your point, how conservative they are about moving forward.

10:57Are there other challenges you'd like to highlight that maybe we haven't dug into so far, areas you want to dive into more? Well, I guess that one of the biggest challenges we see out there is still this huge complex installation and migration process. And that's understandable, right? Because usually the spine of every organization is an SAP server. So if you are doing anything, I don't know, your orders system, your online carts, so all of that is working in a different layer. And after that, that's pushing data inside SAP. So SAP is usually the spine of every organization. So we are very much used to, I don't know, talk about integrations on SAP as well.

11:37One of the biggest challenges we still see, and honestly, I don't know very much about SAP's vision for that for the future, how they're investing on that, is still breaking these huge, I cannot say single, but few points of failure. So still, if you want an SAP, you have to go under a complex project some months for sure, installing things. So you cannot do SAP instead of Kubernetes, as one example. So this common stack I told you, of course, there are examples. Hybris, which is SAP's e-commerce, runs on Kubernetes. That's nice. That's not as simple as installing Sonderkube or Atlassian. these other guys.

12:15So those guys are pretty simple. You just connect there and plug and play to some extent. Yeah. So they're very simple. So SAP doesn't work that way. So this is, I guess, the main challenge that still blocks a lot of things from, I don't know, for even modernizing even more. So you still don't talk about microservices on SAP. And it's a long journey to get to there. So for me, this is the biggest challenge because since we have so few big points, so many few, that's nice to say, so many few big points of failure, breaking those should be simpler. For further modernization. Right, yeah. And it's interesting to talk about this in a chaos engineering context because it really limits the potential of the testing you can do.

13:01What advice would you have for people who either listen to your talk or listen to this podcast, but how to start approaching this and saying, okay, I want to apply chaos engineering techniques to my SAP instance. I want to start trying to extend this. How How would you recommend they get started? This is kind of the mindset that I've been applying since I started to interact with SAP again. I mentioned that I did have some SAP experience, then I shifted into the open source, and now I am an SAP folk again. What I've been doing is, every time that you talk to an SAP expert, you're going to say, hey, I want to do chaos engineering for SAP.

13:36He's going to laugh. He's going to say, are you nuts? What is chaos engineering? First of all, what is chaos engineering? The same subjects, they don't flow with the SAP world as well. Do you want to do a brief definition for folks? We've had episodes on this before, but I want to make sure everyone here has a definition of chaos engineering, just to make sure they have something in their head. So basically chaos engineering is you injecting failures in your servers, in your applications, not just the servers, on purpose, in a controlled environment, in a controlled way, so that you can prepare for the worst scenario.

14:07Some examples, you can inject failures inside your servers, like I said. I don't know, let's break the application. let's kill some process, let's see what happens. You can inject failures on the network, so let's add some latency. And wow, latency in the SAP world is one of the hardest things to spot. Let's inject, I don't know, a system crash. Let's inject some DNS failure, some resolution failure, and let's see what happens. So case engineering goes all around that, injecting things, errors on purpose, and hey, let's see what happens. And once we see what happened, let's come up with a plan so that thing doesn't happen for real when a real scenario is here.

14:43Totally. The thing I really love about chaos engineering is it's such an iterative way to drive continuous learning and to prepare your company, your team, your technology for a worst-case scenario. And particularly when we're thinking about something like SAP that is so integral to the way the world works with payments, it seems like having that injection of failure so you can understand how to improve it and make it safer would be crucial. And so, to your point, you've kind of alluded to this, it's a little concerning that it's so hard to test failure, because that means in a worst-case scenario, which they do happen, even to companies like SAP, there is so much potential for us not to be prepared.

15:23So even on the newest releases of SAP, which is called BTP, I guess it's Business Technology Platform, which is very, very new. I don't know if it was released still in 2022 or 2021, I don't remember. But for SAP, that's extremely new. So a very few number of customers are already using that. So that guy is supposed to have microservices and a more modern approach. But again, it's extremely new for SAP users and it's still blurry exactly what you can do. So I've been talking to some of my personal friends that work on SAP. They told me, okay, yeah, you can do this, you can do that. But still, I haven't seen that working in a real scenario to share more.

16:01And I know, oh yeah, so now that's possible, right? So I've been hearing stuff. So there's potential, that's great to hear. So I sidetracked you a little bit, but I was wondering about the strategy that you'd recommend to folks who are trying to extend the capabilities of SAP and start doing some of those chaos engineering tests. Oh yeah, so that's basically what I've been doing for the last three to four years. If you go to an SAP folk and you say, hey, let's do chaos engineering SAP, he's going to laugh. Yeah, right? If you want to talk to them and say, let's do unit testing, they say, why? So they don't get the importance of that.

16:35So basically what I've been doing, I've been trying to ask the right questions and trying to engage just on the right battles. Because everything that I've been hearing for this operation side, again, So not talking about DevOps inside SAP, about this operation side, chaos engineering, autoscaling, as I said. So some more regular maintenance operations that we have. Whenever I ask a question and they say, yeah, that's not possible, I ask, why? So far, always the response is vague. So basically it's because I don't know or I'm not sure. Yeah, in theory that can work, but I don't know. SAP is not going to support that, things like that.

17:11So SAP not supporting stuff is also another big blocker for that. But again, so my suggestion is let's try to challenge those common knowns that are out there. So usually SAP folks have been working for organizations, I don't know, for 20 years, 30 years. So they are very, again, conservative on the practices that they apply. And sometimes they don't know what they don't know. So if they are not familiar with the chaos engineering concept, they are not sure. How can you tell me we cannot do chaos engineering for SAP, right? So my suggestion is let's challenge those things that we don't know. And because many of them so far, for my researchers, my work with customers, they are possible.

17:57We just are short of people implementing them. And my understanding is you have 12 scenarios of failure that you've identified in particular. Basically, so these are 12 scenarios where you inject failures on SAP. And the example of yesterday was this HANA database to, again, see what happens. So basically there are some commands specifically to SAP like HDB stop, pacemaker. So pacemaker is a library that you use to put the nodes in synchronization. So it's a cluster tool, clustering tool. So basically you inject errors, each one of these 12 items. It's one error that you inject on the server. So first one I said HDB stop, it's not actually an error, but it's an unexpected behavior for the cluster.

18:41So you're basically, you have two servers running and you're going to go there into the primary and you're going to say, hey, stop. Let's see what happens. So the first thing that it's going to tell is if your high availability is set up properly for this scenario. Because this is just one of the scenarios. So all of the 12 go around something like this. The one that I find more interesting is the one that I guess it's so closer to what we hear from Netflix practices. So what do they do for injecting failures on their application? so one of the last one of these 12 is basically you go there into the linux server and you inject an error you basically echo something into a very specific process that happens in linux and that's going to make the instance crash so first the instance freezes and again so that's the scenario i like the most because the instance freezes so the instance doesn't tell its peer hey i'm gonna go down the instance doesn't tell aws hey i'm gonna go down so basically the instance freezes after two or three minutes, AWS realizes that the instance froze.

19:46Then AWS stops the instance and then, yeah, there you go. So this is unexpected. So nobody ever told anyone that, hey, I'm going down. The first instance. A bit of lag time built in there. Yes, exactly. So let's monitor if the second one is going to pick up from there, like expected. And so this is my favorite scenario so far of these 12 you mentioned. Fantastic. It's interesting you mentioned Netflix. Obviously, they do a fantastic job of approaching chaos engineering. I'll shout out that we had Nora Jones, formerly of Netflix's chaos engineering team, on the podcast earlier this year. Great episode.

20:19If you're looking for someone to follow up on this and learn more about chaos engineering, these two will fit really well together. I'd love to dive into how you got here, Gail Arme. How did you decide that this is what you wanted to work on next? How did you end up at AWS doing this? I like to say that no knowledge that we ever got throughout our entire life is useless. So when I started there with my 17, 18s, I was an ABAP developer. ABAP is the language that we used to develop inside SAP. It's proprietary for SAP. So I started like that, so worked like three years doing that. After that, my career just shifted focus into more the open source area, which when I started learning programming, coding, I was in love with Java, so I always loved the open source part.

21:09So I saw some opportunities back at the company that I was. I was an ABAP developer, but I was seeing this big team working in Java, and they were doing great stuff, amazing stuff. So things that were new, I don't know, 10 years ago that we are talking about nowadays, like, I don't know, service mesh, like load balancing, where load balancers were extremely expensive. there was no cloud still let's put it in quotes but the adoption was not as it is today so I was seeing those guys doing this amazing stuff I was excited about that so that's when I shifted my career into 100 % the open source part and when I saw this opportunity at AWS I always loved the company because when we started working with the cloud AWS is the main player so far and hey I want to work on AWS one of these days.

22:02Here you are. So when I saw the listing for some DevOps inside SAP, DevOps SAP engineer, or DevOps SRE engineer, maybe that's something that would be interesting. So I applied myself, I discussed with my previous manager, and he said, yeah, that's basically what you're looking for. The questions I made, right? So we were basically looking for more automations and doing things for customers faster. that's when I landed into AWS and it's been great. Fantastic. Are there other things happening at AWS that you want to mention as far as exciting work being done, things that you see coming down the horizon?

22:40Yeah, so my team works directly, so I'm not from a services team. So everything that I do, you won't find in the console. Like looking for, I don't know, RDS or EC2, I don't do that. So basically I support customers that are doing, I don't know, they are migrating or they are evolving, they are doing something inside AWS. Let's say you're a customer, you're looking to, again, migrate or evolve or do something inside AWS, and you want an expert to know how to best run your SAP inside AWS, that's when my team comes in. Perfect. So the things that we do, they are the most, so again, we try to modernize and to accelerate the journey of customers to go into AWS.

23:17Because again, these projects for migrating stuff, they take years. So they are very expensive, and this is what we've been focusing on a lot. So nowadays we have tremendous levels of automations and we can have just new joiners, guys that have just joined the organization. We have this huge set of automations, Terraform, CloudFormation, Ansible, Bash, whatever, to help them to be faster and to provision things faster and to help the customers faster. So this is something that we are always looking for because this is what customers are most talking about. So AWS is this very customer-obsessed company, as everybody says.

23:55And yeah, so we like to keep hearing what our customers tell us and then invest on that. So again, shortening the time to deploy stuff or to modernize. Right now, so just from some experience, I'm working on this project where we are migrating some applications into Kubernetes. So a few of them that SAP allows, that's pretty nice. Interesting. I'm also curious, given that we're here at DevOps Enterprise Summit, What are the things that are happening in DevOps and around the industry that you think are going to be the next applications of techniques or strategy that you see coming? For me, and I heard that yesterday in many different ways from many different speakers here at the conference.

24:39And this is something that for me is great to hear because it's something that I believe a lot. So something that we were not used to here like 10 years ago was this psychology part, getting more into organizations. So whenever we talk that developers have to be happier to be more productive, whenever we talk that managers, they have to be happier, I don't know, happier is not the word, but more comfortable, more empowered, more safe. There's research that shows that when teams feel empowered and that their work is impactful, they are happier and they perform better. And we're seeing more and more signs of that, to your point, with happiness being such a predictor of retention and success on productivity.

25:26And so this is something that I believe and I keep following, so looking for the stuff that I'm reading and I'm learning, I'm studying about. So this entire, let me use again, the psychology word here for the organizations. So we are taking better care of people, we are taking better care of managers, of individual contributors. So the entire organizations are talking about that right now. And this is like a great beginning. I'm not saying that we are just in the beginning, but this is something already. If we take a look like 10 years ago, it was just like you do as I say. So we are shifting this mindset.

26:02And when this happens, and it's just not IT, right? It's for the entire, I don't know, any other industry. When folks are more comfortable and when they, I don't know, they trust their employer, They do their jobs better. They're delivering better software. Higher quality. And more than that, what catches me the most is that we are investing in that folks, on those folks' lives. Because if you're happy, I don't know about you, I spend, I don't know, 8-10 hours a day in my work every day. So if I'm awake like 16 hours a day, I am almost two - Half of my time, if not more. I am almost two-thirds of my living time working, except Saturdays and Sundays, of course.

Read the full transcript

26:45So that affects directly your personal life. So if you're happy at work, you will be very much likely to be happier, I don't know, with your spouse, with your children, with your family. And then you start thinking broadly. You won't care just about yourself, you care about, I don't know, the environment, you help your community, you help your neighborhood. That's a great perspective. Yep. Yeah, I think there's there's too many folks, particularly years ago, who kind of took this perspective of like, we're going to like use folks up in a way we're going to use their efforts, we're going to push through it.

27:20And I think what we discovered is like, creating a more sustainable social circuitry, that really programs teams to be more successful and happier, is more sustainable for the team. It's better to your point for the individual, and I'm sure their community. And it's also better for the company, because it lets them be more customers obsessed. It lets them dive into their work more and be passionate. They're going to spend that extra care to understand and dive into it. And so I think it's a really wonderful thing we're seeing. And on that note, I'd love to ask you, you're very clearly passionate about chaos engineering, about the work you're doing at AWS.

27:56Are there any closing thoughts you want to share with us? Yeah, so loving the conference. This is a very nice, I don't know, set number of talks here, speeches we've been hearing. And again, I love to hear that we are investing more and more with this culture thing. So just about the book they released yesterday, I was watching the presentation where they mentioned briefly what the book's going to... And just to clarify, this is Wiring the Winning Organization by Dr. Steve Spear and Gene Kim. We interviewed Dr. Spear on the podcast. I'm not sure if it's going to be out already by the time we release this, but definitely check it out if it is.

28:34It's a great, great talk. Awesome. Yeah, so basically that book, from what I understood of their speech introducing the book, they are basically approaching same culture, same concepts that we are used to know, but in a different way. And I love that. How do we program high performance teams? Because there are consistent, like there's consistent high performance across different companies, right? Like companies with the same resources, the same initial approach, build better teams and are more successful. The example Dr. Speer used on the podcast was Toyota, which is vastly outperformed in what employees are capable of because they've intentionally built these happy, high-performing teams that are cut down on friction points and are set up for success.

29:15That's a great highlight. I really appreciate you bringing that up, Gellir May. And thank you so much for coming on the podcast. It's been a fantastic pleasure chatting with you, getting to know you a little bit. Hope to talk to you again soon. Awesome. Thank you very much. Thank you for the invitation as well. It's been great. Thank you for the partnership here at the conference. Our pleasure. If you want to hear more thoughts from incredible leaders like Gil Arame, check out our Substack. We put out articles every Thursday, deep dives into concepts like chaos engineering on SAP. And also our Tuesday editions include articles from our community, our partners, and of course, our podcasts.

29:50Hope to talk to you all soon and thanks for listening.

30:00You

From the publisher

On this week’s episode, host Conor Bronsdon sit down with Guilherme Sesterheim, SAP DevOps SRE Engineer at AWS. Guilherme delves into applying Chaos Engineering and DevOps principles to SAP, a domain traditionally seen as risk-averse and resistant to rapid innovation.

With expertise in both open-source technologies and SAP, Guilherme shares how he’s bringing modern practices to SAP environments at AWS. He explores how Chaos Engineering can be used to test and improve the resilience of SAP systems, focusing on HANA, SAP’s in-memory database. The discussion also touches on the challenges of integrating these practices within the SAP framework and the broader implications for SAP users and the tech industry.


Episode Highlights:

  • 00:20 What does it mean to apply chaos engineering to testing SAP installation
  • 04:05 What does it mean to have DevOps around SAP?
  • 05:58 Guilherme’s approach to DevOps practices around SAP
  • 10:01 The challenge of handling installation and migration
  • 11:50 How to Start Applying Chaos Engineering to Your SAP Instance
  • 16:57 The 12 Scenarios When You Inject Failures on SAP
  • 19:24 How Guilherme ended up at AWS working on SAP
  • 23:14 What’s Next in DevOps Guilherme is Excited About?

Show Notes:

OFFERS

  • Start Free Trial: Get started with LinearB's AI productivity platform for free.
  • Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.

LEARN ABOUT LINEARB

  • AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
  • AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
  • AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
  • MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.

More from Dev Interrupted

All 208 episodes
The 12 Scenarios of Failure: Applying Chaos Engineering to SAP at AWS Dev Interrupted · 30 min
Listen in VO