In short
AWS Podcast Episode #722: The Frugal Architect with Werner Vogels - Summary Notes
Episode Release Date: May 26, 2025 Hosts: Simon Elisha and Werner Vogels Guest: Tom Leaman, VP of Site Reliability Engineering at Warner Bros. Discovery (WBD) Episode Link: [The Frugal Architect](http://thefrugalarchitect.com/architects/tom-leaman-warner-bros-discovery.html)
Episode Overview In this episode, the hosts engage with Tom Leaman to discuss his role at Warner Bros. Discovery, specifically regarding the launch and maintenance of their streaming service Max. They delve into the strategies employed to ensure reliability and cost-effectiveness in a highly competitive streaming market.
---
Key Topics Discussed
- Role of Site Reliability Engineering (SRE)
- Mission: Ensure content availability and efficient user experiences across WBD platforms.
- Importance of observability and operational intelligence for system health and user satisfaction.
- Customer Experience Focus
- Shift from system metrics (e.g., CPU usage) to customer journey metrics (e.g., login success, playback errors).
- Development of Service Level Objectives (SLOs) and Service Level Indicators (SLIs) tied to customer journeys.
- Operational Metadata Schema
- Creation to standardize cloud resource management across multiple teams and AWS accounts.
- Comparison to a standardized mailing address system to improve clarity and accountability.
- Team Collaboration and Governance
- Importance of cross-team buy-in for effective metadata governance.
- Involvement of leadership to foster a culture of compliance and efficiency.
- Cost Management Strategies
- Monitoring costs through metrics like cost per subscriber to balance expansion and operational efficiency.
- Evolution of cost tracking from static figures to dynamic, usage-based metrics.
- Frugality vs. Cheapness
- Emphasis on making strategic decisions that prioritize user experience over mere cost-cutting.
- Concept of frupidity – avoiding decisions that save costs but detract from service quality.
- Global Operations
- Managing services across nine separate AWS regions while maintaining consistent performance and regulatory compliance.
- Market categorization for service deployment tailored to regional needs.
- Error Handling and Learning Culture
- Introduction of the concept of celebration of error to promote learning from incidents rather than just corrections.
- Standardized Cause of Event (COE) reviews to ensure comprehensive analysis and knowledge sharing.
- Future Directions
- Interest in evolving metrics to include more detailed cost-driving factors, like user behavior and streaming patterns.
- Continued focus on aligning engineering practices with business outcomes and customer satisfaction.
---
Key Takeaways
- Customer-Centric Architecture: Prioritize user experiences in system reliability efforts.
- Operational Metadata Management: Develop a standardized approach to resource categorization for enhanced clarity and accountability.
- Collaborative Culture: Foster an environment where cross-functional teams work together towards common goals, especially during high-stakes projects like software launches.
- Balancing Act: Understand the importance of cost management without compromising service quality.
- Learning from Errors: Establish processes that allow teams to derive insights from failures, enhancing future reliability.
---
Conclusion The discussion underscores the critical nature of reliability engineering in the streaming industry, particularly in maintaining high-quality user experiences amidst rapid service expansion. The insights shared by Tom Leaman highlight the blend of technical depth and strategic thinking required to navigate the complexities of modern cloud services. The episode serves as a guide for IT professionals looking to implement frugal yet effective architectural strategies.
---
For more insights and updates, listeners are encouraged to check the Frugal Architect webpage and stay tuned for future episodes.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00This is episode 722 of the AWS podcast, released on May 26th, 2025. Hello, everyone, and welcome back to the AWS Podcast and a very special episode, another in our series of The Frugal Architect. And of course, I'm joined by our CTO and living legend, Werner Vogels. G'day, Werner. Very well. Thank you, Simon. He didn't know I was going to say that. That's what I like about Werner. I can still keep him surprised even after all these years. And we're joined by a very special customer guest today. We're joined by Tom Lehman, who serves as the Vice President of Site Reliability Engineering at Warner Brothers Discovery, WBD, where he leads teams responsible for ensuring the reliability, scalability, and operability, lots of abilities, of WBD's global technology platforms that support entertainment brands such as Max, Discovery Plus, and Bleacher Report.
0:49Before that, he had some really extensive leadership roles at Audible, at Vanguard, et cetera. He's done more in, I'd say, DevOps and site reliability than most of us have had hot dinners. Welcome to the podcast, Tom. Thank you very much, Simon. So actually a former Amazonian. Yes, yes. I had a brief stint at Audible. So you'll notice as we get into it, I'll probably bring up a lot of different Amazonian activities and actions. But it's great to be here, Simon and Vernet. You're amongst friends here. So before we even start, I just want to thank you. I'm in Australia and Max launched recently and I'm a subscriber and it's working perfectly.
1:27So, you know, it's always fun in these conversations, I think, and you'll agree with me, Werner, to dig behind the details of the services we kind of take for granted every day. We're going to do that today, aren't we? Yeah, well, look at the past days in Portugal and Spain. You know, whole countries without power. Yeah. Yet AWS data centers kept on perfectly humming. Nobody can watch, of course, any of these videos. everyone's sitting around campfires but you know fine yeah no database productions or anything like that now and i think uh with with after 20 years i mean we know how to plan and now how to prepare for this and i think actually that's also a very large part of where the conversation is going to be about today because exactly it's not only about planning and execution there it's also keeping costs in mind while you do it well tom let's let's start with that before we get techy and believe me, we're going to get techie.
2:26Tell us a bit about your role at Warner Brothers Discovery and what even does the VP of Site Reliability Engineering even do? Like what does that role entail? Sure. So Simon, Site Reliability Engineering at WBD, our mission is to ensure that our content will always be there, always ready to play the moment any viewer presses the watch button. Our goal is to provide the customer uninterrupted and efficient experiences on all of WBD's digital platforms. And in the case that an issue does come up, that we are able to mitigate it relatively quickly, either through automated means or if manual operator initiative is necessary, that they can go in and fix an issue ASAP.
3:16So to bring with that, having insight into your systems, Because it's not just a matter of operations, there's also a data flow going back to what I mean, frugality, for example. Oh, absolutely. Observability and operational intelligence is absolutely essential to everything that we do. So a big part of what our team facilitates is the standardization of how we visualize all the operational data associated with our global deployments. We have hundreds and hundreds of microservices operating across nine separate regions in AWS right now. And as a part of that process, we need to be able to make sure that we have an understanding of how healthy each of those microservices happen to be, all their critical dependencies, their databases, what connections they happen to have, how traffic is flowing cross-region and within our Kubernetes clusters.
4:10It is a lot of data to synthesize and understand. And so a big part of our job is to understand exactly how that data flows, the overall health of the system, and make that tangible and understandable to a broad engineering organization as well. And it's interesting, you talked about the sort of the scope here. I think there's some elements of scope here. I think it's worth picking up because it's different and somewhat daunting. Firstly, you've got millions of customers, like millions of customers for whom any delay, any breaking transmission is a frustration in what is a highly competitive market.
4:46And it sounds to me that you're really trying to take a standardized approach rather than a firefighting approach. Because I'm guessing at any given time, there's all sorts of weird stuff going on. And you could just get lost in that, particularly in a leadership position like yourself, where you're trying to sort of create this umbrella architecture that can help support that. Help us a bit more around that because you touched on the operational metadata scheme. I think we want to dive into that first because I think there's a bit of a mental breakthrough you've made there in terms of how folks can connect these systems together because I think very often it's spreadsheets and documents and who knows who.
5:22You've gone way beyond that. So to dive into that before we even get to the operational metadata piece, and that's a fun journey to talk about by itself. For us, it all starts with the customer experience. So in order to understand our system, at the end of the day, we're providing a product, to your point, that customers interact with, millions of customers. And for us, what matters is the customer and what they're doing and whether they're getting errors, the string quality of the different videos that they're watching. Not, hey, is container ABC573 running hot on CPU? That can be fine, as long as the customer experience is completely uninterrupted in those spaces.
6:04So what we really take a look at to start off with is what we call our critical user journeys, right? So can a user log in? Can the user log in and then play back video? Can they browse? Do they get recommendations that are appropriate for them? And we start by trying to capture that information. We are in the process now of really structuring how we think about things like service level objectives and service level indicators to be able to map back to those customer user journeys and then how they translate into the actual components and API calls on our back end so that we can really have a good tracing of that customer experience and what might be happening with our back end services, our systems, our databases, and drawing that connection between them.
6:48When we start looking at this concept called the operational metadata schema that we created back when Discovery Plus was in the process of launching originally around three or four years ago or so, the problems that we were trying to solve were a taxonomical problem. We were trying to figure out how do we actually get our arms around millions of cloud resources that had been deployed across dozens of AWS accounts. When we spun up D +, a lot of our focus was how do we engineer quickly, how do we get capabilities out as quickly as possible to release the new D-plus streaming product. Each team had effectively landed on their own way of cataloging their services and systems and isolated in a single location that was fine for their operational needs.
7:36but as we started to wrangle things from a platform and from an overall product perspective that started to fall apart it would be the equivalent of if we use mailing addresses and everybody had a different tone for street or a different tone for city and then populated it with different information based on where they are and when you think about something as simple as a mailing address the the power of that and how we've standardized that in different countries it's amazing we can tie back not only just critical notifications and communications we can tie invoices billing utilities all back to that same structure and the same thing goes with how we identify and organize our cloud and distributed systems so that was really what we wanted to target on is how do we create our own mailing addresses that will be understood and standardized across the organization.
8:28That's what operational metadata started to look like. But of course, you had a system that needed to get out quickly, and such things grew, not necessarily in lockstep with each other. So how did you get everybody to more or less take a step back, stop, and adopt your methodology? So there was a sandwich approach that was applied. There was buy-in from not only our engineers, but also senior leadership. And we had to make a business case as to why this was highly valuable. And a couple of the areas where it really came into play was when we started looking at security and infrastructure vulnerabilities, as well as cost management.
9:13We were partnered directly with our InfoSec team. And as we were going through and trying to identify where different vulnerabilities landed, we would easily have the AWS resource numbers that were available. We could tie that back. But then when we tried to tie that resource to an individual or team to actually take action on it, as many of our teams at that point owned their own infrastructure as code, their own deployments, their own operations for their infrastructure, it became a very, very difficult process to do that. And more often than not, what we landed on was the account owner of an AWS account would get landed with all these vulnerabilities.
9:52But it was actually 15, 20, 30 different teams that all lived within the same account. And the account owner was there saying, I don't know what to do with this. And then security is there saying, we need to fix this. So we had some really good motivating factors across the board from a business perspective, from a security and risk perspective that really clicked with a lot of folks. So we were able to get buy-in from our CTO at the time to sponsor a program. And we were able to start lighting up and provide spotlight into tagging compliance across all of our accounts. So we sent out weekly reports that would provide a listing of those.
10:32Here's what is tagging compliant. Here's what it isn't. Here is the breakdown of the different resource types. And then it made it a lot easier for us to be able to align that compliance and teams were able to start self-serving and build that out over time. So part of the metadata is not only all the resources that are being used in which place, but also who is responsible for it. Exactly, exactly. So when we were building it out, a couple of the key items of metadata included a three-tiered hierarchy of the functionality. Effectively, we landed at a point where one tier of hierarchy saying, hey, this is the user service, for instance, didn't provide enough context to make it useful and also doesn't provide an ability to aggregate data in common ways.
11:24So we effectively created this three-tiered system, starting off with a business service, then a service, and then an individual component, where an individual component would map back to a microservice and its direct dependencies like databases, etc. So anything tagged with a component tag, you could identify this is a microservice, a database like an rds instance that supports that microservice would also have the same component tag and then multiple components would roll up into a service multiple services into a business service and per conway's law our business services services and components more or less mapped to our organizational structure so it was easy to create some organization-based reporting after the fact And then you've sort of, I guess, extended this from just purely a performance slash security landscape to suddenly really being able to get a grip on the efficiency part and the frugality part.
12:21And it's interesting to me that you're establishing a lot of measures already at the outset around the customer experience. But frugality is often or cost is often just a corporate thing that we monitor. How did you start to think about, firstly, what the SRE role could even do in that space? Because their reliability, not necessarily design build side of things all the time. It depends. How did you unpack this? I think it's a fascinating sort of story of discovery. Sure. So some of it was just natural evolution, to be frank. because when we originally started off with the O &D taxonomy, it was really about tagging infrastructure.
13:05And teams could go about doing that in any number of ways through using their own IAC. You could manually apply the tag. Some teams had to rely on that just based on how their systems were put together, but ideally through IAC. And then once we created that, some teams identified, hey, it would be really great If, for instance, we could start applying these tags to our actual code repositories. And then we can align not only the infrastructure, but also the repositories, the code, the configuration that maps to that infrastructure together. So now we can draw that alignment of if I need to update RDS instance or Dynamo table A, I can now map back to the repository that actually drives the configuration for that.
13:52And once we got that mapping together and we started to build out common CICD platforms that everybody ended up aligning to when we were building Macs, we then took the connection of you create a repository, you create the mapping to the OMD schema, and then that CICD pipeline could not only align that OMD information to the eventual infrastructure that gets deployed, but also add that to the actual running jobs. so you can understand what jobs are associated with a component, which jobs are aligned to a business service. And then that got propagated out. It got added to our observability data, metrics, our logs.
14:29So you have this tracing capability associated with our entire ecosystem, where you can map almost any piece of operational data back to that taxonomy. Much of that operational data then translates into your utilization that then drives cost. so as this started to evolve over time it effectively became a no-brainer for us to then start aligning costs and understanding these different efficiencies within the context of the omd schema itself sir i am when you think about your series in the pipeline um do you do then dynamic tagging for those resources that go to say to staging or to testing or to i mean because those are again different resources than the ones you use in production absolutely absolutely so the CICD pipelines which are now shared across all of our services that get deployed to this common SaaS platform that I was talking about earlier now no matter what environment we have those tags applied and when those tags get applied again it gets into the operational data and we can see that across dashboards in Grafana and Prometheus data, our open search logs as well.
15:45When it comes to cost, you and your team did something I think that's really interesting and it's often counterintuitive, is that you chose a deceptively simple measure for efficiency, which I think once you pick away at it, it becomes complicated, as with many things that we do in business. But you chose cost per subscriber as your efficiency measure. Yeah. Unpack that for us. What is a very quick sentence to say? I reckon there's a lot there. Yeah, yeah. Well, part of it is a acknowledgement that there's a unit cost economics associated with any type of cloud platform, any platform, really, whether you're using bare metal in your own data centers or out in the cloud.
16:30It really just depends on how much of it can be utility priced versus how much of it can be statically in fixed costs if you happen to be in your own data center where you have to position your hardware. But ideally, your costs should be driven by some factor, right? Something should be in that driver's seat. And when we started looking at how to drive FinOps initiatives out of the site reliability engineering organization, we realized throwing out static dollar figures would be very difficult in an organization trying to grow its streaming business. Because we knew that as we acquired new subscribers, costs were inevitably going to go up in some way, shape, or form.
17:13It's actually a good thing. It's a really, really good thing, right? So we wanted to find some way to balance and understand what our efficiency was with an understanding that costs were going to grow in some way, shape or form. So after we launched Max in the US back in 2023, we decided, hey, we're going to start tracking costs per subscriber. We had plans for expanding in EU, APAC, and then a number of different countries within those regions in the coming years. so every time we would launch there would be a surge of new infrastructure being built out there would be new subscribers migrated for from hbo as well so we were able to track all of those customer acquisitions or migrations and then that would offset or effectively become that denominator below the cloud cost and then that helped us understand are we actually building a system that is just as efficient if not more efficient than some of the products and platforms that had come before it.
18:17So eventually we were able to track cost per subscriber, not only for Macs, but the predecessor HBO Max, as well as Discovery Plus. And we could do a comparison. Is this new platform that we're building really more efficient than what we've built in the past? Which is interesting, of course, because the economic model for streaming services is subscription-based. It's a fixed price. yet your costs per subscriber are fair. I mean, some of them may do binge watching every night on your service and then there may be others that now and then pick up a sports or a new movie or whatever, things like that.
19:00So your subscriber base must be, their behavior must be quite broad. and mapping that back to a subscriber price, a subscription price, must be interesting. So I think you're hitting at one of the areas and opportunities that we have for evolution in the space burner. So this was our MVP unit price economics that wanted to get out the door. It was simple. It was easy. It was something that we could get that was standardized across all of our platforms out the back. But I certainly have aspirations and a vision to get much more granular with our unit economics so we can get back to exactly what you're talking about as more of a price per string or price per login and really tying it back to specific critical user journeys, effectively using the same type of model so that we can understand what are those critical drivers and then tie that back to active subscribers and things along those lines.
20:03So if I think about the critical user journey at Amazon, the retailer, it's search, browse, checkout, shopping cart, and reviews. Because if reviews don't work, people won't buy either. Yes. Everything else on the site is secondary. I mean, recommendations are important. And who bought this? And X, Y as well. Are there particular parts? Did you apply something like that to your organization as well? where let's say the critical usage really needs to be four nines available, where there's other parts of your organization with two nines, maybe fine. Absolutely. So we have a microservice tiering strategy that factors in a combination of the customer experience, business operations, things like legal, compliance, those lines.
20:54So we have a fairly robust set of effectively questions that we use to be able to evaluate exactly where every microservice happens to fit in that domain. that then translates into a number of different impacts from an operations and process perspective, whether it be how we examine the cost associated with it to how we handle incident response to even what type of gates and expectations we set for deployment frequency and quality of testing that happens ahead of time. So that all factors into the play. And then certainly when we think about the actual operations of visualizing and understanding those services, we give higher priority to those tier one services that have a much stricter SLA than those that have a lower tier.
21:43Is that a conversation that you have with the business as well? Not necessarily only the tech only, but where basically the business helps decide which one should be tier one and tier two? So from a business perspective, we don't get feedback on individual microservices, but we get feedback more on that user journey factor. And then the user journey translates down into the microservices themselves. So effectively, if something happens to be a critical user journey or a microservice that directly impacts that, that's going to be a tier one service in those circumstances. Things that are not critical dependencies as a part of the CUJ, those wind up being your tier two in all of those situations.
22:25Tom, I think it's interesting too. There's an insight here I really want to call out because I think it's vital for our listeners is that both on the performance and reliability side and on the cost side, you're using business amenclature the whole time. It's the user journey. It's the cost per subscriber. It's things that I would imagine you can have a very comfortable and open conversation with the CFO, anyone in the finance department, the head of marketing, whomever. These are not geeky IT, hey, what's the back pressure on this particular a service or how fast is our storage running? There's none of that.
23:00But this ability to have an accessible conversation, I think is vital to getting that buy-in to this whole process you're trying to do. Yes. I'll admit, I love a good quirky microservice name, a good awesomest prime or something along those lines. But at the end of the day, what you said is exactly it. We need to be able to provide understanding of technical systems that are digestible to a large audience. And by doing that, we need to actually describe the feature and function that's being performed. Naming is not easy, and certainly the system is going to evolve over time. And even when you land on functional-based names for the system, there's a good chance that the functionality underneath those names is going to change and twist and be dynamic over time.
23:48But that's where we need to try our best to be able to make that alignment because it makes everything else so much easier. So mentioning things are evolving over time, talk us through the merger and what impact that had on Sonop. Did you get a chance to start from scratch again? Absolutely. So I have to give a lot of credit to our senior leadership when we went into the merger. We didn't build completely from scratch. We actually had a mission of best of both. There was a code word where we were building the Bob platform. It was really great from not only a technical perspective, but also a people perspective, because it wasn't a, hey, we're coming from this particular company, we're just going to use this, land it.
24:34But effectively, the engineering teams from both organizations were brought together and they collaborated and really critically examined the backend and frontend systems, everything from the individual like customer facing services and capabilities to how we manage the actual platform behind the scenes. And it was the engineers that came up with the proposals associated with which direction we should go. Certainly there were certain elements associated with product of, hey, how easy will it be to ship certain features based on these different capabilities? But that joining in best of both really put us in a really good spot.
25:11Now, once we had landed on that, there was still an actual migration because the underlying platform, the CICD capabilities, the actual compute platforms and Kubernetes and how we deploy databases and things along those lines did end up getting built from scratch. So even though we had services that might have done XYZ business capability, they still needed to migrate to some of the new functions. But because everything in both companies was containerized, we were still able to port a lot of the code over and still get that deployed through the new functionality back in platform. So it was a really exciting time, a lot of development, a lot of work and really good conversations across the board were had from both the folks that were originally from Discovery Plus and the folks that were from the Warner side.
25:57It was really a one team moment where we came together and operated together. So some of the key functions that we brought over were certainly the operational metadata format. That was something that had really served its purpose within Discovery+. We had this taxonomy that extended across the entire software development lifecycle, so we were able to more effectively bake that into everything that our engineers did on this new platform. And that helped us get set up right away with easy standardization from everything from how we created our repositories, how we mapped OMD to those taxonomies, to how the CICD platform ended up working how we deployed both containers as well as infrastructure and then how we actually ended up understanding it we could easily create out-of-the-box dashboards and visualizations for individual services to portfolios of services and we could also track the costs incidents alerts etc across a standardized way also when there's a merger especially with two tech-heavy organizations, culture might result in clashes.
27:07And in this particular case, you already had all methodology as well as culture around your OMD or the metadata repository. How easy was that to convince the other team, the other side, to actually adopt your methodology? Yeah, so I think, yes, there can certainly be situations where there can be some antagonism and things along those lines. But quite frankly, with how we ended up integrating and the way that we ended up coming together, it wasn't two separate teams effectively clashing hits during the merger, especially at least from my experience with the platform teams. there were a lot of very good collaborative sessions between the teams where we were truly evaluating the merits of different parts of the system and came up with the recommendations together and certainly there were some disagreements but i'm going to go back to my amazonian days there were there was a commitment to disagree and commit uh without any animosity between you know the engineers on the floor which was excellent to see um and in a lot of ways a lot of our solutions were headed in the same direction.
28:22So if we took a look at the future state vision for what had been built out and discovery in the future state vision for what had been built out for HBO Max, while neither had reached their end state North Star, their North Stars looked a lot pretty similar at the end of the day, including specific technologies that they wanted to use, the ways that they would adapt them. So this was a really great opportunity for platform teams to full-scale propel themselves forward to that shared North Star vision. So it was a great happenstance and great opportunity for that type of collaboration to land on what that future state actually looks like.
29:00Did the fact that you'd been given a nine-month deadline to get everything done, did it help in decision-making? It can certainly speed things up, right? It focuses the mind beautifully. Yes, yes. It was very much a, well, we got to do this. We We got to get it done and get it out the door. So we didn't, there wasn't any time for hashing and rehashing architectures and systems because we needed to get to a point, especially as a platform team, you're the first gate for any of your engineers for the backend, for the front end to actually make progress. So we were in the hot seat for the first two months or so, uh, because everybody was waiting for us to be able to build the capabilities for them to build their software.
29:45And talking about making decisions, you know, one of the things we talk about in sort of the frugality of architecture is this incremental approach to cost optimization, but you can't change everything. And you have a really interesting approach to, I guess, categorizing design decisions that could continuously cost you in the future. Help us unpack how you sort of look at that. sure so i if i'm not mistaken simon you're referring to the the closed door philosophy associated exactly yes yeah so a good analogy that i like to think of is when you're going and buying a house right there are certain things that you can change or update or modify there are certain things that you can't you'll hear real estate agents say location location location conversation and sure it's a stereotype but at the end of the day that's a very important distinction once you buy a house in the land you're not going to lift the land or move the land if you buy a house in a floodplain you have a house in a floodplain and you have to deal with that going forward and that's something that is a closed door you can't change the environment around you unless you have a ridiculous amount of money most of us don't have that type of money to be able to do that so you need to be able to focus on what are the things that are going to be whether they're irreversible yeah everything's reversible it's a question of time and money and usually we're short on both oh some things are just like land i mean if you just sold your car to someone you can't go back a week later say oh sorry i changed my mind want my car back yeah there's one way doors and there's two way doors and and there are some doors that you know they get a little stuck and they need they need a little bit of oil and you can you can budge those free so uh from our perspective a lot of our closed door are very difficult to alter uh in technology decisions come down to things like database choice deciding whether you're going to go relational versus non-relational.
31:49So choosing between RDS Aurora versus DynamoDB, that's a pretty sizable distinction. And if you get into production, you have users. If you're going to switch databases for a critical, let's say a service that happens being our critical user journey, that becomes a big endeavor to make that switch. So we want to be able to make sure that up front those decisions are appropriate for the business use case for regulatory needs for reliability needs and for cost needs right um being able to spend and also maintenance and operational means of course too right we don't want to be in a position where we're creating a database that may only be a hundred rows rarely gets updated and then land that in Aurora RDS because that happens to be a situation where you're going to have to operate and maintain those instances going forward, right?
32:44It's just not the fit. So there's an important conversation to be had with engineers to be able to make sure that they've got the right education of, hey, what are the right ways and what are the right situations to use different types of technologies? And do you document those? Do you use that as a repository of knowledge for of future people coming on board? Absolutely. So we have some clearly defined, at least within the database space, the different use cases for different styles of database. And ideally, we help teams understand and self-serve that information. But one of my teams, the Database Reliability Engineering team, partners with teams.
Read the full transcript
33:22So when they create, we have a document that's standardized in our organization for documenting architectural decisions. And this is whenever or introducing a new microservice, databases, et cetera, into the ecosystem, our DRE team gets their hands on the ADD effectively. If there's a database involved, teams flag it. My engineers get engaged in that and they'll do a once over and just validate, hey, does this seem to make sense from a database perspective, a maintenance perspective, and partner with the team to see, hey, are there other opportunities or different ways of being able to use the right database style different types of deployment methodologies replication capabilities and things along those lines up front before we end up getting anything into a production environment where it could end up costing us a lot of money down the line but hopefully also then the category of decisions that people can just go make without having to talk to others first so so in those spaces particularly when we're thinking about cost right for things like container scaling right we'll typically look at that and if we're in a situation where we understand that a service is inefficient when it's just ready for its initial production launch or we have a major event that's coming up we may say hey it's okay for right now we will take on that actual real financial debt and scale it up horizontally right and we'll keep it scaled up as necessary to handle traffic and load for a short period of time because we know that that is a situation where we can come back we can tune configurations we can update some of the logic within the service itself and then eventually you know bring that back down and we have some good ways of being able to measure understand those efficiencies that we can work with teams and teams have the ability to see that information and pair that down a bit so those are some of those situations where it becomes a yeah go ahead let's just make sure that we keep track of this and we don't lose sight of it in the future.
35:17This leads into that concept that I know you've spoken about previously as well, which is the difference between frugal and cheap. And the wonderful concept that we talk about Amazon too is frupidity. And I think you've touched on that a little bit there, which is, you don't have to save money all the time because it's not always the right thing to do. And there are two angles of frupidity and I can't take credit for introducing it to you know, one of our brothers, that's one of our senior leaders. But I remember we were going down the route of a cost savings initiative and every company goes through these throughout their cycles.
35:52And we were like, all right, we're going to emphasize cost over this quarter. And he looked at me and he just said, Tom, whatever you do, just make sure we don't do anything fruid. And I was like, that has resonated with me and the rest of the company ever since that happen and the position there was don't do something in the name of cutting cost that is going to have negative repercussions for our users but it also has the same implication to simon what i believe you were getting at was hey don't spend tons of time on trying to reduce cost that could be served for improving customer experience right saving ten dollars a month and spending a week on building in that efficiency is probably not the right move when it comes to how we're investing our engineering effort.
36:43Yeah, or when efficiencies are one-offs versus the ones that execute three milliseconds over and out again. That said, you mentioned nine different regions where you're operating in, serving all of the world. Are those regions for you identical? or do you have different deployment strategies and maybe different regulatory requirements where you have to operate in? Yeah, so we try to keep them about as identical as possible with all of those considerations in mind. We have a way of thinking about our architecture where we have different classifications of each of our components and services. One is our market specification.
37:31so we divvy up those nine regions into four markets but really it's three at the end of the day america's basically u.s and latin america emea and apac and each of those are effectively there to serve customers within each of those particular areas but we also have a branch of what we consider global services so these services no matter what market they happen to be deployed to, they are 100 % identical. There is no change in business logic. They have the exact same infrastructure deployment. So if they have an RDS instance per region in the US, they have an RDS instance per region at EMEA, per region in APAC.
38:16For our market-specific ones, there are actually unique databases, for instance, within a particular market. So the US databases are isolated to the US regions, there will still be databases in the EU, but they are self-contained in that area. And typically in those spaces, we'll have a replication strategy within the market, where for global services, there will be global replication across all nine regions, those spaces. So from your operational side and your reliability side, do you look at each of those four markets differently or do you have one, if your operations team has one global view of the whole world?
38:55Yeah, so it's both because we care, again, about both. And because we have services that are spread globally, we need to be able to understand global impact as well as individualized market impact. And the market categorization is a part of the OMD taxonomy. So we have that information aligned to each of the services, their individual deployments, and the measurements across the board. So we can easily pull up and understand, here is all the operational metrics associated with the US. Here's all the operational metrics associated with EMEA. Here's all the operational metrics associated with APAC.
39:34It'll pull that data from those particular AWS regions to be able to bring them up to the forefront. So we can understand that for our databases, our containers, the CUJ performance in each of those domains. But it's not that you have your data flow to one centralized location, let's say in the US, and then aggregate everything there, but each of the regions are responsible for itself in terms of providing that data. In many cases, yes. Tom, I wanted to pivot a little bit. You talked a lot about sort of some interesting things that the team's working on, and things don't always go right. I mean, we wish they go right.
40:08We try them to go right, but they don't always go right. And you talk a lot about celebration of error. and that's a term I hadn't come across before quite frankly we we do correction of error at amazon but celebration of error is an interesting concept talk to us about how it works and what impact it's had even just using it in that way I didn't term celebration of error but I think there was definitely a little bit of a rooting in the the coe acronym there and the desire was to really make it more of a we wanted a positive experience associated with the learning opportunities so using the term celebration in that space felt like an a way by which we could highlight the opportunities and the learnings that come from it certainly we don't want errors to happen the celebration isn't the fact that we had an incident yeah yeah no that's that's not the end goal.
41:05The celebration is really about the shared learnings and how we better understand a combination of our systems, our people, and our processes that are in play. So by doing that and changing the nuance and trying to focus on those three functions that all typically get engaged whenever you have an incident, we've seen a really positive approach to how engineering teams reflect on incidents after the fact. When SRE gets engaged, one thing I should be clear on is that at WBD, reliability is a shared responsibility. My organization's name is Psych Reliability Engineering, but we aren't on point for the reliability of all the pieces of our systems.
41:51The teams that build and deploy ServiceX are responsible at the end of the day for the reliability of service X. SRE helps out and helps provide them tools to make that reliable. But occasionally, we do get pulled in for actual engineering work. And typically, these are some of the hairier, larger scale events that happen to occur. And when we go in, we try to look at the incidents from a number of different angles. We really try to understand the observation, hypothesis, and action life cycle that typically occurs in many incidents. And we try to identify that throughout the entire celebration of error process.
42:29So reconstructing the timeline, tagging and understanding what happened even before we started investigating the incident. So what was the state of the system prior to that point, how that informed how we actually respond to it. When people come in, what is their perception when they enter the incident? What knowledge do they bring as a part of that process? And then as we're going through that, We're identifying action items and observations of what was happening at the time and then use that to inform how we're going to bolster the system and how we're going to bolster the processes and maybe even translate that into better training and material.
43:06on. We don't want to be pigeonholed into let's tune an alert or let's scale up. We want to be able to identify, hey, maybe if this engineer or operator knew about this dashboard a little bit sooner, do we have an opportunity to do some knowledge transfer or knowledge sharing associated with, hey, this deployment dashboard could have shown you the correlation between rebuffering rates or playback failures that went up at the same time, and we could have rolled that back faster in that situation. So we wanted to attack all of those different angles. And then we always share these major COEs with the broader organization.
43:45So we'll invite everybody from senior leadership to level one engineers, and we'll do a review of the COE and err exactly what happened. Do new COEs, in your case, have a particular structure? I mean, at Amazon, we have these five whys and sort of descriptions of things that are teached. There's a fixed format of or there's fixed questions that you have to answer. Do you have something like that at all? Absolutely. Everyone, we have a standardized incident creation process. Some of it end up getting triggered automatically just based on our metrics and configured alerts were manual creation processes.
44:24And if it happens to be what we classify as a SEV2 or a SEV1, it will automatically create a template of a COE. and that template includes number of core sections. The first four or five are really about executive summary and decomposing the customer impact because we always want to understand, again, it goes back to the customer, what were the customers feeling at that point in time? And one thing to call out, it's not always external customers. Our systems support internal business use cases for platform engineering. We have customers that are developers within our organization. so that customer impact isn't just about our customers that have an active subscription they're trying to strain so we try to quantify and understand what that happens to be we have a timeline that we go through for being able to understand again each of those different actions observations and things along those lines of how people interacted throughout the incident and then we have standard questions five wise happens to be a part of it and then a number of templated questions associated with observability.
45:29Was this automatically detected? Is there something that we need to change associated with how we manage or alert as a part of the incident? And then there's certainly an action item section as well. We also have a few sections of what did we learn from this? What went well? What didn't go well? Because sometimes you learn more from what went well during an incident than necessarily all the bad things that might have happened and occurred. So Tom, obviously the future is interesting, exciting, and a little scary as well for all of us in technology in terms of what the future holds. I guess before we wrap up, because time has gone very quickly.
46:05Firstly, the journey the organization has gone on is remarkable. And again, I can't overemphasize how hard it is to keep a streaming service up and running reliably. Like, like streaming, streaming stuff is hard, particularly when you've got hits like succession and you got the last of us, et cetera. Like that's, that's, it hits hard. And so clearly the team has done a lot of work. Are there any last thoughts you'd like to share for others thinking about either who are in this world or even thinking about getting into it that you think are relevant and, and, and specifically obviously around frugality in the way you're, you're thinking about that.
46:40Yeah. So I think at the end of the day, in order to build frugal architectures and enable a culture of frugality, there are some core ingredients to the recipe. First and foremost, you need to be able to align what you have to business. I truly believe that. We've done a lot of that within the WBD space is being able to understand how your costs impact the business, the way that it ends up tying back to certain capabilities within your system. You also need to get buy-in that this needs to be something that's done. I think that's easier than some domains, certainly, because everybody wants to save cost and reduce the cost of operations.
47:26you need to be able to provide teams the right visibility right creating and generating reports on a quarterly basis half year basis yearly basis on where costs are how they happen to be allocated the cycle time on that is too long to really embed in a culture so making it as self-serve as possible and getting that insight into the engineering teams to take action is table stakes in my mind. And then finally, provide the teams the right tools and education to take action on that information. Again, if they have the information and they can't do anything about it, you're not setting them up for success in that space.
48:07So get the mission, get the insight, get the tools and education. And I think teams can make a lot of difference in this space. Makes sense. Tom, thanks so much for coming on and sharing all that with us. It's been really fascinating. Absolutely. Thank you, Simon. That's great. Thank you, Werner. And Werner, always fun to do this. We'll have another one soon, I'm sure. And as always, you can refer to the Frugal Architect webpage as well, get lots of information, lots of tips. It's the sort of resource that you want to revisit regularly because you'll learn something new every time. I know I find that as well.
48:38And some great customer stories there. And of course, until next time and doing it frugally, keep on building.
From the publisher
With only nine months to launch Max, Tom Leaman, VP of Site Reliability Engineering at Warner Bros. Discovery had to move fast to keep millions of viewers streaming smoothly. Learn about their innovative approach to measuring efficiency, managing global operations, and building resilient systems at massive scale with your hosts Simon Elisha and Dr. Werner Vogels.
Learn More: http://thefrugalarchitect.com/architects/tom-leaman-warner-bros-discovery.html
