In short
How AWS engineers learn from large-scale failures (3,000+ postmortems), what makes a good postmortem, why on-call matters, and how engineering practice is changing (serverless/container trends, databases, caching tradeoffs, and AI-driven software development).
Guest
Marc Brooker, Distinguished Engineer at AWS; long-time AWS engineer (nearly 15 years on call mentioned).
Key claims
Hands-on operation beats armchair opinions; great postmortems explain “whys” at multiple levels (code, testing/validation, organizational/process) and turn lessons into services/tools; on-call is valuable mainly for deep investigations and learning, while routine ticket closing should be automated; caching can create “metastable failures” (fast/healthy vs empty/wrong cache leading to stable outages); avoid caching by using complete materialized data or scalable backends (e.g., D-SQL storage tier as a complete cache).
Notable examples
Aurora Serverless and Aurora D-SQL built from customer feedback that relational data doesn’t fit serverless/container paradigms; D-SQL avoids pessimistic locking via multiversion concurrency control plus optimistic commit checks; relational DB outages often stem from long-running transactions holding locks when clients “go out to lunch.”
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Importance of Postmortems
0:45 to 4:00
Discussion on what makes a good postmortem and its role in engineering.
“At some point, when I was a very junior engineer, I looked at the more senior engineers.”
Shifting Trends in Software Engineering
4:00 to 7:30
Exploration of how software engineering is evolving and identifying impactful problems.
“And there's been this incredible explosion around that.”
Learning Through On-Call Experience
7:30 to 11:30
Insights on the value of on-call responsibilities for understanding systems.
“because it has allowed us to and forced us to spend leadership bandwidth, to spend expertise, to spend the time of our best engineers deeply understanding how our systems operate and why they operate the way they do.”
Conducting Effective Postmortems
11:30 to 14:03
Guidance on how to conduct thorough and meaningful postmortems.
“The best estimate I could come up to, and this was about a year ago, was between 3 ,000 and 4 ,000.”
Understanding Postmortem Processes
14:03 to 14:57
Learn how to analyze and improve processes through postmortems.
“And, you know, and then if you're seeing patterns across multiple postmortems, sort of level those up and say, well, clearly there's a hard underlying problem here.”
Designing for Database Resilience
15:01 to 16:22
Discover strategies to design databases that avoid common operational issues.
“And that's a really common cause of operational issues for systems built on relational databases.”
D-SQL and Concurrency Control
16:35 to 17:49
Explore D-SQL's approach to managing concurrency and transaction control.
“How do you prevent misbehaving clients from being a problem for the database?”
Overhead of Storing Data Versions
18:21 to 19:23
Understand the storage overhead associated with maintaining data versions.
“copies of old rows for the sake of those stale reads?”
The Importance of Postmortem Culture
19:27 to 21:16
Learn why a strong postmortem culture leads to better product outcomes.
“and everyone on those teams, the tech leads are asking you, hey, why did that happen?”
Addressing Leadership and Cultural Problems
21:22 to 23:30
Explore how leadership affects the effectiveness of postmortem practices.
“There are places where the details really, really matter, where things like durability are just critical.”
Show all 27 chapters
The Downsides of Caching
23:35 to 26:12
Analyze the potential risks and downsides of using caching in systems.
“them out and say, well, let's spend more time thinking about the postmortem.”
Mitigating Metastable Failures
26:19 to 28:00
Learn how to recognize and prevent metastable failures in systems.
Understanding Metastable Failures in Systems
28:00 to 29:12
Learn about the significance and impacts of metastable failures in large systems.
“getting to the the scale and performance you need rather than putting a cache in front of things So caching isn't a bad pattern, but it is a pattern with some significant downsides that are really best avoided.”
The Future of Software Engineering with AI
29:12 to 31:29
Explore how AI is set to transform software engineering practices and career paths.
“And, you know, and so again, like you might go here, it's operating a system with seeing nothing like this.”
The Evolving Landscape of Software Careers
31:29 to 35:34
Discover how the software industry and career opportunities are changing with new technologies.
“Now, also with those changes, there are going to be needs for us as software practitioners, people who build software, people who love software to adapt.”
Adapting to New Expectations for Junior Engineers
35:34 to 39:43
Understand the shifting expectations for junior engineers in a rapidly changing industry.
“to specification-driven development and a whole lot of other new things whose names we don't even know yet to build software at a speed and a cost that is unimaginable to do with the old techniques.”
Supporting New Engineers in Their Development
39:43 to 42:03
Learn about the importance of mentorship and support for new engineers entering the industry.
“is also much more valuable than it has ever been.”
Supporting Learning in Technical Careers
42:03 to 43:39
Explore strategies for supporting individuals entering technical fields.
“them the right guardrails you know hey that first time that you go out and talk to a customer yeah it's going to be scary.”
Advice for Senior Engineers
43:40 to 45:40
Discuss the need for senior engineers to stay hands-on and current with tools.
“And, you know, we're doing a lot of that thinking.”
The Importance of Being Grounded in Technology
45:41 to 47:54
Understand the challenges of staying relevant in a fast-changing tech landscape.
“You can build such cool stuff during that period of time.”
The Power of Writing for Engineers
47:55 to 53:20
Learn why writing is a crucial skill for engineers and how it enhances clarity.
“I think it is wider than, you know, ever before.”
Purposeful Documentation in Engineering
53:21 to 56:00
Discover the value and purpose of writing documentation in technical fields.
“I had a manager or tech lead who we would write these docs on either designs or the strategy.”
The Importance of Documentation in Technology
56:00 to 57:40
Learn how documenting technical decisions aids future improvements in system design.
“And that's useful in two ways, by the way.”
Balancing Expertise and Visibility in Engineering
57:40 to 1:02:10
Explore the trade-off between hands-on coding and being visible in the engineering community.
“Maybe there was some more advanced thinking.”
Career Insights from AWS's Legendary Engineers
1:02:10 to 1:08:50
Gain insights into the qualities of successful engineers and recommendations for ongoing learning.
“I think in this moment where things are changing so fast, there's so much to learn, swinging a little bit more towards the practitioner side I think generally will help people.”
Advice for Younger Engineers at AWS
1:08:50 to 1:10:00
Understand the importance of being bold and following curiosity in career development.
“And I think I was a little bit more hesitant than was optimal about leaving those teams and looking for the next thing.”
Curiosity and Learning in Engineering
1:10:00 to 1:10:26
Explore the importance of curiosity in professional growth and learning.
“about what am I learning and who am I learning from?”
Transcript
Automatic transcript. May contain errors.0:00Marc Brooker:If you aren't doing it hands-on, your opinion about it is very likely to be completely wrong.
0:06Peterman:This is Marc Brooker. He's a distinguished engineer at AWS, and I interviewed him for technical learnings from his career. 3 ,000 cloud system postmortems. I wanted to ask you, what makes a good postmortem?
0:18Marc Brooker:I could spend a lot of time talking about that.
0:20Peterman:You had a tweet that said that there are cases where caches are bad. I prefer to see the teams around me avoiding caching where possible. We also discussed how software engineering is changing. What is important given that code is kind of flowing like water now? The job changes and you do different work. For someone who's structuring their career, would you say it's better to be overrated or underrated? Here's the full episode.
0:55Peterman:At some point, when I was a very junior engineer, I looked at the more senior engineers. So what is the difference between you and I? I'm working more hours than you. I'm landing more code than you. Why is it that you're so much more impactful than I am? And then I realized that kind of the direction of your work, like what is the thing that you're actually shifting matters more than the volume of your work and your contributions. What would be your advice on how do you find problems that matter?
1:27Marc Brooker:Yeah, I think you have to go super broad. So I think there's a set of those things that come in from customers, from the world, right? Like here is an unsolved problem. I spend a lot of time meeting with AWS customers and listening to them talk about what are the things they still find difficult in our space? What are they investing in? Where are they spending their time? Where would they prefer to be not spending their time? and focus on their core business instead. And so that's one rich seam of ideas and focus on what's interesting. I think completely at the other level is sort of on looking at the technical trends and you can look at just the kind of speeds and feeds like, wow, networks have gotten faster, storage has gotten faster.
2:11Marc Brooker:We've seen this huge explosion in multi-core and now in GPUs. And so there's a bottom-up innovation trend there too, which you can also look at and say, well, this enables all of these new things. And then broadly kind of across the world, like what are the big trends that are going on? What are the things that are changing in our industry? What are the things that are changing in the world? And really it is those kind of moments of change that bring with them the opportunity to build things and to recognize problems. And so to pick one concretely, when I was working in the Lambda team in 2020, and I was talking to a lot of customers about, they were super excited about building on serverless, they were super excited about building on containers, there had been this massive shift.
3:05Marc Brooker:and what people were seeing then was, wow, I love these serverless products. I love building this way. But the world of data and especially relational data doesn't fit super well into this paradigm, right? These relational databases are still very serverful, you know, fantastically powerful products, but not kind of operationally the same. And, you know, that thinking was, you know, just felt super important to me of like, wow, these customers have brought to me a gift. of understanding something that's really important. And so I joined the Aurora team. We built Aurora Serverless, and then we built the SQL.
3:45Marc Brooker:We've been investing deeply across all of our database products to make them a better fit for these serverless and container workloads. And that is an example of a trend that was brought by a customer. but then also these trends that have been driven by kind of architecture or by other things going on right faster networks faster compute faster connectivity and so one of the big technical trends in the database world right now is this sort of block storage becoming the default back end the default durability layer for for databases of all kinds from analytics workloads to online workloads. And there's been this incredible explosion around that.
4:35Marc Brooker:And so if you look at what we did with Aurora D-SQL, for example, that was very much learning from that trend and taking a lead on that trend and saying, well, we're going to make S3, this block store that we built 20 years ago, sorry, object store that we built 20 years ago, the underlying durability layer of this new database. But obviously, it doesn't have the latency properties or the rich interface that an online database needs. And so we're going to build an architecture on top of that that deals with all of these other things in a much better way, but doesn't have to worry about durability.
5:16Marc Brooker:And so that was this perfect collision of a set of things I was hearing from customers and a set of things that were technical trends coming together and thinking, wow, we've got this opportunity to build something now that is going to be a market-leading product, that would be hard to imagine without either of those input signals.
5:38Peterman:I saw something that you wrote. You mentioned that you were on call for 15 years somewhere in there. And I've heard many stories of more senior engineers negotiating out of on call because per unit time, it could be perceived just not that impactful. And so why did you stay on call for so long?
5:58Marc Brooker:I would say that the majority of my in-practice knowledge about how to build distributed systems has come from being on call and analyzing and deeply understanding these post-mortems and COEs. one of the challenges of running a company like AWS and running large-scale systems is that folks come out of college with often great knowledge of computer science fundamentals, great programming skills, great mathematical skills. All of that stuff is fantastic, but without the grounded knowledge of what it actually means to run and understand systems. And on-call is one of the best ways to learn those things, best ways to see, you know, how do systems really run?
6:54Marc Brooker:How do they really behave? You know, how do customers really use them? What happens when customers use systems in unexpected ways? How can we make systems more resilient to customers using them in different ways? And I think that should be almost a goal of on-call, right? If you have folks in your teams who are on call and they're just closing the same ticket over and over and over well you know that's where you need to just build some automation and again building automation is easier than ever it's more powerful than ever fantastic um but where you really want to spend the time of the deep experts on your team is you know here's something unexpected or or unusual that's happened in in the system Let's deeply understand that and let's bring that knowledge back to both improving that system and communicating broadly to the company and the outside community what we've learned from that.
7:53and so one of the most you know one of the most uh powerful things we do at aws is we have
8:01Marc Brooker:this mechanism of a very broad weekly meeting where we all get together you know engineers from across aws leaders senior leaders from across aws and talk about coes these post-mortems that we write and what we can learn from them and how we can apply those lessons across the whole company And I think that particular mechanism, that particular kind of Wednesday morning meeting that we have is one of the things that has been a core, almost causal factor behind AWS's success. because it has allowed us to and forced us to spend leadership bandwidth, to spend expertise, to spend the time of our best engineers deeply understanding how our systems operate and why they operate the way they do.
8:58Marc Brooker:And that level of being just extremely grounded in reality helps you design better products, helps you architect better systems. It helps you think more clearly about the next round of things, helps you fix issues. And so it's this fundamental kind of learning exercise. It's a real blessing. So I would recommend on call to anybody who wants to learn about the practice of distributed systems. And I would certainly recommend spending time reading COEs, reading postmortems and deeply reflecting on not only what can we fix tactically, but what can we fix organizationally and strategically and what kind of tools might need to exist to prevent this kind of thing happening again.
9:49Marc Brooker:And, you know, you asked earlier about, you know, where do ideas come from? This is another, you know, fantastic kind of flow of ideas of saying, wow, you know, we seem to be solving this same problem over and over in different ways and getting it slightly wrong every time. You know, can we extract a tool to do that? Can we build a service around that? Can we build a feature around that to make it easier for us to get right and easier for our customers to get right?
10:20Peterman:Yeah, it's interesting because I think if you ask most engineers, they really avoid on-call. But it sounds like you kind of go towards it and you've learned a lot from it because it's a major source of customer problems.
10:35Marc Brooker:Yeah, and again, I think for me it comes down to optimizing for finding the most important things to work on. and you know if you aren't close to operating your actual system and you don't know how it's actually working how are you supposed to identify what to fix right you can come up with some theories about those but they're probably not going to be right um and again like i i don't think there's a huge amount of value in the rote ticket closing work of on call i think automation should be doing those kinds of work. But I think there's fantastic value in deep understanding, deep investigations, and deep reflection on what you learn from postmortems and COEs.
11:21Marc Brooker:I tried to estimate a couple of months ago for a talk how many industry postmortems and Amazon COEs I'd read over my career. The best estimate I could come up to, and this was about a year ago, was between 3 ,000 and 4 ,000. and so you know even a little bit of lesson from each one and it tends to you know tends to stick.
11:42Peterman:Yeah that was my next question actually. I looked at the slides from that internal presentation and it said I've read approximately 3 ,000 cloud system postmortems from across the industry and my immediate thought was I wanted to ask you what makes a good postmortem?
11:59Marc Brooker:so i think you know what makes a really great post-mortem is first really getting into the details and making sure that you deeply understand what happened rather than just assuming what happened based on on the biases you bring in um and so there's a kind of lesson one there is if you can't understand what happened well that teaches you something about your logging and metrics and observability and simulations and all of these other things. And then once you deeply understand what happened, then the ability, then a great postmortem steps through the whys behind that at multiple levels, right?
12:42Marc Brooker:Like why? Well, yeah, there was a code bug. Okay, sure. Code bugs. Yes, we can fix that. But we can't stop there, right? Like why was that missed in testing and validation, you know, for these reasons, you know, what can we improve? What can we build around those? Okay, next step, you know, why, you know, why was our testing and validation where it was? Or, you know, why did we assume a certain thing about the behavior of the system that we wouldn't have assumed before? And so as you sort of get through these deeper and deeper layers a a great post-mortem not only identifies kind of fixes to the proximal cause but also identifies broader fixes to technology to organizations to you know products and and and so on um and so that's a kind of multiple levels thing right you can't get stuck on you know what is the the most proximal cause of of an incident but you also can't get stuck on this, well, you know, things fail sometimes and what are we going to do about it?
13:47Marc Brooker:And you have to come up with a set of, you know, really concrete action items to fix things at different levels. Fix this particular line in the software that caused something, you know, fix the testing processes that didn't catch that, you know, fix the, you know, maybe social or team processes that led to those technical processes. And, you know, and then if you're seeing patterns across multiple postmortems, sort of level those up and say, well, clearly there's a hard underlying problem here. You know, can we build a service around that? Can we build a library around that? Can we build a, you know, community of practice around that?
14:33Marc Brooker:You know, are there technical changes we can make to avoid whole classes of things? So that's quite a long-winded answer, but I do think it all flows from understanding and understanding at multiple levels, like understanding immediately what happened, but also understanding broadly what happened technologically and organizationally and in context. and then the ability to connect that particular event or postmortem with other ones, you know, and extract those patterns. You know, one of the things that we did in D-SQL was we spent a lot of time as we were designing that looking around, you know, relational database related postmortems and thinking about both our own and our customers and thinking about, you know, how can we design a database that helps people avoid falling into these traps?
15:30Marc Brooker:and you know a really common kind of outage pattern folks with relational databases is you have a client on a distributed system starts a transaction and then goes out to lunch for whatever reason and you know that could be a gc pause or it could be a lossy network or it could be a loss of connectivity and now it's holding locks and so if you look at you know relational databases, they don't tend to be resilient to clients misbehaving in that way. And that's a really common cause of operational issues for systems built on relational databases. And so as we were designing D-SQL, we were thinking, how do we avoid broadly that class of problems?
16:15Marc Brooker:So folks can say, hey, I'm going to build on D-SQL and just not have this whole class of problems. And, you know, I think that's a really kind of powerful outer loop of the postmortem process is to say, how do we turn all of these lessons into new services and into service improvements? How do you prevent misbehaving clients from being a problem for the database? Yeah, so in D-SQL's case, we have no pessimistic locking. And so within the scope of a transaction, everything that happens in that transaction, all of the reads happen using this mechanism called multiversion concurrency control, where every row in the database, we sort of store a history of versions.
17:02Marc Brooker:And so you can read an old version of a row without blocking writers and saying, hey, you can't update this because I just read it. and then you know locally within the query processor that's handling a connection we spool the writes locally and then you get to commit time and we do this optimistic check of you know can i commit this transaction at at the transaction commit time and so combining those two mechanisms of having multi-version concurrency control and and the scale-out storage that comes with it and the commit time optimistic checks we can strongly say that you know there is no way that a reader of a piece of data can block other writers and there's a no way that that a writer of data can block readers um writers can block writers but only um only by changing data not just by looking at it.
17:58Marc Brooker:And so you can say, well, I can cause, sorry, writers can't block writers, but they can prevent other writers' transactions from eventually committing by making a bunch of changes. And that is inherent to the definition of the particular database isolation level.
18:18Peterman:Out of curiosity, in practice, what percent overhead would you expect for keeping copies of old rows for the sake of those stale reads?
18:27Marc Brooker:Yeah, it's actually surprisingly small. And it's surprisingly small because if you look at the access patterns for most online databases, even ones that do a lot of write traffic, that write traffic tends to be quite concentrated. And it's quite unusual for an online database workload or even an analytics workload to make a second version of every row in the database. Typically, what it's doing is making a first, second, third, 100th version of this row and a 50th version of that row, but the vast majority of data isn't changing. And so it's super workload dependent, as is everything in the database world.
19:08Marc Brooker:But the overhead tends to be relatively small. I would say it's unusual for a online database workload for that overhead on storage to be more than about 10%.
19:22Peterman:From my experience, I've seen an interesting dichotomy between teams where some teams, they really understand postmortem culture. They tend to be infrastructure teams. They tend to take it really seriously. and everyone on those teams, the tech leads are asking you, hey, why did that happen? And really follow up and make sure it's not a problem. Then I've also noticed on other teams that is less of a strong muscle. For those teams that don't take it too seriously, what would be your pitch for why they should take it seriously?
19:55Marc Brooker:Yeah, it all comes down to where you want to spend your time, right? Do you want to spend your time improving your product and making it better? or do you want to spend your time fighting the same fire over and over? And really the culture of building great postmortem culture is to make sure that at the product level and at the organizational level, you are fixing known issues and you are avoiding having the same problems multiple times and typically when i see teams that have you know poor post-mortem culture i think they're probably one of two failure modes there you know one of them is a lack of focus on just the outcomes A lack of really, I wouldn't say caring enough.
20:57Marc Brooker:I think that's a little bit too personal. But being really focused on, is this product performing super well? Are we really making our customers happy? And that is fundamentally a cultural and leadership cultural problem of setting the right standards. Oh, and by the way, I don't think standards should be uniform. There are places where the details really, really matter, where things like durability are just critical. And you do need to have super high standards in those places. And places where you want to optimize for other things and maybe have a higher production defect rate. And I think that's okay.
21:46Marc Brooker:As long as that's an intentional decision that's being made. So that's kind of case one, right? Like insufficient focus on the outcome. I think case two, and this is a harder one to change, is normalization of kind of operational heroics. Like we don't need to fix these root causes because our on-calls are super heroic and they're going to stay up all night and they're going to hack around things. and they don't mind being paged 100 times a week. And they can feel from the inside like it's a good culture, right? Like, oh, wow, these people are super strong owners. They're super engaged. They really care.
22:26Marc Brooker:They're really working hard on call. And those are all good signals. But then when you look at it from the outside, it's like, wow, we're not actually fixing the causes of things. We're just doing this fantastically expensive investment of taking all of these people and their strong ownership and their expertise and spending them just on this break-fix cycle. And that's where you need to kind of look at it from the outside and say, well, let's take this energy of this team, fantastic energy, and focus it on improving the service, getting out of the cycle, finding new things to fix, finding new things to build.
23:04Marc Brooker:And that can be hard because it can be hard for those folks who've been in that mode to look at it and say, this feels so good it feels really like we're caring about our customers and caring about our product and caring about our business uh to realize that oh no we're actually caring about it at the wrong level and we're not serving our business in the best possible way by being so narrowly and tactically focused on this break fix cycle and that's where you sort of need to pop them out and say, well, let's spend more time thinking about the postmortem. Let's spend more time thinking about the causes of things.
23:45Marc Brooker:Let's spend more time addressing these things in a more strategic way. And wow, okay, now you've got so much more time to do that because you've broken the cycle and you can improve your product in different ways.
Read the full transcript
23:58Peterman:I mean, since you have worked on AWS for almost two decades. I'm sure you have a lot of experience building distributed systems. And I think one of the most common advice that you hear, I guess this is maybe in the context of system design, is I almost hear almost 100 % of the time people will say, just throw a cache on it. Or you'll have a system design, you say, how do you make it better? Let's put a cache here. Let's put a cache there. And I saw you had a tweet that said that there are cases where caches are bad despite people saying it's best practice. I was curious if you could explain that.
24:36Marc Brooker:Yeah, so caching's good, right? Like it's, hey, I'm going to take these core ideas from computer science of temporal and spatial locality and I'm going to exploit those to make my system faster, scale better, et cetera. And so, you know, obviously very attractive. But the downside of caches, especially in distributed systems, is they have this mode, right? Like they have this, you know, there's a mode where the cache is full and the cache is full of the right data in time and space to perform very well. And there's a mode where the cache is empty or contains the wrong data. And in the first mode, the system is fast and happy and healthy.
25:19Marc Brooker:In the second mode, the system is slow, often down because now the backend isn't scaled to deal with. all of this uncashed traffic, customers are very disappointed. And often it is down in a stable way. And this is this kind of idea of metastable failures where the system has switched from state one to state two. And in state two, it's still stable, right? Like it's still, it's down, but it's not going to come back up under its own energy because, for example, all of this traffic is causing a huge amount of contention in my database. or is saturating the network. And so I can't even refill the cache.
26:00Marc Brooker:It's not even getting the right kind of data in. And so, you know, when I talk about the downsides of caches, it's really about, you know, how do we avoid that modality between, you know, fast and, you know, that value of caches and the, you know, how do we avoid the state where we're down? um and so if i go back to to dsql like our answer there is dsql what we call the storage tier is essentially a cache but it is a complete cache it contains every row in the database um and so it doesn't have this mode where how do i recover from it being empty or containing the wrong data it contains all of the data um similarly if you look at a a more let's say classical relational database design like aurora the aurora leader is constantly telling the potential failover targets here's something you should cache here's something you should cache here's something you should cache so when a failover happens the cache is warm on you know on on the failover target um and so those are the kinds of things that you can do to avoid those modalities but in general um you know and i wouldn't extract this as a rule or or or say that you know this applies 100 of the time but in general i prefer to see the teams around me avoiding caching where possible i prefer patterns where you have a it's a complete materialized view of the data if you need very fast access to it especially if it's slow moving just pull it down onto your local machine and work with it in memory you know if it's only being updated once a week who cares like just make lots of copies of it um uh so that's that's one pattern or you know use a scalable back end you know dsql or dynamo db or whatever your favorite scalable database is and keep your database vendor honest about getting to the the scale and performance you need rather than putting a cache in front of things So caching isn't a bad pattern, but it is a pattern with some significant downsides that are really best avoided.
28:16Peterman:In practice, how often do you see that metastable failure, though?
28:22Marc Brooker:Yeah, you know, it's not super common, right? Like you might go years without seeing something like that. But if you look across the biggest, most impactful system postmortems across the industry, I would say that these kinds of metastable failures have been an underlying cause in probably a majority of them. and it's super important that you know as an industry and as a community of practice we understand those things deeply because the also those cases where these do happen you know tend to be larger scale issues longer recovery time issues and and and more complex to fix issues right where you have to often you know turn it off and turn it back on again which is this very very painful thing for a team or an organization to do.
29:18Marc Brooker:And, you know, and so again, like you might go here, it's operating a system with seeing nothing like this. But if you look at the most impactful issues, it's actually fairly common as an underlying cause for those issues. And so, you know, it's kind of both of these things are being quite uncommon and being rather common.
29:36Peterman:I was reading your blog and you have a series of posts on how AI may impact the future of software engineering. And I kind of want to pick your brain on that. So what's your perspective on how you think AI will impact software engineering and how it will change things?
29:54Marc Brooker:Yeah, I mean, it's, you know, hard, maybe harder than ever to tell the future. And so, you know, this is a set of maybe guesses and predictions about the future. So I'll say the first thing I deeply believe about software is we have only just started to see the impact that software is going to have on the world. There is such an opportunity for more software to exist, bigger software, better software, more personal software, all of these things. And so software has, throughout its 60-ish year history, been supply constrained. And I think that's going to remain true. I think the opportunity for software in the world is just almost unbounded.
30:47Marc Brooker:And that's really exciting, right? It's really exciting to be at a moment when the economics of building software are changing and are changing rather quickly. and that gives us an opportunity to think about what could we do in the world with a lot more software you know a lot more software personalization a lot more just the right software in the right place at the right time and you know that gives me a huge amount of you know excitement about the future of this industry because we have a massive opportunity ahead of us, driven by these changing economics of software development. Now, also with those changes, there are going to be needs for us as software practitioners, people who build software, people who love software to adapt.
31:46Marc Brooker:And that means that software careers are going to look different. They're going to look different early on. They're going to look different later on. I think the software business is going to look different. And the success of people and organizations over the next, you know, next, who knows, five years, decade, is going to be largely predicated on their ability to adapt to that change and to lead that change.
32:18Peterman:You told this story about this guy who bet on analog circuits when obviously we know digital became kind of the more dominant way. Yet he made good money. For the people who maybe don't want to adapt, you could still get by and succeed. It's not going to be like a crazy thing. Is that kind of the takeaway and why you brought up that story?
32:41Marc Brooker:Yeah, I think that's the right takeaway. And so if I sort of break down the the world into three tiers, I think there's going to remain a huge amount of joy in the craft of software, like the craft of joinery with hand saws. It's a nice way to spend time. It's not a particularly economically interesting activity anymore, but not everything we do has to be an economically interesting opportunity. It can just be something I do because I enjoy it, because I I enjoy the product of it because I enjoy talking to people about it. Right. And so there's, you know, I don't think that is going to go away.
33:22Marc Brooker:I think we're going to see, you know, a lot of interest in that. Like there's been interest in retro computing and, you know, people who run an Apple II as their desktop. And like, well, again, it's wildly impractical. It's not economically interesting, but it's fun and something I, you know, could do as a hobby. And so, you know, that's going to be a remaining part of the world of software for probably forever. And then there's this, you know, kind of story that I told in the blog post. And I think this relates to, you know, driving change in the real world, it's always harder than it looks from the outside, right?
34:01Marc Brooker:Like, as you get into the details, things become more difficult, they become more dependent on people, they become more dependent on politics and policy and, you know our various irrationalities as humans and and so driven by that you know there is going to be a huge amount of and a shrinking over time amount but but a huge amount of the software industry that is run in what i might call the old way right past techniques past languages past technologies. And there's real economic opportunity in engaging with that part of the world. As we saw with analog electronics, analog electronics still very much exist.
34:48Marc Brooker:In fact, there are parts of the world like radio and power systems where there's been incredible technological advancement in those fields. But they have become more niche. And so digital became the mainstream. We wouldn't be talking like we are today if it wasn't for this, you know, 12 orders of magnitude or whatever explosion in digital transistor counts.
35:16Marc Brooker:But there's interesting opportunity there. And I think that interesting opportunity is going to change shape and become more and more specialized and more and more niche and great careers to be built there. And then there is the mainstream, which I think is going to adopt these new technologies from agentic development to AI-powered development to specification-driven development and a whole lot of other new things whose names we don't even know yet to build software at a speed and a cost that is unimaginable to do with the old techniques. and I think that is where correctly the majority of the industry is going to be going.
36:03Marc Brooker:I think that's where the majority of careers are going to be built. I think that's where the majority of economic opportunity is. It's the space I'd be in if I was building a company today. It's the space I'm in in my role and the one I would sort of personally be most excited about. But yeah, it isn't the only one. I think there's going to be the spectrum of software practice and And especially where software engages with the physical world, there are going to be some really interesting questions about how do we bring these new technologies, how do we bring these new practices into the various many niches that software is going to and has over six decades kind of wormed its way into.
36:49Peterman:It's interesting. You mentioned joinery. I wonder if down the road we will see apps on the app store that people pay extra for because it's marketed as this was written by a human or it was written by hand. It's a bespoke custom app. Crazy how the world is going to change. But so it sounds like change is obviously the common case. It's the one that we should be thinking about. Maybe we can break up the conversation in two parts. One is for junior engineers, what is important given that code is kind of flowing like water now?
37:26Marc Brooker:At risk of being a bit meta about our past conversation, it really is about finding those problems that matter and doing that early in a career. And, you know, that requires an understanding of customers. It requires an understanding of the business. It requires an understanding of economics and of systems. and that can, I think that's going to move from being almost kind of senior engineer work of like, oh, well, now you're going to go and talk to customers and actually understand the context of the stuff you're building to being more and more part of even the earliest steps of an engineering career, right?
38:09Marc Brooker:Like here's the context, here's the problem, here's the customer, let's go off and work together and solve this problem with all of this context. And I think that's going to be super exciting for one set of folks and a little bit frustrating for people who have come into looking for a pure software development career, right? Looking for a career where they sit down, open their IDE, start typing and and don't stop for eight hours i think that's going to be a mode that we're going to see fewer people in and a mode that's going to be harder and harder to build a career around now the other mode of oh i'm excited to go off and learn from my customers about what they're building and what they need i think that's going to be ever more highly you know highly valuable and so super exciting opportunity to build you know build careers there and then maybe and this might come across as being a little bit, you know, paradoxical, I think there's also a ton of opportunity for, you know, folks who are extremely technically deep, you know, who are, you know, deep on optimization problems, or deep on infrastructure problems, or deep on, you know, various scientific things, or deep on databases, or deep on, you know, one of the many, many topics that are behind our industry because I think the ability to ask the right questions is also much more valuable than it has ever been.
39:49Marc Brooker:And so I think there is a ton of opportunity for people coming into the industry with deep technical or scientific knowledge to now leverage that in ways that maybe were hard before. There was too much sort of boilerplate to really use that leverage that you have. And so I think we're going to see a lot more of those kinds of careers of really kind of building expertise in a technical topic, in a scientific topic, and then be able to turn that into software and software products in a way that was really difficult before and in some cases wasn't possible before and is now vastly easier.
40:31Peterman:If I was to look at a career ladder's expectations, some of what you described of maybe engaging with the customers and understanding the business context. Uniquely in software engineering, it feels like the earliest levels are insulated from all of that. You have your tech lead, tech leads handing out tasks, and then the early level engineers just given tasks, just converted into code. And it sounds like that part's relatively solved and if not now maybe i'd be surprised if a year or two from now wasn't like completely solved um and i i think that could scare a lot of junior engineers because they would think you're going to expect me to graduate from college or start working as software engineer and then i would have the senior engineer expectations what would you say to the the scared software engineer that's just entering the industry with all this change?
41:30Marc Brooker:Yeah, you know, I think, well, I would remind them that, you know, we as people who hire and build organizations of software engineers, and they as people who have are building software engineering careers have really aligned incentives, right? Like, you know, it's not valuable to hire a bunch of people and set them up to fail like that's nobody wants that it's it's it's not an outcome uh that is good for anybody and so yeah we're going to need to figure out how do you support people on on that path how do you help people learn those things how do you give them the right guardrails you know hey that first time that you go out and talk to a customer yeah it's going to be scary.
42:16Marc Brooker:My first time talking to an AWS customer, it was super scary. But I got a bunch of help with that. And I got a bunch of advice. And I got a bunch of mentorship. And I got a bunch of feedback. And I got better and better at that over time. And I think that's exactly what these things look like. You start off and you start small. And you learn you know as you go and and and so that feedback loop goes faster and so i don't expect that people coming in from college or you know will will come in with all of this knowledge i think you know it's never been true that people coming into technical or engineering careers straight out of college know everything or any career for that matter right like you talk to to teachers about you know what they've learned on their job versus what they learned you know studying uh you know they learn a huge amount in things like internships and so on and over the course of a career or doctors or anybody in, you know, in a field like that.
43:19Marc Brooker:And so, yeah, it is going to be about learning. And I think the emphasis on what people learn is going to be different. I think it is going to require, you know, leaders like me who, you know, care deeply about, you know, hiring and developing folks early in their career to be really thoughtful about what, you know, what does that new ladder look like? And, you know, we're doing a lot of that thinking. I think people are doing that kind of thinking across the industry. And yeah, it's changing fast. It's uncertain. It's an interesting time to be graduating. But again, like it's a super exciting time.
43:58Marc Brooker:I think
43:58Peterman:that's just the scale of the opportunity is bigger than it's ever been. It sounds like your advice for senior engineers is different from that of junior engineers. What is your thinking there?
44:09Marc Brooker:Yeah, I mean, I think for folks there, the challenge is how do you retain the value of this incredible experience and knowledge that you've gained over a career while not falling behind, while learning how to best use the tools? And when I look at senior folks, this is a challenge ahead of them. I think a lot of people have found themselves in influence and leadership type positions where they aren't hands-on building every day. And I think it's going to be harder and harder to be in that kind of role and be able to influence and advise in a relevant way, in a positive way. um and so really i think my advice for folks is is is you kind of got to get building like you've got to get back in get back into it you need to deeply understand how the practice of building software and the practice of designing software has changed and and is continuing to change and and so the challenge is how do i you know really take advantage of all of this knowledge and expertise that I've built up in my career and be super curious and be super hands-on and really be in the details.
45:36Marc Brooker:And the good news for that, well, I think there's two bits of good news. One of them is because of these new tools, time spent as a practitioner is so much more leverage than it is today. You can build such cool stuff during that period of time. the amount of kind of wasted time and boilerplate and so on is so much smaller. And so you really do have this opportunity. And the other one is, again, like, why did, you know, why did we get into the space? Well, I didn't get into it so I could go to meetings and sound smart. I got into it because I love learning and because I love building technology and because I love, you know, solving my customers' problems and because I love, you know, learning about new, you know, new technologies and learning new things.
46:24Marc Brooker:And there's more opportunity to do that than ever before. Again, because of this new set of tools and the leverage that comes with them. And so it really is getting back to, why are you here? Why did you get into this career? And I think it really gets us as technology-focused people closer to our original answer to that. It's really obvious to me right now when I speak to practitioners you know who and who isn't using a you know modern set of agentic powered developer practices right and the people who are um have these really interesting things to say about the strengths and weaknesses of those approaches and the work that still needs to be done and the integrations that still need to be done and the things that are working and aren't uh and the people who are you know, using them hands-on, have such a poor mental model of how they work, what they're good at, what they're not good at, that the things they say about them tend to be essentially fiction.
47:31Marc Brooker:And so, you know, I think we are in this minute that if you aren't doing it hands-on, your opinion about it is very likely to be completely wrong. and that takes a level of humility to admit that you know is is tough you know it's tough for folks with fancy titles and it's tough for folks with with distinguished careers uh but i i think it's
47:54Peterman:a must i i feel like there's a common sentiment among software engineers when they when they work with someone who is a quote-unquote tech lead but they're not really hands-on so they've kind of been in the docs for the last five years or or so and there's these minor things they can tell that this person doesn't actually understand the underlying thing and sounds like that gap will widen with these new tools which is if you're you're looking at things from a thousand feet up and you're not actually using the tools that's just another thing that separates you from the people who are actually building where you'll be very out of touch and i i think uh you know when
48:37Marc Brooker:I look at, I think that's always been true. I think it is wider than, you know, ever before. And, but when I look at the, you know, engineering leaders that I've really respected and learned a huge amount from over my career, you know, for example, some of the folks who built S3, you know, 20 years ago, that was such a successful product because those folks were so deep in the details and so grounded on the use cases and so deep in the economics and really just did you know really thought about both the kind of strategic world of like how is this cloud thing going to change the way people want to interact with storage but also the you know minute-to-minute details of what's fast now what's slow what's good what's bad and I think when you think about an extremely enduring product like S3 or EC2, I think it's been that groundedness in the details from early on, from all levels of leadership that has made those things so successful, where other products seemingly with the same amount of early promise didn't turn out to be as successful.
49:59Yeah.
49:59Peterman:I think one of the last topics that I wanted to ask you about was writing. You have a ton of awesome posts on your blog. The style of writing is incredibly clear. And I was curious, why do you write so much as an engineer?
50:19Marc Brooker:Writing and speaking, but especially writing, have this incredible power. and you know for for technical folks it's this incredible multiplier and being able to take these ideas that's in your head and share them with the world and you know you can can take a set of technical ideas in your head and share them with the world by building a great product and that's a fantastic thing to do you can share them in the world kind of one-on-one you know mentorship teach people learn small groups also a great way to spend time but the multiplication factor of doing a talk or even more of of writing something is so much higher right like there are so many more people that you can share that with and it lasts for a much longer period of time and so just having something written on my blog even that I wrote like a decade ago that I can share with someone and say, here's how to think about this problem.
51:22Marc Brooker:Here's an insight that I wanted to share with you, or have people discover that organically is just super powerful. And so writing lets you scale out the impact of your expertise in space and time in a way that's really hard to do in other media. I think with video and with podcasts and so on, we've seen other ways to do that but i think writing remains kind of uniquely powerful and then there's also this idea which is this kind of core belief culturally at amazon and i've obviously been affected by this over the years that you know writing forces a level of mental clarity that speaking making slide decks etc doesn't and you know that's something that has also really been my experience of sitting down to write something down forces me to think that through at a depth that I wouldn't have been been forced to think it through without that um and so I saw one of you know your early conversations with was with Leslie Lamport who kind of takes that a step further and say hey you know it's formal mathematics that is the next step there and I I love that point um but I I think writing is this really accessible thing for for people to do that does force a level of thinking and so i do a lot of writing sometimes just for myself right like i'll write a doc not ever intending to share it with anybody but just to sharpen my own thinking on a on a particular point and so some of that combination of three things right like i just i just have something to say and i want to say it um you know I have something to say and I want to scale it out in time and space.
53:05Marc Brooker:And I want to sharpen my own thinking on a subject or the thinking of a small group on a subject in a way that writing is just a super powerful tool to do.
53:20Peterman:Definitely. Yeah, I remember being surprised early in my career. I had a manager or tech lead who we would write these docs on either designs or the strategy. And he said, even if you just wrote it and you threw it away, it would still be worthwhile because you'll realize things as you're writing and that clarity will save you a lot of time down the road. And it's interesting to me because a lot of engineers, they complain about writing docs. Docs and all the stuff around the code, they kind of hate that. It's a sign of slow, big company processes and what would you say to an engineer like that who's saying, just let me write the code?
54:03Marc Brooker:Yeah, and that's a great point. And I think it really depends on the level of problem you're trying to solve. And so if I look at, I'm going to pick on UML for a minute here, right? Like it's a sort of semi-formal software design process and not one that I've ever found useful because I think it just happens at the wrong semantic level. I think it's bothered with details at a level that aren't helpful. And I think a lot of the let's go off and document this has a similar problem, right? Like, does this actually require that level of reflection and thinking? And so I think what, you know, for me, separates a valuable doc writing and thinking process from a busy work process is understanding what you're getting out of it.
54:55Marc Brooker:and what you're getting out of it might be an artifact to share with the future, which is super valuable. Either your future self, if you've got a terrible memory like me, or new teams, new people, or I want to share something with customers or I want to share something with the world. And so that's super valuable. Or I want to write down something so I can think through a really difficult, often one-way door kind of hard-to-change technical decision or API design decision. And I'm not going to do that every time I make a technical decision. It's not worth it because a lot of those technical decisions are either easy or not as critical or can just be taken back if we figure out they're wrong.
55:43Marc Brooker:But I am going to spend my time that way when there are key decisions to make when there are key um insights to to find and i think you know and so it is that like what is the purpose of writing uh that uh that that separates well-spent time from poorly spent time now there are people who still don't like writing even when it's well-spent time even when it's like you know you have to explain this piece of technology to you know to a future team um i think that's a skill worth developing you know sometimes you you do need to you know uh eat your vegetables uh you know and it's it's it is a skill worth getting good at um and you know especially in documenting the core kind of technical decisions behind a design is is so useful.
56:45Marc Brooker:And that's useful in two ways, by the way. One of them is, as we think about building a big system, we make thousands of decisions. And some of those decisions are very carefully chosen, very particular, and very impactful. And some of those decisions are the best thing we could guess in the moment based on having no data to make that decision. and it's super useful for people who are coming in to improve that system down the line to be able to look at the design and say which of these things were very carefully chosen and thought through and which of these things were arbitrary and because the arbitrary things like okay well I'm going to change that and I'm going to just go ahead and change that because I have better data now I've watched the system run I can go and change those and these other ones are like well well, let me really engage with the reason that we made this decision.
57:39Marc Brooker:Maybe it was non-obvious. Maybe there was some more advanced thinking. And so being able to kind of understand the amount of thought that went into a decision is almost as important as understanding what that thought was.
57:50Peterman:You had a really interesting blog post. This was from a while back. It's titled The Four Hobbies and Apparent Expertise. And you introduced this really interesting idea. It's a two-by-two matrix. And on one side, there's doing versus discussing. And on the other side, there's the hobby and the gear.
58:13Marc Brooker:Maybe I can overlay it for people who want to see.
58:16Peterman:And then later, you kind of liken that to your career and how, I guess maybe we can imagine the hobby is actually coding. And maybe the gear is, let's just say it's like your dev setup or something like that. You talked about these two aspects of being in, depending on which quadrant you are, which is there's this tradeoff between expertise and visibility, where imagine you're really into coding and you're really into doing. You're going to be phenomenal in terms of expertise, but maybe not as visible because you're not talking with everyone about how cool your setup is and all of that. But on the flip side, if you're really into the gear, or maybe you're set up in this case, and you're really into discussing, you're on all the messaging posts and that, you might not actually be that good at coding, but you're very visible and you have this apparent competence.
59:11Peterman:And I thought that trade-off was really interesting because I've seen that so much in software engineering too, is there might be someone who's a really quiet coder. They never write anything, but they know everything. because they've just been in the weeds all the time. And then there are people on the complete opposite end of the spectrum that writing all the time, speaking all the time, but maybe not actually practicing as much. And my question to you is, how do you strike that balance? Because obviously too far in either direction is not optimal. So how do you strike that balance?
59:44Marc Brooker:Yeah, that's something that I reflect on a lot. And I do explicitly think that sort of being 100 % on either of those ends is a failure mode. And I think, you know, I will say that I have a lot more personal enjoyment working with the people that are 100 % on the doing side and 0 % on the talking side. I appreciate and deeply, you know, deeply love their expertise. But I do think that, you know, they could have more impact and leverage if they, you know, swung a little bit away from that. You know, I tend to not enjoy as much interacting with the people who are 100 % on the speaking side. But I think they would, you know, have a lot more relevant things to say, you know, if they, you know, swung a little bit back towards the center.
1:00:45Marc Brooker:The other challenge of being on 100 % on the doing side, it sort of gets back to that, how do you find the really important problems? And, you know, if your head's down in your IDE all day, you could very likely be working on the wrong thing. You know, something that isn't as important, isn't as impactful, you know, doesn't have these properties that people want. So, you know, how do you find the optimal balance? I don't have a recipe for, you know, what really is optimal. I tend to do about, let's say, 75, 25 kind of practitioner versus teaching and communicating, maybe 80, 20 at times. I found that's about what feels right for me.
1:01:35Marc Brooker:I would say that there are great people I work with from sort of 90, 10 on that scale up to about 50, 50 on that scale. I think outside of those, folks tend to get into trouble as practitioners. There are people whose job it is to be communicators, and that's great as long as they have the curiosity and are clear about what they know and don't know. But I found that sweet spot at that 75-25 point in my career, and that's what's worked for me. I think in this moment where things are changing so fast, there's so much to learn, swinging a little bit more towards the practitioner side I think generally will help people.
1:02:28Marc Brooker:But again, you don't want to go too far that way because then you lose the what's important for you that comes with interacting with the outside world.
1:02:37Peterman:on the doing versus discussing axis, I kind of view the doing one as if you were too far, you would be underrated. And if you were too far on the discussing, you would be overrated. And for someone who's structuring their career, would you say it's better to be overrated or underrated?
1:03:00Marc Brooker:I think long-term, if you're using that terminology, it's probably better to be underrated. I think being overrated can feel great in the moment, but it's really sustainable and really sort of gets you to where you need to be. I really enjoy things like sports and these sort of creative hobbies and crafts because it does turn that, let's say, perception and reality knob to very much reality, right? Like as a sports person, you can't fool the world for very long. It very quickly becomes very obvious who can and who can't. I think as a craftsperson, the same, right? It very quickly becomes obvious who can and who can't.
1:04:01Marc Brooker:And I think it takes a little bit longer in a field like ours where there is so much kind of qualitative stuff that goes on. But I think long term, when I look at, you know, careers that I really admire and people I really admire, they tend to be people who are personally very honest about their level of knowledge and understanding and skill.
1:04:23Peterman:So people who walk the walk, not necessarily talk the talk.
1:04:27Marc Brooker:I see.
1:04:28Peterman:Yeah, about engineers that you admire, I'd be curious because you have worked at AWS for such a long time, and you have seen so many legendary engineers. Who at AWS do you look up to and why?
1:04:43Marc Brooker:yeah i mean you know just just fantastic one of the blessings of working at a place like aws is i get to work with so many great people um you know maybe because he's retired i'll i'll talk a little bit about so elver murlen uh was was one of the sort of early engineers at aws and um original say huge contributor to the design of s3 a really big contributor to the design of our lot of a lot of our database services over time um al was actually the the cto of amazon for a for a period of time when he realized i think that wasn't the job he wanted to do um but i you know what i really admired about uh about al from early in my career is you know very clearly he was somebody who deeply understood the things he was he was doing and he could work in these two modes I have a great memory of 2010-ish arguing with Al about some of the edge cases in the Paxos paper.
1:05:50Marc Brooker:He was super deep at that level but could also get up to the executive level and talk about cloud strategy and the way we should be explaining things to people and some of the fundamental things that we need to be building. and I really admired that ability to work sort of almost at every level and I was like wow you know this is this is something I aspire to and you know want to model my want to model my own career after and so you know that's I think that is you know the kind of person I've really you know, really enjoy working with is people who do have that, you know, do have that breadth.
1:06:34Marc Brooker:And I think, you know, one of the other things that is really admirable about a lot of these folks is, you know, they don't want to be celebrities, right? They want to do cool work for, you know, have an impact, do great stuff for customers, you know, optimize for having impact.
1:06:51Peterman:For people who want to continue their engineering education and really remain on top of things, deeply understand the technology, do you have any top technical book recommendations?
1:07:03Marc Brooker:You know, anybody who's building distributed system things, I highly recommend Martin Kleppman's book. I think there's a second edition of that coming out soon. There's a new edition of Quantitative Systems Design book, which I also think is great. Right. Hennessy and Patterson's computer architecture book. This is a super, super useful one that covers a ton of ground. I read a ton of, you know, fiction and nonfiction and mostly papers when I'm reading technical things. I find, you know, I find engaging at that level, you know, more useful for me. um and by the way that's that's become way more uh accessible now you know one of the great ways to dive into a paper is you know hey uh you know hey claude uh summarize this for me and then then i can dive into it and and read you know the author's words and i find that mode is great and it's super accessible for people who haven't been able to read papers uh in the past um but uh you know and then there's also a ton of insight in in some really old stuff too uh for example um you know some of the algorithms that we used in in in lambda uh to to manage traffic and manage bursts of traffic come from erlang's work like a hundred years ago on on managing telephone call centers and his book about that.
1:08:36Marc Brooker:And so, you know, folks also shouldn't think that, oh, well, the industry is changing super fast. And so I should only read recent things like this full incredible insights in some of the, you know, older work and in the foundations of computing and infrastructure and networking and computer science that there's, you know more again more maybe more leveraged than ever before you know deeply understanding
1:09:04Peterman:those topics and then last question for you is if you could go back to your younger self when you just joined AWS and give yourself some advice what would you say I think maybe be a little bit
1:09:17Marc Brooker:bolder I really loved the team that I worked with and and you know you know especially in EC2 and in the early days in EBS. And I think I was a little bit more hesitant than was optimal about leaving those teams and looking for the next thing. You know, as, you know, my own learning and impact kind of, you know, tapered off a little bit in those places. And so, you know, I think I've changed organizations kind of, in a big way, four times in my career, and maybe five or six would have been optimal. Not a lot more, but some more. And so don't hesitate to think about what am I learning and who am I learning from?
1:10:03Marc Brooker:And is there a better environment to do that more quickly and to learn more things? And I'm personally highly motivated by being able to follow my curiosity. And every time I've done that in my career, I've found that a valuable move and something that I've personally enjoyed.
1:10:27Peterman:Awesome. Okay, well, thank you so much for your time. I really appreciate it, Mark. Thank you for sharing with the audience. This has been super fun. Thanks so much. Thank you for listening to the podcast. It's a passion project of mine that I really enjoyed building. Another passion project that I've been working on kind of in secret is building an ergonomic keyboard that I wish existed. And I finally have a prototype, so I'd love to show you what we've built. It's ultra low profile and ergonomic and I couldn't find anything like it on the market so that's why we built it. I'll put a link to the keyboard in the description.
1:11:01Peterman:You can take a look and learn more about the project there. We could definitely use your support. Also, if you have any feedback for me about the show, I'd love to hear it. Comments on YouTube have led to guests coming on like Ilya Grigorik and David Fowler. I wasn't aware of them until someone dropped a comment. Also, feedback in the comments helped me learn to reduce the number of cliffhangers in the intros. So your comments definitely make a difference. Please keep letting me know what you'd like to see more of in the show and I'll see you in the next episode.
From the publisher
In this episode, I talked to Marc Brooker, a distinguished engineer at AWS who started there as a new grad and rose through the ranks. We discussed technical learnings from 3,000+ cloud system postmortems, how software engineering is changing with AI, how to find impactful problems and much more.
🔶 My keyboard Kickstarter: https://www.kickstarter.com/projects/ryanlpeterman/compose-simple-ergonomics-beautifully-done
𝗣𝗼𝗱𝗰𝗮𝘀𝘁 𝗹𝗶𝗻𝗸𝘀:
• YouTube: https://youtu.be/u3GjIXP9N0s
• Spotify: https://open.spotify.com/episode/1qX2GfpbzxzGpGvDZVINdO?si=wsDGZo9PTbCNalKVybFVnA
• Apple: https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835
• Transcript: https://www.developing.dev/p/aws-distinguished-eng-learnings-from
𝗘𝗽𝗶𝘀𝗼𝗱𝗲 𝗹𝗶𝗻𝗸𝘀:
• Post we discussed on hobbies and apparent expertise: https://brooker.co.za/blog/2023/04/20/hobbies.html
• Post on software engineering changing: https://brooker.co.za/blog/2026/02/07/you-are-here.html
• Post about Senior engineers and AI: https://brooker.co.za/blog/2026/03/20/ic-leadership.html
• Post on Junior engineers and AI: https://brooker.co.za/blog/2026/03/25/ic-junior.html
𝗧𝗶𝗺𝗲𝘀𝘁𝗮𝗺𝗽𝘀:
0:00 - Intro
1:27 - Finding problems that matter
11:42 - Learnings from 3000 postmortems
23:58 - Why caches are bad
29:37 - How AI will change software engineering
36:49 - Advice for junior engineers given AI
44:02 - Thoughts for senior engineers
49:59 - Why engineers should write
57:51 - Visibility and apparent expertise
1:04:23 - AWS engineers he admires
1:06:53 - Technical book recommendations
1:09:06 - Advice for his younger self
1:10:37 - Outro
𝗪𝗵𝗲𝗿𝗲 𝘁𝗼 𝗳𝗶𝗻𝗱 𝗠𝗮𝗿𝗰:
• LinkedIn: https://www.linkedin.com/in/marc-brooker-b431772b/
• Twitter/X: https://x.com/MarcJBrooker
• Personal Blog: https://brooker.co.za/blog/
𝗪𝗵𝗲𝗿𝗲 𝘁𝗼 𝗳𝗶𝗻𝗱 𝗥𝘆𝗮𝗻:
• Newsletter: https://www.developing.dev/
• X/Twitter: https://x.com/ryanlpeterman
• LinkedIn: https://www.linkedin.com/in/ryanlpeterman/
• Threads: https://www.threads.com/@ryanlpeterman
• Instagram: https://www.instagram.com/ryanlpeterman
• TikTok: https://www.tiktok.com/@ryanlpeterman




