The Real Measure of Success in Platform Engineering | MassDriver CEO Cory O'Daniel

22 Oct 2024 · 45 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Dev Interrupted Podcast Episode Summary

Episode Title

The Real Measure of Success in Platform Engineering | MassDriver CEO Cory O'Daniel

Episode Description

In this episode, host Dan Lines is joined by Cory O'Daniel, CEO of MassDriver, to discuss the limitations of traditional DevOps methodologies, the evolution of cloud operations, and the potential of platform engineering. Cory shares insights from his provocative article "DevOps is Bullshit," addressing the common pitfalls organizations face in implementing DevOps and how internal developer platforms (IDPs) may offer a solution.

Key Topics Discussed

  • Where DevOps has gone wrong
  • The evolving definitions of DevOps
  • The role of platform engineering in 2024
  • Metrics for measuring success in platform engineering
  • The pitfalls of new hires managing DevOps responsibilities
  • Securing buy-in for an improved platform engineering experience
  • The sustainability and longevity of internal developer platforms (IDPs)

---

Key Insights and Takeaways

  1. Challenges with DevOps
  2. Lack of Progress: Many organizations still operate in outdated ways despite the rebranding of roles from Ops to DevOps and now to platform engineering. This has led to a stagnation in operational excellence.
  3. Skill Gap: There is a decreasing number of operations experts as organizations churn out developers from bootcamps without adequate training in cloud and operational skills.
  4. Cultural Misunderstandings: The idea that everyone should just "do DevOps" ignores the complexities and resources required to implement it effectively.
  1. The Evolution of Cloud Operations
  2. Changing Landscape: The way software is operated has drastically changed with the advent of cloud technologies. Many developers are overwhelmed by the multitude of tools and services that are now integral to their operations.
  3. Need for Clarity: Developers often find themselves burdened with operational tasks that detract from their primary function of writing code.
  1. Platform Engineering as a Solution
  2. Defining Platform Engineering: Platform engineering should ideally relieve developers of operational burdens and enhance their productivity. It is about creating smoother pathways for developers to deliver software without being bogged down by infrastructure concerns.
  3. IDPs: Internal Developer Platforms (IDPs) can be beneficial, but they are not one-size-fits-all solutions. Organizations must evaluate whether an IDP fits their team's specific needs and workflows.
  1. Measuring Success in Platform Engineering
  2. Key Metrics: Success should be gauged through developer happiness and the rate at which they can deliver software. Metrics like cycle time, deployment frequency, and qualitative feedback from developers are vital.
  3. Business Impact: Showing measurable improvements in delivery speed can create a strong case for investing in platform engineering initiatives.
  1. Advocating for Change
  2. Surveying Teams: Before seeking investment for platform engineering, it’s crucial to understand the pain points within the team. Conducting internal surveys can help identify major bottlenecks.
  3. Building Credibility: Start small, fix immediate issues, and showcase improvements to build credibility for larger initiatives.
  1. About IDPs
  2. Potential Pitfalls: While IDPs can offer significant advantages, they should not be viewed as a panacea. Teams must ensure that any chosen platform aligns with their operational needs and doesn't create further complications.

---

Conclusion Cory O'Daniel argues that while concepts like DevOps and platform engineering have great potential, organizations often misuse them, leading to confusion and inefficiency. To create successful engineering cultures, organizations need to focus on practical solutions, foster developer satisfaction, and measure success through relevant metrics.

---

Additional Resources

  • [Cory O'Daniel LinkedIn](https://www.linkedin.com/in/coryodaniel/)
  • [MassDriver Official Site](https://www.massdriver.cloud/)
  • [DevOps is Bullshit Article](https://www.massdriver.cloud/blogs/devops-is-bullshit)
  • [Linear B and Refactoring Survey](https://refactoring.fm/p/public-goals-career-frameworks-and)
  • [DORA Metrics Waitlist](https://linearb.io/resources/free-dora-waitlist)

---

Support the Show

  • Subscribe to the [Dev Interrupted Substack](https://devinterrupted.substack.com/)
  • Leave a review on [Rate This Podcast](https://ratethispodcast.com/devinterrupted)
  • Follow on [Twitter](https://twitter.com/DevInterrupted) or [LinkedIn](https://www.linkedin.com/showcase/dev-interrupted/)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If you would have talked to me like two years ago, I was like, I was sunshine, rainbows, kittens, all the good shit on platform engineering. because I saw an opportunity there. I saw a word that was getting popularized, books are getting written, blog posts are getting written. And I was like, this is the moment where we can take this word and we can market it correctly within our orgs to get the buy-in and maybe actually use that to actually get the f***ing work done. But what's happening is we're just, it's rebranding very similarly to ops to DevOps, DevOps to SRE, SRE to platform engineers.

0:30So you meet some teams where it's like, hey, we're doing platform engineering. We manage all the infrastructure in AWS. It's like, oh, how do you do that? And they're like, we're not software developers, so we just click around the AWS console. Oh, wow. So it's really exactly the same as it was 10 years ago. It's just a different word, right? Tell us if this problem sounds familiar to you. You want to use engineering metrics to improve your work. But every time you try, you realize there's a gap between the metrics world and your team's reality. And even with all the frameworks that already exist, is there enough context to make an informed decision about what metrics to focus on?

1:03That's why Linear B is teaming up with Refactoring. Together, we're creating a comprehensive industry report about how engineering teams use data and metrics to improve their practices. But we need your help. This week, we're launching a community survey to collect real-world stories that will help us build a practical playbook for tech leaders everywhere. The best stories will be quoted and published with the final survey results. To share your story and get a chance to help engineering leaders everywhere, visit refactoring.fm or head to the link in the show notes. Hey, everyone. Welcome back to another episode of Dev Interrupted.

1:42I'm your host, Dan Lines, Linear B co-founder and COO. And today I'm joined by Corey O'Daniel, CEO and co-founder at MassDriver. With over 20 years of experience in DevTools, SRE, PlatformEng, super excited to have you on the show today. Corey, welcome. Yeah, thanks for having me. You wrote this, I was reading it before coming out on the pod, you wrote this killer article. I was actually late to the pod because I started to click around to the different links. And of course, you know, we've talked with a lot of different, you know, pros and DevOps and platform eng and all of that. And I know you're probably doing it on purpose.

2:27You wrote an article called DevOps is bullshit. So, you know, that's me. That's you. I'm sure we'll include it. Where has DevOps gone wrong? You know, how are we failing on this or are we failing to meet business needs? Like what is your like high level take on all this? Yeah. So I think one of the key things, and depending on how far you get into the article, I think you realize this, but I'm on your side. If you've got DevOps in your title, I'm on your side. If you know that DevOps shouldn't be a title, I'm also on your side. Yeah. But the way we've gone about it for the past, what is it, 15, 16 years or so since the term was coined?

3:06I think so. We've gone about it in a pretty bullshitty way.

3:15over the past few years, like looking at like stack overflow survey, looking at number of breaches that we have, like the amount of operations experts or expertise that we have as like a community is going down year over year because we're making so many developers, you know, out of bootcamps, out of colleges and whatnot now. And we're not teaching them the cloud and operations and whatnot. Like it's just kind of, hey, you'll get there eventually. You'll learn it in prod at some point in time. Like you don't come into it knowing it like you do. Hey, I came in from a bootcamp. I know Rails. I came in from a bootcamp.

3:44I know React. I came in from a university. I know C. Well, this thing that is extremely important to how we integrate software, work with other people, we're just winging it when we get there, right? And so that's what the article is really about. It's like, we've been a bit bullshitty about the way that we've approached it. And what you see a lot is people are like, well, you're just not culturing it right. And it's just like, well, yeah, most people aren't culturing it right. You're right. But most organizations don't have the resources to culture it right. And so this whole idea of like, hey, everybody should do DevOps.

4:16We're doing DevOps. You're doing it wrong. You're holding it wrong. Like, I think it just kind of undermines all this doing the right work. And I think we need to figure out a way of getting that experience into developers' hands earlier. Yeah, no, actually, that's a good take. You were saying, hey, you know, so first of all, there's more developers in the world. Everyone's a developer now type of thing. There's boot camps, great career, all of that. very difficult to simulate what it's actually like to work at a company, release code to prod, what it actually takes to do so. I don't know if we'll ever be able to do that in a bootcamp.

4:54Maybe. Right. So I think you're right that most people coming into their first few roles, like all this stuff is new. It's not taught. And I also remember when dev, maybe it's even like, Is it like 20 years ago now? Maybe it's like 15. When DevOps, that terminology first came out, I remember some teams just being like, yeah, we're going to call ourselves DevOps now. I guess we were ops before and now we're DevOps. And so we're like cutting edge. Nice little pay raise. Nice little pay. Cool little title. Yeah. Get a couple of thumbs up on LinkedIn for the role change. It's nice. Yeah. Now you just got to do the work.

5:37Now you got to do the hard part. You got to do the hard part. You got to do the hard part. And when we're talking about like DevOps being bullshit, are you more so saying like, hey, we haven't actually like changed? Or is it like maybe we have changed, but we're kind of missing it? Like what are your thoughts there? I think there's two or three things. So I think one is for many organizations, like much hasn't changed. So, you know, in my day job, I interact with some very large organizations. The company's been around for 50 years, 100 years. And you get in there and they're like, I'm definitely an ops person.

6:16I click a bunch of stuff in the browser for AWS to build it for my engineers. And they ask for stuff through ServiceNow. Like they want a database. They ask me through ServiceNow, bug me via Slack for a couple of months and I'll get there. I'm a dev op. And it's just like, you're not. I mean, your work is important, but you guys haven't figured out how to do that. And it's like, it's still some big companies that are still there. Like we haven't moved much in 15 years. But at the same time, while we haven't moved much, the world around us has just absolutely changed, right? So there's a follow-up article to DevOps' bullshit called Elephant in the Cloud, where I go into this a bit more.

6:55But if you look at what we do as software developers, our job hasn't changed much in 50 to 60 years. We write some code, there's some ifs, there's some variables, there's some loops. Sure, the syntax changes from language to language, but writing software is writing software. If I take a Node engineer and I throw them in a Ruby code base, it'll take them a while to realize that it's sane and there's no event loop. but they'll get it, right? Now, when you talk about the cloud, the way, like everything has changed about the way that we operate software in the past 20 years. So like 20 years ago, let's say 24 years ago, we were on Metal.

7:30And then there was like the VMs started coming out. It's like, okay, we're putting things on VMs on Metal. And then there was like the slice hosts. And then there was the passes. And then there was like this like brief, like containerization Mesos moment. And then there was the Kubernetes moment. And then there's Lambda. And now there's serverless containers. and like everything about the way that we've operated software has changed. So that's interesting. That's changing on us constantly. But also the footprint of our applications has changed significantly. People are always talking about, ah, monolith versus microservices.

8:00But I think the bigger pain point in our applications is even if you have a monolith, there's a good portion of it that is secretly microservices because you're using 14, 15 different cloud services to do your thing. I've got a Ruby monolith, but I'm using SQS. Okay, well, that's interesting because now this configuring a cloud thing just became a part of my job. If I'm doing DevOps, I should be able to do that. Yeah. Okay. Now I do parity. Okay. And now I got to do compliance. Okay. And now I got to do the security around it. Okay. Now I got to become an IAM expert. Okay. And now a whole bunch of stuff got stacked on me.

8:32Yeah. And I'm just trying to build a feature. I'm just trying to fucking process a queue, man. I'm trying to get some stuff out and get some stuff done. But now all this other work came up, right? And so when we look at like DevOps 15 years ago, It was, hey, we got to figure out how to like ship stuff more frequently, get things on some of these VMs. And like the big cloud wasn't a thing then, right? And now where we are today is you have a lot of cloud around you, not just underneath your application. It's integrated into your application that you have to figure out how to manage and deal with.

9:03And that's where I think it starts to get real bullshitty and hand wavy. Like that idea of DevOps where you literally did everything or like you figured out a way to like make your team super harmonious, like that's just hard with how things are today. And I think we need to look at it in a different scope. And I think one of the things that we really screwed up with DevOps is we didn't market it. Everybody just took that title. They took that money. They took the thumbs up on LinkedIn and like some companies got in, they did good marketing. They got the stakeholders to understand why it was important.

9:35They got them bought in. They got the culture to happen. And a bunch of other people were just like, my title changed, right? And I think that is kind of the source of the bullshit. People started to get different definitions of what it is. The entire world's changing around us. And just a whole conglomerate of services are kind of getting thrown at us now to like meet global scale that like what even is DevOps? And you could argue on Reddit all day long about what it is, but no one's going to agree. Yeah, yeah. I mean, no, no, I know. I think it's right. And I remember when it was like 15 years ago, I'm looking at your blog here.

10:10I think it's actually the picture where it was like worked fine in dev ops problem now with a house on fire in the background. Like, I think that's what like the original DevOps, at least to me, the meaning was where it was like developers are throwing things over the wall to ops and then there's no good communication. I don't even think that is really what it is now. Now, what caught on when you were talking is it's almost like you're saying I'm trying to get this feature done, trying to get I'm a developer trying to get to the feature done. But now there's all these other things on my head that I'm expected to do.

10:47Are you seeing out there that it is like on top of the developer to actually like coordinate all this stuff, know all the technology, everything that it takes to get off to like a production? or are you seeing it more like, okay, if you're doing DevOps like properly, in your opinion, it looks different than that. Yeah, I mean, I think that's part of the problem is like so many people, like we see so many different teams that like when they're like, hey, we're, yeah, like I'm the DevOps team or we do DevOps. I have no idea what you mean until like 15 minutes into the call. I have to like, I have to see your world.

11:23I'm like, oh, I get, yeah, we can call this DevOps. That's a little DevOps too. I get it. Like, you know, like you're just doing it your way. Right. And like, and so you do see teams where it's like, I'm responsible for everything. We've shifted it all left. And like, that's, that's the rub is like, there's, there's too much to shift left nowadays. Yeah. Right. And like, and that's kind of where, you know, in the post, I moved towards platform engineering and it's like, now there's caveats there too. Like there's some bullshit in platform engineering. There's a lot of bullshit in platform engineering too.

11:55But like what we need to be doing as teams that have the resources to do DevOps or platform engineering effectively and as providers like MassDriver, AWS, et cetera, is starting to build more of this toil and tasks that devs are kind of getting caught up in into their platforms. Whether you're building it or whether you're buying it, like we need to start doing more of that because just shifting everything left onto the developer isn't working. It's not what you pay them to do. You pay them to build features to build, generate revenue. Yeah. pay a developer make value yeah fiddle with checkoff right like it's like go do something makes value right at the same time we have to remain compliant and secure and all this other stuff like whose job is that right yeah and you might say the security team but now you're tapping and you're waiting on somebody right and so like like the whole thing it's just like like we don't have we got to grease the gears of of the way we're doing this right everybody's becoming a software company now there's more stuff moving to the cloud and we're not building out this expertise.

12:51And we're either saying, slow it down and tap that guy, or it's all your problem. And neither of those I think are effective strategies for scaling in the cloud, scaling you or people in the cloud, not necessarily your compute. I would love to pick your brain on platform engineering. I know it's another word, so it could be bullshit as well, but there's like DevOps, there's platform engineering. When you think about what platform engineering is, can you describe to like what comes to mind for you if you think of like some elite platform engineering team like how does that look to you yeah so this is where the bullshit starts to come in because the exact same things happen with platform engineering so if you would have talked to me like if you would have talked to me like two years ago i was like i was sunshine rainbows kittens all the good shit on platform engineering because i saw an opportunity there i saw a word that was getting popularized books are getting written blog posts are getting written and i was like this is the moment where we can take this word and we can market it correctly within our orgs to get the buy-in and maybe actually use that to actually get the fucking work done.

13:57But what's happening is we're just, it's, it's rebranding very similarly to ops to DevOps, DevOps to SRE, SRE to platform engineers. So you meet some teams where it's like, Hey, we're doing platform engineering. We have this, you know, separate part of the business where it's like, we're building this thing as a product and we're giving it to our developer customers to use. Like they're doing like the platform engineering. And you meet a lot of folks where it's like, oh, what do you do as a platform engineer? It's like, ah, yeah, we manage all the infrastructure in AWS. It's like, oh, how do you do that?

14:23And they're like, we're not software developers. So we just click around AWS. Oh, wow. So it's really exactly the same as it was 10 years ago. It's just a different word, right? And so when I think platform engineering, I'm thinking of the company that you work for has decided that they've gotten to a level of maturity in the DevOps maturity model where they can actually do DevOps well, and they're ready to take that to the next level. We need to pull more of this responsibility off the devs. We need to start building something that feels a bit like maybe Heroku, like it feels like a pass. It feels like a platform.

14:59It feels like I can write code and it just goes and runs. I can build it and run it without taking on 400 other jobs. Like that's the goal of platform engineering. I think that's the distinct difference between this and where we were 15 years ago because we didn't have the cloud then. And we see a lot with platform engineering is infrastructure provisioning, being a part of it, talking about the cloud, talking about getting self-service of cloud resources, which wasn't like the terms that we were using 15 years ago. But these terms really matter now. As a developer, I need a queue. Do I need to know how compliance, security, IAM, doing parity with Terraform work?

15:34Do I need to learn HCL? No, I just need a queue. Why isn't it easy? I can do it in the AWS console. So how do I do it at work without tapping you or going through hell and back to like set up a bunch of CI CD pipelines to like do it right. And I think that's, that's what platform engineering is. We're mature enough. We have the resources, we have the buy-in, and now we're building a product to solve the biggest problems of our developers cadence and delivery. I, that really clicked for me the way that you said it. Cause I do think like, uh, if you have a platform engineering team or you have a dev I've seen, like I always, I'm the type of person that has to go back to like what my mission is.

16:12Otherwise I can't figure out like what, what my purpose is on earth. So I got to know my mission. And what you said that I thought was cool is remember how it feels like to use a pass as a developer. Now, of course you're boxed in more and it's not, you know, you might not be able to use everything that you want to use, but the cool thing about it was you could just code and then like the other stuff just works. Like the app just runs, the code makes it to where it needs to be mated to, the customers get it. It's like super simple. That's the coolest thing about it. And I think you said something like, if you're a great platform engineering team, the vibe that you're giving to your developers that you're servicing, that's kind of how they feel for your company.

16:59They can just do their job. Like that totally resonated with me. I feel like you should be able to take that team and be like, man, this platform engineering team has done such a good job. We're spinning them out as their own business. To me, I'm like, man, that team is killing it right there. If that's the way you feel about them. But there's also many of these teams where it's just like, hey, we manage Kubernetes for everybody. And this is a place where I see like, I'm a Kubernetes fanboy. I develop on Kubernetes locally. I like the thing, but there's a lot of people that are just like, oh, we're platform engineers.

17:32Kubernetes is our platform. It's like, no, no, no, no, no. That's not it. Like if you think about like what Kubernetes is and how big it's getting and how its control plane is expanding into controlling the cloud as well, Kubernetes is just kind of another cloud API, right? It's just another cloud. That's like saying, oh, our platform, we have a platform engineering team. They manage AWS. It kind of starts to feel the same, right? And you can't just like throw a Kubernetes at a developer and be like, we platform engineered it for you. Like there is DevX that needs to happen around it. There is some smoothing out.

18:03Like if your developers are thinking about Kubernetes all day, I'd say that you haven't done platform engineering. You're still in that world of like mixed, like dev-y, op-sy bits, right? Yeah. Right. But that's, but that, that is interesting because it also starts to get, it gets interesting, right? So it's like, you see these companies where it's like, they have a team that just manages Kubernetes. It's their job. And that's a hard job. There's a lot of work to do there, right? But like, what is that team? Yeah. Are they DevOps team? Are they ops team or they, you know, the compute management team, right?

18:33Like, but, you know, calling it platform engineering, we're actually doing engineering work to build out this stuff that's going to serve these engineers well. Like, I feel like that's a, just another misuse of the term, but that's where we are. That makes sense. That's fine. Um, so I'll be writing a blog post here in about two years called platform engineering is bullshit. And I'll have you back on the pod. Yeah. Hopefully we'll find another really good word to use that like actually sticks this time, but probably never. Yeah. I think whatever, whatever word it is, it will be weird. That's the thing is like, you know, it's really hard.

19:07It's so easy to like write a blog post or like hop on a podcast and talk about like the way things should be. But the reality is, is like the real world's messy. Production is messy. Organizations are messy and your teams are probably messy and your code's probably messy. Like there's, there's a ton of mess, right? So like having this, like, like this golden idea or a golden path, it's like everything's fit into this perfectly. It's like, it's so unrealistic, right? So when you start looking at orgs, you'll see orgs that like, you'll see the Googles of the world where they went and did their platform engineering and it came out as something called GCP, right?

19:38And then you'll see other companies like we do platform engineering and it's very much like, A, we write Ansible scripts, right? Like the way this is servicing in many companies is, looks very different, very similar to DevOps, but it's because our companies are different, right? Like think about the average, like Series A company, Series A, Series B, you can't give them a platform, right? I can't go to like a Series B company and be like, everyone's using Heroku now. From now on, that's what you're doing. Sure you could. There's some architecture decisions you're gonna have to deal with, but then your team that's working on like Next.js, Jamstack stuff, they're not gonna have like the greatest experience on Heroku, right?

20:16They probably want something more like Vercel. Your AI and ML team's probably not gonna have a great experience on Heroku, right? Like you're not going to find this like one thing that's perfect that works for everybody. And that's the other thing that's kind of problematic about this is like that happens at the team level, it happens at org levels. And I think that's why you're seeing like these very different definitions of what platform engineering is. It's people sitting around thinking, I got to make something happen and this is the way I'm going to do it. It's going to be, it's just Kubernetes or it's this complicated software thing that we built.

20:45And some teams are doing it and they're doing it all in CI. Like you can get pretty far doing platform engineering and infrastructure provisioning and self-service with just GitHub actions. Like you can do it. It's cumbersome to navigate, but you can do it. You can deliver self-service and DevEx that way. But so many questions now. I mean, certainly the takeaway there is like platform engineering is specific to your organization or like the mission of what your engineering team needs to accomplish in order to sell the product that you're probably creating for your business, like that totally makes sense to me.

21:21And then I think I would say like the level of eliteness that you go, you know, I think it's about the experience. Like I keep going back to what you first said, like if I'm like the head of platform engineering, even if I'm like in a startup or a midsize company, like how am I measuring my success? Like how would I rate myself or like what what what is a metric that i would use in order to say like i'm in the right way or am i just doing bullshit stuff yeah i think this one's this one's an interesting one and my answer is going to sound a little bullshitty unfortunately but bear with me you can go and define golden signals and sre stuff you need to like those are those are things that you have to measure when you're running a platform those are things you should probably be measuring when you're running a product for users.

22:11Like if you got users, you should care about the thing being up. I'm going to kick all that shit to the side. I'm sure people are stoked about measuring that. Two things matter. Dev happiness and how unblocked people are. Like that's in the grand scale of things. I think that's what really matters. Are developers happy with the way that the thing works? Is it solving a real problem for them? And is it improving the cadence and delivery of our software? And that second one's a little harder to measure sometimes because Sometimes you'll have teams that like the stuff's real stable, right? Then maybe they're not releasing 14 times a day.

22:45Maybe they're releasing once a week and they're making billions of dollars, right? So like, what does that stability look like? What does that release process look like for those developers? Is it breaking often? And like, we need to make that a bit smoother. But that's really, I think what it comes down to is like that, the qualitative happiness of the developers using the thing. And then that delivery experience, am I shipping software faster? in my shipping software with less like build failures of getting the thing out. Now you can come in and like come up with like plenty of other ones. I think you could, again, like this comes back to the org.

23:19Like it really depends on like what the mission of this platform engineering team is. And since it can pan out so many different ways in orgs, it's hard to like pinpoint it. Like those are the two that I find most important, but you'll see other people where it's like, hey, one of our goals is like minimizing cloud costs. It's a weird one, but like we've seen it, but it's a real problem. Like developers have just been clicking stuff forever, or you have an ops team has been doing it for devs. But since they're kind of a world apart, like there's a bunch of stuff in prod that they don't need anymore.

Read the full transcript

23:47And it's like, hey, by moving to platform engineering where everything's self-service and like the devs kind of own it, but aren't burdened with the, you know, the actual like monstrosity of it all. Like, can I get my cloud costs under control? Yeah, you probably can. And that might be an important measure for you is to like get that under control and keep it stable or maybe only see some growth that's relative to like your traffic or whatnot. But I think what it really comes down to, like how you know that you're doing a good job is people are stoked and people are doing their jobs better, faster with less breakage.

24:16Yeah, I think it makes sense. I think for our listeners, they already know this about this pod and linear B and all that. But yeah, like on your second point of the measurement, they're coming to us, platform engineering teams. I need to measure wait time for my developers. I need to measure cycle time. I need to see how easy it is for their code to get out into production. What phase is it stuck in? Where's the bottleneck? Like that's the type of metrics that I'm familiar that they're looking at. So that's like the quantitative side. And of course, on the qualitative side, like you said, you know, how stoked are they?

24:53do they want to keep working here? Would they recommend, are they saying like, yeah, other people should come work here. That's like the qualitative, like a survey side. And you put both of those together. And I think you have like a pretty good picture, but I did, I did ask you, cause I do think it's like, I I'm always trying to get in my mind. If you're a listener here to say like, okay, I have a platform engineering team or I'm on a platform engineering team. Like what does good look like? What does good actually look like? And if I know like the criteria to measure, like for example, for us, you know, at Linear B, we want to make sure I think our benchmark is like 24 hours or less to go from coding to production.

25:34Like if I'm coding something and I need to get it out to prod, can I do that in 24 hours or less? Like that's what elite looks like in the industry. What does not good look like? Five days, six days, seven days. That's the type of thing that we're highly familiar with. Yeah. And that's another fun one. Like, I don't know what the audience is like, like what stage of business. All over. Yeah. I think it's like all around. Like you got some listeners in startups, you got other ones that are like, hey, I'm at a big enterprise, but I've been tasked with, I need to start a platform engineering team or a developer experience team.

26:10So what's really interesting is like, you know, I've been in this like DevOps world for 15 years. I've been on ops and the development side. I don't think I've worked at a company in the last 15 years that people couldn't get a feature out on their first day. That to me is like a good dev team. It's like I should be able to run a command or two and like my environment's up locally and I can ship something to prod on day one. It doesn't have to blow the world apart, like away, like just get something in there. So like you feel like you've contributed to this thing on your first day, right? Like that's a great place to be.

26:40Now, when you say like, hey, sometimes it takes five days, six days. Some people might be hearing that and be like, that's forever. Some people are probably listening to this thinking, fuck, I wish we could get it down to five days. Like we're sitting at 30. Those people that are like five days is a long time hear 30 and they're like, 30 is a real long time. How often is this happening? Well, here's one of the things that's really interesting, like going back to like DevOps and bullshit and like how we're still kind of where we were 15 years ago. If you look at the Dora report, it's got DevOps in the title.

27:0950 % of the respondents that are tuned into responding to this report say that an outage can last them up to five days. And the average time of deploying software sits around once every 30 days, 50 % of the people that are tuned into the door report are saying we are not doing this well. Right. And like that, that's, that's one of those numbers. Like, again, we can't sample the entire world. We can't survey everybody. But like, when you start to see these consistent patterns around like just ops skill vanishing. That's, I think, the power of platform engineering. Again, whether it's in your org or whether it's somebody like MassDriver, Humanitech, Covery, like all the people in this space, like we're trying to make these operations teams more efficient.

27:52We need to scale those people. That's our scaling problem in 2024. It's not scaling compute. It might be scaling your GPUs, but it's scaling these operations folk. Everybody needs them. A lot of organizations struggle to hire them. A lot of people are just hiring people straight out of college and saying, you're the DevOps person now. That's such a common trend on Reddit. That's the trend? Saying you're the DevOps person now out of the school? Oh, my gosh. That's scary. Go on Reddit slash DevOps. And the amount of people where they're just like, hey, I just got my first job and I'm the DevOps engineer.

28:23And it's just like, oh. It's like you've never released code before in a professional environment. You're the DevOps person? Yeah. And it's like, hey, we got like 8 million GPUs that we're managing. and it's just like, oh, you got just tossed into a whole world of scale there. I feel like I see that almost every day or two on like Reddit DevOps and somebody's just like, I'm brand new to the career and they put me in this role. Okay, that's, take like a detour for a second. Why do you think that, is it no one wants to do this? Like, why is that happening? No one wants to do this job. Nobody wants, it's like, that is the scariest and oddest thing.

28:59I know I'm more old school because when I was developing, it's a while back, but we definitely wouldn't have been like oh hey new hire you're responsible for operations or like yeah that's like no way so here's here's here's one of the things that we see that's kind of scary in like the hiring so so one of our early go-to-market strategies was looking at people that were hiring their first ops role and so mass driver like we so we help people with cloud operations platform engineering gotta sell that stuff but we were looking at people that are hiring this first role. I'm like, okay, great. Let's go and mine a bunch of job descriptions.

29:33It's a great way to figure it out. Job descriptions are actually really funny. I don't know how many hackers use this, but you can get a pretty good idea of people's infrastructure just looking at their job descriptions, right? And so when you see these people that are hiring their first ops person, nobody's hiring an ops person proactively. Nobody's saying, hey, you know what? We're almost done with our MVP. Let's go ahead and poach a DevOps person from Google. No, they think about that first DevOps person when the world is falling apart or like people are mad and they're like, we should probably get a DevOps person.

30:06So now you're coming into a job. Series A, there's 20 developers and they're like, we got 20 developers and you're the ops person now. We haven't had one. There's no way in hell I'm applying for that job. Like, I don't care how many, well, maybe there's commas that you could convince me, not zero. There's got to be multiple commas in there. But coming into that job description as a seasoned ops person, you're like, I'm going to have 20 customers day one. There's five years this company has been in business, five years of debt racked up. We're looking at some of these job descriptions of what they expect out of you in this first operations role.

30:40It's like, that's a lot of work. I know there's a lot of debt. There's going to be 20 frustrated people day one. And you have to be a certain type of person to be excited about that. And also you see with a lot of these orgs is they're not paying Google salaries. Yeah. Right. It's like, okay. Like they got this DevOps role and it's$85 ,000 a year. Yeah. So it's like, it sounds like it's coming from like a desperation without the proper investment. That's the way that I would, I would put it desperate, but not properly invest. Okay. Here's another thing that was coming to mind. So let's say, you know, we know we have a problem.

31:15developers are complaining it's tough to get code out to like you know some of the basic stuff and i'm trying to advocate for platform engineering and i'm trying to go to the business and saying hey we have a real issue here how do you get buy-in or investment into this like do you have any advice around that i do so first off is i feel like people are reaching for the term platform engineering when they might be still want to be reaching for like that devops mindset first, right? But the unfortunate thing is like, this is the hot term. It's the term that that like mid-tier stakeholder might be like, oh, I heard that.

31:52I saw it in Hacker News. Like, I want to have that on my resume and be responsible for it. That sucks. Like that sucks as like an in for it, right? So like, but that's unfortunately like the term that people are hooking onto. There's a lot of times where you see these orgs and they're like, hey, we need to do platform engineering. And it's like, yo, you got to get just good dev and ops and CICD practices down just in general, right? And so if we got to use the word platform engineering to get the job done, I'm fine with that. But what I hope people don't do is go out and think that they have to build a platform to do that.

32:23Because many times it's just investing in that CICD process. The amount of times we've seen people like, oh, we're moving the platform engineering. Then we get in and talk to them. We're like, what's your biggest problem? And they're like, oh, like our Docker builds take 45 minutes to an hour. And I'm like, it's not a platform engineering problem. That's a, you don't know how to write a Docker file problem, right? Or you don't know how to use build caches problem. It's like, you can do a lot of that work just like fiddling around in Docker and GitHub actions, right? And, you know, you'll see other folks that are like, hey, we're getting into platform engineering.

32:51It's like, oh, what's your, what's your biggest problem? It's like, we got a lot of cloud resources and we don't know what's used and what's not. It's like, okay, well, that's tagging conventions and naming conventions and like, and like generating some reports from the AWS tools. Like you can call it platform engineering. That's what you got to do to get there. But like, Until you've gotten to that point where you're like, we've got ops in a good place. We're fast, but now we need to be faster. We need to automate the automation. We need to stop thinking of automation and start thinking of functional software.

33:21I think that's when you need to move to true platform engineering where you're building a product. But I don't think many organizations have that need. They just need their developers to be happier and move faster. Same goals as platform engineering, but much different effort involved. And so I think the first thing you got to do is take a step back and be like, okay, am I using platform engineering just for the name? Or do I actually need to do like platform engineering? Or do we just need to be better about operations and software development in general? And know what you actually need. Here's how you're not going to do that is you thinking you know what you actually need.

33:57Like step one is take that initiative. And I think that's one of the best things you can do. If you're sitting in this role and you're like, everything's kind of fucked. Take the initiative. Don't go and ask a mid-level engineering manager like, hey, do you think we could get a budget for a platform engineer? Make a survey. Make a survey. Survey some people. The survey part is hard. That's one of the skills that we're missing as operations folk moving into platform engineering. But figure out what your team is stressed with. If you've got five developers, grab them all and be like, what is the thing that slows us down the most?

34:28And just see if you can solve that. Right? And if you want to call it platform engineering to do it, that's fine. But solve that problem first. to make people happier, show that there is success in your initiative and find a KPI that gets the business excited. Hey, we just increased delivery time by 20%. Oh, faster features means more money. Yeah, yeah. Imagine if we could get the rest of this stuff unblocked, if we could move a little bit more of that debt and get these people moving faster. And you still might not be doing platform engineering. Like you might just be doing another small task that you need to focus in on.

35:01Like, I feel like so much of like - You can build up to it. You can build up to it. But I think, you know, as far as like getting in and getting people excited, I think it's, it's really understanding like what your business needs to move to the next level of delivering software. And sometimes that's cadence. Sometimes that's security. Sometimes that's breaches. Sometimes it's outages. Right. And like, that's, that's what I think is most important is like figuring out what that is and how to solve it. Like build up your, you know, your credit. Yeah. Did a good job fixing that one. Like, okay, now, like our biggest problem now is we're moving the microservices.

35:35We're breaking this monolith up into five or six pieces. How much of a pain in the ass is it going to be get all of this stuff moving in the cloud? Okay, now maybe it's time to start thinking about, let's bring in Kubernetes, create a nice little abstraction plane for developers where they're just thinking about their Docker files and not all this YAML. Like baby steps to get yourself there. Like, don't feel like you have to, you know, get a new repo and like build something from scratch that looks like, you know, MassDriver, Humanitech, Recovery, like you can get started on fumes. You know, I, I, what I really, I mean, that, that makes sense.

36:05What I really liked about what you're saying is like, show some improvement before asking like, Hey, this is what I was already able to do. And at the end of the day, you know, from my experience, like the customers that we work with, what the business cares about when you go to advocate for yourself. Well, first of all, one piece of good news, businesses are very hot right now on developer experience and making developers more productive. Salaries are high. Like you can probably get some buy in there, but they care about delivering projects on time. Like if you can ever come and say like, hey, you know, I talked to these five engineers, found what the bottleneck was, fixed the bottleneck.

36:49And they said this was the reason that that one of the biggest reasons the project was successful is you unblocked us. Now you can go to the business and say, I'm Matt. And usually, you know, you could be hiring right now. Hey, you want to bring on 20 more devs? This is like mandatory. We can't bring any more on without doing this. if you'd like to deliver projects on time. Like that kind of stuff, like building up that argument works. But I do agree that starting with like, hey, I did something that helped the developers and I could do more if you'd like it. I think that works pretty well. Yeah, I think there is a challenge in that statement though, right?

37:26And I think there's probably a number of people that are listening to things like, how the hell do you just do it? And this is one of the things I think that sucks about a lot of engineering orgs, right? Like if we're building bridges, right? Like, let's say we're actually, we're real engineers, you know, we're building things and measuring stuff and making sure people don't die under it, right? If I was working on a bridge, if I'm an architect and I'm like, you know what? I work on some like load bearing beam calculations. And like, while I'm looking at these blueprints, I noticed that the bridge is made of sand.

37:57I probably mentioned something. I'd probably be like, hey, that's not right. It might not be my job per se, but my job is to be a professional. My job is to do engineering work, right? And like while we call ourselves engineers in software development, many teams, and this is not a dig at you. I know it's your organization. Many teams are task doers. They're people that write software to do a task that was assigned to them, right? And our jobs as professionals is to do work for business people that don't know how to do the work, right? We're a tool at the end of the day. Like we make software.

38:32They have an idea. We print out some software. And if we're not going about it in a professional manner, like addressing debt and refactoring, right? Like, are you really doing your job, right? And so for the folks that are hearing this and thinking like, oh, how do we just go and do it? It's your job to just go and do it. And I know that might be hard in your org, but like debt and refactoring and solving problems and making your development team faster is a part of your job. And so I have people ask me all the time, like, how do you get time for, like, to address technical debt? How do you get time to, like, refactor code?

39:07And it's like, I just do it. You want me to build this feature. There's something in my way of doing it faster. I'm going to refactor and fix that so that I can deliver this feature. Now, I think the catch there is, there's a lot of people that get caught up in, like, the debt, and they're like, I fiddled with that for two or three months instead of doing what you asked. And I never delivered the feature. Yeah. Yeah. It's like, if you can do what they ask, and in the meantime, like, you've delivered some success, that's good. and talk about it. Be like, hey, I refactored this. I went in and added this caching thing that makes our builds faster.

39:37Like advocate for the work that you're doing and be like, all of this is happening because I did some refactoring or address some technical debt that was in my way. We should be addressing technical debt and refactoring that's in our way. And what you're going to see is you're going to start moving up that DevOps maturity model. You're actually probably going to start doing a little bit of DevOps, right? Like, oh, I went and did some stuff to make this whole thing operate a little better. That felt nice. And advocate from there. I'm not a big fan of, if somebody wants to slap my hand for making a build pipeline faster, slap away.

40:07Because I know everybody else is going to be stoked about it. It's like, I would rather just do it and be like, eh, sorry, I did something you didn't ask me to do, but it made things better. So chill, feature still delivered, back off. You'll get a lot of backing. So it's kind of like a low risk. So we are coming up on time on the pod, but because I was reading your article. It's a good article. Everyone go and check it out. the DevOps is bullshit. I got to ask you a minute about IDPs, internal developer platforms. What is your, I mean, we could do a whole pod on this. So keep in mind, we don't have that much time, but like, I think they're, they're pretty hot right now.

40:46Who I know is using them. It's usually like our larger customers. I don't know if that, that isn't me. I'm not saying that's the only one, but the people that I talk to that are using them, larger customers, more so on the enterprise for us, lots of services, also have large DevOps or platform engineering, whatever you want to say, large teams. But I wanted to just get your take on IDPs. Who do you see using them? Who should use them? Is it all the rage or is it just a fad? Can you give us a little something there? I mean, I think IDPs are important. I think they're valuable, right? I mean, at the end of the day, what they are is they're an internal developer platform.

41:31What is that? It's a pass that you own, right? I mean, and there's a couple of them that are open sourcing and install it in Kubernetes and you get it up and running quickly. And I think that if that thing solves a need for your team, it solves a problem, they're great. I feel like there's a lot of people that are just like, oh, if I get an IDP, that's going to solve my problems. And it's like, no, no, no, it's, it's, it's not like, right. Like it might even create more. And so I would just say like, you know, is that IDP the thing that you're willing to like tie your business to? And the thing that was like a little goofy about it again is like going back to that, let's say you're a, you know, series D company and you're like, Hey, we're migrating the entire company to Heroku.

42:12You're probably going to have a fair number of engineers that are mad about that. Right. And so one of the things that we've always kind of had this idea of when we think about IDPs is like, you need a developer platform potentially per team. Your teams can be very, very different. And so the mass driver approach is IDPs as cattle. IDPs as cattle. Every project in mass driver is its own IDP. And it is very fine-tuned to your team. So your team kind of comes in and says, hey, this is our makeup. we're doing model builds and we're serving our application off of a Lambda. That's great. That's great.

42:50That is your internal platform for how it's going to work. Your services are going to bubble up. You can see them, but we're not jamming that same platform down somebody else's throat. Right. And so when you start looking at some of these IDPs, they're Kubernetes specific, they can run pods, stable sets, workloads, et cetera. That's great. Until you have someone on your team is like, oh, well, we do a bunch of stuff on Lambda. How do I use this IDP? It's like, oh, you don't. Okay. Well, that's interesting because we just moved to platform engineering. We just did this IDP thing. It's going to make everybody more efficient.

43:18But now this whole group of people just has to still solve the problem on their own. Yeah, interesting. That's not great, right? So, I mean, I think that if you've got a pretty like homogenous team and way of delivering software, you might find an IDP that's perfect for you. I think in reality, like as you start looking at your teams, they're going to have very different ways of doing things. And you're likely going to invest in one of these IDPs and then find yourself looking for something else. Yeah. And I say this as a platform that like people have come to us and they're like, hey, we're using this IDP, but like what like we need like our developers need to manage infrastructure, too.

43:53And so then they're like looking at MassDriver and I'm like, why don't you just get rid of the IDP thing? Because like we we do that part, too. And we do the infrastructure management stuff as well. But it's just funny because like people come to us and they're like, we bought this thing and we thought it was what we needed. And it turns out it only solves like 30 percent of our problem. And it's like, yeah. Yeah. So maybe this is then the next pod that we talk about, because, you know, I do think the IDP, like I'm hearing a lot of buzz about them and it'd be great, you know, to get into the details of like what it really solves and doesn't solve.

44:23We know it's not a cure all. But Corey, it's been awesome having you on the pod. I think a really fun conversation. Thanks for coming on the show, man. Yeah, I appreciate it. Thanks for listening to my rambles. and listeners, please subscribe to the Dev Interrupted Substack channel if you haven't already. You'll get access to our weekly newsletter and exclusive articles. So hope to all see you there. And then Corey, one more time. Thanks, man. It's been awesome. Yeah, thanks so much.

From the publisher

It’s time we recognize the idea of a ‘Golden Path’ is unrealistic.

In this episode of Dev Interrupted, host Dan Lines is joined by Cory O'Daniel, CEO of MassDriver, to discuss Cory’s provocative article 'DevOps is Bullshit'. They cover the pitfalls of DevOps, the evolution of cloud operations and whether or not platform engineering is the solution the industry needs .

Cory shares insights on why many organizations struggle with DevOps implementation, the impact of cloud technology on traditional operations, and how internal developer platforms are reshaping the industry.

Topics:

  • 01:03 Where has DevOps gone wrong?
  • 04:19 Have we changed or does DevOps mean something different?
  • 11:51 Platform engineering in 2024
  • 20:10 How can platform engineering leaders measure success?
  • 26:40 Why are new hires being put in charge of DevOps?
  • 29:53 Getting buy in for a better platform engineering experience
  • 39:13 Are internal developer platforms a fad? 


Show Notes:

Support the show:

Offers:

More from Dev Interrupted

All 208 episodes
The Real Measure of Success in Platform EngineeringDev Interrupted · 45 min
Listen in VO