In short
Dev Interrupted Podcast Episode Notes
Episode Title
The Rise of Cloud Costs, and How to Optimize Them | Plaid’s Mark Robinson
Hosts
- Ben Lloyd Pearson (Guest Host)
- Mark Robinson (Infrastructure Engineer at Plaid)
Episode Overview
In this episode, Mark Robinson discusses the challenges and strategies behind managing and optimizing cloud computing costs. He emphasizes the importance of understanding cloud billing, identifying overspend, and making optimizations to achieve significant savings, sharing practical tips from his experience at Plaid.
Key Quotes
- “If you're not careful, cloud computing can lose more money faster than any invention in history.” - Mark Robinson
Episode Highlights
- Understanding Cloud Costs
- 00:56: Discussion on how cloud computing costs escalated.
- 02:34: Importance of digging into your actual costs.
- Cost Breakdown
- 04:55: Accounting for various cloud services and their costs.
- 07:35: Identifying areas where organizations can derive the most value.
- Code Quality Relation
- 12:26: Connection between cloud costs and the quality of code produced.
- Overcoming Blockers
- 16:08: Challenges organizations face in achieving cost savings.
- Leadership Buy-In
- 19:32: Strategies for gaining support from leadership for cost-cutting initiatives.
Discussion Points
- The Complexity of Cloud Costs
- The transition from fixed capacity to flexible cloud resources leads to spending without strict controls.
- Examples of hidden costs due to network fees and data transfer charges.
- Practical Steps for Cost Optimization
- Start by analyzing cloud bills using provided consoles and identifying where the money is going.
- Importance of tagging resources for accountability and tracking costs by individual services.
- Organizational Challenges
- Gaining buy-in from developers and leadership is crucial.
- Highlighting individual contributions can incentivize cost-saving behaviors.
- Technological Solutions
- Existing tools can help analyze costs, but they may be expensive and provide generic advice.
- Mark emphasized building internal tools to derive actionable insights from cost data.
- The Relationship Between Cost Efficiency and Code Quality
- As costs increase, developers might resort to simply adding resources rather than optimizing code, which can lead to escalating expenses.
- Resistance to Change
- Both developers and leadership may resist cost optimization initiatives due to competing priorities (e.g., feature delivery).
- Successful initiatives often require demonstrating clear ROI and savings potential.
- Cultural Shift
- Making cost optimization a company-wide value is key, supported by management commitment and resources.
Final Thoughts
Mark Robinson concludes with the notion that cloud optimization is not just a technical task but also involves a cultural shift within organizations. He emphasizes the need to make the right decisions easy for engineers and to create an environment that rewards cost-saving innovations.
Resources Mentioned
- [Essential Guide to Software Engineering Intelligence](https://linearb.io/resources/essential-guide-to-sei?utm_source=Substack&utm_medium=referral&utm_campaign=202405-SEI-IMC)
- [FinOps Foundation](https://www.finops.org/)
- Mark Robinson's LinkedIn: [Mark Robinson](https://www.linkedin.com/in/mark-robinson-944084b/)
Offers
- [Start Free Trial of LinearB](https://linearb.io/start-free-trial?utm_source=podcast&utm_medium=referral&utm_campaign=devint-shownotes&utm_content=shownotes)
- [Book a Demo](https://linearb.io/book-a-demo?utm_source=podcast&utm_medium=referral&utm_campaign=devint-shownotes&utm_content=shownotes)
---
This summary encapsulates the episode's key concepts, discussions, and takeaways while providing clarity and structure for readers interested in understanding the dynamics of cloud cost management and optimization.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00There's a million little blind alleyways that you're spending on that you don't even realize. Like there's just so much, someone set it up, it must be right. And you go back and look at it and be like, no, that's just wrong. Amazon bills you a lot for networking fees. But by default, they send all your S3 traffic out to the internet and will then bill you for that. But if you flip this one little toggle, it goes internally. And it's such a huge change. It'll save you a fortune if you do a lot of S3. We actually had our account rep call us up and ask if we were down because our traffic drops so much.
0:34That they're like, okay, something must be wrong because obviously it was set up correctly before and now all this traffic is missing. And we're like, no, no, we did this on purpose. Yeah, I actually have a phrase for this I call disabling the suck. Yeah. Are you looking to improve your engineering processes and align your efforts with business goals? Linear B has released the Essential Guide to Software Engineering Intelligence Platforms. This comprehensive guide will walk you through how SEI platforms provide visibility into your engineering operations, improve productivity, and forecast more accurately.
1:07Whether you're looking to adopt a new SEI platform or just want to enhance your current data practices, this guide covers everything you need, from evaluating platform capabilities to implementing solutions that drive continuous improvement. Head to the show notes to get your free copy of The Essential Guide to Software Engineering Intelligence Platforms today and take the first steps towards smarter, data-driven engineering. Hey, everyone. Welcome back to the Dev Interrupted podcast. I'm Ben Lloyd Pearson, the Director of Developer Relations here at Linear B. I'm joined by Mark Robinson, an Infrastructure Engineer at Plaid.
1:40Mark, thanks for showing up today. Thanks for having me here on the Kona Islands. Yeah, the Kona Islands is a great name. We're actually going to have to put that on the sign on the outside. So you bring a wealth of experience to the podcast today. So apart from your time here at Plaid, you worked at places like Uber, among other pretty big names. But your work at Plaid is what really caught the eye of our team here at Dev Interrupted and is why we wanted to bring you on the show. So you spent about five years at the company, right? That's right, yeah. Yeah, and at that time, from what we learned, you helped save 25 % in cost, you would just say?
2:15Yeah. optimizing existing resources and removing waste. In fact, I guess this has become a bit of a personal crusade for you because there's even a quote out there from you that says, if you're not careful, cloud computing can lose more money faster than any invention in history. So let's start there. How did cloud computing get so expensive? It's always these little things. A lot of the way that cloud computing is sold, Like the initial cloud computing was just here's a data center. We'll rent you machines. And then sort of the next revolution was we'll sell you individualized services. But instead of paying like these flat rates, like you buy this much capacity and then that's your limit, it'll just scale up as much as you want.
2:59So there is no forcing function to push people back. So it's, oh, I'll do this. I'll do this. I'll do this. And there's a lot of really subtle bills. Like if you've looked at Amazon's pricing sheet, it's horrifically complex. like it's okay you buy the server and then you buy the data transfer and then you buy the data storage and then you buy the data access to the data storage and it just starts adding up really fast traditionally a lot of places you just have to say like okay we're at capacity it'll take three months to order another server deal with it now it's just like oh spin up more launch in another data center store more data you don't run out of resources so you just keep spending that's fascinating because it's it's one of the probably clearest examples i've ever seen of a slippery slope, right?
3:43It's like, you know, people originally got into cloud computing because it was a way to save costs versus, you know, buying everything yourself. But now it's just the sky's the limit, you know, right? So what made you start this journey into understanding how to save costs? Was it motivated by external factors or was it more of like an internal drive? It was really sort of an internal drive. So several years ago, we would have like regular hackathons. And one day I looked at our bill and like, I bet I could take that down a few percent. So for the hackathon, we just decided like, OK, yeah, let's dig into this.
4:17Got a few other engineers and we started finding just tons and tons and tons of stuff. Like I thought maybe we could do 3 % and we found 15 or so percent without that much work. And when you start going up to management and saying like, oh, yeah, if we bring this down, that's like, you know, headcount to hire. that's other things we could buy. This is just money we're setting on fire. We could have an office party for this. Beyond having a hackathon and making that a personal goal, what is step one in this journey? If you suspect maybe that your costs are getting out of control or maybe you don't actually even know, what's step one in this journey if somebody wants to start finding ways?
5:00You probably want to start by just learning how you're being billed. And like Amazon and everyone else will have, you know, a basic-ish cost console that can go in and see, oh, I'm spending this percent of my budget on computers, as much on databases, as much on storage, as much on network, and just kind of go through and start thinking, is this reasonable for what I'm doing? Should I really be spending a quarter of my budget just moving data around? Is that right? And then from there, you could start saying like, okay, let's break this down and figure out basically where is it going? Are we dropping money on oversized databases?
5:36Do we just store data forever? Are you getting hit by Amazon bugs? Like there are bugs in various implementations that will not delete data until you restart the database. And so you're paying for terabytes of redo logs that you just don't even know. Okay, wow, that's fascinating. This is a common theme that I've heard is that it sounds like a lot of the challenge of fully understanding it is that you have all these different, you have to account for all these different services. Your database is built differently. It's used differently from your APIs versus a deployment service that you use to deploy your app to production.
6:14So how are you able to account for all of those different services? Generally, you want to start trying to break down your usage. So one thing we had a lot of success with was tagging all the resources that are used. So we could say this particular line item is used by like this service, which is aggregated like this team, which is aggregated to this product. And so you get these really fine grained things of like, OK, all these services are roughly about the same, except for this one. It's 10 times bigger and we don't know why. Let's figure out why, what, what that money is going to. Let's talk about tech stack a little bit, because, you know, I think that's sort of the practical key to all of this.
6:53So, you know, when you were looking into this, how much of how much of the solution were you able to find in like out of the box, like products versus like having to build something yourself to analyze this? There's a lot of stuff that's already out there that'll help you out. You can start with just like the basic built in cost explorer. That's not bad. There are existing tools you can buy that'll like introspect and try to give you advice. I'm not super fond of them. They tend to be really expensive, like 1 % of your bill. And the advice you're going to get out of them is going to be pretty generic.
7:26Like it'll say, oh, this machine is like underscaled or overscaled. But there's a reason for that. So you get all this like low actionability advice. They also tend to have these blind spots. So Amazon charges a lot for network traffic, but it doesn't provide great analysis tools. So these services say, well, yeah, you're spending a lot of money on networking. I can't help you, but you're still paying 1 % on that too. So they're not a bad place to start. Like if you have a lot of legacy infrastructure, if your like platform engineering team is like small or overloaded, they're not bad to start with.
8:01But, you know, if you already have a pretty sophisticated engineering team, if you don't have a ton of legacy stuff, I'm kind of dubious about the value on that front. After that, like you can build some stuff that's not terribly hard. The Amazon will give you the raw cost data that you can then slice and dice and analyze. So we built something in a mode dashboard. So we load in our data. We can then start assigning sub costs from like, hey, you're responsible for 15 % of the networking. You get 20 % of the monitoring costs. Here's 20 % of the logging costs. So we can then sort of like charge back and bake in costs that way.
8:37Gotcha. You mentioned something like viewing services that are underscaled versus overscaled. In terms of the magnitude of the challenge, Stuff like that, would you say, is typically less of an issue than just how you build your network services or database services? Which one would you say most organizations would have the greatest value of tackling first? The first thing you're probably going to get a lot of benefit out of is some low-level configuration options, because they're probably set up wrong. Those are relatively easy to fix because they're global, and very few people, if anyone, are going to notice them.
9:15after that you're probably gonna have to go around like the service owners and say you can't keep adding gigabytes of data every time you run out of memory you need to write your software better um oftentimes software is written you know we got the python node ruby like they're great to write in but they're not good at scale like you can make them scale but it does a lot of work and our worst performing services are python and node whereas go scales really, really well. So having to go back to a service owner and say, you kind of need to rewrite this because you're processing too much data and we can't just keep afford to running this huge fleet of nodes for you.
9:56This is actually a great transition into what I want to talk about next. And that's organizational buy-in. And particularly when it comes to the individual contributors, like the developers who actually have to build this stuff. Because at the end of the day, any program like this really doesn't matter if you don't convince developers to adopt it and to follow whatever your recommended practices are. So how do you do that? How do you get developers to, A, care about this, and then what do you do to enable them? People respond to incentives. So one easy way to do that is to highlight people who are doing valuable work in this area and broadcast that to the company.
10:35Say, hey, Bob did this really great thing, all that. You can put it as part of your review specs. like, hey, you did this over the last quarter. You know, that's great on your review. That looks really good. Like really emphasize that this is a company of value. Like, I mean, it's great to say it, but actually like implementing it is where you can start moving the needle. And then you want management to actually buy into this. It's like dedicate time, dedicate people, dedicate resources to this. It's not just please do it and then they wander off. So organizationally, it's, you say it's a company value and then you actually tell people, hey, you're doing really good at this company value.
11:12Yeah, and in terms of enablement, do you give past companies I've worked for, they have happy paths that are like, this is the path that as a developer, if you follow, will lead to the greatest joy and most success. Do you have any sort of enablement around that type of stuff? We're building that out. So the most basic one is when you're building a service, we've got a little calculator you fill out and it says your service will cost this much. and we made it as simple and easy as possible. Like while people are writing up the spec, we fill out this document, it's two minutes. It'll say your service will cost this.
11:47That's your estimate and you can go with that. You don't need to like argue back and forth or finagle with anyone. It's just that's going to be close enough because otherwise you could spend weeks trying to get like these really accurate cost estimates. Yeah, that's great. Have you faced any like challenges with your engineering leadership in terms of like getting them to support this initiative? Initially, they were kind of hesitant about this. Like Platt is, you know, it's a growing company. It's relatively small. They want features. They want customers. And that's not a bad decision at the long time, like historically speaking.
12:24But when you start saying like, there's real money to be saved here, like, you know, we could hire more engineers for this kind of cash. And then as well, some of the economic challenges are like, okay, we're facing stronger headwinds. An easy way to improve profitability is reducing costs because you don't have to find new customers. You don't have to improve sales. You don't have to have salespeople or purchase cycles. It's, oh yeah, we turned off this service we're not using. And that's a lot of money we now have. That's great. Have there been any people who have been a big ally as you've worked on this?
13:00You're always going to find a few people here and there who are like, hey, I love making things run really well, really smoothly. They love digging up bugs and pulling them out. So they're great to find. So if you can identify them, they'll help you out. They will happily get on board. They will dig into just the worst parts of their stack and be like, hey, why don't we just compress our logs? So they're wonderful to find if and when you can find them. And you're not going to have a ton in any company, But if you find 5%, 10 % who are really keen, they're great allies because then they can go to their team and say, hey, for the next two weeks, I'm going to optimize how we move data around in S3.
13:41Or I'm going to expire our old data and things like that. In the past, I've seen that you've promoted this idea that optimizing cloud costs can be a really great way to also improve the quality of your code. I'm wondering if maybe you could just walk us through that relationship a little bit. How do those benefit each other? Generally, people want to write good software, but debugging software is hard. And when something is broken in production, the fastest way to do it is you add resources. Initially, when you're just starting out, that's a good way to solve it. It's, hey, I'm running out of memory.
14:14My database is slow. Just add a bit of resources. But as you grow, those incremental additions start costing much, much, much more. When you're starting out, your mom just signed up, so your user count has doubled. adding a cpu core whatever four cents an hour no one cares but when you have like a million users if you add a cpu core to every service that's a million dollars a year 10 million a year that's that's real money you could hire someone to go and fix that instead and it's not going to be a massive jump all at once it's just like little by little it's oh i ran out of memory here i'll add 500 megs oh this is a bit slow i'll add this oh we're having like latency let's increase the pod count and increase and increase.
14:57And eventually you start saying, okay, we're running one pod on every computer and they're massive computers. We have hundreds of them. It's just like little step-by-steps that you get there. Yeah. Someone who has worked at multiple startups and had to inherit the problems that were created by people that came before me. I'm quite familiar with, you know, walking into a situation like that. And then, you know, right before this, we were talking a little bit about how software deployment also sort of like plays a bit of a factor in this so can you can you explain like some of the work you've done around like helping developers deploy software more efficiently i always want to like emphasize that you want to make the right thing the easy thing so if develop if their natural inclination is to do something that should do the right thing so when you land code you don't want to have to have someone manually deploy you You don't want them building images by hand.
15:49You don't want pushing them images. You don't want them like approving every step. You want that just like land and go. So I actually built a system where once you land code on master, it will build it, push it, deploy it, monitor it. If there's a fault, it'll roll it back because otherwise you're relying on human interaction to go out and grab these things. And we've actually caught some like really serious bugs where like someone misconfigured something. They said, oh, I only want to run like two instances of this service when normally it runs 200. And that somehow got accidentally through. System started crashing.
16:25It rolled it back. And whereas historically, you'd have to wait to get paid. Someone show up, figure out, roll it back. This was, you know, five seconds of downtime. And then we were back. Yeah. And, you know, LinearB, we like really care about analytics and data quite a bit. So with this new work with the deployment time, have you actually measured any sort of improvements that you've made? I don't have any metrics offhand, but just the fact that we don't have all this stale code building up when you go to deploy, it's reduced errors. We've caught some really serious things that could have gone out.
17:02We found smaller bugs that this is hard to test. We don't really have an edge case here. But all of a sudden, the error rates go up, and we can roll that back really quickly. And also, you just aren't wasting engineer time watching deploys. Because they almost always are fine. You don't have to train engineers. Here's how to monitor deploys. You land it and it goes and it works. Nice. So the last big subject that I want to talk about is organizational resistance. Any change is always going to have some level of resistance and it can come at any level within the organization. So what would you say, just to start, has been the largest organizational blocker that you've had to overcome while focusing on cost savings?
17:50So what's sort of like the biggest blockers? There's some people who just don't want to participate. And if the number of them is small enough, that's fine. I say you want everyone, you're never going to get everyone, but if you have most people who are on board, they'll sort of contribute. And that's basically enough. You also want to try to avoid sort of over-eagerness. So we've really tried to emphasize, hey, we're not here to say no, we're here to help you out. You know, come to us if you have questions, we'll give you an answer. Like we really want to emphasize yes, but when we have architectural decisions, like we never want to say like no all the time because otherwise you'll get the shadow IT effect.
18:31People will say, well, I've got this existing service and I can just like shove something in there because I don't have to ask them. So I can't get told no and it's fine. and then you come back six months later and be like, what is this? So if you can say yes to all but the worst requests, you'll be doing okay. Because everyone wants to do a good job, but if you really constrain them, they'll do what they can. As someone whose job used to be maintaining the shadow IT system for my team, that resonates very strongly right now. And I never want to go back to that. It was a thing we did out of necessity, not because we didn't have other options, We did it because we had a job to do and we just needed services that accomplished that.
19:14So do you feel like you get more resistance from developers or leadership? So is it the bottoms-up resistance or tops-down resistance? I'd say it's sort of evenly spread. You have engineers who just aren't interested. You also have managers and leaders who are like, hey, I want features, I want velocity, I want this and this and this. And those are just things you have to trade off against. It's like if you're working on costs, you're not delivering features. And maybe that's okay and maybe that's not. Like it depends where you are in your lifecycle. And you just have to decide, you know, this is an important feature.
19:50But, you know, we don't have any customers for it. So if we delay product launches, that could be bad. Can we launch smaller and fix it later? Do you just say like we're going to invest in this and then we'll fix the cost later once we actually have customers? because at a per customer cost for a service that has no customers, it's very expensive. So you just say, we'll run it for six months and then we'll figure out how to make it efficient. Yeah, and one thing that we found just among our user base is that a lot of people do care about tracking whether or not they're spending time keeping the lights on, adding new features.
20:27And our stance on the whole issue is that it depends highly on the individual product line. Some product lines, you want to build more features, so you might have more time spent there, but you might have another line that's more mature, and at that point, it's more about reducing tech debt, making it more efficient. So yeah, I think that's really great points. So how do you convince your engineering leadership that there's ROI to this? I mean, after the fact, it's obvious. It's like, we saved 25%, but if you haven't saved a dime today, How do you convince leadership that this is something that I should focus on today?
21:05They love hard numbers. If you say, oh, there's some value out there, they'll be kind of dubious. But if you say, here's a project, it'll take a week, two weeks, four weeks, and the savings will be$10 ,000,$50 ,000, a million dollars. That starts to be a conversation. They're like, OK, this is just worth doing because it will pay for itself in a day, in a week. You also can do stuff without getting explicit buy-in. I'm sure there's always slack in your sprint. You could just say like, hey, here's just something we're doing for a day or two to fix this or improve that. And one of the great things about cost is the metrics are already there.
21:42So you can just say like, yeah, the cost went down by this much because we did this. You could see it in the chart. And if you start building up a history of that, eventually management will give you more and more leeway to say like, yeah, go ahead and take three months to rebuild our networking architecture because we're spending a fortune on it. So I've just got one more question for you before we leave. Have you learned anything unexpected or been surprised by anything as you've rolled out this program? There's a million little blind alleyways that you're spending on that you don't even realize.
22:12Like there's just so much. Someone set it up. It must be right. And you go back and look at it and be like, no, that's just wrong. Amazon bills you a lot for networking fees. But by default, they send all your S3 traffic out to the Internet and will then bill you for that. But if you flip this one little toggle, it goes internally. And it's such a huge change. It'll save you a fortune if you do a lot of S3. We actually had our account rep call us up and ask if we were down because our traffic drops so much that they're like, okay, something must be wrong. Because obviously it was set up correctly before.
22:47And now all this traffic is missing. And we're like, no, no, we did this on purpose. Yeah, I actually have a phrase for this I call disabling the suck. Yeah. Like so many, you know, there's lots of products out there that just by default, they come out of the box with these features that you're like, if I was designing this, there's no way I would make that decision. And the first thing you have to do is go through the checklist of like, change this, change that. Historically, you had to do it this way and then it stopped making sense. But Amazon can't change the defaults because there's 10 million customers who are relying on this.
23:18And, you know, they don't really have a real incentive to make that an initiative. Right. Because, I mean, they're going to make less money to do that. Actually, one of the surprising things is Amazon doesn't want you to overspend on their services. They want customers who will use their services forever and will keep growing and keep using it. They don't want people to start, shoot up, overspend, go bankrupt. Because, you know, it's bad for customer acquisition. It's bad for them. They'll get a couple months of unpaid bills. No, they want you around. So they kind of want you like within a range of like, yeah, maybe you're overspending and you should like pull this down because we want you here next year.
23:56Yeah. Okay. That's fascinating. So is there anything else that you feel like we've missed that you think is important to share with our audience? Probably try to reemphasize, you know, make the right thing the easy thing. When you want to like pull up data for this cost stuff, like make assignment and responsibility required. Don't make people like do it after the fact. If you want to run a database, you have to set up, yeah, this is owned by Team X, this is paid for by Service Y, and things like that right at the start. Because if you have to chase people around, it'll be the most frustrating thing in the world.
24:27Yeah, that's great. So if someone wants to learn more about this subject or follow you and your work, where's the best place for us to send them? There's not a lot of resources out there yet. It's very kind of new. I suppose the FinOps Foundation is probably sort of like the leading edge of what this is. Wonderful, wonderful. Well, thank you, Mark, for coming in today. It's been a real pleasure learning about what you've been working on. So yeah, it's been great having you. All right. Thank you so much for having me.
From the publisher
“If you're not careful, cloud computing can lose more money faster than any invention in history” - Mark Robinson, Infrastructure Engineer at Plaid
This week, guest host Ben Lloyd Pearson sits down with Plaid’s Mark Robinson to learn how he helped Plaid save 25% in costs by optimizing existing resources and eliminating waste in cloud computing.
Mark explains the importance of understanding your cloud bill, identifying areas of overspend, and implementing changes that lead to significant savings. From the basics of tagging resources to the intricacies of optimizing network and storage costs, Mark offers practical tips that can help you uncover countless optimization opportunities.
Tune in to learn about the rewards of improving cloud cost efficiency, the role of organizational buy-in, and the benefits of making cost optimization a company-wide value.
Episode Highlights:
00:56 How did cloud computing get so expensive?
02:34 Digging into what your costs actually are
04:55 How can you account for the various services you use?
07:35 Where are organizations going to get the most value out of?
12:26 Cloud costs relation to better code quality
16:08 Blockers in organizations to cost savings
19:32 Getting buy-in from leadership on cutting cloud costs
Show Notes:
- Download your copy of the Essential Guide to Software Engineering Intelligence
- https://www.finops.org/
- https://www.linkedin.com/in/mark-robinson-944084b
OFFERS
- Start Free Trial: Get started with LinearB's AI productivity platform for free.
- Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.
LEARN ABOUT LINEARB
- AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
- AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
- AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
- MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.
