Prometheus and Open-Source Observability with Eric Schabell

15 Apr 2025 · 46 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Prometheus and Open-Source Observability with Eric Schabell

Episode Overview

  • Podcast Title: Software Engineering Daily
  • Episode Title: Prometheus and Open-Source Observability with Eric Schabell
  • Guest: Eric Schabell, Director of Community and Developer Relations at Chronosphere
  • Hosts: Kevin Ball, Vice President of Engineering at Mento
  • Focus: Discusses the challenges and solutions around observability in cloud-native environments, specifically around the use of Prometheus.

Key Concepts Discussed Challenges in Cloud-Native Monitoring

  • Complexity: Modern cloud-native systems are dynamic and distributed, making traditional monitoring tools inadequate.
  • Need for Observability Tools: Dedicated observability platforms like Prometheus have arisen to address these challenges.

Prometheus Overview

  • Definition: An open-source observability tool designed for cloud-native environments.
  • Data Collection Model: Utilizes a pull-based model for data collection, which is more suitable for dynamic environments like Kubernetes.
  • Strengths:
  • Strong integration with Kubernetes.
  • Provides time-series data collection.
  • Offers alerting capabilities and a robust query language (PromQL).

Challenges with Prometheus

  • Data Volume Management: Struggles with large data volumes and lacks cost optimization features.
  • High Cardinality: Problems arise from too many unique labels being tracked, leading to storage and performance issues.

Managing Prometheus at Scale

  • Service Discovery: Automatically identifies targets within a Kubernetes environment for data scraping.
  • Self-Hosted vs. Managed Solutions: Discussed trade-offs in managing Prometheus deployments directly versus leveraging managed services like Chronosphere.

In-Depth Discussion Points Metrics vs. Logs vs. Traces

  • Metrics: Continuous collection of quantitative data (e.g., CPU usage), ideal for monitoring infrastructure.
  • Logs: Textual data providing insights into application behavior or issues.
  • Traces: Records of the execution paths of requests through services, useful for identifying bottlenecks.

Using PromQL

  • Query Language: A powerful, but complex language used for querying metrics in Prometheus.
  • Visualization: Typically done via external tools like Grafana; Prometheus also has built-in dashboarding capabilities but is not commonly used as a primary visualization tool.

Observability as a Business Function

  • Role of Observability: Not just a technical need but a business imperative to ensure uptime, performance, and user experience.
  • Burnout Prevention: Managing observability effectively can reduce the operational burden on teams and improve job satisfaction.

Transitioning to Managed Solutions Benefits of Managed Services

  • Focus on Core Business: Allows engineering teams to concentrate on product development instead of managing observability infrastructure.
  • Reduced Complexity: Automatic scaling, storage optimization, and cost management are handled by the service provider.

Migration Strategies

  • Open Standards: Utilizing open protocols (e.g., OpenTelemetry) ensures smoother transitions between different observability solutions.
  • Pilot Programs: Many vendors, including Chronosphere, offer trial runs to test integrations and understand potential ROI before fully committing.

Key Takeaways

  • Observability is Crucial: Essential for modern software applications, particularly in complex cloud-native environments.
  • Manage Complexity: Transitioning to managed observability solutions can free up engineering resources and improve system reliability.
  • Continuous Learning: Organizations should remain flexible and adaptive, evolving their observability strategies as their needs grow.

Conclusion

  • Eric Schabell emphasizes the importance of observability in driving business outcomes and reducing operational overhead. The conversation serves as a valuable guide for developers and teams navigating the complexities of modern software environments.

Additional Resources

  • Chronosphere Workshops: Eric invites listeners to explore free workshops on metrics, observability, and tools like FluentBit and OpenTelemetry.
  • Follow Eric Schabell: For updates and insights, listeners are encouraged to connect with Eric through relevant channels.

---

This structured note format provides a clear overview of the podcast episode's content, highlighting key concepts and discussions while making it easily accessible for readers who wish to dive deeper into the topics covered in the episode.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Modern cloud-native systems are highly dynamic and distributed, which makes it difficult to monitor cloud-native infrastructure using traditional tools designed for static environments. This has motivated the development and widespread adoption of dedicated observability platforms. Prometheus is an open-source observability tool designed for cloud-native environments. Its strong integration with Kubernetes and pull-based data collection model have driven its popularization in DevOps. However, a common challenge with Prometheus is that it struggles with large data volumes and has limited cost optimization capabilities.

0:35This raises the question of how best to handle Prometheus deployments at large scale. Eric Chabelle works in DevRel at Chronosphere, where he's the Director of Community and Developer. He also is a CNCF Ambassador. Eric joins the show with Kevin Ball to talk about metrics collection, time series data, managing Prometheus at scale, trade-offs between self-hosted versus managed observability, and more. Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space.

1:16Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.

1:35Eric, welcome to the show. Hey, thank you very much. Nice to be here. I'm excited to get to talk with you. So let's maybe start out with a little bit of introduction about you and what brings you here and maybe start touching on Prometheus and Chronosphere. Sure. So my name is Eric Chabelle. I work at the Chronosphere and Observability Company. I have a position that has a pretty long, weird title. So I generally just introduce myself as the director of evangelism. You recognize in a startup, that stuff gets pretty fluid. Every year we want to reset the targets, reset the focus, you know, so things are constantly growing and evolving in different directions and you end up getting more and less and other things under your umbrella.

2:13So I do a lot of stuff around DevRel, basically, in the observability space. It's a good way to put it. Awesome. Well, let's maybe then talk about what is Chronosphere and what is Prometheus, which we came on to talk about. Yeah. As per me, this is the topic we decided to discuss here. I'll kind of twist it in that direction. So ChromeStreet is an observability platform. It's a SaaS offering that ingests open standards. Cloud Native Computing Foundation is something that me and my team are a big part of. Two of us, including me, are ambassadors of the CNCF. We like to help out, contribute, and talk about, and do everything we can to promote open source in the sense that that's, we believe, to be the best way to get started with basically anything.

2:53I've devoted most of my life in the app dev space to it first before I came over to the observability side. So it's a continuation of what I normally talk about. And in relation to Prometheus and OpenTelemetry and Percy's a new project and FluentBit, which is also something that we're very integrated with and a couple of the founder and a couple others work at the Chronosphere. All these things are the standard way of delivering and communicating over telemetry protocols, collecting, storing, querying, that kind of stuff is all integrated into our offering. So anybody that starts out initially at a smaller scale in the open source world, finds themselves growing and scaling and having more and more difficulties managing teams that are growing, that are spending all their time managing the infrastructure and trying to make it all scale.

3:38I find it quite nice to be able to unplug something like that and find a vendor with open standards and is easily able to simulate and recognize the query protocols or the query languages and the things that we use from Prometheus. Cool. So let's maybe dive into what Prometheus is. I saw it described as a cloud native observability platform. What does that mean? Yeah. So Prometheus as a metrics collection project is the best way to do it. That's one of the signals that you would want to collect in an observability platform. There's some very unique things about what they do. So this was originally written back in the day, I believe, by some people like SoundCloud.

4:18And the idea being that it has to be highly performant. It has to be highly scalable. And they want it to be as unobtrusive as possible. So one of the things that it does is it scrapes, which means it goes out and gets the data from an endpoint and doesn't require you to set up things to push data to it. It doesn't involve collectors. It doesn't have agents. It doesn't involve anything like that. They allow you to fine-tune that as you wish as a developer. So you can use auto-instrumentation, instrumentation they call it like you just flip a switch on a java library and it starts spitting out all kinds of java metrics usually not a great experience because java metrics are just a wash list of stuff that's just crazy so you want to get more specific so then you can use that library to trim that stuff down and deploy applications again with the proper instrumentation and they usually expose some kind of endpoint where they publish these metrics on and you set up prometheus to go out there and scrape these endpoints every so many seconds, 10, 5, 15, whatever you choose to do.

5:18You're also dealing with the, you need a backend for this. By default, if you're just playing around with Prometheus, it just does it in memory, but you really don't want to do that in real life. So you have to put some kind of backend storage in there and that's a time series, what you're storing. So normally in a database query, you're, you know, it's a finite set of whatever that you're collecting, right? Select star on whatever table and get this stuff and get a bunch of data. Here, you're basically trying to represent constantly collecting data on a machine. So for example, CPU usage, constantly every second of what's going on is milliseconds or whatever is a little bit hard to store and would probably overload almost any organization's network in a matter of no time, especially when you're doing it across thousands of machines.

6:02So they have a way of sampling the data and then re-representing that when you do queries. So that allows you to do time series queries and figure out what's going on and create the dashboards and all that kind of stuff. It also provides alerting mechanisms. So it has the ability to set up alerts and rules around that kind of stuff that then triggers and sends off and can integrate with things like PagerDuty or send you Slack message or whatever it is that you're using in your org. The query language is also pretty much a standard in the industry. It's called the Prometheus query language, PromQL.

6:33Doing that, you can generate database queries basically against the stored metrics data and create visualizations. And there's an embedded dashboard inside of Prometheus, but that's not really generally how people do that. They tend to query with an external dashboarding tool that used to be centralized an awful lot in Grafana, but there's up and coming projects. The very first visualization and dashboarding project just reached sandbox status this last fall. And it's going to be at KubeCon EU here with one of the Project Pavilion little stands. It's a Prometheus project. and they're basically focusing first on the metrics queering through prometheus instances they're starting to expand into traces and some of the other stuff it's a very young project yeah but that's pretty much the infrastructure what you're going to end up with the prometheus i'd love to dig in on a couple of different pieces there so sure first talking about the sort of piece by piece you talked about one of the things that's distinct about prometheus is It is going out to find the data and things like that.

7:36Can we talk a little bit about that? I saw, particularly in the cloud-native environment, there was some stuff around service discovery and how do you find things within Kubernetes to go and look things up. So how does that all work and what does it enable? So initially, some of the links I sent you, I assume you'll put them in the show notes, is around the workshop that we have. You initially start by doing that statically, just defining where to go, what endpoint, what IP address, where are the targets I'm trying to scrape. And that's fun and games when you have one or two things, but you don't want to be managing a list of thousands of those things.

8:08And let's be honest, Kubernetes is set up to be an environment that you write a description for and say, go out there and run these services. And when certain things happen, react this way. But that's about all the control you have. So you really don't know what their IP addresses are going to be. you don't know where they're going to locate and how many there's going to be. So they have something called a service discovery, and you can integrate various tools that do that. They support several different ones. Zookeeper is a pretty famous one that people know. And what that does is just monitor the space you're running your clusters in.

8:38And so when new nodes, new pods, new containers and stuff come up, it just automatically add those to the list and start looking for endpoints and scrape those also. And generally speaking, you've ensured that those are providing those endpoints. Otherwise, you're going to have a lot of down targets. And yeah, that's kind of how the service discovery works. That makes it all dynamic. And that's much more realistic for a cloud native environment. Yeah. So another thing that you talked about here, and you sort of alluded to, you said these are collecting time series, as contrasted to, you mentioned another project that's doing logging or tracing or things like that.

9:11Can you kind of maybe dive in a little bit on what is the distinction between like time series collection versus what you might get in like a logging or tracing type of observability. Okay. Yeah. So a little bit of the difference between metrics, logs, and tracing. So starting with logs, that's a very standard, it's like text lines, right? Going right down the chain, right? Just a bunch of data that you're storing there and it's log lines. That's pretty normal. Maybe you turned it into JSON with something like Flimbit that parses it into machine readable stuff, but it's still text, pretty normal storage.

9:43Traces is also something that where you're you're setting points in your service calls that you're catching and tracking as it goes by, which is also a pretty straightforward data kind of set that you're putting together. It's not a constant stream of stuff. And the metrics where they have to have a lot of really smart ways to deal with this because of things like cardinality explosions. So if you look at object-orientated programming and you make an object for a person and it has a name and it has a age, has social security number, has whatever an address. These are all metrics with labels. So your person might be the metric and the labels might be all these things underneath it.

10:21And if you happen to put something in there like the IP address from the machine that you're getting this information from, that's a unique thing. And every time this gets sent out, you're getting a unique thing you have to save in the database. That's a cardinality explosion. And the more these things can create that kind of stuff is where you get the mistakes where it gets really hard to create this kind of stuff. And in a time series, you want to be able to have a way to store this kind of streaming data and this massive amounts of data in a way that you can measure it periodically and just connect the lines between those short little periods is the best way I can give you an analogy around that.

10:58I guess then a question becomes, how do you decide when is it you want to use time series collection as compared to sort of full traces or something along those lines? Is it when you start? Well, I mean, don't get too wrapped up in the way it's stored. I mean, that's just time series databases are needed for metrics. They just are. So all of them require this. It's not an option that you can choose one or the other. If you're looking at the signal type that you want to collect, whether I should use metrics, logs, or traces, that kind of has to do with what you're trying to observe and what you're trying to watch.

11:27So if you want dashboards with visualizations on what my infrastructure looks like, you have things like thresholds you're setting and watching for things that get beyond the threshold. When you're looking at resource consumption, when you're looking at network lag and things like that, these kind of measurements only come through metrics. When you're looking to find out where did my calls go through my service network, that's traces. And logs are just dumping information that every developer put into his app. Every application that's running on that machine is logging something, right? It's putting messages in the log that I started up that I'm having trouble, that I'm dying now, or whatever the deal is.

12:08It's three distinct signal types that you're using. Yeah, that makes sense. And each one is probably stored a little different. Yeah. Well, and part of why I'm asking this is I think metrics are much more common in sort of operational monitoring and things like this. And a lot of software developers are thinking, oh, I want my logs so I can debug or trace. So kind of translating that mindset into what happens here. So let's maybe look then a little bit further. You have your metrics. They're stored in Prometheus and you're starting to want to query them or think about them. what types of questions you've already started doing this are best answered by this data and like what does the query language look like to explore them i think you've kind of alluded to it there at the end so i think a lot of devops modern organizations are no longer in that boat where developers are just like i want to see my logs that was when i was a developer that was like 15 years 20 years ago right we used to spend i used to make jokes about this with a couple guys here because i do quite a bit of stuff and flew a bit now and it's a lot of logs and stuff we're were simulating and dealing with.

13:06I used to spend time at the coffee machine because the machines were slow enough to compile the Java code we were writing that you had time to go get coffee and talk about problems you were having. And one of the things we used to discuss quite a bit is like trying to figure out how to make a good Java exception message for the logs. Because some humans reading that downstairs on the third floor, right? That's your ops guy. Now that doesn't work that way anymore. That's not really, really what you're focusing on. And that complete picture of your cloud-native environments are deploying apps and services.

13:33Your team owns a service, and you want to have a visualization of that service and everything about it. And luckily, we have a lot of organizations dealing with things like platform engineering, SRE teams, and all this different stuff that have a really good idea of how they want to do this and provide a developer view, provide an SRE view that's carrying the beeper that has this stuff and know who to call if something breaks. But when you own a service like that, you're not so much digging around and looking at individual metrics, you're having a predefined look at your running object, element, whatever you service that you own.

14:08And that's your landing page. That's your art monitor for what you own. And that's telling you, depending upon what this organization has gone through, they pretty much have a good idea of what they want to know. If I'm the order service, I want to make sure nothing's blocking orders. That's where the money comes in the retail org. So they'll be watching really close the connections to whatever credit card companies they're using or maybe the incoming stuff from the website and something starts slowing down orders and they see that little thing go all the way down like this and maybe they're going to say, hey, you know, it's time to look at our dashboard and they should be able to really quickly find out where's the problem.

14:42Are the orders not coming in? Are they not being processed out? Is the timing out going to the credit card company? You know, that kind of stuff. Yeah. So I think what you end up is with having their own view of their world, right? And that's a combination of all three of those things. not necessarily just metrics or just logs or just trace not saying that that doesn't exist but it's generally you don't want to set up your dashboards with here's my metrics here's my logs and here's my traces and then when sre's beeper goes off he or she is like let's look at a log let's look at no what you're trying to figure out is a high level view you can drill down into to get to the problem so with prometheus then prometheus is still just providing one of these pieces, which is the metrics tracking.

15:24And I think as you highlighted, that's probably why people are then setting up queries to query against it from an external rather than using built-in visualizations because they want to integrate this with their logs, their tracing, things like that. What does that end up looking like for someone who hasn't, for example, written something in PromQL? Like, what is this as a query language? What does it look like? And how does it work? I'm going to say right up front, it's not easy. This is something that a lot of people struggle with, including me when I'm writing labs and digging around and trying to find the right thing to give an example of.

15:57But it's within the workshop that I alluded to. There's about four chapters that are involved with learning PromQL and kind of walk you through all the various aspects of it. There are so many different ways to represent your data. You can do a simple counter. You can do a simple tachometer kind of thing. You can do a graph. You can do histograms, which is, you know, histograms is going in the direction of looking at a graph over time. It's so complex what they're looking at that they're basically bucketing chunks of it. So the first minute is bucketed. The second bucket represents the first bucket plus the second bucket.

16:35And it gives you a different look at your data. You have heat maps. Have you ever watched baseball when they show the pitcher stuff throwing? And what are the pitches that they hit the most? It gets red in the areas where they get the most. you can represent your data or your view of what's going on that way too. You have topologies for your service calls where you're basically showing the network of the services and how the calls are going and the lines get thicker the more calls that go over it. The PromQL gives you the ability to, and some of the tooling embedded right now into the newest Prometheus is really nice where they explain the query.

17:08So you're putting the query together, it has command line completion in there, luckily, helping us find the metrics that are available at the point you're trying to do that. And you can see the result sets. You can dissect them because you're also doing sort of like queries within queries to build up a bigger query. You'll apply things like rate across a query. So to kind of generalize a specific query out over time, maybe you'll only want to look at a small window, only five minutes or only one minute or over 10 minutes. There's so much flexibility in that. And yeah, depending upon what you want to land in.

17:42So generally speaking, say you're a service owner, when you come in there, you have a couple of things that are important to you, probably that it's up. You know, so you'd like to have a green bar when it's up and a red one when it's not, right? Or yellow when it's degrading or things like that. There'll be a threshold when it starts degrading. You're not allowed to get to a certain point and then it's a problem. Things like that are the very high level look. And then when things start going wrong, you have to be able to dig down into it. So if it's a specific service, maybe I want to go look at the traces, see one of those topology kind of graphs and see that, hey, this one is not getting anything.

18:12Click on that and I can go down and look at the trace, and then maybe I can look at the log for that specific thing. And we like to talk an awful lot about observability being the pillars in the old days, where they talked about logs, metrics, traces being three pillars. For quite a few years now, we've been talking about a chronosphere that we think it's aphasis. I think we should speak a business language. It's much more complex than an individual tool. And I like to tell the story that it's just like you're driving around in an old car you just restored. You're really, really proud of this old car, and you're driving along, and all of a sudden the temperature of the engine starts going up and it starts getting kind of weird how it's driving.

18:47You're like, oh, I'm close to the mechanic I use. Let's whip in here quick and see what he says. You pull up and it's getting worse. It sounds bad and you get out and he comes running out with his, he goes, oh, come over here and look at this. And he starts opening up a toolbox and showing you all the tools. Meanwhile, you look outside the window and your car is overheating, catching on fire, and he's in here talking about these great tools he's got. So we like to talk about the phases of observability where you want to know as fast as you can what's wrong. And then you want to be able to triage it as quickly as possible, preferably fix it with remediation.

19:18And you want to be able to go back and look at like root cause analysis and find what the long-term solution is for this thing. And to do that, we don't care whether it's two metrics, one label, three traces and half a log of mine. It doesn't really matter as long as I can get to the answer as quick as possible and fix the problem. And I think that's pretty much modern observability in a nutshell, where Prometheus has a very important role in helping collect, manage, and display these kind of stories. Let's maybe, if we can, make it a little more concrete by going through an example. So let's say, I don't know if we want to use a car example or something like that.

19:57You have a service, you're monitoring. Maybe let's take it from ground up. First, what do you need to configure to get Prometheus starting to track things in there? Is there a thought process around what needs to be tracked or everything or it's there by default? Then how do you design those queries that are going to give you that dashboard that tells you, oh, shoot, my car's on fire. What do I need to do here? And then take it through that route. Right. So for Prometheus itself, there's exporters, for example. So let's say that your service is a node service, Node.js. It's very easy to turn on a node exporter.

20:30It creates an endpoint and generates a bunch of node type of metrics that are pretty standard. You're able to get quite far with just doing that. And what's nice about that is you don't have to redeploy anything. There's no code changes involved. You target the service, you set the exporter on, and you start watching what's going on. Once you start collecting data, you now have all the generic metrics around things like CPU usage, what the memory usage is. I don't know what specific things are in this exporter, but if it was a Java one, it's keep sizes and all that kind of stuff. Everything that you normally would be able to monitor around something like that without having to design it yourself.

21:05So the people that write these exporters and manage these exporters and maintain them are trying to make life as easy as possible for people who just want to kickstart it and see what they like. That's either way how you usually start just so you can see what's available. And then you can start trimming stuff down if it's too much for you. Because one of the things that gets out of control really quickly is it's very easy to collect a lot of data. And if you're doing this in the actual cloud, in and out with your data is water through the pipe is money, you know. And so if you don't have the ability to see it, which is one of the big things that Chronosphere is good at, is using your control plane to see your data coming in and out and tell you when you're not using your metrics.

21:44It turns out that across our customer base anyway, that on average, 60 % of the metrics are not used, that are collected. And that's not a bad thing at this point in time. It's just a big chunk of your bill you don't really want to be paying, right? And so the first thing you start looking to do is how can I trim this down or at least stop it from being stored. You do want to ingest everything because you might need it in the future. But it's not a problem to not pay for storage if you're not using it. If it doesn't show up in a dashboard, it doesn't show up in a query and an alert or anything like that.

22:14No ad hoc queries from the user. Nobody's touching it. What are you doing? So that's kind of what you run into really quickly when you use a standard exporter or a standard library that would just spit out everything. But for us developers, that's a nice way to start, right? And then you say, okay, I only want to know CPU. I only want to know how much memory it's using. I only want to know whatever. And so then you go back and you start instrumenting it using code, going actually in there and saying, I want this, this, and this, and the rest I don't need. Redeploy the thing. There you go. Now we have it trimmed down to just what I want.

22:46And that's the stuff that you're querying to create your dashboards. So I'm saying, I want to see a chart that shows me how much memory consumption is going on in the last five minutes or the last 10 or in the last day or whatever it is. By default, it might be in the last hour, but you can expand that and drill down in it, cut and slice and dice any way you want inside that stuff. So that's kind of the evolution you go through to get to something to make you happy. And trust me, when you start getting somebody carrying the beeper, you're going to start seeing things in there that also include documentation in your dashboard.

23:15So there might even be playbooks or runbooks that they have. when certain things happen, go down this path and do this and call this person, alert that person, look for a feature flag that got changed. Maybe there was a new deployment. Oh, definitely reverse that stuff. That kind of stuff. Yeah, that makes sense. Okay. So just to make sure I'm understanding lifecycle, first, you build this exporter, which you probably can start with just a package off the shelf. That's sending a bunch of data. Prometheus is going to then start tracking that in time series. You look at that and you can start building your dashboard by sending these queries against it.

23:47And I think one of the things you highlighted there that is kind of interesting and might be worth exploring is like what time series enables is querying over ranges. And I did see there looked like there's a core distinction between you can do sort of instantaneous moment in time queries, and you can also do kind of range queries. Maybe we want to look at what differences there are there. You studied hard. Do what I can. So, all right, you build your dashboard on that. Maybe you set up some alerts on that. Question I actually have is, is there any difference between the queries you're doing for dashboard versus alerts does prometheus have native alerting support are you querying that on it's it's a separate binary that you can install but it definitely has an alert manager interesting part about this is and i ran into this quite a bit when i first built the workshop is you tend to forget like it's in memory but it's also writing it to a little file system in your directory there so every time i would like adjust something and restart and things like that.

Read the full transcript

24:44I think, okay, you go look at it and it would start a new graph of, you know, it has standard graphs. I'd be querying, just say uptime. You know, you can just do up and it'll show you all the instances that you're tracking. Are they on or are they off? And you'll see this little thing going. If you query something else that happens to have like a counter that's running or whatever it is, it'll show you whatever the graph is. And it's really hard to get interesting stuff when you just started. So you got to remember that your student is out there doing this. He just started this up on his machine and his graph starts with an hour with this little blip in the corner.

25:14And you basically got to get him to go over here and turn this thing into like a minute. And then it starts looking. But it's very different than what I have that's been running. Because I know this. I let it run for half a day. And then I start working on whatever I'm going to give you an example of. Otherwise, we have no nice examples. Also, what's really, really kind of freaky is you'll go away, come back and redo stuff. And then there'll be stuff with a big blank spot and then stuff way back there and stuff way over here. and if your query doesn't span all that you won't see it all until you do and so it slicing and dicing stuff that is long dead and long gone but it's still in your database because it collected that time series when it was alive so some container that was running or some instance of your service may not be there anymore but you're getting like it feels like false positives you know you're like hey well wait a minute this one isn't running well no that you're looking at legacy data there yeah that is kind of an interesting question especially as we talked about if you're changing your metrics that you're collecting or trimming them down like how do you deal with versioning in this type of metrics database system versioning of what exactly i mean the example you use right okay i had it running and then there's this dead spot and something changed maybe that dead spot is because i took things down and i'm actually changing my collector a little bit and now i have a new set of data from this point forward yeah this is a little bit more of a development environment we're talking about.

26:35So you should never see this in production. You know, a blank spot means it all went south. That means you were offline, your retail store is selling nothing. You know, people can't get to your website, that kind of stuff. It does happen, but that's not really a good sign. You do everything you can not to do that, right? So when you do updates or new versions of whatever, it's a rolling thunder kind of thing, right? So that's why we're in the cloud native environment. So you can bring up a new instance and then take down the other one once this traffic's taken it over. A really good example is how people try to take care of, we haven't really got that far yet, but when this starts scaling, Prometheus, one of its weaknesses, I think, is it was not built for high availability.

27:13That's not really part of the design. So what that means is if I have one instance collecting all my stuff and I have dashboards and I have alerting attached to that, it gets too much traffic, it'll overload and die. Everything dies. No dashboard, you can't query it anymore. You can't, everything kind of goes bad. And if you try to load balance that with another instance, you can't put one alerting and one dashboard behind it because an alert will go off on one of these and it won't see the other one. So then you have to do two alerting. You see the scaling starting to go out. And we're not even talking about putting database funds.

27:48You see this spread out into this wirwar of incredibly complex topologies to try and load balance all this stuff. And one of the things that you saw was that you could set it up as a, I have an instance running, I have another instance on the service, and then I have a third instance on the service. And if it starts really flooding this service, it'll spin off its own instance just to cover that, to deal with that heavy load and keep the other ones running. And that's really nice, but the new one you spun up has no history beyond the point that it came alive. And the other ones lose all history from the point that the other one took it over.

28:22Unless you know that in your dashboards and can account for that and query that together into one dashboard. You know what I'm saying? So there's a real complex problem you start juggling when you're doing this by hand. Let's maybe actually dive down that a little bit because, you know, one of the desirable perks of going cloud native is if you have that big viral hit, you go and it's relatively straightforward to scale, right? You're not dealing with, oh, I have to figure out how do I set up a new server? You're like, okay, take this service and this set of pods and scale it up. Go. what happens to prometheus as you do that it sounds like there's some amount of automatic failover or trying to scale in there but like how does that end up playing out you can define when something gets like like i said you can spin up a new instance with it but the high availability is not built in there's not a feature there's not a function you can turn on that accounts for that so you're getting another non-high available instance of prometheus that starts from that moment onwards collecting data for whatever it's monitoring and you're applying your own tricks to the trade to spin this stuff up and to automate that kind of thing this is where vendors start becoming interesting right because you can already smell and feel if you're any kind of a software developer or a person that's had to manage stuff like this that now i'm starting to do a lot of proprietary work myself and it's going to take more hands and more you know people to maintain this.

29:44And I always do this with like, imagine your DevOps people all over here, your whole team is on the right side of the room, doing everything they're supposed to be doing around developing whatever your business is doing, say a retail site or whatever, managing all the services and having a good time. And then slowly but surely about half the team is on the left side of the room managing your very successful environment, which is now scaled way up and has a whole bunch of metrics monitoring going on and all that kind of stuff. Wouldn't it be nice to get half those resources back on the right side of the room doing what you want to do and not messing around managing the infrastructure anymore.

30:17And that's where vendors start coming into play where you're happy you went down the open source road and you can unplug and reroute your stuff directly into something that looks an awful lot like what you've been doing. You recognize the query language. You recognize the protocols being used. Your dashboards are not, you know, those efforts are not lost. They're easy to replicate wherever you land in something like that. And you take the management out of it. And most of these vendor platforms and like some part of the stuff I just described from the control plane from chronosphere is a big big help you know that's basically quantified all that for you and let you just concentrate on what you really want to concentrate on so let's maybe talk about that then what do you get if you're ready to move from prometheus to your chronosphere or another vendor pact actually maybe even just like how do you know right is it when you start seeing that complexity is it when you first get a spike that overwhelms or Prometheus, like what are the signs that maybe you're outgrowing the manage it yourself infrastructure?

31:13There's several things. They're pretty classic in almost any open source environment, right? So there's a really funny marketing kind of story that goes around where you say, killing your heroes. Who hasn't run into someplace where they've worked, where there's one, maybe two people that are like the big rockstar guys that know it all. Might even be girls, doesn't matter. But I mean, they know everything, right? They were there since whenever they know where all the bodies are buried and what happens when they leave? You know what I'm saying? That's usually the ones that are pretty core to running a complex, highly scaling up open source environment.

31:49Some places are really happy to do it. I know our CEO, both our founders, our CTO and our CEO, were both at Uber in the beginning and spent the first, I think it was three years running. They built the M3 database out from scratch and that's a time series database and set up the whole infrastructure there for uber and ran it for three years he did an article not too long ago about how if i think they have 400 or 450 engineers or something like that they're just doing the infrastructure for the observability i mean who does that apparently they do and yeah i know why they do it because he said how much they would have to pay if they came over and did it at ours it would have been like 65 million dollars so you're like huh we all saw some of the leaked the information that some of the customers that were on somebody's yearly earnings calls.

32:34And you're like, with those kind of numbers, you can pretty much run a pretty nice department. And I've worked in universities where money they didn't have, but time they had. So they didn't care how long it took you or how many hands were involved or whatever was going on to manage the open source stuff. It just couldn't cost anything. And I think what you get to the point in time where you're a CIO or somebody that's responsible for these organizations, head of the observability, central observability team, or you're the head of the SREs, and you just want your guys focused on what they've got to focus on.

33:07You're seeing the burnout. You're seeing the stress. You're seeing too many incidents, things like that. And it's often not hard to figure out where you can start cutting costs and where you can take some of the load off, right? And I think people are getting a lot better. You see the conferences we go to and you hear the talks, The examples from the organizations are setting up pretty good observability teams and environments, and they're doing it at big scales. I mean, good Lord, look at DoorDashes and things like that. They're doing it at a mega scale, and they're not running a team of 450 observability guys.

33:39That's not for everybody. I don't think there's any one specific thing, but we've all been in the environments where you're just firefighting and too much of your team is doing stuff that he doesn't want to do. I used to always make kind of jokes about that stuff when the DevOps first started coming out. Because I was like, I signed up for dev. I didn't sign up for ops. You'd hear that around. And if you're not careful, and even now, if you look on the internet and kind of Google around, I think they say that all these developer reports that come out about what they're doing and what the languages are using and all this stuff.

34:09I think it's like 35, 36 % of your time is spent on actually coding. Think about that. That's barely a day in the week. it's like we did our own research and stuff and 10 hours a week we're spent on this kind of observability problems that's crazy out of 40 you know come on is that what you signed up for and that's why people leave you know if i want to be a developer i want to be a developer i don't want to be a troubleshooter the whole time don't want to carry a beeper that's going off all the time when i'm trying to have christmas dinner things like that so let's be honest it's a complicated thing.

34:43It's not easy to run all these very complex things that we're doing. Yeah. Well, I feel like what you're describing here is there's sort of a curve where you start out and you have more time than money. Maybe you're in a university or you're like in a cash strapped startup environment or something like that. And you're like, okay, great. Open source, get it going. Right. Or just a small environment. It's not a big deal. You get to a point where you are hitting the limits of what open source gets you easily. Right. For Prometheus, that might be, it sounds like, when you go from one instance to having two, and now you've got to navigate all of these, like, am I sharding versus am I just replicating versus like all the different federation questions.

35:19And maybe we can talk a little bit about what, go into a little bit of detail of some of those, though you've covered some already. And you say, okay, hopefully by the time you hit that, you actually have a little more money in the bank and you can pay for a service like Chronosphere to take care of it for you. And I do want to put a pin on this and come back to what does migration look like if I do that. And then at some point, though, once again, you get to the point where you're so big that the costs of paying for the service are high, but you also have so much money, you can pay for a whole department and manage it.

35:48And then maybe you migrate back out. Let's kind of maybe look at that migration path. Well, I think you touched on a really good one. I don't want to let it slip away. I think the big part of open source that we're always chasing is the open standards, the ability to be an architect, whether you're designing apps or designing infrastructure. You want to have the ability to stand the test of time as much as possible. And you also want to have components that can be replaced by another component, but are still speaking the same standardized language. When somebody in all these organizations take enough time and effort to create a standard something, whether it's TCP IP or whether it's an observability protocol like the open telemetry protocol.

36:29If you've chosen that road, that means and I think containers is a great one. I mean, Docker in the early days owned that, right? They had their thing. They were the big cat on the block. And when they started getting approached about, we need to standardize this kind of stuff, they didn't really want to hear it. And so what happens is the open source world gets together and it's a bunch of companies too, you know, but I mean, it's all these people contributing. They sit down there, write the OCI, the Open Container Initiative. And now you can write any engine you want against the containers. I have Podman I have Docker and I have whatever you want that's coming down in the future it's all standardized right Kubernetes standardized YAML standardized you want to have these kind of tools and I think that is what you're doing and what you're positioning yourself for by watching the CNCF those kind of projects and the observability space the Prometheus the open telemetry the Jaegers the whatever you're using that's generally speaking doing their best to try and not, you know, tie your stuff into a knot.

37:28And what that means is, is when you're ready to actually migrate to something that should be relatively painless, right? And you're going to find out really fast if the vendor is compatible or not. You know, one of the things we do as a pilot, so we show it to you. You get a trial run, you know, you get to take your environment, put it in ours and plug it in and see what happens. Just seeing some really neat stuff happen when they do that. People are finding things they didn't find before. They're figuring out that they have so much metrics coming in they didn't even touch, you know, things like that.

37:59It's because we're all so busy trying to do the day and day to day. You don't have the time to step back. And who hasn't been in those positions where you'd love to step back and like do some real strategic stuff, you know, and you just don't get the time. We've talked a couple of times about, okay, you need all these different pieces to have your observability solution. You need metrics. So you're keeping track of what high level is going on with the machine. You need logs and telemetry so you can dive, you know, tracing so you can dive deeper. And in the open source world, you might have spun up some of each of these on your own.

38:30Maybe you're, you know, spinning your logs out to Kibana and you've got your tracing going and you've got Jaeger. So you can do that. And you've got Prometheus covering your metrics. If you're moving to something like Chronosphere, can you pull all of that under one umbrella? Pretty much. Yeah. That's what you're trying to do. And it's not necessary either. It's kind of a funny thing. So say that your organization has been experimenting with this stuff as you go, right? And you might be big on the open telemetry ecosystem. It's just the thing you bought into. It's what you've seen. You got the collectors out there already, fine.

39:01But you also have some legacy stuff over here in the corner that is kind of a problem, right? And expensive and coming due. And maybe you want to get off of it. There's tools like Fluentbit, very lightweight collector, telemetry pipeline, basically. Anything in, anything out is their catchphrase. And so they have inputs and exporters on both sides. It doesn't matter where you get it. And it's very lightweight. It's written for the cloud native stuff. It's a subproject of the Fluent D, which was the monolithic stuff, APMs. The Fluent bit thing can go out there and you can even go to edge cases and really lightweight.

39:34It's able to handle high volumes and high stream. And it's really quick at processing all this stuff and lets you on the edge already take down the amount of volume of stuff you're getting that's coming in. They can expose it as a metrics endpoints. So Prometheus can go scrape it. They can pass it on to a logging backend. They can pass it on to open telemetry. They can use the forwarding protocol they have from Flumebit, or they can turn it into an open telemetry envelope and pass that off to them, that collector understands. So the infrastructure you already have in place can be maximized, but what's out there that isn't able to yet, you can put something like a Flumebit in front of it and easily kind of obstinate that it's not the right protocol yet.

40:15And then work on that on your own time and hopefully get rid of it before you got to renew. And you just unplug all that stuff and you're onwards with your OpenTelemetry. The OpenTelemetry collectors can just be redirected to another destination, right? Which could be Chronosphere and whatever cloud that you need, or it could continue to be some backend you have or whatever the deal is. It's quite flexible. I think that kind of covers the whole roundabout. Yeah, no, that makes a ton of sense. So the one thing that I did want to come back to briefly is we touched a little bit on what happens as you start to scale up these decisions around, okay, in Prometheus, maybe you're deciding, am I doing fall over?

40:54Am I sharding things? What do I need to deal with to scale? Is there a federation type of thing? Maybe we can talk a little bit about what that looks like in the open source world. And then bringing that into Chronosphere, does that whole problem just dissolve? or are there still things to think about even if you're using a managed solution? So when you have to start taking decisions around how you're sharding backend databases or storage, I guess you should call it because it's time series is not exactly the same thing. That's you managing your infrastructure. And I don't know what you think when I start hearing sharding and stuff like that and high availability and load balancers.

41:29It sounds like we're getting complicated. I mean, I'm not running that infrastructure, but it sounds like it's starting to get hard. The whole idea of a managed service is that you don't really care. You know what I'm saying? It's not that you don't care. One of the things that we spend a lot of time on is working with the customer and trying to give them the most optimized whatever we can give them. There was a lot of effort a couple of years ago when everybody was kind of cut back their spending in the marketplace to optimize both storage and transmission and collection and all this different kind of stuff.

42:02I think those are the kind of things you're happy to look at, but you don't want to spend your time on that, right? That's the reason you try to get off of the infrastructure you're managing yourself. And the bill is important, monitoring the bill and providing insights to what does my consumption look like, what teams are using it, being able to balance that kind of stuff. You're starting almost to get into the FinOps kind of view of what's going on in your organization, right? Financial operations. And those things are all baked into a more mature observability platform that can include pipelines, telemetry data, tracing, logs, whatever.

42:36It's events, all kinds of stuff can get integrated into that. The look and feel shouldn't be so dramatically different, which is what's nice about coming into something like Chronosphere. I think there's an awful lot you're going to recognize from the open source world. It's the same query language. It's the same idea of what you're looking at. Dashboards are relatively the same. it's an experience you're looking for how do i get my you know my stuff set up as quick as possible how am i able to dissect stuff there's definitely things under the hood that you will not find in you know the younger open source environments it's stuff that you might want to write yourself but i mean we're talking about serious organizations before you get to that kind of level we have various features like that that take you right down into the to the problem in a couple of clicks got some great quotes from customers when they tried it you know i'm not really trying to sell you anything here, but it's just that that's what the maturity looks like.

43:28And that's the difference between doing it on your own and seeing it scale up and starting to have problems. And to be really honest, everybody's environment's way different, right? Everybody has their own specific problems, their own specific legacy stuff, and their own specific issues. We have some customers that have discovered that they use less than 10 % of their ingested telemetry data. that's extreme but that's a use case where they're very much focused on something specific and that's the only thing they monitor and fair enough if that's the case that's the case but they do it at massive scale and so that leads to a lot of you know automatic ingestion of garbage basically for them but even if you can just have it right get half the chuff out of the way that's got to be got to be an roi you're interested in and we show that you get to actually run it for a couple of weeks and see what it looks like in a pilot.

44:18It's not to alleviate fears and stuff. I mean, it's, but it's to show you what it looks like in your environment, what it can do. Try to use real data and real environments. It's not meant to be just a test bed. So that's kind of the experience that you're looking for, right? Can I get off of what I'm doing and stop thinking about sharding and stop thinking about, oh my God, there's another instance and another one and another one and I got beeped again. Yeah, no, absolutely. It's how do I focus on getting the value out of this thing, not just all the work to keep it running. I was going to say, one of the things we often talk about is like, nobody would care what it costs if you had better customer experience, if I had better on-call experience, if I had happy engineers, if I had more money in the bank and less downtime, right?

45:00And very often, that's not the case. It's just a bucket load of money going somewhere and everything is all a mess. I heard somebody say at some point, they said, you know, data is growing exponentially, but I have yet to find that a company whose data budget is also growing exponentially. Yes. Yes. That's a really good example. Yeah. And we wouldn't care if you're using it, but you're just not using it. That's the problem. So every time I have one of these conversations, I come away being like, holy smokes, all the different things I learned. And hopefully, hopefully folks listening along also have that sense of like, I'm coming away smarter than I was an hour ago.

45:35And I would say if you go take a look at all the workshops we have, you can get hands on from zero to installing it to learning about FluentBit, the Percy's project, Prometheus, or OpenTelemetry. We have all of that online for free. Awesome. Well, thank you, Eric. You're very welcome.

From the publisher

Modern cloud-native systems are highly dynamic and distributed, which makes it difficult to monitor cloud infrastructure using traditional tools designed for static environments. This has motivated the development and widespread adoption of dedicated observability platforms. Prometheus is an open-source observability tool designed for cloud-native environments. Its strong integration with Kubernetes and pull-based data collection model

The post Prometheus and Open-Source Observability with Eric Schabell appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Prometheus and Open-Source Observability with Eric SchabellSoftware Engineering Daily · 46 min
Listen in VO