Grafana’s Approach to AI-Native Observability

2 Jul 2026 · 51 min · 24 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI-native observability and how Grafana is adapting observability for agentic software systems that generate code, deploy changes, and operate autonomously.

Guests

Anthony Woods, co-founder of Grafana Labs; background in systems/network engineering, automation tooling, then software development; helped start Grafana Labs ~12 years ago from Grafana’s open-source project. Matt Merrill, software engineering leader with 20+ years in backend, cloud architecture, distributed systems; architects and leads engineers at Dept Agency.

Key claims

Microservices create “black boxes,” so observability turns them into “glass boxes.” AI increases telemetry volume and investigation complexity; too much data overwhelms teams. OpenTelemetry’s open, vendor-neutral standards help models and tools interpret telemetry consistently. Foundation models already “know” Grafana due to public open-source artifacts, enabling Grafana Assistant without training expensive models. Trust-but-verify is required; graphs should explain AI conclusions. Agents shift observability from humans to agent queries and require new data (e.g., conversation/context) and SLO-based monitoring.

Notable examples

Grafana Assistant MVP built during a two-person hackathon; Grafana Cloud AI observability adds conversation data and token-cost visibility; G6 unified agent CLI; automated rollback PRs after deployments when Assistant detects new errors.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to Observability and AI

0:00 to 1:30

Learn about the growing challenges in software operations due to AI and observability.

“Advanced software systems have long been more complex than any single engineer can fully understand.”

Shifts in Software Operations

2:48 to 3:17

Explore the significant changes in software team operations over the years.

“I've done pretty much backend stuff my whole life.”

Complexity of Microservices and Observability

3:17 to 5:36

Discuss the challenges of managing complexity in microservices with observability tools.

“And so I think where I'll start is when you look at how software teams operated five years ago versus today, what do you think are the most significant shifts you're seeing in software operations?”

The Role of Open Source in Observability

5:36 to 7:50

Understand how open source technology impacts observability and operational standards.

“And then I think the other big thing we're really seeing across the world is open source and the advantages, especially on the operational side.”

Introduction to OpenTelemetry

7:50 to 9:35

Learn what OpenTelemetry is and its significance in instrumentation.

“I can't believe I'm not familiar with that, but that is very interesting.”

AI and Observability in Grafana

9:35 to 11:28

Explore how Grafana is leveraging AI capabilities to enhance observability.

“It was able to understand, you know, what people wanted to do within Grafana and knew how to use Grafana.”

Trust and Verification in AI

11:28 to 13:10

Discuss the importance of trust and verification when using AI in observability tools.

“Like, I don't personally see how that might happen, but I'm curious if that's on your mind.”

Proactive Measures in AI-Driven Observability

13:10 to 14:03

Learn about proactive approaches Grafana is taking to improve observability with AI.

“And so you need to be able to have that information that you can go and look at to say, hey, how did you come to this conclusion?”

Understanding AI in Observability

14:03 to 15:56

Explore how Grafana uses AI tools to improve data quality and observability.

“I mean, everyone kind of now just associates it with LLMs and kind of generative AI, but there's still a lot of other AI practices that we use internally to be able to get better quality data, right?”

The Shift to Agentic Models

15:57 to 18:11

Discuss the transition to agentic software design and its impact on operations.

“I think that's what we're seeing is just that the LLMs are very smart, but you have to give them the right data for them to be able to come to the correct conclusions.”
Show all 24 chapters

Introducing AI Observability

18:12 to 20:02

Learn about Grafana's new AI observability capabilities and their benefits.

“our AI observability capability, which just does this, right?”

Monitoring AI Features in Production

21:23 to 23:40

Examine the challenges and strategies for monitoring AI in production environments.

“And I think you explained a bunch of that.”

The Future of AI in Operations

23:41 to 28:01

Discuss the implications of AI in deployment and operational practices.

“the challenges we have with traces with, you know, this agentic model is just the volume of data.”

Trusting AI in Operational Changes

28:01 to 30:01

Discussion on the cautious approach to trusting AI in production changes.

“But me personally, I don't think we're that far away from it, but I'll caveat that a little bit.”

Building Controls for AI Agents

30:01 to 32:09

The importance of implementing safety checks and controls when using AI agents.

“You know, they can write code, they can go and deploy it, they can observe it and see is it having the desired effect?”

Evolving Roles of SRE and DevOps

32:09 to 34:19

How the roles of Site Reliability Engineers and DevOps are changing in the AI era.

“It also depends on where that change or what change it's making, what it's affecting.”

Training Future Engineers in a Changing Landscape

34:19 to 36:37

Concerns about skill development for new engineers in an AI-dominated environment.

“It's just that because all of the easy things, right, are being kind of taken care of, that's kind of the work that you give your junior people, right, or new grads, where you just have to go and do it.”

Open Source and Community Engagement

36:37 to 38:25

The role of open source in nurturing new tech talent and skills.

“Like we, you know, we'd love to kind of engage with our community.”

Challenges of AI-Generated Code in Open Source

38:25 to 40:45

Addressing the challenges posed by AI-generated code in open source projects.

“So I hadn't planned on talking about this, but I feel like it could be fertile ground.”

The Future of AI in Operational Troubleshooting

40:45 to 42:00

Predictions on how AI will change operational problem-solving in the near future.

“So what's the operational problem that you think AI and these tools are going to crack in the next two or three years that people are still doing manually right now, even right now with the existing tools that we have?”

The Evolution of Self-Healing Systems

42:00 to 43:17

Learn about the trend towards self-healing systems in software and observability.

“And I think that's going to be really important just because of the speed at which software is being shipped, right?”

Improvements in Loki's Log Aggregation

43:17 to 45:11

Discover recent enhancements in Loki's log aggregation system and their impact.

“Going and looking for those kind of trends, those hidden things that a user hasn't noticed.”

Challenges with AI Accountability

45:11 to 47:32

Explore the concerns around AI operation transparency and accountability.

“When you're talking about terabytes and terabytes of data that you want to go and process every second.”

Integrating AI into Operational Strategies

47:32 to 50:08

Understand how to effectively integrate AI into existing operational workflows.

“What happens when you've got an agent in your organization talking to an agent in another person's organization and they're just doing things, right?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Advanced software systems have long been more complex than any single engineer can fully understand. Observability is the established solution to this problem, but with AI agents now generating code, deploying changes, and operating autonomously, the challenge of understanding large software systems is entering a new dimension. Grafana is an open-source observability platform, and one of the most widely used in the world. The company builds tools that help teams collect, visualize, and act on telemetry data across logs, metrics, and traces. They are now extending that capability into the agentic era with AI-powered investigation and monitoring tools.

0:42Anthony Woods is a co-founder of Grafana Labs. In this episode, he joins Matt Merrill to discuss how AI-generated code is straining software operations, why telemetry data volume has become as much a problem as a solution, how Grafana is adapting to a world where agents are the primary consumers of observability data, And what keeps him up at night about where the industry is headed? Matt Merrill is a software engineering leader with over 20 years of experience building and scaling software teams across enterprise and product-focused organizations. His background is in backend development, cloud architecture, and distributed systems design.

1:24He currently architects and delivers software products and leads a team of engineers at Dept Agency. You can learn more about his work at code.theothermattm.com.

1:47Hello, welcome to Software Engineering Daily. I am Matt Merrill, and I am here today with Anthony Woods, the founder of Grafana. Before we start, I am going to let him introduce himself. Thanks, Matt. Thank you for having me on the show. Yeah, so I'm Anthony. I'm one of the co-founders. There were three of us at Grafana Labs. I'm based in sunny Perth in Western Australia. It's a beautiful city. It's just far away from everything, which is okay. I spent a lot of time on a plane traveling around. Yeah, I have a background in tech. I started off as a systems and network engineer, then kept building automation tooling until one day I found I was just writing code full-time and become a software developer.

2:23And so that was a lot of fun, being able to bring together all different parts of a system with tools we can build. And about 12 years ago, Raj, Torko, and I, we started a little company called Grafana Labs, right? Based on Grafana, the open source project that Torko had created. And we took that love of open source and of helping people understand the data and telemetry that are getting out of systems and turned it into a business. That is awesome. I am super excited to talk to you because I am a backend nerd. I've done pretty much backend stuff my whole life. I've always been at the intersection of backend and DevOps.

2:56And so super excited. So today I want to talk about software engineering operations in the age of AI. And I kind of want to go into how you're doing this at Grafana itself, because that's a company at scale and in this age, and also what you're seeing happening with customers, because I'm sure you're seeing a lot of it. And so I think where I'll start is when you look at how software teams operated five years ago versus today, what do you think are the most significant shifts you're seeing in software operations? Yeah, I mean, obviously things are shifting and changing very quickly. That's one of the big changes is just the speed of change, right?

3:37Both from a capability and tooling of what's available, but also just on the demands on engineering teams, right? Like the pressure for them to continually deliver more and faster. You know, we've seen the big shift from building monoliths into, you know, microservices and the benefits that had for teams to have, you know, smaller teams with smaller scope so that they could move faster, right? And that has really worked. But the consequence of that is now we've got these very complex distributed microservice architectures where no one actually knows how the whole thing works together. They just might know their own little piece.

4:10And we've certainly seen that both internally for how we build our software. We find that that model works really well for velocity of development, but also adds that complexity. And so the way we combat the complexity internally and the way we see our customers doing it is with better observability tools. It just means collecting more telemetry data to understand what's happening. We just, we look at all of these small little microservices and they're really just little black boxes that are going, pushed into your production and observability is that tool that turns that black box into a glass box, right?

4:38So when it inevitably breaks and they always do, you've got visibility to kind of see inside, see what happened, what went wrong. And that's becoming increasingly important now with the shift to AI, where a lot of the code is not being written by a human, but it's still going to break. And so you need to have that visibility to know what's going on inside. And this is another thing I think that we've seen, you know, 10 years ago, I remember, you know, maybe the conferences, maybe Monitorama or something like that, we would go to, you know, there was this big push around, like measure all the things, right, just because people wanted to capture more telemetry.

5:07And that really has come back to bite people with just like the costs, right? Like we've got so much data now we're ingesting, it's expensive to ingest all that data and process it. But also, you know, we see so many of our customers who run into problems where now they just have too much data. And when things go wrong, they don't even know where to start looking, They're just overwhelmed by this ocean of data that they've got. So trying to help people trim that down and focus on what's the important data that you need to be able to solve the problem, not being overwhelmed, I think is really important.

5:34And so that's a good problem to have. And then I think the other big thing we're really seeing across the world is open source and the advantages, especially on the operational side. We're seeing technology changes fast. And historically, it was open source is a risk for a business because can you trust it? Do you know what's going to happen? who's going to look after it. Whereas now businesses are looking at open sources as a must have, right? Just because that's where innovation is happening. You look at things like Kubernetes, right? You look at what we're doing with Prometheus ecosystem. You look at open telemetry is a great example, right?

6:05Organizations are realizing that just to future proof what they're doing from the operational side, they need to adopt open source technology, open source standards, and these open ecosystems just to be able to get access to the innovation that's happening, but also move away from those kinds of vendor lock-ins that people desperately want to avoid today. So I am admittedly not familiar with open telemetry. So could you say a little bit about what that is? Yeah, definitely. So anytime you're building an application and you want to kind of instrument it, right, there's a lot of choices for how you go and do that.

6:35How do I, you know, when I'm writing my code, I want to emit logs, I want to collect traces or emit metrics. And so there's been a lot of different ways. A lot of vendors have provided proprietary kind of SDKs and things that you can use. And so open telemetry is the open standard, right? It's been around for a while, but now, you know, we're seeing where it's got the velocity, it's got the input and people are actually, you know, using it and getting a lot of value out of it. And so it's a growing ecosystem. And so it's really just a standard practice for how you instrument your applications and you collect that telemetry.

7:02So it has very good kind of semantic conventions, right? So things look the same. So that way, when you've got different teams working on different projects, building things slightly differently, the telemetry that's coming out of it should all look familiar, right? So if I'm not, if I haven't worked in that team before, but I can go and have a look at their logs or look at their metrics or the traces, it should look familiar. Things should be named consistently where things are understandable. And just having that kind of consistent approach really helps. And obviously the huge advantage here is you don't want to have to be re-instrumenting your applications every time you change your observability vendor.

7:34It's a thing that you just want to do once and never have to think about again, but still have that ability to go and change your observability vendor over time or add additional tools or different capabilities. So having that open ecosystem, open standards and vendor neutral approach is very attractive for a lot of organizations. The HTTP of logging. That's the way I hear it. Yeah, that's awesome. I can't believe I'm not familiar with that, but that is very interesting. And I'm assuming Grafana supports this. Are you supporting the project itself too? Yeah, definitely. Yeah. We're, I think, the third contributor by number of commits that we make to the project.

8:09We're obviously very heavily involved in it. A lot of the products that we build, so within our Grafana cloud platform, we have more opinionated solutions of how to do observability, how to do application observability or infrastructure observability. And a lot of that is built on the open telemetry ecosystem. Gotcha. Cool. So what you said about open source and open standards to me also intersects with AI quite a bit, which is where I want to kind of spend a lot of our time. And so one of the benefits, the many benefits that I see of open source is that it's well documented because it's out on the Internet.

8:42These models can, for better or worse, scrape all this information and support it very easily without some sort of vendor lock in or paywall or something like that. Are you seeing that happening already in this space or perhaps not? Is it not caught up yet? No, definitely. I mean, you know, we're really excited. Certainly Grafana Labs with some of the AI capabilities we've been able to build into the product. So things like our Grafana assistant, where, you know, we had a team, you know, we were worried for a while. Like we saw AI coming, you know, if you'd asked me, you know, 18 months ago, how's AI going to impact operations?

9:14I'm not seeing anything useful. But then we did see that big kind of shift with some of the new models that came out beginning of last year, just like the end of 2024. And so, you know, we thought we were behind and we're like, oh, how would it catch up? And then we had a team of just two people during one of our quarterly hackathons who put together the MVP of what is Grafana Assistant. And it was just amazing at how well it worked. It was able to understand, you know, what people wanted to do within Grafana and knew how to use Grafana. And so we were like, how did you guys do this? Why does it work?

9:44And they kind of shrugged and went, I don't know, it just does. But as we thought more about it, the thing that we realized was we've spent more than a decade building this great open source ecosystem, right, of, you know, we've got 25 million plus. users around the world who love and use Grafana, but also they blog about it, right? They write tutorials, they write guides on how to use Grafana to solve certain things. Everyone's got public GitHub repos, which has got their Grafana dashboards in it, and they've got configuration for how they've employed it. They've got a whole bunch of public information.

10:13And it's this information that the foundation models from Anthropic or Google or OpenAI are trained on, right? So out of the box, the models know our technology. They know what our users are trying to do and they know how to solve those problems for it, right? So we've been able to kind of leap ahead of our competition by not having to go and train expensive models, right? We can just use the foundation models and then just build a tighter integration and build more tools and a more kind of integrated solution, right? Into our product rather than having to do all the expensive part of it. So that's a huge advantage, right?

10:44That we've found and great, you know, minds we had 12 years ago when we decided to focus on open source. It was a great plan that we'd put in place, but it really has helped. And we see this a lot where even simple things like, you know, asking, you know, Claude Coe to instrument, you know, with open telemetry, it knows how to do that, right? It can go and do that for you and, you know, add that to your code and make sure that it's collecting the telemetry that you need. That's amazing. It almost seems like good karma, right? Like you've put out this good thing into the universe and now you're getting something back.

11:11Yeah. I mean, that's it. And it's a big kind of value for our organization, right? Where we just have this, you know, mandate that comes from Raj and stuff. He's like, we just do the right thing, right we do the right thing for our employees we do the right thing for our customers we do the right thing for our vendors we just think that when you do that it comes around right and people will do the right thing for you so yeah i'm a big believer in doing the right thing and and you know it's worked out really well for us used to have a saying at a company i work for is like just do good work and the rest will follow and we knew more of that you don't have to answer this if you don't want but are you afraid that somehow this might turn around and bite you in the butt and like somehow replace the tools?

11:47Like, I don't personally see how that might happen, but I'm curious if that's on your mind. I mean, I'm generally a very optimistic person. So I think there's certainly risk there, right? Like, you know, we can't just sit around and hope for the best, right? We are definitely seeing a change with observability. And I think the main thing, you know, especially for us, which is scary for Grafana Labs is, you know, we've built our reputation on being, you know, the company behind Grafana, the dashboarding and the visualization tool, right? And it's something that users interact with, right? And what we're finding more, both our internal use cases as well as our customers, is that more and more, it's not humans that are interacting with the data anymore, right?

12:23It's agents, right, that are going in, doing queries, looking at what's happening in your environment, trying to understand and kind of derive some kind of insights from it. And so, you know, we are seeing in the, you know, that shift, right, where there's less about the dashboard itself. And so we're definitely making changes within our product to support that. We announced last week at GrafanaCon, we've got our new G6 project, which is a kind of unified command line tool, right? And it's designed for agents to be able to interact with our Grafana Cloud service and all of your data and access it that way.

12:55But that said, I still believe that one of the things that's really important when it comes to AI giving you information is the trust but verify kind of piece of it, right? Where it's pretty smart, it's pretty clever of what it can do, but sometimes it gets it wrong. And so you need to be able to have that information that you can go and look at to say, hey, how did you come to this conclusion? How did you make this decision, et cetera? And often the best way to remember that is still a graph, right? It's like, hey. And so we have that in Grafana Assistant when it'll do an investigation for you.

13:25It'll come back and explain to you, hey, I went and looked at this data. I saw this trend or I saw this spike. And so that's why I've then gone down this rabbit hole to go and look for this problem. And so being able to just explain that with a graph is really easy, right? And that's something that the user is still going to want to be able to consume. That makes sense. It makes a lot of sense that, you know, at the end of the day, mostly what you're looking for is the answer to a problem or a question that's causing a problem or a pattern or something like that. Are you doing things proactively?

13:55Is the tool looking at these patterns proactively at this point? more? Yeah, I mean, so we have a few different things that we do. So like the age of AI, I think we were in it, but there's a lot to it. I mean, everyone kind of now just associates it with LLMs and kind of generative AI, but there's still a lot of other AI practices that we use internally to be able to get better quality data, right? Because the LLMs are great when you can give them a lot of context and kind of quality data, but it's the classic, you know, junk in, junk out. So you want to make sure that they are getting kind of clean data coming in so that they can make the right decisions.

14:27And so we use a lot of kind of AI tools internally to understand your telemetry data. So a couple of things we do that for is we have something called a knowledge graph, right, where we kind of look at all the telemetry and build a graph of all the different entities and how they relate to each other, right? So that you can say, hey, I've got this application. That's great, but I know that it's running on a pod, right? And that pod's running on a node, and that node is in a certain cluster in a certain region. And so you can kind of know the relationships. And that's great when we hand that off to the LLMs, because then suddenly when they do see a problem, hey, I'm seeing a latency spike for this application, they then can traverse the telemetry and understand, okay, well, I can find what node this is running on.

15:03And it's like, oh, look, I can see that that's got a resource contention problem. Oh, look, your latency spike is because you've got a noisy neighbor who's consuming all the resources, right? And simple things like that sounds simple to us, but being able to understand all of that telemetry becomes really important and the speed at which things are changing, right? Like, you know, in the old days, you would just have like a CMDB database and be like, oh, I can describe all my assets here and it's a static thing, right? And the reality is every time you have one of those, they're out of date by the time they get updated.

15:30And so we wanted something that was dynamic and just generated by the data that's actually being ingested in real time. So that way it's always up to date and it's always a representation of what your environment actually looks like. So if I'm hearing you right, like this is different layers and different applications of different machine learning, like AI models that feed into reasoning models that can help you make, identify those patterns and things like that. That's awesome. Yeah. And I think that's really important. I think that's what we're seeing is just that the LLMs are very smart, but you have to give them the right data for them to be able to come to the correct conclusions.

16:04And so that's really what we put a lot of effort and a lot of focus on is how do we make sure that we're curating the data set properly. And we can use a lot of machine learning kind of tools to do that, as well as leveraging things like the open standards, and open ecosystems. Leveraging things like open telemetry means that we know the naming conventions and the semantic conventions so we know what the data is representing, right? And that's kind of public information so the LLMs can understand it as well. That's awesome. So just to pivot slightly, I think one thing on my mind is observability has been about humans understanding what's happening in a system.

16:38But as we use LLMs more in the applications themselves, we start to lose a lot of control and understanding of what's happening there. So how is operations changing with that paradigm shift right now that you're seeing with your customers? Yeah, we're definitely, you know, from our perspective, we see the shift to this kind of like agentic model of software design, right? It's just a new design, you know, framework, you know, we had mainframes, right? And then, you know, we had monoliths and then we've moved to microservices. And now we just see this agentic is just a new way of building software, right?

17:10It's got some different problems that are introduces of how we understand what's happening. But the tools we have, for the most part, get you a long way there, right? Like, as we think about how do we understand what these agents are doing, it's typical kind of, you know, application observability or APM, right? Like you want to be able to collect that telemetry to know what's happening. But then now we've got some new types of data and things that are coming out that we want to collect as well, right? So a lot of like the conversations, right, to understand, so we can go and run our evals and do interesting things.

17:37So, hey, are you actually giving correct responses? And then also when things go wrong, you know, trying to understand, well, how did it get there, right? What was the input that kind of caused it to kind of go off the rails? And so we're seeing definitely demand, right, across the market where people want tools that can just do this, right, that can just build into their agent framework. And so that's something, you know, we've had to do internally as well, right, as we were building out, you know, our agents and capabilities, right, we needed to kind of understand this. And, you know, as we like to do at Grafana Labs, we just, you know, we're a company who just keeps scratching our niche and solves our own problems.

18:07And then, you know, we share that with the world as products. And so we the same thing, right? So, you know, we announced last week that we now have in Grafana Cloud, our AI observability capability, which just does this, right? So it builds on top of open telemetry to collect, you know, your raw telemetry and then collects additional things, right? Like, so the conversations and other data that's really important and then builds, you know, an experience where you can go and have a look at the things that are important to you when you're building agents, which right now is like reliability and quality, as well as cost, right?

18:34How many tokens am I burning, right? For all these conversations, right? And which certain kind of workflows within my stack are the ones that are, you know, using the most tokens, right? And that could be a good thing, right? You know, often we care about, you know, token usage because we want to see adoption, right? We want to see people using the tool. You know, we're in the kind of phase now, the growth phase, certainly for AI, but there's going to come a time in the near future where suddenly people are going to be like, wait, let's think about how we can spend less on our tokens. So we want to kind of give that visibility, yeah, we want to give visibility, you know, to our users so they can understand what these systems are doing, right?

19:06And be able to kind of dive in and when things do not go as planned, be able to kind of have the data they need to understand what went wrong so that we can make changes, right? So that we can improve things over time. Yeah. Most AI frameworks started with voice and bolted on video as an afterthought. Vision Agents by Stream was built video first from day one. It's an open source Python framework that lets you build real-time voice and video AI agents in minutes, not months. With 25 plus integrations for models like OpenAI, Gemini, and Claude, sub-500 millisecond latency on Stream's global edge network, and support for YOLO, Roboflow, and custom CV models, you get a production-ready stack without the infrastructure headache.

19:51Whether you're building coaching tools, multimodal assistance, or real-time security pipelines, Vision Agents handles the hard parts. Get started free at visionagents.ai. Here's something nobody tells you when you're becoming a senior developer. Getting better at writing code won't get you to the next level. The skills that make a great developer and the skills that make a great architect are fundamentally different. The Software Conductor is a new book written by SED host Lee Acheson that explores this gap. It explores the gap through the story of Aaron Blake, a senior developer who meets a symphony conductor and realizes the parallels are almost exact.

20:31A violinist makes sound, but a conductor makes music. A developer writes code, but an architect creates the conditions for great software to exist. If you're ready to make that shift, this book is for you. TheSoftwareConductor.com Available at Amazon.com In mobile application security, good enough is a risk. GuardSquare uses advanced, multi-layered code hardening techniques and automated runtime application self-protection and mobile application security testing, combined with real-time threat monitoring to deliver the highest level of mobile app security. Discover how GuardSquare brings all these together to provide mobile app security for your Android and iOS apps without compromise at www.guardsquare.com.

21:21One of the questions I had was like, what you're seeing your customers, like what does it mean to monitor an AI feature in production? And I think you explained a bunch of that. But what I'm wondering too is, what are those hooks? What are those hooks into something like open telemetry? Are you know, like if I'm sitting there thinking about how I might prompt a model to do something to me. Are people saying like, as part of what you're doing, log out your reasoning steps to open telemetry or something like that, where you have another agent that does that? Or is it more deterministic coding before and after calling?

Read the full transcript

21:57Like, what are you seeing in terms of those? Are they integrated with the APIs themselves in some way too? Yeah. So for a lot of the, there is like some good integration. You think about a lot of the tool calls, right? They're not really any different to any other kind of RPC call that you're making, where you're going to want to measure, is it actually getting a response? Am I getting valid responses coming back? You're going to want to look at latency. How long are these things taking to respond? Because that's going to impact the user experience. And you talk about what does monitoring mean or what does observability mean?

22:25Why do we do it? The main reason for it is to make sure that we're meeting our customers' needs or our users' needs, meeting their expectations. So one of the ways we do that with observability and the simplest approach we have is around the concept of SLOs, right? So service level objectives, right? So you, rather than saying, hey, I'm going to go and monitor my CPU usage, right? Because do I care if I'm using 100 % CPU? Not really, right? Because I'm paying for a CPU, shouldn't I be able to use 100 % of it? What I actually care about is what is the quality of the service I'm giving my user, right?

22:57Am I seeing a latency spike, right? That's something that I care about. And so being able to kind of pick what are those kind of core, you know, user experience things that I want to make sure my application is delivering on. So, you know, reliability is a simple one, but even like latency, depending on what your business is, right, you're going to have your own objectives. And so those are the things that you're going to want to measure over time. And then as they change, that's when you know something's wrong, right? If your users are constantly getting 500 errors, where you know something's gone wrong, right?

23:22Or if you're, you know, maybe something simple where, you know, if you see traffic drop off, right, on your site, maybe that's an indication that something's wrong and you want to have someone go and look at it, or you want to have an agent go and look at it and give you an idea of what it might be. But there's, you know, ways to kind of collect all this data today where you can, you know, a lot of logs, a lot of traces, et cetera, where you can bundle this data. One of the challenges with traces, for example, one of the challenges we have with traces with, you know, this agentic model is just the volume of data.

23:46So like those conversations that you want to have usually exceed what a trace can recently contain, right, within a single span that it can send. So we're having to use kind of supplemental kind of like data stores to store that kind of information. And this is where we see then one of the nice things we'd like to be able to do in Grafana, right, is, you know, we've built this ecosystem around integration and interoperability, right? We call that big tent philosophy, right? Being able to kind of bring data together from lots of different types of data sources and be able to kind of draw those correlations.

24:14And so we're able to do the same thing. We can go and build a new database. It could be a SQL style database that's storing all this information. But within Grafana, we can just stitch that together and be able to have you kind of correlate between my traces, right? Where I've got all of my traditional kind of APM data, but then be able to also tie that to all of the chat history or all of the other kind of context information that you want to go and store and be able to then visualize that in one place. Nice. Cool. So back to what you're doing inside of Grafana itself. So you mentioned that awesome hackathon assistant thing.

24:49In terms of just day-to-day operations, where are you seeing AI have the most operational wins for you. Maybe it's even outside of observability of itself. Yeah. I mean, we're definitely very bullish on AI capabilities across the organization, right? Like we have a mandate of like everyone, you know, try and use AI wherever you can, right? So we can take advantage of it where possible. Obviously we're seeing huge adoption and use within, you know, software development, right? Like it is great for building software, for getting code out there and just being able to innovate quickly. We're actually quite fortunate because we see kind of like the role of a software engineer is changing, right?

25:26It's less about the code and it's more about, you know, engineers are becoming more product managers, right? Where it's really about kind of understanding the user problem and then being able to explain that to your agent and we'll go and build the code for you. And that's something that we're, again, fortunate with at Grafana Labs where we've always had this kind of bottoms up kind of culture where our engineering teams have been responsible for the roadmap, right? Like we have a great product team and great product managers, but they're there to facilitate or they bring the information in, from our customers.

25:53They bring information for what's happening in the industry and we can see and then come and facilitate the conversations. But at the end of the day, we like our engineers to be responsible for the product roadmap, right? Because we want them to own it. Part of that is because we are building products for engineers, right? So they do have good kind of intuition and insights needed. Yeah. And so we've always had engineering teams and had a big focus on giving them autonomy to go and make decisions about product and to think about what are the needs of our customers and how can we go and solve that.

26:20And so that's now helping a lot with the tools that we've got where they are already experienced at being, you know, kind of mini product managers themselves. And so they're able to kind of leverage the tools quite effectively to go and build things and get new products shipped and delivered. The other area we see it is where it's interesting, it's obviously more on the operational side is investigations, right? Like understanding what has gone wrong. So again, you know, I talk about the final assistant started as MVP. You know, one of the things that we like to do with our products, you know, we're scratching our own itch.

26:46And before we obviously release products to our customers, you want to make sure there's good product market fit. And the way we do that is we just make it available to our internal teams. And if they start using it without anyone telling them to, then we know it's a good product. And we saw this with the assistant, right? Soon it was in there, people started using it. They get paid for something and they just send the assistant, hey, can you go and tell me what do you think it is, right? And it's very, like, it's remarkably good at doing those investigations and going through and finding the problems.

27:11So that's been really powerful. So we leverage that a lot to go and understand. It's not always right, but often, even if it's not right, it's going to to kind of reduce the surface area that you need to go and look at, right? Because it will come with some suggestions or it might exclude some certain problems, right? So it really does accelerate our team's ability to go and find that root cause and solve those problems. The other area that we're seeing more and more is, which is a little bit scarier, right? Is, you know, as we see AI shift from just the development side into more of the deployment side, right?

27:38Like, you know, when do we start giving our agents keys to go and just, you know, you've written the code, why don't you just go and deploy it? This is where I get uncomfortable. Yeah. This is where I get uncomfortable. Yeah. I mean, it's a scary prospect, but I don't think we're far away from it. I mean, obviously we see a lot of things in the news cycles, right? With, you know, agents going and deleting people's production databases, right? So people are obviously a little bit scared about these kinds of things. But me personally, I don't think we're that far away from it, but I'll caveat that a little bit.

28:05And I think one of the things that we see is that, you know, we don't inherently trust the AI to do the right thing. That's okay. Right. We shouldn't, right? We should be kind of cautious because if there is an opportunity for it to do something wrong, it will find that opportunity eventually. But for me, I don't think that's any different to how we treat people, right? The number of times I've said to someone, oh, we don't need to worry about the normal kind of process. It's just a small change. It's not going to affect anything. I'll just deploy it in production, right? And then it's like, oh, oops, I didn't expect that was going to happen.

28:35And so we always, we need to protect ourselves from ourselves already today. And so like, you know, certainly at the final labs, we have a whole bunch of tooling in place in our CI city pipelines that put physical gates in place where it just prevents you from being able to do the wrong thing, right? You can't go and do a change where the blast radius is global, right? Like it has to be small blast radius, limited impact. So if something does go wrong, it's not going to impact all of our customers all at once. You know, we have things around cost management, right? When we do a change, we have a CI check that runs and says, this is going to increase spend by 10x, right?

29:07And you can't do that, right? And so we have got all these physical gates in place that prevent people from making mistakes. And that also works then for agents. Once you build these gates in place where they're hard blocks that stop the agent from accidentally shooting itself in the foot, you get closer to this reality where you can start to let it do some things in your environment. And I don't think we're far away from that. But I think that having those controls in place is an organization maturity thing. It's taken us a long time to build that because we add new checks in place. every time something goes wrong that we didn't expect were going to happen, right?

29:42And that's easy enough to do with people, right? Because we move much slower. The scary thing for an agent is that they can move very fast, right? So they can make a lot of mistakes very quickly. And so we're going to have to make sure we've got the right kind of controls in place. And so that's going to be the, I think, the thing that will allow us to give the agents a little bit more control to go and deploy things. You know, they can write code, they can go and deploy it, they can observe it and see is it having the desired effect? Is it working correctly? If it's not, make some more changes, go and deploy it, right?

30:09And just shorten that feedback loop so that they can base their changes based on direct feedback from users and what's happening inside the production environment. But it's a scary thing. But I also think it's exciting, right? I think it's an exciting challenge, right? Because we want to get to that place where we can free it up. Because right now, like as software engineers, you know, often our job is to balance the how do I deliver faster and balance that against how do I keep the system reliable and meeting my customer's expectations? right and so we can write a lot of code right now we can develop a lot of new features with the ai but we're still kind of bottlenecked on how quickly can we get that in production and validate that it's actually working and so i think once we we can get past that like where we can give agents a little bit more control then that whole cycle then accelerates a little bit faster and so we can then ship faster and we can get better insights and kind of keep iterating and going and it's a scary prospect and i think there's going to be a lot of pain along the way but i do think it is coming.

31:02We just have to find ways to do it in a controlled, safe as possible approach. Yeah, there's definitely a risk analysis that needs to happen there and how you step slowly into that. And if you have the right checks and balances, I can't wait for some of the stories that'll come out from this. It's going to be unbelievable. I'll give you an example of what we're doing in Grafinal Labs. So one of the things we're automating is when we see, you know, we do a deployment, right? And be small scoped into a specific set of users, right? And then we can go and look at that and make sure it works. And so we have some automated checks that we can leverage Grafana Assistant to go and say, hey, are there any problems or the new errors showing up, et cetera?

31:40And if it does find problems, we just have it automatically open a PR to rollback, right? And so those kinds of actions we're seeing kind of creep in, you know, we still have a user look at that and approve that PR before it happens, but it's simple things like, oh yeah, this is just a simple rollback, can approve that and all just roll back to the previous release, and we can then go and fix it. And so we are seeing these things where we're slowly giving the agents a little bit of operational control because they can do things much faster than we can. But we still need people in the loop right now.

32:08Yeah, self-healing system. It also depends on where that change or what change it's making, what it's affecting. Like, I don't know how I'd feel about an agent deploying like a data migration on millions and millions of rows of data. But thinking of front-end fix, something like that, something that could be easily rolled back, hey, why not? Exactly, yeah. And I think that's where we're going to start. It's those safer operations. We're going to be more comfortable, right? Giving the AI tools. But, you know, we'll definitely see there'll be a bunch of young startups where they're very risk tolerant.

32:38Like, I just care about moving fast, right? And they'll go nuts, right? They'll just let it do whatever it wants and some databases will get deleted. But that's okay. Hopefully not important ones. So let's talk about the people using these tools, right? Right. So as AI handles more of this operational work, how do you see the role of SRE, DevOps engineer, platforms engineer changing? Do you think they're getting bigger, smaller? I mean, I think right now they're bigger, right? Just simply because the volume of new software getting deployed is just growing so quickly, right? That said, I think they're getting left with, which is actually kind of, well, for me personally, right?

33:16Like they're getting left with the harder problems, right? So a lot of the toil, the simpler work that the SREs used to do can be done by the AI now and be done by the agents now, which then leaves the SREs to focus on more high value kind of pieces of information. So the hard problems, right? How do I go and diagnose this or, you know, thinking more strategically around like, how do I build a more reliable way of doing things, right? More kind of process orientated or, you know, go and find, you know, how that AI is not supporting the needs of the business or, you know, helping you kind of drive things forward.

33:44So right now, I think there's just growing need for it, right? We're shipping a lot of code. It's of questionable quality sometimes. So we just, we do need more of that kind of SRE kind of practices to, you know, make sure we are thinking about reliability, make sure we understand, can we measure it? Are we making sure that, you know, as we're building cool features and new tech, is it delivering the value that we want it to have, right? And meeting our customers' expectations. And most importantly, when things go wrong, do we have the data to be able to understand what happened, right? What went wrong so we can go and fix it.

34:12So I think that's definitely the case. As that changes over time, I think, yeah, we're still going to need those practices. The thing that scares me the most, both from an SRE perspective is also software development. It's just that because all of the easy things, right, are being kind of taken care of, that's kind of the work that you give your junior people, right, or new grads, where you just have to go and do it. And you kind of, as you're doing it, you learn, right? You learn how these systems work, you learn what to do. But as we take that opportunity away, it's like, how are we going to give people the opportunity to develop the skills that they need to become these experienced kind of senior SREs that we're still going to need, but they're not going to have that path, right?

34:48And so that's the thing that scares me a little bit is what's going to happen in 10 years time when we realize, oh my God, we just don't have anyone who has the skills that we need. And so I don't know how we fix that. I think it's the same with software development, right? Like we still need real experienced software engineers, but the new grads and kind of juniors aren't getting the opportunity to develop their skills where they can have that experience, right? Where they've learned things the hard way and have gained that experience, right? They've earned it by making a lot of mistakes along the way.

35:15But, you know, we're taking that opportunity away from people. And I think that's the scariest thing for me when we think about the future. The high-level thinking that's needed just keeps getting more and more and more. And it goes from reactive and detail-oriented things to proactive high-level thinking. And yeah, I'm in the software engineering game. We're seeing the same thing. I saw an interesting, I wish I knew who the author was, but I saw an interesting thing of one potential way out of this is to go back to like a apprentice journeyman master model, which I am a huge fan of software as craft.

35:49And I thought that that was really interesting where it's like the only way that you are going to get this experience is by sitting shotgun to an experienced person and feeling the pain with them and seeing their line of thinking. I think that's especially interesting with SREs because you're reacting to these potentially nasty problems and trying to unravel that. And so I thought that was a really interesting one. I think it is an approach. I think the scary challenge is like an expensive approach where you're not getting immediate value from, right? And I think a lot of businesses and organizations are going to be like, I mean, sure.

36:27Why don't we just let some other company be the ones that train all the people and we'll just keep using the agents, right? So I think it's going to be very difficult to find that. I think one area where we do see this opportunity is still open source, right? Like we, you know, we'd love to kind of engage with our community. We do get a lot of young people who come and contribute and that's our opportunity to kind of coach them, right? And being open source, right? We're happy to invest that time, right? To nurture those people so that we can, you know, help them develop their skills over time.

36:52So I will see there will be more of those like community driven, right? Kind of inconsistent. Because I do think this idea of training people up to be productive so that they can leave and go get a job at another company, right? It's not going to be an attractive option for a lot of organizations. That's a really good point. So I do think more of that, like as software engineers, as a community, I think it's our opportunity to find ways to kind of nurture those people coming through and give them opportunities. I think open source is a great way to do that, right? There's big code bases, right?

37:19There's opportunity to write code. There's ways we can kind of communicate and share ideas and help kind of train and coach people a little bit to give them the experience that they're going to need. Well, also, not all apprentices in the trades are paid either, right? And so we've been spoiled in this tech industry, but high paid internships that lead you to a high paying job, the playing field may be leveled, but yeah, good points. Yeah. But I think even like a lot of the people I talk to, you know, colleagues I've got, you know, even still people we hire today, like I love to go to, we do our onboarding, you know, in person and get people out to go and talk to the new starters of what's happening.

37:53And a lot of them have similar stories, backgrounds to me, right? Where they're very self-taught, right? And I think that's true within this industry where a lot of people love the technology They're excited about it. They love the fact that things are changing all the time and there's always new things to do. And so you have to have this mindset of wanting to go and teach yourself and learn new skills. And so I think that's still going to happen, right? But I think open source is a great avenue for that, where it's an opportunity for you to go and do something off your own bat to develop those skills that you're going to need in the future.

38:21Because the speed at which things are changing in this industry, right? You have to be committed to a life of learning because you blink and there's a new JavaScript framework. So I hadn't planned on talking about this, but I feel like it could be fertile ground. So are you seeing open source projects struggle with the amount of code that agents are producing? Or what are you seeing in your area? A little bit, both from the volume as well as the quality. Yeah. So it's very easy to have an agent just go and build a feature, right? But just because it works doesn't mean it's code that you want to maintain over time.

38:56And so that's, I think, where we're seeing the challenges. is I think this is still where there is, you know, that challenge with the AI assisted kind of code tools where we put a lot of thought into kind of like the architectural design and guiding principles for how we build code, right? And every organization is different, but you end up building your view of the world and how you want code to look, right? So it's got consistency. And so it's much easier for your teams to kind of like people to move around between your teams, right? Because the code bases are all a similar kind of layout. You're using similar kind of practices for how you go and build code.

39:26And so things look a little bit more consistent. Whereas when you start getting the AI in, it doesn't care about those kinds of things, right? It's just going to find a solution that works and go and throw that code in. And so there's a lot more work that has to go on to think about and care about maintainability, right? So that certainly affects open source, right? Where you don't just want to accept any pull requests, right? That adds any new feature, right? You want to think about, is that feature right for this project, right? Is that something that we can maintain over time? How does that impact kind of the rest of the ecosystem?

39:55And so that's where we're seeing the challenge is that, you know, you'll get people that will just submit a pull request, you know, that'll be 20 ,000 lines of code that, you know, the AI is generated for them. And you're like, well, no, right? We're not going to review this. It's too big. It's too much. You know, this is not how we build software. You need to kind of break this into smaller chunks. And so that it can be reviewed and so it can be maintained over time and understood. And so I think that's the challenge that we're seeing is code is a collaborative business, right? Especially for open source projects.

40:23And so the agents have to fit into that collaboration mindset of not trying to overwhelm people with too much change all at once because it's just not going to get accepted. It's too hard to review. It's too hard to maintain over time. And to your earlier point, these things might be learning moments for people that are less senior too. Yeah, definitely. So let's move towards the future a little bit. So what's the operational problem that you think AI and these tools are going to crack in the next two or three years that people are still doing manually right now, even right now with the existing tools that we have?

40:56Yeah, I mean, I think right now we're at the stage where AI can be that first responder, right? When something goes wrong in your environment, right? So an alert gets fired because, you know, you've got an SLO burn. It's like, hey, you know, error rates have spiked. Something's wrong. You're going to page your on-call person, but at the same time, you can go and spin off, certainly in Grafana Cloud, you can go and spin off in investigation. So, hey, go and try and work out what's wrong, right? By the time your on-call engineer gets in front of their PC, there's probably already a response to say, hey, this is maybe what the problem is, right?

41:26It may have worked it out or it may have excluded a whole bunch of things. So having that, I think, is where we're going to see over the next few years. And that's just going to get more reliable, more accurate over time. We're going to build more playbooks. We're going to build more skills for it so that can do more interesting things. But I think that's where we're seeing it is that first responder, as soon as you notice something wrong in the environment, it will go in and it will try and troubleshoot that and find that root cause for you. So that way, instead of it taking an hour, right, to find that root cause and kind of solve the problem, we're now talking in 10 minutes, right?

41:55We can go and get things patched up and move on. We're still going to have the follow-up process, right? Which, you know, is so important for our SREs is to go and actually don't just stem the bleeding, let's go and fix the problem. So that's still going to be important. but again you can leverage your AI tools to kind of process more data and kind of do those kind of things but I think the near term so the next couple of years that will just be standard right like you'll just there'll just be an expectation that 60-70 % of the problems are kind of resolved before your on-call person even gets in front of a computer yeah because right now for resiliency in systems we rely on redundancy and things like that but I think what you're saying is what we're starting to trend towards is actually in a limited sense, but a self-healing system, a dynamic self-healing system, which is pretty amazing.

42:41Yeah. And I think that's going to be really important just because of the speed at which software is being shipped, right? If you want to move fast, right, you're going to inevitably break sooner, right? And more frequently, right? Because you're trading speed for accuracy, right? And so the need for better observability just becomes more apparent, right? You're going to want to make sure you've got the data that you need so it can go and fix. But also we're just seeing that AIs are being able to use that observability data to understand what went wrong. And that feedback can be really fast. And that's just going to help us be able to ship things faster, but also maintain the level of reliability that our customers expect from us.

43:16And I think the other thing that we're seeing AI very good at is obviously processing huge volumes of data, right? Going and looking for those kind of trends, those hidden things that a user hasn't noticed. We do a lot of that internally where we can just say, hey, just go and query a huge volume of logs and find what's interesting. So we're seeing that use case change. We recently made changes to Loki, so a log aggregation system to just be able to do that smarter and better, right? To fit with that use case of what AI agents want to go and do. So we've made a whole bunch of architectural changes and performance improvements so that it can query 10 times faster.

43:49More importantly, when you give it a query, it can reduce the volume of data that it needs to go and scan and look at to go and do that. So that's - I'm curious at a high level what those architecture changes look like. That sounds like a very interesting problem. Yeah, I mean, so one of the things - You can't talk about it. I can, yeah, yeah. We announced it last week, Acrofonicon. So it's in open source, right? So yeah, I can definitely talk about it. I mean, I'm allowed to talk about it, whether I understand it as well as the people who have written it, that's a different question. So what we really think about is, you know, Loki was designed for cost efficiency, right?

44:22So how do we store huge volumes of logs as cheaply as possible? And the way we did that is just don't index them. The use case we designed for was the developer use case where you just want to kind of grep, basically through your logs. Grep-CI. Look at the log message. I'm like, it's not that message. You just keep excluding things. Or you might have a, oh, I'm looking for a certain string within the log mission. And so that's what we kind of optimize it for. But more and more, we see people wanting to do more kind of analytic style queries or needle in a haystack queries more often where they're like, hey, I've got a transaction ID and I knew it happened sometime over the last four days, go and find it everywhere in their logs.

44:57And so Loki was designed, you know, to be cheap to ingest data, right? But we wear the cost on the query side, right? Where we just brute force it. We'll just go and look at all the data and look what it's trying to do. And it can do that quite efficiently, but it still takes time, right? When you're talking about terabytes and terabytes of data that you want to go and process every second. So one of the big changes we've made is we kind of introduced kind of similar to light bloom filters, but it's a slightly different kind of approach where we can then know which portions of the data to go and look at and which ones we don't need to worry about.

45:26So if you're doing that needle in the Haystack query, we could very quickly say, hey, actually, you don't need to go and look at the petabyte of data we've got. You actually only need to look at these 10 terabytes that we've got here. And it can do that very, very fast. So that's one of the ways. The other way is we've just put a lot of effort into the query engine itself, just to optimize it and just make it faster, being able to paralyze it as much as possible, because that's easy to do, right? With compute systems today, right? Where we, you know, we've got a huge volume of compute where we can just go and spread out the workload and get things done as fast as possible.

45:56That works insanely well for us as a cloud provider, right? Where we've got the benefits of kind of statistical multiplexing, right? Where we can, you know, we've got thousands of users all using it, right? We can provision a very large infrastructure estate and, you know, not every user's querying it all at once, right? So we can say, hey, this one user's querying it, this second they can go and use a thousand cores. And then the next second, a different user uses a thousand calls. And so we can get really fast response times because we've got a bigger environment that we're able to kind of build on top of.

46:22Efficiency is a scale. Yeah. Exactly. Yeah. Man, I would love to just talk about that for an hour. That sounds interesting. All right. So a couple of wrap up questions. This one I'm really curious about. So what's an operational scenario that keeps you up at night, especially related to AI, something that you think the industry isn't taking seriously enough right now? Well, that's a good one. I think one of the challenges for me is just, you know, a lot of the AI agents that are being built today are just black boxes, right? And we don't know what they're doing. Sometimes we don't want to know.

46:56So I think that's the scary thing is the speed at which these systems can operate and make changes and do things, right? You know, because we've moved from the chat, right, to this agentic model, right, where it's like, let the AI do things, right? And not just respond to questions. So that scares me, right? When you've got these black boxes that are just, no one knows what they're actually doing, right we have a mandate you know internally where our engineering teams are leveraging AI wherever they came to build software right but at the end of the day the person who merges that they're responsible for that right like they're accountable for that and so I get scared about when that starts to erode right and it's like well who's actually accountable for the things that are getting deployed into our production I was having a conversation someone actually just yesterday right talking about kind of like this governance problem right we're seeing more and more agent to agent communication, right?

47:41What happens when you've got an agent in your organization talking to an agent in another person's organization and they're just doing things, right? And something goes wrong. Who's accountable for that, right? Whose fault is that? Who's responsible for kind of remediating that and take accountability? So that's the thing that scares me the most is around, well, I guess two things. One is the black box aspect of it of like, what are these things actually doing? And the other one is the accountability. Who's responsible when things go wrong if we are giving all this autonomy to our AI agents. Some really interesting legal and contractual things that'll come out of that as well.

48:13Yeah, that's a really good point. It always comes back to humans anyway, right? What's the throat to choke if something goes wrong? Yeah, that'll be the most lucrative career in the future, right? Where it's just AI, you know, scapegoat, right? And cleaning up. Yeah, I'll take responsibility. going. Blame me. AI insurance. Ooh, okay. Hey, you want to make a new startup? No, I'm just kidding. So, all right, let's close with, I always like to leave the audience with something. So, if you could give one piece of advice to our audience about how to evolve their operations for an AI first world, what would it be?

48:49Yeah, definitely. My strong opinion about this is don't go and build an AI solution, make AI part of the solution, right? So, we've got so many great products out there. You don't need to go and write a new product, right? That's AI first or something. Really where we see the value in AI is when it's behind the scenes, right? When it's not in your face, there's just, things just seem to magically work, right? And you don't think about it and you don't know why. You just know that it works, right? And you got the response that you want. As we think about AI, I think that's the approach that we want to have is leverage its great capabilities, but it doesn't need to be the only thing, right?

49:25That you're doing, right? I was a new and shiny thing and everyone's like great it's now internet enabled right now we don't care right it's just an expectation that things should just work and it happens to work because it's leveraging the internet and the technology that we have i think ai is we're in that same face right now we're in that hype cycle everyone wants to label everything with oh it's ai it's ai it's ai we're really users are going to care less and less of that over the time what they actually just care about is does it work right is it doing the thing that i want it to do and does it make So like, I think that's the focus is just making sure that your AI kind of works with existing kind of workflows.

50:01It's coming to meet people where they are and help kind of accelerate what they're doing and not trying to force them to have to go and do things they don't want to do. Well said. Well said. Well, thank you very much. Really appreciate your time. No worries. Thank you. Learned some very interesting things. Thanks for being here. Thanks a lot, Matt. It's been a pleasure.

50:25Thank you.

From the publisher

Advanced software systems have long been more complex than any single engineer can fully understand. Observability is the established solution to this problem, but with AI agents now generating code, deploying changes, and operating autonomously, the challenge of understanding large software systems is entering a new dimension. Grafana is an open source observability platform, and

The post Grafana’s Approach to AI-Native Observability appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Grafana’s Approach to AI-Native ObservabilitySoftware Engineering Daily · 51 min
Listen in VO