Engineering AI Systems for Autonomy and Resilience with Krishna Sai

24 Feb 2026 · 53 min · 22 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Engineering AI Systems for Autonomy and Resilience with Krishna Sai

Podcast Overview

  • Title: Software Engineering Daily
  • Description: Technical interviews about software topics.
  • Episode Title: Engineering AI Systems for Autonomy and Resilience with Krishna Sai
  • Episode Description:
  • Discussion on enterprise IT systems' evolution into complex, distributed environments that include cloud infrastructure and AI-driven workloads.
  • Focus on the challenges of observability and system failures, and the potential of agentic AI systems in enhancing operational resilience and efficiency.

---

Key Concepts and Discussions

  1. Evolution of Enterprise IT Systems
  2. Sprawling Environments:
  3. IT systems have expanded to include cloud infrastructure, applications, and AI workloads.
  4. Observability Tools:
  5. While tools help collect metrics, logs, and traces, understanding failures remains challenging.
  6. The need for advanced systems that go beyond dashboards and alerts is highlighted.
  1. Agentic AI Systems
  2. Definition:
  3. AI systems capable of reasoning about operational data and reducing manual toil.
  4. Goal:
  5. To enhance autonomy and efficiency in mission-critical operational environments.
  6. AI Observability:
  7. Differentiated from traditional observability; focuses on active engagement rather than passive monitoring.
  1. SolarWinds’ Role
  2. Company Overview:
  3. Historically recognized for network and infrastructure monitoring, now extends to modern applications, cloud environments, and AI workloads.
  4. Product Portfolio:
  5. Focuses on observability, incident response, and service management across multiple domains.
  1. AI-Assisted Programming
  2. Investment in AI:
  3. SolarWinds actively integrates AI-assisted coding tools for engineers.
  4. Notable improvements in commit and deployment velocities.
  5. Challenges:
  6. Code review processes remain a bottleneck; concerns about code quality and security persist.
  1. Designing AI Systems
  2. AI-First Approach:
  3. Instead of bolting AI onto existing systems, design from the ground up to incorporate AI-driven components.
  4. Separation of Concerns:
  5. Systems should have distinct planes for data, control, and reasoning to enhance manageability and safety.
  1. Human Interactions with AI
  2. Evolving Engineer Roles:
  3. Shift from being logic authors to context engineers; responsibility is becoming more probabilistic rather than deterministic.
  4. Building Resilience:
  5. Engineers must adapt to the feedback generated by AI systems and cultivate resilience against unexpected behaviors.
  1. Operational Efficiency and Impact
  2. ROI Measurement:
  3. Focus on clearly defined metrics (e.g., ticket deflection rates) to gauge effectiveness and improve systems.
  4. Community Engagement:
  5. Creating vibrant communities for knowledge sharing among junior and senior developers is crucial.

Key Takeaways

  • The complexity of modern IT environments necessitates advanced AI systems that can autonomously manage operations.
  • SolarWinds is transitioning its offerings to focus on reducing operational toil through agentic AI solutions.
  • AI-assisted programming is reshaping engineering workflows, highlighting the need for robust evaluation and review processes.
  • Engineers will increasingly need to focus on context engineering and adaptability, rather than solely on the creation of logic.
  • Effective measurement of ROI and community engagement are essential for successful AI implementation in organizations.

---

Conclusion The podcast episode with Krishna Sai emphasizes the importance of rethinking observability and operational resilience in the age of AI. It outlines the transformative potential of agentic AI systems and the necessary evolution in the roles of engineers within organizations. As enterprise IT systems become more complex, the integration of AI and the development of collaborative workflows will be vital for navigating future challenges.

---

For further details, check out the full episode [here](https://softwareengineeringdaily.com/2026/02/24/engineering-ai-systems-for-autonomy-and-resilience-with-krishna-sai/).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introducing SolarWinds and Krishna Sai

0:45 to 1:25

Overview of SolarWinds' evolution and key insights from Krishna Sai.

“Krishna Sai is the Chief Technology Officer at SolarWinds.”

The Role of AI in Observability

1:25 to 3:32

Discussing how agentic AI systems are changing operational data management.

“This episode is hosted by Sean Falconer.”

Reducing Operational Toil with AI

3:32 to 4:19

Exploring the challenges of operational complexity and the role of AI.

“But all of this is much simpler said than done.”

AI-Assisted Programming at SolarWinds

4:19 to 6:20

Insights on how SolarWinds engineers leverage AI in their coding practices.

“We've all been there waking up at 3 a.m., alert storms, and getting into a war room.”

The Impact of AI Agents in Software Development

6:20 to 8:28

Discussing AI agents' role in enterprise software use cases and development.

“All our engineers use AI-assisted coding.”

Evolution of Operational Systems

8:28 to 14:01

Analyzing the transition from traditional monitoring to agentic systems.

“These types of things are very, very, instead of taking an example of taking a single code snippet, breaking the task down, the coding agent goes and reads the repo, inspects dependencies.”

The Evolution towards Agentic AI

14:01 to 14:55

Explore the transition towards AI systems that can act autonomously in complex environments.

“where there's assistive tooling that is sitting next to all the human decision-making that's happening.”

Data Silos and Operational Challenges

14:55 to 18:10

Understand the impact of data silos on operational efficiency and AI implementation.

“I mean, we saw a similar evolution in the world of biology too.”

The Human Brain Analogy for Observability

18:10 to 20:07

Learn how the human brain's processing can inform observability systems in AI.

“As an example, in our design discussions, we use the left brain, right brain analogy, right?”

AI in Operational Contexts: Beyond Co-Pilot Models

20:07 to 23:59

Discover the evolution of AI roles from co-pilot to proactive agents in operations.

“But it's really about giving visibility in the systems for humans.”
Show all 22 chapters

Architectural Design for AI Autonomy

23:59 to 27:21

Delve into the architectural considerations necessary for effective autonomous AI systems.

“So we have this internally, we think of this as AI by design, which is not AI first or AI everywhere, which tends to get kind of misunderstood a lot.”

B2B vs B2C AI Interfaces

27:21 to 28:00

Examine the differences between B2B and B2C AI applications in operational settings.

“So you talked about these like co-pilot experiences where even a really great co-pilot, you're still finding out after the fact that there's some sort of problem and it has to essentially wait for a user to prompt it.”

Operational Use Cases for AI Systems

28:00 to 29:00

Explore how AI systems can autonomously address operational issues without user prompts.

“Maybe there's an agent behind it, but still just a chatbot waiting for a user to ask a question.”

Architecting AI for Specific Problems

29:00 to 33:00

Discuss the architectural choices necessary for developing effective AI models for specific operational issues.

“Do you need trillion parameter models in that world, or can you get away with something that's more like an SLM that's tuned to the particular problem at hand?”

Refining Data for AI Systems

33:00 to 35:40

Understand the importance of refining datasets for AI systems to function efficiently.

“And I think it's a massive sort of operational efficiency win for the company too, where you don't have to spend sort of human cycles routing tickets.”

Designing for Autonomy in AI

35:40 to 40:00

Learn about creating structured environments for AI autonomy while maintaining control and safety.

“that's kind of purpose-built for the AI system.”

The Evolving Role of Engineers

40:00 to 42:00

Examine how the role of engineers is transforming in the age of AI and automation.

“And there's a lot of activity happening in forums like OpenTelemetry, et cetera, now to support this as well.”

The Evolving Role of Engineers in AI Systems

42:00 to 44:00

Learn how the role of software engineers is shifting from writing logic to engineering context.

“even as engineers, as an example, like how does an engineer's role evolve, right?”

Emotional Resilience in Engineering AI

44:00 to 46:00

Discover the importance of emotional resilience when working with AI systems.

“Yeah, I think that is probably the biggest fundamental shift that we're seeing is sort of what the role of software engineering is.”

Adapting to Agentic Systems

46:00 to 48:20

Understand how junior and senior developers adapt to new AI-driven engineering frameworks.

“I think that goes beyond just engineering as well.”

Building AI Communities for Knowledge Sharing

48:20 to 50:20

Learn about the benefits of creating vibrant AI communities within organizations.

“And I guess, what are your thoughts on that?”

Measuring ROI in AI Tools

50:20 to 51:40

Explore the challenges of justifying ROI for AI tools in engineering contexts.

“And that's a really important part of the process.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Enterprise IT systems have grown into sprawling, highly distributed environments spanning cloud infrastructure, applications, data platforms, and increasingly AI-driven workloads. Observability tools have made it easier to collect metrics, logs, and traces, but understanding why systems fail and responding quickly remains a persistent challenge. As complexity continues to rise, the industry is looking beyond dashboards and alerts towards agentic AI systems that can reason about operational data, reduce toil, and take action when things go wrong. SolarWinds offers solutions to monitor, understand, and remediate issues across complex distributed systems.

0:43The company began as a leader in network and infrastructure monitoring, and has evolved to support modern applications, cloud environments, containers, and AI workloads with a growing focus on reducing operational toil. Krishna Sai is the Chief Technology Officer at SolarWinds. He joins the show with Sean Falconer to discuss how SolarWinds is rethinking observability in the age of AI, what it means to design agentic systems for mission-critical environments, how AI-assisted programming is reshaping engineering workflows, and why the future of operations depends on building platforms where humans and autonomous agents work together.

1:25This episode is hosted by Sean Falconer. Check the show notes for more information on Sean's work and where to find him.

1:43Krishna Sai:hi welcome to the show thanks sean it's great to see you and meet you big fan of the show so thanks for having me here oh well thank you so much that's nice to hear yeah i'm looking forward to this as well so i wanted to start off talking a little bit about you know solar winds and kind of set the stage there? Because I think a lot of people know SolarWinds maybe as a single tool that they used years ago, but you guys do a lot of different things. So given where you are today, how would you describe what SolarWinds actually is today to someone who hasn't looked at the space in a while? No, absolutely.

2:17Krishna Sai:If you take a step back and say how IT and ops teams who typically use SolarWinds products have been using SolarWinds for the past 25 years or so, a product portfolio broadly expands three domains, observability, incidence response, and service management. And to put it simply, IT and ops teams use us to help detect and remediate issues across a variety of workloads in their environments. Network and infrastructure, which is where we started and have been a leader for a very long time, but also applications, databases, containers, ML workloads, et cetera. And so our solutions cover this from a horizontal perspective, meaning give you the ability to look at the general basic health of the typical workloads, compute, storage, network, et cetera, but also vertical cross -cutting concerns like performance, reliability, cost, security, and so on.

3:15Krishna Sai:And what happens is typically IT and ops teams are accountable for SLAs and SLOs, and that kind of drives your day-to-day behavior. More mature teams, of course, manage error budgets at scale, and they have nuances of that same dimension. But all of this is much simpler said than done. I was talking to a CIO who was part of a customer call recently. He's here of a major system integrator responsible for running big managed global services for an organization. And he said it well. He said, I'm responsible for SLAs, but honestly, I can't tell you everything that contributes to an SLA. It is a statement of complexity in these environments, but it's also increasingly these teams have to deal with large microservices, distributed systems, etc.

4:08Krishna Sai:And so complexity is very real. And so what we target is, especially in the context of AI and so on, our goal is to reduce toil. We've all been there waking up at 3 a.m., alert storms, and getting into a war room. And the problem with that is that even today, a lot of the tools just ingest a whole lot of data and show you a lot of dashboards with red lights and so on. But they still finding out why something is red is still a big challenge. So we've been thinking about this challenge. So when we think about AI assisting with this, traditionally, you know, we've gone from statistical approaches, things like anomaly detection, machine learning, basic stuff, to now there's a very clear shift to agentic AI, not just in our industry, but just across the board.

5:00Krishna Sai:And so that's something that we want to focus on and increasingly index on. And the way we talk about that is we just call it Solomon's AI more broadly, but in particular, the agentic portion of it we call Solomon's AI agent, it often gets confused with AI observability, which is something that comes up a lot. And the way we think about AI observability is that as a more of a vertical use case, right, rather than a lot of a horizontal thing. But that's also something that we're starting to do here. Yeah. So I think there's quite a bit to sort of unpack there. I definitely want to talk a lot about how the use of essentially AI agents are starting to sort of impact the types of use cases that your customers are generally interested in.

5:46But one place I wanted to start off with, because I think a lot of people listening to the show, and it's certainly a lot of businesses I talk to are all interested in understanding what are some of the leading technology companies doing when it comes to leveraging AI. So I wanted to talk briefly about AI-assisted programming. So to start off there, how heavily has SolarWinds invested in various AI-powered programming? Is that something where every engineer has now got a junior engineer in their pocket where they're leveraging something like cloud code? Where are you with that?

6:20Krishna Sai:No, we are. We're investing very heavily. All our engineers use AI-assisted coding. Actually, it's super interesting because yesterday we were kind of reviewing the progress that we've been making through the year. And all engineers have enabled using AI-assisted, both in terms of co-pilots as well as agents and so on. And what we're broadly seeing is that in general, we're seeing, of course, increasingly, you know, increasing code being generated with AI, for sure. The percentages vary from organization to organization and how you measure it and so on. but we're seeing increased commit velocity.

6:58Krishna Sai:Our commit request velocity has gone up like 25, 30 % north of it sometimes. But part of it also, what we're seeing is through deployment, frequency has gone up as an example. So we're seeing some improvements in lead times. But what we're also seeing is that tools are maturing, which means that both the acceptance rate as an example of generated code is significantly improved over the last year as both the models have improved and the agents have improved. But the shift is, of course, now more on the code review side, right? Like a lot of a code review portion is still, the bottleneck has moved there.

7:36Krishna Sai:So that's something that we're starting to address and look at how do we make it better and so on. But broadly, I would say the signals are positive. Of course, the concerns are the usual concerns around code quality and flaky test generation and security guidelines. et cetera, like those things still are very much front and center. But, you know, one, maybe this is kind of where you are going, is that when we think about agent engagement, coding is actually a very good baseline because we've all been working as an industry for a long time. And the nuances of, okay, how does the application or the use of coding agents now extend to broader enterprise, like software use cases, is something that's front and center.

8:22Krishna Sai:And that's where we've also been playing both sides of it, which is using coding agents and building our muscle around using how agents work and so on, while at the same time offering them through our products to our customers. So there's that context as well. So one of the things that when I think about this notion of how do these agents apply in the enterprise software use cases, if you take a coding example, for example, If you ask a coding agent, go add an OAuth support to my function or whatever, or refactor this module to be async as an example. These types of things are very, very, instead of taking an example of taking a single code snippet, breaking the task down, the coding agent goes and reads the repo, inspects dependencies.

9:15Krishna Sai:There's a lot that's going on underneath, editing multiple lines, running tests, so on and so forth. So this notion of setting the intent and then the system deciding what actions it needs to take to drive towards that goal turns out to be a very good mental model to baseline on in terms of how we think about agents in the context of enterprise software. Yeah, I think that's a very astute point. I think that one of the reasons, I've been thinking about this a lot recently, I actually wrote an article about it. I think one of the reasons why programming has been such a tip of the spear is, like you said, there's all this sort of history associated with it.

9:55But it's also like a hard truth environment, I would say, where even though the output might be non-deterministic, there's kind of deterministic ways of checking the correctness. And then that becomes almost like a reinforcement learning cycle because I can compile the code. I can run it against unit tests. And most environments don't have that. So I think for other sort of non-coding environments to be successful at the level of coding, you need to be able to create those similar sort of deterministic guardrails where you can actually evaluate the outputs. in some reasonable way so you have confidence that's actually generating something correct.

10:27Krishna Sai:No, absolutely. And I think that's why that analogy kind of makes a lot of sense to me when we try to internalize it as like, what is the intent, right? And what is the intent and how do you measure it? Which is why, as an example, if you, a lot of systems like, which is in operational practices, things like SLOs and SLAs come in super handy because at the end of the day, what you're really trying to drive is towards a certain healthy operational state, which is actually well-defined in the non-agentic world, a very human-driven world as well, because a lot of your practices, et cetera, incident response is an example, traditional setup, a threshold fires, page goes out, engineer wakes up, engineer does a series of things, checks dashboards, pulls logs, inspect traces.

11:16Krishna Sai:It's like a very, very disciplined type of an approach that the industry itself has matured along those lines. But the way for an agentic system to kind of then, let's say, mimic it, so to speak, and to make it a lot more effective, efficient, autonomous later on when we talk about, one approach is to say, how do you not remove that logic or the set of practices, but how does an agentic system say, absorb it, so to speak, right? And that's where I think a lot of the implementation design challenges, et cetera, come in. So a system, the way we think about a typical, like an agent, as an example, right?

11:54Krishna Sai:If you have an SLO, a system could be observing raising error rates, as an example. Notice that, okay, this is isolated to a specific service. And then it correlates that to a deployment that happened 10 minutes earlier. It observes a trace pattern that happened during a previous incident, concludes that this is a bad config change or whatever. All of these sets of steps is very, very, I would say, there's a lot of historical knowledge and actions that an agent can learn from. And I think that's where I think the analogy with a lot of what happens or how do coding agents work in a coding use case versus an operational agent working in an operational use case, there's a lot of similarities there.

12:40Yeah, absolutely. And you have a tremendous amount of experience in, I guess, like traditional sort of enterprise infrastructure, having worked at a number of large, successful organizations. How much of you see building sort of agentic systems as something that's brand new versus, I don't know, like a rebranding of some of the typical things that we would do with any software application?

13:07Krishna Sai:Yeah, no, that's a great question, actually. And if you think about this evolution of operational systems, if you think about it, like traditionally they were monitoring systems and then monitoring just things that were polling and observing the state, so to speak, manual thresholds and UAC, like in the first generation, so to speak. Then they evolved into things like observability, where the system expressed its state through multiple ways, and then there was a way to correlate them across these multiple signals. Then there was this concept of AIOps, essentially, which is essentially using these signals and coming up with ways to correlate those signals and making decisions.

13:51Krishna Sai:still very, I would say, very early on, but that's where it started. Then we had started to see cases of like co-pilots emerging, where there's assistive tooling that is sitting next to all the human decision-making that's happening. And now we're starting to see the early green shoots of agentic AI, where there are agents that can actually act, even autonomously at times, within boundaries and so on. So there's an evolution of how this industry has gone through all of that. And some of that is, I would say, the natural push and pull of technological evolution. But also, a lot of what has made that almost an existential need is just the sheer exploding complexity and tools brawl and data.

14:40Krishna Sai:And it just, at some point, we all realize that it's just a human is not going to scale in terms of maintaining the health of these complex environments, right? Like that's the evolution that I've been seeing. Yeah. I mean, we saw a similar evolution in the world of biology too. There's just the sheer amount of data that exists in biology. Like people started sequencing the DNA and stuff like that 30, 40 years ago at this point. So technology has served a massive role there. And a lot of people say that the 21st century is going to be the age of sort of biology because of the fact that now we have powerful enough computers.

15:19We have these really powerful models to kind of assist in the data crunching involved in essentially evolving that science. Because no one human, no matter how gifted you are, could possibly keep all those things in your head. And I think we're seeing a similar evolution in technology because there's just any complex enterprise environment. There's literally thousands of different data systems that might be sending important signals all over the organization. It's like, how do you sort of start to be able to parse that? And most businesses are sitting on these terabytes or heaps of unused data where they hold on to it because, you know, they might be useful some day, but they don't have sort of a way to unlock that use.

15:56Krishna Sai:That's right. That's right. And a lot of that, if you think through, for example, if you extend that and then you ask yourself, why? If you think through that, in an operational context, this was an aha moment that we had probably a couple of years ago. We've always kind of loosely had this strategy, but it came more front and center a couple of years ago. If you think through how the operational systems have evolved in digital environments, there was a monitoring observability industry, which was just only focused on getting signals, showing dashboards, and giving alerts, as an example. And then you had this incident response or more DevOps-y types of environments that came where you really saw all those signals, but then it adapted to how teams were operating with incident response, being able to maintain SLOs of services, manage error budgets, and so on.

16:49Krishna Sai:And then you had the IT systems off on the side where there were very, very ITIL-driven, service delivery-driven, very, very, shall we say, structured processes that enabled enterprises to scale at these practices. But what all of that did was create all of these silos in terms of IT operations management, IT service management, et cetera. And the silos happened in organizations, but they also happened with data. And to your point, when you have all of these massive amounts of data ingestion, which data ingestion is something that we've mastered very well now, but they ended up creating all these massive data silos.

17:29Krishna Sai:So what happens is when you want a Gentex system or any kind of an AI system to work, when you have data silos that becomes incredibly complex and expensive, results in you have this third wheel of separate AI systems that then ingest all the data and then having to process data on top of that. And then you have multiple dashboards and consoles. And so when you're dealing with operational situations and war rooms, There are just so many different consoles that are spread around everyone's desktop. All of these challenges, which really almost became important, critical, that we really take another look at how these things work internally in SolarWinds.

18:15Krishna Sai:As an example, in our design discussions, we use the left brain, right brain analogy, right? Where if you take the human brain, it's actually one of the most, we're talking about biological systems, right? It's the most wonderful biological system for observability ever created. We walk around the face of the earth ingesting all kinds of signals through our five senses. And you can be in a crowded mall with a lot of noise and someone says, Sean, and you're instantly paying attention to that, right? With your name. And the subconscious seems to have this phenomenal mechanism of being able to process all those signals, decide what's important, what's not, and being able to surface that, shall we say, actionable signals to the conscious where you as a Sean can say, hey, is that really meant for me?

19:05Krishna Sai:Oh, maybe that's a different Sean. I can go about walking around, getting cookies in my mall, so on and so forth, right? So you can do all those things. So when we think about, you know, extending that analogy to an observability use case, you have this system of observability on the one side, which is optimized for your mean time to detect, so to speak. And then you have those right brain systems, which were conscious systems in the past, where there were actions and runbook automations and workflows, etc., all optimized around remediation. And these two come together as a unified system. And then you have kind of the analogy of the subconscious and the conscious, where in the subconscious is all this processing that's happening and the conscious is where the human element comes in and increasingly the conscious actions which are the actionable things are also becoming dimensions or various dimensions of autonomy right so this analogy of the human brain is something that we talk about a lot internally when we think about these systems how does that impact the way that you think about using agents for observability you know i think traditionally, as we've talked a little bit about, like observability is a lot of dashboards, it's metrics, it's maybe some level of statistical-based ML to highlight certain issues like anomalies and so forth.

20:28But it's really about giving visibility in the systems for humans. But now with AI agents potentially playing a role there, it's really about giving inputs to machines, not necessarily dashboards for people. So how does that kind of think about how you would build these systems and how is SolarWinds approaching this problem?

20:45Krishna Sai:No, I agree. I mean, there's a lot. Maybe it may help. It may be useful to maybe baseline on a couple of different examples of how we've seen this evolution take place. And then maybe I'll go into more of the design choices that we've had to make. So if you think about, for example, in this evolution of how Gen AI essentially, right, started to help out with the problem solving, diagnostics, resolution, so on and so forth. We saw this phase where there was this co-pilot phase. where there was this Gen AI essentially becoming a very capable interface layer sitting next to the system, not necessarily inside it, but next to the system.

21:26Krishna Sai:So when you're debugging a production issue, the co-pilot then helps you summarize things and explain an error pattern and help you write a query across multiple metrics and traces and so on. It was still very useful, but fundamentally still very reactive, right? So, for example, we have a couple of examples. Just to illustrate this, we have a configuration agent as an example. And the configuration agent is interesting because it is one of those configurations, one of those highest leverage, highest risk surfaces in modern systems. If you think about the number of outages that have been caused by poor configurations, it's pretty massive.

22:06Krishna Sai:And the other thing that's super interesting about that is, if you think about DNS misconfig, which we hear about every other week these days, are certificate expirations and overly permissive security rules and so on. What happens is a lot of them don't necessarily result in a crash, so you don't get an exception that you can go look at, but there's a subtle degradation of surface behavior. You don't realize that the outage actually happened because nothing specifically crashed. The system is just executing to your configuration. What typically happens is in a co-pilot world, that configuration failure would be detected, discovered after the fact.

22:45Krishna Sai:An engineer would get paged and a co-pilot will summarize all of that. But what we're increasingly doing is changing that behavior to where a config agent is continually looking at how the service itself is degraded. And then when there's a service degradation, having all the information that is required to be able to go and correlate that to a config change and then be able to make a very effective choice about whether I want to do a rollback or something else, right? Like I'll bring a human in the loop, so on and so forth. So we're starting to, I would say, see these types of use cases increasingly.

23:26Krishna Sai:Now, what happens is when we think about engineering for a lot of this, there's a lot of, I would say, very important considerations that come, right? And the hardest problems actually tend to be architectural in nature. And that happens, especially when you're dealing with production systems, mission critical systems. How do you think about building AI software that can act in real-world environments and not just a co-pilot? How can you kind of start to build that autonomy? And that becomes a very, very important kind of design choice. So we have this internally, we think of this as AI by design, which is not AI first or AI everywhere, which tends to get kind of misunderstood a lot.

24:11Krishna Sai:But how do you design a platform from the beginning with the assumption that AI-driven components would exist, would evolve, and eventually operate autonomously. I think the mistake that a lot of teams end up making is to treat agents like a feature that you bolt on after the fact, rather than... Then what you end up is you have these powerful models, which people have built billions of dollars building, sitting behind super brittle guardrails. And then you have this problem where, hey, I have this best model. Why is it not giving me the results that I want? Then you realize that you don't have the basic system that is really designed for these things to act.

24:55Krishna Sai:So we have this statement that we use, the model can propose, but the platform must dispose, meaning treat the model for what it does. It's a great reasoning component, but make sure that the platform's a safety boundary. And so when you start to build out these types of systems, then you have to have specific architectural platform components in place, right? And these become very concrete design choices. So, for example, LLM gateways is a great example, right? Like early on, when we started experimenting, you know, teams were experimenting with like wiring logs directly to an LLM as an example, right?

25:32Krishna Sai:It works great in a demo. Executives are super impressed. But then you start to say it immediately runs into problems the first time you try to put anything close to it in production. Because cost spike unpredictably, sensitive data gets into problems. Different teams hard code different models. And then suddenly you can't change providers, enforce policy and so on. So one of the initial design choices that we had to make was actually bring all of that together in a platform service, which is an LLM gateway, which then handles everything from model selection, abstraction, PII masking, rate limiting, auditability, and so on and so forth.

26:12Krishna Sai:So a lot of those shared concerns across expanding your EI cases are kind of isolated. So that's a very good kind of choice. The other one for us is also around, you just can't throw melt data at LLMs, right? Like the logs thing that I mentioned about, it's one thing about generating a whole lot of logs. Logs tend to be super noisy. And so we have to think about how do you feed logs to an LLM or a model for decision making? You know, you can't have, so you need to really think about logs. Are you going to compress them, deduplicate them, summarize them before you ever expose them to a reasoning layer in your platform?

26:54Krishna Sai:So instead of asking a model to read like a half million lines of log lines, you have a compact representation of what changed, what's anomalous, you know, what's new. And then that's a kind of a classic systems approach that we've applied in other use cases, but then you need to bring that into something like an AI system. The same thing kind of also is, yeah. Yeah, so I wanted to go back to a couple of things that you said there at the beginning. So you talked about these like co-pilot experiences where even a really great co-pilot, you're still finding out after the fact that there's some sort of problem and it has to essentially wait for a user to prompt it.

27:35And it sounds like you're trying to move to a world where you have kind of these more ambient agents that are always on. And I really believe I'm a big proponent of this. Like, I think that's going to be the next evolution of the use of agents. You know, there's been a lot of success in the B2C world with AI, especially with ChatGPT. And I think that has been fantastic, but it's also locked a lot of companies into thinking that the only way these interfaces work is it's a chatbot, you know, it's some sort of co-pilot. Maybe there's an agent behind it, but still just a chatbot waiting for a user to ask a question.

Read the full transcript

28:07But in these like operational use cases, I don't want to have to ask if there's a problem. I want the system to know that there's a problem and then do some work on my behalf and then loop me in when human decision-making is required. So it kind of sounds like you're thinking about that in a similar way. And I guess one of the questions I had is, I think there is, in a lot of ways, a substantial difference from both when we think about chatbots in especially the B2C world versus the B2B world of solving specific operational use cases. If I want to constrain this to how do I figure out whether we have some sort of DNS config issue or a certificate has expired?

28:48The shape of that problem is very different than being able to go to a chat interface where anybody can ask any unbounded, unlimited set of questions. So I'm curious, how does that shape your thinking when it comes to the types of models that you might need to use or even the way that you architecting this? Do you need trillion parameter models in that world, or can you get away with something that's more like an SLM that's tuned to the particular problem at hand?

29:13Krishna Sai:Right, right. No, I agree. I think the model definition is definitely very, very important. And we do that. And that's why this example of LLM gateway is a very good one, where depending on the type of use case and the consumptions, you can actually make that choice at the gateway level rather than downstream consumer having to make that choice. And you're absolutely right. Like what we initially saw is that LLMs, as they get bigger and better and so on, they're able to pretty well handle a lot of the generic use cases. And then marrying that with a RAG system as an example. And even there, when you think about RAG systems, there's a lot of design choices that you have to make.

29:58Krishna Sai:But marrying that and then having a system that needs to be able to work in the background, all of those things need to tie together, which is why initially bolting on a co-pilot on top of a set of existing APIs, exposed via MCP as an example, will get you going initially with that very, very basic assistive tool calling. And when you wake up at the 3 a.m. in the morning and dealing with alert storms, you have no idea where to start even, right? Like having a co-pilot which has a few prompts and give you some contextual prompts based on the problem you're looking at are alerts make a lot of sense.

30:36Krishna Sai:but immediately the problem shifts to, okay, what now, what next, and how much of this can I do? And that is where I think a lot of these type of ambient agents, as you call them, come into the picture. And you're absolutely right. Like there are agents where you do need, let's say the heavier agents, like a lot of use cases, for example, even within, for example, we use a cloud models a lot. We use other models as well, but cloud models a lot. And we use the Haiku models for some very, very specific use cases. For example, in our ITSM product, we have this agent. One of the agents that we have in our service management product is when a ticket comes in, so typically, and it gets assigned to a human agent, and a human agent has to go review the ticket, and oftentimes a ticket gets forwarded 15 times before it ends up being the right person.

31:27Krishna Sai:So when an agent comes in, you know, an agent gets a ticket, you have all this history behind it. And there's a lot of cognitive loads. One of the agents that we have is one that will go look at your incident data, previous incident data, and go through that whole process of managing that and generate a response, give you a context summary. So when an agent comes, they not really say, like, this is what's currently going on. This is a customer you're dealing with. This is a sensitivity. And here's the summary of everyone. Sean's looked at it. Si's looked at it before. And here's the summary of what they thought.

32:02Krishna Sai:And by the way, here's a suggested response, and you can quickly edit it as a human in the loop and respond to it. So these types of use cases and experiences. And if you think through this, the beautiful thing about that example is there's so many things that come into the picture. You know, you have your entire system that is working in the background. You have the system of all the architectural choices you have to make around LLM gateways and model choices, RAG systems that process that have to deal with chunking, similarity measures, embedding dimensions, and so on. It has a user experience dimension of actually extending an existing use case so that a human is not completely having to deal with completely new use cases.

32:49Krishna Sai:And an agent first example of the agent's done all the work for you and is looping in a human right when you need it, when that highest point of decision is required. So a lot of these design choices come together in a very, very elegant use case, I would say. Yeah. And I think it's a massive sort of operational efficiency win for the company too, where you don't have to spend sort of human cycles routing tickets. It's probably nobody's favorite job to do. It's super, I mean, this was one of those initial use cases that we rolled out, but the amazing thing was the take rate on something like this was instantaneous, right?

33:28Krishna Sai:And the MTTR, you know, goes up by 30 to 50 % just overnight using a very effectively well thought out use case that comes into your natural workflow, right? I think that's why I'm bullish on a lot of these things, because if you do take the care to have a systems approach to this problem, rather than bolting on flaky agents on an unstable kind of an environment, if you really design it, engineer it, build a platform, build these platform primitives, and take a user experience first approach, then you actually can leverage a lot of this. and talking about ROI and so on, you can see instant ROI in these types of agentic use cases right up front while still maintaining all the concerns around boundaries and choices and so on.

34:22Yeah. It's also easier to bound the problem when it's sort of more use case specific. If you're looking for config issues, for example, then you have more of a bounded input data set to build guardrails around versus just like, hey, someone could ask the weather support for my upcoming family vacation, as well as to do log analysis on our systems logs. That's hard to build test cases for, essentially.

34:51Krishna Sai:That's right. One of the things you also mentioned there earlier was this idea that you don't necessarily want to just take your raw logs and kind of like fire hose it into a model. And it's going to be hard for a model to be able to interpret it. It's probably also going to explode your token cost to some degree because you're putting in a lot of input tokens. And there's just going to be a lot of noise in that data set. And I think for a lot of use cases, it's not really, and I think where companies kind of struggle sometimes is because you have data all over the place, sort of the easiest thing to do is try to attach essentially the agent to be able to pull in the raw data from all those different systems.

35:29And that might work okay in the demo, but realistically, you don't want the raw data. you need a refined data set, essentially. The data set needs to be massaged into something that's kind of purpose-built for the AI system. Just like you don't take a bunch of raw ingredients and call it a meal, you refine your spices and your salt and pepper, your raw ingredients, and you cook it together and you make a meal, and you call that a meal. I think you need a meal for your data sets as well

35:56Krishna Sai:for these AI systems to be successful. So what are some of the things you're thinking about there, and how is that being reflected in some of the tools that SolarWinds is building. No, I agree with that. And I think, you know, in terms of recipes, thinking about that, right? Like you need to have a system in place. And this is where I think going back to the platform choices and decisions that we had to make, especially for a data heavy system like ours, we tend to very, very broadly, we think about our platform in three different planes. And this is not unusual. There's a data plane, which is where we deal with a lot of your, what we call as a melt data metrics events.

36:34Krishna Sai:logs, traces, topology ingestion, and normalization. A lot of that happens in the data platform. And we have the idea of a control plane, which of course deals with all your policies, actions. And then we think about what we call as a reasoning plane, which is something where the intelligence and the agents tend to operate. And that separation, I think, is very intentional because it mirrors, number one, how big systems like Kubernetes, service measures that are already built, but it also ends up, you know, goes back to what you were talking about, which is you're dealing with a lot of very, very operational concerns like distributed state, partial failures, you know, components that should never have implicit authority, and so on and so forth.

37:20Like a lot of those types of things come, which is why having an explicitly

37:24Krishna Sai:permission by design reasoning plane is something that's a first-class choice decision, design choice that we had to make, right? And that distinction actually matters a lot because the aha moment for us was you want models in the reasoning plane to do what they're good at, which is really be expressive, but not be dangerous, which is like the key distinction, right? Like you want your models to do what they're really good at. So by actually separating the concern and saying, let the reasoning plane go do what is really good at, explore hypotheses. propose actions, but the execution of those actions, et cetera, when things have required mutation or action, they always happen through a controlled interface that enforces things like least privilege, et cetera, right?

38:15Krishna Sai:Like that's a very, very critical, I would say, design choice that we had to make. The second one was that when we think about autonomy, don't think of it as flipping a switch, think of it like a tiered capability, meaning that having an autonomy level into every type of action is something that needs to be baked in and built in. So you can think of it as an autonomy level. You could start with something like recommend only, no mutations, right? And then you could evolve to something like execute with approval. And then later on, you can go to execute autonomously, but with certain constraints, right?

38:53Krishna Sai:For certain low-risk, high-performance use cases. And then you can also tune it based on action type, by environment, by service, by team. So having all these types of bells and whistles as you build autonomy into the reasoning plane for your system, I think is super important. And then the other one that is very important in terms of actually realizing these things in production is to make sure that you have first-class transaction traceability for all your agentic actions. Because things will go wrong, and when they go wrong, you need to have the ability to go back and look at the why and what and whose authority did an agent do certain things.

39:35Krishna Sai:And these things are important. You don't realize it's at front, but these things scale very fast. I mean, the moment you give your engineering team the ability to say, go use these LLM gateways and use these agentic frameworks for your use cases, overnight you will see 20, 50, 100 different agents being built and contributed. And unless you have these baseline things in place, it's super important. And then last but not the least, being an observability vendor ourselves, we always keep reinforcing the fact that you have to make observability for your autonomy a first-class concern. And there's a lot of activity happening in forums like OpenTelemetry, et cetera, now to support this as well.

40:19Krishna Sai:So these are, I would say, some of those, I would say both they cross the boundaries of not just being guardrails, but also good sound platform decisions that help you achieve these things at scale. Yeah, I think it's really important to have kind of that running decision log from these systems. And like you mentioned that, you know, when you turn these things on, suddenly you're going to have like, I think, a huge volume of use. in some ways, I made this analogy recently, and I think it's kind of reached true, is it's like when banks suddenly had mobile apps, and then they went from a world where in order to check the balance in our bank account, you had to go to an ATM machine and pull it out, where maybe somebody's doing that once a month if they're really motivated.

41:03And then suddenly it's like, I can do that a hundred times a day, and their APIs are just getting hammered as a result. And they had to scale those systems massively. I think it's similar when you start to turn these things on is the usage is just going to really escalate. You have to think about that as well from a systems design perspective.

41:18Krishna Sai:Exactly. No, I agree with that. And that's where I think things like token usage, et cetera, like a lot of those experiences, compressing logs, these types of things, having good design decisions and saying how you're going to distribute the design choices when you're building a large scale systems greatly help with how you think about all of those things. You know, looking sort of, you know, longer term, how do you see human role evolving in the space? Like, are they going to, you know, be supervisors or auditors, collaborators with agents? something else? It's an excellent question. It's front and center for all of us.

41:58Krishna Sai:The way we think about, I think, a lot of, even as engineers, as an example, like how does an engineer's role evolve, right? As an example in a lot of this. The way we internalize it is what we're seeing is more than just a tool shift. Like a tool used to be able to do X, Y, Z, and it's now becoming this. What we're seeing is that there's a responsibility shift, meaning the role of an engineer or an operations person is clearly changing from or will change from having to be the author of the logic to be how are you going to engineer the context. So the shift from writing logic to engineering a context is a very, very important shift and we're starting to see that and that's why things like coding agents are super useful because you're already starting to do a lot of that in your day-to-day, I would say, example.

42:50Krishna Sai:Like now in a traditional world, responsibility was super deterministic. You wrote the code, something broke, and you're responsible. But with the agentic system, that responsibility actually moves earlier and becomes more probabilistic. And engineers, just our nature, we are super uncomfortable with that. We want everything to be very, very deterministic and want to encode logic. And I think learning to do that is a very subtle, nuanced shift that will happen across the board, right? And so if you think about that, like, and that overcoming that uncomfortable nature of learning that engineers have to go through and thinking about agentic systems, which force you to think in terms of probabilistic engineering.

43:41Krishna Sai:And so that's a very big shift that I think we all need to go through. But I think the shifting of that responsibility from building business logic to engineering context is a very good way to frame how I think our responsibility shifts in this process. Yeah, I think that is probably the biggest fundamental shift that we're seeing is sort of what the role of software engineering is. is historically, it really is about writing these precise rules and logic to be able to represent the business logic. And now essentially your program in many ways almost becomes like your data pipeline. It's that context engineering funnel into the model.

44:23And the business logic is essentially a by-product of the execution of the model. And so part of that is non-deterministic, which is a hard thing to sort of wrap your mind around and get comfortable with. And then it introduces these new things that like we've been talking about that become front and center of observability, tracing, decision logs, evals, replace unit tests and things like that.

44:44Krishna Sai:And you know, the other thing that as you were talking about, the other thing that occurred to me was one of the things that we don't talk about enough is having a sense of like emotional resilience when you're dealing with these systems, right? Because these systems can be super frustrating. There's that initial, aha, like I have the superpower. And then soon they start to get very frustrating because they hallucinate, they make surprising choices. And we've seen that the engineers who actually succeed are the ones that who are able to treat that as feedback rather than failure, right? And it's like the classic, the engineering, you have to kind of go back to the core of your engineering first mindset.

45:25Krishna Sai:And we see that, for example, with engineers and teams that are doing this better. And so a lot of teams will look at that initial results and say, this is garbage. I'll just write the script myself rather than saying, hey, that's interesting. I wonder what context is missing, right? And what constraint did I forget to encode in my prompts or in my system, right? So building that mindset, I think is very critical because agentic systems improve through iteration and they're not like one-off correctness type of systems, you know? Yeah. I think that goes beyond just engineering as well. I think that's true of people are using LLMs for generating market material and things like that.

46:08Like I've certainly seen my fair share of people in that world where that are more resistant to it, they like put one prompt in and they're like, oh, it didn't work for me. I didn't get like a, you know, the perfect pressure release on the first attempt. And they don't know how to kind of take that signal as feedback and figure out how to provide the right context.

46:25Krishna Sai:Right. Or you end up dealing with AI slop, right? Which is essentially the difference is that different systems have different challenges. Writing some content for SEO is one thing where you can probably get away with not getting the exact thing right. But dealing with mission critical systems is a whole other thing. So there's a big spectrum there. And I think going back to how we used to think about even levels of engineers, right, will change a lot, right? It's no more about you knowing the syntax and you knowing about styling guides and effective principles of basic writing software, but it's more about how you're able to architect the systems and how you're able to work the systems to do what you want to do, which is really build great products that people love to use, you know?

47:13Yeah, exactly. I think there's going to be a shift in terms of where we've historically rewarded like really deep expertise in a particular domain. And now, at least for certain classes of problems, I think you can kind of get away with having more of a generalist sort of wide view of how to build things because the LMs are so good at being, having that deep expertise in like the syntax of Rust or whatever programming language you happen to be working in. It actually, one thing along these lines is what are your thoughts on now this kind of, I think, perceived risk of what does this do to the junior developer?

47:47Because we are working at a higher level of distraction now. And I think that initially, a lot of the co-pilots, the perception was, these are great for junior devs, like this kind of levels them up a little bit. But now I think with the Identic Systems and this change of context engineering and really needing to understand how the pieces of the puzzle go together, that's something that generally, more senior people in the organization are good at and are now getting a lot of value out of these models. So of course, the risk, the challenge is how do you get the junior people to senior people

48:18Krishna Sai:if suddenly the junior person isn't getting sort of that hands-on training? And I guess, what are your thoughts on that? And is that a problem that people should be concerned about? Or are we just kind of working on the wrong track? And this may be one way. What we're seeing is, of course, this is all super early and we're learning a lot as we roll out these systems and so on. What we're seeing is both are true, meaning the junior developers tend to be a lot better at adapting and learning how these systems work and being able to get the system to do what you want. So they get really good at that.

48:50Krishna Sai:The more senior devs are really better at understanding, hey, agentic systems don't replace your basic engineering discipline. They actually expose and amplify where a mess exists already. So teams, for example, which have clear ownership, strong data, having good operational fundamentals, then are then able to use agents as force multipliers. And the teams without those foundations tend to struggle and tend to experience agents as unpredictable and challenging and so on. So one of the things that we have done actually internally, which is actually a pretty good practice, is to build these AI communities and have a very, very active, vibrant community where you have senior devs and junior devs actively engaged in sharing knowledge because there's a lot of learning that's happening on a daily basis.

49:44Krishna Sai:and being able to share that knowledge and be able to articulate, hey, here's what worked, here's what didn't work. And by the way, as we are going through this, here are some architectural foundations and best practices that we need to put in place on which we can build greater production systems and so on. I think having doing some of these things greatly tend to help is what I've seen. Yeah, absolutely. I think that one of the misses sometimes companies make is that they buy these tools and then they think that overnight There's going to be 100 % efficiency gain from it. And they forget essentially that it's a new skill set that you have to give people time to sort of train up on.

50:22And that's a really important part of the process.

50:24Krishna Sai:The other thing which is interesting that you mentioned, because this is the other challenge that most companies that are adopting these, is having to justify the ROI for spending a lot of these tools. Yes. And that's a common challenge. And the way we try to balance that is there are like focus on the cases where ROI is very quickly, shall we say, well established, number one, as a system. For example, if you take customer service, right, like they all they have they're highly metrics driven. They know what works for them and they know what doesn't work for them. So you can actually measure your ticket deflection rate or your time to first response or time to resolution.

51:06Krishna Sai:and these things can be actually measured. Engineering productivity, on the other hand, as you know, is not at all well-defined. There's a lot of, we talk about late times and commit velocities and all of that, which is good. I mean, we try to somewhat early, but what we are really poor at is being able to try a lot of that to business outcomes. So again, going back there, for example, we had to do a lot of work in terms of it's early, but being able to quantify certain metrics in terms of, you know, whether your lead time is improving, as an example, your change failure rates decreasing, etc.

51:47Krishna Sai:Like a lot of those types of things, having that in place also tends to help, right? So you know that you're trying something and you're learning, but over a period of time, you're starting to get those signals or green shoots of, hey, this is actually helping me do things better, work better. Yeah. I mean, I think the reality is it doesn't matter how powerful the models are. You still need to have, do the work essentially around putting the right metrics in place to measure what success looks like. I think sales is another, similar to customer service is another place. So it's like, if a sales rep, they're going to use whatever tool allows them to hit quota more often.

52:24So if you can unlock that, there's a clear ROI for the business. Well, Cy, I want to thank you so much for being here. I really enjoyed the conversation. And the last thing here, Is there anything else you'd like to share?

52:35Krishna Sai:No, it's just thank you again for the opportunity. And we're super excited about where all of this is going. And we're taking a very comprehensive approach to how we look at this. And we're definitely excited about moving the needle, so to speak, in terms of how a lot of these things continue to add value for us, even as we adapt and learn. So thanks again for the opportunity. Great. Well, cheers. Thanks. Cheers. Thanks, Sham.

53:03Thank you.

From the publisher

Enterprise IT systems have grown into sprawling, highly distributed environments spanning cloud infrastructure, applications, data platforms, and increasingly AI-driven workloads. Observability tools have made it easier to collect metrics, logs, and traces, but understanding why systems fail and responding quickly remains a persistent challenge. As complexity continues to rise, the industry is looking beyond dashboards

The post Engineering AI Systems for Autonomy and Resilience with Krishna Sai appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Engineering AI Systems for Autonomy and Resilience with Krishna SaiSoftware Engineering Daily · 53 min
Listen in VO