In short
How BMC Helix uses agentic AI to automate IT service management (“service ops”) by detecting issues from telemetry, performing root-cause and impact analysis, and generating economical remediation plans—initially with human approval, moving toward more autonomous execution.
Guests
Arhan Giral (spelled “Arhan” in transcript), runs the AI Office for BMC Helix; background in monitoring, early work on monitoring applications/infrastructure, reasoning over high-throughput machine data streams. Ryan Manning, co-developer applying AI to service operations workflows at BMC Helix.
Key claims
AI agents will shift IT from ticket-driven workflows (e.g., Jira/ServiceNow) toward real-world, first-person learning via “gyms” (safe data-center labs). Agents reduce noise in observability data, ask “why” repeatedly to reach actionable mitigation, and use “fingerprinting” to distinguish recurring incidents from novel ones. Enterprise-specific “reasoning” is achieved via fine-tuning and mixture-of-experts; plans are optimized to be lower-risk and lower-cost. Execution is assistive now (human-in-the-loop) with causal traces to combat automation bias.
Notable examples
anomaly detection (e.g., unusual ATM network spike at 8:30 a.m.); root cause like a busy router due to a full queue/broker; deep root-cause postmortem agent that delegates to log/metrics/topology sub-agents.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Changing Landscape of IT
0:00 to 0:34
Discussion on how automation is transforming IT staff roles and workflows.
“It sounds like it's going to get increasingly automated.”
Understanding Service Ops
0:42 to 2:08
Overview of service operations and its importance in digital business.
“I work with Ryan on applying AI to various different problems in service ops space.”
Evolution of Service Management
2:08 to 4:25
Exploring the changes in service management with AI agents.
“Can you begin, one of you, by talking about how the landscape has changed with the introduction of AI agents?”
The Role of IT Service Management
4:25 to 6:40
Explaining the critical nature of IT service management in enterprises.
“And you guys, we'll give the background of BMC first.”
BMC's Historical Context
6:40 to 8:10
Review of BMC's evolution through different eras in service management.
“Yeah, I like to think about BMC in four eras.”
Integrating Service and Operations
8:10 to 11:11
Discussion on how BMC integrates service and operations for efficiency.
“And there's the service management being the front door to IT can encompass a lot of different requests that need to get fulfilled.”
AI and Anomaly Detection
11:11 to 14:01
How AI helps in detecting and analyzing anomalies in IT systems.
“CMDB where they store all their assets with asset information.”
Understanding IT Problem Detection and Impact Analysis
14:01 to 16:52
Learn how systems detect IT problems and assess their impact using real-time data.
“But it's like there's always this cosmic background noise that you have to essentially first detect and then push away to then focus on the real signals in the data to understand what the root cause is.”
Leveraging LLMs for IT Reasoning
16:52 to 18:50
Discover how LLMs analyze IT issues and provide actionable insights.
“and then we ask LLM, okay, I mean, we ask various questions to LLM, but most importantly, we ask the question why and what needs to be done.”
Implementation of Mixture of Experts Architecture
18:50 to 23:54
Understand the mixture of experts architecture and its role in model training.
“they become extremely, extremely capable.”
Show all 27 chapters
From Problem Identification to Actionable Plans
23:54 to 27:37
Explore the process of turning identified problems into actionable plans in IT management.
“so that whatever generates plan or output summarization or what have you is now affected by that training.”
The Role of Human Oversight in Automation
27:37 to 28:00
Examine the necessity of human oversight in automated IT systems and the concept of automation bias.
“as they respond to tickets and leave little digital traces of their actions.”
Automation Bias and Incident Fingerprinting
28:00 to 30:02
Learn about the concept of automation bias and how fingerprinting techniques help identify issues in IT.
“because it's worked the last 10 times, and so the 11th time you don't even think about it.”
Root Cause Analysis with AI Agents
30:02 to 32:01
Discover how AI agents conduct root cause analysis by formulating hypotheses based on data.
“And that is actually in IT, in this sort of problem space, it's very important to know whether you are dealing with a brand new problem versus something that happens with some frequency.”
Collaborative Hypothesis Testing Among Agents
32:01 to 34:38
Explore how different AI agents collaborate to test hypotheses and refine their findings.
“a consensus about, you know, which, what is the optimal output or plan, whether that's a plan in action or, yeah.”
Modular Approach in AI Service Architecture
34:38 to 37:18
Understand the modular architecture of AI services and how users can customize their experience.
“but I didn't find that, but I found this other thing which suggests your root cause analysis was misinformed, meaning I found more evidence for you.”
Deployment Flexibility and Data Privacy
37:18 to 39:42
Learn about the deployment options available for AI solutions and the importance of data privacy.
“They say, hey, I'm trying to do this, I'm trying to do that.”
Containerization and GPU Requirements
39:42 to 42:00
Find out how containerization is used for deploying AI models and the necessary hardware requirements.
“So for that reason, we have, you know, we built a lot of scaffolding around these models so that A, they run very efficiently.”
Understanding the TX6000 Inference Server
42:00 to 43:55
Learn how the TX6000 system transforms into an inference server leveraging AI.
“and nowadays our TX6000s are pretty good.”
Metrics and Long Sales Cycles in IT
43:55 to 45:55
Explore the metrics companies use to evaluate IT solutions and the lengthy sales processes involved.
“you periodically, let's say every 24 hours, your model revs up.”
Shifts in Software Engineering and IT Operations
45:55 to 48:25
Discuss the recent disruptions in software engineering and their implications for IT operations.
“so we do have a portfolio of savings, measured savings that we share with them.”
Creating AI Training Environments
48:25 to 50:26
Learn about the development of 'gyms' for AI agents to enhance their training through real-world exposure.
“And that's why that was the first thing that got heavily, heavily automated now.”
Innovations in AI Model Training
50:26 to 52:30
Discover the approach taken towards developing foundational models for AI and the importance of empirical testing.
“I keep a record of how things are progressing and whether the remediation works.”
The Future of IT Staff Roles
52:30 to 56:00
Examine how automation in IT will change the roles of IT staff and their required skill sets.
“to hear that that research is going on are are you developing your own models in in uh for that or do you use, you know, models from one of the foundation labs and then fine-kill it?”
The Role of AI in IT Management
56:00 to 56:58
Discover how AI agents assist IT managers by handling repetitive tasks.
“managers of agent systems as opposed to having their fingers in the code?”
Gaining Digital Capacity with AI
56:58 to 57:55
Learn how AI technology provides digital capacity for organizations to innovate.
“They say, well, you know, this saves me a lot of time.”
Impact of AI on Workforce Dynamics
57:55 to 58:58
Explore the economic shifts and changes in workforce dynamics due to AI integration.
“So we see this, you know, a certain cohort in every organization sort of really shining, you know, with the help of these technologies.”
Transcript
Automatic transcript. May contain errors.0:00What's that going to do to IT staff? It sounds like it's going to get increasingly automated. That space is changing pretty quickly and that workflow is changing. Egentic AI pretty dramatically. Right now AI, if you think about it, AI is only learning through someone else's description, right? People will start exposing these models to the real world more and more so that they can create, they can gain that first person view of things. Once we detect something, we always ask the data, why, why, why, until we get to a point where we can take an actionable step to mitigate or perhaps even remedy that issue.
0:34Why don't you introduce yourself to listeners? My name is Arhan. I run the AI office for BMC Helix now. I work with Ryan on applying AI to various different problems in service ops space. So service ops is essentially you can think of that as any digital business these days is really a conglomeration of IT services you can imagine. And service ops is essentially an attempt to automate operational activities around these services as much as possible. so that we look at enterprises as a combination of software, infrastructure, and human beings operating on these assets. And we provide various products and nowadays agents to automate different workflows around these services.
1:32My personal background is in monitoring space. I've done early work on monitoring applications and infrastructure and then combinations of them, which means you have to process a lot of machine data. You have to be able to reason about high throughput streams of data that these data centers and these machines are constantly emitting. So, yeah, that's what I do in Helix. We're going to talk about Helix and IT service management. Can you begin, one of you, by talking about how the landscape has changed with the introduction of AI agents? What the process for service management was before and then what it's migrated to today?
2:27Yeah. Thank you.
3:06yeah uh and and you know just thinking back in the old days it was jira right you would you'd open up a jira ticket uh it is uh it is does uh your solution work alongside jira replace it just for people that aren't that familiar with. Yeah.
4:04Thank you.
4:25And you guys, we'll give the background of BMC first. You were talking about that before we started.
4:50Yeah. Yeah. So, I mean, service management, the discipline is complicated. The explanation of what they do is pretty simple. It's the front door to IT. So if, you know, I'm a new employee trying to get my laptop provisioned or if I'm having an issue with an application, I might go to a portal, send an email, call the service desk and try to get that remediated. And that space is changing pretty quickly and that workflow is changing with Agentic AI pretty dramatically.
5:54Yeah, no, that's, Jira does have it. So our competitors, you know, go beyond service now into Atlassian, into Freshworks. The market is a lot of players. The interesting thing about the service management market is that unlike the CRM market, to broaden it out for a more general audience. Describe what IT service management is and how critical it is to enterprise and the economy in general. In the high enterprise where we operate, we'll say the Fortune 2000, you submit a ticket, you're either submitting it to ServiceNow or to BMC Helix. Down market, Jira is a very popular tool. Freshworks is a very popular tool.
6:44I volunteer a few others. Yeah, I like to think about BMC in four eras. Era one, remedy rules. A lot of folks know remedy. That was an acquisition that BMC made years ago and was the ServiceNow and Service Management before ServiceNow. It was the on-premise ServiceNow. Era 2, about 2010, the music began to slow a little bit, lost our way a little bit in terms of hopping on the cloud train. The company ended up getting bought by Bain and taken private. It sort of continued the same mission. then there were three KKR came in and said hey ServiceNow still only has one competitor it's VMC Helix what if we took a different approach what if we brought in a technical team in the operations of those companies it's kind of the critical Broadcom I'm in Syed, Ali Siddiqui were asked to come in and rebuild the platform from the ground up you know, AI first.
8:09And that, that journey, you know, took us to close to where we are today where we, you know, went through the natural progression of the heartaches for our customers and for ourself and rebuilding a platform and still servicing some of the biggest companies in the world to now being able to innovate and separate ourselves from the competition with the innovation that we got access or permission to build through KKR's ownership.
8:51Yeah. Yeah. And there's the service management being the front door to IT can encompass a lot of different requests that need to get fulfilled. Some of them are very simple you know i need a new iphone um some of them are hey the quoting tool is down it's the end of a quarter um we're in trouble and there's all these workflows that sit behind that request some of them more complex um when those more complex general model issues come in talks to the engine room i think everyone everyone keeping those services up what we've been trying to do is there's a process to remediate that issue and so what we did in that transformation i talked about is we brought those two disciplines together service and operations naturally in category service ops because when you because when you end up sharing data ryan was talking about communication you have to be able to quickly process all the telemetric data You don't have to worry about lots of outages taking too long to restore that outage.
9:58Understand that graph and those relationships. And certainly less regression to the same thing happening again. Natively. And in IT, it's very common that you have to constantly deal with deep relationships. Because something that runs on a computer turns out it's actually not a real computer. It's a visual computer that is resourced by a physical computer somewhere else. So there is all these resource dependencies, resource and transactional dependencies. in the fabric that you have to understand before you can be, you know, you can produce anything actionable and useful. Because on the operation side, what we want to create is not, you know, on the service management side, a lot of disruption, what we want to eventually create is we want to be able to create new execution plans that other agents, that other missions can pick up and then one of Microsoft teams type something in a virtual agent.
10:51And that does require to be able to do that in a reliable sense. You have to create repeatable, predictable infrastructure from these models. And there is so many relationships between investment in that regard to keep that service option and running together and viable. CMDB where they store all their assets with asset information. Can you walk us through an 18 million relationship amongst those assets? And so going to that space, it's not like I was just plugging LLM in. And Iran, hopefully you can expand on this. And you get your answer and you get that remediated. From outside in, it's actually very simple.
11:32So you first need to understand, hey, do I have a problem? To get to root cause. Why is that happening? Should I mobilize anything at all? Do I have a problem? Awareness and identifications of problems is one use case we constantly deal with. And that requires our agents to stream this telemetric information. This can be your machine logs, time series information, alarm information that they emit. So something needs to be there and constantly monitor these things. And we have a funnel that deals with that. Let's take the observable data and run that through various machine learning models so that we can understand, oh, that spike that you see right now at 8.30 a.m.
12:17on a Monday on the bank's ATM network is actually kind of unusual. We don't expect that spike to appear at 8.30 perhaps. Maybe it's supposed to come in later. so something needs to essentially constantly reason about, hey, you know, what's unusual or given the time we are in or given the workload we are dealing with, given the circumstances in general we are. So that's one part of the big use case we try to do. And of course, once you detect an anomaly, once you detect an availability problem, you then first ask, okay, so why? You know, what is the root cause of this issue? You know, why is this router so busy, you know, where it should be, you know, close to idle state, for instance.
13:01So then we call that the root cause analysis. So essentially something that needs to dig into all that telemetry data, dig into all that observability data to say, ah, you know, it's actually, it's not that, that's a symptom. what's really going on is such and such file system or such and such network queue, such a broker somewhere is now full and dropping the traffic or what have you. So we always ask, once we detect something, we always ask the data, why, why, why, why, until we get to a point where you can take an actionable step to mitigate or remedy that, or perhaps even remedy that issue.
13:46And while doing that, of course, you know, data centers and IT is very target-ish. All kinds of stuff happens all the time. So you need to be able to do this by sifting away all the noise, all the background noise that typically happens in that environment. But it's like there's always this cosmic background noise that you have to essentially first detect and then push away to then focus on the real signals in the data to understand what the root cause is. And then we also do impact analysis. Like, okay, so is this a big problem? Because sometimes technically challenging problems are just that, you know, no one cares about.
14:29Maybe it's a QA environment, maybe it's a staging environment that's being modified at the moment. So essentially from an economist's point of view, we always ask, okay, we have a problem, but is this an impactful problem? Should we mobilize human beings? Should we spend more resources in mitigating this issue? So our system constantly makes these decisions in the background by looking at real-time status of systems while also constantly comparing the state to its past self. And we also make plans of action based on plans of actions that were executed in the past by human beings or by bots. And this, well, first of all, this is all being done within an agentic layer.
15:24It's not, and with reasoning models, to look at the problems and try and figure out. It's not kicking questions out to the human team who then sit around and try and figure it out. It's doing this internally, right? That's right. So some parts of what I just described that is actually taking that gigabytes of gigabytes of monitoring data and then reducing and then essentially pushing that noise away. That's typically we employ proprietary technologies to do that because there's so much data to process. We couldn't really hope to expose all of that to the generative models all the time. So we take all that sparse data and run through various machine learning and statistical analyzers first to basically find the needles in the haystack, if you will.
16:29But once those needles are found, they are collated, combined, correlated, causally analyzed. And then the LLM, you know, the LLM is given a pretty comprehensive, causally described, like a patient chart, like a set of x-rays of the system. and then we ask LLM, okay, I mean, we ask various questions to LLM, but most importantly, we ask the question why and what needs to be done. And that level is completely agentic, just like you said. And there we use reasoning models, just like you said, maybe with a twist, because we want our models to reason just like one of the employees of our customers.
17:25Because, for instance, if you go to an open AI model or a generic off-the-shelf model, they'll always give you plausible responses, accurate responses based on the documentation, based on what they know about the space. But it's often the case that when you go to a big organization, the rules of engagement around these systems are very different. Meaning you can say, oh, go to this file and make these edits and everything should be fine. but then that bank's knock, you know, network corporation personnel will probably say something like, well, you know, that sounds plausible but that's not how we do things.
18:02You know, I need to first, you know, run these additional tests. Maybe I need to talk to some other team, get their approval. You know, I need to make sure, you know, there's a backup procedure. I need to estimate this. There's lots of lots of things that they need to account for and then that's often specific to the enterprise, to the day they work. So what we do is we take the reasoning capabilities of these models and we fine-tune them so that they can reason just like that employee of that particular enterprise. So we found that when you tune these models, especially the planning aspects of these models around the subject matter experts' practices, they become extremely, extremely capable.
18:56They really produce actionable and relatable results for the users. So just to point to your question about reasoning, yes, we rely on reasoning, but we try to guide that reasoning towards economic ways of solving that problem. That's another thing great. So a lot of times when you ask complex IT questions to these models, they'll come up with plausible plans. But then is the plan actionable? Essentially, if the solution is, you know what, tear down everything, rebuild everything, this is a rewrite situation, then the customers won't be happy. They'll say, that's too expensive. I can't really do that.
19:39And so then we train our models and we reward them in ways such that more economical, more expedient ways of lower risk solutions are preferred over others. So there's a lot goes into the guidance of that reasoning. But at the core, just like you said, we do rely on LLM's reasoning capabilities to figure things out, to do the orchestration especially. Yeah. And so that whole reasoning workflow, is that one module? I mean, we were talking before we started recording that you guys build your modular. So you have these microservices that work together that presumably then you can configure or pick and choose among them, depending on the use case.
20:39Is all of that reasoning one module? Yeah, yeah. So we follow the similar philosophy of divide and conquer, you know, create good encapsulations around capabilities. So we follow an architecture pattern called mixture of experts. Yeah. What this is, is it allows you to, you know, these pre-trained models are trained on piles of piles of publicly available textual data, if you're talking about language models. and what we found was actually what we not we didn't find this but a meta sort of led the way into this architecture in our in our adoption they they described they had this wonderful people paper to talk where they talked about you know how how these models can be further tuned for different purposes while encapsulating that training in a specific set of weights and biases called experts.
21:42So what that is, is you take one of the open weight models. Open weight model means a model that you can basically download your computer and run as much as you want based on, of course, depending on the licensing agreements. So what you do is you take one of these open source models, and then you identify a problem data set and a reward function. So you essentially, you basically devise a clever way of signaling the network, hey, you know, here is, when you give me a good response, here, you know, I'll give you cookies. If you don't, if you fail, you know, you'll get the stick. So you basically train these models on additional data sets that allow you to focus on that problem.
22:33And then you capture these as what we call experts. So experts are, just like you're saying, they're mini, I mean, I wouldn't say they're services, but they're encapsulations of that training. They are the artifact of that training cycle. and within this mixture of expert architecture, you also specify gate activation values during your training, which means the model then knows when it's prompted with a question whether to activate that training set or not. Like, for instance, if you ask our model, you know, what's the weather like in France today? It will give you an answer. It will do some tool calls and figure out the answer.
23:15but in that prompt, nothing will signal that, oh, they're asking me a question about like a root cause analysis on a mainframe Z system. So it will detect that and it won't activate that part of its training. But if the question is, hey, I'm in big trouble, how do I restart, how do I reinitialize this L part on my mainframe? Then that activation layer kicks in and says, oh, okay, so it looks like they're talking about this other thing that I was trained on, and it loads those weights and biases that you create during your training. And then those weights and biases sway the network just enough so that whatever generates plan or output summarization or what have you is now affected by that training.
24:09We found, and it's not just us, but also this is throughout the academia and now also in the industry, the best way to generalize a data set, a training data set, is really to go through these trainings so that the model can reason about the native in a model-native way. So that's why we've been pursuing this for some time now. Yeah. And then so once the reasoning is done and the system has identified a problem and designed a plan of action, is that then passed on to an agent that executes that plan of action? Yeah. So when we first started, we just started with a textual recipe that we would hand off to the user and say, OK, so here's what you need to do.
25:09This is the recipe I need you to follow. And maybe there's like eight to 12 steps in there. And then we call it a day. But of course, you quickly realize, well, OK, so first of all, some of these actions are very automatable. Like once we break it down to a problem, a problem into little steps, it already, you know, automatically suggests automation. So in that regard, what we first said was, okay, so a lot of these are first, let's verify the problem. Let's do additional diagnostics. It's actually a lot of read-only operations. So we categorized, we thought, you know, we trained our model so that it knows how to categorize and label these actions as, well these are diagnostical so you'll never break anything by just doing these so we started automating those as you can imagine some of the actions are on the remediation side of things which means you know oh you need to go to this configuration and actually change that port number or go to this machine and allocate more file systems so it's actually things that touches the system We still require a human in the loop for them to approve the plan before any such automation is actually called upon.
26:31So right now, essentially, we are in an assistive capacity, which means these human beings are now fed these recommendations and AI-generated plans of action. and as they do things and as they adhere to the plan or deviate from the plan, we actually track that lifecycle, Craig, so that if any of our recipes are executed and everything is fine, that's great. That's great feedback for us. But we actually learn from our mistakes more, meaning if, let's say, for instance, we tell them, hey, you need to do A, B, and C, but then if they do A, B, and X, then we actually learn that once that ticket is closed so that we can, even if that subject matter expert is not really donating and writing a lot of descriptions of things we want to essentially understand their intent and the real way they fix the issues and we train our models as that flywheel turns as they respond to tickets and leave little digital traces of their actions.
27:45The human in the loop. With a lot of these autonomous systems, there is something called automation bias where people become accustomed to accepting the machine's recommendation because it's worked the last 10 times, and so the 11th time you don't even think about it. You just say, go ahead. I mean, is there a point at which this will be fully automated? And, yeah, and the human users will sort of be monitoring outcomes but not necessarily approving every action. Yeah, that's certainly where it's going. There's definitely huge economic pressure in automation, of course. One thing that you said is automation bias is very true.
28:54to address that we have built various quite interesting fingerprinting techniques of incidents and issues essentially when something goes wrong we keep a very detailed record of what the machine signals were what human beings have done and locality of that issue how that issue transpired like what was the first domino that toppled over what was the second domino. So we create these causal traces of incidents and issues in the enterprise so that we can say, oh, you know, what's happening right now is something that just happened last week. And it seems like every Friday you guys are going through this.
29:40Hence, it's highly likely that this is that issue. So it's actually to combat that and also give confidence to people that this is something they can actually tackle with automation. We have done a lot of work on fingerprinting, fingerprinting issues. And that is actually in IT, in this sort of problem space, it's very important to know whether you are dealing with a brand new problem versus something that happens with some frequency. There's tremendous difference between them. Essentially, being able to say, you know what, I don't know, I'm throwing the towel in, you really need to pay attention to this is a great asset.
30:32That's not a failure of the product. When it does that, we are extremely happy that we've been able to do that because one thing that you've noticed probably, LLMs will never tell you, oh, you know, I don't know something, right? They're very positive. They love to generate. So it's actually to get them to say something like, well, you know, this is really out of my bounds is difficult. And that's something we've been working on ever since we started this. But of course, if you can sufficiently fingerprint these incidents and issues and can sufficiently confident to say, hey, this is actually something we dealt with and here is how we solve these problems, then there is a lot of economies there.
31:27So you can imagine that the pressure in the market is always towards, okay, so can we fix these reoccurring problems, which tends to be like 80, 85 % of the problem space? What can we do to automate these as quickly as possible? That's essentially what people like Ryan and me think about all day. Yeah. Well, one of the things that people have been doing is they build this kind of a committee of LLMs that check each other's work and then come up with a consensus about, you know, which, what is the optimal output or plan, whether that's a plan in action or, yeah. Are you doing anything like that?
32:18Yeah, we absolutely do, which is interesting. So we have this new agent called Deep Root Cause Analysis Agent. So this is something that you unleash on big problems in a postmortem sense, meaning problem already happened and you're trying to understand why that was. And you have a lot of time, so you have a lot of research time. It's actually, it's not a, you don't have to be, you know, counting milliseconds necessarily. So when this thing is launched, it quickly formulates an opinion, a theory about what it is based on all the available data, all the snapshot data that was handed to it. But it has agency.
33:06So, in fact, when we say agentic AI, what that really means is, hey, do you have a system that can take initiative to dig into data and come up with an outcome? Come up with an outcome that is towards, that is in the trajectory that you think it should do. So, for us, that's root cause analysis, meaning we task our agents to say, hey, here's all the telemetry data that we have. Here are all the data connectors that you might want to pull in, you know, in an agentic manner if you want. And I need you to get to the bottom of this and do a deep root cause analysis on this. And when you test this agent, like I said, it formulates a hypothesis.
33:55And then this hypothesis is then passed to other subagents in the system. We have an agent that is expert in log analysis, analysis of log streams and log data in general. So when this hypothesis is passed down to that agent with the task, because our planner says, okay, so I suspect it's going to be the load balancer again, so somebody needs to go study the load balancer logs, and it hands off to that agent. And that agent says, okay, I guess this is what I do now. It goes to the log source and then starts analyzing the data. But we give that agent enough creative space such that if it finds a refuting evidence or supporting or refuting evidence in the data, it can signal that back to the other agents and say, hey, you know, what you told me was to look for this.
Read the full transcript
34:58but I didn't find that, but I found this other thing which suggests your root cause analysis was misinformed, meaning I found more evidence for you. Perhaps you should reconsider. So we don't have a peer-to-peer-to-peer agent forum like maybe you're imagining, but we have a hierarchy of like a principal researcher versus individual clerks that are very good at researching different types of machine data, like logs and metrics and topologies, like graphs of things. So those are individual areas of expertise. And they know how to test the hypothesis and also create new insights that the master planner can then reason about.
35:46And they constantly back and forth, do this until they converge on a response, like everybody's happy. or exhausted. Yeah.
36:00We were talking about how you're containerized, and these are built as microservices, and you have a list of these services. From the users, the customers' point of view, how does that work? Do they choose to use some and not others, or is that just architected that way because it's easy for you to swap modules out as you improve things? Yeah, so when the platform was first designed, it was just like that. It had that a la carte feeling of, okay, so here is what I need. Here is, let's say I run a SRE team, site reliability engineering team. So for my team, I need a workflow like this. I'm going to need this service to be modeled and then I need this workflow to be defined.
37:00So it certainly was like that so that you could cherry pick the services that the platform offers you. But in the agentic world, we are quickly migrating everything towards agents nowadays. This actually is a lot smoother and transparent from the point of view of the user because they just naturally conversate. They say, hey, I'm trying to do this, I'm trying to do that. and there is sufficient intelligence baked into the agentic layer of the platform that it knows how to delegate these to all the other bots and models and what have you. So we have a big umbrella term. We call it Helix GPT, which is essentially all the generative, all the reasoning, all the analysis tasks that the platform is generating is then routed to either a model provider of the user or one of the models we host for the user either on their on-premises or in their private clouds versus maybe a cloud offering of choice.
38:16So this is the flexibility Ryan was talking about. We actually, you know, we are able to deploy the components pretty much everywhere, including on-premise. So we have a rather sophisticated routing layer that knows what workload goes where, and it does all the metering and all the compliance and whatnot. There's something about what you guys have built with Helix that he was saying ServiceNow isn't able to do that because we didn't get to the why. I would say we have, you know, you can run us in any cloud or hybrid or on-premise scenario. So from that perspective, we do have a unique stance, I suppose.
39:09Meaning a lot of organizations are pretty nervous about, you know, sending their most intricate IP and data to third-party AI vendors. because they see how capable these models are, they can quickly learn things. And so there is a lot of interest in being able to run this level of intelligence in an air gap manner, if you will, so that nothing leaves that organization's boundaries. So for that reason, we have, you know, we built a lot of scaffolding around these models so that A, they run very efficiently. You know, we don't, we actually, we look for parameter efficient models as much as possible, meaning I need to be able to run all this intelligence on highly available, not so cutting edge hardware.
40:12And that does require some work on, like I said, being able to,
40:20train the model so that you don't essentially activate necessarily the entirety of that model for every single token. Also being able to support multiple tenants perhaps on one installation, one single tenant. So we've done a lot of work to make Generative AI not just possible and exciting, but also economic to run. So from that perspective, I think we do have, like I said, I don't do a lot of competitive analysis. It's part of my job. But from that perspective, I know we have pretty unique features and advantages. Yeah. And if somebody wants to run this on-premise or part, if it's a hybrid solution, you know, on-premise and in the cloud, how do you deliver the model or the modules to them?
41:19Is that just they download it and install it on their side? Pretty much, yeah. So, you know, we talked about containers and containerization. So this happens to be yet another set of containers you deploy as part of the product. The only perhaps maybe slight difference is the machine that's going to run this container will now need access to a decent GPU. And we support pretty old hardware all the way from NVIDIA's L4s to A100s and nowadays our TX6000s are pretty good. So we just require them to procure or lease or find a GPU and then we give them the software architect that basically is part of our platform really.
42:21And that thing turns our model into an inference server that all the other things in the system can then start utilizing. And also the training. We have this optional tenant-specific training pipeline also, which what that is is, you know, they just go to our administrative UI and tell us where their sources are, where their ticket sources are, or their knowledge articles, runbooks, and what have you. And what we do is we harvest all that, and then we analyze the data to see what can be learned from them. So that's also pretty, I would say, interesting because that allows us to keep up with the organization because no IT organization is constant, right?
43:11So they deploy something, maybe their network is all Cisco hardware, but all of a sudden somebody introduces new Juniper hardware somewhere. So they shift, right? And as they adopt new technologies and new components, all the failure modes change. they bring their own problems they bring their own comparabilities so there's always movement that we have to deal with so for that reason we tap into these sources and periodically pull the data out to see what needs to be generalized and what's learnable from them so anyway so that's also another component that goes with this generative AI inference server and once you do that you have you periodically, let's say every 24 hours, your model revs up.
44:03So it might decide to learn new things and update itself. That way it's never behind. So, yeah. Yeah. Yeah. Do you have metrics that demonstrate how this Helix performs in a large organization? I mean, these are big decisions by organizations. I imagine that much of your customer base has been with you for a very long time. But to get somebody, Fortune 2000, to switch, I mean, they've already got a solution to switch to Helix is, I would imagine, a long sales cycle. So what do you present to them as evidence that you guys save them time and money? Yeah, yeah, yeah. So, yes, we always, you know, we have to refer to some measured numbers, some customers.
45:21But what happens, Craig, at the end of the day, it always comes down to a bake-off of sorts. So they, just like you said, these organizations don't buy these tools just to use for a year, right? They typically plan for three to five years out. and it's often the case that they ask us, okay, so what can you do now? But we also need to understand where you are going to see if our goals align. So just like you said, these are actually really complex, long sales cycles. What we tell them is, so we do have a portfolio of savings, measured savings that we share with them. So based on the use case, this typically anywhere from 25 to 40, sometimes 50%.
46:12So we give them these numbers. But of course, at the end of the day, we have a lot of benchmarking information, both synthetic and open source and third party. We have competed in some competitions to show people what these models can do in isolation. So we have a lot of proof data. But at the end of the day, like I said, often what happens is you get deployed in a pre-production environment or some QA environment that they're choosing, sometimes together with your competition with other tools. and then they just measure. They just say, okay, so given an operator and this tool, then they maybe sometimes inject a problem and then they see, okay, did you find it?
47:09Did the operator gain any insights from you or so on and so forth? So essentially this sort of software, I've never seen being sold unless you actually do a proof of concept so that you collect some data yourself. And on where you're going, I mean, what is on the roadmap? So you, I mean, you see how software engineering got disrupted, right? Software engineering, I would say within the last six to eight months, is now different, fundamentally different. We had tectonic shifts in software engineering. And we believe a similar disruption will also happen in operations, in ops. And the reason is software engineering is a very verifiable problem.
48:07So you can generate a piece of code and then compile that code and say, okay, does it compile or compiles? Does it run? You can run some tests against it to see. So it's like you can imagine you can create training flywheels for software pretty easily. And that's why that was the first thing that got heavily, heavily automated now. We believe ops are also close to this. I mean, it's the verification cycle of software is maybe measured in minutes. For ops, that's typically hours. It may be perhaps days. But at the end of the day, it is still a verifiable problem because these are digital systems at the end of the day.
48:52And IT system is, it might look complex. I mean, it's a complex system, but it's not a chaotic system, right? So at the end of the day, these are still computers running software. And, you know, it's a deterministic system. system. But the question is, can you verify this quickly enough to be able to train your model? And that's exactly what we are doing right now, Craig. So right now, most of our research effort is going through building gyms or labs, if you will, where the agents can go and get trained on. So large language models were a function of text data that was available on the internet or offline, right so they're they've gone through piles of piles of text data but we are probably at the limit of that data so nowadays in the future what's going to happen is people will start exposing these models to the real world more and more so that they can create they can gain that first person first person view of things because uh right now ai if you think about ai is only learning through someone else's description, right?
50:05Text means, you know, someone wrote something and that's their interpretation of reality, right? So now everybody's trying to cut that off and expose the AI directly to the reality of whatever domain they are in so that they can, you know, they can do things and learn from that. And in our space, that means creation of these, what we call gyms. So gyms are essentially, imagine a data center that just runs a whole bunch of different applications, but maybe nothing is sensitive or created for the purpose of training in the sense that I can let my AI go there and break things. I can say, okay, so create like a chaos monkey agent that just goes and randomly breaks things and say, ah, you know, I messed up your application, which then my other AI comes in and says, okay, So let me first verify what's going on.
51:00Let me try to fix this. And it tries different things. And I log all of this. I keep a record of how things are progressing and whether the remediation works. Is it expensive? Is there a better one? So I can do all these in a 24-7, lights-out manner in a data center so that I can create a bespoke agent that knows all the intricacies of whatever application it's responsible for. So I think that's where we are trying to go. I think that's where the industry is also going towards. So you'll see more and more, you'll hear about these world models and whatnot. That's also sort of philosophically aligned with what I just said.
51:43But what I can confidently tell you is these things are a function of their training data. And good training data, we exhaust textual space. So most of the best data will probably come from environments that are created so that these AI agents can just go and experience themselves and learn from those mistakes. yeah yeah and that's fascinating i mean i've talked to uh a number of people about world models and and direct uh uh learning from uh reality as opposed from text uh so it's exciting to hear that that research is going on are are you developing your own models in in uh for that or do you use, you know, models from one of the foundation labs and then fine-kill it?
52:49Yes. So we have always used a foundational model both in our research and in our shipping functions
53:01and because they are just marvelous and good enough on their own. And also, training a foundational model is pretty expensive and slow operation. So we always started from foundation models, but we have a pretty advanced workflow now, I would say, that allows us to try out new ones all the time. So I have a radar that sort of, you know, we constantly try, you know, different, you know, new foundation models from different vendors, benchmark them, train them, and to evaluate them. These days, we are really fond of the Quinn family from Alibaba. They do great, great frontier work. And also Gamma 4 from Google is other.
54:01So we have two editions of the same model. So we have a coin-based model and then a Google-based one. And we offer them both. And the foundation model, the way we are architected is swappable. So as these things develop and advance, we are not really married to them very tightly. so we can easily swap them out in and out.
54:35It's interesting, Craig, that it's a very empirical space. So you can't just look at the architecture. You can't read the paper and say, oh, that's a better model. You really, at the end of the day, you have to sit down and test these, try these. There's all kinds of reward hacking models do when you're trying to test. I mean, sometimes they detect that they're, you know, oh, it looks like I'm being tested right now. So it takes a different rigor to formulate an opinion about a foundational model. But we are all learning that. You know, I think everyone in the industry is learning some techniques around being able to evaluate this and essentially right size them.
55:17Because you'll see that one model comes in a whole bunch of different sizes and resolutions and quantizations and what have you. And so, you know, we did have to tool ourselves so that we can make these judgments, you know, in a data-driven way. But like I said, we offer two foundational models today. So this is the future of IT service management. And it sounds like it's going to get increasingly automated. What's that going to do to IT staff? Does it make them change what they need, the skills that they need, or are they going to become agentic managers or managers of agent systems as opposed to having their fingers in the code?
56:18So the good news is I think what's going to happen is the quality will go up because what these systems, these agents really tackle, target is repetitious things that shouldn't even have happened for the first place. You know, there's a lot of, you know, we expend a lot of calories just to fix the same problem again and again or things that could have been prevented to begin with. And when we work with our customers or our own IT personnel, you know, who are, you can imagine they're quickly adapting all these technologies also. They actually, so right now my reading is they're quite happy to have all this assistance.
57:03They say, well, you know, this saves me a lot of time. I can go home at a predictable time now. I'm not swamped of all this garbage firefights. I can now concentrate on actually making the organization better or making my such and such projects better. So this sort of technology buys you additional digital capacity, as Ryan calls it, to then invest in other new things. I think the smart growing organizations will take this digital capacity and then redeploy it in solving the customers' or end users' problems better or offer them more services and more digital goods. And what does this, what means for this individual is people who have agency is now able to move a lot faster.
58:06So we see this, you know, a certain cohort in every organization sort of really shining, you know, with the help of these technologies. So I think some people will move up in the chain. Some people will be disillusioned and will say, oh, you know what, I'm done, I'm retiring. It's certainly a big economic labor shift. You know, it's happening in front of our eyes. But my, like I said, my personal opinion of this is this is great for the society. I think everything will get higher quality and will become cheaper because of this. and for the people who are involved, I think they'll just tackle higher order, higher value adding problems as these technologies get bigger and become more ubiquitous as they will.
59:02Yeah. Okay.
From the publisher
Most enterprise IT teams spend the majority of their time fighting the same fires repeatedly. BMC Helix is building the AI system that handles those fires automatically, detecting anomalies, tracing root cause through millions of asset relationships, generating remediation plans, and learning from every incident it resolves.
Craig Smith sits down with Erhan Giral, VP of AI Strategy and Innovation at BMC Helix, and Ryan Manning, Chief Product Officer at BMC Helix, to explain how agentic AI is transforming IT service management from a reactive, human-driven process into something closer to a self-healing system, and why doing that at enterprise scale requires a fundamentally different architecture than most AI deployments attempt.
The most technically interesting part of this conversation is where BMC Helix is headed: building "gyms", synthetic data center environments where AI agents deliberately break things and learn to fix them overnight, 24 hours a day, generating the bespoke operational training data that text-based foundation models can no longer provide.
Erhan describes an architecture of specialized sub-agents, anomaly detection, log analysis, root cause analysis, remediation planning, that work in a hierarchy, passing hypotheses between each other until they converge on an answer, fine-tuned to reason the way a specific enterprise's best IT engineer would rather than the way a generic documentation page reads. For customers, the results are measurable: 25 to 50% cost reduction, fewer recurring outages, and IT staff who can finally go home at a predictable time rather than spending their nights firefighting problems that could have been prevented.
Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.




