In short
The episode discusses Virtana CEO Paul Appleby’s “AI factory reality check,” arguing that enterprises are investing in AI data centers faster than they can govern and operate them. Key claim: 6 in 10 enterprises can’t automatically identify the root cause across AI infrastructure domains when AI workloads fail, leading to higher failure rates, poor GPU utilization (throttling/idle GPUs), and inefficient, higher-energy operations. Appleby says root cause is harder in AI factories because failures can originate anywhere across data pipelines, orchestration, and compute/network/storage layers, and many companies stitch together only partial telemetry.
Guest
Paul Appleby, CEO of Virtana (observability for business resilience/operational efficiency). Background: previously at South Horse and Elasticsearch (president); joined Virtana after it was formerly “Virtual Instruments,” led by former Microsoft chairman John Thompson.
Notable examples
trading platforms/internet banking, airline scheduling and baggage handling, retailer supply chain/back office, healthcare remote diagnostics/records collaboration; cybersecurity anomaly detection as a secondary use case.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOChallenges in AI Workload Management
0:00 to 0:34
Learn about the difficulty enterprises face in identifying root causes of AI workload failures.
“Six in ten enterprises cannot automatically identify the root cross across AI infrastructure domains when AI workloads fail.”
Introduction to Paul Appleby
0:34 to 0:47
Meet Paul Appleby, CEO of Virtana, and learn about his background.
“So let's start by having you introduce yourself to listeners.”
Overview of Virtana and Its Purpose
0:47 to 1:49
Discover how Virtana aims to protect technology and infrastructure for businesses.
“Yeah, I'm Paul Appleby, the CEO of Vertana.”
Paul Appleby's Background and Role
1:49 to 3:10
Explore Paul Apple's journey in technology and his role at Virtana.
“And how long have you been with Virtana and what was your background before coming?”
What is Observability?
3:10 to 5:24
Understand the concept of observability and its importance in AI operations.
“In Virtana, we were talking before we started recording, is an observability platform for large-scale operations or technology operations, not necessarily AI.”
AI Factory Growth and Complexity
5:24 to 7:59
Learn about the rapid growth of AI factories and the complexities they introduce.
“And before we get to the study, you were saying that this is a layer that sits above the technology stock, even above the orchestration layer.”
AI Workloads and Governance Challenges
7:59 to 11:35
Discuss the governance issues facing AI workloads and the need for better controls.
“So supporting huge remote and hybrid workforces is key as well.”
Root Cause Discovery Challenge
11:35 to 11:51
Discover why 6 in 10 enterprises struggle to identify root causes in AI systems.
Complexities of AI Workload Failures
11:51 to 14:00
Delve into the complexities that make AI workload failures difficult to resolve.
“identify the root cause across AI infrastructure domains when AI workloads fail.”
Efficiency in AI Data Centers
14:00 to 14:46
Learn how efficient operation can reduce AI data centers' environmental impact.
“and environmental implications of AI data centers.”
Show all 23 chapters
Challenges in AI Data Management
14:46 to 16:40
Explore the difficulties companies face in managing AI data effectively.
“What do people do without a platform like Vertana?”
Vertana's Platform Overview
16:40 to 19:39
Understand how Vertana's platform captures telemetry for AI infrastructure.
“because if you've got a full system view, you could then use agents to accelerate getting to causality and then actually automate remediation.”
Anomaly Detection and Cybersecurity
19:39 to 21:49
Discuss how Vertana's technology aids in detecting anomalies for cybersecurity.
“And why it's so powerful now in this AI era, because we can bring those core capabilities into these so-called AI factories and play a really important part in helping govern them.”
The Role of IT in AI Infrastructure
21:49 to 23:43
Examine the responsibilities of IT departments in managing AI infrastructure.
“So that has certainly happened, but it's not really the core of our business today, although it is a very powerful engine.”
Executive Disconnect in AI Decisions
23:43 to 27:41
Analyze the disconnect between IT operators and business leaders in AI decision-making.
“I mean, this is not another alarm dashboard because most of the remediation is being done autonomously.”
Cost Management in AI Deployments
27:41 to 28:13
Learn how to handle GPU spending as AI transitions from pilot to production.
“How should CEOs and CIOs think about GPU spend differently as AI moves from pilot to production?”
Understanding AI Cost Dynamics
28:13 to 30:56
Explore how AI token usage impacts costs and the implications for ROI.
“And, you know, as we were saying, you know, just as we were chatting beforehand, things are changing on almost a daily basis.”
Governance Challenges in AI Investments
30:56 to 34:20
Discuss the necessity of governance in AI amidst rising investments and complexities.
“And so we're entering a phase, what you're saying, we're entering a phase where AI, ROI will be judged or AI will be judged less by model performance and more by operational efficiency, ergo cost.”
Hybrid AI Infrastructure Management
34:20 to 38:21
Learn about managing complex hybrid AI infrastructures for enterprises.
“And I think that what's happening, Craig, is the kind of gold rush to be in the AI investment cycle is overriding some of the good governance principles that companies have.”
The SaaS Model and Customer Relationships
38:21 to 40:56
Examine how a subscription model fosters ongoing customer engagement.
“Yeah, but the platform resides with the enterprise, or are there instances of it in every place that you have data flowing?”
Future of AI and Data Center Infrastructure
40:56 to 42:00
Speculate on the future of AI infrastructure and the potential overbuilding of data centers.
“And they're not growing year on year for any other reason than that they're growing and scaling their usage of us across their infrastructure because we're solving big, important problems for them.”
The Future of AI and IoT Integration
42:00 to 43:54
Explore the potential impacts of AI on IoT and manufacturing processes.
“How long do you think we've got to answer that one, Craig, in one minute?”
Wrapping Up the AI Discussion
43:54 to 44:26
Reflect on the key points about AI, its myths, and associated risks.
“Is there anything I didn't cover that you want listeners to hear?”
Transcript
Automatic transcript. May contain errors.0:00Six in ten enterprises cannot automatically identify the root cross across AI infrastructure domains when AI workloads fail. Why a root cause is so much harder in an AI factory than in traditional enterprise IT to discover? What you're saying, we're entering a phase where AI ROI will be judged, or AI will be judged less by model performance and more by operational efficiency, ergo cost. And business impact. You know, it might not just be a cost dynamic. The kind of gold rush to be in the AI investment cycle is overriding some of the good governance principles that companies have. We'll see the equilibrium come back.
0:46So let's start by having you introduce yourself to listeners. Thanks, Craig. Yeah, I'm Paul Appleby, the CEO of Vertana. And we're a company that exists really for one really important purpose, where, you know, companies across many industries, whether it's banking, telecommunications, healthcare, retail, airlines, whatever, are so reliant on their technology. In fact, you know, a lot of chief risk officers say that the single biggest point of catastrophic risk of failure for a business now is their technology and infrastructure. Votana is in the world of trying to protect that. We live in this world called observability, and that's all about business resilience and operational efficiency.
1:36How do we make sure those services that are critical to your business and your customers stay performant and available? and even more relevant in this era of accelerated AI adoption, of course. Yeah. And how long have you been with Virtana and what was your background before coming? I've been here now for a couple of years. Inherited an amazing business that in the past, in fact, under its prior name of virtual instruments, was run by John Thompson, the former chairman of Microsoft. So the company's been around for some time and has an amazing history with its core technology. But I came to the company two years ago with a charter from the board to really lean into this world of, you know, the broad digitization of services and the scale out adoption of AI.
2:34Prior to that, I've been working in technology for years, sometimes in startups because I love that whole idea of the scale up. But sometimes in much larger companies, through phases of transformation and growth, like South Horse, where I was for a number of years, and also companies like Elasticsearch more recently as their president. So I've got a long background in enterprise software and enterprise technology and growth and scale of businesses. Yeah. In Virtana, we were talking before we started recording, is an observability platform for large-scale operations or technology operations, not necessarily AI.
3:29I mean, it existed before the generative AI boom. Is that right? Yeah, I think it's a couple of really important things to say. I mean, this class of software, Craig's been around for a long time. For as long as technology's been supporting critical business services, there's been a need to monitor that infrastructure to make sure it stays available and performant. But interestingly, if you think about observability, what really is its purpose? Its purpose is to identify threats and risks in real time. um you know identify what the cause of those was and remedy those so as a consequence we've been building deep ai and ml capabilities yeah for over a decade so ai is not new to us it's really part of our core reason for being a you know streaming in all of that event data in real time and doing you know the the mapping and correlation needed to identify you know risks and threats um what What we've done more recently, of course, is lean in heavily to a lot of agentic capabilities to support automating a lot of these IT operations.
4:45But the other thing that we've done is recognize that as companies scale out the true industrialization of AI with massive AI data center investments that the market loves to call AI factories, is build full end-to-end observability for the AI factory. So what we've essentially done is taken an incredible legacy of observability and fully embrace this whole AI era to support companies as they scale up their AI investments. Yeah. And before we get to the study, you were saying that this is a layer that sits above the technology stock, even above the orchestration layer. Correct, Craig. And, you know, what I'd say is if you think about any critical business service, let's call it a trading platform or internet banking is, you know, something that, you know, is core to any retail bank.
5:50You know, we all do most of our banking online these days. In fact, I do most of mine on my mobile phone. And, you know, what supports our ability to interact with our bank is a really complex system, if you like. um, uh, of, of, you know, data and applications and networks and compute and storage infrastructure, some of it potentially in the cloud, some of it in legacy infrastructure, some on-prem and really what observability needs to be is something that views the whole system. Cause I like to think about it as a system that's incredibly, you know, in most cases, heterogeneous in most cases um you know hybrid um which is the the reality for most companies today um and and the ability to you know dynamically um observe that whole system and and manage that system is is where our class of software sits and um the um it it's becoming increasingly critical this observability as these systems scale right and and they're scaling not only to keep up with i mean some are scaling to keep up with customer uh customer base but some are scaling simply because the business is growing and becoming increasingly complex i mean you're not only looking at customer-facing or public-facing applications.
7:27You're looking at any system that has AI in the stack. You're absolutely correct. So you think of any system, even for large retailers, some of their most important systems aren't just the customer-facing systems, their website or their point-of-sale environment, but it's their whole distribution supply chain infrastructure as well all of the back office functions of a retailer um the same with an airline you know obviously we experience the airline through baggage handling and ticketing and and you know all of those things but there's all of the crew scheduling and flight scheduling and everything else that exists and you know the thing that you know obviously makes that even more complex is this world of remote and hybrid work.
8:17So supporting huge remote and hybrid workforces is key as well. So when we look at a large company, we're thinking about all of these critical systems that are key to the operation of the business. And, you know, you think about, you know, healthcare companies, you know, the way we experience healthcare today, you know, with things like digital medicine and remote diagnostics and remote visits and all the rest of it. It's wonderful as a consumer to be able to do that. But the medical practitioners are consuming all of that as well as they're able to access your electronic health records and collaborate with other practitioners.
9:00So it's being able to support all of that infrastructure for a big healthcare org or a big airline or a big bank. That's where we sit. Right. And the study that you've done, the AI factory reality check, I believe it's called, you're finding that these AI factories, to use that term, but these systems are growing faster than there is ability to operate them. or maintain them? Yeah, I think that's a great way of putting it, Craig. I think that, you know, I think we'd all agree that there's been a wild explosion of investment in AI infrastructure and AI data centers and, you know, even above and beyond, you know, these AI platforms that we're all very familiar with.
10:04As we think more about large enterprises and governments adopting AI, the clear narrative is that most of this AI infrastructure or a lot of it is going to be built on-prem to support these huge AI use cases for these companies. And what our study has found is that whilst companies are leaping into investing in AI data centers and, as you said, AI factories, which is a phrase that's bandied about a lot by the likes of NVIDIA and Dell, these AI factory investments, we're not seeing the governance that needs to come with the kind of controls. to optimize the resilience and efficiency of those data centers.
11:02You know, the investments in governance and controls aren't running at the same pace as the investments in the infrastructure, which increase risk. And that's what we're calling out with this study. It's like, hey, checkpoint here, you know, huge opportunities in so many different industries for adopting AI at scale. make sure the governance and controls that you've got built into all of your, you know, legacy infrastructure get replicated in, you know, this AI data center world. Yeah. And I think one of the top findings is that nearly well over half, I think six in 10 enterprises cannot automatically identify the root cause across AI infrastructure domains when AI workloads fail.
12:01Can you talk about that, why root cause is so much harder in an AI factory than in traditional enterprise IT to discover? Craig, what I'd say is that root cause is hard to discover either way, but it's much harder in an AI factory or a massive scale AI data center because of the complexity of that system, if you like, from the data pipelines to the AI workloads and AI orchestration into, you know, the infrastructure layer of computing and network and storage. storage, et cetera, if there is either a slowdown or a failure of an AI workload, the challenge is where did that occur? And it's really quite fascinating is that the vast bulk of these failures don't have an identified brute cause.
13:10Now, the problem with that, of course, is if you don't know what causality is, you get to remediation. And that leads to, you know, a really, really high percentage of AI workloads actually failing. Now, that doesn't make sense. If you're going to operate at scale with critical services, you can't have, you know, a really high percentage of AI workloads failing. And that's just one element of it. The second element of it is, you know, incredibly powerful compute infrastructure and capability with these GPUs. But if AI workloads are getting throttled and GPUs are sitting idle, you've made these massive investments and you're not getting a return on the investment.
13:53So how do you make sure you're maximizing utilization and throughput to get great ROI? And then the final thing is we're all aware of the massive energy consumption and environmental implications of AI data centers. So if we're actually operating them efficiently, we can actually significantly reduce the, you know, the energy requirements and ultimately the footprint and impact that AI data centers have. So it's really at multiple levels that we need to be thinking about this whole system. And unfortunately, right now, Craig, what, you know, what exists for many companies is, you know, bits of telemetry out of each level of the stack that they're trying to stitch together to identify, hey, what's really going on and how do we actually better manage this environment?
14:43And that's not the answer. Yeah. Well, as a matter of fact, that was going to be my question. What do people do without a platform like Vertana? Because this isn't a new problem. It may be increasing in its the risk as things scale, but what have people been doing or what are they doing as they scale? They can't just be holding their breath and hoping everything's okay. Well, they're not just holding their breath. There's clearly data coming from certain elements of that AI factory stack. And what it's left to is that the IT professionals whose job is to run and operate these data centers to try and stitch that data together.
15:32That's why when you have a service failure or a service disruption, you end up with a group of people, historically it would have been physically in a knock, but now it's a combination of people in the knock and a bunch of people on a call like this, trying to work out what the hell is the problem and how do we fix it? And it's trying to stitch together all of this data. Now, that's the reality for most companies today. In fact, you know, they've got four or five, sometimes six different sources of data that they're somehow trying to weave together. And I think the clear message moving forward as we think about critical services at scale in this AI world is that you need visibility into the full system.
16:26You have to be capturing rich telemetry and doing that dynamic mapping and correlation in real time to be able to really truly govern these systems. And as we were chatting about earlier before you know, this conversation started, it's then you can actually apply agentic capabilities, because if you've got a full system view, you could then use agents to accelerate getting to causality and then actually automate remediation. How does Vertana's platform work? I mean, do you have multiple models that are cross-referencing each other? Do you have sensors in different servers that are watching things?
17:17And then all of that is fused and analyzed. So yeah, can you just talk about the architecture of it all? Yeah. The short answer is yes, yes, and yes. But let me talk a little bit more about what I mean there, Craig, because to be able to do the kinds of things we're talking about of getting to true autonomous operations, you need to capture telemetry from the whole stack. So over a period of many years, we've developed deep integrations into every layer of the stack in both traditional infrastructure and now, of course, in this AI infrastructure. And we're capturing 20 ,000 different metrics in a sub-second time frame and correlating and mapping those metrics in near real time.
18:21So it's by capturing all of those metrics and correlating and mapping those and using our models to see patterns that identify causality. And of course, the patterns that recommend remediation are really core to what Vertana does. But it really is a really important thing to understand. And one of the things I understood by having worked in other organizations in this observability sector is that you can't really apply agentic capabilities and expect to drive true autonomous outcomes unless you've got system-wide visibility and that kind of real-time dynamic mapping and correlation because it just doesn't work otherwise.
19:10You've got some automation to a part of the stack as opposed to the full system. And that's really what Vatana does. And it is a really, really tough problem to solve, Craig, that ability to stream and ingest that volume of metrics and then analyze, correlate, understand, and recommend dynamically is really core to what the Tana story is all about. And why it's so powerful now in this AI era, because we can bring those core capabilities into these so-called AI factories and play a really important part in helping govern them. Yeah. And we also talked briefly about cybersecurity, whether this enhances your cyber attack monitoring, because this is LLM adjacent, right?
20:14You use models to do some of the analysis. are most of the anomalies detected things that you've seen over and over again? And so the system has a pretty good bead on why that's happening, or does it find new things that no one's seen? And in those cases, is it capable of addressing them autonomously? Yeah, I mean, a lot to unpack there, Craig. So let me take them piece by piece. You know, on the question of cyber, you know, we don't position ourselves as a cybersecurity solution. We're primarily focused on the infrastructure. But the point that we were talking about earlier, of course, is that if you've got this ability to ingest all of these events and analyze them, correlate them and identify anomalies, and it's such a powerful engine, it can be applied for cybersecurity use cases.
21:34So in fact, although we don't position ourselves as a cyber security company, we've got a bunch of customers who've gone, wow, this data is so rich. We can actually augment what we're doing in terms of our cyber capabilities by using the ability to detect anomalies. So that has certainly happened, but it's not really the core of our business today, although it is a very powerful engine. As it relates to your other questions, yes, we do see because we have so much rich data and because a lot of these events follow similar patterns, we can identify those patterns and get to causality really quickly.
22:19Um, but the other fascinating thing is because we're capturing so much rich data, even when we're identifying, you know, new or newer anomalies, we're very, very quickly able to learn what those patterns are and then start recommending remediation. remediation. Even if it's, you know, and depending on a company's policies, you know, we can either automate that remediation or, you know, pass the, you know, the recommendation together with the supporting evidence around the recommendation to a human operator to then act on as well. So, you know, really depends on the, you know, the company, where the company's at on their AI and automation journey and, you know, you know, what policies they want to adopt associated with that.
23:09But what we found in going into a number of these AI labs of some of the larger, you know, technology providers and some of the larger customers is that we're identifying anomalies and usage patterns that, that they didn't even understand about their technology themselves because of our ability to do exactly what we've been talking about earlier. Yeah. Well, I have so many questions. So who owns this? I mean, this is not another alarm dashboard because most of the remediation is being done autonomously. But as you said, there are times when you have to pull a human into the loop. Who is that? Is that in the IT department?
24:02And then the other question is, your report shows a disconnect between executives and practitioners. Yeah, what does that say about how AI infrastructure decisions are being made or the purchase of Vertana, for example? Maybe the guys in the IT department are saying, we can't handle all of this. we need vertana but the but the c-suite is skeptical i mean yeah can you talk about that craig craig you're the master at asking multi-layered questions so i'm gonna i'm gonna try i'm gonna try and um and and pass them all there um you know what i'd say i'll tell you a story i can't name the company for obvious reasons but i visited with a very large company here in the US recently.
24:57And I met with the person tasked with this. And in this case, it was the senior vice president responsible for IT operations. So that kind of senior vice president of IT operations is really a critical role in companies. And the exec that I was chatting with had been with the company for many years, a couple of decades, and said when they first joined a company and were in that role, they reported out to the CEO and the C-suite about what they're up to and creating rigidity and resilience once a year. And then it became once a quarter. And then it became once a month. And what they shared with me is that they report out once a week on where they're at around providing resilience, rigidity, and operational efficiency to the CEO and the CEO's leadership team.
25:53And what that says, and the reason why I wanted to share that anecdote is that this role of the person who's responsible for ensuring the resilience and efficiency of IT infrastructure and the services it supports has now become a C-level, virtually a board level issue. So it's that persona that we sell to. But beneath them, it's obviously all of the IT operators, the specific role is called site reliability engineer, and their job is to ensure the reliability of that infrastructure. It's that group that we sell to. And I think what I'd say to the other layer of your question is that the disconnect between the IT operators and the business, if you look at any business today, every board is saying, what's your AI strategy and how are we embracing AI to stay relevant and compete and all of those things?
26:54So this huge focus on getting in and investing in AI and being an AI forward company. So that pressure is coming from the business. the the IT operators you know these VPs and senior vice presidents of IT ops are sitting there going but hey we need the governance and controls to make sure that you know we can you know protect the the resilience of the the business so that's kind of the disconnect the pressure coming on investment and and investment at speed and then that kind of you know IT and the professional IT executive turning around and going, hey, we're increasing our risk profile here, and what are we going to do about that?
27:39Yeah. The token prices are falling. How should CEOs and CIOs think about GPU spend differently as AI moves from pilot to production? And I think I've mentioned to you, I've been seeing people warning that the drop in token costs is great, but it tends to then people deploy all these agents. The next thing, your cost is going through the roof, even though the token cost is lower because you're using more tokens. Spot on, Craig. And, you know, as we were saying, you know, just as we were chatting beforehand, things are changing on almost a daily basis. So, yeah, sure. Token is going down and will continue to fall, you know, for the foreseeable.
28:41The token consumption is going through the roof. and you know it's really just even looking at at labor arbitrage and that question of you know what is you know where where is the work done most efficiently there's been conversations over the the last few days about you know is the labor arbitrage conversation headed in the right direction and and in fact will ai be more efficient and i think that is going to lead to some really important decisions that companies have to make about what use cases and business problems um we need to solve leveraging this kind of infrastructure um that are going to make a huge impact and you know there i can think of some great examples you know like accelerated drug discovery, reducing the time of bringing new medicines to market, making them far more highly targeted and improving the efficacy of those is like, there's no argument about that.
29:48You know, using AI to, you know, better analyze, you know, medical imaging, using AI to better identify, you know, fraud, you know, criminal financial transactions and those sorts of things. You know, I could go on and on of really powerful use cases for AI. But I think what companies are going to do, and I think they are doing right now, is looking at where do we deploy AI that's going to have a meaningful impact on customers, customer experience, and drive increased ROI. That's going to be key. But I think that a lot of companies are rushing headlong into the investment right now without thinking through the implications for ROI.
30:39And there's a lot of studies being done on this, the hundreds of billions of dollars that are being invested right now. And when the curve between the investment and the revenues will actually cross, and based on the data I've seen, it is not in the near future, Craig. Yeah. And so we're entering a phase, what you're saying, we're entering a phase where AI, ROI will be judged or AI will be judged less by model performance and more by operational efficiency, ergo cost. Yeah. And business impact. It might not just be a cost dynamic. It could be a revenue dynamic or it could be a customer experience and retention dynamic, but there'll have to be measurable impacts that's commensurate with the level of investment that we're making.
31:34And as I've said, there are so many use cases where AI is just the only answer, but it's about being really smart about how and where we deploy AI, building the right governance and controls around those investments, and continuing to scale thoughtfully. yeah uh uh and does vertana help in tracking roi metrics because the report says i think nearly a third say they need clear metrics roi metrics before confidently scaling ai further so yeah so what what metrics do you think matter most and and how does vertana handle that Yeah, Craig, great question. What I'd say is that if you think about any large scale IT organization, they're in a pretty tricky position, not just as it relates to AI investments, but just IT investments in general, because they're being asked to do a couple of things.
32:51support more and more complex services, AI being one of those things with more and more complex infrastructure with incredible resilience. But they're also being asked to do that efficiently. So, you know, I haven't met a CIO yet that hasn't told me that they've got pressure on cost and operating costs. So it's like there's this duality that's kind of, it's competing, you know, priorities. And so if you're going to play a role in supporting enterprise resilience, you've got to play a role in making sure that you help customers use their infrastructure efficiently. So that whole area of understanding and managing control for a cost for both on-prem or physical infrastructure as well as cloud infrastructure is a core part of our solution.
33:45The only thing that we don't include is the business metrics because we're not bankers. We don't run insurance businesses or whatever, but we can provide all of the cost analysis and metrics that then get combined with the business metrics to determine return on investment and payback and the like. Yeah. The study found, though, that a lot of this governance work is being deferred just as complexity is increasing. So why is that and how do you see that impacting enterprises? Yeah, it's interesting. And I think that what's happening, Craig, is the kind of gold rush to be in the AI investment cycle is overriding some of the good governance principles that companies have.
34:48I think that we'll see the equilibrium come back. And in fact, that's why we published the study to show that there is a delta between what the business executives think and what IT operators think. there is a huge failure rate for AI workloads. And that, you know, we didn't publish the findings to be naysayers around AI, because we're not. But what we are saying is, hey, as we make these investments, we got to catch up with the governance. I think we'll see that happen quickly, Craig. You know, especially as people see the data and are confronted with the data, and the reality of putting governance and controls in place, particularly as it relates to AI, is not only in the public interest, but it's in the business interest as well.
35:40You've just extended AI factory observability to Dell AI factory and AWS Bedrock and Nutanix. uh what does that say about where observability needs to live on-prem at the model layer yeah or yeah um you know i i think um there there again a couple of elements to that craig the first thing that i'd say is that what we're seeing in terms of the you know ai data centers that are getting deployed is that they are leveraging as companies architect these solutions, they're leveraging what they believe is the best solution for every layer of that stack. So you end up with this heterogeneous environment of different providers providing components, if you like, of the AI factor that come together.
36:43So deepening our relationships with the likes of AWS and with Dell and with Nutanix and many, many others, NVIDIA, your deep integration into the NVIDIA stack and their GPUs, of course. It's all about ensuring that we're capturing the richest telemetry. So it is really, I think, the first thing to note is that heterogeneous nature of these AI data centers pulling on the best technology for each layer of the stack And then for observability companies, it's our job to capture all of that telemetry and weave that together for our customers. I think as it relates to, you know, on-prem or in the cloud, I think, you know, like we're seeing with, you know, a lot of enterprise technology, we're going to see a hybrid world.
37:38you know the the the world of hybrid is the reality for just about every large company on the planet certain workloads and data will remain on-prem and and that's not going to change certain workloads and data will be in the cloud and and it will be a kind of a dynamic you know model where where we'll see you know sometimes you know workloads and data shift depending on some kind of price arbitrage or, you know, speed to outcome or, you know, data protection and privacy issues. So what's the big responsibility for organizations like ours is to make sure that we can, you know, manage that complex hybrid world and take that complexity away for our customers.
Read the full transcript
38:26And, you know, hence the investments in, you know, clearly significant on-prem infrastructure capability, as well as partnering with the hyperscalers and neoclouds to ensure that we're supporting customers wherever they need to be. Yeah, but the platform resides with the enterprise, or are there instances of it in every place that you have data flowing? You know, the one thing that we've also tried to do is ensure that our technology can be deployed anywhere. So our technologies are deployed both on-prem and in the cloud. So there are instances on-prem, there are instances in the cloud. And in fact, with a number of providers who are providing, you know, as a service infrastructure, like Hitachi, our platform is embedded in their platform.
39:28So, you know, the interesting thing is Vatana shows up in all sorts of different places, you know, really depending on, you know, how you're consuming the service. And the whole idea and the whole philosophy around that was, you know, providing ubiquitous access and not going, hey, we are purely going to reside in the cloud. And that is our vision and strategy because that's not where our customers are. Our customers leveraging the public clouds, they're leveraging service providers, managed service providers and consuming services there, and they've got their own on-prem infrastructure that they're managing themselves.
40:12We need to be available across all of those environments. And so we've made a choice to do that. And it's ended up being the right choice because it allows us to be wherever our customers need us to be. Yeah. And this is a SaaS offering. I mean, it's not a single purchase. And then they're, yeah, no, where, you know, where, where, you know, we, we, the great thing about being a, you know, a subscription software company, Craig, is that, that you have to continually earn the right to do business with your customers. And that's what I've tried to encourage my team is that that philosophy is that we have to continue to earn the right.
40:55And, you know, the amazing thing for us and that I'm incredibly grateful to all of our customers for is that almost every single one of our customers grows year on year. And they're not growing year on year for any other reason than that they're growing and scaling their usage of us across their infrastructure because we're solving big, important problems for them. I want to ask kind of an off-the-topic question, if I may. You've got me worried now, Craig. No, no. With the economics of deploying AI changing, and there's a lot of work on pushing inference out to the edge, and there's a lot of work on making it more power efficient and cheaper.
41:50Are we overbuilding data centers, do you think? I mean, that would argue to me that you can do more with less.
42:03Wow, that's a big question. How long do you think we've got to answer that one, Craig, in one minute? You know, I think that all of those things are true. if we think about, you know, the promises of IoT, I mean, we stopped talking about the internet of things. But if you talk about, you know, city scale automation, you know, large scale adoption of autonomous vehicles, you know, massive scale manufacturing, we talked about physical AI earlier and the role of AI and manufacturing processes. You know, a lot of this has to occur out on the edge um there is no doubt about that um but you know what i would say is that that that's going to be additive i think that the the physical build out of ai infrastructure is going to be absolutely key particularly as companies and government adopt ai they're going to be building out more and more their own ai data centers ai factories and and you only have to listen to Michael Dell talk about the explosive growth that Dell is seeing in the adoption of their AI factories to realize that companies are doing that.
43:23So I think that'll happen in parallel, Craig, and maybe we'll see ultimately the promise of large scale autonomous vehicle adoption. We'll So the advent of truly distributed power generation and smart utilization of energy in the home and all of the things that have been promised for a long time, I think, could potentially be deployed with AI out at the edge. But that's, I'm sure, a much longer conversation. Okay. Okay. Is there anything I didn't cover that you want listeners to hear? Um, no, I think Craig, we've, we've, we've touched, we've touched on all of it. It's been a great conversation and it's fun chatting about what's going on out there in the world and kind of, you know, debunking some of the myths around AI, but also raising consciousness and awareness around the risks.
44:23And I think it was a really great conversation. Yeah.
From the publisher
Companies are spending billions building AI factories, but most of them can't tell you why their AI workloads are failing, whether their GPUs are actually being used, or what their infrastructure is going to cost them when agents start running at scale. Paul Appleby, CEO of Virtana, joins Craig Smith to discuss the findings of their AI Factory Reality Check study, a research report that reveals a striking and underappreciated gap between the pace of AI infrastructure investment and the governance needed to run it safely and efficiently. Six in ten enterprises, the study found, cannot automatically identify root cause when an AI workload fails, a problem that compounds fast once you're running critical services on AI infrastructure at scale.
The conversation covers the mechanics of Virtana's observability platform, capturing 20,000 metrics per second across the entire AI stack, correlating them in real time, and increasingly using agentic capabilities to remediate failures automatically, but its most important insights are structural. Appleby makes a sharp observation that cuts through a lot of AI optimism: token costs are falling, but token consumption is exploding, meaning the total cost of running agentic AI systems is still going up even as the per-unit price drops. He also tracks a cultural shift inside enterprises - IT resilience reporting that used to happen annually now happens weekly - as evidence that technology risk has become a board-level conversation in a way it simply wasn't before. The result is a conversation that's less about the promise of AI and more about what it actually takes to make it work at production scale.
Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.




