In short
NVIDIA AI Podcast - Episode 293 Summary: Building AI Factories
Episode Overview In this episode of the NVIDIA AI Podcast, host Noah Kravitz is joined by Chris Wright, CTO of Red Hat, and Justin Boitano, VP and GM of Enterprise Computing at NVIDIA. The discussion focuses on the transition of enterprises from AI pilots to full-scale AI factories, exploring the concept of the "five-layer cake" AI factory stack. The conversation delves into the necessary infrastructure, security, and governance to ensure successful implementation.
Key Concepts
What is an AI Factory?
- Definition: An AI factory is a structured approach for enterprises to utilize data and convert it into actionable intelligence.
- Purpose: It aims to enhance productivity by enabling organizations to leverage digital intelligence efficiently.
The Five-Layer Cake AI Factory Stack
- Hardware Infrastructure: Essential data centers with powerful chips and rack-scale infrastructure.
- Software Infrastructure: Tools required to orchestrate operations and manage data efficiently.
- Models: Algorithms and frameworks developed to process data and generate insights.
- Applications and Agents: End-user applications that leverage the generated intelligence.
- Governance: Standards and protocols to ensure responsible and ethical use of AI technologies.
Benefits for Enterprises
- Efficiency: AI factories enable organizations to process data and gain insights faster.
- Productivity Gains: AI can lead to significant improvements in business operations, potentially doubling productivity in some cases.
- Data-Driven Decision Making: With structured data handling, enterprises can create specific use cases that drive business outcomes.
Current Trends in AI
- Adoption Rates: A small percentage (1%) of organizations have optimized AI factories, while over half are still in early transformation stages.
- Investment Growth: Global AI investment is projected to exceed $1 trillion by 2029, with a large portion focused on agentic systems.
- Shift to Agentic Systems: AI agents are evolving to perform more complex tasks autonomously, moving beyond basic functions like chatbots.
Key Challenges and Considerations
- Legacy Systems Integration: Enterprises must find ways to modernize existing systems while integrating new AI capabilities.
- Security and Governance: Ensuring data privacy, compliance, and governance is crucial as AI is deployed more widely.
- Risk of Over-Analysis: Companies should act quickly to implement AI without getting bogged down by exhaustive analysis.
Practical Steps for Implementation
- Initial Setup: Start with a solid hardware foundation, ensuring data centers are equipped to handle AI workloads.
- Utilize Proven Blueprints: Adopt validated designs and blueprints for rapid deployment and initial wins.
- Iterative Development: Focus on small, manageable projects to demonstrate value before scaling.
- Involve Security Teams: Ensure security protocols and data access controls are established early in the process.
- Monitor and Evaluate: Continuously assess the impact of AI implementations on productivity and adjust as necessary.
Future Outlook
- Autonomous Agents: The future of AI factories will heavily feature agents that can operate independently, handling complex tasks with minimal human intervention.
- Integration Across Business Functions: AI will become an integral part of operations across various sectors, transforming traditional processes into AI-driven workflows.
Conclusion The episode emphasizes the importance of building AI factories to harness the full potential of artificial intelligence. The transition requires a thoughtful approach to infrastructure, integration, and governance, ensuring that enterprises can scale effectively while maintaining security and compliance.
Resources
- Learn more about the AI Factory by visiting the [Red Hat AI Factory with NVIDIA](https://www.redhat.com) and [NVIDIA's AI Factory](https://www.nvidia.com) websites.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding AI Factories
0:45 to 3:02
Exploring what constitutes an AI factory and its importance for enterprises.
“Thank you so much for taking the time to join us.”
Current Trends and Challenges in AI Adoption
3:02 to 5:12
Discussing the rapid changes in AI and how businesses can navigate them.
“Right now, as we record this, there's a lot of talk about OpenClaw and Autonomous Agents and kind of long-running agents.”
The Role of AI Factories in Business Growth
5:12 to 8:02
How AI factories can bridge traditional and modern business practices.
“But at the same time, projections have global AI investment exceeding a trillion dollars total by 2029, just a few years out.”
Advancements in AI Agents and Their Impact
8:02 to 11:30
Examining the evolution of AI agents and their practical applications.
“Well, I got to say, what's interesting is in the last three months, it feels like the market has really started moving even faster.”
Building an AI Factory: Key Considerations
11:30 to 14:05
Essential capabilities and strategies for establishing an AI factory.
“I don't know that it has, you know, a hole in it waiting for a prompt injection attack or whatever the case may be, right?”
Establishing Governance in AI Factories
14:05 to 15:11
Understand the importance of governance and evaluation in AI projects.
“And you can start in this dev environment with, I'll call it, narrow use cases that are aligned to your core business goals and then scale as you start to see success.”
The Role of Inference in AI Factories
15:12 to 18:12
Learn how inference serves as the production environment for AI.
“So training, whether it's pre-training or post-training, those are things that happen pre-production.”
Building Initial AI Factory Infrastructure
18:13 to 20:44
Discover the foundational components for starting an AI factory.
“and can scale up to the future and really help transform companies into AI natives, as we've been talking about.”
Key Strategies for AI Factory Deployment
20:45 to 23:35
Explore strategies for successful AI factory deployment and scaling.
“and if we're world-class at those, then we can be world-class in market.”
Practical Steps for the First 90 Days
23:36 to 28:00
Get practical tips on structuring the first 90 days in an AI factory.
“where you've normalized all your data and everything is well-defined, you'll spend all of your time doing that and you'll never be able to get to showing some business value.”
Show all 14 chapters
Building an AI Factory: The Iterative Process
28:00 to 29:29
Learn about the iterative process of building AI factories and the importance of starting small.
“I don't think that's particularly useful.”
Implementing Guardrails for AI Success
29:30 to 31:08
Understand the essential guardrails necessary for scaling AI implementations in organizations.
“What are the things you can lay down kind of from the beginning to make sure that, you know, technical and process and governments and, you know, the guardrails are in place for these kinds of things?”
The Future of AI Factories and Agentic AI
31:09 to 35:00
Explore predictions about the evolution of AI factories and the increasing sophistication of AI agents.
“Chris, I'm going to turn this one to you first.”
Transforming Work with AI: Productivity Gains
35:01 to 36:51
Discover how the integration of AI in various roles will boost productivity across industries.
“And so it's doing very long-running thinking and work that is the work of many, many, many, many software engineers, I'll just say.”
Transcript
Automatic transcript. May contain errors.0:10Justin Boitano:Welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz. My guests today are Red Hat's Chris Wright and NVIDIA's Justin Boitano, and we're talking AI factories. Why should enterprises build AI factories, and how can they do so with confidence in building AI factories that they can trust? By way of introductions, and I'll keep it brief because both of these guys' work speaks for itself, really. Chris Wright is Chief Technology Officer and Senior Vice President of Global Engineering at Red Hat. And Justin Boitano is Vice President and General Manager of Enterprise Computing at NVIDIA. Gentlemen, welcome to the NVIDIA AI Podcast.
0:46Justin Boitano:Thank you so much for taking the time to join us. Thanks for having us.
0:50Chris Wright:Thanks for having me, Noah.
0:50Justin Boitano:Let's get right into it. And Justin, I'll start with you, but always both of you guys feel free to jump in. you know, as the spirit moves you, so to speak, as we go. But Justin, why don't we start with you? Can you talk a little bit about, well, maybe first give kind of a working definition of what we mean, what you mean when we talk about an AI factory, and then get into kind of at a high level, why would an enterprise be interested? Why are enterprises building AI factories? And what are some of the tangible benefits that an enterprise can expect to see from an AI factory?
1:22Chris Wright:Sure, Noah. Yeah, you know, and I think it's important to understand kind of the context of where we are as an industry. And, you know, building digital intelligence to power the productivity of organizations is going to be as critical in this decade as, you know, energy in running our companies. This is the next industrial revolution and companies are always asking us, you know, how do we build these factories that basically take data in and then produce the intelligence that, you know, helps them run their businesses more efficiently. And so as we talk about, like, what is an AI factory? We think of them as really kind of five layers of technology that need to come together.
2:02Chris Wright:At the base layer, you've got to make sure that you have got the data centers with power to bring into these factories. You've got to have chips is the easy way to talk about it. But we're at this point of building rack-scale infrastructure that's six chips with extreme co-design to build the best token efficiency from the power available to you. The next layer, You typically want to have the software infrastructure to orchestrate everything. And then you want to have models that run that intelligence. And then ultimately, the apps and the agents on top. And so what every business needs to do, though, is take this intelligence and build use case specific business outcomes that help them drive innovation, build products faster, and ultimately grow revenue top line through deploying this intelligence at scale.
2:54Justin Boitano:Right. Right. And so these five layers you're referring to, this is the cake, right? The five-layer cake? That's right. This is a five-layer cake. Excellent. Chris, the world is, I feel like we can say this so often, but things are changing so quickly. Right now, as we record this, there's a lot of talk about OpenClaw and Autonomous Agents and kind of long-running agents. Can you speak kind of a little bit sort of to that and how Nvidia and Red Hat are working together to help enterprise and enterprise IT departments kind of step into this new world? Yeah, actually, OpenClaw is a great example because there's so much enthusiasm about, I guess, what's possible, what you could do.
3:34Justin Boitano:It's captured the kind of the builder's imagination, but also built quite quickly, certainly leveraging AI to help produce code quickly, but not with the enterprise in mind. So when we think about what Justin was describing, that kind of data in to a factory context that produces business value as an output, we're talking about enterprises. That's their data. Those business outcomes are really either driving net new growth or focused on the productivity and efficiency. All of that needs to be done responsibly, safely, respecting access controls, delivering audit trails, things that are maybe not as fun in the builder world, but fundamental to the enterprise world.
4:22Justin Boitano:And so a lot of what we're doing is taking these building blocks, the layers of that five-layer cake, and making them accessible to the enterprise together. So obviously, NVIDIA's got world-class hardware. We're bringing a software layer that enables the higher levels of that cake. And then we're building the right guardrails and security considerations into this combined solution so that our customers can then feel confident about bringing this into their enterprise as they're all trying to figure out how to do AI transformation, go from a traditional company to really an AI native company. and in that context, not introduce undue risk or, you know, essentially undermine the core of their business.
5:11Justin Boitano:Right. So there's research that shows that only 1 % of organizations right now have reached the stage of an optimized, AI-fueled, AI-native, as you were talking about, Chris, enterprise, while over half of organizations still remain in the early stages of transformation. But at the same time, projections have global AI investment exceeding a trillion dollars total by 2029, just a few years out. And of that trillion dollars, these projections are saying agentic systems are going to account for roughly half of that spending. That's a big shift from, you know, a year ago, two years ago. You guys know the timeframe better than I would.
5:50Justin Boitano:But, you know, when agents were kind of this buzzword that, you know, nobody necessarily knew, there are all these different definitions, et cetera, et cetera. And now we're talking about all of this, you know, resource and spending going in specifically to agentic systems. Chris, what can we glean from this? And I know you spoke to it a little bit just now, but what are the kinds of things that the AI factory can do for an enterprise, infrastructure-wise, but confidence-wise, as you were talking about, when it comes specifically to figuring out how to deploy and integrate these agentic systems?
6:24Justin Boitano:Well, if you think about that notion of transforming the enterprise and leveraging internal data and focusing on your core business, how do you improve it or grow it? There's a whole set of things that are underneath that. Obviously, the data piece that we talked about, but also it is the existing tools that operate your business that are not going to just go away. They're fundamental. They're the baseline, the business as usual components, pretty critical and fundamental. So part of this is how do you carry that forward and really modernize your entire infrastructure to bring these two worlds together, this highly modern AI native world and the traditional set of applications that literally run the business?
7:08Justin Boitano:Because you need to bring AI capabilities, not just in the net new, but also in the existing content that runs all the enterprise. And to me, that's exactly what the AI factory does. it helps bridge these two worlds. I mean, in the end, we've got models, but we also have, as Justin described at the beginning, agentic content or AI-enabled applications, and then also the traditional applications. So bringing all of that together and then doing it in a context with consistency across the enterprise so that you're not asking every team to go figure out their own, choose your own adventure path forward.
7:44Justin Boitano:And that consistency, you build best practices across your organization. And then you're ultimately improving your chances for success and reducing the failure rates. There's so many studies that suggest a lot of AI projects can fail. There's a number of reasons for that. One of those is having the right tools and having the best practices, access to the data, and essentially combining forces as a company to produce an output rather than devolving into sort of the next generation of shadow IT and everybody building their own thing and creating this highly fragmented internal environment, which is then kind of difficult to get your arms around if not just produces very little success.
8:32Justin Boitano:Yeah. Justin, are you seeing similar things?
8:35Chris Wright:Well, I got to say, what's interesting is in the last three months, it feels like the market has really started moving even faster. I'll just say this. You look at coding companies.
8:46Justin Boitano:Every one of you guests who comes on this podcast says that same thing, moving faster, moving faster.
8:50Chris Wright:Well, but you can actually really feel it now. And I say that because the first area of agents, like product market fit, really was in software development. And we see it really as a software company ourselves. We can feel these agents doing so much more work for our developers and running longer, more complex software tasks. So you give them design goals and they can work towards those goals. And at the same time, like you said, this moment of Claws came out. And Claws basically take it to a new frontier of full autonomy. And so we're getting to this point where agents are going to have a lot more agency within our enterprise.
9:38Chris Wright:A lot of those studies that you mentioned where people were having a hard time getting AI to work, I think was at a previous era of the world where people were trying to do chatbots, just very basic chatbots. And that was before reasoning, and it was before this level of autonomy that I'm talking about. And so I feel like a lot of what enterprises might have been experimenting with might be a couple generations behind where state of the art is right now. And so as we deploy agents internally now that can use a very, I'll call it, deep agent-like reasoning framework. They can plan and reason and act across many different business systems to do deep research, as an example, to understand the intent of what a user might be asking and help them get to the information across the enterprise in a way that's faster and more efficient than ever previously thought imaginable.
10:31Chris Wright:And the nice thing about running this on a factory, an AI factory within the context of an enterprise, as Chris mentioned, delivers data privacy and security by running that all across open models in this on-prem world. And then you can do things where you still potentially use the frontier models, but you can use the frontier models in a way where you might only use it for the planning stage of the agent, and all the search and summarization is using open models. And so that drives a lot of cost efficiency. In some of our newer blueprints, we see a 30x cost reduction, by doing a hybrid model architecture across your private unstructured information.
11:09Chris Wright:And so that is a use case. The enterprise search, I think, is a broadly generalized use case that gets us from, say, these early adopters that were seeing the benefits of agents for coding into really how knowledge workers are going to start to use agents to help them do their jobs in a much more productive and efficient way.
11:30Justin Boitano:As somebody who sits more on the knowledge worker than a software developer side of the fence myself, getting me more towards that and away from Vibe coding is probably a good idea, but that's just my own sort of personal use case there. But that does make me want to double click a little bit on, you know, on security and governance and things like this, which, you know, I think, Chris, you mentioned at the top with the advent of, I mean, joking aside, with the advent of, you know, coding tools, Vibe coding tools and these more advanced Agenta coding tools in the hands of anybody, including folks like me, it's easy to spin something up.
12:08Justin Boitano:I don't know that it has, you know, a hole in it waiting for a prompt injection attack or whatever the case may be, right? And get into that shadow IT world, Chris, you were talking about. So I want to ask you both, and Justin, I'll start with you because you were talking about it a little bit just now. When you talk about planning and building an AI factory, what are the non-negotiable capabilities that have to be built in, that the enterprise must have to move from, you know, kind of first experiments and prototypes with AI to production, getting into industrial-scale production AI use cases with confidence.
12:42Justin Boitano:And, you know, Justin, you mentioned some of these, but there's security, there's governance, reliability, obviously moving to scale. Can you talk a little bit about some of these factors?
12:52Chris Wright:Yeah, and I think I'll say in the software development world, we're really good at separating the notion of development versus production. And I think that's obviously the best practice as enterprises get going is to separate the two. And on the one hand, you want to help your internal, I'll call it AI development teams, do discovery in a development environment, but separate that access control from production data until you've basically proven the verification or done the functional verification of the outcome that you're trying to get to. you've QA'd it, you've pen-tested it. It's got things like role-based access control so that if a user is using that agent, it inherits their permissions to access business systems.
13:34Chris Wright:And you're going to promote the agent from this development environment into a prod in that way. And so I think the worst thing that enterprises can do is overanalyze this, though, and try and get to, like, how do I prove the TCO up front before I start to make the investment? You've got to believe that AI is this new frontier and the companies that are able to harness it and put it to work for them are going to have a massive competitive advantage. And so the sooner you get going, the better. And you can start in this dev environment with, I'll call it, narrow use cases that are aligned to your core business goals and then scale as you start to see success.
14:16Chris Wright:But to your point, you want to make sure these agents have a clear set of governance. There's clear ability to trace the data systems that they access and that you can continuously evaluate them against known business outcomes that you're trying to achieve. And then that accuracy against certain use cases is what allows you to promote it then into production.
14:42Justin Boitano:Chris, can I ask you how things like, well, inference, obviously we did an episode recently about, it was energy focused, but talking about the coming wave of inference and the shift of the load moving to some extent from training to inference, maybe in this calendar year or whatever kind of the next wave is. But talking about things like high-performance inference and also hybrid cloud agility, how does the AI factory sort of figure in and support these two things in particular? Simply put, inference is your production environment. So training, whether it's pre-training or post-training, those are things that happen pre-production.
15:18Justin Boitano:and inference is where you're bringing this intelligence to life. So scale, efficiency, security, robustness, reliability, compliance with policy, compliance with SLAs or SLOs, these are like the table stakes. And an AI factory is a significant investment for an enterprise. The expectations are it produces significant business outcomes. And so we're focused on optimizing that production of outcomes, which you could back up and say those are business intelligence, or you could back up a little bit more and say it's simply tokens. Optimize that throughput of tokens in the context of cost, in the context of power consumption, because we're also power constrained.
16:06Justin Boitano:And so how do we do that? That's through this scaled out inferencing, which is part of the AI factory. It's really the core underlying platform that you're running all of your models and above that, the agents and AI applications on top of. So to me, it's the critical substrate and the agility that comes with flexibility and choice of where and how you deploy your models or your workloads, that notion of pre-production environments and production environments and where production data versus non-production data is used, you get some choice in where you deploy. And that, to me, is really the hybrid cloud.
16:49Justin Boitano:You've got optionality. There's cloud environments. There's enterprise environments. There's even edge environments where you may want to deploy your workloads. And taking advantage of all of that with a consistent footprint, like we're building with this AI factory, it gives you the best of all of your alternatives. And so, you know, I think we're bringing the efficiency, we're bringing the flexibility, we're ensuring that we have those confinements, whether it's confidential computing or guardrails or any kind of sandbox technology that I think becomes really critical as we're building and delivering these new capabilities.
17:33Justin Boitano:And if you go back in time before the focus on AI, we developed through decades of experience what Justin highlighted, that pre-production, dev, test, prod kind of best practices. There's a whole set of learnings and rigor and discipline that we've built in building and delivering applications into production that we're bringing as part of an AI factory for building and delivering AI applications into production. I'm speaking with Chris Wright of Red Hat and NVIDIA's Justin Boitano, and we're talking about the AI factory and how enterprises can build AI factories that they can go to production with with confidence and can scale up to the future and really help transform companies into AI natives, as we've been talking about.
18:23Justin Boitano:I want to get into a little bit about specifics, infrastructure and software and platform components. And Justin, I'll start with you. For customers who are thinking about an initial AI factory footprint and might want to start small but have that ability to scale as they scale, how should those customers think about sizing and selecting NVIDIA infrastructure and software?
18:46Chris Wright:Yeah, I think, you know, as the customer starts to try and build the AI factory, they got to think through the five-layer cake that I mentioned previously. So it's where do I have data center power? What is the power density of the data center? Do I want to run air cooling or liquid cooling? That seems to be a decision point right now. A lot of enterprises still run air-cooled data centers. And so, you know, platforms like our RTX 6000s give you, you know, very good price performance. that's kind of a general purpose GPU to do experimentation with. So if you don't know where to start, that kind of gives you a great platform for many different use cases.
19:21Chris Wright:And then from there, you start to ask yourself, well, what's the orchestration management platform that I want to run my business on? And that's why we worked very closely with the Red Hat team. Red Hat's AI factory takes care of really the next few layers of the technology stack from software orchestration and management, model delivery, all the, I'll say, commercial security patching, lifecycle management of all of that open source software so that you can run it with confidence and kind of get the factory up and running. And then you get up into the application layers. And the application layers, the way we try and make it easy for customers to start is we provide reference blueprints, which are examples of proven use cases that even we run on our AI factories at NVIDIA for things like enterprise search that make it easy to then connect into your enterprise documents and do document ingestion and then start to provide benefits to your users.
20:16Chris Wright:And then from there, you can start to expand into, you know, your own developed use cases and such. But that thinking through that full stack is really the easiest way to get going. And then I think taking some of these proven examples is like kind of the quick way to get an early win with kind of your executive leadership team. Yeah, yeah. With them the benefits. And then from there, usually you pivot into, you know, what's the most important business outcome for the company to be competitive? You got to ask yourself for NVIDIA, we're a chip company, we're a software company, and we're a supply chain company when you really boil it down.
20:49Chris Wright:And so we then go super deep into those use cases to make sure that we're enabling, you know, tens of thousands of chip designers, software engineers, or all the people dealing with all the components that allow us to have supply and availability of building this rack scale infrastructure. and if we're world-class at those, then we can be world-class in market. And I think that's generally how companies should think about it.
21:10Justin Boitano:Chris, on the Red Hat side, what are the key platform components that you see as foundational for this first AI factory deployment? Thinking about things like OpenShift, Red Hat AI Enterprise, AI factory with NVIDIA. You know, when thinking about this first enterprise AI deployment, what are the key platform elements to start with? And also, how should customers think about sequencing them? Yeah, I think for us, the stack starts with hardware, hardware enablement, and then the distributed nature of Rackscale architecture. How do you get access to that whole distributed system? And then going up from there, we start getting into specifics of models and agentic applications and AI-enabled applications.
21:52Justin Boitano:So the bottom of the stack, very clearly, that's the world of Linux, right? Hardware enablement, device drivers, low-level system software, and near and dear to our hearts, we spend a lot of time in that space and making sure that we work closely together with NVIDIA to do that first phase, right against the metal enablement. The next layer above that is the distributed layer. Bringing that rack-scale architecture to life includes a distributed system like Kubernetes. Kubernetes is tried and true in the application space, and it's supporting well delivery of agents or models or other content as containers on this distributed system with access to all the accelerators down at the bottom of the stack.
22:36Justin Boitano:And then that Red Hat AI Enterprise layers on top. And this is where we start to integrate directly with some of the key capabilities that NVIDIA brings, like optimized models like NemoTron or some of the NIMMs. And that's where we bring that distributed inferencing stack that is the foundation for intelligence for the business. So we sometimes call this the metal-to-agent stack and starting with that layer right above the hardware, building up through inferencing and then supporting the models is what we're building together to enable those key reference architectures that Justin highlighted or the validated blueprints or the reproducible plays that you want to bring into the enterprise.
Read the full transcript
23:23Justin Boitano:because I think it's important to have those early wins. Justin highlighted that. It's important to have those early wins. It's an interesting tension. Perfection is the enemy of good enough. So if you have this perfect view of your future world where you've normalized all your data and everything is well-defined, you'll spend all of your time doing that and you'll never be able to get to showing some business value. But if you over-rotate to the easiest thing to do, the flashiest thing I can show, it might not have much business value. So picking those right first key use cases and also having in parallel this long-term mindset of it's a pretty fundamental shift in how we operate.
24:03Justin Boitano:Living in that duality, that's the future on building the right stack to support rapid movement, consistent, reproducible or replayable plays, and building from infrastructure that IT operations teams already understand. They know Linux, they know Kubernetes, they're learning a lot of new things in this context, so we'll give them as much stability as we can along the way. Yeah, no, it makes sense. Going along those lines of getting those first wins, right, which is a great strategy for lots of workplace projects to take on, but I think talking about such a big shift to the AI way of working, if you will, let's look at those first 90 days then, And can you lay out some kind of practical first steps?
24:48Justin Boitano:We've got the elements of the joint stack laid out, the hardware, the software, the metal to agents, as you called it, Chris. What are some practical things that folks listening to the podcast, enterprise leaders can do and kind of structure their first 90 days to get some wins and really start building that AI factory that can grow? Yeah, I'll assume the data center infrastructure is built out.
25:11Chris Wright:Okay, fair enough. Let's assume the data center is built, the infrastructure is built out. So what we publish is what we call validated designs that sort of walk you through a lot of the design decision points of the software. And you have to think of how do I bring all of my software into this factory? How do I make sure I do security scanning? If I'm going to want to rescan everything and operate it, how do I have automation to stand it up? And then quickly, how do I get these first, we call them blueprints, but think of them as like Kubernetes services that you deploy. on the clusters to then get users on the system.
25:49Chris Wright:And then ultimately what we do is we have, we call them like user acceptance test teams that we will roll an application out to to have them use the application. So they can start to, you can start to survey them and understand how are they doing work now versus how did they do it before? How much time are they saving versus how they did it before? And really that time savings is the productivity gain that you're going after. And you can really quickly get to, from time savings across a user group to productivity gains. And so if you can get a 2x productivity gain across a big population of users, then you know you're onto something really big.
26:27Justin Boitano:Absolutely, yeah. Chris, anything, Dad? The learning that you'll gather along the way, I think, is really important. And so the notion of starting with a focused, have a hypothesis and a focused outcome, and also iterating as you go. So it's about how quickly can you move forward? I think that's really important. Our experience internally is reinforcing that. And we started with some really focused examples of data that we want to bring together within Red Hat, the research we wanted to do across that data and having evals. I can't understate the importance of evals. I think that it's an often overlooked part of the stack because they help you ensure the quality of what you're trying to produce.
27:19Justin Boitano:And so building iteratively towards improving your evals, we see this in the public with Frontier Labs focused on benchmarks and evals, but they're just as important within the enterprise. And that iterative process of refining any portion of the stack, it could be your prompting, it could be how you're managing the data sourcing, It could be even the scoping of the problem that you're trying to solve. I think that's really important. And that notion of picking something that's real, so it's not so artificial that you can just show it. It's flashy. You get high fives all around, but it doesn't really change anything internally.
28:00Justin Boitano:I don't think that's particularly useful. So focusing on those things that are real, but again, not making it too big. So it's the right sizing and the iterative process of learning as you go is how you start building the thing that ultimately is quite big. But I think it's starting small and iterating, which we do a lot in open source. We do a lot in software development. And having a little bit of diversity, we have touch points across every different function in our organization. There's different personas, but there's also different use cases. It's more software development oriented. It's more finance oriented.
28:39Justin Boitano:It's more sort of sales and pipeline oriented. Each of these brings a little different dimension that is, again, is helping you flesh out your end-to-end view of what's needed to go through whole-scale AI transformation and be operating with a full-tilt AI factory powering your business. Yeah. And as you get these first projects going, and not to, you know, sort of skip all the hard work in between, but as you mentioned, thinking about, you know, getting something going with an eye towards building out to scale and transforming the whole org. Looking at it from the other perspective, what kinds of guardrails would be, not just could people put in place, but what kinds of guardrails would you recommend would be appropriate kind of from the get-go to make sure that as things scale, as things expand, as, you know, wins are won and people get excited and want to use this stuff more and go faster, What are the things you can lay down kind of from the beginning to make sure that, you know, technical and process and governments and, you know, the guardrails are in place for these kinds of things?
29:43Chris Wright:You know, I think, so one thing that we did up front was we made sure, obviously, our security teams were deeply involved as we did this. Just to make sure, I think you learn a lot about your organization as you start to put AI to work. And you'll find AI is really good at doing discovery in business systems that it has access to. And you might realize you've got user permissions overscoped in areas. And so having the security teams understand, are we allowing too broad of access to what we want to keep confidential within the organization? You'll discover as you start to connect agents into your business systems.
30:20Chris Wright:But there's all kinds of techniques for guardrailing data access. Once you do find systems that it might have access to, you're going to realize that you've got to change permissions of many different business systems. Ultimately, what you're going to want to do is scope the agents potentially as users. Think of them as digital employees. So where a lot of people start is they scope them to the user that's using the app, their permissions, so they see access to the information that they've been granted as an employee in the organization. But as we go forward, these agents are going to start to work more and more autonomously, and we're going to have to treat them almost like contractors we bring in.
30:59Chris Wright:You give them least privilege access into your business systems, and then they got to come back to you and check in with you and ask you for access to more business systems. And you're going to have to have a process in place where you can slowly grant them more access to do the job and fully onboard them into the job that we're asking them to do.
31:17Justin Boitano:Chris, I'm going to turn this one to you first. But, you know, Justin, you can be thinking in the background about your answer. We like to end these. The more time passes, the more I feel like it's an unfair thing to ask at the end of the AI podcast, what's the future going to look like, right, for obvious reasons. But if we look ahead, Chris, a year or two, maybe even three years down the line, if you're feeling really bold, what does the AI factory look like as, you know, agentic AI develops and models keep developing and the infrastructure keeps developing? but especially as more enterprises put these systems, build these factories and use them and put them to use solving real problems and driving new ways of working.
32:01Justin Boitano:What do you think the AI factory looks like a couple of years hence? I think you'd take it from a few different points of view. The one angle would be the layer cake picture. And that one, we have a pretty good understanding of the layer cake. So while there might be some subtleties, certainly in terms of specific tools that will come and go over time, that layering of what we're describing from hardware up through AI-enabled applications, I don't think is something that will fundamentally change. So you look a few years ahead, we'll see something that looks quite similar. How it's used by the enterprise, I think is what's going to shift completely in that timeframe.
32:45Justin Boitano:Today, a more sophisticated enterprise has some agents in production, but they're not entirely agentic. And it's not translated into the core of their operations, essentially. And so that, to me, is the shift that we should anticipate. The autonomous nature of agents and the scoping of tasks will continue to grow. So initially, it was the simple chatbot, which is just essentially fetching information. Then you got a little more sophisticated with stronger and stronger recommendations. You could call that some kind of an assistant. The doing phase of agents and total autonomy, we're seeing that time horizon just stretch.
33:35Justin Boitano:It feels like almost daily, stretch out to be longer and longer. So you can, in the coding context, you can give coding agents very sophisticated tasks, and they will spend hours and hours producing very sophisticated code as a result. That's just the coding example. It's a language that's well-structured. It's a good template for how we should think about the breadth of the enterprise. And so in the end, the AI factory, the layers look similar. The sophistication of the tasks grows and it becomes the core of the business. It becomes the place where we do our, we build our operational practices around.
34:17Justin Boitano:And so in the end, it's not that we're going to go through and kind of augment each of today's processes. Because if you just think of it like that, you take a bunch of questionable, in some places, even stupid processes and automate them. And then you get an automated stupid process. It's really redefining how we work together completely end to end and where agents take on critical tasks in the business that I think is that future view, which again, you put some timeframes on it. We're not talking decades. We're talking quarters away, which is itself kind of phenomenal. But yeah, I think that's, to me, that's that future outlook.
35:00Justin Boitano:Well said. Justin, your thoughts?
35:03Chris Wright:Yeah, I think the way Chris framed it is right is software development, even I'm saying in the last six months has evolved where you can give AI, I'll say, almost like a design document and let it go off and think and produce the code and then do, I'll call it, functional verification of that code to make sure that it's accomplished its task before it comes back to you. And so it's doing very long-running thinking and work that is the work of many, many, many, many software engineers, I'll just say. And I think in the software engineering world, like I said, we've seen this product market fit where we're seeing a 2 to 3x productivity gain with software engineers that can use these long-running agents.
35:44Chris Wright:And if you extrapolate that out, the productivity gains for the whole software industry is massive. But we're now seeing that move into this knowledge worker world and CAD designers, engineers across every industry, where they can do the same thing, where they can start to give design document goals to long-running agents, where they can basically explain the exit criteria and give the agent the tools to do the functional verification and say, come back when you're done. And so I think that's what the future of work is going to look like in two to three years. You're going to have different agents working for you that you give these more structured, long-running tasks to.
36:24Chris Wright:They go off and think and do the work, and then they come back to check in in a period of time. And that will make us all infinitely more productive than we are today. And searching through UIs, trained to find information on our own. And so I think, you know, we're going to live through a big, you know, change in how we work in the next couple of years. But every company across every industry and every job function will really be transformed with the use of an AI factory. Perfect place to leave it.
36:54Justin Boitano:Chris, for listeners who would like to learn more about your work, the work that Red Hat is doing, places online, they can go, obviously, website, social media, technical blog, other places? Where would you direct a listener to learn more about what Red Hat is doing with AI Factories? The easiest one would be learn more about the Red Hat AI Factory with NVIDIA. So that's sort of an easy thing to search. And you'll find information from redhat.com. You'll find more information together with NVIDIA on the NVIDIA website. And that's a really easy place to start digging into the Red Hat view on all this content.
37:34Justin Boitano:Fantastic. Chris Wright, Red Hat, Justin Voitano of NVIDIA. Again, thank you both so much for taking the time to come on the pod and talk about AI factories and really the future of work as we landed on Justin. It's an exciting time to be alive. Thank you, guys. Thanks, Noah. Thank you.
38:02Thank you.
38:29Thank you.
From the publisher
Enterprises are moving from AI pilots to full‑scale AI factories that turn data into trusted digital intelligence. Red Hat CTO Chris Wright and NVIDIA’s Justin Boitano unpack the "five‑layer cake" AI factory stack, from accelerated hardware and hybrid cloud infrastructure to models, agents, and production‑grade governance.




