From AI-Curious to AI-First: Engineering Production AI Systems

28 Jul 2025 · 36 min · 20 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Engineering realities of building production-grade AI systems; why “AI-first” is an engineering discipline, not research or demos. It covers five pillars: software-first foundations (CICD, FastAPI, async Python, Docker), true AI agents (planner-memory-tool orchestration loops), enterprise RAG beyond naive vector search (contextual chunking, hybrid retrieval, re-ranking, retrieval evaluation), LLMOps (prompt/version management, evaluation, monitoring/observability, governance), and crossing the production chasm (cost/latency, security/compliance, resilience, integration, change management).

Guests

Armand Ruiz, VP of AI Platform at IBM, leads 1,000+ engineers building production AI systems.

Key claims

75%+ of orgs use AI but ~80% see no material EBIT impact; demos create “slideware mode” and deployment bottlenecks; durable moat is superior system engineering around commoditized models.

Notable examples

Stripe fraud detection (59% to 97% accuracy); Netflix LLMs over a content knowledge graph using Metaflow; Spotify AI DJ with domain-adapted models, human-in-loop feedback, VLLM/quantization/prompt caching (4x engagement; +14% vs base).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Reality of AI Adoption

0:45 to 2:39

Discussing the challenges organizations face in adopting AI effectively.

“We're using the insights of Armand Ruiz, right?”

Building Production-Grade AI

2:39 to 4:04

Exploring the first pillar of Ruiz's framework: starting with software.

“Now, for a lot of people, Gen AI feels like this magical new thing.”

Technologies for AI Systems

4:04 to 6:22

Overview of crucial technologies Ruiz recommends for AI systems.

“What specific technologies does Ruiz point to for building this foundation?”

The Role of AI Engineers

6:22 to 7:51

Defining the emerging role of AI engineers and their importance.

“The AI engineer is exactly that, this emerging hybrid discipline.”

Understanding AI Agents

7:51 to 8:08

Discussing Ruiz's insights on the nature of AI agents and their architecture.

“It's a function of your engineering culture and your engineering capability.”

Components of AI Agents

8:08 to 13:04

Exploring the essential subsystems necessary for building AI agents.

“This one tackles a concept that I think is widely misunderstood, the true nature of AI agents.”

Retrieval-Augmented Generation

13:04 to 14:01

Examining misconceptions about retrieval-augmented generation in AI.

“This one tackles another area ripe with misconception, retrieval augmented generation, or AA.”

Understanding ARAG Limitations and Solutions

14:01 to 18:08

Learn about the limitations of naive ARAG systems and how to engineer sophisticated alternatives.

“Plus, relying only on semantic similarity, what sector search is good at, often misses things that require exact keyword matches.”

Engineering Discipline for LLM Systems

18:08 to 22:45

Discover the principles of LLM operations and the importance of composition and orchestration in AI systems.

“We're past the point where just prompt engineering and maybe some fine-tuning are enough.”

Navigating Production Challenges in AI

22:45 to 28:01

Explore the key economic, security, and integration hurdles businesses face when deploying AI systems.

“OK, brings us to Ruiz's fifth and final pillar, which is super pragmatic.”
Show all 20 chapters

Challenges in Production AI Systems

28:01 to 28:22

Explore key engineering strategies to address challenges in production AI.

“Security means input filters, output scanning, preventing prompt injection.”

Case Study: Stripe's Payment Model

28:22 to 28:44

Learn how Stripe's model improves fraud detection accuracy significantly.

“That's a great summary of the core challenges and the key engineering strategies to mitigate them.”

Building on Transaction Data

28:44 to 29:36

Understand how Stripe trained a model on payment transactions for better insights.

“Stripe's work, particularly their payments foundation model, is a really fascinating case study.”

Case Study: Netflix's Knowledge Graph

29:36 to 29:56

Discover how Netflix integrates LLMs with their content knowledge graph.

“basically overnight when they deployed this system.”

Leveraging Software Foundations at Netflix

29:56 to 31:04

Learn about Netflix's Metaflow platform's role in powering recommendations.

“They use LLMs with their massive content knowledge graph, right?”

Case Study: Spotify's AI DJ

31:04 to 31:41

Examine Spotify's AI DJ and its engineering-first approach to music recommendations.

“They're AI DJ and the contextual recommendations they provide.”

Engineering for Real-Time Recommendations

31:41 to 32:38

Understand how Spotify optimizes AI systems for real-time user engagement.

“They also implemented a human-in-the-loop feedback system to continuously refine the AI's conversational style and recommendation quality, directly tackling that subjective output evaluation challenge from LLM EPS.”

Synthesis of Case Studies

32:38 to 33:18

Synthesize insights from Stripe, Netflix, and Spotify's engineering approaches.

“You know, looking at these case studies together, Stripe, Netflix, Spotify, they paint a really consistent picture, don't they?”

Pillars for Building an AI-First Organization

33:18 to 34:11

Learn key pillars for cultivating AI engineering culture within organizations.

“It's a fundamental transformation of your organization's capability, its processes, and its mindset.”

The Importance of System Composition

34:11 to 35:17

Explore how the composition of systems provides sustainable competitive advantage.

“So individual teams don't have to reinvent the wheel.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Welcome to the Deep Dive. business value most companies are seeing. The potential is, I mean, it's undeniable, but the reality on the ground for a lot of organizations is quite different. So our deep dive today is all about the engineering realities, you know, what it actually takes to build production-grade AI systems. We're using the insights of Armand Ruiz, right? He's the VP of AI platform at IBM, leads over a thousand engineers doing this stuff in the real world. Exactly. And his core assertion is pretty blunt. Becoming truly AI first is fundamentally an engineering discipline. It's not just research anymore.

1:01And the data absolutely backs him up on this. There's a 2025 McKinsey Global Survey. It shows over 75 % of organizations are using AI. Sounds great. Yet a staggering 80 % of those same companies report, well, no material contribution to their bottom line, their EBIT for their generative AI work. That's a huge production gap. It really is. So what's going on? What's causing this massive disconnect? Well, a big part of it is how easy it seems to create a compelling demo. You know, hook up a language model to a vector database maybe, and suddenly you've got something that looks impressive. Right.

1:36The dangerous illusion of progress, as Ruiz calls it. Precisely. It makes everything look much simpler than it actually is under the hood. He calls this getting stuck in slideware mode. Slideware mode. I like that. And if you zoom out, there's this 2025 survey from Rider.com. It revealed something pretty alarming. 42 % of C-suite execs admit AI adoption is actually tearing their company apart. Wow. Tearing them apart. Yeah. Mostly due to internal friction and siloed development. 72 % of executives reported that as a problem. Get all these siloed demo grade projects popping up and they mask the huge engineering effort needed for real reliable production systems.

2:15It's this weird paradox, isn't it? We can prototype incredibly fast, but that speed is creating a deployment bottleneck down the line. That's a great way to put it. So our roadmap for today's deep dive is basically to dissect Armand Ruiz's five-point strategic framework. He really believes this can bridge what he calls the great divide between AI hype and shipping code. Sounds like a plan. All right, let's cut to the chase then with Ruiz's first pillar, and it's maybe the most challenging one. Start with software, not just research. Now, for a lot of people, Gen AI feels like this magical new thing.

2:49So why is Ruiz so insistent that it's really about disciplined engineering? Why is that view so vital? Well, it's because these AI systems, especially the ones using large language models, they behave very differently from traditional software. Traditional code is deterministic, same input, same output, every time. Right. Predictable. Exactly. But AI systems are probabilistic. They're non-deterministic. Their behavior is learned. It isn't explicitly programmed step by step, and it can even sort of drift over time as the underlying data or model changes. Which means your old ways of testing just don't cut it anymore.

3:22Pretty much. Traditional QA methods are insufficient. You can't just write a simple unit test expecting one exact output string. Quality in AI requires continuous evaluation, constant monitoring, and importantly, iteration. Okay. And that's precisely why having a mature CICD pipeline, continuous integration, continuous delivery becomes absolutely indispensable. It automates the testing, the packaging, the deployment of the entire AI system. Not just the code, you mean? Not just the code. The prompts, the configurations, the evaluation data sets, all of it. It's the whole package. So it's less about picking a specific programming language and more about committing to this whole architectural philosophy, this engineering discipline.

4:04That's it, exactly. What specific technologies does Ruiz point to for building this foundation? And more importantly, why are they crucial for AI specifically? Yeah, his recommendations are very intentional, very much tied to the nature of AI systems. First, he mentions FastAPI for building APIs. Okay, APIs. Why FastAPI? Because an AI model itself is kind of amorphous, uncertain. But to be useful inside a company, it needs to be exposed as a reliable service. Fast PI helps create these stable, documented contracts. It effectively decouples the model's inherent uncertainty from the other applications that need to use it.

4:42Got it. Creates a clear interface. What's next? Second, asynchronous Python or async Python. Think about it. Large language model inference, getting an answer back, it can take several seconds sometimes. Yeah, they're not instant. Right. And if your application just stops and waits for each call, it grinds to a halt. It can't handle other users. Async processing lets your system handle thousands of concurrent requests efficiently. It switches to other tasks while it's waiting. It's really an architectural necessity for scalable AI. Okay, that makes sense for responsiveness. What else? Third, Docker for containerization.

5:15You know the classic, it works on my machine problem. Oh yeah, the bane of developers everywhere. Well, it gets amplified in AI because you have all these complex dependencies, specific library reversions, system packages, maybe even GPU drivers. Docker packages everything up the code, the libraries, the OS bits, ensuring you have identical execution environments everywhere, from the developer's laptop right into production. It eliminates a whole class of deployment failures. Consistency. Crucial. And the last one. And finally, back to CICD. The experimental nature of AI means you need to iterate rapidly.

5:51Automated CICD pipelines give you that velocity, but also a safety net. They test the whole system automatically, catching issues before they hit users, and dramatically speed up how often you can safely ship updates. It's really fascinating how these foundational software engineering principles become even more critical in this probabilistic AI world. It's not like AI replaces them. It demands them even more. That's a perfect summary. And this blend of software engineering rigor with machine learning understanding it seems to be creating this demand for a new kind of role, doesn't it? Yeah. The AI engineer.

6:23Absolutely. The AI engineer is exactly that, this emerging hybrid discipline. They bridge the gap between, say, data scientists who are focused more on experimentation and model building and traditional software engineers who build the application logic. So they're the ones actually making it work in production. They're the ones translating those promising models out of the research notebooks and into reliable, monitored, scalable API services that the rest of the business can actually use. Ruiz calls them the most critical talent profile for any company serious about being AI first. You know, there's this constant debate around Andrew Neng's point that the real value isn't necessarily in building the next biggest foundation model, but in the application layer built on top.

7:06How does Ruiz's engineering first philosophy connect with that? Why is building around the model the durable advantage? It connects directly. As these powerful foundation models become more available, more commoditized, maybe through APIs, the real lasting strategic moat for your organization isn't going to be owning the absolute smartest model off the shelf. Right, because someone else can just use the same model. Exactly. The advantage comes from having the superior engineering capability to build an efficient, reliable, scalable system around that model. That system includes your proprietary data pipelines, your fine-tuned agentic workflows, your optimized infrastructure that's much harder for competitors to copy than just switching to a different LLM provider.

7:50So the ability to ship shippable AI, as Reese puts it, that's the real differentiator. Ultimately, yes. It's a function of your engineering culture and your engineering capability. That's the durable moat. Okay, that commitment to shipping shippable AI really leads us perfectly into Ruiz's second big insight. This one tackles a concept that I think is widely misunderstood, the true nature of AI agents. He argues pretty strongly this isn't just a fancy chatbot. It's a fundamental architectural leap. What exactly defines that leap? Yeah, this is key. A production-grade AI agent isn't just about conversation.

8:25It's really a cognitive architecture designed to mimic human problem solving. It's not just a single monolithic call to an LLM. So how does it work differently? It operates on an iterative loop. Think, thought, action, observation. The agent thinks about the goal. It decides on an action using its available tools. It executes that action, and then it observes the result. And that observation feeds back into its next thought process. Okay, so it learns or adapts based on what happens. Exactly. It's this iterative, stateful process that fundamentally distinguishes a true agent from just a reactive, stateless chatbot that mostly just answers one question at a time.

9:03So it's not just talking. It's doing things and adjusting course. What are the core building blocks, the components, that make up this cognitive architecture he talks about? To build one of these, you need several distinct subsystems working together. First, you need a planning engine. It's like the agent's strategic brain. Strategic brain. Okay. What does it do? It takes a high-level goal like summarize these reports and email the key findings and breaks it down into concrete, executable steps. But crucially, a robust planning engine supports dynamic replanning. Meaning if something goes wrong.

9:35Precisely. If an API call fails, or maybe some data it needs is missing, the agent doesn't just stop. It analyzes the error and generates an alternative plan. It needs to have fallback logic built in. This is absolutely vital for production systems because you can't have silent failures. Right. You need resilience. OK, so planning engine first. What's second? Second are memory hierarchies. And this goes way beyond just remembering the last few lines of chat. It's more like how our own brains manage information. OK, intriguing. How so? You can think of it in layers. There's short term memory, like working memory, holding the immediate context, what the user just said, what the agent just did.

10:14Then there's midterm memory, which is more like an agent's personal journal or episodic memory. It stores a persistent record of past interactions, conversations, task outcomes. So we can learn from past mistakes or successes? Exactly. Avoid repeating errors. Remember what worked. And then finally, you have long-term memory. This holds the agent's enduring knowledge base, maybe distilled insights, learned procedures, and user preferences for true personalization over time. There's even academic research on things called memory operating systems to manage all this. Wow. Okay. Planning, memory. What else?

10:48Third is intelligent tool orchestration. Tools are how the agent interacts with the outside world. APIs, databases, search engines, code interpreters. They're the agent's hands. Makes sense. But it's not just about picking the right tool. It involves meticulous parameter preparation, formatting the data correctly for the tool's input. And maybe most critically, it needs robust error handling. What happens if that API is down or returns an unexpected error? Back to the resilience point. Absolutely. A production agent has to have predefined retry logic, maybe fallback strategies if a tool fails, and mechanisms for graceful degradation doing the best it can even if some tools are unavailable.

11:27That is definitely a far cry from just plugging an LLM into a chat interface. Yeah. And Ruiz even talks about taking this further, orchestrating armies of agents, mentioning frameworks like Microsoft Autogen or Crew AI. Can you give an example of how that might work? Yeah, absolutely. Think about a complex task like software development. Instead of one giant agent trying to do everything, you might have a team. A planner agent outlines the features based on requirements. A coder agent takes those specs and writes the actual code. Then maybe a critic or tester agent reviews that code, runs tests, and suggests improvements.

12:03They collaborate. It's like a virtual dev team. Sort of, yeah. And Ruiz's own example of agentic rag is another good illustration. He describes a meta agent that receives a complex query needing information for multiple documents. Okay. Instead of trying to read everything itself, it orchestrates specialized document agents, each an expert on one specific document or data source. They each provide their piece of the answer, and the meta-agent then synthesizes those partial answers into a coherent final response. That sounds much more scalable and robust for complex multi-source information. It really is.

12:36And the key takeaway here is that the intelligence or the autonomy of the agent isn't just inherent in the LLM itself. It emerges from the careful composition and orchestration of all these components working together, the planner, the memory, the tools, the control logic. So the LLM is like the reasoning core, but the architecture around it determines how effective it can be. Exactly. You can have the most powerful LM in the world, but if it's embedded in a poorly designed, brittle architecture, you'll get a poor agent. The system design is paramount. Okay, moving on to Pillar 3. This one tackles another area ripe with misconception, retrieval augmented generation, or AA.

13:13Rees argues the common view that ARRI is basically just vector search is an oversimplification that leads to, under quote, silent and catastrophic failures at the retrieval stage. Why does that naive ARRI, the kind you see in quick demos, fail so badly in real enterprise settings? It fails for a few key reasons, especially with complex enterprise documents. The typical approach is chunk the document into pieces, create vector embeddings for each piece, store them, and then do a vector search based on the queries embedding. Right, the standard demo setup. But arbitrary, fixed-size chunking often breaks things.

13:50It can sever critical context. Imagine splitting a sentence or paragraph right in the middle. The retrieved chunk might contain the right keywords, but lack the surrounding information needed to make sense. Ah, okay. Context gets lost. Exactly. Plus, relying only on semantic similarity, what sector search is good at, often misses things that require exact keyword matches. Think specific product codes, technical terms, or names that might not be semantically close to the query's phrasing, but are lexically identical. So it fails on both context and precision sometimes. It can, yes. And the dangerous part is the LLM is often so good at generating fluent text that it can take these irrelevant or incomplete chunks and weave them into an answer that sounds plausible but is actually factually wrong or misleading.

14:35The failure isn't the generation, It's the retrieval feeding it bad information, garbage in, garbage out. Okay, so if that's how naive ARAG fails, what does a sophisticated enterprise-grade ARAG system look like? How do you engineer it to overcome these limitations and actually get reliable results? Getting ARAG in the enterprise requires moving up the ladder of sophistication with more advanced retrieval techniques. First, it starts with thoughtful chunking. Not just arbitrary splitting. Right. Moving beyond fixed sizes to contextual chunking splitting documents along logical boundaries like paragraphs, sections, or even table rows.

15:11An even more advanced method is something called sentence window retrieval. Sentence window? How does that work? The core unit you retrieve is a single sentence, which gives you high precision. But then you also provide the LLM with a window of sentences around that core sentence to give it the necessary context. Best of both worlds, potentially. Interesting. Okay, better chunking. What's next? Second, and this is a really significant leap, is embracing hybrid search. This acknowledges that pure semantic search or dense retrieval using vectors isn't enough. Neither is pure keyword search or sparse retrieval using older methods like BM25.

15:48So you need both. Exactly. Hybrid search combines both approaches. It runs dense and sparse retrieval in parallel and then intelligently merges the results. This combination handles a much wider range of enterprise queries, robustly queries that need conceptual understanding, and queries that need precise keyword matching. Okay, hybrid search makes a lot of sense. What else improves it? Third, and this is critical for quality, is adding a re-ranking step. This usually involves a two-stage pipeline. Two stages. Yeah. The initial retrieval, hybrid search maybe, is optimized for speed and recall finding all potentially relevant documents, even if some aren't perfect.

16:24You might get, say, the top 50 candidates. Then you use a second, more powerful, but computationally more expensive model, often a cross encoder, to re-rank just those top candidates. The cross encoder looks at the query and each candidate document together in much more detail to assess true relevance. And pushes the best ones to the very top. Precisely. It dramatically improves precision, ensuring the handful of documents actually passed to the LLM for generation are the most relevant ones. This significantly reduces the chance of the LLM getting confused by irrelevant information. Got it. Chunking, hybrid search, re-ranking, anything else?

17:02One more absolutely crucial piece, systematic evaluation. You cannot treat our gag as a black box. You need quantitative metrics specifically for the retrieval component. Things like hit rate, precision at K, mean reciprocal rank separate from evaluating the final generated answer. So you can diagnose where the problem is, retrieval or generation. Exactly. This allows you to systematically diagnose and fix issues in the retrieval part, which is often the root cause of poor IG performance. Wow. That is a lot more involved, a lot more engineering than just embedding some chunks. It really sounds like there's this whole rag sophistication ladder that organizations need to climb.

17:38Absolutely. The source material actually describes it that way, a ladder from level zero, which is that naive reg up through levels incorporating hybrid search, re-ranking, contextual chunking, all the way to level three, which is agentic reg, where the agent itself might refine queries or orchestrate retrieval. So it really drives home the point. Effective IoT is an incremental engineering process. Requires genuine information retrieval expertise, not just plugging in a vector database. Couldn't have said it better myself. All right. Pillar four. Sure. This focuses on LLM system design, and Ruiz makes a strong claim here.

18:12We're past the point where just prompt engineering and maybe some fine-tuning are enough. He says it's now fundamentally about composition and orchestration. What is this new engineering discipline he points to, LLMOPFs? What's it all about? Yeah, LMOPs. It's essentially a specialized domain that extends traditional MLOs and DevOps principles, but specifically tailored to the unique challenges that come with large language models. Okay, so how is it different? What makes LLMs need their own ops? There are a few key differentiators. One is prompt management. Prompts in these systems are almost like source code.

18:46They dictate the model's behavior. So LLM ops involves version control for prompts, testing prompts, managing libraries of prompts. Treating prompts like code. Makes sense. Then there's non-deterministic output evaluation. As we discussed, LLM outputs aren't always predictable and quality can be subjective. So LLM OAPS needs sophisticated evaluation pipelines. These might combine automated checks using other models, rule-based heuristics, and crucially, human in-the-loop feedback systems. Right, evaluating fuzzy outputs. Tricky. Very. And the third big piece is composition and orchestration. LLM OAPS isn't just about deploying a single model.

19:23It's about managing these entire chains or agents or complex workflows involving multiple LLM calls, RAG systems, and tools working together. Okay, that sounds like a significantly broader scope than just deploying a single ML model in traditional MLOps. What are the key stages or components in this LLLF lifecycle that teams really need to master? Managing these complex, composed systems involves several critical parts. It naturally starts with data management and preparation, both for any fine-tuning you might do and for feeding your RG systems. Got it. Data first. Then, as we mentioned, prompt engineering and versioning, storing prompt templates, versioning them like code, running tests against different prompt versions.

20:06Then comes evaluation and quality assurance, which has to be a continuous process. You need to track performance against benchmark data sets to detect a model drift when the model's behavior changes over time, and often use A-B testing to compare different prompts or system versions in production. Continuous evaluation. Yeah. What about when things go wrong? Debugging these complite systems sounds like a nightmare. It can be, which is why monitoring and observability is perhaps one of the most critical and often overlooked components of LLMOs, especially for those distributed agentic systems. The key here is distributed tracing.

20:39Distributed tracing, like tracking a request across multiple services. Exactly, but tailored for LLM systems. It captures the entire execution path of a request. The initial prompt, the calls made to the ILM system, which documents were retrieved, any tool calls that were made, intermediate LLM reasoning steps, the final response. This lets you pinpoint the root cause when something fails or produces a weird result. It directly answers Ruiz's question, how do you debug them? This is how. That level of visibility seems essential. Okay, what's the final stage? Finally, there's deployment and governance.

21:15This includes the CICD automation we talked about, but also implementing crucial guardrails. Guardrails? For what? For security things like detecting and preventing prompt injection attacks or scanning outputs for sensitive data leakage. For compliance ensuring regulatory requirements like GDPR or HIPAA are met, maintaining audit logs. And also for cost control implementing budgets or rate limits. It really sounds like those orchestration frameworks we hear about, like Langchain, Lomindex, Semantic Kernel, must be pretty central to making all of this LMO-ed stuff actually work in practice. They're the glue holding these composed systems together, right?

21:49They absolutely are. Frameworks like these act as the essential glue. They provide the higher-level abstractions needed to build, manage, and reason about these complex chains and agents. They really embody Rui's principle of composition. They bring the software engineering structure that's needed to turn these ideas into maintainable applications. And this leads to a really profound shift, doesn't it? You said the deployable asset isn't just the model anymore. Exactly. The thing you deploy, the unit of change, is the entire system or workflow. It's like a directed graph of components, prompts, models, retrievers, tools, decision logic.

22:25Even a seemingly minor change to a prompt template is effectively a new deployment that needs to go through the whole testing and validation pipeline. Which perfectly explains why this new dedicated engineering discipline, LOMOs, isn't just, you know, nice to have. It's absolutely necessary if you want to do this seriously. Couldn't agree more. Organizations that keep thinking they're just deploying a model are going to really struggle to build reliable, maintainable and scalable AI applications. OK, brings us to Ruiz's fifth and final pillar, which is super pragmatic. He calls it crossing the production chasm.

23:01This is basically where most promising AI initiatives die, right? Because those flashy demos conveniently ignore all the messy real world constraints. That's exactly it. The demo works fine on a laptop with 10 users, but production is a different beast. So what are the biggest, maybe harshest economic realities that companies slam into when they try to move from that demo stage to actual production deployment? Two big ones often bite companies unexpectedly. Cost and latency. Let's talk cost management first. Unlike traditional software with fixed server costs, LLM costs are often highly variable and usage-based you pay per token, essentially.

23:39Which can get expensive fast if you're not careful. Very fast. So mitigation strategies become crucial. Things like intelligent caching of responses for common queries, implementing model-tiering routing, simple, low-stakes queries to smaller, cheaper models, and reserving the big, expensive ones for complex tasks. and adopting fine ops financial operations principles like close monitoring of usage and setting budgets or quotas to prevent runaway bills. Okay, managing the variable cost. What about latency? Yeah, latency optimization. Users, especially in interactive applications, expect near-instant responses.

24:14But, as we said, LLMs can take seconds to generate complex outputs. Which kills the user experience. It can, yeah. So you need mitigation here, too. We already mentioned asynchronous architectures to avoid blocking the whole application. But also using optimized inference engines, Ruiz mentions VLLM, which is a popular high-throughput serving engine specifically for LLMs, and techniques like quantization, which basically optimizes the model's calculations to run faster, often with a small tradeoff and accuracy, to speed up that crucial time-to-first token. So managing cost and speed are paramount.

Read the full transcript

24:47Beyond the economics, though, there's this huge trust and safety imperative. What are the unique security compliance and just general robustness challenges that AI brings into production environment? Building and maintaining trust is absolutely critical. On the security front, you have novel vulnerabilities. Prompt injection is a big one where malicious user input tries to hijack the model's instructions, maybe making it reveal sensitive information or perform harmful actions. Data leakage is another concern. So how do you mitigate those? Mitigation involves robust input validation and output filtering those guardrails we mentioned, scanning user prompts for malicious patterns, scanning model outputs to ensure they don't contain sensitive data before showing them to the user.

25:30Okay. What about compliance regulations? Big one, especially in finance, healthcare, etc. Compliance means ensuring data privacy requirements like GDPR or HIPAA are met. It means having comprehensive audit logs who did what, when, which version of the system was used. Often, for highly sensitive data, it might necessitate deploying models in a private cloud or even on-premise rather than relying solely on public APIs. Right, keeping data secure and auditable. Yeah. Just making sure the thing doesn't break. Robustness. Exactly. Robustness and failure handling. Production systems have to operate in an imperfect world.

26:07APIs go down, databases get slow, networks glitch. As Ruiz emphasizes, these AI agents must not fail silently at 2M. Love that phrase. It means building in comprehensive error handling, automatic retry mechanisms for transient failures, and designing for graceful degradation. For example, if your ad database is temporarily offline, maybe the system can detect that and fall back to giving a simpler, non-RAG response based only on the LLM's internal knowledge, rather than its crashing. That kind of resilience is a hallmark of professional production engineering. And then there are the challenges that are maybe less glamorous, but just as critical.

26:44Actually plugging this AI into the rest of the company's systems and getting people to, you know, use it and trust it. Oh yeah, the integration and human side. The legacy system integration is often the hidden iceberg. Connecting your shiny new AI service to decades-old databases, internal APIs, complex existing workflows that take significant specialized software engineering effort to build adapters, data transformers, and so on, it's totally invisible in a demo but can consume huge amounts of time in reality. The unglamorous plumbing. Exactly. And then there's the human challenge. Successful AI adoption isn't just about the tech.

27:20It's fundamentally a change management problem. Remember that survey saying AI adoption can tear companies apart? Yeah, the internal friction. Right. Overcoming that requires a formal, top-down AI strategy from leadership. It needs clear C-suite buy-in and sponsorship. And you need to empower internal AI champions, people within business units who can help drive adoption, tailor the AI to specific needs, and gather crucial feedback from end users. Technology alone isn't enough. Okay, so just to quickly recap those production hurdles and how you tackle them, because that was a lot. If you're facing a high cost, you're looking at caching, model tiering, fine ops, high latency.

27:56Think optimized engines like VLLM, async architectures, quantization. If you're worried about inconsistent responses or hallucinations, that's where advanced RA techniques and output guardrails come in. Security means input filters, output scanning, preventing prompt injection. Integrating with old systems needs dedicated adapter work. And if you can't figure out why something's broken, you need that deep observability through distributed tracing. Does that cover the main points? That's a great summary of the core challenges and the key engineering strategies to mitigate them. It really highlights that production AI is systems engineering.

28:32Okay, fantastic. Now, let's make this concrete. These principles aren't just theory. Ruiz points to leading companies actually putting this engineering-first approach into practice. These case studies really drive home his thesis. Let's start with Stripe. Yeah. Stripe's work, particularly their payments foundation model, is a really fascinating case study. It aligns perfectly with Ruiz's idea of AI becoming core infrastructure, Pillar 5. How so? What did they build? They trained a massive transformer-based model similar in style to an LLM, but specifically on tens of billions of anonymized payment transactions.

29:07This created a general-purpose vector embedding, a numerical representation for every single payment. So not just for text, but for transaction data. Exactly. It moves way beyond simple vector search on text. They then built a specialized real-time classifier using these embeddings. You can think of it as a highly specialized agentic system focused purely on fraud detection. And the results. Pretty staggering. They reported that their detection of sophisticated card testing attacks, where fraudsters test stolen card numbers, improved from about 59 % accuracy to an incredible 97 % accuracy, basically overnight when they deployed this system.

29:43Wow, 59 to 97. That's huge. It really demonstrates the power of using these AI architectures as core infrastructure, as Ruiz suggests, combined with that sophisticated, specialized, agentic design from Pillar 2. Okay, impressive. What about Netflix? They use LLMs with their massive content knowledge graph, right? Yeah. Sounds like a different kind of challenge. It is. And Netflix's use case really exemplifies both advanced RG, Pillar 3, and the absolute necessity of having robust software foundations, Pillar 1. Their knowledge graph is this huge structured database about all their content, actors, genres, relationships.

30:19It's critical for powering recommendations. And how do LLMs fit in? They use LLMs for sophisticated semantic inference on the graph. For instance, understanding nuanced relationships or inferring thematic connections that aren't explicitly coded. This is really advanced RG in action, grounding the LLM's reasoning with the structured data from the knowledge graph. But the key is the platform underneath. Absolutely. What makes this possible at Netflix's scale is their reliance on their mature internal software engineering platform called Metaflow. It handles the immense scale, the parallel processing needed, strict version control for experiments, reliable deployment, all the hard engineering stuff.

30:57It's a direct validation of Ruiz's first pillar. Start with software. Build the foundation. Right. The platform enables the advanced AI. Okay, one more. Spotify. They're AI DJ and the contextual recommendations they provide. How does their approach align with this engineering first philosophy? Spotify's AIDJ is a really compelling example that touches on almost all the pillars, especially the full LLMO's lifecycle, Pillar 4, and dealing with production deployment challenges, Pillar 5. Okay, how so? Well, first, they didn't just use an off-the-shelf LLM. They heavily domain-adapted Meta's LLMA models specifically for music and podcast recommendations.

31:33That required significant investment in curating specialized training data and building a robust internal training ecosystem. So tailoring the model itself. Exactly. They also implemented a human-in-the-loop feedback system to continuously refine the AI's conversational style and recommendation quality, directly tackling that subjective output evaluation challenge from LLM EPS. And what about getting it to run fast enough for a real-time DJ? That was a huge focus on Pillar 5 deployment optimization. They integrated high-performance engines like VLLM. They used techniques like prompt caching and quantization heavily to ensure they could achieve the high throughput and low latency needed for a seamless user experience at massive scale.

32:16Did all that engineering effort pay off? Apparently, yes. They reported things like a 4x increase in user engagement with the AI-generated explanations for recommendations, and their domain-adapted model performed significantly better, like 14 % improvement compared to the base model. It's really a masterclass in applying rigorous LLM apps and solving those hard production deployment problems. You know, looking at these case studies together, Stripe, Netflix, Spotify, they paint a really consistent picture, don't they? Success didn't come from just, you know, bolting chat GPT onto a database somewhere.

32:48It wasn't a quick hack. Not at all. It came from a deep strategic commitment to engineering excellence, building robust infrastructure, designing sophisticated and composed systems, implementing mature operational practices like LLMOPs, and really tackling the hard, unglamorous problems of deploying and scaling AI reliably. So if we synthesize the core conclusion from Marisa's insights and these examples, it really comes down to this. Becoming an AI-first company isn't just about adopting a new piece of technology. It's really about cultivating a new, more advanced engineering culture. It's a fundamental transformation of your organization's capability, its processes, and its mindset.

33:26And for you listening to this, maybe as a technology leader or someone involved in AI strategy, this translates into some clear, actionable pillars for your own organization. First, based on everything we've heard, invest in the right talent. You've got to prioritize and cultivate that AI engineer profile, that crucial blend of software engineering discipline and machine learning understanding. These are the people who will actually build, ship, and maintain these complex systems. Absolutely critical. Second, think about establishing an LN OLMF's center of excellence, or at least formalizing those practices.

33:59Build that paved road for AI development across the company. This COE handles the standardized tooling, the CICD pipelines, the observability stack, the prompt versioning systems, the security guardrails. So individual teams don't have to reinvent the wheel. Exactly. It lets other teams build and deploy AI features with speed, consistency, and safety, because the foundational infrastructure is handled. Good point. Third, really adopt a ship to learn mentality. We heard how AI is probabilistic and iterative. The only way to truly understand how these systems perform and deliver value is to get them into production, even if it's a small contained feature first.

34:37Foster a culture of rapid iterative deployment, creating tight feedback loops with real users and real data. That's where the real learning happens. I can agree more. You learn by shipping. And finally, and this might be the biggest cultural shift, think systems, not just models. Champion the idea that the LLM, however powerful, is just one component in a larger system. The real, defensible intellectual property, the durable value, lies in the unique composition of the entire proprietary system you build around it, your specific data pipelines, your custom agentic workflows, your optimized retrieval strategies, your robust infrastructure.

35:12That system is the moat. Well said. The value is in the composition. So maybe a final provocative thought for you to ponder building on that idea. in an age where the underlying large language models themselves are becoming increasingly powerful, but also increasingly commoditized, maybe even available through simple APIs. Yeah, where does the real advantage lie? The true, sustainable, replicable advantage for any organization probably isn't just what AI models you have access to. It's how masterfully you build, ship, monitor, and operate the unique systems that bring those models to life within your specific business context.

35:47That engineering capability itself becomes the core competency.

From the publisher

This discussion emphasizes that an AI-first organization is fundamentally an engineering challenge, not merely a research endeavor. It argues that a significant "production gap" exists, where many organizations experiment with AI but fail to achieve tangible business value due to a lack of operational maturity. The text presents a five-pillar roadmap for building production-grade AI systems, focusing on treating AI as software systems, understanding the true architecture of AI agents, mastering advanced retrieval (RAG) techniques, establishing LLM System Design (LLMOps) as a new discipline, and prioritizing deployment realities like cost, latency, and security. Ultimately, it contends that shipping reliable AI is the crucial differentiator, requiring investment in AI Engineer talent and a "ship to learn" culture.

More from Best AI papers explained

All 475 episodes
From AI-Curious to AI-First: Engineering Production AI SystemsBest AI papers explained · 36 min
Listen in VO