Building Durable AI Agents

9 Jul 2026 · 47 min · 22 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Building durable, reliable AI agents for enterprise/cloud use. Hamza argues that classic ML pipeline principles (graphs, productionization, retryability) still apply, but agents add cycles, real-time branching, and non-determinism. He explains “harness vs agent”: the harness is the runtime program that maps LLM tokens/tool calls into actions; the agent is the combined harness+model. Key fragilities include lack of durability when agents move off laptops into cloud sandboxes, state loss, hard-to-update harnesses, long-running tasks, and infrastructure issues (queues/workers, idempotency, observability, checkpointing).

Notable examples

coding factories needing sandboxes (E2B/Daytona/Modal), tool-call failures mid-loop, and replaying traces with different models/tool calls. Kitaru (ZenML) adds an open runtime/SDK+UI for harnesses with checkpointing, replay, and adapters (Anthropic/OpenAI Agents SDK).

Guests

Hamza Tahir, co-founder at ZenML; background in MLOps and productionizing ML workflows/DAGs across compute backends.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Background on ML and Agents

1:18 to 2:12

Discussion on ML pipelines and the transition toward AI agents.

“I think you are out at the AI Engineers World's Fair, right?”

The Evolution of Agent Workflows

2:12 to 4:36

Hamza discusses the evolution of workflows in AI and the role of agents.

“Now you're getting into agent, agentic things.”

Challenges of Durability in Agents

4:36 to 8:14

Exploration of the fragility and reliability issues in current AI agents.

“So yeah, I'm happy to chat deeper if you're interested in a particular.”

Future of Multi-Agent Architectures

8:14 to 11:52

Discussion on the future scalability and complexities of agent architecture.

“Yeah, I would be curious to know maybe as we get into those things, maybe just some of the stories or types of failures that you see.”

Understanding Agent Harnesses

11:52 to 14:00

Analysis of the 'harness' concept and its significance in AI agents.

“And that migration is well and truly underway right this year.”

Understanding Structured Outputs and Tool Calling

14:00 to 14:59

Learn how structured outputs and tool calling enhance AI agents.

“with structured outputs and tool calling.”

The Evolution of AI Agents and Harnesses

15:00 to 15:57

Discover the progression of AI tools and the significance of harnesses.

“you literally have the, literally do a exec, like an eval in Python, which converts the string into code and runs it.”

The Role of CloudCode and Model Coupling

15:58 to 16:45

Explore how CloudCode integrates with AI models and the implications of coupling.

“I will obviously be very embarrassed in a year when I'm listening back to this.”

The Tension Between Open and Proprietary Harnesses

16:46 to 18:28

Examine the debate over open standards versus proprietary harnesses in AI.

“And Opus 4.8 by now, unlike Opus 3.5, understands what CloudCode itself is.”

Agent Architectures and Operational Challenges

18:29 to 20:04

Understand the architectural challenges faced when building AI agents.

“Yeah, I would guess that there's going to be a mix for some time.”
Show all 22 chapters

Framer's Integration with AI Agents

20:05 to 21:02

Learn about Framer's approach to integrating AI agents into their platform.

“So somehow like Peter Steinberger and Mario and Armin, like those guys, they had better ideas how to manage that program, which just caught fire.”

Designing Agents for Real-World Applications

21:49 to 22:46

Discuss key considerations for building durable agents for various applications.

“Well, Hamza, it was really good intro to, I guess, just how to think about or categorize some of these things and the shifts that we're seeing in our mind.”

Infrastructure for Scalable Agent Platforms

22:47 to 25:17

Explore the infrastructure necessary for developing scalable agent platforms.

“So I think first, we have to understand the different types of agent workloads that can run in the background.”

Challenges in Agent Task Execution

25:18 to 28:00

Understand the complexities involved in executing tasks with AI agents.

“Like, what are some of the main sets of things that like need, I need to be thinking about, I guess?”

Exploring Complexities in Workflow Systems

28:00 to 30:50

Learn about the challenges of managing dependencies and long-running tasks in multi-workflow systems.

“So yeah, I mean, there's a whole plethora of problems that start happening from the infrastructure perspective to make things really reliable, and that you don't just write defensive code all the time.”

Infrastructure Challenges in AI Agents

30:50 to 34:18

Discover the infrastructure complexities and developer concerns associated with updating AI agents.

“I mean, imagine you have a coding factory, like this is the buzzword for AIE right now, right?”

Introducing Kitaru: A Solution for Agent Builders

34:18 to 37:58

Understand the Kitaru framework and its purpose in enhancing the experience for AI agent builders.

“given the right ability to go back in time and introspect and try to run some of these experiments.”

Optimizing AI Processes and Future Goals

37:58 to 41:24

Examine strategies for identifying bottlenecks in AI workflows and the vision for future enhancements.

“And I would have gotten a better and cheaper result.”

Reflections on Current AI Industry Trends

41:24 to 42:01

Gain insights into the current state of the AI industry and the ongoing challenges in MLOps.

“And then maybe something like you said, oh, if we could solve this, if we could move this big rock, like that would open up a lot of opportunities.”

Optimism and Concerns in AI Development

42:01 to 43:30

The discussion covers the current state of AI, exploring both challenges and advancements in the industry.

“you know, you catch me at a very opportune time.”

The Rise of Open Models and Internal Infrastructure

43:30 to 45:15

The conversation shifts to the potential of open-source models and the importance of internal infrastructure for competitive advantage.

“We have amazing open efforts with open source models.”

Encouragement for Open Projects

45:15 to 45:52

Listeners are encouraged to explore open-source projects and tools for AI development, emphasizing accessibility.

“Yeah, that's, I think, a great perspective.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:02Welcome to the Practical AI Podcast, where we break down the real-world applications of artificial intelligence and how it's shaping the way we live, work, and create. Our goal is to help make AI technology practical, productive, and accessible to everyone. Whether you're a developer, business leader, or just curious about the tech behind the buzz, you're in the right place. Be sure to connect with us on LinkedIn, X, or Blue Sky to stay up to date with episode drops, behind-the-scenes content, and AI insights. You can learn more at practicalai.fm. Now, on to the show.

0:41Daniel Whitenack:Welcome to another episode of the Practical AI Podcast. This is Daniel Whitenack. I am CEO at Prediction Guard and really excited for today's episode because it fits right in the theme of our show, which is practical AI, focusing on some things that are actually useful and practical. Have with us today Hamza Tahir, who is co-founder at ZenML. And they have a new product, a new project out, Kitaru, which is focused on agents and making agents durable, which is super interesting. And Hamza is joining us today. I think you are out at the AI Engineers World's Fair, right? Yeah, I am. It's like 7 ,000 people.

1:27All of our crowd gathered in one small, like, big hallway. So it's just fantastic to be in San Francisco when the energy is so high.

1:36Daniel Whitenack:Yeah, that's awesome. Always inspiring. And really cool to see also growth in that from Swix and others who've really built up an amazing community over time. Friends of the show. So if you haven't checked it out, go ahead and check out what they're doing over there. But yeah, excited to dig in today, Hamza. Maybe just to set the stage. I know your co-founder of ZenML. There's kind of some background with that. project and product around ML ops. Now you're getting into agent, agentic things. I love your perspective on maybe first off kind of the world that you have been inhabiting around ML and ML pipelines as now we're all thinking about agents and generative AI and all of these things.

2:31Daniel Whitenack:Like what, from your perspective, before we get into agents specifically, like what role does the more traditional ML models, training pipelines, et cetera, play in our world moving forward from your perspective? Awesome. That's, I think, a great one to start with. And thank you for inviting me on the show. I appreciate the opportunity. So I co-founded ZenML about five years ago. So this was really almost at a point where MLOps was really reaching fever pitch on, And, you know, there was all sorts of chatter about how to productionalize AI and machine learning workloads. And I had done four or five years of that in my previous job where I was co-founding another company trying to deploy ML models in disparate, you know, compute backends and all over, especially out of Germany where I'm based.

3:29So that led me to having a framework internally that we used that you could write workflows and DAGs and you could deploy them on these different backends. And that turned out to be ZenML. We open sourced it. We got a bit of traction at the beginning and a few projects and revenue and we raised. And that's been the story so far. And smack dab in the middle of this, from then to today, we had the agent renaissance. And it felt a bit funny because in MLOps, it was like DevOps reinventing itself. And with agents, it's like MLOps reinventing itself. And at the end of the day, it comes down to these very basic principles of how to write good software engineering code that runs non-deterministic code in a way that's safe and reliable and retryable.

4:18tribal. So I think if anything, even if you throw away every other tool that we ever used in MLOps, the principles and the learnings that we took from productionalizing these applications at scale still translate and are being rediscovered. Even at the AIE WorldSphere, I sometimes hear docs. I'm like, I seem to remember I've heard this talk before in the MLOps conferences. So yeah, I'm happy to chat deeper if you're interested in a particular.

4:43Daniel Whitenack:It's interesting, like you're talking about the workflows, DAG DAG pipelines, there was very much this phase, and at least this is how it occurred to me. I don't know if everyone had this perspective, but there was this phase with generative AI around workflow automation, and there was this very much workflow focus for some time with things like N8N or whatever. And those tools are still very useful, of course. But there is like this DAG focus. And now it seems like people have thought, well, and I remember having Jeffrey from Noose Research on the show and he's like, well, with like their Hermes agent or whatever, it's like, well, I don't want to impose my workflow into this, but there's still a workflow under the hood.

5:33Daniel Whitenack:Like there's decisions made, there's a workflow executed. It's just like you're not defining it. And so from the human perspective, you actually don't see that that workflow in like a visualized DAG, but it it sort of exists under there. Is that partially why you think like some of these principles carry carry over or reinvented in a new way? Because like at the end of the day, there is some flow of things being executed. Right. Yeah. I mean, a graph like all software is essentially a graph, right? Like when you write code, you have if this, else that, execute this, execute that. That's like the sequence of steps you're doing in a graph.

6:13I think the way we thought about workflows before were more deterministic, obviously. For example, I come from the world of machine learning pipelines, so it's like load your data, preprocess your data, train your model, evaluate your model. So we sort of knew the steps ahead of time, and that lended itself to a graph structure. And, you know, graphs introduced order to the chaos of just willy-nilly scripting things. And then, obviously, DAGs are directed acyclic graphs, but they are graphs, so they're directed in one direction. They're acyclic. They don't have cycles. They don't loop back. So in the world of agents, the acyclic part gets very tricky because it cycles all the way down, right?

6:56It loops. So it's where we had to reinvent ourselves as well. So back in like 2023, four users started hacking our pipeline engine to run agents like dynamic steps, conditional branching, state through artifacts, store workarounds. And then, you know, we were like, OK, they're starting to fight our abstraction. And so we introduced a new dynamic mode. And I think the main, the key difference here is really, as you said, everything is a workflow. So everything is an agent, in my opinion, is just an unrolled graph. So it's like LLM call, tool call, LLM call, tool call. Sometimes you do the tool calls together.

7:37So you're just, it's like a tree structure. So the trace, I guess, is a graph. And we just needed to make our system more capable of having graphs that are defined in real time versus statically compiled at the beginning. And we did that quite early. And since then, the same abstractions have worked, except, yeah, and we can get into this. There's this whole new workload that we need to think about different things about, like durability, state management, retries, and how those things work and replays. And those are the things I've been working on in the last two years.

8:14Daniel Whitenack:Yeah, I would be curious to know maybe as we get into those things, maybe just some of the stories or types of failures that you see. There's like one piece, which is like how we handle those, how we instrument things, how we, to your point, make our agents more durable. But that's assuming like they need to be made more durable. They are currently not durable. they're fragile, right? So what are the, I guess, help the audience understand some of those main categories of why agents in our world today are not durable or reliable or however you define that. Yeah, so I think in order to follow that thread, we need to see where the world is going, right?

9:07So if you look at how agents, most people, when they think of agents, honestly, still outside of our little bubble of these 7 ,000 people in the Moscone Center, they think that agents are these local cloud code instances or, you know, like Hermes or something. Yeah. And I think that because we come from this revolution of local processes that run on your computer and you're token maxing, you know, in your local machine, I think that it's very hard to then understand durability because that durability, I guess, and recovering from failure in your local machine turns out it's a simpler problem than when these things migrate out of your little machine and go into a sandbox or somewhere in your computer, in computers running in the cloud, defined by or controlled by your company that are executing arbitrary code and doing all sorts of things like MCP calls and, you know, like loading skills.

10:05Daniel Whitenack:And would this be like, to kind of go off of your point, This would be certainly people are running cloud code or whatever on their local machine, like you're talking about. But ultimately, at least the way I've heard companies express this is to really lean into this element of how the future is going to play out. there will be this digital workforce, however you want to think about it. Agents that are operating and actually taking action within your enterprise infrastructure, right? That aren't tied to someone's laptop. And they maybe then eventually are not just acting, you know, in a single agent type of way, but eventually there are many of these agents that are operating, which I'm sure adds another level of complexity.

10:59Yeah, these fleets or swarms or, you know, the multi-agent architectures that keep coming up, I think they're getting very real now. So, I mean, the companies certainly that we work with are deploying these things. And what ends up happening once you do that is there is no limit to the scale that you can reach technically, right? A disregarding token spend. And, I mean, why would you not arbitrarily scale that to 100 ,000 agent executions per, I don't know, hour, if you could afford it, once that you loosen the yoke of the laptop, right? So I have the feeling that, you know, as things, as infrastructure gets more and more mature and architectural practices get more mature, we can't really predict how much volume of agents would be running once they're not running locally.

11:52And that migration is well and truly underway right this year. And people who get there tend to just it's like this sort of slow adoption. And then suddenly, boom, there's like a crazy increase. And then at that scale, of course, you have different sorts of problems with agents like and I guess that's what we we can spend a bit of time talking about.

12:12Daniel Whitenack:And these agents that would run in your cloud environment, you know, disconnected from your local environment, you know, people might have some ideas of some of the local ones, like a Cloud Code or OpenCode or OpenClaw or whatever we're talking about. Could you give us some examples of kind of the other type of agent? Like what are they built on top of? Like what are these? Do you envision these as mostly kind of software vendors that are building vertical focused agents and they're deploying, you know, into companies or companies building them themselves? I'm just trying to get like a bit of a vision for people of like what these things might be.

13:00So I think this is a big battle in the industry right now, who owns that part of the stack. So, you know, let's start with the most easiest part of the stack. It's the harness, right? So the harness and the model providers have been for the last year tightly coupled. And I like maybe we should spend a little bit of time talking about what the harness is. Yeah, yeah, go ahead, please.

13:23Daniel Whitenack:we've mentioned it on the show but i think always it's like this is one of those concepts that it has popped up and it is often very subtle for people like what we're what we're talking about right it's very ephemeral anyway because what used to be an agent is now a harness yeah so uh the definition i cling on to in um july 2026 is is basically um you know we we we had this a notion of an LLM model, right? The LLM model is simply a token generator. So in itself, it doesn't do anything. It doesn't do action. So the way we did action was we started three years ago with structured outputs and tool calling.

14:02And suddenly, what is structured outputs and tool calling? That's simply imposing policies on what type of tokens can the model generate to predict the next action inside the environment in which it's playing. So you can, ahead of time, you give it like five tools and you say hey um you have get weather or you have like ping hamza or like ping dan and like you have all these bunch of tools to send emails or do different things and here's the definition of the parameters of these tools and then whenever you want you think that you need more information given that description you can return me a bunch of tokens which i can parse and actually map that to my code and execute it and that while loop like that's the agent, right?

14:47And suddenly we started having a program that runs on your machine at the beginning, which mapped those tokens back to the tool calls. And turns out that's harder than it looks. So it's, I mean, I still remember when I did it the first time, like three years ago, it was like, you literally have the, literally do a exec, like an eval in Python, which converts the string into code and runs it. I guess that's still what's happening under the, under the hood. But basically three years of software development later, we had things like CloudCode, which started doing more things while having that while loop.

15:22So, you know, things like compaction when the context window gets too big or things like ensuring that the tool goals are, you know, have the right parameters. Indexing memory. Yeah. Indexing and memory. So that software program is called the harness. So it's basically the thing that gives your brain a hands, like a body. So your brain is just spewing out tokens and you're converting those tokens into actions. And the outcome that comes out of combining the harness and the model is the agent, right? So that's how I think about it. I will obviously be very embarrassed in a year when I'm listening back to this.

16:03Daniel Whitenack:For what it's worth, I mentioned Jeffrey and Noose, like he gave the same, a very similar metaphor of like brain and body. So at least we're in, yeah, we're in safe, safe, safe territory for now until the industry decides to flip this. Yeah. So, so, so, so obviously the most famous harnesses being like some, some interesting phenomena started to happen because what we ended up figuring out is that the models started to asymptote. performance and what the model providers quickly found out given their valuations is that if you just make this program better and you manage context better and you manage memory and all those things general purpose tasks get easier to solve and therefore you could just make a better program right the better program will win given given the model stays static and and then they started a weird rl loop so what they what they did was they like cloud code for instance has a very you know, CloudCode is the harness and underlying it is Opus 4.8.

17:07And Opus 4.8 by now, unlike Opus 3.5, understands what CloudCode itself is. So it's self-aware in the way it's that it's running inside CloudCode. So when it calls a tool like edit file, it actually uses certain parameters and the tool calling is very accurate. If you drop GPT 5.5 into CloudCode, the harness, that would not be as accurate. It wouldn't perform as good because simply the two things have coupled together. So the harness with the reinforcement learning loop that has gone on for the last year and a half has coupled deeply with the models. Now, on the same time, there has been this renaissance of open harnesses, right?

17:50So we had Pi. We have frameworks like Lengraph or Piedantic AI. And those harnesses are of the opposite opinion that we should have an open standard that shouldn't be tied to the model because we don't want to be tied to the model. I mean, if the US government decides to ban and unban models, that shouldn't affect our business outcomes.

18:10And that has been an underlying tension in the industry for the last year and hasn't resolved yet. So I wouldn't know whether, my intuition says that an open harness will win, which standardizes those things. But on the other side, I mean, you know better than most people that, I mean, a cloud code works, right? Why would I use anything else? So that's another thing.

18:32Daniel Whitenack:Yeah, I would guess that there's going to be a mix for some time. You see this also with the existing software vendors who are trying to figure out how they show up within agents, like a NetSuite or a Salesforce or whatever. You see on the one side them officially supporting MCP interaction with their platform because they see maybe, well, the value of our platform is how we manage this data, how we create a functions on top of it, the ability to do actions within this environment and the interface through which people do that, like an agent, really like the value is in the platform and that data, not so much the web interface through which people interact with the product.

19:24Daniel Whitenack:But then on the other side, you see the same individuals, you know, the same companies creating proprietary agentic things in their web, you know, web app, which you can't swap out to your point. You can't swap out the, you know, in many cases, you can't swap out the underlying endpoint to change out the brain to your point. Yeah, because the brain ultimately is commoditized. It's like electricity. So what else are you going to make the next whatever? a trillion dollars on um and and that's the whole stack right and we we haven't even gotten to the point where so this is all local right so this this revolution is happening december 2025 we come into this era of suddenly this works right excuse me for the language um and then it's like okay so what has started working ah so we have pi and we have open claw and the innovation is that it's running all the time and it has a heartbeat and it's like it's managing memory somehow better and okay, what did we actually change?

20:22Oh, we changed the RNS. So somehow like Peter Steinberger and Mario and Armin, like those guys, they had better ideas how to manage that program, which just caught fire. And yeah, and you know, we haven't even gotten to the point where let's get these things out of the computer and into a Kubernetes cluster or something.

20:42Daniel Whitenack:Yeah, yeah, super, super interesting. Sometimes the gap between AI-generated ideas and production-ready work can be really frustrating. Agents can produce things that aren't editable and don't live in your core workflows. That's why I really love what our partner Framer is doing with the way they're integrating agents into their website platform. Agents work in the same place where the real site is designed, managed, reviewed, and published. It lands on the canvas, stays editable, and can be published when the team is ready. Framer is a complete website platform, not just a builder. So teams can launch and keep improving their sites in one place.

21:25Daniel Whitenack:Learn how you can get more out of your website from a Framer specialist or get started building for free today at framer.com slash practical AI for 30 % off a Framer Pro Annual Plan. That's framer.com slash practical AI for 30 % off. Framer.com slash practical AI. Rules and restrictions may apply. Well, Hamza, it was really good intro to, I guess, just how to think about or categorize some of these things and the shifts that we're seeing in our mind. Let's go under the assumption that, you know, maybe there's even people listening to this podcast that have envisioned products that have agents in them.

22:11Daniel Whitenack:They're maybe building agents that they intend to sell into the enterprise, or maybe it's part of people in a larger company that are building their own agents for operational efficiencies or internal tools. And they are picturing that future that you described, which is, hey, I actually want these things to run off my laptop. I want to be able to take my laptop, go into a meeting and not have to keep it up all the time for my agent. So in that sense, then, assume we have some of those agents running in that type of environment. What are some of these, I guess, fragilities or points about agents in terms of how people are architecting them now, how they're building the agent harness piece that make them fragile, not durable?

23:01Yeah. So I think first, we have to understand the different types of agent workloads that can run in the background. So if you're talking about a chatbot interface, or a voice agent, that's largely different from a personal assistant that's running or a deep research or an auto research that's just doing a single loop and always on and achieving some outcome. So I think that for all of those different types of use cases, there's different infrastructural pieces and different ways to think about it. So I would say that if it's a simple agent that's running online somewhere, I think that I would personally default to something which is very easy, whether that's using the model providers like Anthropic has Anthropic managed agents.

23:54they have inbuilt like routines and all those things inside the harness now that you can that you can deploy and they take care of all of that infrastructure for you. Now when it starts getting more complex where your workflows start looking a little bit more business process-y meaning you have a workflow which is doing things like okay fetch an order here and do some processing with an LLM which is an agentic loop here and then do some post-processing there. So then it starts looking a bit more complicated. And if you do that at scale, then you might want to own that infrastructure. So the first thing you do is sort of have an agent platform internally as an enterprise.

24:34And this, I highly encourage people who are scaling beyond single teams to really invest in. Like similar to MLOps, where the people who build the best MLOps platforms like Uber, they won their markets. I'm not saying the MLOps platform was the reason Uber won the taxi game. But it was a big reason that they made, you know, things like search pricing and, you know, they were the best product out there in the market for a while. And I think the best companies will be the ones that invest internally to build platforms that can run these things at scale. Because simply the act of doing that informs you so much internally of what works and doesn't work for your particular business context.

25:15Yeah.

25:16Daniel Whitenack:And in that, I guess the set of things that you need to have in place to support that, let's say I did want to go down that path, I want to own it. Like, what are some of the main sets of things that like need, I need to be thinking about, I guess? Yeah. So the first thing is you need to pick your harness or you need to give your teams the ability to create, you know, using an agentic framework, whatever, let's say you're using LandGraph or PyDantic AI or something like that. Then you need to deploy it onto some compute target. Let's say you picked something, probably you have other applications running.

25:54So probably you use the same thing, whether that's ECS on AWS or whether that's a Kubernetes cluster running somewhere. And now you have an API, so you have a REST API. So I'm more of a Python guy, so I'm going to give Python analogies. So let's say you have a fast API application. You have a post REST call that kicks off an agent. And obviously, the first failure mode that will happen is that, you know, you can, these things are not quick requests, RESTful, that just execute in milliseconds. These are very stateful processes that run for a long time. So they can't run in process, right? This is where you start to farm out things architecturally to things like workers and task queues.

26:36And when you start thinking infrastructurally like that, then you need to think about, okay, what is a task queue? Is it a pull-based system? Is it a push-based system? So this is just where in the realm of how do we actually execute this so that at a certain scale it keeps working? then obviously you have this a very simple task you would be so imagine you have an API and you just directly from the API you spin up a worker like a salary worker well you don't want to be doing that because what if the worker goes down because at scale the workers can go down for a number of reasons like network failures or maybe the pod didn't exist at that point because somebody was using your underlying compute for something else.

Read the full transcript

27:24So a very normal one-on-one architecture is just putting a message queue, a message broker in the middle of that. So you have a bus and the fast GPI server, which is the entry point, puts events on a bus and then these events are durably persisted. And then you have workers that can spin up and down and you can do an intensive task like an agent tick loop inside that worker. from there you have this problem that oh this system is very hard to update it's it's you know you have to ensure things like idempotency um and you know you need to make that system more observable so for for example if you're um processing a file upload and suddenly you have a multi-workflow system like your task consists of two steps rather than just one thing like maybe you're calling something at the beginning and then you're doing something else suddenly you have dependencies between the workers so suddenly this complexity starts to explode from a very simple oh i just need to slap a queue in front of my fast api to okay i need a dag like a workflow execution a thing right and then and then you're like looking at it like okay how do i make this dag orchestration thing work with my harness and how do i how do i make things like in-flight updates like what if i my agent up like what if there's a long-running task for 30 days right and you know that's running for 30 days i mean most tasks are not running for 30 days but they can very soon right and what if then the the next person kicks off another 30-day task um but that's a different you've updated the code so what happens in the 31st day should i use the latest code or should i use the old code that's version to that agent system and you know what happens in within the 30 days uh the llm model gets banned by the u.s government i'm going to keep saying that because I'm in the US now.

29:13So yeah, I mean, there's a whole plethora of problems that start happening from the infrastructure perspective to make things really reliable, and that you don't just write defensive code all the time. Like you don't want your developers to be writing defensive code. You want them on the offense writing use cases.

29:28Daniel Whitenack:Yeah. And what about, I guess, and I really liked how you framed this when I was looking through the Kitaru pages and docs is like there's a lot of questions that you could could be asking like what if the tool calls time out what if I don't use a re-ranker can I use a cheaper model for there's a lot of I guess developer side questions that come up certainly related to some of those things like the supply chain thing you mentioned about you know a model being all of a sudden identified as a supply chain risk, which I think is a very real one. But there's all these other things. There's so many possibilities of how you could update your agent harness.

30:16Daniel Whitenack:It's very much, I mean, I remember just when I would teach workshops and just talking about ML models or single LLM calls, I would get pushback saying like, well, how do we test these things like rigorously, right? They're non-deterministic and they could return anything. And having a background in physics, I was like, hey, we wouldn't know a lot about the universe if we weren't able to test things that were non-deterministic, right? Look, this is, I mean, you're so right, because I was lingering too much on the infrastructure side probably like the developer side the things that you're talking about is post factum that once those things are running yeah so you have some notion of state yeah then you have it's a multi-layered problem yeah yeah there is the infrastructure and there it's like both interact in interesting ways like in in the way how long something takes or how it times out or whatever certainly related to network and infrastructure but then there's all of these choices that you can make within the agent, right?

31:23Yeah. Yeah. I'll give you a good example. I mean, imagine you have a coding factory, like this is the buzzword for AIE right now, right? Coding factories. So we're going to automate software engineering. And imagine that an agent is running in such a system that I described in workers. And suddenly it obviously has to execute code, right? So it needs a sandbox. And this is why the sandbox providers like E2B, Daytona, Modal are so popular nowadays, because you need a sandbox to execute arbitrary code. So that's another problem suddenly that appeared on the infrastructure side. And now what if the file system that you're editing the code, you fail before you can commit the code?

31:56Like, how do you get back to that state? When, you know, 20 ,000 tool calls in, Cloud Code is almost about to finish the feature and it just suddenly fails. And are you mounting the file system into the pod? Because this is not your machine anymore, right? Yeah, yeah. Typically when pods die. And then how do I, like, even if you get to the end and it finishes, like, what if somebody looks at that and says, this is a feature. Like, this is not what I wanted. So how can I go back and evaluate, to your point, like, how can I go back and evaluate, hey, could I have done this faster, cheaper, better if I had done an experiment of using GLM maybe instead of GPD?

32:35Or maybe I wanted a different tool call. Maybe I should have given it so many tool calls. It got confused. Maybe just let's give it two calls and start again.

32:43Daniel Whitenack:Yeah, and I think that would be especially true if you are moving from this, like this is my personal agent, which I can look at the output and then maybe provide, you know, I can actually commit to memory or skills or that sort of thing. Like this is how I want it to go. This is how you should do things. This is, but it's another thing if these are more autonomous, they're operating in the enterprise environment or like I'm a software vendor and I'm creating my own, you know, my own agent, which is my product, my IP. And this shouldn't just be like a general purpose agent, right? Like there should be an opinionated take on how to do these things in this vertical that operate better because I'm infusing actual opinions into the harness, right?

33:33Yeah. Actually, that is the biggest unsolved problem, right? Dan, like, I mean, you've probably been building agents for years now. So I still am terrified updating my agent in production. I'm terrified because I have no idea. I have literally no idea given the entropy in that system.

33:49Daniel Whitenack:Yeah. How I can even adding a word to the system prompt, what would happen? And that terrifies me in a world where you have hundreds and hundreds of millions of these things running an enterprise in flight, right? and suddenly you have to task the poor agent developer to crock all of these context states. It's a very stateful application and try to make an educated guess as to how to update this system in a way that it wouldn't break for all your customers. It's almost impossible, right? Without given the right observability, given the right ability to go back in time and introspect and try to run some of these experiments.

34:26It's like you need to have some sort of a simulated, I keep saying the physics words, entropy, It's just simulations. So, yeah.

34:35Daniel Whitenack:And that's some of what you're doing with Kideroo, right? Which is this, could you explain kind of, maybe just introduce a little bit of the idea behind that? And I know there's some really interesting things around replay, for example. But, yeah, would love to hear more about that. So, yeah, Kideroo comes from this concept of, you know, so it still uses ZenML. So ZenML is our open source product that's been running in production for enterprises for five years now. So we've really, we sort of know how to do workflow orchestration now. So we, our task this year was how do we convert that into using that engine in a way that's ergonomic to agent builders and sort of try to answer, try to take opinions on some of the things that we just spoke about.

35:20So Kitaru is built on top of ZenML. it's a new SDK and a new UI that works almost from the harness backwards. Because I think very important is the harness layer, because that's what most people are, is the entry point to agent building. And a lot of the opinions taken in that harness layer, actually, as we just talked about in this episode, it do matter a lot at the infrastructure layer. So our goal is to build an open runtime that allows you to take any harness and deploy it in a way that you don't have to think about some of these problems and then on top it gives you all these goodies of replay and how we end up doing that is um turns out if you hook into the harnesses there's certain checkpoints in state that you can snapshot um and store while you're running it on let's say a kubernetes cluster that will make life a lot easier once those are inevitable failure modes to kick in.

36:15So the first order of business was how do we get the best world-class adapters for all the popular harnesses like Anthropic Agents SDK, OpenAI Agents SDK, that you can just drop Kitaru as a runtime in and it gets it from your laptop out to the production. And then once it's running, it makes it resilient to tool calls failing in the middle of an agentic loop by storing that state in an external database or a blob storage. And once we have that state, then we have all sorts of interesting questions like, you know, like we already talked about, what if you want to mock a tool call or if you want to change the replay, like replay a different model and drop that in the middle of your trace.

37:01And this is a very interesting problem anyway, like from a scientific perspective, because if you change the model midway in a multi-turn conversation or a multi-turn situation then what ends up happening is that you never really know um that the model that the agent would have ever gotten to that point if you had started with a different model anyway so it's anyway a broken experiment and also the problem is that when you replay from the middle um you have certain like if you swap out glm with cloud code like the harness differences also make it just hard to just start again from that point anyway so so there's a lot of code that we have to write in order to make that experiment somewhat, you know, feasible.

37:42And again, it's not perfect. And we're still working with our customers to make it good. But given that you're grounded in your production traces, production executions, and you have a big enough sample size, you could then, like we have seen early signs of where you could just say, oh, I should have just used an open source model or a smaller model, or I should have just swapped out some of the tool calls and made it a bit easier. And I would have gotten a better and cheaper result. in place of that.

38:10Daniel Whitenack:Yeah. Do you view this as, what's the way to put this? So like I think of, you know, taking it out of the AI agent world into like manufacturing, you have like a pipeline in your manufacturing or a line in your manufacturing plant. There's always going to be a bottleneck in there, right? And so often what people recommend, right, is you find the bottleneck, you ignore the other things, you address that bottleneck. and then you have a new bottleneck, right? So one way to approach this would be that way. You could also approach this and look at, well, there's all these things happening, right? And it's really, yeah, there's of course all sorts of optimization theory on like how to optimize things, but like there's all sorts of things that are actually coupled together, right?

38:58Daniel Whitenack:If I optimize this, then that gets faster, but creates a different problem. So like, what have you noticed after actually having, which I think a big piece of this, to your point, is like actually getting visibility into what's happening. But then comes the next question of, well, now what, right? Like what are the best, and maybe it's a part of like, we don't know the best practices around addressing some of these things yet. Yeah, yeah. I mean, exactly how I imagine it as well. It's like playing a game of whack-a-mole and trying to sort things out. So when you fix one thing, another thing breaks.

39:38I think the first order of business, as I said, checkpoint everything. So that's the durability aspect, right? So, I mean, obviously there are latency concerns. You need to be efficient about it. But checkpoint everything so you can look back. And then I think it's almost humanly impossible to do all the experiments. So you've got to get agents doing that, right? So I think our goal really in the future is how can we close the loop from like once, because you sort of know the outcome, right? You want a cheaper, better, faster model agent at the end of the day. So we have the production executions, the traces.

40:09And the first thing I ask people to do when they're using GitHub, I'm like, run it for a week. And then after a week, just filter for the most expensive traces, which were successful, watered by your customers. Go through the checkpoints and see the bottlenecks, just like you said, and figure out the common failure modes. And then, you know, go from there. But obviously, my goal is not that humans do this. our, you know, we shipped an MCP CLI day one of Kitaru because we know that eventually that problem needs to be solved by other agents that sort of are embedded in that loop. So my ultimate goal is you have every time you launch an agent, you have a companion, like a, like a companion, a nurse agent.

40:54Daniel Whitenack:Trainer, or yeah, it's a loaded term, but. Yeah, exactly. I mean, something like a trainer that's, you know, that's constantly looking at, okay, what is this guy doing wrong? And then running experiments in the back, replaying things and constantly editing that thing. I mean, we're very far away. I mean, to be clear, this is a, this is a, this is a dream because then, you know, your agent builders can start from a really state and get to a very efficient state very quickly. But I think we need more tooling and infrastructure to make that possible because that is eventually the future. Yeah. And I guess that does get to, you know, a really good place for us to kind of start wrapping up the conversation, which is as you see the current state and you look to the future, maybe on both sides of this, like what's one area where you're really encouraged and, you know, it excites you of like, hey, things are going in this direction and I'm really happy about that.

41:49Daniel Whitenack:And this is maybe what we could expect. And then maybe something like you said, oh, if we could solve this, if we could move this big rock, like that would open up a lot of opportunities. Any thoughts? So I'll start with the thing I'm less optimistic about for now as an industry because, and you know, you catch me at a very opportune time. Right after this, I'm going to go to the AIE conference, right? So I'm going to be at the World's Fair and walking around the expo. It strikes me how many people have different takes on solving the same four or five problems. So I think we're at this fever pitch of, like we have an MLOps in 2021 when Chip Hu and like maybe we can drop that article in the show notes.

42:30It's a very good, it's a very good, yeah. Like when we did explosion of tools, when we had these MLOps problems and investor money came in and then, you know, it was just extremely confusing to figure out the modern MLOps stack or on top of the modern data stack. And I think that this is why, I mean, we're sort of contributing to that noise, right? By being, by taking opinions. But I feel like we are very early in deciding the canonical ways of separating things like the harness and the infrastructure and the deployment paradigms. And I'm less encouraged by that, by the model providers taking such aggressive stances on how to deploy things and how aggressively they make things hard for other vendors to be interoperable.

43:14Although there are some good signs sometimes, but it's just they're pushed into the harness. And when you're pushed into the harness, you want everything running on your infra. And I think this is a slight negative of economical impact. But on the other side, okay, here are the things where I'm very optimistic. We have amazing open efforts with open source models. Take Minimax or like take the Kimi models or take something like GLM coming out of like China right now, but hopefully also the US and Europe with Mistral that have made it very economically viable to start, like the enterprises have started to really take it seriously to replace that, the bigger model providers with their own systems.

44:06And I think that's going to be eventually where when they start thinking, okay, we need to have system engineers that think of that problem and start playing around with and fiddling with the other pieces like the harness and the durable runtime. And I'm very encouraged by the fact that we have now performing models that are just, this is like two weeks old, right? The GLM is as basically 95 % of Opus 4.8, which is just phenomenal. Like we were at a trajectory that that was very far away. But in a world where you can have sort of the same performance with open models, then I think you see more investment in creating internal platforms that can deploy those agents and then And open harnesses will be a huge thing, in my opinion, because a thing like Pi or a thing like OpenCode or something like that, or even specific harnesses for law or for science or something, will probably explode in growth.

44:57And suddenly you have this thing where people will start realizing that actually investing in internal infrastructure at the enterprise level is probably going to be the competitive differentiator in a world where the tokens regress down to the cost of electricity and the models become commoditized. And this is, I'm extremely, extremely encouraged by the recent progresses around that.

45:19Daniel Whitenack:Yeah, that's, I think, a great perspective. And I think it encourages people also that are listening to the show, please check out these open projects around open harnesses. A lot of them are even linked in the ZennML website and integrations that they have. I encourage you to check out the Zennemel website, try some things in your own environment. It's never been easier to spin up some of your own tools and infrastructure and try things. And I would encourage that. You'll find some links in the show notes. Thank you so much for joining us, Hamza. It's been great. Thank you for having me.

46:02All right, that's our show for this week. If you haven't checked out our website, head to practicalai.fm and be sure to connect with us on LinkedIn, X, or Blue Sky. You'll see us posting insights related to the latest AI developments and we would love for you to join the conversation. Thanks to our partner, Prediction Guard, for providing operational support for the show. Check them out at predictionguard.com. Also, thanks to Breakmaster Cylinder for the beats and to you for listening. That's all for now, but you'll hear from us again next week.

From the publisher

What does it take to move AI agents from demos to reliable production systems? In this episode, Hamza Tahir explores how MLOps principles are shaping the future of generative AI, covering workflows, agent harnesses, fleets, and the infrastructure needed to build durable, scalable systems.  The conversation dives into open source tools, production challenges, and how ZenML's new project, Kitaru, helps developers build resilient, replayable, and observable agent systems.

Featuring:

Links:

Sponsors:

  • Framer: The enterprise-grade website builder that lets your team ship faster. Get 30% off at framer.com/practicalai

Upcoming Events: 

More from Practical AI

All 157 episodes
Building Durable AI AgentsPractical AI · 47 min
Listen in VO