Practical workflow orchestration

15 Oct 2024 · 58 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Practical AI Podcast Episode Summary

Episode Title

Practical Workflow Orchestration

Episode Description In this episode, the hosts discuss the challenges of workflow orchestration for data scientists, especially in the current age of AI. They are joined by Adam Azzam from Prefect, who shares insights about Prefect's open-source Python library designed to improve orchestration and visibility in Python-based pipelines. The conversation covers concepts such as agentic workflows and introduces tools like Marvin (an AI engineering framework) and ControlFlow (an agent workflow system).

---

Key Participants

  • Adam Azzam - Principal Product Manager at Prefect
  • Chris Benson - Principal AI Research Engineer at Lockheed Martin
  • Daniel Whitenack - CEO at Prediction Guard

---

Key Concepts Discussed

  1. Challenges in Workflow Orchestration
  2. Workflow orchestration has historically been a complex task for data scientists, compounded by the current AI hype.
  3. The need for handling arbitrary workflows that can fail in various ways is critical.
  1. Prefect Overview
  2. Prefect is an open-source Python library aimed at simplifying workflow orchestration.
  3. Its main functionalities include:
  4. Retrying tasks: Automatically retrying failed tasks with customizable limits and timeouts.
  5. Caching: Holding onto outputs to avoid recomputation.
  6. Transactional Logic: Ability to roll back tasks if necessary.
  1. Marvin: AI Engineering Framework
  2. Marvin serves multiple purposes: a mascot, an internal Slack bot, and a framework for writing LLM (Large Language Model) workflows.
  3. It integrates with Prefect to provide an ergonomic way to handle tasks related to AI and workflows.
  1. ControlFlow: Managing Agentic Workflows
  2. ControlFlow allows for the management of more complex agentic workflows, where workflows can adapt dynamically.
  3. The distinction between deterministic workflows (where the logic is defined upfront) and agentic workflows (where the LLM can create and revise plans) is emphasized.
  1. Future Directions
  2. A focus on enhancing workflow orchestration as LLMs demand more complex infrastructures.
  3. The need for effective management of resource consumption and potential bottlenecks when multiple agents access shared resources.

---

Key Takeaways

  • Orchestration Simplification: Prefect aims to make workflow orchestration more accessible for developers without requiring extensive DevOps knowledge.
  • Resiliency in Workflows: Emphasizing the importance of treating failures as first-class citizens in workflows to improve reliability.
  • Dynamic Workflows: The shift toward creating workflows that can adapt to changing conditions and inputs using LLMs presents unique challenges and opportunities.
  • Community Engagement: Marvin's integration into the community provides personalized support and guidance for developers utilizing Prefect.

---

Conclusion This episode of Practical AI offers valuable insights into the complexities of workflow orchestration in the age of AI. With a focus on tools like Prefect, Marvin, and ControlFlow, the discussion highlights the strides being made to streamline these processes and the potential future developments in the orchestration landscape. For anyone interested in practical applications of AI, this episode is a must-listen.

---

Additional Resources

  • [Prefect Official Website](https://www.prefect.io/)
  • [Marvin Documentation](https://www.prefect.io/marvin)
  • [ControlFlow Overview](https://www.prefect.io/controlflow)

Feel free to explore these resources for more detailed information on workflow orchestration and the tools discussed in this episode.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Welcome to Practical AI. near your users. Learn more at fly.io.

0:37What's up, friends? I'm here with a friend of mine, a good friend of mine, Michael Greenwich, CEO and founder of WorkOS. WorkOS is the all-in-one enterprise SSO and a whole lot more solution for everyone from a brand new startup to a enterprise and all the AI apps in between. So, Michael, when is too early or too late to begin to think about being enterprise ready? It's not just a single point in time where people make this transition. It occurs at many steps of the business. Enterprise single sign-on, like SAML, Auth, you usually don't need that until you have users. You're not going to need that when you're getting started.

1:15And we call it an enterprise feature. But I think what you'll find is there's companies when you sell to like a 50 person company, they might want this. They actually, especially if they care about security, they might want that capability in it. So it's more of like SMB features even if they're tech forward. At WorkOS, we provide a ton of other stuff that we give away for free for people earlier in their lifecycle. We just don't charge you for it. So that AuthKit stuff I mentioned, that identity service, we give that away for free up to a million users, 1 million users. And this competes with Auth0 and other platforms that have much, much lower free plans.

1:49I'm talking like 10 ,000, 50 ,000, like we give you a million free because we really want to give developers the best tools and capabilities to build their products faster, you know, and to go to market much, much faster. And where we charge people money for the service is on these enterprise things. If you end up being successful and grow and scale up market, that's where we monetize. And that's also when you're making money as a business. So we really like to align, you know, our incentives across that. So we have people using AuthKit that are brand new apps, just getting started companies in Y Combinator, side projects, hackathon things, you know, things that are not necessarily commercial focus, but could be someday.

2:24They're kind of future proofing their tech stack by using WorkOS. On the other side, we have companies much, much later that are really big, who typically don't like us talking about them. They're logos, you know, because they're big, big customers. But they say, hey, we tried to build this stuff or we have some existing technology, but we're sort of unhappy with it. The developer that built it maybe has left. I was talking last week with a company that does over a billion in revenue each year and their skim connection, the user provisioning, was written last summer by an intern who's no longer obviously at the company and the thing doesn't really work.

2:56And so they're looking for a solution for that. So there's a really wide spectrum. We'll serve companies that are in a, you know, their office is in a coffee shop or their living room all the way through. They have a, you know, their own building in downtown San Francisco or New York or something. And it's the same platform, same technology, same tools on both sides. The volume is obviously different. And sometimes the way we support them from a kind of customer support perspective is a little bit different. Their needs are different, but same technology, same platform, just like AWS, right? You can use AWS and pay them$10 a month.

3:24You can also pay them$10 million a month, same product. Or more, for sure. Or more. Well, no matter where you're at on your enterprise ready journey, WorkOS has a solution for you. They're trusted by Perplexity, Copy.ai, Loom, Vercel, Indeed, and so many more. You can learn more and check them out at WorkOS.com. That's W-O-R-K-O-S.com. Again, WorkOS.com.

4:08Welcome to another episode of the Practical AI Podcast. This is Daniel Whitenack. I'm CEO at Prediction Guard, where we're building private secure gen AI. and I am joined as always by my co-host Chris Benson, who is a principal AI research engineer at Lockheed Martin. How are you doing today, Chris? Doing fine, Daniel. Maybe a little bit too much coffee on my side. I'm just ready to go. I want to get into the show, man. You're fired up. I'm fired up. I'm on it. Yeah. Well, might you potentially need to take all of those things that you're fired up about and orchestrate them? Oh, yes. It may be into some type of workflow.

4:50Most definitely. Most definitely. Well, we have a great guest for you today, then. We have with us Adam Azam, who is a principal product manager at Prefect. Welcome, Adam. Hey, Chris. Hey, Daniel. Thanks for having me. Yeah. I mentioned in the pre-show that this is the second episode we're recording in a row based based on some shout outs from our friend Bing Sun Chua, who was building some awesome broccoli AI things, practical AI that's healthy for your organization. And he mentioned Arjila, which we just recorded with, but also Prefect. And yeah, we're just really excited to hear about such a practical thing that definitely overlaps with hopefully a podcast that's trying to make some of this stuff practical.

5:42So yeah, I'm wondering if you could tell us a little bit, maybe personally, it sounded like you kind of had a background in wrestling through some of these workflow orchestration things, you know, even before joining Prefect. So how did you kind of come into thinking about these types of problems and what's needed there practically from the developer standpoint? I stumbled into workflow orchestration by accident. It was a total necessity for a previous startup that I was working on. I was working on basically a career co-pilot or like a job search co-pilot in early 2023. And there we were using medium to large language models to basically help job seekers conduct job searches in the background.

6:30And then we would let them have like conversations with jobs. And we would also be able to explain and deduce why they were good fits for particular roles that we matched them with. And what that meant from an engineering point of view is I had to take your like routine call out to a large language model. And then now I had to do this millions of times a week. And when you do this at the scale of millions of times for different types of pipelines, whether or not it's extracting like schemas from particular applicant tracking systems or job postings, whether it's conducting a conversation with a job seeker, you just start seeing exactly how these things fail and break down at scale.

7:12And so when I went to my brain trust and was like, look, I'm having a hard time dealing with resiliency issues at scale. What should I use for something like this? You know, I heard the usual player of like airflow. So I got that spun up and it felt like I had to learn a completely new language. And when I got stuck debugging a particularly nasty airflow job, I turned to a good friend of mine who was running the conversational AI team at Square. And he was like, oh, cool, I'll help you debug this. And then he pip installed Prefect. And so we spent that entire afternoon just rewriting all my pipelines in Prefect.

7:47And that's really what kind of, you know, you speak about like sort of broccoli AI, That is what turned a very well-meaning, well-designed consumer AI app into something that, from an engineering perspective, was durable and resilient. I don't want to presume that everybody knows what workflow orchestration is. It was a thing that I didn't really know what it was until I needed it. I think we need a definition from you, actually. I like to think about workflow orchestration as if you tell us what to run, where to run it, when to run it, and what to do when it fails, Prefix makes it extremely easy to author those rules on your typical script that you've got running locally.

8:32So if you've got your Hello World script, you've got your call out to OpenAI script, you've got your go scrape this web page and then go extract a bunch of data from it, we give you the tools to say, great, this works locally. But now if I want to run it on a massive Cates instance, where I want to run on AWS, or I want to run it on Azure, you can do that in a line of code. If you want to run it every single, every third Thursday when there's a full moon, we make it incredibly easy to express. When you want to say, if you start getting a bunch of 503 errors because the resource is down, or maybe you start getting 429 errors because LinkedIn is onto your scraping org, you're sort of scraping, let's say, pipelining off of it, you can control the custom behavior of saying, well, pause this workflow or maybe start using a different IP address.

9:22So we allow you to bake in different handlings of failure and sort of native resiliency into that. So retries, item potency, caching, transactions. So before we dive into kind of the golden path, can you... I can see you're intrigued, Chris. I am intrigued. Our audience may only be listening to us, but we're all looking at each other in video and stuff. And I guess I definitely have I am intrigued. So go back for just a moment and the painful parts of workflow orchestration. Talk a little bit about like, because I get that the airflow, sorry about the airflow folks out there. Airflow wasn't the right thing for you.

10:01You went to Prefect and it solved that. But before we make that, like what was the pain? And I'm not trying to pick on airflow, but just in general with workflow orchestration, what is it that hurts so bad for people to understand the kind of the pain you came through? Yeah. So when you're at a startup, you have to do every single job. And so let me talk about pain number one. Pain number one is I'm like a trained data scientist. I had to overnight become like a machine learning engineer and a DevOps engineer just to get stuff out into the world. And the how to sanely go from something that I have working locally to something that can operate at real scale in the cloud was, you know, there are specialists who can make that work really well.

10:48At the time, there weren't a lot of tools to make that intuitive and I had other things to worry about. So a workflow orchestrator with a good like ergonomic interface to infrastructure was absolutely key for me being able to do essentially the design work of the product to actually getting it live and accruing value to users. That's step one, which was like, I was a big infrastructure dummy, and this allowed me to approximate a very smart infrastructure person with a few lines of code. I would say that from the actual workflow side, infrastructure aside, is that when you're orchestrating, say, LLMs in particular, they can fail for a whole host of reasons.

11:28At the time, when you didn't really have good structured outputs. I'll say what that means, which was, you know, at the time, if I had tens of thousands of job posts that I was trying to extract information from, even though they were just one single blob of unstructured text, you know, you would say, here's the JSON blob that I expect out of this. I expect, you know, keys that tell me what the title is, what the location is, the salary. And so I would take a document and try and ask OpenAI and say, look, here's the schema that I expect out of this. Can you extract this information? And it would fail for a whole host of reasons.

12:07One, the API for OpenAI was brittle at the time. So it would fail because the resource was unavailable. And then sometimes it would fail because the information wasn't available in the actual document I was trying to extract from it. Then you would have all sorts of parsing issues where what it was returning to you wasn't valid JSON. So there was this whole host of cascading errors of either the actual resource was down, there was problems with the infrastructure I was working with, the data quality was bad, the generation from the LLM was bad, or there were like syncing issues at the end when I wanted to go put that into a data warehouse at the end of it.

12:47And so we saw so many different cascading layers of failure that it wasn't going to come up every single time. You wanted to react to them differently and you wanted to be able to express those contingencies in code. And when I wasn't working with a workflow orchestrator, if the machine that I was on failed over and I was 10 % or like say 90 % of my way through a very expensive job, I would lose the entire state of my workflow. And I would have to start over from scratch. And when things are linearly ordered, nice and serialized, maybe that's not that big of a deal. You're saving records as you process them.

13:28So you know how to restart over. You've got a cursor that you built yourself with your database. But for things that require like, first, I'm going to do these four things, and then they're going to map into these two jobs, then I need to go do 1000 things that come from that, reduce them to two things, keeping track of the state of where you are in that pretty, say, complex graph of dependencies. If you fail on a single node, where do you restart from? And depending on how big those jobs are, it can get really expensive if you don't treat failure as a first class citizen. And so I would say that was the big thing, which was, when I was building this, I was spending about 5 % of my time writing code for the upside.

14:09Like if everything works, this is how the code should run. I had to spend 90 % of my time handling failure. And since workflow orchestrators allow you to handle that failure gracefully, and as a first class citizen, it made my life a lot easier. Does that answer your question, Chris? It does. No, it's a great answer. Yeah, maybe just a follow up on that. So you mentioned some of the things specifically with open AI, kind of failure modes, different things that were happening early on. I'm wondering if you could talk specifically, so workflow orchestration is a larger idea than maybe just machine learning pipelines or AI related workflows.

14:50But what is it specifically about maybe these sorts of workflows that a lot of us are trying to build now that maybe further stress some of these issues and or maybe, you know, even providing some examples within that would be helpful. Like what makes machine learning workflow orchestration or AI workflow orchestration different and the same than just sort of like general workflow orchestration? If I return back to, you know, I gave like five or six theses of like things where I was I was encountering failure a lot. And I think that some of those things are sources of failure that many folks are familiar with, right?

15:34Like I'm calling out to an external service and the service is flaky and it's bad. And that's like existed forever, right? As long as people are building data pipelines, upstream sources being flaky, that makes sense. Hitting deterministic errors of like, I'm ingesting data, but somebody added a new field or they removed a field and I'm dependent on that. And now my pipeline's broken. Or I'm scraping something and target.com, instead of labeling the name of their product with a div whose name or ID is this, they've changed it to that. And now all of my data is corrupted because I couldn't detect that in real time.

16:10And then the last piece is the loading part where you've gotten, you've cleaned all your data, now you want to go put it in a persistent place where you can go query or do analytics on. So that's like classical stuff, right? The extraction, calling out to an external service, the transformation that you're doing deterministically in the loading. So the ETL, like the sort of ETL business of all of that. That's like a persistent problem that exists far before Gen AI or ML workflows. And that's sort of what's been category-defined for workflow orchestration. That's like the single case that people usually break out an orchestrator for is when they're doing ETL-type jobs.

16:47I would say that what's unique these days is that since workflows are now in LLM land are more dynamic, so you really can't plot out every single thing that's going to happen from the start. They are now basically we're dealing in English now, or you may not know the full space of their responses before, like at the beginning of a workflow. So I think that LLMs introduce a dynamism component that's hard to reason about and is kind of escaped classical workflow orchestration. I think the second piece is that the nature of errors here just feel totally new. So the fact that you can ask an LLM for a particular shape of a response, and then you can get parsing errors out on the end of it, that's a new source of failure that's now buttoned up with, say, with some commercial LLM providers that give you very structured, guaranteed outputs.

17:45So you don't see as much parsing errors. But now, I like to joke that you can lead an LLM to JSON, but you cannot make it think where like you can say like, look, I've got this job description and the title is, I'm going to give you a schema that says the job title, the location or whatever. And sometimes you'll say the title it's required. It has to be a string. And the response you get out is I'm sorry, but I could not find a title. And when you're doing this at tens of thousands of jobs, now you also have to reason about, okay, I got now the parsing error, but the error was pushed down deeper in the stack.

18:21Now there's data quality errors that I have to reason about that I didn't really have to account for in last generation's ETL. Things were much more deterministic. We had stronger contracts about what you were going to get. And then I would say the last piece is what makes this harder is that so far, and I hate to keep throwing out random definitions. I hate being a merchant of complexity and talking about why things are super hard, but trying to at least motivate why this is a new source of difficulty. We've got tools to handle this, but why do we have these tools in the first place? And the last piece is around agentic workflows.

18:58Now, this is a buzz term. So what do I mean when I say agentic workflows? Everything that I've talked about so far of like, I get a document and I want to extract stuff from it, or maybe I want to classify it, or I want to summarize it. These are all sort of modern takes on classical ML problems, right? It just, you don't have to bring as much training data. You don't have to train a model first. You're basically throwing the weight of the compressed internet at every problem that you come across. But with agentic workflows, what I mean by that are those are things that operate in a loop, are able to call out to external tools should it choose to, and can create, refine, and reflect on its own plan, which means when you're orchestrating agentic workflows, you have to do this interplay between who's doing the orchestration.

19:51There are some times where I'm coming correct with a plan and I'm saying, first, extract this topic, then classify it, then write an email and send it off. But now I have to be able to add resiliency to a workflow that I'm unaware of at the beginning of it. So if it decides call out to this tool, call out to this API, I now need to be able to reason about resiliency for a workflow that I don't have any visibility to at the beginning. And so I would say that those last three pieces around parsing around like not knowing your full decision space at the beginning and then how that feeds into now having to hand off some bits of orchestration to the LLM to create its own tasks that you now have to execute.

20:34I think that's what makes the new generation of orchestration a much harder and a much more interesting problem.

21:05You know, when we started podcasting back in 2009, an online store was just the furthest thing from our minds. Now we have merch.changelog.com. And you can go there right now and order some t-shirts. And that's all powered by Shopify. What do we do before Shopify? I'll tell you, we did nothing. We couldn't sell. There were other ways, of course, but they were very hard, very difficult. Shopify let us build out an entire front end, obviously branded like changelog is. It's amazing. merch.changelog.com. And our favorite feature is we use their API to generate a new coupon code, a personalized coupon code for every guest that comes on our podcast and they get a free t-shirt from our merch store.

21:47And that's so cool. They choose the shirt they want, they use the coupon code, it arrives free of charge to them and life is amazing. But also, you can go there right now to merch.changelog.com and buy some threads yourself. And that's awesome as well. So upgrade your business and get the same checkout we use with Shopify. Sign up for your$1 per month trial period at shopify.com slash practical AI, all lowercase. Go to shopify.com slash practical AI to upgrade your selling today. Again, shopify.com slash practical AI.

22:45so going back that's a great explanation there before the break um and and earlier i was driving you back to the pain and you've kind of carried us forward so now i want to dive into kind of how you get it all fixed. Wondering if you can start introducing us to Prefect Core and talk about the open source aspect of it, how it's fixing, and then we'll continue from there later. But I'd really love that to, you know, give me an intro to it now. So Prefect is a open source Python library for workflow orchestration. You can pip install Prefect today. Last month, we put out our last major version. So Prefect 3 is out into the world and it's free to use.

23:27So what Prefect enables you to do is when we talk about like, where did all this pain come from? Where the pain came from were classic problems like, I've got this flaky resource, but I know that it's alive, right? It's not like the resource isn't dead. It just, I knew that when I tried to access it, it was overwhelmed. So I need to try again. Prefect makes it incredibly easy to say, look, if you ever see this task fail, retry it. Here's how many times you can retry it. If you want to go really deep on it in the same line of code, you can say, here's how long you should wait between retries if you want to be respectful of the resource.

24:08So Prefect makes it really easy to add retries. Prefect makes it really easy to do caching. So what I mean by that is often when you're building LLM workflows, we talked about if you have really, really complex dependencies between different parts of your workflow. If something fails over and you need to run it again, well, you can either like try and pinpoint the exact dependencies that you have, blah, blah, blah. Or you can say, well, I've already successfully executed this stuff and it has all the same inputs. So I'm just going to hold onto a reference to what its output was. And then now I can just zoom back to where I was.

24:44So I don't have to sort of, you know, we keep track of the state of your workflow, but we make it really easy to say when you want to recompute an answer or when you want to just sort of recall it in a very inexpensive way. And also make it very easy to add transactional logic. What I mean by this, if you've used, like if you're a database nerd, you'll know what transactions are. They're just a way of like co-locating work with each other and then being able to control what happens if something fails. So maybe I want to have task one and task two. But if I see a failure in task two, I want to be able to express like how to undo the fact that task one already executed.

25:23So Prefect makes it super simple to say, if I see a failure in a group of tasks, here's I want to go and undo maybe some writes that I did, which is really helpful when you're trying to build applications that are like building up knowledge about a particular user, where you get to the end of something, you realize that you've made a reasoning mistake. And now you need to walk back through the previous stuff that you've committed to your, say, your knowledge graph or your vector database or what have you. Those are like three nice features that we make it really, really easy for folks to express.

25:55We have handles for like timeouts. If you've got like a really sensitive operation and you want to put an SLA on it that say this thing can't take more than 10 seconds to run, we can fail it out. And then we give you really good ways of custom handling those errors. So if something fails for one reason, say OpenAI's API is down, you can just cancel the rest of your workflow. If it's not down, now you can reschedule it to retry it. And I think that may happen because, I mean, everyone just kind of walked out the door a few days ago. I'm sorry. Yeah. Hot takes abound today. I got to say, as somebody who built on OpenAI in the early days, I'm still such a fan.

26:37They're like wishing the best for that team for sure. But yeah, it's, I think they've had a rough couple of months. But so I'll say that that's like classical stuff where you've got your like, hello world Python function. This is not complicated. It's literally like retries equals three. It's timeout equals 50 seconds. It's you slap a decorator on top of your function. And then it basically gives your Python code superpowers. And this is stuff that normally you would have had to write your whole like own framework to do. And here you can do it in a couple lines of code. I would say that the other piece around Prefect, the other two core value propositions are around infrastructure.

Read the full transcript

27:14So if you have something that's working locally, you don't have to be like a DevOps engineer to get it working on Kubernetes if that's your flavor. But honestly, if you're like, look, I've got this thing working in this Docker container locally on my machine. I want to get it working on Amazon's ECS. It's really like a one-click deploy in order to get it working somewhere else. And the last piece is around observability. We haven't talked about it, but observability is really this pair that comes with orchestration, where if orchestration is really the practice of like, crap, how do I get this stuff to actually run and how to react to it when it fails?

27:52There are some times when it just cascades and fails all the way down. And no matter how much you account for everything, everything goes wrong. And you need to be able to like, see what went wrong so that you can replan, right? Like, you're not going to be able to handle the entire universe's amount of complexity ahead of time. So when it does fail, you need to be able to learn from those mistakes as a human being and actually write in the logic to handle that for next time. So observability is really this element of this thing failed. How do I provide the breadcrumbs to figure out why it failed?

28:25Which data provider was the one that caused me to fail? Was it a parsing error? How often am I seeing that parsing error? If it's OpenAI's fault, now there's a bridge between observability and orchestration. If OpenAI has failed more than 80 % of my requests in the last 10 minutes, now I'm going to switch to Anthropic and we make it easy to switch those two things out for each other. And so I would say that between just adding your most native retry caching transactions, making it dead simple to submit to infrastructure, and then the last piece around really having a clear understanding of how things are working, and importantly, how to figure out when they fail, those are really the core, I'd say, value adds of Prefect.

29:07Yeah, and I would definitely recommend people check out the Prefect docs. There's a quick start in there. Just to give people like a visual of this, I see you have, you know, this kind of converting your Python script into Prefect workflow type of thing. That's really helpful. there's sort of these definitions of like a workflow related to getting information about a GitHub repository, like contributors and repo information. And then, you know, those are in Python functions, you know, like definition, get contributors function. And then above that, you kind of put these decorators, task and flow decorators with options to convert this into a Prefect workflow.

29:52So part of my thought on, you know, follow up on that just to kind of really hone in practically to give people a sense of what it's like to do this. So let's say that I've, I've converted my Python code into some, you know, one or more Prefect workflows. I can run that in my understanding, if I'm looking at the docs, right, I can run that just Python, my workflow. But obviously, like you said, there's a deployment element to this. So could Could you talk a little bit about, you know, I know one of the things that's always a struggle in my past is maybe getting something working locally. And then because I'm going to run it in a different environment or something like that, you know, then I deploy my workflow on top of other infrastructure, right?

30:42And it's hard to just connect that like local development, debugging and development staging production environments. So practically, what does that look like in the Prefect world in terms of going from that Python command to run your workflow locally to, you know, watching things flow through your workflow and a nice dashboard that you have pulled up in production? So I'll do my best to do this without a screen to share. Yeah, it's hard. So you had said, talk to me the experience of taking my Python code and converting it to a Prefect workflow. For folks that are listening, it's as simple as from Prefect import flow and then at flow on top of your function.

31:27That's all the work it takes to take a Python function and give it superpowers, basically, to add all the stuff that we've talked about. And that doesn't take away any sort of attributes of that function. If I take my hello world or my ETL flow that runs locally and I start adding prefix things on top of it, it still runs locally. And so it's not like we don't sort of require additional infrastructure just to have your stuff run. We detect whether or not you're running it locally or not. And if you're running it locally, we execute it as if it's regular Python functions. Now, there's kind of two stories of, well, this thing is working on my machine.

32:09How do I get it to run somewhere else? We have two ways of doing this. And I'm sorry for the vocab lesson, which is, let's say that what's the crawl, walk, run for anybody trying to get something to run remotely? it's you start off on your machine and then it's like well i'm gonna go get a server somewhere and i'm just if i can close my laptop and go to bed and this thing is still running on my server that is what counts for me as remote as running it remotely uh that's like you know the next step in in the hierarchy of needs so for that if it can run locally on your server right like you You go to EC2, you SSH in, you've got your console, you pull your GitHub repo, you pip install your requirements, right?

32:54You're still able to execute that flow on your machine. And now if you execute that function, you hit, say, like.serve. Now in the same way that you would like start a fast API application, it is now running. You can specify a schedule that it runs on. it exposes an HTTP endpoint that you can call out to if you want to invoke it on demand. And it listens to events. So maybe you don't want to just run it on a schedule. Maybe sometimes you just want to be able to hit a specific endpoint to manually invoke it, maybe from your Django app or something like this. Or maybe you want to emit an event into the world that says, look, when my user signs up, when my data set is ready, I want to have this ETL flow go and operate on it.

33:40Because my data is not always ready at 6am. Sometimes it's not uploaded till 6.02. So sometimes you want to have things be a bit more dynamic. That's what happens in a flow.serve world. So where are we in the hierarchy of needs? I had it on my laptop, but it stopped running when I shut it. I go and I get an EC2 instance. I have it running there. Now I can shut my laptop. It's still running? And then now how do you do this at a massive scale? And how do you do this auto scaling with respect to the amount of work that you have, right? If your EC2 instance, you hit it 10 ,000 times, and all of those things, you know, every invocation requires downloading some data, that machine is going to fail over quick.

34:18So what happens when you want 10 ,000 things to spin up just as many machines and you want to fan out, that would look like my flow.deploy. So I'm here my laptop. I've got my hello world flow. I would literally just say, if name equals main, right? Like I'm writing any other script, flow.deploy. And the sort of autocomplete in your IDE is going to tell you like, all right, what do you want to name this thing? Where do you want it to run? Do you want it to run on ECS, in Amazon? Do you want to run in Google Cloud Function? Where do you want it to run? Can you point us to, do you have this in a Docker container somewhere?

34:53If so, we'll just like pull it from whatever Docker registry that you have, or do you want to point us to a GitHub repository? And one of the nice things is when you build a tool like Prefect, that is treats failure as a first class citizen, you see the failure modes of tens of thousands of people trying to get their code to run. And so this really allows us to build kind of an ergonomic experience that's like, look, here's the minimal stuff that we need to know in order to have like a high reliability guarantee that this is going to work when you shut your laptop, when you submit it to a Kubernetes cluster, what have you.

35:26So the experience is really like pretty dead simple. It is, if it's running locally on your machine, you can hit.serve. Now it starts a process on your machine. Or you want it somewhere else, it's.deploy. And we can guide you through exactly how to point it at remote infrastructure. Does that answer your question, Daniel? I hope I'm not glossing over anything. It does. Yeah, yeah, that's great. I just have always felt this as a pain in my own life. So I often like to ask the selfish practical questions on this podcast when I get the chance for them. Yeah. One thing that you had said, which I did gloss over is like, you had said, what's the experience of deploying?

36:07And then like, what's, you know, tell me about like a shiny dashboard that you see or something like this. So once I go through all the work and I have my workflow running out in the world, I have this really beautiful UI that I can log into. And when I access that UI, what it displays for me is it basically tells me like it's almost like standing on the platform above the factory floor, right? You can see everything that's in progress. You can see the box of widgets that everything has produced at the end. And you can see everything that was broken in process. And so if you're the type of person that treats your workflows as cattle and not as pets, it's very easy to see like, yeah, you had 10 ,000 jobs that succeeded.

36:47Here's the 10 that failed. You can easily click in, see, OK, the 10 that failed. Why did they fail? If you're on Prefect Cloud, which has a very generous free tier you can run a business on, we do AI summaries of the errors. And so we will summarize like, look, these workflows are failing for these reasons in natural language so that you don't have to dig through like a 10 ,000 line stack trace just to find out that you had an out of memory error.

37:25well there's no shortage of helpful ai tools out there but using these ai tools means you gotta switch back and forth back and forth between yet one more tool so instead of simplifying your workflow it just gets more complicated but that's not how it works when you're using notion notion is the perfect place to organize lots of stuff tasks tracking your habits writing beautiful docs collaborating with your team knowledge bases and the more content you add to notion the more this cool thing called notion ai can personalize all of the responses for you unlike generic chatbots notion ai already has the context of your work plus it has multiple knowledge sources.

38:09It uses AI knowledge from GPT-4 and Cloud, and that helps you chat about any topic. And here's the kicker. Now in beta, Notion AI can search across Slack discussions, Google Docs, Sheets, Slides, and even more tools like GitHub and Jira. Those are coming soon. And unlike specialized tools or legacy suites that have you bouncing between different applications, notion is seamlessly integrated infinitely flexible and beautifully easy to use so you are empowered to do your most meaningful work inside notion from small teams to massive fortune 500 companies these teams both small and large use notion to send less email cancel more meetings save time searching for their work and they reduce spending on tools which helps everyone stay on the same page you can try notion for free today by going to notion.com slash practical ai that's all over case notion.com slash practical ai to try the powerful easy to use notion ai today and of course when you use our link you're supporting our show and i know you love that again notion.com slash practical ai

39:43So, Adam, I want to ask you a question. I understand that you have a friend named Marvin. I'm wondering if you can tell me a bit about Marvin. Yeah. So Marvin means a lot of things at Prefect. Prefect, if folks don't know, is an homage to Ford Prefect from Hitchhiker's Guide to the Galaxy and the genius alien that, well, I think Marvin's a bit more of the genius, but the alien at the beginning of the book that's just like, come on, let's get out of here, introduces him to the world or to the universe. I just reread that last week, by the way, just completely randomly there. I just had to say that.

40:21Amazing timing. Yeah. For folks listening, this is not like paid actors, I promise. And then Marvin is his like terribly depressed Android companion that's like just kind of has the curse of knowledge, knows everything. Everybody hates me. Exactly. And so Marvin started off as a mascot at Prefect. And we were first to ducks. So we had Marvin the Duck, which was this lesson of like, often the fastest way to triage failure is by rubber ducking. So we would send like ducks off to new customers to put on their desks so they could literally rubber duck with Marvin. And then Marvin was the name of our first internal like LLM powered Slack bot, which is still alive today.

41:09It serves a community of 30 ,000 data engineers that show up in our community Slack every day who have questions about Prefect. So, you know, if somebody is trying to configure particularly complex behavior, maybe that's a blind spot in our docs. Marvin is hooked up to all of our docs, all of our GitHub issues, everything that's ever been written about Prefect. Marvin is hooked up to. And so in Prefect today, if you're a community user, you can show up and you can say like, hey, Marvin, I saw this stack trace and I thought I'd configured this correctly. What's going on? And Marvin would be like, oh, well, actually, if you configure it like this, here, maybe you want to check out this doc or this video.

41:50And so Marvin has really created this personalized interface to learning prefect that docs never can be, right? Docs are always kind of shooting for like the middle, I don't know, 80 % of technical prowess, the middle 80 % of use cases. And LLMs, just as they've done for pretty much every other product, create a human interface to PETA entry, right? So it's somebody can show up as they are and get a personalized interface to our docs. The last piece is Marvin is our opinionated LLM framework. So Marvin is incredibly popular project. It's got 5 ,000 plus stars on GitHub. And it got started in early 2023, where we had just built our internal Slack bot and we had felt that like existing options in the world didn't allow us to write LLM workflows in a way that felt particularly natural or ergonomic to us.

42:49And as a company that more or less was like, look, we think that writing dynamic workflows and configuring infrastructure, those are hard problems, but we should take on creating a good ergonomic interface to them And so you don't have to face that complexity. Prefect is obsessed with giving people access to sort of the complexity of the world without having to face it themselves. They can opt into it whenever they want, but you shouldn't have to have like a CS degree just to get a script running on another machine. And similarly, you shouldn't have to like, you know, import a thousand agents and a thousand different data loaders.

43:27And you shouldn't have to learn like a new common expression language just to get LLMs to do the work that you want. And so Marvin is like, you've got a Pydantic model. Guess what? You decorate this thing. And now this thing that was responsible for validation can now, you put a document into it, and now it can extract data from it. You've got a typical Python enum. You decorate this thing. And now you've got a classifier. You've got a function. You write in a doc string that says, here's what this function is supposed to do. Now you can get strongly typed outputs because we're just reading from the function signature, its return annotation.

44:05We're basically trying to figure out how to make writing LLM workflows feel very Pythonic and very ergonomic. And it's been a blast to build and it's just welcomed so many AI engineers into basically our ecosystem who saw a really sane way of writing Pythonic LLM workflows that they just didn't have in any alternatives. And I guess if we take that maybe a small step beyond, I know that you also have control flow as another piece of this stack. Could you kind of highlight maybe the differences, overlaps kind of distinction there? I know it's really interesting, I think, what you said before in the sense that a lot of people are trying to build these agentic workflows.

44:52And a lot of those things are very flaky and hard to debug. Right. That's the main thing that I hear in the main, which is often why I think people are decomposing their agentic workflows into maybe static kind of workflows that aren't really agents anymore. more, but they can be debugged and, you know, they can check errors and that sort of thing. So yeah, I would love to hear about your kind of approach on the agentic side, given the sense that you're treating those sort of failure states as first class citizens, that sort of thing. Yeah, yeah. No, so I think one thing I want to talk about is like, what's the value of agentic workflows?

45:29I feel like a lot of folks just hop into it being like, cool, now I can make this thing agentic, but for what reason at all? So just like little context here is like, I truly believe that LLM workflows and just to make sure we have the same vocab, it is like very deterministic workflows that say, first, I'm going to call out, I'm going to extract this data. If I see the extracted data, now I'm going to classify this piece. And then if that now I'm going to, I don't know, send an email. Those are things where the human is really writing the logic. If we were to take this workflow and we were to go back to like 2016 or 2017, it would look the same.

46:08But instead of calling out to an API, we would say, now I'm going to have like some hidden Markov model that's going to go and do my entity extraction. That's my step one. If I detect this entity, now go to my next thing. Now I might have some like interpretable logistic regression model, try and classify the output coming from it before. So LLM workflows are really like, how do I take traditional workflows that existed in the world of ML, but now instead of having to go and train models, now I can essentially wish models into the world through natural language. And those workflows are amazing.

46:44That's like, frankly, I think that they solve most business problems that you need to. That's like a very reductionist, like broad brush that I want to paint with. And they're incredibly easy to debug. They're incredibly easy to observe. because the logic is something that you wrote. If your first if statement fails or something like this, now you're able to say like, well, I knew exactly where in my logic this thing failed. And so now I can go debug this tool that I called out to. It basically, it's easy to debug because it's not changing our mental model of how to build data pipelines. And so all of the tooling that we built in the last decade around like building robust evaluations, collecting data to try and figure out how often they're succeeding, these are tried and true.

47:29So we're able to basically like do our own bit of transfer learning from, you know, decades of ML research and evaluating ML pipelines onto LLM pipelines. Now, there was a paper that came out from Alpha Kodium a few months ago that was basically like, look, if you want to do, say, like build your own automated software engineer, there is only so far you can get with like, say, deterministic LLM workflow that's built on top of like GPT 3.5. They were able to show that, you know, even though it's all just kind of like function calling under a big while loop, right? They were able to show that if you adopt this paradigm where you start turning over some of the planning logic to the LLM itself, now it's able to outperform the like top of the line frontier models on specific tasks.

48:17So for things where the decision space is very large, like what piece of code do I want to run next? Stuff where the information space is very large. I've got a giant code base to reason about. Agentic workflows really tend to thrive because if the decision space isn't known to you at the very beginning, agents are able to discover that dynamically, create and refine on a plan that you as a human, it would take a lot of ifs to get to the same style of performance. So that difference between LLM workflows, which are basically prompt engineering, calling out, calling a function, single shot stuff for single shot tasks.

48:54That's what I call LLM workflows. The stuff where you turn over behavior or orchestration logic to an LLM itself to decide its own plans, maybe even create its own tools, that's agentic. And that's really that distinction in problem space is also how we distinguish between Marvin and control flow. I promised I would get to control flow. So Marvin is really good for saying, like, you want to call out to an LLM and extract something. We're going to make it feel the absolute most natural that you can. And we're going to make this so that if you know Python, you know Marvin. If you know Python, you know how to extract something super easily.

49:29If you know Python, you can classify something super easily, summarize things super easily. So Marvin is really a prompting library. It's a way of creating dynamic prompts using classic Python functions or Python objects. Control flow is really now when you start turning over the wheel to LLMs to formulate their own plans, this is where we can't rely on the tools of the past around observability. we can't rely on the tools of the past around evals you can still build evals that say like i put in this thing at the beginning i got this thing at the end how often did it succeed on my test set but if you're trying to figure out like how often did this thing take the specific path through reasoning space to get to this outcome that's uh now much harder to design tests around and it's harder to debug if something fails and you didn't write in the orchestration logic you You can't really point to a place and say, ah, this is where I messed up in how I designed this thing.

50:24And so control flow is really how do you ergonomically express agentic workflows? How do you express dependencies between tasks? How do you tell an LLM when it has the wheel? And then how do you explicitly tell it that it no longer has the wheel? And now it's going to cede back to human written logic. That is what we make it extremely easy for you to express in control flow. And since we have a pretty opinionated developer experience of how to express this, we're able to give folks much better observability. And so this is built on Prefect3, which means that you have an orchestration library sitting behind every single action that an LLM takes.

51:09If something fails, you've got retries. If something takes too long, you've got timeouts. If something fails and you don't want to go through all the work again, you've got caching. if your agent needs to create a sandboxed code environment to run untrusted code, it can now deploy its self-created function to remote infrastructure. So it's really how do you give agents durability and resiliency? Because those are often the biggest reasons that they fail. And so that's how we distinguish between it. But I realize I'm monologuing a little bit. So I'll turn it back over to you, Daniel, if there's some stuff that you want to double click on.

51:46It's all good. I'll jump in. We've covered some ground here, and I appreciate that very much. With Prefect Core producing the open source that you're building around, you guys have the managed workflow orchestration platform that is Prefect Cloud. We've talked about ControlFlow. We've talked about Marvin. That's a lot. What are you thinking about next as you're looking in the months and the next few years to come? you've accomplished so much, but as you look at what you think the space is going to turn into, because it's evolving so fast, do a little bit of crystal ball gazing for us. Right or wrong, tell us what you think the future is going to hold and how you'd like to play into that.

52:33Yeah. So it's a great question. And it's something I obviously have to try to think a lot about. I would say that right now, there is way too much emphasis on single machine, local LLM or agent workflows. Here's what I mean by that. It's like you pull up any framework in the world to run LLM things or to build an agentic workflow. And it's like happening in a local process on your machine. it spins up like an input field in your terminal, and then you have to like type your answer to what your favorite color is. And then it goes and writes you up like a poem or something like this. But businesses who are building LLM workflows or building on agents at the model level, we see that frontier models are accounting for, sorry, frontier model providers are accounting for a lot of the core sources of failure.

53:29So structured outputs from OpenAI managed to wipe out just a whole host of your classic resiliency issues. And the resiliency issues, I think, are going to be around planning. And so the ability to add item potency and transactions to LLM workflows is really what we would invest more into. You had a plan that executed 10 things in a row, you found out at the 11th step, the plan was bad, but you've already done 10 things. How do you walk that all back? You would do that with transactions. Where I am really interested in long-term is when Daniel asked earlier, like, talk to me about how I as a human go from like a locally running function to something that's running elsewhere.

54:14I think that the me as a human part of that is going to be, is to go away, right? And that it's often going to be an LLM in the course of solving a problem decides that it needs to create and massively parallelize a function to call out to. And so now you need to be able to give an LLM on demand the ability to create and provision infrastructure and submit a whole bunch of jobs to it. And so while we've built Prefect to be something so easy a human can understand, that's going to play into a lot of strengths of how do we expose an API that can really be taken advantage by an LLM, if that's the intended audience of how to provision infrastructure.

54:55The third piece is at companies now, you'll have so many teams that are basically building LLM workflows in parallel to each other. So you'll have like the conversation team that's like trying to build in LLMs into how it talks to customers. And then you'll have like the platform team, which is trying to use it to like give feedback on internal pull requests. You'll have teams that are trying to use a bit more of like, you know, in human in the loop for like commenting on like a design doc, something like this. But fundamentally you have tens of thousands of parallelized executions or calls against LLM APIs.

55:36So how do you solve that coordination problem across the programs that are trying to invoke LLMs. And so really trying to figure out if I have 10 ,000 agents all trying to access the same resource at the same time, how do you govern that in a way that doesn't cause them all to crash? If I have 10 ,000 engineers that are all trying to access the same LLM API, how do I make sure that they don't all consume the same token budget all at the same time? So that's a bit more on the practical side of like, you know, future of agents, future of LLMs, Aside, there is still a very, very interesting fundamental engineering question, which is how do you get tens of thousands of concurrent API calls to OpenAI or Anthropoc all behave sanely and not lean into a commons problem of everybody cannibalizing the same data resource?

56:31And so that's like the very boring thing that I like to think about it. But orchestration is one of these like very fun, but ultimately like worst case scenario disaster planning style things. Well, we love the boring stuff here, which is actually not boring. And I think a lot of people want to think more about because they do have the desire to put these workflows into production, right? which is, yeah, of course, on your all's mind and what you're digging into. And yeah, just really appreciate you taking time to join us, Adam. It was a great conversation and I would encourage people, we'll include the links in the show notes, but go check out the docs for Prefect and Marvin and Control Flow.

57:15Try out some of these things on your own. Like Adam said, you can pip install Prefect and be off to the races. So no reason not to try it out. And yeah, thank you so much, Adam, for taking time and joining us. My pleasure. Thanks so much, guys.

57:37All right. That is Practical AI for this week. Subscribe now. If you haven't already, head to practicalai.fm for all the ways. And join our free Slack team where you can hang out with Daniel, Chris, and the entire ChangeLog community. Sign up today at practicalai.fm slash community. Thanks again to our partners at fly.io, to our Beat Freaking Residence, Breakmaster Cylinder, and to you for listening. We appreciate you spending time with us. That's all for now. We'll talk to you again next time.

From the publisher

Workflow orchestration has always been a pain for data scientists, but this is exacerbated in these AI hype days by agentic workflows executing arbitrary (not pre-defined) workflows with a variety of failure modes. Adam from Prefect joins us to talk through their open source Python library for orchestration and visibility into python-based pipelines. Along the way, he introduces us to things like Marvin, their AI engineering framework, and ControlFlow, their agent workflow system.

Join the discussion

Changelog++ members save 9 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • WorkOS – A platform that gives developers a set of building blocks for quickly adding enterprise-ready features to their application. Add Single Sign-On (Okta, Azure, Google, Microsoft OAuth), sync users from any SCIM directory, HRIS integration, audit trails (SIEM), free magic link sign-in. WorkOS is designed for developers and offers a single, elegant interface that abstracts dozens of enterprise integrations. Learn more and get started at WorkOS.com
  • Shopify – Sign up for a $1/month trial period at shopify.com/practicalai
  • Notion – Notion is a place where any team can write, plan, organize, and rediscover the joy of play. It’s a workspace designed not just for making progress, but getting inspired. Notion is for everyone — whether you’re a Fortune 500 company or freelance designer, starting a new startup or a student juggling classes and clubs. 

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Practical workflow orchestrationPractical AI · 58 min
Listen in VO