Orkes and Agentic Workflow Orchestration with Viren Baraiya

2 Oct 2025 · 47 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Orkes and Agentic Workflow Orchestration with Viren Baraiya

Episode Overview In this episode of Software Engineering Daily, host Gregor Vann interviews Viren Baraiya, founder and CTO of Orkes, discussing the challenges of orchestrating modern software systems composed of microservices, and the solutions provided by Orkes' agentic workflow orchestration platform.

Key Concepts Discussed

  1. Modern Software Challenges
  2. Microservices Architecture: Modern applications consist of independent microservices including frontends, backends, APIs, and AI models.
  3. Orchestration Needs: Coordinating these microservices efficiently is complex and requires a reliable orchestration framework.
  1. Orkes Overview
  2. Enterprise-Scale Solution: Orkes builds on Netflix's open-source Conductor project to offer scalable, secure, and resilient orchestration for microservices.
  3. Key Features:
  4. Integrates AI agents, humans, and APIs.
  5. Focus on security, governance, and support for long-running workflows.
  1. Viren Baraiya's Background
  2. Career Path: Baraiya has experience at notable companies like Google (Firebase and Google Play) and Netflix where he contributed to the development of Conductor.
  1. Netflix Conductor
  2. Origins: Developed to manage the complexity of orchestrating numerous microservices at Netflix, using chaos engineering principles to ensure resilience.
  3. Open Source Transition: Initially developed at Netflix, Conductor was later open-sourced to foster community contributions.
  1. Workflow Orchestration
  2. Definition: Orchestration coordinates tasks across different systems, crucial for managing complex business processes.
  3. Traditional vs. Programmatic Approaches:
  4. Rule-Based Systems: Simplified workflows via predefined rules but limited flexibility.
  5. Agentic Orchestration: Allows developers to create workflows programmatically, empowering greater control and flexibility.
  1. Agentic Orchestration
  2. AI Agents: Defined as autonomous entities that can plan and execute tasks based on their programming.
  3. Multi-Agent Systems: Incorporating multiple agents (e.g., LLMs and humans) allows for complex decision-making processes within workflows.
  1. Trust and Safety in Automation
  2. Guardrails: Mechanisms to ensure safe operation by allowing human supervision before executing critical actions.
  3. Transparency: Detailed execution graphs provide visibility into decision-making processes, enhancing confidence in automated systems.
  1. Orkes Development and Future Directions
  2. Enterprise Readiness: Focused on compliance and security for regulated industries.
  3. Evolution of Agentic Systems: Aims to simplify complex workflows and improve the integration of AI into business processes.
  4. Future Features: Continuous investment in trust, safety, and collaborative tools for developers and business users.

Getting Started with Orkes

  • Developer Edition: Accessible at [developer.orkescloud.com](https://developer.orkescloud.com) for rapid onboarding.
  • High-Impact Use Cases:
  • API orchestration using pre-built system tasks.
  • Building conversational agents using language models.
  • Implementing business-process templates (e.g., order management).

Conclusion The dialogue between Gregor Vann and Viren Baraiya illuminates the critical role of orchestration in managing modern software architectures and the innovative solutions offered by Orkes to tackle these challenges effectively.

--- This summary encapsulates the critical discussions and insights from the podcast episode, emphasizing the significance of orchestration platforms in contemporary software engineering.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Modern software systems are composed of many independent microservices, spanning frontends, backends, APIs, and AI models, and AI models, and scaling and scaling and scaling. reliably is a constant challenge. A workflow orchestration platform addresses this by providing a structured framework to define, execute, and monitor complex workflows with resilience and clarity. Orcas is an enterprise-scale agentic orchestration platform that builds on the open-source conductor project, which was pioneered at Netflix. The platform coordinates AI agents, humans, and APIs, with a focus on scalability, compliance, and trust.

0:40It further expands on the conductor core by adding features like security, governance, and long-running workflows. V-Ran Barea is the founder and CTO at Orcas, and he's the creator of Netflix Conductor. V-Ran joins the show with Gregor Vann to talk about his building Conductor at Netflix, the challenge of orchestrating microservices, rule-based versus programmatic workflow orchestration, agentic orchestration, MCP integration, and much more. Gregor Vand is a security-focused technologist, having previously been a CTO across cybersecurity, cyber insurance, and general software engineering companies.

1:22He is based in Singapore and can be found via his profile at van.hk or on LinkedIn.

1:40Hello and welcome to Software Engineering Daily. My guest today is Varane Barea from Arcus. So we're going to be talking all about what Arcus does and especially Arcus Conductor. And some of you may already know the word conductor from some other companies and we're going to be talking about that. But yeah, welcome Varane. Thanks Gregor for having me here. And I'm very excited to be here. Yeah. So as we do on this podcast tradition, we just like to kind of get an understanding of your background. We've definitely worked at quite a few interesting companies that I think our audience will be fairly familiar with.

2:13But yeah, just walk us through what's your kind of career path, I guess, to Orcas. Yeah, I started Orcas about close to getting four years now. And prior to Orcas, I spent like in a very different set of industries. So before Orcus, I spent almost a few years at Google, mostly working on developer products, so Firebase and Google Play. This is one place where I got to kind of work with a developer as one of the audiences. How do they interact with systems and how do you build for them? Which was also kind of one of the reasons why, you know, which kind of motivated me to build Orcus. more interestingly before google i spent my time at netflix which is where i was part of an infrastructure team that was responsible for building the platform for netflix's back then very ambitious you know studio project right basically the goal was to build the largest production studio in the world and to enable that there is of course the whole production side but at the same time there was the engineering side in terms of building out the products and tools and my team was responsible for building out the platform.

3:20This is where Conductor originated amongst the many other things that we built. And interestingly, before I moved to Netflix, I was in a very different kind of industry, right? So I spent almost six years at Goldman Sachs working in the investment banking technology side of it. That was a totally different experience altogether. So I was in East Coast, moved to West Coast, all the way from investment banking to entertainment and then internet consumer and then now SaaS. Yeah. Yeah, that's very interesting. I don't think we have many guests who have actually been in like investment banking tech and then managed to move into, I guess, sort of Silicon Valley tech.

3:55So that's super interesting. And obviously Netflix, especially the time that you were there, which years you were there, that was such a pivotal time for that company. and speaking of Netflix so some of the audience may be familiar with Netflix Conductor and I think we should kind of talk about that to begin with so you were working on that part of the technology but let's just talk about like what was Netflix Conductor and what was the challenge it was solving and sort of what was that all about? Yeah so if you look at Netflix's history right Netflix has historically kind of been the pioneer when it comes to some of the frontier technology, right?

4:34Like they were one of the very first big tech company to be completely on cloud. They invested very heavily into microservices. And when I joined, one of the things that was very surprising was that it really had embraced microservices, right? So there were a lot of microservices and that enabled teams to kind of move very fast. At the same time, one of the biggest challenge was how do you essentially coordinate the work across different microservices, right? Because by definition, a microservice is not going to implement the entire business flow. It just implements part of it, which means that you have to stitch together multiple microservices.

5:09And the very kind of traditional way of doing that is through some sort of eventing system, right? Back in the day, there used to be message buses, service bus, and so forth, right? and given Netflix and its scale, we had our own internal implementation of the service bus, you know, pushing through, you know, billions of messages a day and that worked very well. However, like, you know, the major issue that you run into with something like that is not when things are working fine, but when things are not working fine. And this is where you start to see the kind of brittleness of the system. What happens when something is not working?

5:46Now to understand like exactly what's going on, you have to kind of either dig to the record, talk to 20 different people, manage, you know, hundreds of different queues and things like that, right? So that is where the motivation for Conductor started to say that we want to continue investing into the same principles of building through microservices. We want and we like this whole loosely kind of coupled aspect of building distributed systems. What we don't like is a coordinating directly through code and let's have an orchestrator do that thing for us at the same time we wanted their orchestrator to be working at netflix scale built using chaos engineering principles and that's how we started working on conductor and that's what conductor started to do right and the whole idea was you know let's tame all the microservices bring the order to the chaos without losing and sacrificing what it delivers at the end of the day right yeah and yeah for those that maybe aren't familiar chaos engineering is the sort of principle of just switching things off and not telling anyone and seeing what happens what sort of how our system fails and how it's responded to so i think it'll be interesting maybe to get into that maybe a bit more as we go through what orcas is doing and with netflix am i right in saying it wasn't originally open source but then it was open sourced at some stage is that correct yeah so when we started of course like you know we started building it.

7:09At one point in time, what we realized was that, you know, the product was good. There was a lot of adoption inside within the company. And what we saw was that like, you know, if we open source this, we get the benefit of community contributing to it. We can also leverage it as a way to kind of recruit good people, build kind of collateral for the team itself. And more importantly, Netflix has been always very active about like, you know, pushing things back to open source, especially things which are not proprietary, not key to the business, right? And Connector is a very generic piece of software.

7:43It's nothing to do with encoding or, you know, streaming or anything. So we decided to kind of make it open source. And since then, it has been open source for quite some time now. Yeah. Yeah. So we'll come back to kind of then the transition from it being kind of associated with Netflix and then sort of what happened there and then Orcus coming out as a company on its own, so to speak. But I think we should also just take a step back. You already mentioned orchestration there. So let's just kind of give a baseline of what is workflow orchestration? I think that's a great question. And the reason is that orchestration is a lot of things.

8:21In the end, the moment you start coordinating work across different systems, that's orchestration. But then you could have orchestration of containers, if you are talking about infrastructure. It could be messaging, it could be data pipelines, services in our case, for example. And in general, like if you look at the workflow orchestration and workflow engines, it's not a new concept, right? Like it's been around for very, very long time, decades probably, if not longer. And the primary reason for that is that when you look at the business processes, everything is a workflow. And one thing that I like to say is that whether you like it or not, whether you know it or not, but everybody is building a state machine one way or other.

9:01And the hardest part of building a system is not how to implement a particular business logic, but is how to maintain the state in a way that it remains consistent and coherent with the business goals. So if you can offload that entire responsibility to a workflow engine, then how you develop systems, how do you add resiliency and scale and everything else becomes much easier to kind of deal with. In the end, like if you think about it, right, if you have a serverless Lambda or a completely stateless service, if you want to scale up, you can just horizontally scale because it maintains no state.

9:35But the moment you add a state, now it becomes challenging, you know, because if there are failures, you have to recover from the state failures. When you are scaling, you have to ensure that like, you know, state can be managed in a distributed environment. So scaling stateful systems are much harder. And this is where like, you know, one of the things that we have also seen is developers end up spending a lot of time. So workflow engine kind of solves that problem. The right ones, especially, because you also don't want workflow engines to be single point of failure because if the workflow engine is down, then everything is down.

10:06So that itself has to be resilient. Yeah, absolutely. So yeah, we have this concept of like a workflow engine. And before we then get to pure conductor, there are kind of two sort of ways you could still go about a workflow engine. Is that right? And we've got like rule-based and conductor style. So maybe just walk us through what rule-based, I assume the rule-based came before conductor style and maybe people are more familiar with that, but let's just talk about that and maybe it's limitations and then we can sort of then go on to conductor. I think the rules for the rule-based workflow engines are in the fact that like, you know, if you are able to define a set of rules, how the process should be orchestrated and then like, you know, somebody who owns the business, you know, this could be a product manager or a business analyst, they can come and define the rules.

10:50which works well, I can know if that is how in practice things happen. But it also simplifies things because like now you can write rules based on your business specific needs. Where it starts to kind of break down is that one, you are constrained by the rules that you can write and underlying implementation, right? So if you were to make systems a bit more complicated, then that starts to become a challenging. The second is that like, you know, for simplest cases, that is completely okay. but as it gets more complex it becomes much harder to understand what is going on here and then you are constantly translating those rules into the code right to see you know how this is going to get executed so that kind of choreography is constantly happening in your mind when you are debugging things when you are trying to understand what's going on there so that becomes a bit of a challenge and the third thing that we realized was that oftentimes what happens is that in theory it looks great that like you know as a business owner you can come and define rules in practice they will write up a doc and a developer is supposed to you know then go and implement the rules and write the rules as a developer you deal with the code and not the rules so Conrad takes a different approach to say that like you know let's keep that idea of being able to define orchestration as a drag right a direct acyclic graph but instead of making it a very business specific or rule-based system make it programmatic so you know just like you write code you should be able to write a workflow and it should follow the same principles and the same kind of semantics like a code in terms of being able to run things in parallel decision cases loops so as a developer the way you would think about it you can just write the workflow as it is and the fundamental principle and the thumb rule that we follow was that if you can write code you should be able to define a workflow it should be one-to-one there should be no missing cases there and it should follow the exact same principles as variables states and things like that essentially making your code completely durable that has worked very well that has worked very well because one as a developer when you are implementing it you are not constantly fighting against a different type of system you are doing exactly how you think about it when you're debugging it becomes pretty clear as to what you're debugging and then to give the business owners the visibility you can always transpose that into a business specific set of dashboards so you know it keeps everybody happy so yeah because i mean when you mentioned or when you talked about role-based it sounded like the focus was more for like non-technical people ultimately and conductor style is is where actually the developer has a lot more control over that process that is correct yeah yeah awesome so let's kind of I guess look at how Orcas kind of came to be I believe there was something around Netflix decided to stop supporting or stop contributing I guess to Conductor as an open source project but I think there was a group of you that weren't actually working at Netflix by that stage but then sort of came back together so what was that story?

13:45Yeah so when we started Orcas we weren't at Netflix right like I had led Netflix for about four years back then I was at Google and when we started Orcas, one of the things that we did was we started working with Netflix. So, you know, we started contributing back to Conductor. At some point in time, we started to become one of the larger contributors compared to Netflix. And it just made sense that like, you know, let's take it out of Netflix umbrella and put it into its own project repository, right? Giving community more control over the project and increasing the velocity. Because in the end, the motivation for maintenance of the open source project between Netflix and an open source community and us is going to be very different.

14:23And we have a much bigger team in terms of being able to support and take this forward. So that's where we kind of work with Netflix to kind of have them kind of archive the repository. We took over the code base and we did it in such a way that like, you know, we are not losing all the previous contributions, contributor history and everything. So it remains that way. And that has been kind of the case and it remains completely compatible because it's the same source code in the end. And we see that as an evolution of the open source project, right? Many of them kind of go through the similar case where they start at a particular company like kafka is a good example right going from linkedin to apache and then mostly shepherded by a concluant same little bit of databricks and spark yeah exactly we've seen this sort of as you say across a bunch of open source projects that sort of yeah usually come out of of a company and then yeah there's just various reasons why it makes sense to well either to just kind of fully open source and say hey we're just not going to support this anymore and whoever wants to come in and work on it.

15:18So that makes a lot of sense. Let's then talk about Orcus. And if you could explain what Orcus is, and maybe let's start with like, what is it ultimately, what has it built on top of Conductor and the Conductor project? Yeah. So when we started Orcus, right, the primary motivation was that we have the open source project, there is a very clear fit between the project and the market. Because, you know, we had seen by then like you know thousands of companies some of the very well-known companies as well using them in their productions flows personally i was getting a lot of things on linkedin sometimes people asking me to review their pr or asking for some help and in the end we felt that the market was ready typically what tends to happen is that you know companies like netflix they are on the bleeding edge of the innovation so the problems that they see and they solve the industry starts to see them three four years later so in a way kind of the timing was very correct that like no this is the time where companies are going to start looking at it and realizing that as you move your systems to cloud you need to break down your monoliths at the same time cloud yes it gives you the elasticity of infrastructure but at the same time unlike data center you have to start thinking about resiliency aspect also and this is where they will start thinking about workflow engines and we were kind of right in that sense also So, you know, that's how we started kind of the Orcas as a company.

16:40And in terms of the business model and like how we go about it by this time around other open source companies had already paved the way in terms of, you know, how should you think about building an open source project and monetize that? How do you kind of differentiate and things like that, right? So that kind of become the foundation for how we think about Orcas as a company and the conductor as an open source project that we kind of monetize on. and then bringing back to your other question is like you know how does it differentiate right so what we have been doing is that you have the open source project we use the open source so in the end orcas is the open core right where i think enterprises wanted us to be able to support them was in terms of adding enterprise features because if you want to run a project inside your company let's say a bank or a healthcare company you need to have things around security compliance governance those things open source i would like to think about it more like a linux kernel right you can take a kernel and build your own distribution but as a company you probably want to get a distribution that is vetted by a vendor and has all the security features and everything so that is exactly kind of way we did it the other part of conductor is that connector is a very plug-and-play system so you know it supports multiple different backends and what we do is you know we take the right ones we optimize for the performance and the cost and everything and the manageability also on top of it.

18:03And that's essentially what we deliver, right? What our customers are essentially paying us for is in the end, how do we take the project and run it reliably in their environment? Because that's the key challenge that we solve for them. If you were to give them three nines or four nines of availability, that's one thing that we can deliver them without them having to worry about it. Yeah. So yeah, again, kind of walking quite a, well, I think now we have seen it's quite a familiar path where an open source project can just benefit hugely actually if there is a sort of commercial arm around it where as you say it's able to provide the security the stability the compliance especially you know in the enterprise setting which is what orcas really caters to having all that just kind of taken care of as you say there's a huge need often for that on the basis that original project has got a bunch of people using it and yeah i track back to sort of it was interesting sort of in hacker news at the time there was a lot of comments people say oh this is awesome it's so great to see that some people are taking this on as like a proper company we can just kind of buy it from them as opposed to even try and run it ourself now so yeah yeah yeah as a matter of fact like our first few customers were the open source adopters was like you know glad that you guys started the company can you help us so yeah and especially because quite a few of you were very much the the core contributors in in the first place so like that's awesome to see so let's move on to Orcus maybe walk us through so how has Orcus evolved because I think we're talking sort of back in 2022 ish is kind of when that started what we've just been talking about so maybe just walk us through like how has Orcus evolved we're going to get into agentic orchestration and AI because that's where it can help in a big way in things that are very pertinent now but maybe just walk us through kind of has the product kind of evolved yeah yeah so I think when we started our initial focus was that hey here's open source how do we make it enterprise ready run it in cloud so we kind of focused on that one right and of course our customers tend to be mostly in very regulated industry quite a few of them right which means that like you know you have to support different modalities in terms of whether it's running fully hosted by Orcas or is it like bring your own cloud or in some cases running in data centers so that is one area where we spend time and making sure that like, no, we can take the software, we can run it at a highly reliable scale for the customers.

20:27And then we started to kind of start thinking about as a company, if you are leveraging something like Conductor, they don't want different tools for different problems. When you think about workflow orchestration, you want to be able to do a number of things with it, right? So we started to kind of add some of the features that we got as clear feedback from our customers. Sometimes they had kind of built out their own internal versions of it but they wanted us to kind of support them by adding it as a proper feature inside conductor so you know some of the things that we did was like workflow engines traditionally are asynchronous orchestration meaning the workflows can run anywhere from few minutes to hours to days we added support so that like now your workflows can run for much longer period of time as well like months and months or even years for some cases and we do have some use cases like that.

21:15And then on the other extreme was if you are orchestrating services, you have HTTP services, you have gRPC services, and you want to orchestrate them, those are going to not run for seconds. They are going to probably finish the entire flow in tens of milliseconds. So how can we run a workflow synchronously, right? So true microservices orchestration, but very much synchronous and very low latency. So that's another area that we focused on. And I think that's one of our key capabilities that is very unique to conductor that you typically don't find in other workflow engines and then as kind of the industry was starting to think about ai and llms that's where we started to kind of invest into how can we let workflow engines orchestrate language models i mean today i think that has become like a very common place that like you know you need a workflow engine to you know orchestrate your agents but you know back in the day people were still writing python code to just call ellms and that's where we started to kind of build integration suites and everything right like we core to our nature right like in the end we are not a solution we are a platform which means we want people to be able to kind of use it in whatever ways and format they want to use it so one area where we focus on is that like you know let's start to integrate and provide support for pretty much every foundational model that is out there And today, I think we support pretty much every possible model out there.

22:39You can switch back and forth. You can run them together in the same workflow and things like that. So that's one of the areas where Orcas has evolved into a true LLM orchestration platform, right? If you have multiple agentic models. And that then allows you to kind of build. If you think about traditional workflows, those are deterministic flows, right? You could have switch cases which could take different paths, but in the end, it is still very deterministic. given the right input, it is always going to produce exact same output path. Now, if you add language models and LLMs inside that, you start to kind of see the non-determinism aspect of the workflow because even for the same input, it could take a different path.

23:19And then we started to kind of support those things. And I think that's one area where in general industry also is moving towards and we are continuing to kind of, you know, invest into that area as well. Yes. I mean, maybe again, just to sort of baseline this, I'm sure majority of the audience are kind of familiar with what AI agent is, or sort of what they maybe think it is. But at the same time, I think it's always helpful to kind of get your definition as well, because I think you could probably pull up five different sort of definitions of what an AI agent is, and especially in the orchestration sense.

23:52I mean, for example, are we talking when we say agentic orchestration, are we saying, well, these are multiple agents that get orchestrated or are we saying that ultimately an orchestration could be termed as an ai agent or you can help me out here i think yeah that's a good question like you know i think agent is a very confusing term because it's a very general purpose thing right pretty much anything can be thought of as an agent but in the end like i think the textbook definition of agent is that agent is something which is an agency it has its own autonomy in terms of how can it plan and execute its goals.

24:25And now if you translate that into agentic systems, it means, I think, three different things, right, in my opinion. An agent essentially could be purely an orchestration where you have language models deciding the path. There has to be some sense of autonomy. Otherwise, you just have a very deterministic system. So agents by definition has some level of autonomy and therefore or a non-determinism kind of built into it. Now, you can think about a workflow with a single language model or an LLM that is either running in a loop or in a single execution path. In that case, you have a single agent that is operating inside that workflow.

25:05Now, we have heard a lot about humans in the loop and guardrails, right? As humans, you can think about humans also as an agent. So the moment you put a human inside a workflow with an LLM, you are starting to think about multi-agent systems where now you have two agents and they have very clear responsibilities. Maybe the LLM has a responsibility to come up with a plan and human has a responsibility to kind of vet that plan or approve or reject the plan and then continue executing on that one. Similarly, you could add more LLMs and build true multi-agent systems where LLMs are participating and each one has a pretty well-defined role.

25:40I think a very good example, I would say, is what we see with AI coding tools, right, like Cursor and Windsurf is, you could think about an agent, one of the agents which takes your instructions, generates the code. A second agent could be actually responsible for compiling, and third one could be responsible for testing against and checking against your input goals, right, and they are all coordinating, running in a loop until it achieves the goal. So that's a true multi-agent system in the end, and as a human, as a developer, you are also an agent who is kind of saying, yeah, this looks good, approved it commit the code so that's now the true multi-agent system but in the end agents are if you think about a heuristic workflow right the way i would like to think about is that you don't have a very set defined part but you you have a very high level definition of this is how you should do will you do it or not it depends upon like you know how lms are thinking about doing it yeah i think that's really helpful and the code orchestration example is a good one i think also through the sort of Orcus website and sort of there's like examples of flows and I think this is an example that sometimes you pull out around inventory management or like claims management for example maybe that would be quite interesting to sort of understand now how does it differ compared to say like a traditional rule-based system like what are the things that can be done differently and better I guess when we're now talking about AI agentic orchestrated and then apply to these kind of quite clear, you know, business use cases?

27:09I think, see, the biggest thing that I think that can be done better is if you have a non-agentic system, and because by definition, it is a very deterministic system, every time you have a different use case, you have to build a new workflow, a new system around it, which essentially creates an explosion of different use cases, and which is what you see is that, hey, if I were to approve a claim, for example, right, and depending upon different requirements, you have different claim systems or different parts of the claims and things like that. But if you were to add something new, again, like, you know, you go back to development mode, rebuild or build a new feature and it takes time and things like that.

27:46With agentic system, I think the biggest change is that instead of writing the entire system end to end, you focus on writing tools. A tool can be something that sends an email. A tool can be something that looks at the claim information and pulls up the customer information or the claimant information or looks up the policy. Now, if you think about, right, you can put those tools in any particular combination. So now we are talking about combinatorial explosion, right? If you were to build deterministic systems, you end up building a large number of different use cases and paths, which is why most of the software projects takes months and months to develop, because you have to cater to all different possibilities and everything.

28:27but if you break it down to say that like i have got m tools it can be used in any combination and an llm can decide which one to use now you're thinking changes right like you're no longer thinking about putting them together by yourself you are building stateless tools very similar to microservices if you think about it but instead of as a developer you kind of putting them together and llm is taking your input and deciding on the fly how should i do this which means that like you know your development process becomes simplified you can introduce a new tool without having to change everything and start incorporating them so you know the way it differs from traditional rule-based systems is that it now allows you to go from zero to one and one to n very quickly by you know just incorporating more and more tools but you are no longer catering to kind of the combinatorial explosion of different use cases right you can just do it out of the box and we are starting to see that right that's how i would say a lot of new systems are starting to build out is through agent.

29:25Agent can do pretty much anything as long as they have the right tools and context given to them. Yeah, I mean, I think to sort of use a slightly overused term is sort of this idea of basically setting the first principles of what can be done. Yes. And then letting the orchestration aspect kind of then deal with how it wants to then go about that. And you touched on it, obviously the deterministic or non-deterministic, especially in this case, aspect. And I think that's something probably a lot of the audience is curious about is how does that then work? You know, because that's basically the kind of the crux of all this is how do we allow the system to take its own decisions and what sort of constitutes this, say, a first principle that can be laid down and then the rest is allowed.

30:11Yeah, how does Orkis like deal with this and how does somebody using Orkis, I guess, how can they kind of feel confident that the non-deterministic aspect is kind of taken care of, I guess? I think the analogy that I like to think about it is when you have a car that can do self-driving, there are two aspects of it, right? One is the notion of control that like, you know, I can take on the steering wheel at any point in time and do whatever, right? So, you know, guardrails, humans who can be in the loop wherever you need to be. Second part is, which is, I think, more critical. Like if you think about it, right?

30:47Like if you just read RLM as a black box and say, here are the tools, just go and do it and come back with a result. How do you know what was the thought process there and what did it do? So second part is basically showing me what it sees, saying, this is what I'm thinking. This is my plan. And this is exactly how the graph of this execution is going to look like. So as a human, now I can look at it and say, this makes sense that you're going to execute step one, two. And based on the output of step two, I can take three or three prime and then go and execute step number four. So now this graph is something that I can see and say, this is what you are thinking about doing it.

31:24This makes sense for me that, you know, you should do it this way and go and do it. So, you know, humans in the loop becomes critical along with that entire aspect of being able to visualize the execution graph. I think that's tremendous because now you start to build confidence that this works. The second part is when to apply guardrails. A good example that I like to give here is if I'm building a DevOps system and I'm using an agent to manage my Kubernetes clusters, when it decides to execute an operation to get the list of pods and deployment, yeah, nothing bad is going to happen. So just do it.

31:59Even if you execute that command on a wrong cluster or a production cluster, I mean, you are just going to execute a read operation, nothing bad is going to happen, and that's completely fine. but if you are going to destroy a cluster you better check with me first maybe you send me a slack message or an email and let me approve it because you know you might hallucinate you might end up taking wrong decision or a typo and destroy a production cluster so i don't want you to do it so when to apply a guardrail is another aspect this is where we we are spending a lot of time to say as a builder of the agent you should have full control so instead of saying you know here is the llm you give them the tools and let it execute everything our approach is fundamentally different in the sense that here is the llm we give llm saying here are the tools that you can use tell me what you are going to use and then based on the outcome i can decide and build that inside my workflow so now the workflow becomes a combination of some set of algorithms and some set of non-determinism right you add determinism when you need and otherwise let non-determinism and take care of everything else.

33:00Yeah, and I believe exactly that Orcus is really focused on sort of this trust aspect, you know, because I think that is what everyone is, I say everyone, but you know, especially enterprise is sort of concerned about like the potential productivity gains around allowing agents to run a bunch of stuff is in theory fantastic. It is just kind of that, well, that developer example you just gave of, you know, a cluster being destroyed or being some typo somewhere. That's, I think, what a lot of, especially I would say the non-maybe technical folk in companies are very concerned about. They're sort of like, this all sounds great, but there's no way that this can actually do it reliably.

33:40So can you maybe just walk us through maybe a few mechanisms or sort of how does Orcus, or maybe I don't expect it's all kind of solved today, but like how is Orcus actually approaching this and like what kind of tools and mechanisms are there to help the developer and then like, What could that then help the developer say to the non-business person to sort of help them feel more at ease about all of this? The way we are approaching is, as I was kind of trying to explain why, like in two ways. One is being able to add guardrails and be able to add guardrails when you think this operation is going to be something that you want someone to take a look at it.

Read the full transcript

34:17And guardrail doesn't have to be necessarily a human, right? There are a lot of systems for, you know, automated guardrails. You can also use agent as a guardrail. So, you know, you can delegate it to another agent. which can get tricky because what if that also hallucinates and two of them agrees and does something bad. But depending upon the use case, depending upon the contextual need, we allow developers to put the right guardrails and adding guardrails is a deterministic step. We are not asking LLM to decide when to use guardrail because that then brings back a cyclic dependency, right? And trust aspect.

34:50Instead of that, we let developers to add and say that this is where you will add a guardrail and that's a very very deterministic step that you know if you have a guardrail set up for a specific tool it will get executed so that is one part of it that takes away the whole aspect of llm doing something bad without your approval the second part is understanding what really happened so it's completely possible that the llms did exactly what it was supposed to do no hallucinations and everything but then there are questions about why and that's important right in terms of, let's say, if it's a claim processing system, and if I approved or denied a claim, and if there's a question as to why, the answer cannot be that, you know, because my AI said so.

35:32It has to be that, hey, this is the thought process. This is how we evaluated the claim. And therefore it is. So it's less about LLMs making decisions, it's more about LLM defining the flow, but you need to have the complete visibility. So other aspect that we give is that we give the full blown graph of exactly what happened, step by step. Every step, what was the input given what was the output that came out of it and exactly what was the decision made based on that so now as an operations person or a human you can look at it and explain exactly what happened and why it happened so you know that takes away the other aspect of is as to i can't explain what happened now we can completely explain what happened you can control and these two things combined we think kind of gives you enough guardrails and of course like one other aspect of conductor and workflow engines in general is that it keeps trail of everything So every conversation, every execution that happened is captured, stored, and can be kept in the storage for whatever is your retention policy.

36:27So if you were to go back and see what was happening, how things were happening, those things can be later carried on. The other part is access control. One thing that is built into Orcas is just because you have a tool does not mean anybody can use it. So even to use a tool, you need to have the right access control. so which means a good example i'd like to give here is that if you are building an agent that can do all stuff hr for you you should not be able to you know ask agent to give yourself a promotion unless you are an hr admin and it goes through proper approval process right so that is also built into orcas so you know with the right level of access control visibility and the human guardrails i think we think that that's going to be enough for someone to say hey we can trust the system yeah and yeah you've mentioned it good explanation the graph sort of aspect of it does it present kind of the same and let's just mix this in with the access control for a second does it present the same to sort of across all types of person or is it you're able to present different kind of views that make the most sense of the explanation is this an ops person who's going to understand what happened in this way which is quite different at times to a developer who wants to kind of see it in a slightly different way.

37:42But maybe this has been solved in one pane of glass. Again, to use a slightly overused phrase, but yeah, tell us about that. Yeah, I think the short answer is yes. Slightly longer answer is that as a developer, you are able to see every step. But depending upon how you construct the whole thing as an ops person, either you can look at the high level blocks saying, you know, step one, step two, step three. Step two could be a lot more complex, which as an ops person, you may not need to understand and know. And of course, because one thing about Conutter is that it's pretty much API-driven system.

38:11You can then go and build very business-specific views of it, which might make sense for your business users and follows your kind of process flows and definitions and everything around it. So it kind of decouples those two aspects. As a developer, you don't have to think about building everything with a rule for the business. And as a business user, you don't have to think about, I don't understand this. Somebody come and explain to me. Yeah. So bringing this kind of back and forward, I guess, to where the developer sits, It's one bit that we haven't touched on before we kind of get on to sort of just getting up and running, so to speak.

38:42But one thing we haven't touched on is actually how MCP and MCP servers come into this. And I believe you have open sourced your MCP server for Conductor. I think this was the thing when I was sort of getting my head around what Orcus and Conductor does in the first place, I instantly started to think about MCP because I thought, well, isn't this what MCP is sort of for? So maybe talk to us about where the intersection is and how they kind of work together. Yeah. So MCP focuses primarily on how do you expose your tooling and API, something that LLMs can understand. And that simplifies a great deal in terms of LLMs being able to call the tools.

39:22And what we do is that we allow developers to bring their own MCP servers. We are also working to kind of bring in most of the common ones as a part of the out-of-the-box capabilities inside our enterprise edition. And then it can basically, you know, use them as tools. So, you know, if you want to send an email and decide, LM decide that, no, I need to notify the user. And if you have an integration through MCP via Outlook or Twilio, it can send you an email. So that's the primary role for MCP. The conductor MCP server does a very similar thing, but it acts as a tool to generate the execution graphs, to say that, hey, I have this goal.

39:59Can you give me a conductor workflow for this that I can then go and execute? So that's like a stepping stone for us to build a fully autonomous systems. Because one area where if you think about today, right, like LLMs, essentially what they do is like in the programming terminology, they do a look ahead of one. They look at the current context and say, what's the next set of tools that I'm going to execute? Where we are going with is I can look at the goal and say, I can define the entire execution graph with a look ahead of, you know, N, N being a recently finite number. so that improves both performance the cost aspect of also because you know you are making less llm calls and more importantly reliability because that output can be pretty much deterministic or deterministic enough for multiple iterations so that's the primary goal of the mcp server got it and i mean it is effectively completely optional it's not sort of yeah it's not sort of required in terms of and i mean are there i mean mcp is obviously it's a protocol that was ultimately developed by Anthropic.

40:57Are you looking to do any support for any of the other competing, shall we say, protocols or sort of does MCP make sense as the one to kind of sit with? I think MCP is a great one for being able to call the tools. We added some features that like are kind of gaps or like, you know, some of the things that MCP as a protocol definition lacks, things like access control, right? It does now support a notion of authentication, but the odd z is other part that we have added then the other one is a2a when you start thinking about multi-agent protocols i think a2a is coming out to be something that people are starting to think about as agent coordination so that's another area where we are going to add support pretty soon nice awesome so let's just sort of talk about i guess sort of up and running first of all where does developer kind of go and then maybe could you just talk us through what is the sort of high impact say first 10 minutes of getting started with orcus i believe there's like some kind of template type workflows you can kind of run out the box yeah so what's that kind of yeah what's like a high impact 10 minute place for someone who's never used this or and let's just say for argument's sake they've never even used orchestration before this is the first time they're actually approaching this like what does that look like so i would say there are three main categories right one is like you know if you are looking to orchestrate apis you can create a workflow, a conductor has notions of system tasks, so things which are pre-built, you don't have to write code for it.

42:26Even if you write the code, you're going to do the same thing. So you can orchestrate multiple HTTP endpoints and see for yourself. There are a lot of example API endpoints available on the internet. You can just put them together and see how it orchestrates them and gives you the visibility. So that's the API orchestration use case that you can very quickly test it out. Second part is if you are trying to build an agent, you can try and build out a simple chat complete agent. Like, you know, you can put a loop and chat complete inside it. And, you know, it will keep on running until your loop terminates.

42:56And you can actually put two agents, like you can take two chat complete, two agents and, you know, give them some instructions and you will see that, you know, they start talking to each other in a conversational way, right? That's pretty fun to see, you know, and quite interesting sometimes. And the third part is if you are building a workflow, you can take an existing business process that you have, like, let's say, order management or claim processing, right? And we have templates for it. You can try it out, mock up the actual implementation and see for yourself, you know, how easy is it to like change, modify, get the visibility into it.

43:28But I would say those are some of the things that can be done in the next, like, you know, in 10 minutes. And we have a developer edition. So, you know, So anybody can go to developer.orcuscloud.com and get started pretty quickly without having to worry about how do I download run locally? And which if you want to do it, you can always do it. But nothing beats like, you know, one click, go to this URL and start working on it. Yeah, exactly. I think that's kind of where I went. So that's developer.orcuscloud.com Orcus spelled O-R-K-E-S. So yeah, head there. Yeah, it's kind of pretty foolproof. You could just either choose templates or you could just hit start from scratch, sign up and then off you go so awesome so from what you can share you kind of touched on like where you might go you say agent to agent things but like from what you can share just before we wrap up like what does the next say like six six months look like for Arcus and like what are you kind of looking to add or develop as well I think I would say that the industry is slowly moving towards like agentic workflows right everyone is thinking about how can they incorporate language models into their business processes and leverage them to kind of accelerate the pace at which they can innovate, get ahead of the curve.

44:40And that's one area where we are focusing on it. And most importantly, as I said, like trust and safety aspect is the most important one. That's what enterprises care about more than anything else. And that's one area where we are like spending a lot of effort and see how can we simplify those things. And the other part that is coming up pretty quickly is when you start thinking about agentic systems, traditionally, when you think about software it was like as a developers you will build end-to-end stack that role might shift towards as a developer you will build tools and the agents will be built by the business users going back to our original discussion about rule-based for flow engine side i think they're coming back and i would say this time with a vengeance saying hey we are going to let you do it but now no more dsls no more quirky you know rule engine but rather just describe what you want to do and i'll figure it out and do it for you so i think that's going to be a pretty interesting area to see and that's one area where we are also investing to see you know how can we bring business and developers together to you know accelerate the speed at which they can innovate for companies you know yeah yeah awesome sounds really powerful i mean especially as you've called out given that orcas is super focused on enterprise enterprise grade reliability and trust in this sense this is kind of where if it's going to be possible to do it in an enterprise setting then this is kind of the place to come to and obviously it's you know it's been it's been proven as of base level given it came out of places like netflix and sort of used in big settings big companies so yeah very exciting so yeah well thanks so much for coming on varane i think we've learned a lot and yeah again just for anyone who wants to just kind of get up and running that's developer.orkiscloud.com just head there and give it a try so yeah varane thank you so much i hope we get to catch up again in the future.

46:29Yeah, thank you. Thanks for having me here.

From the publisher

Modern software systems are composed of many independent microservices spanning frontends, backends, APIs, and AI models, and coordinating and scaling them reliably is a constant challenge. A workflow orchestration platform addresses this by providing a structured framework to define, execute, and monitor complex workflows with resilience and clarity. Orkes is an enterprise-scale agentic orchestration platform

The post Orkes and Agentic Workflow Orchestration with Viren Baraiya appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Orkes and Agentic Workflow Orchestration with Viren BaraiyaSoftware Engineering Daily · 47 min
Listen in VO