The era of durable execution (Interview)

10 Apr 2025 · 1 h 40 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The Changelog Podcast - Episode Summary

Episode Title

The Era of Durable Execution (Interview with Stephan Ewen)

Podcast Description The Changelog is a podcast focused on software development, particularly open source projects. It provides weekly news briefs, technical interviews, and discussions with industry leaders and builders.

Episode Overview In this episode, hosts Jared and Adam interview Stephan Ewen, Founder and CEO of Restate.dev. They discuss the concept of resilient applications and delve into the technical aspects of implementing durable execution functions, idempotency, and the significance of stateful durable execution in software development.

---

Key Concepts

  1. Resilient Applications
  2. Definition: Applications that can withstand errors, network failures, and other hiccups without losing functionality.
  3. Examples:
  4. Avoiding duplicate orders.
  5. Ensuring chat history is retained.
  6. Handling temporary outages effectively.
  1. Idempotency
  2. Definition: The property of an operation whereby performing it multiple times yields the same result as performing it once.
  3. Importance: Crucial for preventing double execution of transactions (e.g., charging a user twice for the same order).
  4. Implementation: Use unique identifiers in requests to verify whether a similar operation has already been completed.
  1. Durable Execution Functions
  2. Definition: Functions that can maintain state across executions and recover from failures without losing progress.
  3. Benefits:
  4. Resilience in application workflows.
  5. Ability to recover from temporary failures seamlessly.

---

Discussion Highlights

Importance of Durable Execution

  • Stephan emphasizes that the current state of backend development for applications with complex state management is unsustainable. He argues for a better solution to manage distributed systems effectively, leading to the creation of Restate.

The Technical Journey of Restate

  • Background: Prior to creating Restate, Stephan and his team were involved with Apache Flink, a stream processing framework. They recognized an industry need for better tools in managing distributed transactions and stateful coordination.
  • Development Philosophy: Restate was built from the ground up, focusing on low latency and high throughput without the complexity of traditional database management.

Observability and User Experience

  • Observability is a key feature of Restate, providing insights into function executions and errors through an integrated SQL query engine.
  • The goal is to deliver a user experience that simplifies the development process for engineers, allowing them to focus on building resilient systems without the overhead of managing complex states.

---

Use Cases and Practical Implementation

  • Example: The discussion includes a hypothetical implementation for a podcast episode management workflow using Restate, where video uploads trigger a series of transcriptions and metadata processing steps that benefit from durable execution.

Comparison with Competitors

  • Stephan discusses competitors like Temporal and NATS, highlighting Restate's unique approach to combining durable execution with service orchestration and state management.

---

Conclusion The episode concludes with an emphasis on the growing importance of resilient applications in the evolving digital landscape, particularly as more organizations move towards distributed architectures. Stephan encourages developers to consider Restate as a robust solution for building durable execution functions.

---

Key Takeaways

  • Durable execution enhances application resilience.
  • Idempotency is critical for transaction safety.
  • Restate offers a streamlined approach to managing stateful functions and workflows.
  • Observability features enable better monitoring and debugging of applications.

Additional Resources

  • For more details, visit [Restate.dev](https://restate.dev) for documentation and further exploration of durable execution concepts.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:07Okay, friends, it's time for your favorite podcast. Welcome to the Change Law, where we feature the hackers, the leaders, and those who are building durable execution functions. Today, Jared and I are joined by Stefan Ewan, the founder and CEO of Restate. Talking about the coming era of resilient applications, the meaning of, and what it takes to achieve item potency, this world of stateful, durable execution functions, and when it makes sense to reach for this tech. A massive thank you to our friends and our partners over at Fly.io. That is the home of changelaw.com. Learn more at Fly.io. Okay, let's get resilient.

1:00Well, friends, before the show, I'm here with my good friend, David Hsu, over at Retool. Now, David, I've known about Retool for a very long time. You've been working with us for many, many years. And speaking of many, many years, Brex is one of your oldest customers. You've been in business almost seven years. I think they've been a customer of yours for almost all those seven years to my knowledge. But share the story. What do you do for Brex? How does Brex leverage Retool? And why have they stayed with you all these years? So what's really interesting about Brex is that they are an extremely operational heavy company.

1:32And so for them, the quality of the internal tools is so important because you can imagine they have to deal with fraud. They have to deal with underwriting. They have to deal with so many problems, basically. They have a giant team internally basically just using internal tools day in and day out. And so they have a very high bar for internal tools. And when they first started, we were in the same YC batch, actually. We're both at Winter 17. And they were, yeah, I think maybe customer number five or something like that for us. I think DoorDash was a little bit before them, but they were pretty early.

2:00And the problem they had was they had so many internal tools they needed to go and build, but not enough time or engineers to go build all of them. And even if they did have the time or engineers, they wanted their engineers focused on building external facing software because that is what would drive the business forward. Brex mobile app, for example, is awesome. The Brex website, for example, is awesome. The Brex expense flow, all really great external facing software. So they wanted their engineers focused on that as opposed to building internal CRUD UIs. And so that's why they came to us. And it was honestly a wonderful partnership.

2:34It has been for seven, eight years now. Today, I think Brex has probably around 1 ,000 Retool apps they use in production, I want to say every week, which is awesome. And their whole business effectively runs now on Retool. And we are so, so privileged to be a part of their journey. And to me, I think what's really cool about all this is that we've managed to allow them to move so fast. So whether it's launching new product lines, whether it's responding to customers faster, whatever it is, if they need an app for that, they can get an app for it in a day, which is a lot better than, you know, in six months or a year, for example, having to schlep through spreadsheets, etc.

3:09So I'm really, really proud of our partnership with Brex. Okay, Retool is the best way to build, maintain, and deploy internal software, seamlessly connected databases, build with elegant components, and customize with code. accelerate mundane tasks, and free up time for the work that really matters for you and your team. Learn more at retool.com. Start for free, book a demo. Again, retool.com.

4:06We are joined today by Stefan Ewan from restate.dev. Stefan, welcome to the ChangeLog. Hey, thanks for having me. It's a pleasure. Pleasure to have you as well. Adam, how are you doing, man? So good, Jared. How about you? I am doing well. Always excited at the beginning of a conversation to dig into something new, something different, and something called Restate. This is supposed to be the simplest way to build resilient applications. This is a requested show, Stefan, so we do take episode requests. This listener would like to remain anonymous. However, they say that Restate is a super exciting approach to managing distributed systems.

4:55And they say that we should get you on the show. And so we just take orders around here and our listeners often give what they want. And so that's how we found you. What's the real listener request? Awesome. That's very cool to hear. Is that open source communities at work and all that? That's right. So Restate. let's not get into restate itself at first let's talk about resilient apps first because this is called tagline the simplest way to build resilient applications let's talk about that what is exactly a resilient application in your in your estimation okay yeah so in in the in the way we think of it in the context of restate um we're talking mostly about the the back ends of application the sort of coordination and orchestration logic.

5:44A resilient application would be an application that doesn't accidentally drop your order, that doesn't accidentally place it twice if you hit F5 at the wrong point. When you're in the browser, that doesn't, you know, accidentally book your Uber for two people instead of for one. that doesn't, you know, just disconnect you from your chatbot, lose the history, make you start over, all these kind of things. This is what we call, like what we're talking about when we mean resilient apps in the context of restate. So basically apps that tolerate all sorts of hiccups, errors in the infrastructure, unavailable endpoints, you know, network failures, process failures, temporary outages, but also like you know types of programming glitches that cause like requests to fall through and having to be retried in order to you know be processed reliably but then not get duplicated but the system understanding how to either potently treat them as a retry and not as a second request this is sort of the the bigger picture of what we mean here with resulting applications Can you take a moment to demystify that term that you just said, idempotently?

7:01So everyone's on the same page. What does that mean, idempotently? Yeah, I think idempotently, you can think of it as just like understanding that a repeated request is actually not a new request, but it's the same request again. You're just sending it again because you were maybe disconnected from the original request or you got an error back. It's the thing that if you don't do it correctly, that is actually accidentally placing an order twice when it did try to place it once. It's the thing that many applications don't get right, and that's why you still see too many websites saying, don't hit F5 while this is showing.

7:43Do not reload while we finish this transaction. Exactly. Because it doesn't understand if you submit the thing again. This is actually supposed to be the same thing. It's just like another submission of the same request. It has no way of identifying that, so it might accidentally treat it as a second one, or it has just like a very rough way of doing this. It's a surprisingly complicated problem, and there's lots of applications that don't get it correct, and there are lots of weird ways applications work around this. As a fun fact, I think the first bank where I was a customer, they did only allow you to wire to a certain recipient a certain amount once per day.

8:25If you're trying to wire the same amount of the same person a second time for the day, they just wouldn't allow it because they didn't know if that was a retry from your browser or if that was something stuck in a queue at some point in time. That's just how they duplicated it. And yeah, so generally, out of potency is deduplication in a meaningful way, understanding retries versus new requests. Yeah, that's called overkill, I think, when they did that. They're like, how can we just use a blunt tool to solve this tiny little problem? Let's just not let you do more than one per day. Not ideal, surely, for lots of uses.

9:02Okay, so same request twice, operates the first time, won't operate the second time. generally speaking how do you achieve this i mean you could just limit to one request per day but if you're not going to do that how do people usually implement or ensure item potency in their applications yeah so i'm in a way you basically have to find a way to anchor the identity of requests all the way through there's different standards for doing that um the like the http standard has actually a new has it has to find a header where you can put in all the potency key in the same servers when they support this are supposed to understand if that key is set and a previous request with the same key has come in the same parameters that this is that this is a you know a duplicate request for the same operation but then you know down the road you basically just try to to to anchor requests and and the processing of different steps in each other when you do message queues you place correlation ids when you're working with databases maybe you try to address primary keys or transaction ids or uh or leases or tokens this there's tons and tons of tricks people do but it's ultimately still very hard uh hard problem to do um if you want to do it end to end it's kind of a mindset you have to set out to do it don't you?

10:31Otherwise, I think you have designed from this, from the start for it. Like if the weakest link in the chain that breaks, it's one of these things. Yes. Or you can just say, do not reload your page while you're hitting this API. Exactly. It's the simplest way to do it. Just like make it the user's problem. Don't handle it in your infrastructure. Yeah. Also as an engineer, when I've achieved item potency, I know it feels very good when you're like, okay, I'm for sure not going to do double execution in this particular code path, that feels good. And as an end user, I'm also happy when I know that I'm not going to get charged twice, for instance.

11:05That's right. So you can actually see good APIs design this in from the start. I think the first API where I came across that where I thought it was really, really well done was Stripe's payment API. You can really see. I think that's also why it got so insanely popular so quick because the way they made that handling just seamless for folks that embedded this into their code to really understand how do I make sure I sent this request once and deal with all these things. That was really stellar. All right, so here's a harder question. Where does the word idempotency come from and why do we use that to describe this thing?

11:46It seems unnecessarily verbose and jargony. Do you know? You should spill it first, Jared. i-d-e-m-p-o-t-e-n-c-y would be the adjective do you know if you don't know you can just say i have no idea i don't really know but i'm my guess is it's a latin word it comes from latin it doesn't sound like yeah it's um i don't put it i don't know it's all greek to me maybe we can get a real-time look up for that and follow up on it from some sort of LLM. Just prompting Adam behind the scenes to prompt his name or LLM. Wikipedia as we speak. Okay, good. I was stalling for you. You know, I don't really have the details here for you, Jared.

12:29I'm sorry I can't LLM quickly enough for you. But it says item potence is the property of certain operations in mathematics. All right. I went straight to the LLM and I got the answer to my question. So the term item potent comes from Latin roots. So Stefan, excellent call there. Item meaning the same and potent meaning having power or being able to. Put together, item potent roughly means having the power to remain the same. So there's the actual word. And then, yes, mathematics, blah, blah, blah. I've stopped reading now. So hopefully that wasn't a hallucination and we can all move on. That sounds about right.

13:11Like if it's an hallucination, it's a good one. Yeah. Yeah. I D E M plus potence, same plus power. There you go. It's the same power. Very cool. Well, uh, thank you for scratching my itch, my curious itch there, both of you and chat GPT, I suppose pitched in on that one. Um, what else? So we're talking resiliency. I'm curious, obviously to have a resilient app just as like good, right? Like who wouldn't want these things? and just taking one of those things, idempotency, and realizing it's hard to achieve on your own throughout especially more complicated applications. Why did you all set out to solve this problem for folks to build restate, and why did you feel like I'm the guy for the job?

14:00Yeah, that has a long answer and a short answer. I can start with the short answer. Okay. The short answer is it's really, I would say, the state of the art to build backends that are supposed to, backends that do any non-trivial state management and coordination is completely unsustainable, I think, though, the way we're building this today. Just to give you an example, so let's actually start to example. Let's stay with LLM, because we just talked about it, right? So let's say you're building a chatbot, you're submitting something like a message there this thing in the end has to it has to reach the LLM but it has to look up the context in which that chat happened before it has to make the call has to go back store the context you don't want it to like just lose everything if you know if you lose your connection in the middle of you let's go with that five thing again like you don't want it to actually trigger the the same the same request tries or you know like lose the entire session make you start over so you're you're probably just putting this as an asynchronous request that runs in the background that you're, you know, you're sending from your chat session from your browser, but it's a separate, like, asynchronous request that runs that talks to the LLM.

15:15You want it to be actually retrying in case something fails or is overloaded and it's throttling you back, and then you want to be able to reconnect to that task or request in case, you know, something goes wrong or in your browser or you accidentally hit the back button or whatever. Like, just implementing this is a surprisingly complicated thing where you start to stitch together like probably a queue, a database, and a bunch of tasks to manage that. To give you another example, we just talked about Stripe, right? So let's say you're sending a request there for a payment, and sometimes they tell you, look, this is good or bad, like we accepted or didn't.

15:50Sometimes they tell you, I don't really know, like off-road detector is still running, or we have some weird thing in the background that, you know, we're still asking and it hasn't told us. So I'm going to send you a webhook in a moment to tell you whether this went through or not. And now you have like a synchronous request there. And then somewhere else an asynchronous request coming up. You just want to make those two reliably meet. Even if this one fails, you want it to sort of like recover somewhere, understand where to, you know, reconnect with the webhook that you're awaiting. And like this little piece, like it's really just one case handling in the back end where Stripe says, okay, I'm processing instead of like yes or no.

16:28This is actually many days of work to make that look reliable, to make that work reliable. And it's like lots and lots and lots of things like this, that just like get in the way with so many moving pieces, so many APIs to talk to, so much work, so and so much more work than originally happening asynchronously in separate requests than just in the synchronous user interaction. Like just gluing this all together has become such a complicated thing that we felt this does need a better solution. This is like the motivation more from the, let's say, from the use case side. I can give you a motivation more from the, like, why we actually ended up doing this.

17:10I think this is a motivation that probably lots of folks stumbled across ultimately and said, like, okay, this needs a better solution. I think there's, like, different projects approaching that problem. Why are we approaching it the way we're doing this has to do with, like, where we come from? Before we worked on Restate, we were building Apache Flink. It's a different system. It's a stream processing framework. It's basically events and analytics. So, you know, you have these events coming in, often through a message queue, and you want to, you know, aggregate them, join them. Just like a few examples where this is used is fraud detection in banks.

17:51Like some payment events go in, you aggregate feature vectors, some that through a fraud model or um i think you know things like the tick tock recommend recommend look flink to i use flink to actually um join information uh from users and interactions together in real time and and understand like how to how to create how to update how to update the features that will go into the into the recommendation model um if you i think i think companies like uber use it for like determine pricing and traffic models in eta so it's whenever you have events and you want to analyze them in a way that you aggregate them into some sort of typically statistical value or a materialized view this is what we were building before so it's an analytical framework what did actually happen then is at some point in time we saw folks were using that thing to solve so distributed transaction processing the types of things that where you would say, hey, let's assume an order processing service that takes the event, check out that order, and it has to do a bunch of steps.

18:59Let's say update the inventory, trigger payment, call the service to prepare logistics, maybe call another service to put this in the user's history, maybe more steps and so on. And we started to see folks using Flink for that because it had this interesting property that it had sort of this baked-in way of reliable communication and state management. It was all built for analytical use cases, but they found this such an interesting property that they started to apply this to transactional use cases as well, like auto-processing, just because they found that this is otherwise way too complicated to build and way too easy to build it in such a way that it is brittle and not scalable, that it has corner cases where it violates a lot of these properties that were just set.

19:48And when this started happening repeatedly, we thought, okay, but apparently there isn't really a good tool out there yet. And apparently this property of like correct, stateful coordination is something people really appreciate. They feel like it makes their life easier to build these type of frameworks. And then we set out to build a solution for that, and that became Restate. It's in many ways actually from the way it approaches things, from its architecture, it's inspired by our work on Apache Flink. But it's almost a complete mirror image implementation of it. It takes almost the opposite design choice in most aspects because it's really optimized for load-adgency transactional processing rather than high-throughput analytical processing, which was Flink.

20:37but but what we retained from from this idea is yes uh stateful orchestration an event-driven foundation and so on this is this is something we should we should build and we should be working on and yeah and that became restate what's the timing all this like when did flink start how mature was it or is it and when did restate be born out of the idea like what's give us a context in time Yeah, so this context is measured in decades, I think. Okay, that's useful. So, yeah, I mean, Flink was officially founded in 2014, but the work that became Flink is back to 2010 when I was still in university.

21:21So it was like four years of academic work, and then Flink started in 2014. I worked on it until 2022, so eight years after it became an open source project. And then I left Flink because I needed to work on something else, like for a change. And then we started working on Restate. Like end of 22, early 23. So Restate is a bit over two years old now. Flink has had its 10th anniversary last year. Flink is super mature. I think it's used by thousands and thousands of companies at absolutely insane scale. The probably largest installation of Flink that I know must be Alibaba, who runs tens of thousands of cores for like a single processing pipeline to like live compute their e -commerce search and recommender continuously.

22:24Reset is not quite as old and quite as mature as Flink. Yeah, two years. It's been two years. but we've released our stable 1.0 last summer and we recently released our first distributed, replicated, highly available architecture and have a bunch of folks that actually productively use that. So I would say we're on course to get there. Nice. Was leaving Flink difficult, sad, joyous? Like what was that like when you left? Because that's a long time to work on one thing. and then to move on to that thing. I mean, you become pretty attached, don't you? Yeah, yeah, absolutely. So it is definitely a difficult thing.

23:08And it was not just leaving Fling. So we created Fling as an open source project, but we also built a company around that. And the company went through an acquisition, but we actually still stayed there, started building and growing the team. So I was simultaneously leaving the open source project and the company I built and everything. And absolutely, it's a difficult thing because that becomes like your baby. It's actually two babies in a way, right? The project and the company. But yes, I feel that at some point I felt like there's this, I don't know how to say this in English, but I felt I was getting this like tunnel vision on problems.

23:47I've been working on this for so long. And I've seen so many sort of repeated things that like whenever I had a problem, I was just like putting it into this category that I knew. And that like, it started happening that I did that with things. And then later realized like, oh no, that wasn't actually the right thing for that problem. That might've been for the last nine, but this one should actually have been different. And so you get this kind of, if you work too long on the same thing, you're starting to not see the forest for all the trees. And I felt I was reaching that point. So it sounded like, yeah, this, I should probably start doing something new.

24:21It's kind of like familiar grooves in your brain. If you just go to the same thing over and over again, it's like your brain gets new grooves and then different problems fall into those familiar grooves and they just slide in there. Sometimes when they don't even fit or they aren't a good fit. I certainly understand that. I think when you've focused on one thing for a very long time, it is hard to think outside that groove. and it's interesting that you're building a similar system spiritually similar but i guess you could say radically different architecture is that the way you described it like it's inspired by flank but it seems like it takes the opposite approach of the world yes i think you could you could call it like that it is and it is under the hood still an event-driven system still as as flank is but it it just does it just built for completely different traders just to give you an example the the core of link is is the way it is this exactly one stateful processing so it basically has the status streams that it keeps moving operations that you know do stuff stateful stuff with events count joint and so on and then there's this asynchronous process running in the background that takes these consistent snapshots so if something goes down you can restore the state from the snapshot and sort of like just start start the flow from there and has this like kind of clever way to do this um in a way that maintains consistency across all the parallel machines and it is like efficiently incrementally frequently and so on but it's a very throughput optimized thing so it stays off off the critical path if you wish like runs in the background so it's a yeah it's really good for throughput but it you know it does this like persistence operation once every couple of seconds if you know if you to turn it to be like very very frequent most people actually run it more in the order of minutes right so when when something goes down and that in a pipeline you just like replay the last minute of data which typically doesn't quite take a minute because replies faster than the than the original uh the data rate at which the events are produced but still like it takes you back a certain amount in time which is usually okay for analytics the worst thing that happens is like okay you know like this this feature here in that vector that goes into that traffic model or that recommendation like is maybe a few seconds older than it would have otherwise been it's like not such a big deal typically they're going to transactional uh processing side imagine you have this multi-step process you want to like you want a really fast checkout process and you say you know i do want to to just start the next step um after i know the payment has gone through like before that i'm not updating inventory, I'm not kicking off any of the other processes, then you really need actually like a persistent step to be recorded, ideally in milliseconds.

27:13Maybe it's not that critical for audio processing, but we're building this for like even more low latency use cases, like payment processing and settlement and so on. And there you really just want to kick off the next step after, you know, like the previous step is persisted, like possibly in a multi-data center replication way and only then do i start the next thing so you have to really design this completely differently it's it's it's completely optimized for low latency transactional durability rather than analytical throughput so it's yeah it's a complete completely different design even though both ultimately are event-driven architectures right but your your atomic unit now is like is compute right like it's logic and perhaps data that comes from that it's a transactional step and so you can't just skip a transactional step because there's a workflow here and certain things rely on other things and so the way that you think about durability as opposed to analytical data is like i said earlier radically different that makes sense to me yeah exactly the atomic step in in restate is is extremely fine-grained right like we we're really building this in such a way that you should feel comfortable in a program that you do to use to use restate to persist fine-grained steps, state updates.

28:35It uses internally, actually, this durable mechanism for leader election to understand that it can lock and fence off different retries. And this is such a fine-grained nature. What's really important that recording a durable step has the lowest possible latency. Whereas in Flink, the atomic step is like a couple of million events being aggregated together in some state distributed over 10 machines. So that's one atomic step. It's completely different, yes.

29:16Well, friends, I'm here with a good old friend of mine, Terrence Lee. cloud native architect at Heroku. So Terrence, the next gen of Heroku called Fur is coming soon. What can you say about the next generation for Heroku? Fur represents the next decade of Heroku. You know, Cedar lasted for 14 years and more. Still going. And Heroku has this history of using trees to represent ushering in new technology stacks and foundations for the platform. And so like Cedar before, which we've had for over a decade, We're thinking about fur in the same way. So if you're familiar with fur trees at all, Douglas Furs, they're known for their stability and resilience.

29:55And that's what you want for the foundation of a platform that you're going to trust your business on top of. We've used Stacks to kind of usher in this new technology. And what that means for fur is we're replatforming on top of open standards. A lot has changed over the last decade. Things like container images and OCI and Kubernetes and CloudNave, all these things have happened in this space. And instead of being on our own island, we're embracing those technologies and standards that we help popularize and pulling them into our technology stack. And so that means you as a customer don't have to kind of pick or choose.

30:28So as an example, on Cedar today, we produce a proprietary tarball called Slugs. That's how you run your apps. That's how we pack to them. On Fur, we're just going to use OCI images, right? So that means that tools like Docker are part of this ecosystem that you get to use. So with our Cloud-nated BuildPacks, you can build your app locally with a tool called Pack and then run it inside Docker. And that's the same kind of basic technology stuff we're gonna be running in FIR. So you can run them in your platform as well. So we're providing this access to tools and things that people, developers are already using and extensibility on the platform that you haven't had before.

31:00But this sounds like a lot of change, right? And so what isn't changing? And what isn't changing is the Faroku you know and love. That's about focusing on apps and on infrastructure and focusing on developer productivity. And so you're still going to have that get push Heroku main experience. You're still going to be able to connect your applications and pipelines up to GitHub, have that Heroku flow. We're still about abstracting out the infrastructure from underneath you and allowing you as an app developer to focus on developer productivity. Well, the next generation of Heroku is coming soon.

31:29I hope you're excited because I know a lot of us, me included, have a massive love and place in our heart for Heroku. And this next generation of Heroku sounds very promising. To learn more, go to heroku.com slash changelogpodcast and get excited about what's to come for Heroku. Once again, heroku.com slash changelogpodcast.

31:55This word durability is being used a lot. Durable execution, durability. What exactly is durability? Doesn't fall down, doesn't break, always good? Yeah, something like that. I think durability is probably the same as persistence, maybe with a bit of a stronger emphasis on this really doesn't get lost after it happens. So durability is the D in asset when it comes to databases. Databases say we're giving you atomicity, consistency, isolation and durability. Once you do an update, we're not going to lose it. No matter what crashes, the database has a mechanism to bring that change to the database back.

Read the full transcript

32:37if I told you I've recorded that row, I've recorded that change, it will be there no matter what. And in the context of restate, that doesn't mean, for example, if you – the core building block of restate is a stateful durable function. You can think of it like that. And a stateful durable function, when you schedule an invocation for that or as you go through the code of that stateful durable function, has like multiple steps, recording a step. Whenever you go beyond a step that you asked Restate to treat as durable, you know that no matter what happens, you will never re-execute that step. You'll never come up with a different value.

33:21Like if your machine goes down, the reset server goes down, if you deploy it across availability zones, the data center goes down, the network gets partitioned, whatever, You'll never, ever go back and re-execute that step if it once told you that it's done. That's sort of the meaning of durability. Once it says it's there, it's always going to be there. And I think this is, in a way, I'd say almost like one of the magic ingredients. The way Reset looks at making distributed application development simple, I'd say there's two core pieces that you need to think about. One of them is the durability.

34:09Make durability extremely fine grained and extremely cheap. Because if you can apply durability in fine grained steps, you always have to worry about very little after failure. Let's say your durability is coarse grained. Let's say the order workflow is just like one durable step, right? And it crashes in the middle. It gets retried. It's up to you to figure out, well, did I actually process the payment already or not? Maybe there is a way to just like assume, okay, it's idempotent. I can send it again. Or I might even not be able to ask the service, did I do that or not? Did I actually decrement the available kind of product already or not?

34:45Maybe I have a way to, again, make this durable or not. I don't know. These things tend to be harder than one thinks because sometimes the API gives you, you know, it might have given you an error back the first time and you thought I didn't do this and followed some control path flow. And then the next time you actually get not an error, but the real result. And then you follow a different path. So people mess up this all the time. It's really hard to reconcile if you have these multiple steps as a coarse atomic unit. What did I do? How did I do it the last time? How do I recover from this? But if you have extremely fine-grained durability, if you're recording every individual step as durable in the system, and when it comes back, it can tell you exactly like this was the last step that you record it then you just have a very small amount of uncertainty okay here's this one thing that I might have tried already I have to just worry about that bit instead of the whole history and and possible control shown all the choices how I might have ended up here that I need to reconstruct in order to proceed consistently from there so just like very fine good durability is extremely powerful and it's simply fine things I'd say the second magic ingredients is then how do you anchor this in the whole retrying and, you know, resolving potentially inconsistent situations with partitions, with timeouts, with zombie processes and so on, so that there's always a very consistent view of what the last durable step was.

36:04I think that's the second sort of ingredient of restate. It's not just durability, it's actually durability and consensus and giving you a very, very crystal clear view on what, like where you left off, where you need to continue from. I think if you take those two things in conceptually, you've simplified the problem massively, and the rest is almost API sugar that you built on top of that. That's the, I would say, the magic that happens in the restate runtime. It's a very low-latency, durable consensus log that fuses queuing, state management, locking, fencing, creating futures resolving futures like all these scanner operations that tend to be part of a distributed coordination process and yeah when you say the restate runtime what can you liken that to for those of us who don't know what a restate runtime is yeah like is it like a node js thing is it like a database is it like a like what is that yes um so using reset is a bit like I would say somewhere in between using a database and using a message broker.

37:15So you write your program pretty much as code, however, how you write it before. But you're using the reset SDK. Think of it a bit like your database driver in order to sort of wrap certain operations as, okay, this operation here should be recorded as a durable step. or attach this state to the invocation transactionally or create this future, complete this future, and so on. So you do these operations through the Reset SDK. Reset itself is then like maybe message broker is the best comparison. It's on the level of the message broker. So when you invoke your code, you're not calling it to that directly.

38:00You're actually calling it to Reset, which makes the invocation of your function on behalf of you. The programming model that we try to provide is you're writing a service that looks like an RPC service, like you're writing handlers, RPC handlers. And then Reset almost looks like a reverse proxy for you. So the other services, instead of calling the code directly, they call it indirectly through Reset, Reset proxy in the call. And it puts itself in the middle with its durable consensus log. and when it forwards the request to the service it just isn't forwarded naively as an HTTP request but it actually uses a like an invocation protocol it uses like HTTP 2 or another type of like streaming connection holds on to that connection allows the service to sort of synchronize fine-grained steps it will it will when it forwards the an invocation for example tell it exactly what the supposed state of the world should be as in here's the steps I know you should treat just completed here are the ones that here's where you should continue and it will then allow to use that connection that that sort of lifeline to to let the let the application you know create durable actions so the yeah it's like on the level of a broker or database looks like a reverse proxy to the invoker looks like a maybe almost like a database to the to the service that uses it And when would somebody reach for this?

39:24Now you said distributed systems, but some people think every network attached system is a distributed system. So, I mean, if I'm building a web application, let's call it a monolith that answers HTTP requests and has a database backend, a Ruby on Rails or a Phoenix or a insert your Django, insert your backend framework here. are those folks pulling in restate and using it for certain aspects of their workflows or is that not necessary for them because they are kind of a monolith like do i have to be building a services-based architecture like where does it fit in i am i would very much be with you on like almost any system we build is a distributed system yeah completely so so i think it becomes useful very very quickly um maybe one way to think about this is your your back end where you do where you maintain state and update it and run operations and changes.

40:21It usually has a database that has the sort of core business state. Some operations coaches like purely straight to the database. That's all they do. That's fine. But any time you have to do something that's not straight against this, like your core database, but something that goes against like different APIs, something that runs in the background. Yeah, something that is asynchronous work that goes beyond just touching the database. I think you already are at the point where it's starting to become useful. Then, you know, if the only thing you're doing is maybe forwarding one call, yes, maybe it's overkill.

40:58But I think the usefulness starts much, much sooner than lots of folks realize. I would say every time you think about pulling in a message queue, you should probably start to think about pulling in something like Restate. Because it gives you a way to do the things you're probably trying to do with a message queue, but in a more high level, in a more well-defined, concrete way. You're not treating events, but you're dealing with stateful, durable invocation, stateful, durable functions all of a sudden, which is very often what you really want. If you're putting something off the synchronous path with a queue, you very often want to say, okay, here's something where I really care about that this happens.

41:52It shouldn't get lost, right? That's why I'm putting it in a queue. and then you you probably care about this this thing happening once having having reliable retries you quickly reach the state where the the processing of this operation is actually multiple steps and then you're again in the okay how do i do reconciliation of multiple steps if it failed somewhere in the middle and i don't know what i already completed or not so i would say the moment you you started pulling a message queue you probably should think about something like like we said the point comes very quickly yeah but then it's it's that's the that's the simplest use case i would say the most the most complicated ones that we see people build with this right now is using this to replace complex a complex choreography of like multiple kafka topics and rabbit mqqs and session servers and workers and so on um or even like a distributed sort of payment ledger keeping system so the it's really a very broad spectrum the i would say in a in a way you could think like all the type of work you do in the back end that's not the central database that keeps your business state i think is ultimately it's ultimately where we sit comes in yeah a little about a scenario where it's publishing i'm thinking like tiktok or youtube for example as a creator we will upload videos to youtube there's a process that happens there's a certain orchestration that happens it has to be compressed it has to go through certain filters maybe there's even a content filter has to go through a copyright filter is that an example of where you would use something like restate where you want it to go you want the user to be able to upload properly and your server capture the data and you know all the good things but then you got to run it through a process of saying okay this is now content that can be seen by what we call the world because it's been blessed by the copyright filter, et cetera, et cetera.

43:55Is that a scenario where it makes sense? Yeah, absolutely. This is basically a workflow, again, if you think about it, right? You're uploading the video. Let's say maybe the upload first puts it into some cloud storage. But then, as you said, you first pass it to the content filter. Then you have maybe a few steps that even run in parallel, like recoding it for different resolutions, optimizing it to be served through the CDN, and so on. Then you're, I don't know, running it through a system that tries to figure out what's the best sort of title frame to display and like all these different steps that you do.

44:27And they take potentially a long time. So it's a long running process. There's a fair chance that the container goes down and the metal wants to be migrated. And when it comes back up, you really want this to understand where did I leave off? Like what are the processes I should reconnect to that are doing the encoding or the analysis? Like this is exactly the orchestration of that process is where we said would come in. You wouldn't feed the video frames of the system. That's like overkill. You don't need to feed the video frames through transactional log. Like it's just that you put them in whatever cloud storage or so, but the, like the orchestration of the process of the workflow of the pipeline that does that, that's, that's a very good to reset use case actually.

45:04Yes. That's a great example, Adam, because it definitely makes it easy to think through. I guess as people who upload to YouTube, we are intimately familiar with all the different steps. That's why I enumerated very well. And it is asynchronous because you can go about doing the other things while It's working on, you know, the long running tasks, for instance, and somebody coded up some nice orchestration behind that sucker to keep that thing running. Yeah. And, uh, Google has the engineers to code up reliable orchestration flows, even in a way that, you know, they're nicely observable. You can, you can reconnect, uh, to them.

45:40They, you know, they're efficient. They know how to parallelize step and synchronize steps and so on. it's a much harder thing to do for many companies who don't hire the same type of engineers as google does and i think i think for those reasons it actually makes these type of things much more achievable than if you try to to embark on that yeah on that journey without even when you were sharing how it worked early on you were saying that it seemed at least from my perspective it seemed like it was user born every time every new application every new scenario every new job every new, you know, what have you, maybe even in your boring scenario where you were sort of focused for a bit there, you know, you keep recreating this, uh, durable invocation world over and over and over again.

46:29And why not turning into like you have done here with a server and a client and SDKs for different languages and a flow that every developer can grab. Is that kind of where what landed you to this point here was that frustration of the repetition and repeating and rebuilding every single time you build an application? I think that's a great way of putting it, yes. I would say the number one alternative to restate that people do or use is roll your own, absolutely. And it's a very repetitive process. And most of the time, I would say folks don't realize really all the edge cases that existed, what they do.

47:08they just maybe don't even solve them so the this this is basically half baked roll your own and it's yeah every time again and again and it's it's very often it's very often very similar problems that you're solving like let's say you're taking a message queue to say an action that I triggered should run asynchronously and it should you know chat have reliability retries and then I'm pulling in another store like redis or other key value store to to record different steps then i might be pulling in something like zookeeper etcd to place a lock on certain operations so no no they don't happen concurrently like i shouldn't be updating certain i don't know should a retry shouldn't work on the same payment id um if if the lock is still being held by the original processes so and then you're trying to sort of going back to idempotency trying to make an update to another system and understand how do how do you actually anchor the idea of that processing forward into the update to that other system like and you're recreating exactly you're recreating that type of pattern over and over and over again and i think this is where in a way workflow engines were originally born if you wish like enterprise workflow engines which try to say okay let's try to define a flow where we can have steps following, you know, a certain predefined control flow graph.

48:35And we have the workflow engine giving you the guarantees that step B, that follows step A, really only starts after step A is done. And step A is transactionally persisted before B starts and so on. They tend to be extremely heavyweight and flexible. Yeah, it's just not – they break all the tools and everything when you want to interact with them. And what Reset does and the durable execution space, what they're trying to do in general, is kind of bring this level of guarantees in a very, very lightweight way into almost arbitrary programs because it's just such a useful power to have to kind of define these durable steps, especially if you don't have to branch out in a different domain-specific language or graphical way to define them.

49:24But if you could just like write your regular code, but have it treated, have it executed with the same sort of guarantees as if it was an enterprise workflow. It seems like you might have a large education challenge in front of you because there's so much thought that has to go into this kind of architecture. I think the fact that your largest competitor is roll your own means most people don't know. Like, like we kind of all discover this pattern slowly over time inside of our own daily work. And so I'm just curious, like how you think you can attack that or are there other people like, is there a common thread or movement that you could attach to or create in which people are like, yeah, here's this new style.

50:11I thought of this because you mentioned like workflow. And I think that's sort of in the wheelhouse or in the ballpark of what restate is. uh message queue i mean what is there like a simple idea or concept pattern of which restate could be one or maybe restate is the the brand but have you thought through this because you have a marketing problem here or a challenge i should call it yeah so um i think i think that is that is very true in many ways the okay i think there's like lots of layers of answers to that I would say in the simplest way you can actually explain it reasonably simple if you just start from, you know, it's stateful, durable functions, which have guarantees that they execute, they run to the end, they are able to record steps, they're able to basically do these sort of asynchronous building blocks that you have in your usual programs, calling other functions, creating promises, resolving them, updating state, making calls, and so on.

51:14just in a fine-grained, persistent way that knows how to recover. This is sort of the basic building block, a stateful durable function. Now, the harder thing is actually, in a way, making people realize that they should be using something like this rather than roll your own. It's not uncommon that it's mostly on the junior side of engineers that you talk to them, they say, like, I don't get it. Like, I know how to write a retry loop. Like, what is this? And then, you know, it's a journey from there. the interestingly i think the most enthusiastic audience is often the the engineers that have been burned before that that know okay i know how to build distributed systems but holy cow i know how hard it is and like even though i'm really good at this i tend to overlook still two out of 10 corner cases and i get paged sunday night or so right those those are the ones that that really that often go like yes i know why i want to use this because i know how much time i would otherwise spent on solving all these things if I have to do it myself right um so I guess I guess that's right like there's definitely an education challenge there and I would say a very a very sort of like in your face example of this is if we look at the AI space and agents uh right now um I think every AI company is like reinventing workflows in the context of like agents and agentic point yeah agentic workflows like everybody's building and I would say slowly rediscovering like all those things when there's been an entire sort of industry that has been working on this for like wait i mean we've been working on this for two years but if you ask ibm they've been working on this for probably 30 years or something like this um i mean in a very different way right but but still and i feel like the all the ai companies are kind of like bit by bit rediscovering this and like when you when you start talking to them i think some of them understand okay even where if you're building agents, if you deploy them, how they ultimately end up having to solve these problems again.

53:12Imagine you have a chatbot that does your flight booking. There's something you have to do to make it not rebook your flight twice if the agent just crashes on the wrong point. They're ultimately going to the same problems. They actually have a perfect foundation to build on with these systems being built today. But yes, I think they're not aware yet that this is something they run eventually into. So yeah, I think you can see this in many places that the industry is rediscovering work in different sort of like subfields that other fields have been done just because information flow isn't perfect.

53:50Minor off topic rant. Why are all of the AI agent, like hello world examples, why are they all booking flights for us? it's like do you want some undeterministic half-baked language model booking your flight that's like a very difficult thing to roll back you know like i just don't that's gonna be one of my last human out of the loop like ai agent moves like can we start with something a little bit less critical i don't know about you adam but i get like serious heart palpitations thinking that someone's gonna book a flight for me and you don't get heart palpitations jared you're a pretty chill dude i am i'm pretty sure but i just feel like gosh you know how hard it is to roll back a flight i mean come on oh yeah well i think it depends i mean i don't mind i think it's the human dream to have somebody or something take that kind of action right that specific action like book me a flight let's simplify it how about you just like appointment give me a yeah exactly a restaurant reservation right you know because worst case scenario i ghosted and feel bad but if I don't show up for my flight I lose my 400 bucks or whatever you know yeah maybe this is so much accumulated pain from people waiting in the like call centers for airlines that I know all these people feel like oh that's a perfect example for a shed but like people will want to use it because they will not want a single other minute to spend on the phone with these calls yeah perhaps

55:31Well, friends, I am here with a new friend of mine, Scott Dietzen, CEO of Augment Code. I'm excited about this. Augment taps into your team's collective knowledge, your codebase, your documentation, your dependencies. It is the most context aware developer AI. So you won't just code faster. You also build smarter. It's an ask me anything for your code. It's your deep thinking buddy. It's your stay in flow antidote. Okay, Scott. So for the foreseeable future, AI assisted is here to stay. It's just a matter of getting the AI to be a better assistant. And in particular, I want help on the thinking part, not necessarily the coding part.

56:07Can you speak to the thinking problem versus the coding problem and the potential false dichotomy there. A couple of different points to make. You know, AIs have gotten good at making incremental changes, at least when they understand customer software. So first, and the biggest limitation that these AIs have today, they really don't understand anything about your code base. If you take GitHub Copilot, for example, it's like a fresh college graduate, understands some programming languages and algorithms, but doesn't understand what you're trying to do. And as a result of that, something like two-thirds of the community on average drops off of the product, especially the expert developers.

56:44Augment is different. We use retrieval augmented generation to deeply mine the knowledge that's inherent inside your code base. So we are a co-pilot that is an expert and that can help you navigate the code base, help you find issues and fix them and resolve them over time much more quickly than you can trying to tutor up a novice on your software. So you're often compared to GitHub Copilot, I got to imagine that you have a hot take. What's your hot take on GitHub Copilot? I think it was a great 1.0 product. And I think they've done a huge service in promoting AI. But I think the game has changed.

57:21We have moved from AIs that are new college graduates to, in effect, AIs that are now among the best developers in your code base. And that difference is a profound one for software engineering in particular. You know, if you're writing a new application from scratch, you want a webpage that'll play tic-tac-toe, piece of cake to crank that out. But if you're looking at, you know, tens of millions of line code base, like many of our customers, Lemonade is one of them. I mean, 10 million line mono repo as they move engineers inside and around that code base and hire new engineers, just the workload on senior developers to mentor people into areas of the code base they're not familiar with is hugely painful.

58:02An AI that knows the answer and is available seven by 24, you don't have to interrupt anybody and can help coach you through whatever you're trying to work on is hugely empowering to an engineer working on unfamiliar code. Very cool. Well, friends, Augment Code is developer AI that uses deep understanding of your large code base and how you build software to deliver personalized code suggestions and insights. A good next step is to go to augmentcode.com. That's A-U-G-M-E-N-T-C-O-D-E dot com. Request a free trial, contact sales, or if you're an open source project, Augment is free to you to use.

58:42Learn more at AugmentCode.com. That's A-U-G-M-E-N-T-C-O-D-E dot com. AugmentCode.com.

58:56i'm going to go out of a limb to bring us back into uh somewhat left the center but basically center please do and i'm going to say that this is the year 2025 is the year where durable execution of things is more important than it ever has been oh really it's always been important but more and more people are leveraging apis they're building out this agentic world we keep hearing about right and i think you keep having more and more people program against brittle apis brittle latency of networks databases etc and you need that promise i'm gonna say that this is the year where the marketing problem that you have that jared alluded to is still there i'm sorry but it's less and i'll tell you why it's less because render i just talked to anwar goel ceo founder of render and this is on their radar so they're building an application for developers did a whole show on this and during that conversation he mentioned a brand at least i think i did actually i mentioned a brand that sponsors us not this show but has been and i think still is a sponsor into q2 and maybe q3 and that brand is temporal so i'm gonna ask you to to sort of help me understand the difference between temporal nats synadia restate your open source flavors in your cloud what render may be doing for application developers it seems like this durable execution retry model doesn't live in the language itself it's something you have to build every single time that sucks and it seems like more and more people are trying to solve it so break down all those for me temporal and that's and nadia yourself what render's doing and anything else that may be doing i mean flink but you know that's a different world yeah there's another one called uh resonate you know that one stefan do you know yeah i know resonate i know the guy behind it um it's pretty new but anyways there is definitely like you said there's other people trying to solve this problem yes exactly um i think this starts from from the same observation like the state of how things are built if you don't rely on one of those tools it's like it's almost unsustainable it's hard to build it's hard to hand it over to another person there's often so much implicit and brittle assumptions and how this works so folks have been trying to come up with solutions from from the ones you mentioned temporal is absolutely the closest maybe yeah between to Paul and resonate I would say those are those are the closest to to restate so I would actually focus on those I would say nuts goes more in the rest and the direction of like flexible persistent messaging together with like some state management blended in and so on.

1:01:47But you can already see like folks are trying to just like figure out what are the different aspects we need when building applications and sort of like make them tight, work together with each other in a tighter way. And if you wish, I think this is the business for me reset is There's a couple of things that make it unique but I would say two things stand out first. I'd say the model goes a bit further than every other system. So Restate is, if you look at Temporal, Temporal is workflows. That's really what they implement, workflows and activities. So it's like durable steps and then, you know, with sleeps in there and signals and so on.

1:02:27So like the full-fledged workflows. It's actually fairly flexible if you're a power user and know how to use that. Restate goes beyond that by saying, we're not just looking at a workflow. So I'd like one durable execution of multiple persistent steps. But we're sort of generalizing this almost what Temporal does for workflow, which we're trying to do this for a distributed service architecture consisting of like multiple stateful services that interact with each other. And that can see this from the fact that Reset has like persistent messaging and RPC built in. it has state built in that lives across a single durable execution.

1:03:09So again, 10 pole terms, the workflow is done. The workflow is done like, you know, it's sort of a self-contained unit within the workflow across the durable steps. It remembers context, but once the workflow is done, it's done. And then Reset is a stateful model where you could almost think of the activities or like decoupled from the workflow. The activities can be stateful services and entities that live for a very long time, and then you have durable functions that interact with them. It's a much more flexible and powerful model to build things like distributed state machines. We have folks that actually start ditch certain elements of databases to put their state and restate because that is transactually integrated then with the durable steps and out of the box consistent.

1:03:50So let's say number one, sort of think of the temporal model, but generalized into distributed services to include long-lived states, include communication like between microservices. It makes for like more powerful, more flexible box. That's the one thing. The second thing goes a bit back to what I said earlier. When we started this project, we set out with the following. You can implement durable execution. I think it's not terribly complicated to implement a durable execution API on top of a database. If you make it very simple, have a step, write it to a database on replay. just like query the database what are the steps that are already in there it has a lot of holes but you know like it it gets you started but then okay let's talk about the holes right like all of a sudden you you have a problem with like long-running processes that suspend for a long time scaling this to zero you have a problem that you have to worry then in your library you have to implement your own distributed locking uh and mutex with you know in case you have timeouts and zombie processes and so on and um so when you when you try to make it a really good experience you quickly come to the point, okay, we actually have to go a lot further than building a library on top of database.

1:05:06Then you start, you know, maybe we're building a big orchestration server that still uses a database in the background. And then you really come to the point of, if you want to make durable execution so lightweight that you can use it almost pervasively, how low latency do you have to make these steps under load? What is the best you can actually do if you deploy this across multiple data centers, if you deploy this across multiple regions. And then you come to the point that, you know, a distributed database across multiple data centers and regions, there's a lot of coordination back and forth because the database model, it needs to, you know, guarantee integrity.

1:05:42It does a lot of like transaction time stamping back and forth and round trips. On the other hand, if you build this on a log, on a like on a like optimized transaction log, you can get as good as make one flexible quorum right across your across your different data centers and you have the step persisted and you can continue, right? So this is kind of going to the point where it's saying, if we want to make this extremely fast, so low latency that it has, you can actually start to use it in places where you didn't think you could use durable execution before because it becomes so cheap, so low latency.

1:06:18How would you have to build a system to do that? And that's where we went. You'd have to build it from first principle, starting with a low latency replicated log. On top of that, build it like end-to-end event driven. So you don't do like batch queries on a database, but you do the most low latency thing you can do. You do fine grant messaging and event pipelines and basically layer from there. And then the other thing is like, okay, let's not just make it really low latency, but at the same time, it has also to be an extremely lightweight thing because somebody who, you know, we just said, what's the simplest use case?

1:06:56Like, when should you actually start looking at reset only when you have a distributed ledger to build? Or do you want to do this if the only thing you want to do is like put your asynchronous email sending in the background, but reliable. So the next thing is how do we actually make this extremely lightweight? What's the most lightweight package we can give that thing? And the most lightweight package is single binary, zero dependencies. Just download that thing. it has its log built in, its orchestration layer, its metadata consensus module, everything in a single binary. Just download one command starts in a second and you're done.

1:07:27Like there's literally nothing else to do. And then you can take this thing actually and start scaling out just by adding more nodes. If you want to migrate it, let it take a snapshot to an object store, start deploying the data, send us, resume, go from there. So what's really the experience that durable execution needs if you want to be able to take it from the point that it's so lightweight you almost want to embed it with almost any application to this thing powers like distributed multi-regional payment processing. What's the architecture you need for that? So that's what we started building in Restate.

1:08:00So the second thing, that was a very long way of saying the second thing is Restate is really sort of a durable execution stack built from first principles. for low latency, serverless operations, high throughput, and just like really, really nice operations from the small to the large scale, rather than saying, let's start with whatever database we have. I think in Temporal's case, when I came out of Uber, they started with Cassandra and said, let's build a server that sort of like sits on top of Cassandra and like stores all the state that it needs for coordination in there. And then, you know, you have like different pieces that you need to scale.

1:08:37you have a database that does actually a lot more than you really need for durable execution but on the way also sort of sacrifices the the potential for optimizing those are the differences i would say gotcha built for speed basically that's what you're saying built to be lightweight scale down yeah scale down scale up and yeah lightweight simple to operate i usually don't don't like to do this. I usually like to talk more about like what makes reset great and what makes other systems not great. Adam pushed you on the spot. Well, you know, I think it's important. Well, if I'm going to say, if I'm going to go on a limb and say, this is the year, then you have to follow me.

1:09:17Okay. Yeah. You have to follow me and you have to answer my questions because I'm, I'm reducing your marketing churn for you just by nature. So I just say, if you look at the way reset is built and it allows you to, to get, get started and scale from there. If you say, okay, I care about self-hosting this because what I pipe through this is like critical data. It's not something I trust with a managed cloud. It really has to run in my account. So I think the experience you get out of reset is vastly different than what you get from many other systems. And that's because it's just been this very thoughtfully crafted stack from the very beginning and not sort of incrementally evolved from, yeah, from this database and then that server.

1:10:08And yeah. Right. If you're directly comparing to Temporal, which is an incumbent, which was spun off from, as you mentioned, Uber and has different principles for which it was built on, you went back to first principles and said, okay, if you want to get to the point where you can put this in almost everywhere you want, you have to be low latency. You have to be fast. You have to, these first principles you built on have to be there. Yeah. and you can't have the requirement to first install a distributed database before you get started. What is the requirement? That's where I was going to go to.

1:10:37So it seems like it's client, which is an SDK essentially inside your code base making calls to a server. What is the architecture, the infrastructure required? So the reset server, which is where the low latency consensus log lives and the thing that basically becomes the reverse proxy for your service. That thing has not really any requirement if you want to get started. So self-contained binary, it embeds its own distributed log. Roxy has a storage engine. It's our consensus engine. The only thing, if you want to run it like as a single node, is you need to give it a persistent disk. It's almost like, let's say, if you would want to run like SQLite or Postgres, It's a little bit like, let's go back to the good old days where you download one binary.

1:11:29I just started and it's actually running. It's actually good. There's nothing else you need to do. But then at the same point, it's also able to go from that single process that you start with a single binary to actually cluster up and build a distributed cluster. And there's a very interesting architecture in there. in the we built it basically for the cloud data age where you would say any system that you that you run at scale should not really store its own data but it should just you make use of object stores for as much as it can because s3 and these these systems they're these like bottomless insanely durable and insanely cheap storage systems so make use of that as much as you can to put a lot of your a large a large chunk of your data so that that means while you may be working with the data on your individual nodes you're not really required to safeguard it on the notes because you can recover it from from s3 or an object store right so what reset then does is it it actually implements its log in such a way that it only uses this for to really give you the very low latencies for the durable steps and then in the background it incrementally moves moves data to s3 which makes the individual nodes like fairly lightweight to operate so um to go go back to your question like what are really the requirements when you want to run it if you want to run it on a single node none or a persistent volume if you actually want to run it like in production if you want to run it in a distributed setup give it an s3 bucket those are the requirements if you want to if you want to use it from your code the requirement in your code is to use the SDK and to to basically create a reset entry point that reset can connect to and where can use it sort of durable invocation protocol that understands how to decode that this sort of entry point is it's mimicking the the popular frameworks like you know it's relatively relatively close to Express.js if you're talking in the JavaScript world.

1:13:37In the Java world, it looks more like Spring Boot and so on. And then within an individual durable function, service handler, you need to use the restate context to say, okay, I want to run this step and record it as durable. I want to create this as a durable promise for a persistent callback or so. But otherwise, the structure of your code is very much the same as it used to be so it's supposed to you know be as little invasive or as little to get as little in the way of how you used to do things as it can just sort of changing the paradigm as in because it has this fine good durability for these operations you can get rid of a lot of the sort of unhappy path code like there's still cases you need you need to treat but mostly you still need to treat sort of like persistent errors that come from the application and where you say like okay you're making a call to an API where you're not authorized there's not really a way you can recover from this it's trying to do something you're not supposed to and you know handle this but don't worry about handling process failures network failures rate rate limits where that bounce you back don't don't worry about many classes of race conditions about you know like the state being maintained in the database versus the logic that interacts with it in a function that can, you know, where you don't know really did this go through or not.

1:15:05Just if you put the state at the reset handler, it's just going to be consistent for you. All of those things. While keeping the structure of the code like close to what you used to write. So I made you talk about things you don't like to talk about, except for maybe the architecture. That seems kind of fun to you. What is it that you do like to talk about when it comes to defining and describing restate and why developers should consider it? So now what I don't like to talk about is competitors in the sense of I don't want to say, OK, I don't like this about that competitor. I don't like that about another competitor, because number one, I'm not an expert in those systems.

1:15:41I try to be honest. I look at them to the extent I need to look at them, but actually no deeper than I need to, because I found this very liberating to not have my judgment sort of clouded or pre-biased by having looked at something. I feel if you look, for example, if we look in depth at how Temporal would build their APIs and so on, there's a very good chance that, oh yeah, I get it. This is why they did it. This makes sense and so on. And there's a good chance you'll probably do it the same way just because you've sort of seen this example, decoded it, understood it, and you're preconditioned exactly.

1:16:14If you don't do this... It's like Trillinger's cat. is the cat alive or dead in the box we won't know until we look maybe yeah but if you don't live you actually have a chance to do something to come up with your own creativity possibly possibly do something better right so that's one of the reasons why i don't like to talk about them so much because i'm absolutely not an expert i look as much as i need to but i don't i usually don't try to go super deep into these systems um and the second thing is that i don't know there's so much i'd rather talk about good things than than bad things it's like more of a it's more more fun to say nice things than bad things i understand your your discomfort then now definitely it's always uh it can be tumultuous talking about competitors and what they do and what they don't do i think in the context the reason why the question is pertinent is because whenever like to jared's point you have a marketing challenge ahead of you and i think it's because Because the idea of durability and item potency is mostly well-known, not always easily implemented.

1:17:15And there's options out there. And so when you sort of look at that challenge, you think, well, what could someone reach for? When would they reach for it? When does it make the most sense to reach for it? And does it actually fit whenever they do try to implement it at scale across different boundaries and whatnot? And so I think when you compare that and you look at like, well, NATS, NATS is a whole different scenario. But they do similar things. It's kind of funny because when you mentioned Flink, you were like, well, it does this in a different way. And then you got to restate because of your experience there and whatnot.

1:17:44Same thing with Nats. It's like Nats does a lot of the similarity things where your brokerage messages, there's a lot of retries, there's a lot of key value storing in there. It's like a lot of that same principles, but it's not about durability. It's not about retries. And then you obviously have Temporal. You have Render who's trying to or going to do something like that in the same platform, which I just had that conversation with. honorag and then you obviously have restate and how you went back to first principles versus being spun out of something or what have you so i think from a a guide standpoint you're the most you're the you're the best suited guide in this conversation to explore those because jared and i can't do that for us yeah absolutely so the if you um like if you want a quick summary I'm very biased, but I think there's almost no reason to not use for reach for restate.

1:18:38I think it really is this solution from first principles with amazing developer experience with a very powerful abstraction that allows you to build what you can build with workflows and signals, but also so much more. and yeah just the just the journey from the beginning downloading the binary then you know migrating um getting out it's an it's an absolutely it's a it's a great experience and the i mean the project is newer than other projects so it will have a rough edge here there but it's also moving moving very quick it's very good at reacting to community feedback fast so i think it's a good choice.

1:19:20It hasn't made a lot of users happy so far. Could we use maybe, Jared, an example from our own application to consider how we would pick up Restate? I know we publish episodes, right? We publish episodes. We often will have scenarios where the slug isn't right. We've had different scenarios where we had to do things in prod to fix something. It could be metadata and you've got different checks before the publish process. Is there a way, knowing what you know now about Restate, how you would consider implementing something like that to safeguard publishing episodes in a durable way? I've never really used one of these tools before, so it's difficult for me to say that.

1:20:00I do know, just at a technical level, that I do not believe Restate has an Elixir SDK, so we might be out in the yard. Elixir one, that's a good ask. Okay, maybe I can help you, come up with an example here. Let's say you're recording the episodes and then every time an episode is done, let's do an AI thing here. So you're building your chat where you can chat with an episode. Like, okay, show me, or tell me when did they talk about this or tell me what episodes talked about these topics and so on. So what you're doing whenever an episode is done, you're feeding it first through a model that transcribes the audio then you're chunking it up feed it through embeddings models or maybe in a vector database and then you have kind of a rack style way of you know when when a query comes create the embedding look up the similarity search in your vector database feed it to the model to yeah to get the to get the answer for something like this if you let's say um let's say you started just building the flow in say in a node.js application like in a simpler way just said okay you know here's the episode i have a i have something like it gets uploaded let's say you're uploading it to an s3 bucket and there's like an event whenever something you know gets uploaded to this to this bucket you have an event that represents this and then it starts like a node.js script or something like this and this script is of the type that you know if it fails you know somebody would have to restart it and now let's say you're trying to implement that with ReState.

1:21:43I would say approach it the following way. Like the first thing is get a handle of ReState itself. Like there's a cloud service that you can use on our site, which has a free tier. Either go there or just use one of these ways to run it yourself on like a single machine with an EBS volume. then you have the server there. Then put your Node.js script, maybe you can actually put it on something like Lambda ECS, just like use a serverless option to host this, and then tell Restate, use the Reset SDK to define the entry point and tell Restate, okay, hey, here's the service that you now should sort of durably manage.

1:22:25So Reset will then go there and discover this and understand, okay, hey, there's this, you know, like what do we call it, like video transcriber or video embedder service. and then reset knows about this and then you can you would go to to your to your Amazon console and say okay for this type of event I want to create a webhook to restate so that it makes an invocation to reset says okay this thing has been updated you know the kind of event that would previously call directly your node.js process or script you know you actually make it an HTTP call to restate and reset with that call yes your process you've already gained one thing right away.

1:23:07You've now basically have a reliable queue in front of it, right? Like just that if you don't do anything special. So when the web port comes, it's going to be acknowledged back and reset has this thing, it's of your process of the crashes. It will retry this. It will actually give you a nice observability, like much more than you would get from like your average message queue about like individual retries, configuration about timeoffs and backoffs and timelines and so on. As the next step, you would actually then go into your script and say, okay, let's actually identify the steps where if something fails in this step or after that step, I don't want it to go back.

1:23:42Like, let's say, forking the process that does the transcribing or like calling the LLM to create the embeddings. You introduce then the reset context that you get by using the reset SDK and just say, okay, let me wrap these API calls just with reset or run. And that will capture the results of this durably and basically turn, you've now turned it basically into a workflow. Let's say you want to do something like, let's say you want to do something like parallelize the different steps. You know, maybe just typing this one by one through this embeddings model is a little tricky thing. So let's, you want to fan out.

1:24:30You would then, you could then go and say, let me try and do the exact same thing I do in a regular node process. Just make a bunch of function calls, record, like remember the promises, sort of a way to promise.all for those in the end, join the results, put those in the database. You can do exactly that in your code. Just again, anchor this in the reset context so you get like this durable parallelization, durable sort of like scatter, gather, and so on. And so you would then incrementally sort of like rewrite your code to say, okay, let's actually make this step durable. Let's make that step durable and that step durable.

1:25:00Say as the next thing, maybe one of you folks wants to approve it before it really goes out. So you then possibly, let's say, let's do that in the simplest, most possible way, which is create an awakeable or a durable promise in Reset and say, okay, somebody needs to complete this actively. like send an event, make an HTTP caller to complete this and say, okay, this is approved, go through, or no, this is not approved like abort. You can then use, for example, you could put the result of transcription just in restate state. Somebody could look at it from the UI and then say, okay, yeah, I'm making an API call here to approve this and continue.

1:25:45So you can then incrementally build your process, rebuild your process into durable steps. as the next thing you could then for example take it and migrate it from a long-running process to a lambda function because one of the nice things you have with durable execution is when it's waiting for something else to happen it can actually just make this thing go away because with durable execution it knows how to recover it to the back to the place where it was by you know replaying the history of durable steps so you could then say you know if if you're on vacation and you approve it a week later, you don't have like some process running and waiting for it.

1:26:22It's just like it's going to go away. And when the approval finally comes, it's going to come back, use the durable steps to replay back to the point and then do the remaining steps. And so you typically folks would incrementally then rework their non-durable services. First connect them to reset to basically get the equivalent of a durable queue and then like incrementally rework it and say, okay, I want to ask durable steps here, maybe parallelization, maybe a signal. And I think that's typically how you'd approach it. that makes a lot of sense I do see also you have some guides on the website about how to implement certain things I'm curious about the observability bit is that a part of your hosted offering is that a part of the open source project how does the business and fit in and is observability part of that open core sort of thing yeah so at the moment at the moment the like what you get in the open source is is very broad you get in the open source compared to the hosted offering pretty much everything except the fact that you would have we are self hosted and you don't like the whole day whole authentication and API tokens and so on that that exists only in the in the managed offering but other than that we've we've started with an open source first approach so the open source has pretty much the full suite at the moment.

1:27:46The observability, there's two things about observability in restate. Number one, it can actually give you an amazing amount of observability itself out of the box because it funnels all these durable steps through its consensus logs. It has all the information of what happened. It's not just function call, but to the granularity of here's a step that happened or it's actually failed. This is the last step that completed in this failure and since then I've retried so and so many times and this is the last exception I've seen. It has all that information available because it also is connected to the service and understands, okay, what type of errors are happening?

1:28:22Is this like a retryable error or not? And it gives you access to all that observability data in its own UI. It's actually a fascinating way that this is implemented if I can throw in one or two technical details. So there's this durable log that recalls all the actions then everything is indexed into into roxyb instances to sort of retain it in a scalable way and we've built a sql query engine using the data fusion project around this that allows you to basically do sql queries against all of that invocation and transaction journal state and so on so what the ui actually does is it's basically issue sql queries it's almost it's almost like back to the good old days whenever all your state was in a single postgres database and if you wanted to find out why your application is stuck, you just would open the SQL shell and start querying.

1:29:12And we kind of lost that because we went into distributed microservices. And if you want to find out what happened, you now have to do a murder mystery with 20 services. And bring it back. And we're kind of bringing it back, like, yeah, SQL query for the win, for your distributed application state. So this is one of the things. You get an amazing amount of insights right out of just the restate journal. The second thing is because like all the operations go through there um reset can also just uh out of the box generate open telemetry traces and spans for you so if you give it an hotel endpoint that it should push those to um to just give you the the traces right away without um without you needing to configure anything you can then extend it and augment it with your own traces but yeah so those are the two things um the two things you can do and so the business end is basically cloud hosting for restate the business ends going to be a lot more than that but this is gonna look like that's I don't think we can go into this yet interview us again in six months and then it's again in six months it's not ready to be to be announced but what's available right now is most of it what we have built is also the open source so yes on the on the business side we do currently only only only hosting for the next six months or so maybe fair enough cool well i think it sounds like a really cool system i'm excited about this new world of durable execution functions and some way to slap a name on that that brings all of the junior engineers to the yard, along with us seasoned engineers who have felt these pains for all these years.

1:31:03You know, like the serverless folk did. You know, they just said it's serverless and they're like, oh, okay, cool, serverless. Maybe I should try it. I feel like Restate and friends need some sort of a marketing term to just simplify the overall concept of what you all are building. But I do think it's very interesting tech and very promising. I do like the term resilient apps so i think maybe something with resiliency involved but that's all for me adam any other questions from you before we let him go i was liking durable personally i don't know if you want to wordsmith that a little bit here but i like durable i think that yeah that seems to have so if you're interested like we've actually gone through a few iterations we started with something just calling a durable async await because in many ways that's what is underneath the hook.

1:31:52It's like durable asynchronous operations, like a function vocation as an asynchronous operation made durable a step, like a sort of an asynchronous API call with the durable result. And there were some sort of expert programmers that thought, oh, that's really cool. I get it. It's like distributed durable event loops. It's very cool. but then 90 % of the folks did not get that. And we just went with durable execution. And then it turns out that's a term that's maybe increasingly more recognized, but it also undersells a little bit what we said because folks actually think like, oh yeah, so it does the same thing as Temporate, but actually does a bit more.

1:32:31So yeah, we're still on the wordsmithing side. Like yes, distributed durability, resilient apps, resilient distributed state management, like there's so many things on the table. At the moment, I used this earlier in the talk, like stateful durable functions is something we've used. I think this is maybe increasingly getting recognized because of a lot of the efforts that, let's say, Cloudflare does with its work or like durable objects. There's a construct in Reset called a virtual object that is surprising amount of similarities with durable objects. and we would have called a durable object if that term hadn't been trademarked by Cloudflare.

1:33:16And Azure Durable Functions is probably even closer to Reset than Temporal. So if you, I think you can actually think like Temporal, Azure Durable Functions. So Reset is a bit more than that. I think it combines a bit more of like orchestration and stateful logic, even in a more flexible way than Durable Functions does. yeah um but yeah so stateful durable functions is where i've currently landed at um but look it's it's it's a journey i think honestly even temporal hasn't figured that out after five years i think well that's why i said like it's a challenge that i think maybe you know restate shouldn't solve it but i feel like maybe everybody who's in this category like you need it there's a missing category it's almost like a a style of application or an architecture you know where it's like well what architecture is as well it's model view controller okay it's mvc i can build an mvc style app you know where this is like i don't know what i don't know what to call it i got i'm missing a word but it's almost like you know it's restate style or something like maybe you have to term it after yourself if you want to really own the market that's durable function style or it's uh yeah durable function doesn't speak to me personally at all sounds really boring but that's just me and maybe it's working on other folks adam likes durable durable to me just sounds like cool it's not gonna be full durable functions that's what you said is that right that's what i said earlier yeah stable durable functions sdf sdf style an acronym something like that i don't know yeah tdd sd right mvcs yeah i mean model view controller it doesn't have any sort of appeal to it either on its face so maybe that wasn't a bad example anyways we could uh continue to workshop it till we're blue in the face but obviously you've been working on it longer than I have.

1:35:00Hey listen this is the year this is the year of it. The year of the what? The year of whatever this is. The durable function. Whatever this is it's the year. Yeah I feel there's something about durable in itself that's not recognized by lots of folks I think you actually asked about that earlier as well like what does durable really mean like maybe a stronger emphasis on persistent or so um but i i'm i think there's a there's something to be said about resilient yeah i think resilient is a much more attractive word generally speaking and one that to me calls and says like is your app resilient i'm like oh i don't know if it is i want a resilient i want resilience so because durability can show somewhere whereas resilience is like you know what no matter what happens i'm gonna succeed i'm gonna i'm gonna try until i bounce back yeah and I think you're right in the sense of durability as a means to an end like it's very implementation detail if you wish the restate achieves resilience by doing a lot of fine-grained durable operations which make it easy to bring it back to a consistent state and that drives resilience there you go now you're getting your rest of the team down love it now we're deep listen to say you know we have a fun place to hang it's called Zulip oh that's true go to changelaw.com slash community doing this in there and then if you have some ideas about this name or this wordsmith this with us this world that uh stefan is is creating then you know pile on share your thoughts all the good stuff what else what's left anything left unsaid about this durable resilient world we're going to live in what's unsaid about the durable world that we live in um i think it's i think it's inevitable to to come the question is mostly in what in what shape is it coming I think there's it's been worked on actually from multiple dimensions they're like folks like us that work on this from the reset side like here's the lightweight durable lock that sort of is easy to integrate with other functions I think that sort of the serverless folks and the wasm folks are working on that from a from a different side of from a different side like saying okay hey, let's compile everything to WASM and let the system sort of use the WASM interpreter to snapshot things.

1:37:23Then there's, I think there's folks that kind of use container engines to implement this. The thing, what they all share is just understanding of it's completely unsustainable to not have anything like this. It's harder. It's hard. It gets increasingly more important the more moving parts and the more asynchronous processes you have. And I think if we all believe what the AI people tell us, that like 80 % of all this is anyway, it's going to be some agentic stuff in two years. And I think you've just like created an even bigger problem and an even bigger need for this type of systems. So I think this is coming in one shape or the other.

1:38:02This is our sort of... It's inevitable. Our bet of how to best achieve it. And it's going to be fun to see. It's going to be fun to see what happens. there you go all right stefan well thank you so much for sharing the journey sharing the love sharing the things appreciate you thanks for having me cheers

1:38:26okay very fun conversation today with stefan very big idea of resilient applications love the idea of restate the idea of stateful durable execution functions is awesome but it's just four words too long for me i do agree with jared on the marketing challenge ahead but resent applications i'm down with that i think you are too if you haven't yet check them out restate.dev okay big thank you to our friends over at augment code our friends over at retool and of course our friends over at heroku heroku.com actually i think it's heroku.com slash changelogpodcast. If you want to use the URL they gave us, there you go.

1:39:12But the next gen platform fur is coming and I heard it's awesome. Okay. BMC, thank you so much for those beats. You are awesome. And we'll see you soon.

1:39:38Thank you.

From the publisher

Stephan Ewen, Founder and CEO of Restate.dev joins the show to talk about the coming era of resilient apps, the meaning of and what it takes to achieve idempotency, this world of stateful durable execution functions, and when it makes sense to reach for this tech.

More from The Changelog: Software Development, Open Source

All 232 episodes
The era of durable execution (Interview)The Changelog: Software Development, Open Source · 1 h 40 min
Listen in VO