In short
How cloud platforms should evolve for AI-generated code and AI agents that deploy/debug infrastructure directly; why Kubernetes-based setups lead teams to rebuild internal tooling; and how Render adapts with higher-level primitives, guardrails, MCP integration, undo mechanisms, and durable workflows.
Guests
Anurag Gohl, founder and CEO of Render; previously around the 8th employee at Stripe, where 15–20% of engineers managed AWS VMs. Host: Sean Falconer.
Key claims
As LLMs generate production code with limited human review, workloads run in sandboxes and need stronger security/governance. DevOps teams can’t keep up as AI accelerates app creation, pushing teams away from Kubernetes complexity toward application-level platforms. Token economics and correctness favor higher-level interfaces over low-level infrastructure control. Agents will prefer fast, deterministic “single-call” outcomes with cost and severity guardrails plus human-in-the-loop approvals.
Notable examples
Stripe and OpenAI as Render customers; Base44 (now part of Wix) running web compute on Render; Render Workflows for heterogeneous long-running agent tasks; MCP used mainly for deploy debugging and log/metric correlation; undo via delayed deletion/replay using infrastructure-as-code Blueprints.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Evolution of Cloud and AI Integration
0:45 to 2:00
A discussion about the shift in cloud management from developers to AI agents.
“Anurag Gohl is the founder and CEO of Render.”
Anurag Gohl's Journey and Render's Mission
2:00 to 4:30
Anurag Gohl shares his experiences at Stripe and the founding of Render.
“So by the time you left, how many engineers were managing AWS?”
Challenges of Kubernetes at Scale
4:30 to 6:40
Insights into the complexities teams face when using Kubernetes as they scale.
“And that's what Render ended up becoming.”
Understanding Infrastructure Management
6:40 to 8:40
Exploring the intricacies of managing Kubernetes and the related challenges.
“And then you have to kind of manage those bottlenecks by understanding exactly what that is.”
The Impact of AI on Code Generation
8:40 to 10:00
Discussion on how AI affects code generation and the abstraction of infrastructure.
“Because something will stop working at 2 a.m.”
Internal PaaS Development Trends
10:00 to 12:00
Understanding why companies continue to build their internal platforms despite available solutions.
“I think there's a little bit of, well, this is how we used to do it, so why change?”
The Future of Platform as a Service
12:00 to 14:00
Anurag discusses the sustainability and growth potential of Render as a PaaS.
“So it's actually really interesting to see so many more companies come to us now that already have existing DevOps teams.”
Scaling with Render: Customer Examples
14:00 to 17:08
Explore how different companies utilize Render's services for scalability.
“And we've been able to do that with a lot of our customers.”
Progressive Complexity in Development
18:43 to 21:08
Discuss the balance of complexity and user experience in software development.
“And by the, you know, I don't know, level 15 or something like that, you're an expert, but not everybody's going to get there.”
Utilizing AI in Infrastructure Management
21:08 to 22:48
Examine the role of AI tools in managing infrastructure effectively.
“I mean, I think that there's always going to be a spectrum, right, of the types of applications that we're building.”
Show all 22 chapters
The Role of User Interfaces in Developer Tools
22:48 to 27:00
Understand the importance of user interfaces alongside command-line tools.
“but everyone will have to build as agents start using these applications more and more using their API or CLI or MCP server.”
Control vs. Cost in Workloads on Metal
27:00 to 28:00
Explore the considerations of running workloads on bare metal for performance.
“I mean, I still think the UI is the store window of the application where people kind of discover things.”
The Shift to Bare Metal for Performance and Control
28:00 to 30:25
Learn why moving workloads to bare metal enhances performance and control for cloud applications.
“Because in that scenario, what do you mean by control?”
Introducing Render Workflows for Durable Execution
30:26 to 33:16
Discover the motivation behind Render Workflows and how they simplify long-running tasks.
“which is durable execution for long-running stateful processes.”
Comparing Render Workflows to Existing Frameworks
33:17 to 36:28
Explore how Render Workflows stack up against existing durable execution frameworks like Temporal.
“can call a function and it can control the state and recover from failures and so forth?”
State Management and Recovery Mechanisms
36:29 to 37:58
Understand how Render manages state and implements undo mechanisms for agent-based infrastructure management.
“In many ways, it becomes like a Postgres database.”
The Need for Fast Infrastructure Deployment
37:59 to 40:46
Learn about the importance of rapid infrastructure deployment for agent-driven applications.
“And then, yeah, there are places where we use Redis as well, but really I think Postgres is the bulk of it.”
The Future of Cloud Infrastructure with AI Agents
40:47 to 42:03
Speculate on how AI agents will influence the future of cloud infrastructure and application development.
“And you might even need to spin up a whole Kubernetes cluster to do what you're trying to do on Render, especially if you're doing some kind of like private networking.”
The Future of Token Economy and AI Governance
42:03 to 44:10
Discussing the rising costs of tokens, the need for governance, and future infrastructure dynamics.
“I do think the kind of free launch era of tokens is going to be coming to an end relatively soon.”
Dynamic Infrastructure and Security
44:10 to 45:24
Exploring the implications of dynamic infrastructures and security challenges in the cloud.
“And so it's just going to become everything is just going to operate at a much faster rate and it's going to become much more dynamic.”
Democratization of Application Development
46:12 to 48:08
Exploring how new tools are empowering a wider range of people to build applications.
“Is this truly democratizing access to anybody with an idea, essentially being able to actually build something, at least POC or experiment with it?”
Ephemeral Applications and Modern Development
48:08 to 49:35
Discussing the trend of creating short-lived applications and their benefits.
“on a new kind of cloud as opposed to AWS back in the day.”
Transcript
Automatic transcript. May contain errors.0:00For two decades, the cloud has been shaped by human developers writing code and managing its deployment. Now a growing share of production code is generated by LLMs with little human review. Because that code is not fully trusted, it increasingly runs in isolated, sandboxed environments. Meanwhile, AI agents are starting to operate infrastructure directly by spinning services up and tearing them down on their own. Together, these shifts raise the question of whether the cloud needs to be rebuilt for machine operators rather than humans. Render is a cloud platform designed for application deployment by handling scaling, self-healing, and security to reduce operations work.
0:44Render has been adapting to the AI era by building tools that let agents deploy and debug applications directly, with added guardrails and security. Anurag Gohl is the founder and CEO of Render. In this episode, Anurag joins Sean Falconer to discuss why so many teams end up rebuilding the same infrastructure on top of Kubernetes. What changes when AI agents become first-class users of infrastructures and the guardrails that shift demands? And why the economics of AI are pushing developers towards higher-level platforms that trade fine-grained control for speed and safety? This episode is hosted by Sean Falconer.
1:25Check the show notes for more information on Sean's work and where to find him.
1:41Anurag, welcome to the show. Thank you, Sean. It's great to be here. Yeah, absolutely. I've been looking forward to this. So I wanted to go back a little bit in time and ask you a little bit about your history and maybe how some of those experiences led you to the founding, render, and everything that you're doing there today. So you were roughly, I think, around the eighth employee at Stripe? Yes. Yeah, absolutely. So by the time you left, how many engineers were managing AWS? And I guess, like, what did that experience kind of teach you about cloud complexity and the actual cost companies are putting into just simply like managing this beast of the cloud?
2:18Yeah. So at all times when I was at Stripe, around 15 to 20 % of the engineering team was simply managing VMs on AWS and all the complexity around that storage and networking and load balancing. and it didn't really get easier. And obviously as Stripe's scale grew, the complexity increased and you needed to continue to hire people to continue to help you scale. And it just never stopped. I also think, you know, if you think about like Stripe or even other examples of great engineering organizations like that, they're probably one of the best engineering organizations in the world. And if they're running into this problem, clearly it's a problem that's going to essentially impact everyone if they can't figure this thing out.
3:09Yeah, and just to be clear, I do think that Stripe did the best they could at that time because that was really the only way for them to run in the cloud reliably and scalably. But as technology changed, as containerization took hold, it became clearer to me, and I'm talking about the 2017, 2016 era, when a lot of people were containerizing their apps, but then they had to deal with Kubernetes complexity to actually run these containers in production in a way that was scalable or secure or reliable. It became clear to me that you could give app developers, and specifically app developers, not DevOps engineers, because Render has always focused on application developers.
3:57You could give application developers the same powers that they would get from the most accomplished, world-class internal DevOps team. And you could do it in a way that was entirely self-serve for them. And they could simply focus on pushing their code and then get out-of-the-box primitives for scaling, for self-healing, for all kinds of state management. And they would get, you know, effectively one of the best DevOps teams in the world productized. And that's what Render ended up becoming. Yeah, I guess before kind of getting into specifics on Render, what kind of breaks down, like you have people using containers, using containerized orchestration platforms that are well sort of established in the industry like Kubernetes.
4:49But at the same time, it seems like at scale, that's a very hard thing to just manage. So given that there's all this industry investment into Kubernetes? What is it that specifically kind of breaks down as teams start to try to scale the management of that? Yeah, so we do manage a lot of Kubernetes clusters ourselves, and we see this firsthand. I think it all depends on what kind of application architecture you're working with. If it's a really simple application, then it can run anywhere. It can run on Kubernetes, it can run on Heroku, it can run on other cloud providers. But as you start to become bigger, if your business is growing, you start needing more configuration around which applications can talk to each other.
5:36You start needing the notion of CICD needs to work in a certain way. You need to have certain environments be protected. You need certain services to be not available to the Internet. You need observability across all of this. And then you need to be aware of when something goes down, you need to get alerted, and then you need to go solve that problem yourself because you have to then understand, well, you stood up on Kubernetes and you need to understand sometimes exactly how Kubernetes works under the hood to know, well, this network policy change is what caused this DNS error. And with Kubernetes, you have to manage all of these things yourself.
6:17You spin up, even on quote unquote managed Kubernetes clusters, you have to write a lot of YAML by hand that describes how to deploy your applications. But that really what that YAML is doing is describing a lot of the internals of the Kubernetes architecture and the Kubernetes concepts themselves. And then as your application scales, things become bottlenecks. And then you have to kind of manage those bottlenecks by understanding exactly what that is. You might never have encountered it, but suddenly you have a lot of internal DNS traffic because your application grew and now your core DNS isn't working anymore.
6:56Well, who takes care of that? And then you have to go figure out how to configure core DNS to serve all that traffic. So it's every layer of the infrastructure stack within Kubernetes and even below it, because you also have to figure out how to keep your VMs patched when things happen. And you have to deal with node pools and you have to configure different nodes for different things. And you have to make sure that your nodes are also well utilized. You have to pay for a lot of extra capacity because that's what Kubernetes likes to do. And you end up spending way more on just managing lots of unused Kubernetes capacities just so you can deploy your application the next time.
7:34Yeah. And then I think companies end up spending a lot of time on COGS optimization because their cloud bills escalating and getting out of control. I made this point recently about sort of data infrastructure where we have all these great data tools today, but the complexity of actually managing our data has not decreased as we've gotten better tooling. it's actually gotten harder because there's just simply more data, there's more complexity involved, and now we have a bunch of tools that we also have to understand how those things fit together. It kind of sounds like the same problem when it comes to managing network infrastructure.
8:09We have all these great tools, we have elastic scalability, but there's a lot of complexity making this sort of juggernaut of different services work together in a reasonable way, especially at scale. Yes, exactly. And different parts of your application might scale differently. And so you just have to manage all of that across the board. And even understanding Kubernetes is not something you can sort of just do. You can't really build your Kubernetes YAMLs using AI because AI makes mistakes and you have to understand when it does make a mistake. Because something will stop working at 2 a.m. at night and then AI won't really be able to tell you what it should change.
8:51because you just wrote it with AI. You have no idea how the thing works. And then you don't really want to be vibe coding your infrastructure when something's down. Yeah. Yeah. And I would think that I'm sure we'll touch on AI because you can't really have a conversation of technology without talking about AI these days. But as people are using AI to generate more code, we're creating essentially more abstraction from the details of the code. So there's going to be, I think, more buffer between really understanding the deep internals of how that thing was built and run in the interest of essentially speed, which then when things go wrong, of course, makes it even harder to kind of debug and understand what's going on.
9:30Exactly. So I've heard you mention that you've observed that every company that runs on Kubernetes ends up building some version of like an internal pass that all kind of look roughly the same. Why is it that teams end up insisting on building their own thing? It kind of reminds me a little bit of in the early days of auth, Every engineer wrote auth eight different times across every company. And eventually, services came out that take that thing off of people's plate. But why is it that we're still in this place where people are writing these internal passes? I think there's a little bit of, well, this is how we used to do it, so why change?
10:09And there's also a generational thing. A lot of people learn to build cloud infrastructure using the tools that AWS gave them. And then Kubernetes was sort of this state that we ended up in. And you just sort of assume that that is how you always do it or that that's how you will always do it. But I also now see from Render's perspective, we actually see teams moving from large Kubernetes clusters to Render because they get all the functionality that they're looking for without having to deal with the complexity of Kubernetes. And interestingly, that's happening more and more now, just partly because, and I'm sorry to bring up AI again, but so many more people are building so many more applications with AI.
10:57that this whole notion of statically managing your cluster using a DevOps team is breaking down. It's breaking down very quickly because the DevOps teams are stretched thin beyond capacity everywhere. And they're just trying to keep up with the core application. But then everyone else is like, well, I have this other application I want to stand up. Well, good luck because you can only do so much with a small DevOps team or even a large DevOps team. Yeah, I mean, I think that if you look at sort of the traditional software development lifecycle of discovery, design, implementation, testing, productionization, that kind of thing running in a loop.
11:31Historically, the coding part, the implementation piece was like kind of slow and hard. The other things around it could kind of be relatively slow because it was not as slow as the thing in the middle, which was actually writing a code. But now we've compressed that time. So if people generating, instead of having two apps, they're generating two apps a week. The crippling slowdown, the company suddenly becomes the DevOps team, and how do you actually run the platform? Yeah, and that's exactly what we're seeing, and it's why Vendor has grown more in the last two years than in the previous six.
12:02So it's actually really interesting to see so many more companies come to us now that already have existing DevOps teams. That was not the case two years ago. Yeah, it's interesting. So you mentioned that in terms of why is it that companies are still continuing to build their own internal past is perhaps because of the history, essentially, of people being kind of trained to do it. And this becomes sort of the default of like, oh, this is how you have to do it. One of the knocks I've heard historically against platforms as a service that have tried to ease the challenge of kind of running infrastructure has been that at some point, the company outgrows that and you have to get to a place where you're managing the knobs directly.
12:42Do you think that is going away? Or do you think that's an unfounded reflection on the reality of a platform as a service. Yeah, for the majority of applications out there, you really don't need something like Kubernetes. And Render, especially as we continue to add more functionality, more customization, more flexibility, more power to Render, we continue to raise the ceiling on what's possible. And we continue to do that as a platform. Now, I know I've seen other platforms where you hit that ceiling much, much, much sooner. For example, with Heroku, you can't get more than 14 gigs of RAM if you wanted to build an application that did that, or your application was restarted every 24 hours.
13:32And there was just a very small number of compute plans that you got out of the box. Private networking isn't available in the self-serve version. There's no way to store data on a disk. So there were so many reasons for why you grew out of Haruko really quickly. And Render's always built for people who are building complex, ambitious applications. And we want to make sure that you never outgrow the platform because we think that we can continue to raise the ceiling on what a platform can do. And we've been able to do that with a lot of our customers. And some of our, I mean, OpenAI and Stripe are amongst Render's customers.
14:09Obviously, they use other cloud providers, but we also have companies like Base44, which is one of the largest Vibe coding platforms in the world. They're running all of their compute for their web apps on Render. And they're able to scale with Render. They're now part of Wix, and Wix has some infrastructure on AWS, but Base44 runs really well on Render and they want to stay on the platform, even though now they have access to this massive DevOps team inside Wix that has access to AWS. What are the things that you have to do specifically to make it so that people don't outgrow you? And how do you think about sort of balancing, I guess like blurring the lines between automating it, making it really easy and hiding some of the complexity while also giving enough controls to your customer to not outgrow you or not get frustrated with some opinionated view of how this should be done?
15:04Yeah, that is really the crux of how we think about product and specifically progressive disclosure of complexity. So to give you a really simple example, something that happened recently, we gave a customer a specific control that was only available in our REST API because they were creating services through the REST API. They're not using the dashboard. Their use case is actually quite complicated. and they want to use Render just like they would use AWS where they would spin up a VM and install things on it and spin up an application there. With Render, you can do all of that in a single API call.
15:44But most of our users don't need that specific thing that they want. And so we're not going to expose it to everyone in the dashboard. And so it's stuff like that where you kind of think about who needs something and what's the surface area of how they're using Render and how do you expose that feature in a way that is fully productized, fully supported, but it may not be something that everyone has to think about all the time. Agents are getting smarter every day, but even the smartest agents get stuck without the right context and the right tools. That's where Notion comes in. With the recent launch of custom agents, Notion became the collaborative AI workspace where teams and agents work side by side.
16:24And now their new developer platform is turning that workspace into infrastructure developers can build on. What stands out to me about Notion's developer platform is how much they've built underneath the surface. The platform ships with a set of new primitives that change what you can actually do. Workers are Notion-hosted sandboxes where you run custom code, database syncs, agent tools, webhook triggers without spinning up your own infrastructure. The CLI lets developers and coding agents read from Notion, take actions, and deploy workers through one interface. And with the external agent API, you can bring agents like Claude or Codex into Notion as workspace participants with real permissions and triggers, not just a sidecar tab you're copying from.
17:00Pulling in research sources tracked externally, syncing them into Notion, and having an agent help triage what is worth covering, all in the same place the team already works. Learn more about Notion's developer platform today at notion.com.s-e-d. That's all lowercase letters, notion.com.s-e-d, to try Notion's developer platform today. And when you use our link, you're supporting our show, notion.com.s-e-d. If you're running Postgres in production, you've probably felt the moment analytical queries start fighting your transactional workload. Most teams end up adding a second database and all the pipeline complexity that comes with it.
17:35Tiger Data, creators of Timescale DB, takes a different approach. We extend Postgres with hybrid, row, and columnar storage so one table handles both writes and analytical scans. Native compression cuts storage costs up to 95%. Continuous aggregates keep dashboards live without bash jobs. And it scales to petabytes without you re-architecting. Companies like Cloudflare, Octave Energy, Schneider, Axpo, and FlowCo run production workloads on Tiger Data today. No stale data. No second system to operate, just Postgres. Managed for you, ready for the workload you're building toward. Try it free at tigerdata.com.
18:09You're shipping faster than ever with AI coding agents, but those agents don't vet the packages they pull in, and they don't have security contacts built in. Ori by Endor Labs fixes that. It plugs directly into your editor via MCP, catching vulnerabilities, blocking malicious packages and flagging exposed secrets in real time. No separate tool to switch to, no dashboard to babysit, security that fits how you actually build. Teams using Ori see 10 times fewer security tickets and six times faster fixes. Free for developers. Get started at www.endorlabs.com slash A-U-R-I. I guess it's a little bit like you mentioned, sort of progressive disclosure.
18:46It's kind of a little bit like how video games work, where you start on level one, you have kind of limited amount of functionality, And then over time, as you get more comfortable to controls, they add more and more complexity to it. And by the, you know, I don't know, level 15 or something like that, you're an expert, but not everybody's going to get there. You know, some people just want to come into level one and have a good time. Yeah, exactly. And the other thing that is changing is a lot of people are now using Claude or Codex to manage their render deploys. And we have skills and MCP servers that allow them to do that.
19:20and the complexity is hidden away a little bit further from them, even though Cloud and Codex can actually look at all the different things in our API and make the right calls and they can create our infrastructure as code format called Blueprints and they can describe the application of Blueprint so you don't have to understand what a Blueprint even looks like. But the main thing is that Render operates at a much higher level than, say, AWS. We operate at the level of the application. So you don't need to, even if you're using Cloud and Codex, you don't really need to understand a lot of complex infrastructure concepts.
19:57You really need to focus on how really your application runs on render, which is a very small surface area typically, depending on the complexity of your app. And you can debug it fairly easily given the tools that we have. And these days, again, because of the tools we've built for AI agents to debug render, people are able to do a lot more themselves. And again, I should clarify that we're building all of this for application engineers. We're not building for DevOps people. So the kind of tools that DevOps engineers might want, we don't worry too much about not having them. But if you're building a standard AI native company or a SaaS or whatever it is, there's a 95 % chance that you don't need to go to AWS.
20:41And there's a smaller number of companies, like if you're building an infrastructure company yourself, if you're building, say, a database company that needs to really go down to the metal and figure out exactly how the database clusters are going to work together and use some esoteric AWS incantation, then you should not be using render. It's not for you. But if you're an application developer, you can really use render and continue to use render as your application grows. Yeah, it makes sense. I mean, I think that there's always going to be a spectrum, right, of the types of applications that we're building.
21:12Are you building core infrastructure? Are you building more of an application? And do you need scale for users? You mentioned AI management of infrastructure there. How do you think about, as people leverage tools like CloudCode, Cursor, anti-gravity, whatever it is, more and more. I think the tools that are going to be really successful in kind of this agentic engineering era are the ones that are easy to use as a default experience from the agentic engineering experience itself. So has that been something that you've thought about at Render? How do we make ourselves essentially usable by AI and AI agents to be successful in this new wave?
21:49Yeah, we think about it a lot because our users are already using us with Cloud and Codex and Antigravity. They prefer it because they're already in there. They're already coding using these tools. And so why not deploy using these tools? And why not debug your deploys using these tools? So we've made it so that our MCP server can give you all the information that you need, that we expose the right kind of tools that allow you to spin up new applications. But we're also now thinking about how you can build systems that can be reverted if your agent makes a mistake without losing data or without causing some sort of massive disruption to your infrastructure.
22:32We're also thinking about how if we can identify that it's an agent user, then you can say, well, if it reaches this level of severity or this level of cost, then I need a human in the loop to approve it. So those are some things that I think not just render, but everyone will have to build as agents start using these applications more and more using their API or CLI or MCP server. Yeah, I mean, a lot of people talk about like guardrails around these things, but I think from infrastructure, like guardrails around the cost as well, because you don't want to accidentally spin up something that suddenly you have a huge bill for because some agent made a decision that was not necessarily the right decision.
23:10So for MCP, was there specific things that you had to do design-wise that was above and beyond what the render APIs were capable of? Yeah, so when we first launched our MCP, there were certainly some things that we did not put into the MCP that you could use our API for. But in general, our MCP is able to do a lot that Claude and others wanted to do. One of the big things that we did was the ability to really pull logs and deploy data in a way that can then help you troubleshoot failing deploy. And you can get metrics from the MCP as well. And so what we find is that a lot of people use our MCP to debug what's going on, especially when they're developing the application and they're in this loop of quickly pushing code and seeing what happens.
24:01And so it makes it really fast for them to develop on render because all of that stuff is available as tools in Cloud and cloud can correlate deploys with what happened in production. Yeah, I mean, that's a common pattern that we see at Confluent, which is where I work my day job as well. I think the number one use of our MCP server is probably outside of exploring the data that you have would be around debugging. And I've even seen recently some clients out there putting into their contracts with vendors that they're considering purchasing from within the contract, essentially need to know what is their support for agent debugging.
24:41Yeah, it makes sense. Yeah, I mean, it's just the way the world's kind of going. Like, why leave the tool that you're in all day to go and log in some dashboard? Yeah, exactly. And I think dashboards are great for maybe when you're first trying to explore a new tool and you want to understand the different concepts. But once you have a good sense of what's going on, and maybe these days you don't even need that because you can build the exploration into your API or your MCP server, just like any MCP server can expose the tools that are available to the LLM. How do you think about balancing, just from a product standpoint and a resourcing perspective, where to put your resources?
25:23If the world is kind of largely going towards some sort of unified experience from the terminal, from the CLI, interacting through these agents, Does that mean that you put less time and effort behind the user interface, your web-based interface into Render to put more time behind your MCP and your APIs, which are going to be more of the native experience that agents are interacting with you? I don't think that we're treating it that way. We still want to continue to improve our dashboard because I think people do like a visual experience for certain things. And again, if you want to see a lot of information in a very specific way, then often very purpose-built info-dense views that are designed with that workflow in mind might be more efficient than using Cloud or an MCP and might be faster.
26:20So I don't think that we can do either or. So we're actually increasing our investment in both. We're hiring people, we're hiring design engineers to continue to grow the dashboard with all the new tools that we're supporting. And I don't think developers are saying, look, I'd never want to use the dashboard. They're just adding another interface. And it depends on maybe where they are in their journey. If they're coding and pushing code, then sure, they need that MCP server. But sometimes when something goes wrong, sometimes they do need to go into the dashboard or they do need to see all the metrics in one place or they do need to look at all the logs manually.
Read the full transcript
26:58So that's not going away. And there's an element of discovery that I'd say is still a lot easier to do visually than through Cloud or through tools that are fundamentally text only. I agree. I mean, I still think the UI is the store window of the application where people kind of discover things. And the other thing too is sometimes an open-ended chat is not necessarily the best user experience. It's great if you kind of know what you're looking for, but if you don't, then it's easier to be sort of prescriptive and point people in the right direction when you have a user interface. Yes, absolutely.
27:32And you can really focus on the workflow that the user is trying to go through and make your dashboard very specific to that. And that's a lot easier to do and it makes the user more productive than giving people a generic MCP server to do whatever. Yeah. You started moving workloads to bare metal, I think in mid 2025. And I've heard you say that it was more about control than necessarily cost. Because in that scenario, what do you mean by control? Yeah, so I do want to maybe issue a correction there. So we didn't actually move customer workloads to bare metal until just a month or so ago. Yeah, so we've been waiting, but it's happening now.
28:22And the cost versus control thing is really interesting because we are finding more and more reasons to run things on metal itself because of the nature of untrusted code and its need to run in micro VMs. and micro VMs certainly run better on Metal, even though they can run in a nested VM scenario, the performance characteristics are different. And they're almost always going to be much more performant running directly on Metal. In the case of like Firecracker with a KVM hypervisor, it's a lot easier to do that on bare metal than on some like nested BERT situation. So the control that we get is we are able to tweak the hypervisor.
29:07We're able to control the different elements of the machine in a way that is very specific to the kind of workloads that Render runs. And you could get metal machines in the cloud from AWS or others. And I am pretty sure we're going to utilize those machines as well. But in the long run, I think that we want to be able to look at even things like placement of your application and your database. maybe if we can, even on the same rack. Now you don't get rack level placement controls as far as I know on AWS. So it's stuff like that where you can really hyper optimize things and go down to the level of, okay, well, this is the best possible performance or cost structure or the kind of CPU.
29:58That's really the kind of control you get when you own your own metal. And with AWS Metal, you also, I think you get some, but it's just, it's the metal that AWS has. And you don't get, for example, a concept of different machine sizes. AWS machines are all the same size. The largest possible plan is the metal plan. And you can get a metal plan that's either larger than that or smaller than that. Okay. I think you also recently launched Render Workflows, which is durable execution for long-running stateful processes. What was the motivation around that? I feel like durable execution probably in the last year is finally having its moment.
30:38Is that being driven in part by the investment companies are making around AI or is this something else that's sort of driving this interest in durable execution? Yeah. So what we're seeing is applications themselves are now very dynamic in nature. So when you build an agent, for example, and your user asks the agent a question, the first question might pull up 10 records, the next one might pull up 400. And then the tasks themselves could also be really diverse. It could be a web scraping task or a headless browser thing, or maybe just data crunching that requires 128 gigs of RAM. So instead of what used to happen where these asynchronous tasks were as simple as sending an email or processing a payment, our asynchronous tasks have become very, very diverse and heterogeneous.
31:32And as a result, it's become very hard to build these asynchronous processing systems using standard background workers and queues. And that, I think, is why durable execution is becoming really popular. People want to run a series of tasks, and then they also want to be able to provision different kinds of compute for each task. And with Render Workflows, all you need to do is define your task in code. And so any task that you want to run, you can define it in code, and then you can trigger it in code as well using Render's SDK. And Render just takes care of the execution without you having to provision any compute.
32:20yourself. So you don't have to stand up workers, you don't have to stand up queues. And we give you all the flexibility around how many times you might want to retry the task or the concurrency level for each task that you want. And in many ways, when you try to do this with a queue-like system, you end up building very diverse sets of worker pools, each with its own memory or CPU requirements. So you can process the diversity of your tasks. And with durable execution on render using workflows. You don't need to worry about any of that. It's much simpler. And we take care of the compute for you.
32:55We give you a great observability view into exactly how your tasks are running, which ones are failing, which ones are being retried, the inputs and output to each task, for each task. And it's a really great experience for people who are simply trying to run non-deterministic flows in response to user requests. Would this be something that's directly competitive with the durable execution framework set out there, like a temporal, restate, orcs, or is it more aligned with both like GCP, AWS have now like stateful functions, for example, where I can call a function and it can control the state and recover from failures and so forth?
33:36In some ways, I think it's the best parts of both. Because for Temporal, for example, yes, you have the control plane. But the biggest issue that we hear about from people who have used Temporal is at first, it's really hard to get into. That's a very steep learning curve for Temporal. You have to change your application to match exactly how Temporal works when really all you need is retries or you need a certain level of observability. and you need to make sure something runs and then you just need to be able to define a sequence of tasks. Temporal can also just really become overkill. And then Temporal doesn't actually manage the compute for you.
34:16You have to spin up your own worker pools and you have to design your worker pools so that, again, you kind of almost run into the same problem, like how big should your worker pools be? And if you're keeping them running, are they being utilized or not? And so it's the same problem that you run into with Kubernetes where you just have these things running all the time, but it's not clear if they're being utilized. So there's fundamentally some of the same issues with Temporal that you see in some ways with Kubernetes. And with Render, it's different because Render is executing your tasks and behind the scenes, Render is making sure that we only charge you for the task execution time.
34:56You don't have to worry about worker pools. You don't have to worry about running out of memory. You can define your memory at the task level and that's what you get. And over time, the other thing we want to do is also allow you to scale up in memory dynamically so you don't have to predefine it. So let's say you're processing a really large document for a task that typically takes six gigs of RAM, you might want it to spike up to 12 gigs of RAM and Render will do that for you for that task and only charge you for that task consuming 12 gigs of RAM. So it's much more flexible. It's a lot more accessible as well because you define your tasks in code simply by adding a decorator or a pure wrapper over your existing functions.
35:40So in that way, it's actually as simple as something like Celery or Bullend Queue Management, but you don't have to manage queues, obviously. So we've tried to build something that is incredibly accessible, but still gives you what you want out of durable execution, which is this notion of maintaining state across a sequence of tasks that could execute for days or weeks, and then also letting you make sure that it is actually durable, that you can retry these things, you can define concurrency, you can define the inputs and outputs in a way that are all typed. So I think it's going to be really interesting to see how people prefer, and we already have people moving to us from Temporal.
36:19I think it's going to be really interesting to see when we launch workflows in GA in a couple of weeks, by the time this comes out, maybe it's already in GA, what people use Temporal for versus render workflows. Right. And then do you think as you support this and even the hyperscalers are supporting some version of durable execution, at least like a lightweight version, is this something that just becomes kind of like the default experience for developers building any of these types of applications where, you know, regardless of where they're building, there's some form of durable execution available for them out of the box?
36:52Yeah. In many ways, it becomes like a Postgres database. So you just have to have a primitive in whatever cloud you're using to do this kind of work because it is becoming so common now. And it's a completely new pattern of how people are building applications, which is why we invested in it so early. And we think that it's really important to offer it as a platform primitive, as opposed to you trying to glue together open source things to run them on render. And then was this from a ground up project, Not forked from anything? No, no. Everything was written in-house. We haven't used anything else.
37:30We're not even using a temporal or anything underneath. Everything is built end-to-end using our stack. What are you using for state management, if you can share? We are using, in some cases, we're using Postgres. And I think so far, Postgres has actually been relatively good for what we need. we haven't had to use like a more complex state management primitive yet. Okay. And then, yeah, there are places where we use Redis as well, but really I think Postgres is the bulk of it. So we touched on this earlier, but the idea of as agents start to use infrastructure, you need guardrails, you need these cost controls.
38:15And you mentioned undo mechanisms as part of this as well. So what are some of the new things in terms of being able to undo stuff And how does that work in like a safe way so that agents can actually start to manage infrastructure in a way that you could trust? Yeah, that's an interesting one, because let's say that an agent wants to delete a database. If the database isn't around, no one's using it, you could show the agent that the database has been deleted. but then you keep the database around for 30 minutes or maybe 24 hours so that if someone decides that the agent makes a mistake, then the agent can recover and issue a recover command on the database.
38:55So just keep things around longer. That's the simplest example I can think of for undo. There are other patterns like, okay, well, what if the agent decides to delete your web service and this web service has a custom domain on it and suddenly your traffic is gone? I think even there, you can make recovery a lot easier by having this notion of application infrastructure state that you can replay back really quickly. So that's where infrastructure as code becomes really useful. And so Render's infrastructure as code format lets you define all your services in just a few lines of code. And the agent will then be able to simply add the service back.
39:38So in some ways, it's very similar to making a code change that leads to a bug and reverting the code change. Right. Do you have to also think about, and obviously this is something that you'd probably want to do anyway, but does it become even more something you have to think about is the speed with which you can stand up pieces of infrastructure and tear it down when you might have agents that are trying something, deploying it, realizing it's not the right deployment model, tearing it down and then redeploying and doing that sort of like continuously almost in a loop. Yeah, absolutely. I think the reason we've invested so much over the last year in making builds and deploys much faster is partly this because these things need to be able to come up much faster when agents are trying them.
40:24And it's good for humans too, but agents are just operating at a very different speed, right? And so the faster you can make actions in your system, the more agents are going to prefer it. This is also why I think even with agents and maybe agents writing Terraform, agents will always prefer higher level systems like Render to spin up a web service or spin up a workflow. Because trying to do that on AWS will take much longer because you have to wire up a bunch of things together and each thing takes a while to come up. And you might even need to spin up a whole Kubernetes cluster to do what you're trying to do on Render, especially if you're doing some kind of like private networking.
41:02And so I'm not worried about agents when people say, oh, agents are just going to make AWS easier to use. Well, guess what? I think agents are going to prefer the most token efficient way to get to a certain outcome. They're outcome driven. And if Render can serve that outcome in a way that is much faster, it's a single call. It's much more deterministic for the agent because it doesn't have to configure 5000 variables. then that's what they're going to prefer over something that operates at a really low level, at the level of a VM, where the agent might need to decide what kind of VM it needs to spin up, and then what kind of plan on the VM, and then what kind of, you know, just like there's a lot of stuff that the agent would have to do on AWS.
41:43And maybe more back and forth with asking the... Much more back and forth with the user, many more chances of error, because there's so many more steps. And so the more steps it has to go through, the more likely it is to make an error in any one of those steps. Yeah, I think that's definitely true. Especially, I think that this might not be the case for every company right now, but you mentioned the token cost. I do think the kind of free launch era of tokens is going to be coming to an end relatively soon. Oh, it has already, yes. Yeah, I think companies, you see this in the news now, companies are blowing through their yearly budget of tokens in like a month and then not necessarily being able to map that to the ROI of what value they're getting out of it.
42:25So I think companies are going to be looking for ways of how do we optimize our token costs. And the things that are greedily chewing up tokens, people are going to be looking for alternative ways to solve those problems. Yeah, I think token economy combined with correctness makes the case for much higher level interfaces, which is why I think Render is seeing the growth that it is. Yeah, so if you're right about a lot of this, and the next version of the cloud doesn't necessarily look like AWS, agents really change our interface to infrastructure. What do you think things look like if we fast forward three or five years, which is a long time in our industry, but what is your sort of vision of where we're going with all this?
43:06It all starts from the applications that people are building. And these days, a lot of applications are AI native and they're agents and agents require a lot of different primitives. I think a lot of code being run in the cloud is already untrusted. I think more code three to five years, there'll be more code run in the cloud that has been generated on the fly without any human input, which means that it'll need to run in sandboxes or isolated environments all the time. And you will require much more governance across the board. And I think that in our current cloud, because of humans trusting other humans to kind of do the right things, we just don't have governance across the board.
44:00And the kind of abuse mechanisms that exist today are still, this is changing, but they're still limited by human actions, but that's going to change as well. And so it's just going to become everything is just going to operate at a much faster rate and it's going to become much more dynamic. And what that means is that the infrastructure underlying all of this has to become more dynamic as a result. It has to become more maneuverable by agents, obviously. And it has to also become more resilient against abuse and security issues. Because again, we're seeing that today. A new kernel exploit is coming out every week.
44:42And you really can't protect yourselves against AI-driven exploits unless you have AI-driven protections yourself. That's kind of where we are. And that's going to become a much bigger deal. I mean, we already see all the supply chain attacks, but what we're not seeing is the next level of dynamic network level attacks or some kind of other attack that hasn't even been invented yet that attacks a different part of the infrastructure stack. So the cloud is going to become more scary, but there will be good actors as well. And a lot of the code that's being executed in the cloud is going to be generated on the fly and people are not going to understand it all the way they understand it today.
45:23and I just hope that it all still works somehow. Yeah, absolutely. Well, Anurag, this has been fantastic. Anything else you'd like to share? Well, I can talk a little bit about Render. We're seeing just the most growth we've ever seen in our history in terms of revenue. We have more than, as of June, we have more than 400 ,000 developers signing up every week. We have nearly 10 million live services on the platform and we're able to help so many more people and organizations move faster and bring their applications and grow their applications and make them really capable in the cloud. And I hope that if you've never tried Render, that you will.
46:03It's just render.com. And we're looking forward to helping the cloud evolve in this new era that I just talked about. Yeah, fantastic. Actually, one question and follow-up of that, with all this growth in the signups, and obviously more people building applications, are you also seeing a persona change in terms of the types of people that are building applications? Is this truly democratizing access to anybody with an idea, essentially being able to actually build something, at least POC or experiment with it? Yeah, there's definitely a huge movement of people who are not quite the consumer by coder.
46:40The people who are technical and who can kind of understand pseudocode, which means they can kind of understand what's happening with Cloud. But then we're also seeing people using Cloud and Codex to teach themselves coding because it's now become a lot easier. You can just ask Cloud, hey, you did this, explain this to me. And I think Codex also has, or maybe Cloud too, they have like a learning mode in them. But instead of just coding it, it will actually teach you along the way. So in many ways, I think the definition of a software developer is expanding. And I'm not talking about people who just want to buy code on places like Lovable or Base44.
47:18I'm talking about people who want to build apps for their business or internal apps for their business use case or even external apps like product managers are the perfect example where they're technical. They've had maybe a computer science degree, but they haven't been coding professionally. Well, now they can and they are and they're building applications that would be impossible for them to build earlier. So the whole industry and nature and the number of applications is expanding really, really rapidly. and I think Render's actually kind of fortunate because we're in the right place at the right time to take advantage of this wave of people who just want an easy place and a fast way to deploy these applications and bring them to the people who want to use them because none of these people are actually going to use AWS to deploy their applications.
48:04That's not what they're looking for. So many new developers are really being trained on a new kind of cloud as opposed to AWS back in the day. Yeah, I mean, I think that's exactly right. Even within an organization I work for, Well, one, I'm generating more code than I have in a very long time. And then I think that's across the board with all the PMs that I work with. But also, I think that we're kind of creating a world where a lot of software is like ephemeral because it's easy to spin up stuff. Like I created like a dashboard the other day to measure engagement with a particular product that I could have had someone within the company build in Tableau or something like that at some point.
48:41But I can just spin that up easily using something like Cloud Code, run it on Render, and have that thing ready to go for two days. And then when I'm done with it, I can throw it away. And it's not the same heavy investment that I spend time on. And it's interesting. We're seeing that what you just said about ephemeral applications. We're seeing that people are creating a lot of apps that live for shorter durations. Yeah. And Render is obviously great for that because you only pay for when the application is live. You don't have to worry about all this other infrastructure just sitting around waiting for you to deploy another app on it.
49:19But then a lot of people are also able to create long-running applications that are helping their businesses. And it's really great, really exciting, motivating to be at the center of all of it. Yeah, absolutely. Well, again, thank you so much for being here. This was fantastic. thank you for having me this was a great conversation all right cheers
From the publisher
For two decades, the cloud has been shaped by human developers writing code and managing its deployment. Now a growing share of production code is generated by LLMs with little human review. Because that code is not fully trusted, it increasingly runs in isolated, sandboxed environments. Meanwhile, AI agents are starting to operate infrastructure directly, by spinning services up and tearing them down on their own. Together these shifts raise the question of whether the cloud needs to be rebuilt for machine operators rather than humans.
Render is a cloud platform designed for application deployment, by handling scaling, self-healing, and security to reduce operations work. Render has been adapting to the AI era by building tools that let agents deploy and debug applications directly, with added guardrails and security.
Anurag Goel is the founder and CEO of Render. In this episode, Anurag joins Sean Falconer to discuss why so many teams end up rebuilding the same infrastructure on top of Kubernetes, what changes when AI agents become first-class users of infrastructure and the guardrails that shift demands, and why the economics of AI are pushing developers toward higher-level platforms that trade fine-grained control for speed and safety.
Sponsorship inquiries:
sponsor@softwareengineeringdaily.com
The post Rebuilding the Cloud for AI Agent Code appeared first on Software Engineering Daily.
