In short
How Vercel’s open-source Eve framework enables cloud-native, horizontally scalable AI agents for business workloads (thousands of concurrent, isolated, durable sessions), plus declarative agent definition, evals, permissions, and self-evolution.
Guests
- Andrew Barba (Vercel technical staff; joined as an early enterprise customer in 2020; worked on CDN/security then billing; helped build Eve after a Feb opportunity to deploy agents at Vercel).
- Shar Dara (Vercel product lead for Eve; 10-year billing background; joined Vercel ~1.5 years ago; co-led Eve product direction).
Key claims
- Eve is “cloud-native” vs laptop agents: scales to 1,000 simultaneous prompts with isolated, durable, recoverable sessions.
- Agents are defined declaratively (directory + instructions.md); compiled into infrastructure-as-code.
- Evals are first-class primitives (deterministic checks and LLM-judge scoring) and Eve uses its own eval suite to ship multiple times daily.
- Permissions: channel-layer identity, tool approvals (human-in-the-loop), and DefineDynamic for caller-specific tool exposure.
- Self-evolution: memory slots and PR-based code changes via a coding sub-agent.
Notable examples
D0 (Vercel’s internal data agent) as the origin; Slack-first conversational agents; background report agents; web-app-powered agents (open-source CRM lead research; AI music maker collaboration).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOBackground of Guests
1:42 to 2:48
Discover the backgrounds of Andrew Barba and Char Dara and their roles at Vercel.
“I'm excited to get to know you guys and to learn a little bit more about EVE and this new Agentec framework.”
Overview of the EVE Framework
2:48 to 3:22
Gain insights into what the EVE framework is and its cloud-native design.
“And Char, how about you, your background and how you ended up here?”
The Importance of Cloud-Native Agents
3:22 to 5:12
Learn how EVE differs from traditional agent setups and its advantages in scalability.
“Yeah, I mean, so it is a cloud-native agent framework.”
Building on Prior Art in Agent Frameworks
5:12 to 7:44
Explore the origins of EVE from Vercel's previous projects and thought processes.
“So let's maybe then go one level of detail further.”
Core Features and Primitives of EVE
7:44 to 11:22
Understand the main components of EVE and how they work together for agent deployment.
“Well, and I go back to, I remember a conversation I had, I think, with Tom Okino at some point about Vercel's vision of automatically provisioning infrastructure and being able to infer that from the application.”
Setting Up Evals in EVE
11:22 to 14:03
Learn about the eval system in EVE and its role in testing and monitoring agents.
“So we have this concept of sub-agents, which in its most simplest form is effectively context management.”
Exploring EVE's Eval Framework
14:03 to 17:04
Learn how EVE's eval framework provides guardrails and testing capabilities for AI agents.
“And you basically get this test function that you can write what you want.”
Exploring EVE's Eval Framework
17:08 to 19:10
Learn how EVE's eval framework provides guardrails and testing capabilities for AI agents.
“AI is writing more code than ever, which means GitHub Actions is running more than ever.”
Managing Permissions in EVE
19:15 to 24:22
Understand the built-in permissioning and authorization features of the EVE framework.
“And one of the things I'm thinking about here is similar to skills, I feel like you probably want some sort of progressive disclosure or something like that, because otherwise the model may get overwhelmed.”
User Audiences and Adoption of EVE
24:22 to 28:00
Learn about the diverse user base and the adoption patterns of the EVE framework.
“different sets of connections based on the current caller.”
Show all 21 chapters
Types of Agents in Development
28:00 to 29:00
Explore the three types of agents used in software applications.
“Or do their agents look very similar to the engineering developed agents?”
Integrating Agents with Web Applications
29:00 to 30:20
Learn how to integrate agents into web apps and their development lifecycle.
“So that makes me want to, I'm going to send this back maybe to Andrew then.”
Collaborative Interfaces with Agents
30:20 to 32:20
Discuss the potential of collaborative interfaces enhanced by agents.
“So the way we think about it internally is EVE is a back-end framework.”
Eve's Stream Protocol and API
32:20 to 34:20
Understand Eve's low-level stream protocol and its APIs for interaction.
“And I think the Vercel Lab team is doing really interesting stuff here, I think, with tools like JSON Render, and taking these lower level protocols and producing UIs out of them.”
Handling Cloud Scale with Eve
34:20 to 36:20
Learn how Eve manages cloud scale requests and session durability.
“So now you've suddenly got to wait for that whole block to do it.”
Self-Evolution and Memory in Agents
36:20 to 38:20
Discover how agents in Eve can evolve and maintain memory over time.
“So we run on serverless functions, which historically have 15-minute timeouts.”
Memory Management and Tool Calls
38:20 to 42:00
Examine how memory is managed and accessed in Eve's agent framework.
“which can literally just update its code in place and then restart itself effectively.”
Memory Management in Agent Frameworks
42:00 to 46:02
Explore how memory is handled in agent frameworks, including different provider implementations.
“grabbing memories from the wrong people and this and that.”
Contribution and Management of Open Source Projects
46:02 to 47:35
Learn about the challenges and processes for accepting contributions in open-source projects.
“Yeah, no, I actually, I was going and looking.”
The Future Vision of EVE and Agent Building
47:35 to 49:50
Discuss the vision for EVE and how agent building plays a crucial role in company development.
“and then also anything that we haven't talked about that you think would be important, we can bring that up.”
Hosting EVE-Based Agents and Infrastructure
49:50 to 51:07
Understand how to host EVE-based agents and the infrastructure requirements involved.
“So how does one host an EVE-based agent framework or set of agents?”
Transcript
Automatic transcript. May contain errors.0:00Most AI agent setups today are built around a single session, where one user interacts with one agent at a time. However, that model breaks down when an agent has to serve a business, where thousands of requests can arrive at once and each session needs to be isolated, durable, and recoverable. Getting agents to run reliably at that scale has meant a lot of hand-rolled infrastructure beneath the agent itself. Eve is an open-source, cloud-native agent framework from Vercel that removes much of the agent-scaling burden. In the EVE framework, an agent is defined declaratively through configuration files, and these files compile into infrastructure as code, so the platform provisions only what the agent actually uses.
0:45Andrew Barba is a member of technical staff at Vercel, and Shar Dara is the product lead for EVE at Vercel. In this episode, they join Kevin Ball to discuss what it means for an agent framework to be cloud-native, why they chose to express agents in plain English. and their view that company building is becoming agent building. Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space.
1:24Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.
1:42Andrew, Char, welcome to the show. Thanks so much for having us. Happy to be here. Yeah, thanks for having us. I'm excited to get to know you guys and to learn a little bit more about EVE and this new Agentec framework. But let's start with you. So we'll go to Andrew first. Can you tell us a little bit more about your background and how you ended up working on this project? Yeah, so I've been at Vercel for some three and a half years now. I was actually a Vercel customer, maybe one of Vercel's very first enterprise customers back in 2020. So a bit of an unconventional path to Vercel. But yeah, I spent my first year and a half or so doing all CDN works.
2:16A lot of deep systems, distributed systems, edge networks, firewalls, security stuff, things like that. And then basically went over to the other end of the business and worked on billing for a good year, which is where I met Char. So Char and I were leading the billing team at Vercel for most of 2025. And then, yeah, February came around this year and we had a really interesting opportunity to kind of look at what does it take to deploy agents to Vercel. And we came up with Eve. Awesome. And Char, how about you, your background and how you ended up here? Yeah, I've been at Vercel for a year and a half.
2:53I was in the billing space for a decade before. So that's where I joined Vercel on the billing team where I worked with Andrew. And earlier this year, Andrew's like, you want to join, you know, build Eve. And I, of course, I loved working with him. So great opportunity. Awesome. Well, let's start then with a little bit about Eve. You know, it's brand new. A lot of folks probably haven't heard of it. What's the kind of high level of what this thing is? And then we can dive into the details in a little bit. Yeah, I mean, so it is a cloud-native agent framework. And the reason I say cloud-native is I think most people, when they're using agents, they're probably running a process on their computer.
3:33And so they're familiar with working with these things locally. Very familiar. So your codexes, cloud codes, open codes, things like that. And these are obviously very, very powerful harnesses that people are interacting with on their machines. But we really wanted to look at how do you use something like this in a multiplayer cloud environment? and what does putting these things in the cloud actually need to do? And it turns out there's a lot of different products that go into making that happen. And so EVE is a framework that basically stitches together with very good defaults, all of those types of products to give you a very robust, out-of-the-box harness that lives in the cloud.
4:10That makes sense. So one thing that immediately strikes me to follow up on there is like, okay, what makes for it to be different if it's cloud native, right? And I'm thinking about all these folks who got excited about OpenClaw and they got their own Mac Mini and they're running it. And it's not exactly cloud since it's running on their desk somewhere, but it is isolated off in this place. But what does it mean to be cloud native in this context? Yeah, great, great question. The example I love to give is, let's take an OpenClaw and let's take an Eve agent and let's let 1 ,000 people send a prompt at the exact same time and what's going to happen.
4:45And so the difference between the two is Eve on Vercel or another platform that you host on will scale out horizontally to meet that demand. We can spin up these harnesses very, very quickly in parallel. We can isolate these sessions. They're durable. A whole bunch of mechanics behind it for recovery and retries and error handling and things like that. And so it's a very different use case than just you communicating back and forth as one person with one agent. This is really meant to drive businesses, whether that's internal agents for a 10 ,000-person enterprise or if you are some B2C company and an agent is your primary product, EVE will support that type of multiplayer type of model and it can do this very, very scalably.
5:30Nice. Okay. So let's maybe then go one level of detail further. And let's start with, as you mentioned, there's a lot of folks doing agent frameworks out there, many of them focused at running on a laptop or a single machine or things like that. But what kind of prior art were you able to bring into this and what were the things you said, hey, this is actually something brand new, we got to solve this? It started from our own data agent that we built inside Vercel. This is called D0. And this was really the first agent inside of Vercel that the entire company was interacting with. And it was a lot of work, a lot of hand-rolled systems to go put D0 onto Vercel.
6:11We had to wrap a lot of our products like our servos functions and sandboxes and our workflows and various other Vercel products. And it was a ton of work by that team, basically just getting it to run correctly before even thinking about what D0 does. So we felt like there was this very large layer sitting below D0 that basically every single agent that we would then deploy to Vercel was going to need. So that was the first piece. And then when we actually looked at D0 and what it was, so much of D0's functionality actually came from markdown files, YAML files, other types of spec files, because D0 being a data agent, it's basically trying to provide this semantic layer.
6:52When you say how many monthly active users are there, maybe that's too obvious, but this semantic layer is telling the agent what these types of words mean. And so when certain people ask questions, they're using different acronyms and things like that. So there's this entire layer that's basically educating the agent. And that idea is really where we started. So we kind of took two approaches here. We said, well, it has to become way easier to deploy an agent to Vercel in the same durable fashion that D0 was doing. And we thought that a declarative file system was really, really interesting. We think a lot of other agents can almost be entirely expressed in English.
7:27And so we took this very dramatic stance in the beginning where we basically said there was no code. E was going to be purely English, markdown, text files. And so, yeah, with those two extremes, that was kind of like the very, very early idea behind what EVE is now today. Yeah, that's super interesting. Well, and I go back to, I remember a conversation I had, I think, with Tom Okino at some point about Vercel's vision of automatically provisioning infrastructure and being able to infer that from the application. And that requires having applications with certain sets of constraints and structure and all of those different pieces.
8:03So what does this end up looking like for someone who wants to develop an agent? What is the Eve? Is it an SDK? Is it something where it's like plug in these markdown files in this place? What does this actually look like? Yeah, it is literally a directory. And the most simple Eve agent could actually just be an instructions.md file. And that's it. And so, yeah, it's, of course, a framework because the grammar behind Eve and our slots, those are all compiled down into this manifest. and that manifest is basically the infrastructure as code that you kind of just referenced where we're provisioning things based on what you have.
8:40So if you add skills, for example, then you get a sandbox. If you define schedules, you get our event infrastructure for schedules. But at its core, it is effectively, we call it chat GPT in a box. If you drop it in instructions.md, you are getting a frontier model as trained and you're getting a cloud-hosted version of it. Interesting. So this reminds me back of like the original page router days of Next.js or something like that. So what are the conventions? You said instructions MD, you've got a directory, you put things, you referenced here schedules or skills. Like what are the different primitives you have to play with here?
9:18Yeah, great question. So yeah, instructions is the identity, of course. So I think most people are familiar with this. Skills is the things that you want to progressively disclose. So I think skills have been around for maybe a year, year and a half now. And I think people are generally pro-adding skills, but I don't think necessarily they understand how it works behind the scenes. And so the key thing to understand with Eve or any of these hardnesses is that is your progressive learning mechanism. So skills are basically giving a hint as to when they should be used. And the models are really good at understanding when they need to learn more about something.
9:53And so they can take that hint and then basically pull in additional context. And so D0 is like 90 % skills. And then tools, of course. So you most likely want to run deterministic code, pull in a deterministic set of data from some outside source, maybe like a Postgres database or some sort of open API connection. And so tools kind of facilitate that. And then we have a specialized version of tools called connections. And connections are for things like MCP. And so Eve has a first-class way to just drop in an MCP URL, and we automatically expose those tools for you so you don't have to think about it.
10:33Those are the core primitives. I think that's the brain and the operations of Eve. Then the next layer that is fairly critical is how do you invoke it? You need to start the agent somehow, and this is what we call channels. And so channels are basically all your entry points into your harness. And so we have a default channel, which is just some HTTP route, and this is what our Tui uses. If you were to hook it up to a Next.js app, you're basically using that default EVE channel. But the more interesting ones are things like Slack and Teams and GitHub. And we basically support all the first-party events that you would expect from those platforms.
11:08And we come up with some opinionated defaults for how we think you probably want to use your agent inside of those platforms. Then we, of course, let you customize it beyond that. So that's really the core primitive. And then it expands quite a bit out from here. and much more advanced use cases. So we have this concept of sub-agents, which in its most simplest form is effectively context management. You're kind of taking some larger set of tasks that you don't want to blow in your parent agent. And so the sub-agent can go take that and work with it on its own and then kind of report back. But sub-agents also enable completely different ways of kind of building teams around agents.
11:44So like internally in Vercel, we use sub-agents where like literally the entire teams own a sub-agent. And then we have kind of another agent that interacts with those. So D0, for example, is actually a sub-agent of another one. And so the data team owns that. We have a help, Eve agent, for example, that knows all about how to build Eve. And so the rest of the company can invoke that. And that's effectively a sub-agent of our larger agent as well. So yeah, sub-agents enable really a different way of organizing who owns what. So those are pretty interesting. And then something else that I'm very excited about that is first-class in Eve is evals.
12:20And so we think it's very, very important when you build a non-deterministic system like an agent that you can actually test it and monitor it over time and see how it's performing, right? And that could be based on changes that you're making. It could be testing a new model or different levels of reasoning and things like that. So evals is like a very first class primitive inside of EVE. I want to double click on that because evals is something I think everyone is talking about and trying to do, but at least in most of the folks I talk to and certainly what I see just diving in with people in places, it still feels a lot harder than it needs to be.
12:54So how do you set up an eval in Eve? Is it a sort of static test case? Are you doing evals on live queries? What does that whole lifecycle look like? Yeah, so we have two different types of evals. So the first type is more deterministic. It's probably more what you're familiar with, where you're looking for conditions to be met, and these are fairly deterministic things. And Eve's example of that is you might want to send a query and then make sure that a tool was called. So not perfectly deterministic because the models might not call it, but in general, you can say, okay, I have this tool defined.
13:28And let's say you're building a weather agent. When you ask, what is the weather? You would like to make sure that the get weather tool is called. So those are a little bit more deterministic, very straightforward to write. And then we take this one step further where we let you do, I believe the term is LOM as judge. And so you can actually reason about responses and grade and score these types of responses. And so for D0, this is where things can get really interesting and much more complex, where we can actually grade the SQL query that might be generated and make sure that it's looking for the right tables and things like that.
14:02So yeah, and eval in Eve is literally a define eval function. And you basically get this test function that you can write what you want. It's effectively a full interface to your agent. So you can start turns via prompts, you can trigger things like human in the loop. And so basically all the functionality of EVE, you can programmatically write and set up an eval. And you can do as much or little with it as you want. What I think is particularly interesting about the way we have the EVE repo set up is our test suite is in fact a full set of evals. So we actually use our eval framework to test EVE.
14:38And that is basically what gives us the confidence to ship new versions many times a day. yeah no that's super cool okay i'm gonna push on this you may not have anything i haven't seen anybody really do this this yet but it's a thing i'm kind of minorly obsessed with so what we just described is essentially evals as guardrails right like we have something we believe to be true we're using evals to make sure it remains true it's kind of a ai or l non-deterministic variation on unit and functional tests a kind of direction that i'm interested in is this concept of essentially turning prompting into an optimization function, right?
15:15So your evals lay out what should be true as an outcome, and then you use maybe an LLM or something else in the loop to actually iteratively evolve your agent to get it to behave the way you expect. Is that something that you've played with at all in Eve? Totally fine if not. I don't know anyone has, but when you started talking about this, I was like, oh, does this make it easier to create that optimization loop? Well, yeah, this is really interesting. So I've never thought about it in this way, but if you use a coding agent to build your Eve agent, I think you'll find that it will actually write and run evals while it's iterating.
15:50And part of this is because of the way we expose our docs inside of the package. And so we hint that you test Eve with evals. And so the coding agents actually understand this. And when we work locally on Eve, we have a bunch of fixture-style apps. The coding agents are running those evals. So I think it wouldn't be asking too much to actually prompt your coding agent to get what you want out of here. And I think it would do a pretty reasonable job at using those evals to iterate. I've never quite seen it write an eval first and then go build the implementation. But I think you could probably get something pretty close to that.
16:27Nice. That's cool. And I really like that idea of thinking, I think more and more frameworks are starting to do this, where they think of our documentation needs to be designed not just for humans, but for coding agents, right? To expose the right things, put the right guardrails in place. Yeah, we had a lot of learnings from Next.js in that regard where the models are trained on much older versions than Next. So it's like, how do you get it to learn the latest thing? And obviously EVE doesn't exist as far as these models are concerned, right? So yeah, so we spent a lot of time making sure that the package had to come with LLM-friendly docs.
16:59And I think we've seen really, really good payoff there. This episode of Software Engineering Daily is brought to you by Warp Build. AI is writing more code than ever, which means GitHub Actions is running more than ever. Your GitHub Actions bill is now a function of how much AI code you generate. And every engineer knows the feeling. You push a commit, and then you wait. Warp Build makes GitHub Actions twice as fast at half the cost, with a one-line change to your workflow. Linux, macOS, and Windows Runners, in Warp Build's cloud or your own, enterprise-ready, SOC 2 Type 2 attested, and trusted by teams like Sky from Comcast, Bitcoin, and Braintrust AI.
17:36Get started with$50 in free credits at warpbuild.com. Every visitor on your platform looks the same at first. Can you trust them, or are they a threat? Fingerprint lets you answer this question right away. With one API call, no model to maintain, and implementation in minutes, Fingerprint recognizes bots, AI agents, VPNs, tampering, and over 100 more unique signals to give you durable risk scoring and threat detection in real time. Even as human and agentic activity blends together, you can identify suspicious traffic faster and build safer and more secure sites and apps. Over 6 ,000 companies rely on it every day.
18:20Join them now at Fingerprint.com. Think about your mobile app's source code. Once it hits the App Store, it's out in the wild. And without the right protection, decompiling is easy for malicious actors looking to steal your IP or tamper with your software. That's where GuardSquare comes in. GuardSquare provides the highest level of mobile app security for Android and iOS applications and SDKs. Their advanced tools integrate seamlessly into your CICD pipeline. We're talking polymorphic, multi-layered code hardening techniques and automated runtime application self-protection, paired with mobile application security testing and real-time threat monitoring to deliver the highest level of mobile app security without compromise.
19:05Don't leave your hard work exposed. Secure your mobile applications today. Go to guardsquare.com to learn more. Curious, actually, any lessons learned there? And one of the things I'm thinking about here is similar to skills, I feel like you probably want some sort of progressive disclosure or something like that, because otherwise the model may get overwhelmed. Like you don't want to stuff the model with all the context about your whole framework all the time. So like, how did you design those to make that work well? Yeah, two things. So our scavody includes an agents.md file, which of course points at the docs.
19:42And we also include our changelog in the package. And this is really, really useful because the agent basically knows that it only has to look at certain things, like when you start upgrading the e-versions and things like that. So it doesn't have to start from a clean slate. It's really good at knowing, like, I'm on version 0.2, and I see the changelog is on 0.5. Let me just look at the diff between that. And then in terms of the docs themselves, the coding agent, like Codex and Cloud, whatever, they're so good at knowing that there's tons of files in a directory and not to scan them all. So we didn't really do much to optimize there other than just picking a folder structure for the docs that we think would make sense.
20:20But the changelog is really interesting, and especially for a framework that's pre-1.0, it's critical because we're breaking stuff every week. And so that's kind of the main way that we can make sure people are upgrading without throwing a fit. Yeah, no, that's super useful. Actually, that's a technique that I think I've found for a lot of code bases. When you have multiple people in a code base at all, having that changelog really helps the agents keep up with what's going on. Yeah, yeah, exactly. Different direction here. So still in this like mental model of, okay, we're moving from my pet agent living on a box somewhere to these cloud-based agents.
20:56One of the big challenges I feel like a lot of us are grappling with is how do you manage things like permissions, authorization? Like what is this agent allowed to see depending on who's interacting with it and that types of things? So is that something that Eve has primitives for? How do I think about data access control and stuff like that in the EVE framework? The funny thing is, when we started EVE, we were pretty set on enterprise right off the bat. EVE was, we knew it was going to be code first. And we knew that, or thinking about use cases like DZero, we expected companies to go build their own EVE agents to run their businesses.
21:37And so now we're actually starting to double back on more of the personal use cases. But the reason I said it is a lot of the permissioning and whatnot, that was baked very early on. And yeah, there's quite a bit built in in terms of where we look at permissioning. Because there's quite a few layers, right? At the channel layer, we basically are saying, who are you? And you define what that is. So for Slack, for example, the who are you is going to be your workspace ID plus your Slack user ID, right? And Evo is designed to be multiplayer from the beginning. And so we give you things like, and so basically everywhere you write code in need, you have a context object.
22:13And that context object is going to give you things like, who is the current user in this turn? Who is the user that started this turn? They can be different people. Different people can follow up at different times. And that kind of gives you the ability to do basic type of authentication or permissioning. What I would say is the more traditional permissioning, where if you're writing any typical backend REST API, You're looking at some off-header coming in, you're authenticating it, and then you're saying, okay, you can do XYZ. And so E facilitates that basically by resolving that identity of the channel layer.
22:46Where it gets more interesting is, can the agent perform some action? And so we have this concept of tool approvals, which basically surface as these human-in-the-loop style messages, and it's completely dependent on the channel that invoked it. So all you have to say is, always require approval on some tool. And if the invocation comes in from Slack, we can render, that channel knows how to render human in the loop. If it comes in over SMS, it could take on a completely different form. And so those approvals give you a lot of control. You can look at, again, who asks the question and who's asking for approval, who is then giving the approval.
23:23Maybe it's not the same person, so you may decide that that's okay, you may decide it's not. And so this is where Eve really shines as a code-first framework. So all these cases, we can't necessarily predict how people want to run these types of things. And we basically just say, hey, go write the code to make the experience that you want. And so yeah, so approvals are baked in. And then the last piece, which I think is maybe most interesting, and something that we haven't even necessarily solved ourselves yet, but you hinted at it, is how can you see what you're supposed to see? And this creates a whole bunch of really interesting problems.
23:59So for example, let's take D0. D0 effectively has access to all of Vercel's data lake. And so if Guillermo is asking a question, he may have access to things that when I ask the question that I shouldn't. And so how do you resolve my credentials versus his? And so we basically have a pretty advanced API, which we called DefineDynamic. And this lets you resolve different sets of tools, different sets of connections based on the current caller. And so you could do things like literally not expose a tool when I'm the one invoking a turn versus, say, Guillermo. And so that's, again, a very advanced use case.
Read the full transcript
24:36It is extremely well documented. So the agents do a really good job figuring this out. When you start prompting for this type of behavior, it will learn that this is basically the way to do it inside of EVE. That's super interesting. Before we go deeper on the tech side, something you mentioned. So you talked about going enterprise and then maybe dialing back to personal use and things like that. I'm going to direct a question over to Char. Who are you seeing as the key user audiences for this? and what is the sort of rollout and usage look like? Yeah, as Barbara said, we did initially target enterprise, given that that was where most of the requests came from.
25:13And we have the infrastructure, so it was kind of a no-brainer. And also, we were scratching our own itch to an extent. So we actually use EVE internally for all of our internal agents. And now we have hundreds of them running on EVE. So that was the initial target audience. But very soon after we launched, we saw thousands of our users actually adopting it. And those are not necessarily just within the enterprise segment. You can see them around our hobby users or pro users building. So the range of use cases we've seen go all the way from business agents, what we call what enterprise is launching, where you automate a business function down to a hobbyist building a iMessage multiplayer agent on a weekend.
25:58And we want to enable basically all of those use cases and let you extend the framework to your needs. And so I love that you started scratching your own edge, right? I think especially one of the advantages working in a developer tooling company is like you can really do a lot of that. I'm curious, you know, if you were to look at the distribution of internal agents that you're building, how much of those are engineering related versus other functions within the business? It definitely started within engineering. And engineering is the function that actually spun up a lot of these agents initially.
26:33But as we saw the adoption over time, it became go to market, move everything they had on EVE. Then marketing is working on automating content creation, socials and so on on EVE and so on the rest of the team. So you see billing that help EVE was an example. So you see examples of product teams that give the product documentation to the agents to answer questions both on the docs and internally and sometimes share those with customers directly on Slack. And so the distribution is still heavy on the engineering side where those are the main users and builders of the agents. But I think you're seeing that distribution kind of get more neutralized as it gets easier to spun up agents and build them as we get to self-evolution and so on.
27:22Yeah, that makes sense. The advantage of the approach you guys are talking about in terms of like, hey, we make this really easy for coding agents to work with is increasingly coding agents are available to everyone. right? Like it's not engineering specific. Interesting. So I guess a question, are there any patterns that have been emerging in terms of how folks end up adopting this? And particularly once again, thinking outside of the engineering context, like I immediately am like, oh yeah, I'm used to working with agents, I'm doing things and all of that. But as you start to see go-to-market adoption, as you start to see marketing adoption, are they approaching this differently?
28:01Or do their agents look very similar to the engineering developed agents? Yeah, there are three types I've seen, I would say. The most common one is the conversational agents where the agent is channeled through Slack, mostly. We are a Slack company, so most of the communication happens in Slack. That was the initial use cases. So like DZero, for example, which was the initial use case, did primarily work on Slack. The second one was kind of background agents. And these are mostly the agent's job is to run some sort of a report and also send it to Slack or send it through email to internal folks.
28:40And the last one, which GTM, for example, is a big user of this, is actually applications. So actually build an agent, but your front-facing is a web app that you're working with. And that web app is actually powered by an agent, and you also have a chat interface with the agents on the web app. So those are kind of the three different types we've seen that come out of it. But yeah. So that makes me want to, I'm going to send this back maybe to Andrew then. So in the case where I'm building a web app of some sort, but it's got an agent powered backend, how does that interact with Eve? Because you mentioned this is directory structure focused.
29:17Do I embed these directory structures inside my web app directory structure or are these independent repos running in different ways? How does this integrate with the rest of the development lifecycle? Yeah, so we wanted a pretty opinionated story with Next.js. So Next has the pages directory and then more recently the app directory. And so we knew the agent directory had to just work. So there's first-party support for that type of setup where the agent literally lives next to your Next.js app. But it does not have to be that way. And I think in many cases, there's still lots of teams that are deploying a separate front end from a separate back end.
29:52And Eve can work absolutely just like that, where you deploy your Eve agent as its own project, and then you just connect to it over HTTP. And that's with that default. We give you the Eve channel, which is basically our HTTP protocol, our stream protocol. So it has things like sending messages and canceling and things like that. And then ultimately, we give you back this durable stream, which you can basically use to power any UI that you can think of. And so yeah, it really is just an API. So the way we think about it internally is EVE is a back-end framework. It is not a front-end framework.
30:28And so yeah, it works with your front-ends the way any of your other back-ends would. Okay. On the application side, we've actually seen some really cool examples of our users shipping web apps with EVE. Two of the examples we saw, one of them was an open-source CRM. So what happened was the user goes in, adds a lead, adds a name and email, and the background agent goes in and researches the lead and completes the whole thing behind the scene. And the second one was, I believe, an AI music maker where you kind of go in and add the tempo, but then the agent kind of collaborates with you on the site and completes the music.
31:05So you kind of are seeing these agentic native applications that are being built. and the Eve agent really is like the back end brain behind the scene. Yeah, no, I think it's really interesting. And this is one of the spaces I feel like is under explored yet. Like a lot of generative UI interfaces are like, let's throw a chat bot in this thing or let's asynchronously generate a report. But I think there's a tremendous untapped opportunity for almost collaborative interfaces. Like what does collaborative document editing look like when you have an agent involved? Right now, you can sort of do that with Claude and their artifacts.
31:45You could try to do it in Google Docs with Gemini, but I wouldn't recommend it. So I think there's a really rich space here for easy to interact with agentic pieces. So I'm curious. So, Andrew, you said basically you're exposing this durable stream and now you can do whatever with it. Is there any sort of primitive that is missing there? I feel like that's pretty low level, I guess, in terms of interacting, which is fine for first generation. But I wonder if there are also opportunities there in terms of what are the abstractions that make sense in interacting with an agent inside an app context.
32:23Yeah, it's interesting. And I think the Vercel Lab team is doing really interesting stuff here, I think, with tools like JSON Render, and taking these lower level protocols and producing UIs out of them. And Eve, the stream protocol that we have is very low level. You were literally getting delta chunks of reasoning and things like that. So it is a lot to work with. The TypeScript client makes this stuff much easier. So we have really nice APIs for iterating over the stream and strongly typed phases and events and things like that. And so you can definitely work a little bit higher up. But you're right, not much.
32:58It's still a fairly low-level API. So yeah, I'd love to see a lot more done here. I do think Eve is a lot more than just interacting with a chatbot, like you said. We have cool features where when you send messages, you can actually specify an output schema, for example. So this is really interesting because now, if you know your UI takes some shape of data, you can agentically generate that thing. So this lets you do really completely different experiences than just asking a question and getting some English response back, right? Oh, interesting. So question on that immediately. Do you include built-in validation of that schema and serialization of it?
33:35Or is this just... Yes. Okay. Right. So what does that look like? Tell me more. Yeah. So we take the output schema, and we basically plumb that all the way through down to the... It was built on a tool loop agent from AISDK. So basically what we end up doing is we actually produce a tool called submit result, and that validates your schema. And it's retribable, of course. So a lot of times these agents will call a tool like that, and they produce the wrong thing, and then we can hint back to the agent basically saying, oh, you missed this required property, and it'll do it again. So it's not guaranteed to get your schema, but especially with these frontier models and the way they are today, it's quite good.
34:14Yeah, they're good at that. Now, the challenge with tool calls is that you lose your streaming once you go to tool calls, right? So now you've suddenly got to wait for that whole block to do it. No, not exactly. So the stream will produce all the reasoning in between, right? Because it's just the final step. Instead of getting a message, right? It's some English message out. You're just getting that final tool call, but you get everything in between. So, you know, if you're using some high reasoning effort, and if you want to, you can inspect all of that, even though you're waiting for some sort of output, you could still see kind of everything, everything in between.
34:47Totally. No, no, no. I guess what I was saying is, yeah, you're streaming up until that tool call, but the tool call introduces a synchronous block, essentially, where you're not getting pieces of the tool, like, you're not getting a partial schema results along the way, which influences the UI you want to build, right? Right now you have to have like a waiter or something like that rather than streaming out the pieces of it as you go. But super interesting. Let's talk a little bit more infrastructure, getting back to this like cloud scale thing. So you mentioned a little bit of like, hey, if I got a thousand requests, suddenly I can just like spin up a thousand different processes to answer this.
35:21Very interesting. That now introduces questions around durability, resumability, all those different pieces. Like how does Eve handle that? Yeah, so I mean, it really started where we knew we wanted to build EVE on top of workflow, which is basically our version of Temporal, I think a lot of people are familiar with. And so the way we modeled this very early on was a single workflow is a single EVE session. And today it's actually two workflows, but you can think of it as one. And so the partitioning is basically every session is completely separate from every other. And there's ways to share data between them.
36:01Or I should say we provide guides on how you might want to share data between them if you bring your own storage. But that is the partitioning. And so a thousand sessions versus a single session, you shouldn't have to think about that. The cloud should do that work for you. But yes, as soon as that turn starts, you are initiating a workflow. And then we do some clever things. So we run on serverless functions, which historically have 15-minute timeouts. These agents obviously can run for a lot longer than 15 minutes. And so we can do some really clever things here where every time we invoke a tool, we can move over to another function.
36:35And it just continues the workflow. Because it's another synchronous break. Yeah, no, this makes sense. Exactly. And in the future, we might decide to, if we have a lot of time allotment left, maybe we run multiple tool calls on the same function. But today, we actually moved to a new function for every single step. And so as long as workflow gets faster and continues to be performant, you don't notice this, and the agents can effectively run indefinitely. That's cool. What about things like durable memory and stuff like that? So one of the things that's interesting working with something like an OpenClaw or a Hermes or one of these things is that they have as a part of their built-in loops some amount of, oh, I'm going to remember this.
37:19I will write my own file that I can reference, or I will iterate on this skill that you maybe gave me a first version of, but I can actually modify my own prompts and sources based on your feedback. Are there mechanisms for that type of learning agent within EVE? Yeah, this is probably the biggest thing I'm working on right now. So all of memory is open NPRs, which you can go inspect if you want to right now. But yeah, so I'd say all of this is, and this actually goes back to what I was saying before, where we started enterprise. And funny enough, these types of things aren't necessarily your enterprise use case, but they are very personal agent use case.
37:57And so now we're starting to double back and build up these features for the personal case. And self-evolution is a really, really important part of that for all the reasons that you just said. The idea is you just deploy the out-of-the-box agent and it just starts getting better as it works with you. And there's different forms of self-evolution. And I think it's also different. Eve is going to do it a little bit differently than maybe something like Hermes, which can literally just update its code in place and then restart itself effectively. So on Eve, we're thinking about self-evolution in a few different ways.
38:29Memory is obviously one of them. And we are doing a first-class folder called Memory. And you will be able to define different memory slots. And so you can actually, you would be able to have, let's say you wanted like workspace memory for your Slack channel. So this is like literally global memory that's shared across all your users. But then you could define an additional slot that is just memory for yourself. And so now the agent actually has multiple memory banks that it can pull and save to, depending on what scope you define. And so it's kind of like a mix of, we'll give you a really easy personal memory provider, but we think it has the primitives for much more advanced enterprise memory use cases.
39:08So hopefully it threads the needle on both sides there. But yeah, so memory is one form. And then the other one is obviously modifying yourself. So can you add a new connection to some MCP server or write new tools, produce new skills, things like that. Skills maybe is a special case or it might actually be closer to memory. But for other things, you literally have to write code to go make it happen in EVE. And for now, we're going to lean pretty heavily into Git. Our customers are using Git. And so what we want to do here is open PRs. And so we will do a coding sub-agent that you can define, and that is basically your self-evolution.
39:46And so we'll kind of do all the hard work of when you deploy this thing, it'll have a sandbox for you with your code base. We'll have the right prompts so it knows when it has to invoke the coding agent to change something. But the goal is to effectively just open PRs against itself. And so that's kind of our model of the VPS world where they just get to change their files. Yeah, no, that's honestly a really nice, clean model because it allows you to define whatever level of oversight you want on those. Exactly. You want to let it self-merge? All right, you can build those automations. You want it to have heavy human review on everything?
40:21You can configure that the same way you would anything else. So yeah, that's really nice. Yeah. So you used a term slots, and this is, I think, the second or third time you used it in this conversation. It may not be familiar to everyone. Can we define what you mean when you say, oh, we're going to define a few slots for this? Oh, yeah, I use it specifically in memory. So yeah, when I generally say slots, I'm talking about basically our top-level directories, right? So skills, tools, sandbox channels. So memory is going to be, if I was talking to my coding agent, I would call it an authored slot.
40:51This is how I would refer to it. So it will be a top-level directory. And then inside of memory, we're basically giving you, it's effectively kind of like memory-like drawers almost, or shelves that have different drawers inside of them. And so one drawer might be global, one drawer might be user-based, one drawer might be channel plus user. Maybe you want memory just in specific Slack channels. So it'll give you the control over how you want to scope those memories. And this lets Eve do interesting things where we can control. So the classic example is you almost never want a tool call that takes in a user ID, for example.
41:26You don't want the agent to guess or think it knows the user ID. And so with these memory slots, do is they're lowering tools that are bound to those things. So it cannot guess the wrong slot, right? Yeah, yeah. You want all of those permission and other related things to be deterministic under the layer of what the agent does. It just asks for a thing and you say, oh, that is scoped to this user. We're going to include that in the actual functional call. Yeah, exactly. Yeah. And so we went back and forth on whether we wanted memory to be first class or not. And the more you look at cases like that and you realize that you can start grabbing memories from the wrong people and this and that.
42:02When we see problems like that, it kind of hints to us that, yes, this does belong as a first-class thing because we can provide the right guardrails and the nice APIs around them and things like that. Quick detail question on memory. How do you expose that to the agent? Is it through a tool call? Is it embedded in the context all the time? What does that look like? We took a stance here where we have this concept called provider and the provider can give whatever tools it wants. And so Eve actually makes, we have no opinion on the tools that should exist when it comes to memory. And so the model is actually, we determinously invoke three functions, recall, save, and this tools function.
42:46And the provider, it's entirely up to it on what it wants to do there. And so the provider that will ship in the framework is a file provider. And this one is very, very simple. Recall, it returns the file in full as context. Save, it completely ignores. So there's no implementation of save. And the tools it provides is basically add memory and forget memory. And so the agent can add a single memory, in which case we will resave the file to blob storage. And then forget memory, same thing. We remove the memory and resave it. And so that's the file memory implementation. And so just to make sure I understand then, this is a deterministic thing.
43:27It's going to be called and the result would be whatever gets returned by recall is embedded in the context of the agent from the start and it has access to these tools. So in this version, the contents of the file are going to be in every agentic prompt and it has access to save or forget particular memories within it. Exactly. That's like the most basic version. This is very similar, I think, to like Hermes has like a user.md file, I believe. And so this is kind of our version of that. But the APIs we're providing let you go way beyond this. And so we really wanted to support something like SuperMemory, where all of the recall and saving is done on their service.
44:05So Eve has no opinions on those things. And so the SuperMemory provider actually looks quite different, where they might provide no tools. They don't want the model calling save memory or forget memory, whatever. They don't want to do that because what they want to do is, we'll call save at deterministic times. they actually take that message history and they will go run some agentic process and figure out what to save and so it doesn't want to expose tools right because they're doing that in their own service and then their recall can do something much more interesting than like the file case they can actually look at the inbound message and say let me recall agentically based on this inbound message and so it's a completely different recall interesting so recall is called in every turn of the loop so you could actually change what your sort of embedded prompt looks like based on nice okay exactly yeah and so we call recall there's a few phases that we call it in obviously turn start is is the main one but we also call it after compaction for example because we can't guarantee that compaction keeps you know your memories around so we'll call recall again after compaction and someone like super memory they might actually look at the compacted prompt and say oh no we're good, it has what we need, or no, we have to re-inject.
45:17So they can do really sophisticated things, and Eve doesn't care. In that world, sorry, I'm going way down in the guts here, but are these each turn recall, is that being injected as a new message at the end of the stream, or are you actually editing the system prompt? So we give you the choice. The default is append, so we don't want to bust the prompt cache if we don't have to. So we use the role user for memories, but you can override this, and you can do a role system. And so for the file memory provider maybe role system actually does make sense. But the downside of that is every add or forget does bust the system prompt.
45:52Yeah, that makes sense. Awesome. That's super fun. Thank you for geeking out with me on that. Yeah, it's all open in PRs. So hopefully it launches soon. We'll see. Yeah, no, I actually, I was going and looking. You've got a lot of PRs open. So actually, that's kind of an interesting question. So this is a Vercel project. You all are open sourcing it, it looks like. What's the contribution model? What's the ongoing management ownership model look like? Yeah, so we have a bit of a funny problem right now where we want to accept outside contributions. But one of the problems we actually have is because our test suite is hitting models and is actually deploying projects to Vercel.
46:37So it's using a bunch of secrets, basically. So we actually have this problem where we can't easily test outside contribution PRs. So we actually have to carefully pull them into a sandbox, make sure that they're not doing anything malicious, and then we have to basically recommit them under our accounts with a co-author, and then we can actually run the suite. So yeah, this is definitely a bummer, and we're trying to figure out a better process here. It hasn't been the easiest thing. But yes, we want the outside contributions, and we've been trying to take in as many as we can. We obviously have EVE agents internally that are actually analyzing these PRs daily to look for the ones that we should be pulling in.
47:13But yeah, the more contributions, the better. And the EVE team was really two people full-time from March until about maybe mid-July, I want to say. And we just brought on four more people that have been doing a great job. So yeah, it's about six people full-time on it now, engineering-wise. Well, we're getting closer to the end of our time. I have one more question I'm going to put to each of you. and then also anything that we haven't talked about that you think would be important, we can bring that up. But the question for each of you in turn, and maybe we can start with you, Andrew, and then Shara, you can close us out, is like, what's the vision for where this is going?
47:50What do you see coming down the road, the next three, six? I don't know if we can project out nine months in today's AI age, but what's coming down? What are you excited about? Yeah, for me, I mean, one of the very early principles behind EVE was we're going to bet on the models getting smarter. And the way we designed the framework was really meant to take advantage of this. I think a lot of other agent frameworks that we had seen, they scaled vertically, where they just appeared to get more and more code. And eventually it kind of doesn't look like an agent anymore. You have a lot of determinism.
48:23It's almost like these workflow style APIs. And what we really wanted with Eve was to go horizontal and basically just say, we are going to bet on the models making the right decisions. This is also generally where the folder structure and things like that came from, where we just want to provide more things for the model and let it make the right decision. And so Eve is going to stick to that principle. It's clear that the model intelligence is not slowing down anytime soon, and we think Eve is going to take advantage of that in a great way. Awesome. What about you, Shar? For me, I think the way we're thinking about it is that company building is agent building.
49:03So the idea is that Eve is more fundamental than your certificate of incorporation. So your agent predates your websites, your domain name, and even incorporating the company. So you start with the agent and the agent incrementally builds the software factory for you, including your complete software development lifecycle. And your job is then to actually fine tune that factory. And we believe that Eve is positioned to basically become that brain of your company. moving forward. That's kind of where we're heading. Well, that makes me think immediately, what's the hosting story outside of Vercel?
49:40I mean, obviously, it's great that you guys can manage all of this and all this sort of thing. But if I'm building my company on top of this, I want to know that I can move it where I need to move it whenever that might happen. So how does one host an EVE-based agent framework or set of agents? Yeah, so I mean, this is actually very important to us. And we do have customers internally that are running EVE on Kubernetes, on their own hardware even. And part of this is taking advantage of what is called a world inside of Workflow. And this is part of our end-to-end test suite where we test different worlds, like a Postgres world, a local world, and things like that.
50:15And so as long as you can define a world for your infrastructure, and many of them exist, like there's worlds for tons of other providers out there, that is really all you need to host EVE yourself. And so, yeah, there's a lot more work that we're going to be doing here to make that experience much better. And the key with things like memory and self-evolution and these other types of things is there has to be adapter contracts. We can only do so much in the framework without bringing in external things. Dynamic scheduling, for example, that will be first class on Purcell, but you will need to bring in something.
50:51But the key is just making sure that those APIs and those adapters are thought of early on in the process and not some afterthought after it already works on Purcell. I love it. That's a pretty good cut unless you guys have something else you want to talk about. No, nothing for me. That was great. Yeah, that was fun.
From the publisher
Most AI agent setups today are built around a single session, where one user interacts with one agent at a time. However, that model breaks down when an agent has to serve a business, where thousands of requests can arrive at once and each session needs to be isolated, durable, and recoverable. Getting agents to run reliably at that scale has meant a lot of hand-rolled infrastructure beneath the agent itself.
eve is an open source, cloud-native agent framework from Vercel that removes much of the agent scaling burden. In the eve framework, an agent is defined declaratively through configuration files and these files compile into infrastructure as code so the platform provisions only what the agent actually uses.
Andrew Barba is a Member of Technical Staff at Vercel, and Shar Dara is the Product Lead for eve at Vercel. In this episode, they join Kevin Ball to discuss what it means for an agent framework to be cloud-native, why they chose to express agents in plain English, and their view that company building is becoming agent building.
Sponsorship inquiries:
sponsor@softwareengineeringdaily.com
The post Scaling Agent Workloads at Vercel appeared first on Software Engineering Daily.
