The Agent Harness: Building Secure Sandboxes for Autonomous AI Workloads

14 May 2026 · 1 h 5 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that AI agents are “digital knowledge workers” and therefore need their own isolated “sandbox” computers to safely run tools, access data, and execute code. It explains why sandboxes are different from stateless cloud infrastructure, how they fit into the agent stack (models, tools/MCP, memory, orchestration, observability), and why Daytona built its own infrastructure (including scheduler changes) rather than relying on Kubernetes. It also covers scaling constraints like fast sandbox spin-up, CPU/GPU utilization, and potential global CPU shortages.

Guest

Ivan Burazin, CEO of Daytona (agent-infrastructure startup). Background: previously co-founded Code Anywhere (cloud IDE) starting in 2009; built developer-tool infrastructure for ~16 years; learned enterprise/PLG distribution and developer marketing through conferences and events.

Key claims

every agent needs at least one sandbox (chat-only may not); sandboxes enable secure, killable, account-scoped access (e.g., 2FA bank access without granting real spending); hyperscaler app platforms are stateless and don’t match agent needs; firecracker micro-VMs are fast but limited (e.g., no GPU/Android).

Notable examples

agent logging into a bank via separate Daytona account + phone-based 2FA; background agents (Harvey/Perplexity-style) using headless code/command execution or browser/computer use; RL researchers needing thousands to millions of concurrent sandboxes for GPU utilization.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the AI Super Cycle

0:00 to 0:40

Learn about the current AI super cycle and its implications.

“We are a part of this like super cycle right now.”

The Concept of Agents as Digital Workers

1:34 to 2:15

Discover the analogy of agents as digital knowledge workers and the need for computers.

“So you have said that every agent needs its own computer.”

Defining Sandboxes for AI Agents

2:16 to 3:18

Understand what a sandbox is and its importance for AI agents.

“So as an introduction to this whole conversation, what is a sandbox?”

The Risks and Setup of Digital Agents

3:19 to 5:00

Explore the setup and security risks associated with giving agents their own machines.

“Basically, you want to be able to kill it if it goes rogue and therefore you can unplug the Mac Mini the way you could sort of kill a sandbox.”

Stateful vs Stateless Architectures

5:01 to 6:59

Learn about the differences between stateful and stateless architectures in agent development.

“The database might change, information might change, but you don't want the app to change, right?”

Evolution of Sandboxes in Development

7:00 to 8:29

Hear about the history of sandboxes and their evolving role in the developer landscape.

“So sandboxes are fundamentally new primitive.”

The Need for Computers in Agent Work

8:30 to 10:38

Understand why agents require computers and how this impacts productivity.

“When we think about agents, so one, if you, even if you do run an agent on your computer, like you want to do that, up to you, go ahead and do that.”

Use Cases for AI Agents and Sandboxes

10:39 to 14:03

Explore different use cases for agents utilizing sandboxes in various environments.

“So if you do a tool called coding, like any of those actions, then you need a sandbox.”

Understanding the Agent Stack and Sandboxes

14:03 to 16:41

Learn about the key components and the concept of sandboxes in AI agents.

“of the sandboxes to do those two different things are different.”

Memory and Learning in AI Models

16:41 to 18:50

Explore the challenges of memory retention and learning in AI models.

“The tooling, I don't know if I would spend too much time.”
Show all 31 chapters

The Role of Sandboxes in AI Development

18:50 to 21:30

Discuss the importance of sandboxes and their future in AI systems.

“But solving it itself is outside of our, let's call it mandate.”

Overview of AI Provider Strategies

21:30 to 23:08

Understand how different AI providers fit into the landscape of AI agents.

“where do the model providers like the big AI Frontier Labs fit in this overall picture?”

Founding Daytona: Lessons from Past Ventures

23:08 to 24:48

Ivan shares his entrepreneurial journey and key lessons learned from previous startups.

“That's the sort of like overall landscape around the agent stack.”

Building a Brand and Distribution in Tech

24:48 to 28:00

Discover strategies for effective distribution and brand building in technical ventures.

“teachings of Heroku, which definitely went into that direction.”

Overcoming Stage Fright Through Experience

28:00 to 29:20

Learn how overcoming stage fright led to valuable insights in event organization.

“as a child and the the stage fright terrible terrible like pitching my first startup like I would be so nervous.”

Understanding Conference Dynamics

29:20 to 30:50

Discover the essential elements that make conferences successful and the role of sponsorships.

“And it's also a problem because when we said we're going to do a conference, I didn't understand what a conference was.”

Evolving Go-To-Market Strategies

30:50 to 32:10

Explore how iterative cycles and feedback loops improve marketing strategies.

“But we'll do these little things because it's easy.”

Key Differentiators in Product Marketing

32:10 to 33:40

Learn about the three critical factors that influence product selection by consumers.

“Like, if that doesn't happen, there's no way they can pick you.”

Enhancing Customer Experience

33:40 to 35:30

Understand the importance of customer service and experience in driving growth.

“even if our product is equal, that we can supersede that.”

The Role of Social Media in Awareness

35:30 to 37:30

Investigate how social media engagement impacts brand visibility and customer perception.

“For me, my belief is that that is a key driver to continue growth.”

Effective Customer Support Strategies

37:30 to 40:00

Learn simple yet effective strategies for providing excellent customer support.

“So we have found that people that even are completely opposed to the posts end up being customers, assuming that they need the product line, not everyone is there.”

Technical Insights on Sandboxes

40:00 to 42:01

Dive into the technical aspects of sandboxes and their significance in computing.

“There's obviously, there's like key things if you're actually down and don't work and the person can't work, that's a completely different thing.”

The Importance of Speed in AI Agents

42:01 to 44:28

Learn why speed is crucial for background agents, especially in user experiences and AI research.

“So if you're like a long running background agent, again, everyone prefers it to be fast.”

Infrastructure for AI Agents: Sandboxes Explained

44:28 to 46:07

Understand the components of AI infrastructure, including primitives, tooling, and sandboxes.

“The other two are the primitive and the tooling.”

Understanding Firecracker and VM Technologies

46:07 to 50:04

Explore the differences between various VM technologies, including Firecracker and their use cases.

“So most sandbox providers are firecracker VM, micro VMs, most.”

The Role of Scheduling in AI Operations

50:04 to 54:35

Discover how scheduling affects the performance and management of AI sandboxes.

“So if you are an app layer company and you have 10 million users and your 10 million users spin up, let's just say 10 million sandboxes, just to make the math easier, these things cost, right?”

Challenges of Long-Running Sandboxes

54:35 to 56:00

Learn about the technical challenges and solutions related to running sandboxes for extended periods.

“where our scheduler can be attached to any CPU machine, any server that we think is good enough, and it becomes part of our cloud, essentially.”

The Challenge of Managing Sandboxes

56:00 to 58:08

Learn about the complexities and requirements of managing sandboxes in computing.

“And so at some point in time, there's nothing else on that machine.”

Performance and Complexity in Sandbox Environments

58:08 to 1:00:46

Discover the performance differences between various sandbox setups and their implications.

“Listening to this whole technical deep dive part of the conversation, the thought crosses my mind that a lot of people seem to think that they can create their own sandbox and it is not that hard.”

Future Trends in CPU Demand and Supply

1:00:46 to 1:02:46

Understand the emerging trends in CPU availability and the implications for AI.

“and then when you get to scale and then you need performance because now, oh, agents can do a lot of things, then it starts like breaking that sort of mold.”

The Future of AI Models and Their Capabilities

1:02:46 to 1:04:48

Explore the potential future developments in AI models and their implications.

“How do you think all this agent stack evolves?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We are a part of this like super cycle right now. And the super cycle does not last forever. And so if you're going to pause by the super cycle, you are seeding market. Like that is what you're doing. I would ask, can you go fetch the data from our bank? And then it was like, oh, yeah, just log in and give me access. I'm like, log in, give me. No, I will not give you access. Right away, fundamentally, for me, that was like broke the entire thesis of it. So you give it its own machine. When I think about agents, I think of them as digital knowledge workers. And to do anything as a knowledge worker, you do need a computer.

0:36My argument is that every agent will need at least one sandbox, sometimes more. Hi, I'm Matt Turk. Welcome to the Matt Podcast. My guest today is Yvonne Borosin, CEO of Daytona, one of the most talked about startups in agent infrastructure. If you've been hearing the word sandbox in every AI agent conversation and quietly wondering what it actually means and why it suddenly matters, this episode is for you. We go from first principles, why does an agent need a computer at all, all the way to the deep technical end, why Daytona had to throw out Kubernetes and write their own scheduler, and why a global CPU shortage might be coming faster than people think.

1:11Ivan also unpacked the full agent stack as he sees it. Models, sandboxes, tools, MCP, memory, orchestration, and where each piece is heading. Along the way, Ivan shares some really interesting lessons on go-to-market and distribution for technical founders that he has learned through 16 years building developer tools startups. Please enjoy this great conversation with Ivan from Daytona. Hey, Ivan. Welcome. Great to be here. So you have said that every agent needs its own computer. What's the simplest way of explaining that idea? Well, when I think about agents, I think of them as digital knowledge workers.

1:49And to do anything as a knowledge worker, you do need a computer. Or I should say anything sophisticated. So you and I can be in a conversation, we can get something done. but we usually need some sort of tools in our world as usually a computer to get higher productivity and to be able to do things. And so I think of that in the same lens. So for me, agents are literally just digital knowledge workers. Great. And so it's on computer. That's the whole concept of sandbox. So as an introduction to this whole conversation, what is a sandbox? Absolutely. So sandbox is, although the term sandbox first comes from isolation.

2:26So making sure that there's a secure place for, in this case, an agent to run and do things. But on the other side, it is essentially a computer. It's a full-on computer that an agent has the ability to install tools, access the web, run scripts, run code, whatever it needs to get its job done. So the shortest answer is a sandbox is essentially, we call them composable computers for AI agents. Yeah. And is in some ways like this whole OpenClaw and Mac Mini, like a good analogy for what a sandbox is, like the Mac Mini sort of being the sandbox? Exactly. So I think the OpenClaw Mac Mini thing helped a lot of people understand what we actually do.

3:08It's like, oh, I get it now. I get it. Right. So it needs a computer, in this case, was a Mac Mini to be able to do different things. So, yes, that helped a lot with awareness for sure. Yeah, because, again, to unpack it, like the OpenClaw as a framework and system to help an agent do all sorts of things on your computer. Basically, you want to be able to kill it if it goes rogue and therefore you can unplug the Mac Mini the way you could sort of kill a sandbox. Is that kind of? So there's a couple of things that I personally thought about. And so my OpenClaw runs on a not a physical Mac Mini, essentially, but a sandboxed virtual Mac Mini.

3:46And the reason why I didn't, so the way people usually run these things and Cloud Code or OpenClaw is usually on their own computer because it helps them, you know, organize emails, you know, search whatever documents they have and whatnot. There's a one, there's a high security risk there because at one point we were doing our board meeting presentation and I would ask, can you go fetch the data from our bank? And then it was like, oh, yeah, just log in and give me access. I'm like, log in and give me a no, I will not give you access. Right. And so right away, fundamentally, for me, that was like broke the entire thesis of it.

4:27So you give it its own machine. I personally gave it its own like Daytona account, gave its own phone number. And the reason I have to give its own phone number is because it has to do 2FA to get into the bank. There's no other way except for 2FA with the phone number for this particular bank. and so it has to have all these things like a employee like a digital employee to be able to go and access these things and so the risk now that now that it has its own computer it has its own account these accounts have the limitations so it can only it can only look at the data for my bank it can't spend the money for my bank and so except for the credit card that we did give it which has a you know a you know a hundred dollars a day or whatever it may be and so the only risk you essentially have at that point if it's in a sandbox is will it take this data and now leak it somewhere and that's something we can talk about later but essentially the worst thing can happen is that and to your point you can kill the whole machine if you need to kill machine yeah and there is a phenomenal concept of stateful versus uh state less uh can you can you unpack that for us so that basically comes when we talk to people about what we are building they're like oh doesn't this already exist in any of the hyperscalers right and so the answer is no and the answer is know is is because everything that they were built for which was deploying apps they were they were stateless like if you have let's let's pick any of your websites or whatever you want web apps you do not want that to change on the fly you if you if you are the company will pick like ebay because they're in this building so they're they came to mind and so if you're ebay and you're an engineer there you're like oh you have this new you know update a button is here or does this thing, you want that state that you don't want that to be changed on the fly, right?

6:08The database might change, information might change, but you don't want the app to change, right? And so with those things in mind, that is how people had built the hyperscalers. Like that is the fundamental architecture you came into and built on top of that. And the simplest analogy I give people is, let's say you're building a truck, right? You are a factory for a truck. There's like the weight, the type of the engine, the chassis, all of that is made to very slowly but surely securely transport some sort of goods, right? On the other hand, you can have a sports car, which is still, you know, it has four wheels, it has an engine, whatever, but it's made for different things.

6:48It's made to go very fast. So the way the chassis is built, the way the, you know, engine, the weight balances is completely different. And so you can do a lot of things with both, but they're not fundamentally the same thing. And as a company trying to build out both of those, they're separate platforms completely, right? So sandboxes are fundamentally new primitive. Is that correct? Is it like a history of sandbox? Did sandbox exist before? So I would argue, I would say that the person that kind of nudged, although he said he did and someone else did. But anyway, there was a company called Code Sandbox way back in the day when we used to compete with our company Code Anywhere, which also was in the realm of like Replit.

7:27It's all these cloud-based IDs. And so they called it Code Sandbox because I believe it sounded cute. It's like a box where your code for ID lived. And they were actually one of the first. So the team there that actually used micro VMs, did snapshotting, forking, all these things that we do use today. and so that is sort of and this is maybe a decade ago but the utilization or the usage from or the value that it was giving human developers was not there there's this whole article about like the end of local hosts and people have been talking about this for i've been talking about this for 20 years um and basically developers would say like you'll take my local hosts like out of my cold dead body but now that agents are here local hosts no longer actually, one, you don't want that for a number of reasons.

8:17And so we're finally getting to that. So that technology and that thesis around sandboxes originally now seems to be coming to fruition. And the big acceleration around sandboxes is agents, right? That's the thing. Yeah, absolutely. When we think about agents, so one, if you, even if you do run an agent on your computer, like you want to do that, up to you, go ahead and do that. There's a bunch of problems. One is, let's say you're on your laptop. there's this whole thing on Twitter now where people are holding their laptops open. Yeah, yeah. What is that? I saw you tweet the... Oh, that's because I always hold my laptop open.

8:50That's just like my vibe without that. But I've been doing that for a long time. But people do that because they want Claude or OpenClaude to finish. Yeah. Because if you close it, it's... Yeah, yeah. You terminate. You terminate, yeah. Or pause or stop, whatever. So the problem is the ability to work nonstop is not there as long as your laptop. So that's one thing. The other thing is you can't do concurrency. So you can do multiple to the amount of compute that you have in your laptop, which will be quite limited. And so you might want to spin up 10 or 20 or 50 or 100 or 100 ,000, whatever that number might be.

9:20And so that's very hard. So ideally, you actually really want to remove it some remotely. So you can start on your laptop, you continue on your phone. It's the same computer and same agent that is doing their thing. Right. And so that is a takeoff of we have now decided that it's absolutely OK, that it's no longer on our local host. but again it is still a local host to that agent so it's still a computer to that agent what a laptop is to us do all agents need a sandbox or is that a specific category my my argument is that every agent will need at least one sandbox sometimes more and we get to that again there are places where you don't need and i analogize agents with humans for the most part and again we can have we can do productive work without computers and so there is and that was the original sort of like chat bot where basically you would just talk to an agent and it would inference so it just think and give you value so i don't know you a lot of people use it for like emotional support or whatever for the most part it doesn't need a computer it has enough data from you going back and forth and then it can sort of understand and give you feedback but that's a smaller subset if we think of where all productivity gains are among the biggest verticals, healthcare, financial services, whatever, most of that is done via a computer.

10:37And so if you want agents to do all these things, then agents will need these computers to do. So if you do a tool called coding, like any of those actions, then you need a sandbox. If you just chat, you probably don't need it. Yeah. So like, even if you have a chat that we can get it to deal, but if you chat and it has to search the web, it still has to open a browser or it has to use something like, you know, Parallel or Exa or whatever to get to that. That would be a tool call. So it depends on the use case. But the thing that's quite interesting about the world that we live in is one, let's take it a step back.

11:13One, I firmly believe that in due time, all the tools will be headless. Now we'll be inside of a sandbox or outside. That's another thing. All of them will be. And that's the most efficient way for an agent to work. But most of knowledge work is still locked into legacy apps inside of Windows for the vast majority, like the absolute vast majority. So if you want an agent today to do a job end to end, you literally have to give it a computer. And so an example of this is again, the board report, which is like, oh, can you pull out this report and for our bank has an API and it can pull it out but it only has spend on the API it doesn't have incoming revenue on the API it's just not exposed or through the mcp tool actually it's not exposed and I'm like then the agent was like oh I can't get to it I'm like dude log in download it okay I'll go to it and then you can see it opens up a browser it goes down or whatever Same thing with, I don't know, with any other data source that we have.

12:17If it can pull it from the headless, it'll do it headless. If it cannot, then it sort of logs in and do that. And so if we want to give them power today and get value, we have to enable them to have these tools. You had a great tweet the other day where you were talking about the actual use cases that you see as a provider of sandbox of agents. Do you want to go through this? You had a code command execution, computer use browser, use an RL environment infra? Do you want to unpack that? Sure. We basically, I've made it, I think I've structured it better since that tweet, which is right now we have two major use cases or two types of customers consuming Daytona and different use cases.

12:58And one is on the researcher side. So it'll be, you know, RL evals benchmarks. And the other will be on what we call background agents or long running agents. And so when you think of background agents or long running agents, the most popular, those are where a human is the end consumer the human talks to a let's call it app layer service that has an agent in that and then the agent will call in the sandbox so things will be you know think of like harvey or perplexity or whatever as or lovable as these types of background long running agents and so they can both be sort of headless so code and command execution and or computer browser use so depending on what they need to do and the same thing is on the researcher side where it's like rl and evals and whatnot you can do rl and evals for you know coding and for that it's basically just headless like commands and command execution in there or you can actually teach it to do things in the real world and then it does have to fire up a you know windows a mac a linux sort of desktop or a browser to go through the thing so code and command execution and browser computer use are like two ways an agent can work and then it's to basically the consumption of the sandboxes to do those two different things are different.

14:10So that's a great sort of sandbox one-on-one introduction to the concept. Help us understand where sandboxes fit in the overall picture of this emerging agent stack that I think everybody's trying to figure out at the same time. So there's different components, there's file systems, there's orchestration, what are the different pieces? I try to think about everything. To me, it's actually quite interesting where, and we'll get this a bit later as well, is like a lot of this all exists in real life today. And so people are like overthinking this. I'm not saying that there's not going to be new products and solutions and technology to solve it, but it's like, it's not a new fundamental way of work.

14:54And so when you think about the agent stack itself, it's like, okay, you first have the models and the models are essentially the brain. That is what it is of equivalent to what a human's brain is sort of. So you tell it something, it replies and understands and whatnot. And then under that, it's like, oh, what are the tools it can do to get things done? Right. And so that can be anything from like any of the MCP or tool calls can do. It could be the sandbox, the computer, like whatever we as humans, we also have a bunch of tools that we use. Everything from a hammer to a computer and everything around that.

15:28Right. So all of those things exist as well. Then there is memory that exists. Can an agent remember these things similar to like, do you remember these things that are there? People that talk about orchestration of agents, it is like management. You manage people, I manage people. Managing agents is similar, is not dissimilar to managing humans. Now, what tools will you use to manage them? Will you manage them in the same way, just send a Slack message or you do linear or whatever it is? we will see or it's gonna be a net new app we'll figure that out but it is of that and of course it's like observability can you see what your teammate or colleague has done and verify that that job is good right and so there's that's sort of how i think about that on a very sort of higher level and then you can break down to different types of solutions people are creating for this yeah let's do some of that the things that i think about like there's also like things that are more solved and less solved where like the models, I'm not saying they're solved, that they won't continue to progress.

16:31There might be different versions of models, but like it's very clear that there is some sort of brain that is there. And we have the leaders and the frontiers and we have people trying to do different things. And that is all said and done. The tooling, I don't know if I would spend too much time. That's all we're talking about now, which is like the sandbox, the computer, the MCP tools. There's like a bunch of them that are being there. the thing that's not solved very well is one so there's memory itself and so memory itself like how do you do this right now the best solution is just dumping things into md files and then you know sort of having access to that and then you know compressing that and how you're going to get that i i we talked about this the other day which was the book of why we sleep right um and in that book of why we sleep it's basically a lot of the how you do data retention and things like that and that's sort of what we're trying to solve for agents agents themselves into that segment and so memories there the thing that is actually quite interesting for me is actually the learning of the model one thing that i didn't understand internalize is that models actually don't learn right now so you use a model and if even if you solve memory it has memory of things So it has context of these things.

17:44And so it can, oh, here's the context. And so it can have a better answer because it has the context, but it doesn't actually learn on the job that it is done yesterday, right? And so who is solving or how do we solve that? Does that, one, does that require something like, you know, constant RL learning, post-training for this? Or is there something fundamentally that we change in the model so the models can do that? I don't know, but it's definitely something that will that's hindering progression today and it's not absolutely clear how we get to that point because it is sort of weird to work with someone that actually doesn't get smarter right like you get a new model but like it's like oh it's like a new one that comes out of it but like you do something five times and it can mess up it literally screws up the sixth time like right so like every single time really and so you can try to prompt her better you can have more context so it doesn't do it but doesn't actually learn i think that is something that's going to be quite interesting to see sort of how that that sort of progresses yeah and from your vantage point memory um is one of those macdan files or like a combination of those versus a database seems like macdan files are just incredibly elegant and simple but somehow feel less robust than what a database would offer i mean not from a personal but i might change this for a personal perspective i also would tend to agree with you that like markdown seems quite trivial for the for the problem that's there and the compression and and how you get to all these things inside of that so the thing that we try to do as a company providing infrastructure for these things is to make sure that we can expose the ability for the agent to access these things in a very simple way and add to them.

19:33But solving it itself is outside of our, let's call it mandate. And in this whole kind of agent harness, I guess, everything that you described kind of falls under the current concept of harness. Where do sandboxes fit long term? Do you think a lot of this gets eventually built into the sandbox or does the sandbox remain the execution layer for it all? Two things. So if you look, there's like the sandbox itself. And again, restating, I'm saying it so many times in this conversation, but if we take like an average worker in Goldman, for example, right? You have the worker, the person, which let's call it the model in this sense.

20:17You have the computer, which it logs into. and the harness is a set of the way it shapes that model that it can and cannot do things and track things so it's almost like hands sort of to speak of that a bit deeper but basically the sandbox does support it and so if you think of a computer again i'll pick on i'll pick on goldman for example i've never worked at goldman but it's my just assumption i've worked at a bigger company so my assumption is that when you log into that computer there's so much software in that computer that logs what you do, restricts what you do, make sure that you don't leak data and do these things.

20:53Again, no prior knowledge of Goldman, just assumptions on these things. And so those are the types of things that we as a sandbox provider will most certainly incorporate into that. But that harness still has its function in that, which it guides the model like, oh, I now know how to interact with this machine. Oh, I know how to do a tool call. oh, I know how to, this is how I work with memory, this is how I change with a model. And so there is value in that. And so we don't take on the harness. That is something we're pretty sure that we never do. But there's things that we do in the sandbox that does help or support that entire system.

21:35Great. We talked about models. where do the model providers like the big AI Frontier Labs fit in this overall picture? OpenAI recently had an agent's SDK announcement. What is that? Yeah, so, I mean, they have a big push into this space where like Cloud Code and their agent SDK has been. And so it's a focus for them to essentially catch up on that segment of the market there. So there's two things, right? There's OpenAI, Agent SDK, and then Cloud Managed Agents. They're different. Yeah, they're different. Unpack that for us. The Anthropic Managed Agent, it is a managed service where you essentially have the model, the harness, and the sandbox all wrapped into one, basically.

22:26And so you have that all managed for you as a service, that entire stack. So that is there. Whereas if you just have an agent SDK, you can use that and run that into any sandbox provider or your own or any machine that you want to run it. And it can connect to, depending on the licensing or whatever, you might connect it to that model provider that gave you that harness, or you might be able to interchange those things. So basically the one is just a harness, the other is the model, the harness, and the sandbox altogether. Okay. And I believe Daytona was a partner to OpenAI for Agents SDK, right?

23:04Like one of the sandbox providers. Exactly. That's part of that original framework. Yeah. Okay. That's the sort of like overall landscape around the agent stack. We're going to go much deeper into sandboxes and how that works from a technical standpoint in a minute. But as a quick detour, let's talk about your story and Daytona's story. story so you this is not your first venture talk about the the prior thing that you did you alluded to some of it like uh that was a little a little bit like replet yeah so we started the cloud ide space basically in 2009 both me and co-founder and so in 2009 for those who are you that were around there was like no docker no kubernetes no vs code and so we had to build the entire stack which was something that we learned how to do and something that we applied today so like it was actually very very very early the only i say that we started it because the company that started before us was heroku which became heroku and killed the id which we probably should have done sooner um as well so but we were the only one that sort of like started and finished finished in that sense later on you had other ones like code sandbox like replet and like others that that had had joined and um stack plates which is now bolt and whatnot and so replet also now changed and bolt into their other directions and Repli's doing really well, we decided to go in a completely different direction when we decided to do their next company.

24:32So we had learned a lot of things on how to build this entire stack. But when we decided to kick off Daytona V1, which is different than it is today, we decided not to go AppLayer, but just be infrastructure. Probably like on the teachings of Heroku, which definitely went into that direction. It's like, oh, we learned how to do all these things on the orchestration level underneath, that seems to be very, very valuable. So let's push on that. And that is how we kick that off. You tweeted, my first startup taught me exactly what not to do at Daytona. And to talk about like selling to developers, no on-prem solution.

25:11Like talk about some of those learnings. Yeah, so we learned a lot of things. One thing we learned is timing. I never understood like timing of the market. Like you have to pick. It's very hard to time the market, but you know if you're in the right time, like fairly soon. ideally you want to be just before the time but that's like very very hard you definitely don't want to be completely wrong we were like two decades wrong or decade and a half so definitely wrong on that one the other thing is like i did not know the difference between a user and a customer and so selling to daytona now is a plg motion 100 the code anywhere product was also a plg motion but the code anywhere product had no enterprise use case value it was all single developer value and single developers do not want to pay for these things whereas a product like what we are it is a plg motion it gives value to a single developer but a single developer probably working inside of a company and so they will pay with their company's card which is very very very different but you don't know when you start out so um those are definitely things that we um understood when we're creating this company and the difference we didn't get into detail but difference between Daytona V1, which was managing, it was an infrastructure product for human engineers in very large enterprises.

26:25It was an enterprise play. Whereas Daytona V2, the sandbox is a, we can call it a Neo cloud, like it's a new cloud. Although Neo clouds are usually for GPU clouds. We did understand that there was an enterprise play to be there. So everything that we had learned of like on-prem, multi-cloud, different ways of managing observability, audit logs, all these things that you need for enterprises, we had already either baked in or understood that that would be needed. So we prepared the product for that originally. Great. You're also pretty amazing at just distribution in general and building a brand and developers, which is fascinating, right?

27:08Because that's one of the typical issues with very technical ventures. people tend to be excellent and deeply thoughtful about product and technology and then distribution comes as an afterthought. How did you become good at it and what are some lessons you can share for technical builders? I mean, it's all about like, I probably sucked at all that. Terrible, terrible. Like I was, so let's take, there's so many ways to say how we did. So one thing is, I think I very well understand humans at scale. like my like what the market wants and or needs and or feels um and so when you do that it's like i did not know that inherently but you sort of start learning that you understand humans i it's very hard for like single humans not so much but more like larger masses humans yeah aggregate humans much better um so that's one thing but the other thing is i was like super shy geeky whatever as a child and the the stage fright terrible terrible like pitching my first startup like I would be so nervous.

28:12I had to talk on stage. Five minutes before stage, I'd go to the restroom 10 times. It's just the pressure of these things. And the thing that happened quasi-randomly is with our first venture, Code Anywhere, we did pitch around all these conferences around the world and whatnot. And it was such an interesting thing that I decided in like with one of the founders of one of these conferences at the after party, quite intoxicated probably. He's like, oh, do you want to do one in Croatia where I was living at the time? And I'm like, fuck yeah, let's go do this. And so I go do, set this up, this conference.

28:49It's in Croatia. 250 people come, which was really big for us at the time. And I had hired an MC that had bailed last minute. And so there was no MC. I have stage fright. You're it. I'm it. I have to go. And so I was literally the MC for two days in a row. And you broke, sorry, it's like you break. It's like you just have to do it. But so what I'm trying to say, this is very different from the go-to-market, but it builds on, it's like, okay, now that I no longer have this stage fright, you start teaching, start understanding what interactions excite people, less excite people, how to bring people and whatnot.

29:23And it's also a problem because when we said we're going to do a conference, I didn't understand what a conference was. And so it's like, how do you break down what is a conference? Okay, a conference is, it's entertainment, first and foremost. And so you have, you know, the show, the show are the speakers, how do you get the speakers? then you have the, for the speakers, you have to sell the audience. Who is the audience? Who's coming there? How do you get them there? And then someone has to pay for everything. The audience does pay for tickets, but the vast majority is under sponsorships. And then how do you sell that to sponsors?

29:49And so you break those things down and then you start understanding incentives of all these different parties and you start putting that together. And so our go-to-market motion, when we, and we did conferences for a decade, more or less. So you sort of like, you have iteration cycles of a year then we started doing two a year we started doing three a year and then every conference got better because the iteration cycle was was much faster to get the feedback loop was much faster and then when we decided to do like daytona and our whole go-to-market strategy was originally around in real life events and so we did dinners meetups drink ups hackathons but all of them were simpler versions of conferences and mind you the conference that we had was like 4 ,000 people.

30:32So it's like, it's a fairly big, it's not 100 ,000, but still fairly, fairly big. And I remember telling one of my teammates that works on the conference business with me now at Daytona, I'm like, there's no way in hell I'm ever doing a conference, ever not doing it. That's the most stressful job in the world. Never doing it again, never. But we'll do these little things because it's easy. And then people will come up to us, how do you guys do these events? How do you do these meetups? Well, when you understand how to do a big conference, like what are the interested parties, then you know how to do a very small meetup.

31:00But the same things apply. It's just like quite a lot smaller, a lot easier to do. And we ended up doing a bunch of these. They were all, we didn't sell. Obviously, these events don't sell, but sold out in the sense of like there's no more space in the rooms that we're doing. We now like attract like partners that, you know, co-sponsor these events for us. Just a great motion for us. And then at some point, my colleague was like, we have to do a conference. Like, no, we have to do it. And so we ended up doing that in the Chase Center. But we ended up doing it in the Chase Center in San Francisco two months ago, I think.

31:33You were there as well. You helped out. And the whole... Which was incredible, by the way. So when you think about marketing as, you know, an effort to make a company, I mean, create a brand and make a company look, you know, big and powerful. That was, you know, a masterclass in how to do that. Thank you. But all of it, all of the entire GTM that we have, we have no other things that we do, like Twitter is a go-to-market. We don't have, we have zero salespeople and whatnot, but it is definitely. So the way I think about this, and I stole this from, or paraphrasing it from David from Century, which is like, there's basically three ways that people pick your product.

Read the full transcript

32:14And one is awareness. Do they even know you exist? Like, if that doesn't happen, there's no way they can pick you. Two is preference, which is, you know, the pricing, the brand, the person, the different features, whatever, like it's preference. And the third thing is like, is there a deterministic thing that you offer that no one else has? So the easy example for the third one is like, do you have FedRAMP, right? Do you have that certificate? If you have like, I can only, the customer can only use you and no one else. but the two above are quite interesting, which is like how to make sure that everyone knows who you are.

32:49And then you work on the, um, then you work on the preference. And so for me, the all, I don't, it's not all I think about a large part of what I think about is if all things were equal. So if our product is equal to everyone else's product, what is the differentiator to that? And so like one, can more people know about you than others? Like is the branding there? Is the, um, the, um, the experience there, experience is a nuance, which some people don't think about, some people do think about, but it's like, do you prefer the feel of this product versus that product? And these are all things that have nothing to do with the actual technical capabilities of the product.

33:29And so my co-founder, Vedran, our CTO, his job is mostly to make sure the product is better than anything else in the market. And my mandate is to make sure that even if our product is equal, that we can supersede that. And so that's how I think about the go-to-market in general. Do you think of customer support and customer service as... All of that. It's all go-to-market. All that is go-to-market. And we've seen this, we've chatted about this, where we've been getting users and customers just because we are so good at that. And so all of it is an experience. So if you think of any experience as a human, you go to like store, restaurant, whatever, it is the entire experience.

34:12It is from your, the way, what is a brand? It's the perception of that brand itself. It's just a perception. And the perception is like, if you go into whatever store, pick your brand you want, the smell, the music, the people, the smile, that whatever you get, all that together is the perception of that brand. And so if you think of that through the lens of a product, which might sound counterintuitive, I don't know, or non-obvious to people. Like I think about that altogether. Let's say we were selling sneakers, right? Like we don't sell sneakers, but the, the, the sneaker, which we won't, man, we won't say other brands, but if you do, you can always pivot to GPU.

34:52Yeah, exactly. Like if you go to the store here in New York, it's like a beautiful store. The people are very nice. Every, everything's aesthetically pleasing and you just like enjoy that entire experience. And it's a good sneaker. Right. And so you have to have that all together. And so that's how I think about this as well, which is, you know, you have to have a good product, you have to have all these things, but all these other things have to co-alide with that. Now, the risk is, and we've seen this with a bunch of startups, like there's no sustenance. So the product is not good and you have everything else.

35:21And that is sort of, that's when you get into trouble. But if the product is good or at least as good, if it's better, it's amazing. And you have all these things. For me, my belief is that that is a key driver to continue growth. great uh you mentioned twitter and x uh a minute ago to which extent is that part of the whole uh the whole it's hard to measure yeah like i've worked in bigger companies and if someone if i was working for my former company they're like how do we measure that i have no idea they can measure that but what i can say is that i was never active on twitter very i have an account since 2009 but never very active until like this holiday season where there was like this semi-viral tweet that went out.

36:04What did you say again? So the tweet was, if you're taking a break these holidays, you're NGMI. You're not going to make it. Yeah. Which I honestly believed. Yeah. Like the tweet was basically, I wasn't even thinking about, there's like a typo in the tweet. Like it's not, people are like, you're rage baiting. It's like, no, I was, my thought is we are a part of this like super cycle right now. And the super cycle does not last forever. And so if you're going to pause by the super cycle, you are seeding market. Like that is what you're doing, right? And also it's like, it's very much towards, you know, founders and executives and tech leaders.

36:39It's less about the individual person that can't have an impact on all. In that, like people have to take their breaks. And so people took that in all different ways. They're like two poops, two poops. Like you should die, you're a capitalist, you're whatever, all these things. And the reason why it became so popular, like it was kind of popular, which was some person said, dude you never i've never heard of you or whatever and then i said basically that's why yeah to keep working and that was perfect perfect and now again not i was like instant reply on that and so basically not talk too much about that is that from then i had noticed that people even if they like your tweets or don't or are completely opposed or not has no bearing on i should say it has that positive bearing on their thoughts of your company.

37:32So we have found that people that even are completely opposed to the posts end up being customers, assuming that they need the product line, not everyone is there. So it was quite interesting to me that, oh, just because someone doesn't agree does not mean that they won't like that because it's all generally awareness. and so that has been something that I've spent a bunch of time trying to catch up to you on the Twitterverse. That's super interesting for any AI builder or AI founder. So the PLG motion ultimately to create inbound is a combination of all the things. So we talk about conferences and meetups.

38:15Twitter is important. Is there anything else that people should know? We talk about customer service. You mentioned no salespeople as of now. We have no salespeople now. Yeah. So I think like the, how do you experience the product? Right. So like once you've seen the product, how do you experience the product? It's like your door to the product, like the website, the login, the whatever. We can fix a lot of these things to be very clear. I'm not saying we're the best at least, but there's that. And then what are the feature sets are in there? Can I get things easily? Our SDKs are really, really, really good.

38:45Like people really like, like the ergonomics of them. So it's like really good. That part is great. And then the thing that, and this is even public on Twitter, all our case studies that we've done, and we outsource the case studies to third parties, so we're not part of this, is that we reply very, very fast. And so this is a very public thing, and it's something just core to who I am and how I learned to work. Because one of the, let's call it first real jobs I had was a system admin. So my job was like to fix the printer and computer and whatever. And so one thing that I learned, and I try to talk to my entire team is like, one thing you have to do is like, one is the first response very fast.

39:24Just the first response very fast. Like people are then calm. They know they have transferred their problem to someone else and someone has acknowledged that. And so when they do that, they feel rested, right? And then you just have to promise that you will solve or get back to them at X amount of time, whatever that is. And they were calm till that moment. And the thing that you have to do is one, either solve the problem until the given time or call them, message them, whatever, and track them before that time and state a new time. That is the solution to support. That is it. There's nothing more than that.

40:00There's obviously, there's like key things if you're actually down and don't work and the person can't work, that's a completely different thing. But every other non-critical problem and you have the most happy customers in the world because they don't have to think about their problem anymore. You think about their problem. And so you can like keep that on until you get it done. But the key part is if I said that something's going to be fixed or deployed or done or whatever in two days, a day and a half later, if it's not going to be in two days, a day and a half later, I said, hey, this is going to be delayed two more days.

40:32And they're fine. You thought about it before they thought about it. And it's not magical. I mean, it's a magical experience, but it's not very, it's not a complex thing. but I found that people don't understand that intuitively. And so that is what I learned back in the day. And that is what we do in the company today. And that's why we have very, very happy users and customers. Okay, fascinating. Now let's switch back to the more technical stuff. So we talked about the concept of sandboxes. Just walk us through what a sandbox is technically. Yeah, there's a lot of things. There's a lot of things there we can talk about.

41:08The way I think about Sandbox is the ergonomics of consumption of compute. And so, because there's a lot of different layers on that. I basically put it into three different layers, which is the infrastructure, the primitive, and the tooling. Those three things together. And what I mean by infrastructure is like, does it spin up very, very fast? So like we spin up in 60 milliseconds. Can you spin up a lot of them at once? So the rate of creating them. And so we can spin up 50 ,000 in 70 seconds. So a little less than a minute and a half. And then once they're up, how many can you keep running?

41:48Like we can have, we have customers that have billions a day. So those are like the infrastructure parts that are non-trivial. Why does it matter how quickly you can initialize a new sandbox and how many you can run? Again, it depends on the user and the use case. So if you're like a long running background agent, again, everyone prefers it to be fast. Like no one wants to wait. Like you don't want to wait for a reply. Everyone wants to be fast to be very clear. So the faster, the better. But generally there's an actual reason why you want it very, very, very fast. And that is especially if you so for I was gonna say for a background agent, a background agent might work for like 10 minutes or an hour or whatever.

42:23So the incremental millisecond might not matter. but I still believe that from a user perspective, even a second of waiting is kind of uncomfortable. You don't want that. And so it's 60 or 90, maybe less so, but there you want that one, two seconds for sure under that. But the more interesting part where that is really, really important is for the researchers, where when you're doing reinforcement learning, you basically have a lot of GPUs and the GPU is more expensive than the CPU. So the vast majority of sandboxes are CPU boxes, to be clear. So they're the computers that we all work on. There might be a graphics card in there, but basically it's the RAM, the CPU, and the hard disk that's in there.

43:02And they are cheaper or less expensive and easier to get, at least for now, we'll see for how long that lasts, than GPUs. Which means you want your GPUs always to be at maximum utilization and the CPUs can then idle if need to be idled, but you don't want the GPUs to idle. And so what that means is you want to make sure that the CPU machines, the sandboxes, spin up so fast because you don't spin them in a training run you don't spin them all up you spin them on like depending on how many you have it's like a thousand at once they do the task then do the next one the next one next one and so between turning off and on the cpus you want that time to be as short as it can so that the gpu utilization does not go down right and so that's why that's very very very important and obviously the number that you have concurrent so depends on so on the background agents, let's say just for the sake of people know lovable.

43:56So the amount of users that they have is astonishing. And so each user for every task, and they can have multiple tasks, multiple agents, they need a sandbox. And so imagine, I don't know their user number in the millions, whatever. So you have to have millions of sandboxes up and running. On the RL side, you can have people, smaller labs that will have concurrent 5 ,000 or 10 ,000, But we have a request today for 5 million concurrent looks. So 5 million at one point in time, right? So being able to handle those types of things is part of that infrastructure segment. The other two are the primitive and the tooling.

44:31And so the primitive is essentially, let's call it the computer itself, the isolation, the VM, the container, the micro VM, the isolate, the whatever. We can get into details what those are, but there's different shapes of them. And also what feature set do these things have, right? And so for the most part, a VM that you would find in, you know, AWS or whatnot is usually much slower to start, much harder to get those spikes. But also when you get a certain size, they're usually that size forever and forever. And then we just give one example, whereas in Daytona, you can define a size, the CPU, the RAM, the disk, and then while it's running, you can resize that.

45:15So if an agent gets to, you know, use the entire memory or 100 % CPU, the sandbox would, or VM would die. In our case, we can expand that. And there's other things. It's not, for us, it's not just a Linux CPU box. It can be a Windows, a Mac, an Android. It can be, you know, they can have a GPU in there. It can not have a GPU depending on what you need. And so that's the primitive. And the last thing is the tooling. The tooling is, the tooling is either there to support the agent, to be better at its job, or to be guardrails on the agent to stop it from doing dumb things. stuff basically so you can think of it we have a bunch of tools so headless you know terminal and file operations and other things that enable it to use less tokens get the job done faster but also guardrails like secrets manager and a firewall and all these other things and so the combination of those three the infrastructure primitive and tooling essentially creates what we call a sandbox so what's ultimately what's the difference between a sandbox or a container VMware, VM, micro VM.

46:16All the things, all the joy. So most sandbox providers are firecracker VM, micro VMs, most. Let me remind people what firecracker is. Firecracker, it's an isolation privilege. So it is essentially a type of VM, a virtual machine, that's very stripped down. It was made by the AWS team. It's used inside of Lambda. And so the idea, which are like functions, and so the idea was to have something very very fast that can spin up and stateless that that was the entire time and they were very very ephemeral and so they were used for you know when your website has a large if it's black friday a bunch of people hitting your website you can spin up a lot of these very very fast like handle the load and enable all these humans to to look at the website and buy whatever they were buying right and so that technology is there and it's a very stripped down version of a full on VM.

47:14So it has less features there, but the things that it does, it does very, very, very, very well. So you can spin them up fast. It does that because it can sort of lazy load things into memory. You can do point in time snapshots. You can do all these nice things. The things that I can't do is it can't run a, let's call it sandbox now. It can't run a sandbox with a GPU. It just doesn't work. So you can't have a firecracker. So if you want a GPU, you have to do something which is a cloud hypervisor or a keymu spelled Q-E-M-U, which are two different types. Keymu is almost a full VM. So it has all these different things inside of there.

47:49So you can't do that. But let's, for example, if you need to run, we were solving this for one customer earlier today is they need a machine that also has an Android device inside of it. And so the only way we could solve that was inside of one of these key moves. Like it can't work in a firecracker, it can't work in a container, it can't work in anything else. And so there's different types of there. These are like the micro VMs of the world. There's also containers. And so container, most people know Docker, which is there. Daytona originally started running Docker containers. We do have them.

48:18The problem with Docker is that they're much less secure than a VM. The thing that we do in Daytona is we harden that with Sysbox, which is VM-like isolation that wraps around that. It's really good for handling density of these machines. It's very fast. It allows us to run Docker and Docker. There's a bunch of features that it's very, very good for. And so there's differences there. There's also like isolates, which we've talked about, or like other app containers, which is abstractions of these containers, just you're running a single app. And so these different things all exist in the world and they can all be theoretically sandboxes with these other two things that are there.

49:06And our original idea is to get back the original conversation, most sandbox providers started as firecrackers. But as we start seeing that agents have different needs in different use cases, they are going to need all these different shapes consumed through one interface. And so we at Daytona now support containers and all the micro VMs, depending on your use case. The user doesn't know the same ergonomics. It's like, oh, I need a Windows, it'll spin up this one. I need a Linux, I'll spin up that one. And so we will continue to add these things inside of the same infrastructure. Sometimes it'll be slightly faster or slower, but the concurrency will be there.

49:47All the tooling will be there. And so everything that you would come to know and like about a Daytona will be there, but it'll be different sizes and speeds and configurations depending on your use case. Why are snapshots and orcs important? There's a lot of reasons. Let's get into the commercial one, which is probably the one that people most think about, which is pausing a sandbox. So if you are an app layer company and you have 10 million users and your 10 million users spin up, let's just say 10 million sandboxes, just to make the math easier, these things cost, right? Because they're running, they're using CPU, RAM, and hard disk the entire time.

50:26The most expensive thing is the CPU. Second, the RAM, disk is almost free. It's very inexpensive. And the user sends the agent to do something. the agent does something and now it's waiting for a reply it could be waiting for the human or it could be waiting from a service depending on what it's trying to do you ideally don't want that sandbox to run idle because you are now regardless if you're a customer Daytona or running your own like there is a cost to running these things and so what you want to do is pause that and wait for a reply either from a service or for a human and then resume and the feeling and experience is that it never turned off so you feel like it never turned off but it actually did turn off.

51:08And so that one is from a compute management perspective, if not from a cost perspective is probably the first reason. The second thing is you can enable your agent to take multiple paths. So your agent can say, oh, at this point in time, I'll take a snapshot and then I will either continue and be able to roll back to this point, like a point in time, or at this point in time, I will replicate the sandbox and have two sandboxes and I can try two at the same I mean, obviously it can be more, but right. But those are the reasons why you would have that there. You mentioned performance and speed.

51:43And in connection with that, you said that you had to rebuild your own scheduler. Yeah. So what is it? First of all, what does a scheduler do? And then how did you go about it? Every cloud that exists today, NeoCloud, Hyperscale, or whatever, they are all built on servers, right? Like metal machines. And like historically way back in the day, we actually stacked these data centers, these data centers, me and my co-founder a long, long time ago, but they're all just like servers, like CPU, RAM, disk, they're computers, basically. And on top of these computers, there is a software stack that everyone has built for their own reasons.

52:19So AWS has their software stack. Cloudflare has theirs, which is very different. Everyone has their own on top of that. And so basically what you're trying to do when you think of these machines, these servers, basically what we do is we cut them up into small little machines and then give you that sandbox. And so you have these small, on these big servers, you have these small little machines that can run for a minute, a second, three hours, whatever, and you don't have to worry about this. And so basically the scheduler or the orchestrator is the one that basically says, oh, after you send me a request, I send you to this server and turn on that sandbox, a little computer there, and then I turn it off or I kill it or I snapshot it.

53:01It's the management of all these things there. And most of our other companies in the space basically have off the shelf schedulers. So it can be like Kubernetes or Nomad or whatever. And when we decided to build Daytona, the sandbox product, we inadvertently with very naivete were like, oh, this stuff doesn't work because we had worked with all we had built our own. Then we have used Kubernetes. We're like, we've seen all these different things and we knew what was good and was bad or what was use case or not. none of these were made for these like super fast, stateful, long running machines.

53:35And so we're like, we have to do this again. We have to build it again because the way we thought about it, which was not known at the time is, because at the time sandboxes were very ephemeral, similar to a Lambda function. And we're like, no, why would your sandbox be ephemeral by default? Like your laptop, you don't want it to die. Like you want it to work until it's done and it has to be very fast. It has all these things. and so my co-founder mostly built this sort of initial version of our scheduler and that scheduler of ours is the basis of everything that we do and it gives us a lot of things like outside of just like the performance it also enables us to do things like have four different isolation provide like we are the only company right now that you as a user don't know but like we have depending on your use case we will spin up any one of these firecracker cloud hypervisor whatever to get the job done.

54:27So whatever is needed, we'll get there. The other thing, the benefit is we built this, we have co-location providers, so we have like bare metal machines. So you can think of us as our own cloud, but we're more and more akin to now, like the NeoClouds, like the base tens and like the fireworks, because for us, and we've got to this question yesterday because we have a large request of number of CPUs, is like, we can't get enough. And so how do we solve that problem? Well, we've already solved it. where our scheduler can be attached to any CPU machine, any server that we think is good enough, and it becomes part of our cloud, essentially.

55:05And so why I say Base 10 and Fireworks is that both of them run on more than a dozen compute providers, depending on what you're going to call them. And so we can do something similar because we've created this in such a fashion. So that has given us, I believe that has given us sort of an advantage there. and is there a fundamental technical difference between ephemeral and long running so if you have a an agent that runs for 24 hours how does that translate in terms of sandbox requirements the reason most sandbox environments do not run forever or even it is because it's a technical problem if you think about servers underneath these servers also have to be managed right and maintained and so if your sandbox can run forever that means that you can never reset you can never reboot the underlying server you can't update it you can't patch it you can't do all these things without turning off all the sandboxes and so the way you solve that the easiest way to solve that is having sandboxes that have a termination time it's like they will only last an hour 24 it doesn't matter pick your time and then if you decide that you have to do something with the underlying machine, you just flag that machine as non-schedulable.

56:22And so at some point in time, there's nothing else on that machine. You can do whatever you can fix it. You can like reboot it. You can do whatever you want. Easy peasy done. Right. And so because historically, most of the workloads were fomeral, you didn't have to try to solve that problem because you didn't care. Like most workloads, like a Lambda function usually runs what? Five minutes, 10 minutes, like whatever. It's not a problem, right? You never had that restraint or constraint. now that you do, to have something that can run forever, you have to be able to live migrate the sandboxes itself between the machines so that you can reboot and manage these machines.

56:58And I understand that most people, it's interesting, even technical consumers are like, oh, all computers elastic. Well, yes, your consumption of it from us or from AWS is elastic, but someone actually has to like screw together a server, turn it on, boot it, make sure it works, right? And have enough of that there. so that is the difference from from our perspective and why most start off ephemeral is just it's an easier way to do it we believe and we've seen this i think two and a half percent of our revenue now comes from or two and a half percent of our sandboxes run longer than 24 hours but it's a large percent it's like 20 percent of revenue so like it's not an insignificant part to be there just because it can run for a much, much, much longer time.

57:44On the user perspective, just to finish on that, sometimes users actually want it to be ephemeral. One is they don't want any data retention. So they want it to run and die. Like whatever was run inside of that, they got the data they wanted and they want that data dead, dead. And so they flag it with us as ephemeral, which means that we won't force it to die. But when they kill it, it's no longer it goes into the either. It does not exist anymore. Listening to this whole technical deep dive part of the conversation, the thought crosses my mind that a lot of people seem to think that they can create their own sandbox and it is not that hard.

58:22But again, listening to all of this, I'm like, good luck. We've seen this multiple times. And I'm sure you've seen this across other companies as well, where, you know, everyone's like, oh, I can spin up a firecracker on a machine. Yes, you can spin up one, but spin up a million is a problem. Adding all these features is a problem. Having throughputs, different use cases, they're all different problems. And what we already see is like companies that we're like, oh, we'll do ourselves, come back three months later, six months later, eight months later. It's like, oh, actually, actually now we can't do X and Y.

58:52And so there's a bunch of things on technical side, which we can like address. And I think we can finish with that is that the thing that Daytona does is we, if you think of most sandboxes, this is a performance thing, Most sandbox and most VMs actually, they use the CPU RAM from the server, and then they use the hard disk from a network drive. You do that for a bunch of reasons. Your life is so much easier if you manage that. Why is it easier? Well, it's easier because if a server dies, then all the data is on a shared drive, and I can reboot that server really fast without maybe even the user noticing.

59:33And there's no data loss. maybe something in the memory, but nothing there is lost. So it's very, very easy to manage. Whereas what we do is we actually use the CPU, RAM, and hard disk of the machine, which means we have so much more overhead of if that machine dies, how can we reboot it back? You have to do backups that users don't see. You have to do all these things that happen there. But what that means is that the IOPS, IOPS, the speed with which you can move data to the hard disk in which the CPU and RAM interacts with. From a network drive, just to give people numbers, is in the hundreds of thousands.

1:00:15If it's a local drive, it's in the tens of millions. So it doesn't matter if you know what the numbers are. It's just extraordinarily large. It's very, very different. And so you have use cases that no one really cares about that speed, which is fine. We have new customers that have a use case. They're like, it has to be that fast. Right. And so when you get to the point, it's like, oh, I can use this, but wait, now I need this performance. Oh, now I can't do that. Right. And so to the complexity of sandboxes, if when people are starting out, their use cases are like, oh, I just need to run code.

1:00:49You can run that on anything. It's very easy to do, except scale. and then when you get to scale and then you need performance because now, oh, agents can do a lot of things, then it starts like breaking that sort of mold. And security. Yeah, because we didn't even touch the security part, but yeah. Great. Zooming out as we close this conversation, you mentioned something that caught my attention. So we have been in a very well-documented GPU shortage, but it seems that we might be heading towards the CPU shortage. Yeah, it might be happening. So one of your former guests and was at our conference, Dylan Patel, like sent me an analysis.

1:01:27They wrote the report. And I believe it's somewhere like October. We have no more CPUs. And people hadn't thought about it. Like it's all GPUs, all GPU intensive, then memory. But basically now, now that agents need all these computers, like it is very much needed. now that RL is the way that we've gotten most of the progress in models over the last 18, 20 months, you need all the CPU machines to do this. Now that the GPUs are more powerful, they can do in parallel more of these CPUs, CPU boxes. And then you see this, there was a tweet that I put a while back and is like, you should definitely invest, not financial advice, you should invest in Intel and look at their stock price, right?

1:02:11And so like, there's definitely, need for that. And it's something that people didn't understand or intuitively that you would need there. So it's something that we think about quite a bit is like, can we, alluding to also how our product is created, how do we make sure that we can match the demand of customers for the CPU that they have? Because I don't know it goes to the extreme to where GPUs are, because that is very, very, very extreme. but it is quite high probability that there will be shortages of CPUs going forward. How do you think all this agent stack evolves? Does it feel like we have all the core pieces together or do you think like in two years the overall architecture may look very different?

1:03:00So we talked about the stack there and things that have to be done. The identity part does need to be solved. Again, it's basically solved, but not solved. So there's a lot of things to do there. I'm actually also wary, wary. I think about people think that the models that we have, the technology for the models that we have is the end state. And I really, really don't think that is true. Like I think, and that would probably be the most surprising thing is like we have a new type of model that is just better or different than what we have now. that either one needs less compute or doesn't use GPUs or doesn't, like, I don't know.

1:03:43There's things that we've seen where, you know, we've seen things with wetware where people can have like fake human brains and it learns how to play Doom. You have the, what was it? The fly brain that was replicated inside of a computer that then works like a fly. Like, can you replicate a human brain? And I say this not to replicate myself, but can you make a generic human brain? And then it's like an, it's an AI that is essentially equal to a human AI. And so what does that mean for like all technology? I mean, I'm saying, I'm now saying extremities for people not getting me wrong. I don't think that happens tomorrow or like that I'm completely crazy, but like, I don't think that this is the end all of the state.

1:04:24The question is, does that happen sooner or later? Or do we keep our transformers, the vast majority of intelligence going forward? And if it is, then it continues probably going directionally where it does. But is there something that happens that changes that fundamentally, something that I think about? Because then all our bets, yours and ours, are very, very different, right? Okay. Well, Ivan, that was fabulous. Thank you so much for spending time with us today. Thank you for having me. We're great. Hi, it's Matt Turk again. Thanks for listening to this episode of the Matt Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from.

1:05:09This really helps us build a podcast and get great guests. Thanks and see you at the next episode.

From the publisher

If AI agents are the new digital knowledge workers, where exactly do they do their work? In this episode of the MAD Podcast, Ivan Burazin joins us to unpack the emerging infrastructure stack for AI agents and explain why every agent needs its own secure, stateful "computer." We explore the technical realities of sandboxes, dive into why legacy, stateless hyperscalers weren't built for these new workloads, and break down the mechanics of microVMs and custom schedulers alongside a contrarian prediction on an impending CPU shortage. Finally, Ivan delivers an absolute masterclass on product-led growth, community building, and go-to-market strategy for technical founders.


(00:40) Intro

(02:13) What is an AI agent sandbox?

(03:17) Security risks of running agents locally

(05:17) Stateful vs. stateless hyperscalers

(07:04) The history of cloud IDEs and the end of localhost

(09:45) Do all AI agents need a sandbox?

(12:26) Sandbox use cases: RL evals & background agents

(14:10) Unpacking the emerging AI Agent Stack

(16:20) The unsolved problem of agent memory and learning

(19:37) Where sandboxes fit in the agent harness

(21:35) OpenAI, Anthropic, and agent SDKs

(23:06) Ivan's founder journey: From CodeAnywhere to Daytona

(26:59) GTM strategies and building developer communities

(33:48) Why customer support is your best GTM strategy

(35:34) Leveraging Twitter during the AI super cycle

(40:50) The technical anatomy of a sandbox

(41:53) Why fast spin-up speeds maximize GPU efficiency

(46:09) Firecracker, QEMU, and isolation primitives

(49:58) Why sandbox snapshots and state forking matter

(51:40) Why Daytona built a custom scheduler from scratch

(55:24) The challenge of long-running stateful sandboxes

(58:10) The build your own sandbox trap

(1:01:03) Why AI agents might trigger a global CPU shortage

(1:02:46) The future of the AI Agent Stack

More from The MAD Podcast with Matt Turck

All 44 episodes
The Agent Harness: Building Secure Sandboxes for Autonomous AI WorkloadsThe MAD Podcast with Matt Turck · 1 h 5 min
Listen in VO