In short
Docker Sandboxes (SBX) for AI coding agents—isolating agents in micro-VM “sandboxes” to control filesystem, network, and credentials, addressing security threats from agents that can mutate environments.
Guests
Mark Cavage, President/COO of Docker; previously worked at Stripe, AWS, and Oracle; long history with AWS (generation one). Gregor Vann, security-focused technologist; former CTO across cybersecurity/cyber insurance/software engineering; based in Singapore (van.hk, LinkedIn).
Key claims
Containers’ immutability breaks for agents because agents need runtime mutation (download packages, write files). SBX preserves Docker ergonomics but uses micro-VMs (emulates hardware, separate kernel) for stronger isolation than shared-kernel containers. Threat model focuses on preventing host compromise and data exfiltration; secrets are never exposed inside the VM via proxy “placeholder” tokens.
Notable examples
“YOLO mode” agents that run without constant permission prompts; prompt-injection leading to key/file theft; Slack read/write risk; agents transferring money without keys; micro-VM startup ~1s on laptops and ~100ms in cloud.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to Coding Agents and Security
0:00 to 1:35
Learn about the capabilities and risks associated with coding agents and Docker Sandboxes.
“The most useful coding agents can mutate their environments by downloading packages, writing files, and connecting to services across the network.”
Mark Cavage's Journey to Docker
1:44 to 3:10
Discover the career path of Mark Cavage leading to his role at Docker.
“And Docker, I'm sure a lot of our audience know pretty intimately.”
Understanding Docker
3:10 to 4:47
Gain insights into Docker's history, technology, and ecosystem.
“but at its root it's containers and basically it's as i always tell people if you go back to the very first genesis when Solomon Hikes, who founded Docker, solved a problem.”
Docker's Expanded Product Suite
4:47 to 7:55
Explore the range of Docker products and recent developments in their offerings.
“So that's maybe a longer answer than you were looking for.”
The Importance of Supply Chain Security
7:55 to 9:21
Learn about the rising threats to supply chain security and the significance of developer laptops.
“So, and then there's AI, which we'll talk about for, I'm sure, oh, quite a bit.”
Introduction to Docker Sandboxes
9:21 to 14:00
Understand the concept of agent sandboxes and their applications in Docker.
“Today, yeah, we are going to speak a lot about AI.”
Understanding MicroVMs vs VMs
14:00 to 17:52
Learn the key differences between micro virtual machines and traditional virtual machines, including their efficiency and resource management.
“I mean, it's in the word there is micro perhaps.”
Docker Sandboxes and Agent Architecture
18:50 to 22:28
Explore how Docker sandboxes operate, including network and file system management for agents within a secure environment.
“which is the idea that you do actually get, like, rather your agent does actually get its own Docker or daemon network and file system inside that sandbox.”
Cold Start and Agent Initialization
22:28 to 25:56
Understand the cold start process for Docker agents and how it is optimized for user experience.
“that kind of all work yeah so the thing when we're talking about the architecture i admitted this The SPX command today that you download off the website is clearly optimized for the laptop case.”
Product Development and Cloud Code Integration
25:56 to 28:00
Learn about the strategic decisions made during the development of Docker and the integration of cloud code solutions.
“The core thesis from us is, again, like most people using Docker today, when they think of it.”
Show all 22 chapters
Introduction to OpenCode and Agents
28:00 to 28:30
Learn about the ethos behind OpenCode and the current state of AI agents.
“So it's open any model, any harness, any agent, any operating system, any cloud.”
Understanding Threat Models
28:30 to 31:28
Explore the key threat models concerning AI agents and their security implications.
“I think a lot of people would be quite interested in this one.”
AI Governance and Company Policies
31:28 to 36:38
Discuss the importance of AI governance and how companies can enforce security measures.
“And now hand in hand with that is, okay, you're a company and you actually have many people.”
Secrets Management in AI Agents
36:38 to 39:03
Delve into how secrets and credentials are managed to ensure security in AI interactions.
“It's just a fancy way of basically saying, don't do really stupid stuff, user.”
Open Source and Product Development
39:03 to 40:18
Explore the balance between open source and proprietary development in AI.
“By definition, we are not able to stop that, at least not today.”
Docker's Internal Use of SBX
40:18 to 42:00
Learn how Docker is implementing its own products for internal use to improve testing and development.
“But the answer to your question is today, it is not a hundred percent open source.”
Building Autonomous Agents
42:00 to 44:01
Learn about the integration of autonomous agents in software development workflows.
“We run all the things we, as a company, use as an MCP through that.”
The Future of Hosted Solutions
44:01 to 45:27
Explore the potential of hosted versions of development tools and cloud integration.
“As we sort of start to vaguely sort of cruise to the end of the episode, I'm curious about the future of this.”
Governance in AI Usage
45:27 to 48:03
Understand the importance of governance and safety in AI agent deployment.
“I think many things actually need access to the laptop that are like, oh, say the microphone or the camera or again, Excel files or whatever else it is.”
Mobile Integration Challenges
48:03 to 49:10
Discuss the feasibility and challenges of running Docker on mobile devices.
“So that's the digest radio safe version of the roadmap.”
Focus and Roadmap Decisions
49:10 to 49:50
Reflect on the importance of focus in product development and roadmapping.
Mobile Integration Challenges
50:02 to 50:16
Discuss the feasibility and challenges of running Docker on mobile devices.
Transcript
Automatic transcript. May contain errors.0:00The most useful coding agents can mutate their environments by downloading packages, writing files, and connecting to services across the network. However, that freedom also presents dangers and promises to usher in a new wave of security threats. Docker recently announced Docker Sandboxes, which gives each agent its own isolated micro-VM while preserving the familiar ergonomics of a container. A standard container shares the host's kernel, but a micro VM emulates hardware and runs its own kernel, giving a stronger security boundary around code that cannot be trusted. Mark Cavage is the president and COO of Docker, and he previously worked at companies including Stripe, AWS, and Oracle.
0:44In this episode, Mark joins Gregor Vann for a wide-ranging conversation that includes why agents break the immutability assumptions containers were built on. how micro VMs differ from both containers and traditional VMs, and the still unsolved challenge of giving agents scoped, trustworthy access to sensitive services and data. Gregor Vand is a security-focused technologist, having previously been a CTO across cybersecurity, cyber insurance, and general software engineering companies. He is based in Singapore and can be found via his profile at van.hk or on LinkedIn.
1:34Hello and welcome to Software Engineering Daily. My guest today is Mark Kavage. Thank you for coming on today, Mark. Yeah, thanks for having me. It's great to be here. Yeah, so super exciting. You're the president at Docker. And Docker, I'm sure a lot of our audience know pretty intimately. We'll get into what Docker is for those that don't know in a second as well. As we like to do, you've had a pretty like storied history in the tech world or through tech companies. Could you just give us like a kind of brief story of how you got all the way to Docker? Absolutely. Well, it's a funny story. There's two other characters here, Don and Tushar.
2:11We were part of the, let's call it generation one of Amazon Web Services back in 2005. Oh, wow. And we're all Office mates back when 80 of us was, I don't know, somewhere around 100 to 200 people, somewhere in there. I don't, time all blurs it together. We've gone off and done different things over the years. We had also gone to Oracle and built Oracle's cloud, which was, I always tell people it's a weird thing to do voluntarily, but we did that. At some point in the last couple of years, Docker was looking for new leadership. And as I was joking, fortunately for Docker, they found us. And so Tushar had found his way to Docker two and a half years ago, three years ago.
2:42And then as the band of Merry Brothers do, sort of called on his friends and here we are. So then Don and I joined February of last year. yeah let's start with docker is and then actually might just kind of pull back for a second after that but yeah what is docker full stop for those that don't know yeah docker is a ubiquitous brand i think many of us grew up on it over the last decade docker equals equals containers is the shortest way to put it but there's docker the company there's docker the technology there's almost like docker the community and ethos it's like it's many things roll in one word but at its root it's containers and basically it's as i always tell people if you go back to the very first genesis when Solomon Hikes, who founded Docker, solved a problem.
3:19It was, the cloud exists. I have code on my laptop. I would like to get my code from my laptop to the cloud. This is really hard. Docker is a unification of packaging technologies to go make it easy to build and then ship software reproducibly from one place to the other, as well as a runtime, which we'll talk about in a bit, that allows you to run those said containers. And then cleverly packed in, there are several interesting primitives. Historically, it has leveraged OS virtualization in Linux. So you have a container is really, at the end of the day, a Linux namespace and gives you sort of isolation at the kernel level.
3:54That's sort of like multifaceted packed in there. Anyway, there's Docker technology. There's Docker Hub as well, which is the ecosystem where people produce all kinds of packages. We Docker see the ecosystem with things like Node.js and MySQL and Python and Postgres and every other popular open source thing you can think of. And then tens of millions of people go and publish to Docker Hub every day and produce software that other people can depend on and build stacks and so on. And so now, as you might imagine, a decade later, there is a very rich and vibrant ecosystem. There's a fascinating stat I had read.
4:27It's not ours. I wish it was, where somebody had done a survey and concluded, I think was like 92 or 93 % of every company on earth at this point runs containers in production somewhere. So it truly is now this ubiquitous way that people build and run software because it's kind of the lingua franqua. It gets you portability. It gets you from laptop to cloud. It gets you from cloud to cloud. It runs everywhere. So that's maybe a longer answer than you were looking for. No, no. There's never such a thing as a too long an answer I see daily. So yeah, I mean, I think it's one of these in computing, we have names for things that we always try and sort of analogized to something in the human world, if you want to call it that.
5:02But a container really is that. I think it's such a great metaphor or thing, which is you have this portable, literal container. Think of container ships that you can move it from ship to ship and it's the same system. And that was always the thesis here, right? Yeah. And importantly, you can pack all your stuff into it and it gives you a way to get your thing from point A to point B. And what it lets you do is encode all of your dependencies and take no, you don't need to worry what the host runs. In the old days, I am old, you know, it was always this like protracted system to keep the build system up.
5:35And like, oh, it worked on my laptop, but failed in CI. And then it worked in CI, but then I failed in prod because there's some variances across whatever the kernel is and whatever the operating system is and whatever glibc is there. I'm not dating myself at this point, but these things are all still with us. But that problem really has been largely just solved by Docker because it allows you to encode all the dependencies all the way down to effectively the base Linux layer and just not worry about that anymore. So you put all your stuff in a container and then you ship it wherever you want on whatever boat you want or whatever crane you want.
6:07So, yeah. So we're going to get, I mean, today's topic is specifically is a newish product that you guys have produced or on agent sandboxes. We're going to get to that very shortly. Could you maybe just also walk us through like, what is the general Docker products suite if you like nowadays? Like I think for those that maybe kind of started with just like a Mac app and then maybe haven't appreciated there's a whole bunch of our bits to Docker now. So like just before we get into the latest product. Absolutely. I mean, we'll definitely spend the majority of time, I think, talking about AI. As you might imagine, that's where a lot of the focus goes.
6:41But the company, the core Docker desktop and CI tooling, we've been around for a decade. We have a rich ecosystem of things. The biggest thing above and beyond Docker desktop, which I think most people are, when they think of Docker today, I think they think, I downloaded Docker desktop and I did my stuff. And there's all kinds of things that are in there around AI assistants called Gordon. There's Scout for doing dependency scanning, et cetera. The biggest thing I think we've done in the last year besides the things around that is really what's in the, I'll call it the hardened image space. So, you know, again, we have in Docker Hub, again, the hosting package manager for the ecosystem.
7:13We see the ecosystem for the last decade with several hundred images that help people again, get started with Java, Python, Node, MySQL, et cetera. The wonderful thing with open source is everybody can contribute to it. The problem with open source is everybody can contribute to it. And then some people contribute malware. And so many companies struggle to manage dependencies and manage their open source dependencies specifically. So Docker hardened images is a way to go. It's a very opinionated, stripped down, Debian and Alpine derived image where the contract from us is we'll never run as rude.
7:45We do all the things you'd expect someone to go do to run software securely. but importantly we also there's an sla available with it where we actually are able to keep cvs out of it and promise you that within seven days of a vulnerability we're able to go patch it and so that's probably the biggest i'd say the biggest thing besides ai that's been new that we've done in the last you know year to year and a half has been really docker hardened images because it's sort of above and beyond the tooling and the runtime has been sort of the okay this is the content that really now i put all my stuff in the container and all my stuff is actually much safer than being, you know, it can make it clean, not dirty effectively.
8:17So, and then there's AI, which we'll talk about for, I'm sure, oh, quite a bit. Yeah. The hard and containers piece, super interesting. We've done a couple of episodes on that sort of concept and it's a very interesting space in terms of supply chain attacks and them only increasing is what we see. So, yeah. At this point, I think the funny thing, everybody goes through these, the industry goes through these trends and like four or five years ago, I swear, if you asked anybody, they'd say, oh, the laptop's dead. We're all going to use cloud environments, et cetera, et cetera, et cetera. Nobody cared.
8:45You couldn't get the time of day to talk about it. And now I think the developer laptop point, in fact, is the most juicy target for hackers on earth because it's like, great, if I get my keys there, I can go get massive distribution through the open source ecosystem. And so, yeah, just supply chain security all up. I think it's a very broad and sprawling topic. You could spend two hours talking about all the ins and outs of it, but it's a real problem. And And when we look at the last six months of attacks, they practically speaking all started with got access to GitHub Actions or got access to laptop.
9:16And then basically once there, got credentials and the world's your oyster. Yeah, it's kind of scary what we've seen over the last, especially a couple of years, just some of the sophistication of how people have gotten to others' laptops to do such a thing. Today, yeah, we are going to speak a lot about AI. I always like trying to time it like how far do you get through an episode before saying AI now. But today's topic is that, and we're going to be talking about Docker sandboxes. So I think to kind of set the scene, though, we're talking about the concept of an agent sandbox in general. This isn't as a concept is not unique to Docker, but so could you maybe just set the scene?
9:53Like what is an agent sandbox at all? Basically. Great question that we get this all the time in particular, because Docker being such a ubiquitous word, Docker, let's call it historic Docker gets used for a lot of agent sandboxing. but we've clearly done something very different and I'll explain why. So at the end of the day, the sandbox is yet another container, for lack of a better word, that I can put my agent into and bound what it does and control all the dependencies it has and control the environment has. So fundamentally, it's a virtual machine or a container or something. Now, our flavor of it is actually based on micro VM.
10:28So it is actually the funny thing we've done is tried to, it's not a funny thing, but maybe hopefully we think the clever thing, give it the ergonomics of docker but importantly have a different construct on what the runtime is and the reason this is important is it's full micro vm base what you get every time you run there's a command line you download off of docker called it's in brew and all the obvious places if you run sbx our sandbox command sbx run claud sbx run gemini whatever you want open claw you get a dedicated micro vm per thing that looks feels acts kind of works the same way as a starts up in a second or less, like it's all feels, oh, this feels like a Docker thing, but it is fundamentally a VM.
11:06Now, the reason this is important, Docker containers have been around for a decade. And the ecosystem, I think, settled very early on in a very key primitive or very key property by not necessarily by force, but by convention of immutability. So when you think about Docker containers, I build my stuff, I put it in the box, and the box now goes from point A to point B, and I get everything about it outside the box. And then there's no logs, there's no nothing. The thing is meant to be immutable. And indeed, if you look at the way that people deploy at scale, you know, that people run something like Kubernetes and they'll look at, they'll actually have drift detection alerts and all the rest.
11:39And the moment that the container is different at all or changed from the signature that was deployed, it's considered a problem. And like in many cases, it's just like, that's it, shoot it in the head, start over and we'll go like redeploy the world. The problem is that's not how agents work at all. You run cloud code or open claw or whatever you want. The very first thing it wants to do is like, oh, thanks, Mark. thanks for that great question. I've got you. You're so right. I'm going to go off and I'm going to go download 17 Python packages and write a bunch of temp files. And I'm going to do all kinds of crazy agent things that they do.
12:08They fundamentally, they need to mutate their environment. And I think that starts at a very core place of, well, often they're very good at writing code. And so whether you're a software engineer or not, many times the best way to answer a question to solve the problem is the agent needs to write code. But in many cases, as people build more and more and more sophisticated agents, you really can't plan ahead what access it needs to something. It becomes no different than a person, where at runtime, you want to do some dynamic decision of what it needs, what it can access, and so on. So for us, this Docker sandbox or this SBX command really is a micro VM with this core, very breaking primitive, like it's a very breaking model for what Docker containers have been, is meant to provide the ergonomics, meant to provide sort of the brand and the ethos of everything about Docker, but for that agent world.
12:55And then what it is, is it's a bounding box and you get to have deterministic controls over a non-deterministic thing. You keep all your secrets out of it. You can very carefully specify what files and file systems you want in and out of it. And you can manage everything about the network on it. We're doing more things beyond that. But for today, that's the core shape of it is it's a box that you can observe and you can control and you can be a god of your own little agent world. And you can keep it from ever seeing anything in a very deterministic way, which is very different than every time I interact with agents and a new update to auto mode or something comes out, I'm like, oh, let me go see if I can make it do the thing that it shouldn't do.
13:30And it takes me indeed like four minutes or less. It's just not that hard to go make them do something that you don't want them to do for the generation of technology we have. And sometimes I think you need determinism to actually help make the non-determinism safe. So sandboxing from us, the docker, sandboxing, I think is an overloaded word. The computer science definition is it's a place to run untrusted code in the agent world that practically speaking means sort of vm boundary for us this micro vm technology is something we've taken all the dna of docker and all the things we've been doing for decades we make it work on mac windows and linux repeatably you can work anywhere you want you get to package up your stuff actually as a docker container slightly confusingly for convenience and ease but it runs in that bounded way that you can then observe and control so one area like just to sort of pick on there is is is actually the the micro microvm versus vm so obviously we're talking about virtual machine just for anyone not familiar with the acronym microvms in sort of my world have just been you know exploded in the sense like i work people know i work at super base and super base we interact with a lot of companies who want to have these agents doing things and running things and microvm has just become this exploded term what exactly how do you describe a microvm in comparison to a vm because i think a lot of people previously said well if i want to run something sort of quote safely i create a virtual machine and then I do on that.
14:49But what is the distinction? I mean, it's in the word there is micro perhaps. At the end of the day, I mean, this is where it's like, we used to call things unikernels. There's always these words that are for SE Daily, it's like still kind of half marketing. At the end of the day, it is still a virtual machine, but it has some properties of, well, we decide the kernel for you. We pick a bunch of things and we super optimize that it can be, you can run a lot of them. They run very resource efficiently and they start up very fast. The way to think about it today, we have many layered, it's a fairly complex architecture, as you might imagine, to do all the things we do on the host, like systems programming is back in a very exciting way.
15:23But the core, core thing of start VM of any type, we can do in under 100 milliseconds, like just to get the most raw thing up. We don't expose that because it's not terribly useful to the average person that wants to run a cloud code. But that core thing is meant to be super, super lightweight and super fast. and then we do just a bunch of optimizations to make it where historically think of VMs and it's like oh I run VMware Fusion or Parallels or something and allocate nine gigs of memory plus two CPUs like you just don't do any of that and it's meant to be this is not a terribly technical answer but it's meant to convey really it's a very fast thing where we've made a very lot of opinions around the kernel and everything else so you get an environment that has the properties of a VM, but we've just made it very small and very fast.
16:10Yeah, diet VM. Yeah, diet VM. That's a great explanation in terms of understanding a lot of people probably sort of, yeah, exactly. A virtual machine to me is just this big heavy thing exactly that takes up all this memory, but then a micro VM, as it sounds, sort of super optimized. At the end of the day, a virtual machine emulates CPU, memory, network, IO, which is contrast to like something, as I was saying at the very beginning, with Docker, with OS virtualization, where you have a shared kernel and you're basically emulating that kernel, this is actually full on. I emulate hardware. The contract you get is a hardware interface.
16:42And then we just tune a lot of stuff to make it small and fast. Yeah.
17:06with a one-line change to your workflow. Linux, macOS, and Windows Runners in Warp Builds Cloud or your own, enterprise-ready, SOC 2 Type 2-attested, and trusted by teams like Sky from Comcast, Bitcoin, and Braintrust AI. Get started with$50 in free credits at warpbuild.com. You're shipping faster than ever with AI coding agents, but those agents don't vet the packages they pull in and they don't have security contacts built in. Ori by Endor Labs fixes that. It plugs directly into your editor via MCP, catching vulnerabilities, blocking malicious packages, and flagging exposed secrets in real time.
17:41No separate tool to switch to, no dashboard to babysit, security that fits how you actually build. Teams using Ori see 10 times fewer security tickets and 6 times faster fixes. Free for developers. Get started at www.endorlabs.com. Think about your mobile app source code. Once it hits the App Store, it's out in the wild. And without the right protection, decompiling is easy for malicious actors looking to steal your IP or tamper with your software. That's where GuardSquare comes in. GuardSquare provides the highest level of mobile app security for Android and iOS applications and SDKs. Their advanced tools integrate seamlessly into your CICD pipeline.
18:24We're talking polymorphic, multi-layered code hardening techniques and automated runtime application self-protection. paired with mobile application security testing and real-time threat monitoring to deliver the highest level of mobile app security without compromise. Don't leave your hard work exposed. Secure your mobile applications today. Go to guardsquare.com to learn more. That's quite a nice sort of segue into what makes, I think, Docker sandboxes different, which is the idea that you do actually get, like, rather your agent does actually get its own Docker or daemon network and file system inside that sandbox.
19:05Could you maybe speak a bit to kind of that architecture? Yeah, the way it works is you get this micro VM. And in that it is a full blown Linux OS, you can root around anything like we have a bunch of them built on our Docker hardened images, images, you can build them on Ubuntu, you can build them on Red Hat, build them wherever you want. And by the way, for agents, there is an SPX run shell. So you too, as a human can just get a shell and do whatever you want. Now in that, you get a network proxy and an HTTP proxy where we are able, and IO proxies, where we enforce everything that comes in and out of the machine.
19:36So by default, you can start this up and it has no access to any host system, file system mounts, et cetera. Sometimes that's useful. You just want to make a throwaway thing that maybe goes and downloads some shady code off of GitHub and tests it, inspects it, whatever. That's what you want. More commonly, you're working in something like a GitHub repo or project or get project, excuse me. And you want to start either a work tree clone, or you want to start literally with that project mounted into the file system so the agent can actually work on your behalf and you can launch many of them. Thing one.
20:06Thing two, you have network controls. So you have on the outside of this, you have an L4 terminating proxy that inspects everything. And you get to, when you start it up, you get to pick the level of security and safety you want. You can be incredibly locked down and say no to every network call ever. Again, if you're trying kind of like, I don't know, introspect some Bitcoin code out of a banned country or something, maybe you should do that. In practical terms, people wind up with something like an allow list or denial list and just block bad things versus common things and allow things for their company.
20:36And then there's the credential injection. So along with it, it leverages an OS keychain, and you put something like your Anthropa key or your GitHub key or your Docker hub token and so on and so forth in a vault. And then inject it into the machine is actually a proxy parameter that is not the real credential. So inside the box, as long as you're using this correctly, the agent or the software can never see a real credential and everything gets intercepted by either the network proxy or by an HTV proxy. So we'll get into all the threat models and all the rest of what it can and can't do, et cetera.
21:08But with those primitives, you have the ability to keep only the things you want in the machine allowed, only the network calls you want to happen. And then you have a bounding perimeter around breach and how hard it is to roll your credentials because your credential was never shown to it this effectively hostile adversarial kind of collaborative cooperative agent with that you then get this environment with all those controls around it and then there's we've called them kits which are in a way for you to go package up your own flavors of clog code or pi or pick whatever you want you want and you can run and create and share agents with other people that do what you want.
21:46Like I have my own that I'm a big fan of the pie agent now. I've come to love this thing. And mine is this like crazy setup where I've forked in a bunch of ways. I've changed a bunch of the hooks and parameters. I have, you know, all my skills and my MCPs and everything else like loaded into it. And then it's mine and I can share it to other people inside the company and they can go use it if they want. And people can do the same thing. So. Yeah. And yeah, as you say, threat model is something that's super interesting and we will definitely sort of get onto that i had a feeling you might ask yeah you definitely get onto that just sort of sticking on i guess sort of architecture and that kind of thing for a second the cold start piece to this how does it actually work in practice you know it's like these things as they sound they need to spin up from effectively nothing but at a speed that you almost don't realize so like how does that kind of all work yeah so the thing when we're talking about the architecture i admitted this The SPX command today that you download off the website is clearly optimized for the laptop case.
22:44I'm a human being. I'm using my laptop. I want to run cloud code. I would like it to be safe and not be dangerous. I've been optimized for that. That's what it's for. We have this coming soon to a cloud near you or to a Docker website pricing page near you is a Docker cloud behind that that allows you to then work with those agents behind the scenes. And so you'll get different flavors of it. And we have a flavor working with people to, again, just like Dockers, we're doing everything to put the Docker ethos. Allow that SBX agent to run, or daemon, sorry, to run anywhere, Kubernetes clusters, your environments, et cetera.
23:14So with that as a thing for us, where we're explicitly making a runtime that is secure and cross-platform and has all these controls, to your startup question, what literally happens is when you take your laptop case, this in mind, it doesn't matter if it's Mac or Windows, that when you want to go launch, say, Claude Code or OpenClaw or even a shell, the very first thing it will do is go look at a registry, that registry might be Docker Hub, look for a name and pull it down if you don't have it locally. If you don't have it locally, it will pull down this kit image that I've described, which is in fact just an OCI or a Docker container as the actual bits and packaging of your stuff, what's in it, like what skills, what agent, et cetera.
23:54And then it will create a new virtual machine for you, or you can obviously attach to ones that are running, but it will create a new for you, boot up into that micro VM, and then start that software image as the actual agent, and then present to you back as a human. Here's your cloud code interface, and you can just now type to it the way you would. And you as a human, your workflow is unchanged from using cloud code or open code or anything else on your laptop. It's like bit for bit identical, we just literally give you the screen. We give you the screen with a box around it or the logical box around it.
Read the full transcript
24:25Now, in terms of the startup time for it, the lowest level primitive is sub 100 milliseconds to get up. In practical terms, in human experience time, I should do this before this call where it's like, okay, I've already downloaded off my laptop. I'm not sitting on my airplane Wi-Fi or something. I have a shell image. A small image that doesn't have an interactive agent starts in around a second or just a little underneath one second. So very fast for the human to get up. In the cloud case, it actually is optimized for closer to that 100 millisecond mark because now you're talking about programmatic construction and spin up and spin down and very bursty, scale-out, ephemeral transient workloads that need to go happen.
25:01So that same runtime is optimized for both of these. The actual SPX command that you download off the website today is very much optimized around the human laptop case. And so that's closer to a second because, well, when you get a Claude code or a Codex image from us, it's like four or five gigs because it's not just Claude. It's like, okay, I've packed in the Rust compiler, I've packed in the Go compiler. We've basically packed in a dev box that's useful for the coding agent to go do stuff. You can tune it to your heart's content to change that. The one you get from us is meant to just air quotes, just work.
25:30So it's around a second to go get the human facing one up, which is practically speaking for most people. Okay, so what? So every time I start Cloud, it takes it several seconds to go back and forth with the Cloud API, which may be erring and overloaded at the moment and whatever it does. So for the human, our goal is to get you the machine and get you into the agent, which is then operates at the speed of, you know, inference in human time, which is very different than something like programmatic scale out time. So yeah, that's very impressive. I'm just sort of zooming out to a product question, I guess, for a second, which was when this was being conceptualized at Docker, thinking about cloud code, that kind of thing, was that always a sort of first class primitive that had to be part of the product?
26:14Yeah, I'm just teasing this. We have a cloud coming soon. The core thesis from us is, again, like most people using Docker today, when they think of it. They think of, I have Docker desktop or have a Kubernetes engine or I have Docker images, and that's how they think about it. But that's what most people associate Docker with. It's a natural thing for us to go do is like, okay, I solved the same problem. Because we got asked over and over and over again, hey, how do I run Cloud Code inside of Docker? And or, hey, Cloud Code is using Docker. What do I do? And we ourselves use agents everywhere inside the company at this point.
26:44So as we went through the journey ourselves, we're also like, well, okay, we're hearing from the market and we ourselves are trying to do this. We kind of solved the problem for us and for everyone, but it is, yeah, it was very much built around. There's something like the chatbot that you would build for a customer support site or so your quintessential examples you have as these agents that run for consumer facing things that are practically speaking still look like normal software because they have a bunch of stuff to go with them. And then there's the human optimized case of trying to work with the coding agent.
27:09And so long-winded answer to, yes, it was a very deliberate decision to go optimize for the popular common agents that people are trying to run day in and day out. And so I think there's 12 or 15 that we ship and then a bunch of people have obviously done their own things on top of the ecosystem. So yeah, that makes sense. And this is like a sidebar and I disclaimer, I have almost no knowledge of this area, but I'm just curious the whole open claw thing was that also does that apply here at all? I often joke all the time. I keep seeing everybody like when there's like this Mac mini shortage for buying open claws.
27:43Yes. You can put them in a little virtual machine thing. You just put them in that. And so for us, we made it an agnostic machine. What's the ethos of Docker? It's build runs, run anywhere. We're Switzerland. I mean, without getting into geopolitics, because especially as an American, I feel like I can't say anything anymore. But anyway, we're Switzerland and we work with anything. So it's open any model, any harness, any agent, any operating system, any cloud. That's our entire ethos. And I think we're very credible to be the company that keeps that true. And so cloud code and codex, and if you go back in time last year, it's like, oh, everybody's on cursor.
28:16And then cloud code came out and then we changed that. And then the open claw moldbook thing all happened. And so now I think you have this explosion of a bunch of popular agents and open code is one amongst many. So sorry, open claw. Yeah, that makes sense. So let's get onto threat model. I think a lot of people would be quite interested in this one. So what is it? I mean, in terms of what have you had to think through and step through in terms of making this secure and then we'll get to like out of scope as well because that's every threat model you have to decide what's in and out so what is in effectively yeah well i think there's two things to this of what's in which is first the sandbox itself and then if you were a company or a team of people with multiple people because again you have multiple problems but start with just the i'm a single person i'm running something on my laptop what am i worried about again threat number one if you're running something like what It doesn't matter if it's Clawed or OpenCode or OpenClaw or whatever.
29:09It is not terribly hard to go trick the agent into doing something that it shouldn't do, regardless of what the permission setting is. And that's true if I'm running something like MCP and I get managed to like, you know, there's this hack last year of people doing Google Calendar invites that had, you know, bad prompt injection, et cetera. Great. I get the agent to do something, read something that is innocuous. It gets injected. The agent says, oh, okay, great. I should go do that. and it goes off and reads your keys, reads your files, does whatever. Big threat number one is actually protect the agent from the host.
29:40And for the thing we talked about at the beginning, where the developer environment is now like the juiciest target for supply chain, that's a big one because we've actually just taken all the keys outside of the box. No matter what you tell the agent to go do, no matter what the permissions are set on. By the way, when you run these, I omitted to say this, Claude Code or something, run it air quotes YOLO mode, where effectively that's the contract of the box, which is it's actually safe to run on a YOLO mode and doesn't sit there and nag you every two minutes of like, hey, Mark, is it okay to do this?
30:06Hey, Gregor, can I do this? Hey, can I do this? Hey, can I do this? Like, Jesus, shut up, just do it. And so the productivity gain you get is great. Now I can actually run a YOLO mode, but it can never, ever, ever go do bad things. It can't read my keys. It can't RMRF my file system. It can't do bad things to my environment that I don't want it to go do. I can very carefully keep things out of it. So big threat number one is keep the agent from being runaway and doing bad things to access to the host that's on host. Two is, similarly, the other big case you get is network. The threat model there crosses secrets and files together and can solve that problem on the host.
30:43The threat model on network is, well, okay, it's not that hard to go make the agent again. Even if I've constrained it down to the project that I want and I can't access all my things, it almost certainly can still ask something that I care about because otherwise, what's it doing? And so in that case, now you have the other problem of we have exfiltration of, okay, I can now get the agent to send something that it legitimately has access to that I must give it access to. I can get it to send something to somewhere else. And this is where the network firewalls and the secret, the HTTP proxy firewalls come in and basically allow you to go prevent the agent deterministically from doing something that it can't do.
31:18So the two big threat models by far are keep it from doing bad things on the host, keep it from exfiltrating data. That's the biggest ones from the individual's perspective. We'll talk about what it doesn't do in a second. And now hand in hand with that is, okay, you're a company and you actually have many people. And unfortunately, many people, the bigger the company gets, the less you can trust everybody. And so along with this, we have sort of the commercial, the SBX, I've admitted to say this too, the SBX product is free for everybody to go download. You can just go use it. It's no strings attached.
31:48Go forth and make your life better. Commercially, we have this AI governance package that we sell behind the scenes, which is we go to a company and say, okay, I can give you very rich, expressive policies and controls on MCPs, on agents, and on what people can do. And so if you're a platform or an IT or a security team or something where you're actually trying to manage the threat of your company of who can even use what agents, what tools can they have access to, what are they allowed to let them call, you need to go enforce that for the company. So it's kind of the same threat models at scale or the same threat models just applied at scale.
32:23The one additional one that comes with it, and we are working on making this available for everybody just today, it's enterprise only, is actually richer controls around MCP and data exfiltration. I've heard this, I don't know how many times now from enterprises of no financial data or people data about my company can ever go to the model provider. I don't ever want that to happen. And so we actually can do that from an MCP perspective and a tool calling perspective, but we can have a sidebar at some point on MCP versus CLIs and all that. but we actually can allow organizations to take control of that and actually prevent things like that and we're working on a whole lot more things there and then they get clearly there's no threat model on earth that works like trust but verify you still need observability and audit and all the rest and so they can actually get back out of this now they can set controls they can enforce these things and then they get back centralized observability and audit around it that's what it does and it's in scope at least today we are working on more stuff but that's what i would on the tin today.
33:19And I've actually, it was on somebody else's podcast and they were literally sitting at Scott Hanselman who was sitting there trying to live hack and break out of the SBX session. Like every five minutes was like, oh, I tried this. It didn't break out. The VM actually, we feel quite confident is very well hardened and very well contained. And then that thread as I've described it is there. Clearly you have to configure it for yourself, for your own environment, but it gives you a very deterministic hardened contract around those things I described. Yeah. Unfortunately on this podcast, we're not live hacking live coding but i am interested to get on to after this i believe the way that docker has been dogfooding this product specifically is super interesting so we'll get on to that in a second but yeah let's run this out as well so out of scope what things are you saying this is just not something that we are considering as protected against at the moment i'll be very precise in terms of at the moment because there's stuff we're working on like at the end of the day i think the hardest problem is when you think about agents you can't talk about this without talking about the data and the semantics around the data where data could mean, you know, cloud services too.
34:23Like very concrete example is I have my own agents set up to read Slack all the time. Then they tell me all kinds of useful stuff. It's incredibly helpful. I have a love hate relationship with Slack. My nightmare scenario is that it ever writes to Slack and starts sending different things to private channels or like the wrong, I mean, I'm in all these workspaces and all these other companies. Like, could you imagine telling one about the other. So I'm like, I have very, very, very hard boundaries that I like basically double check once a week to make sure it's still in place around what I can and can't do.
34:51But what I really want is that I want an agent that I can control and say, okay, I want you to use Slack. I want to be able to write controls that say you agent can write to this Slack channel, to those people with this types of data. That's the actual thing I think that we need to get to as an industry. I don't think anybody has solved this yet. That to me is where we're not all the way there to that point yet. Actually, to be clear, this is a software engineering podcast, not a marketing podcast. In terms of what it does at the moment, we don't do that. And so you still need to apply good judgment in terms of how risky you're willing to be with your agent and what you want to let it have access to.
35:26I was reading a pretty funny Reddit post I read not too long ago from somebody who was really, I like the open claw forums and nano claw forums and stuff. And there's somebody in open claw on Reddit that was like, ah, this thing has been amazing for me. It's been buying my groceries for four months until last week. And then it bought 50 pounds of garlic. Whenever it had done, it had gone off and, you know, hallucinated like a multiple of times 10 or times a hundred or something. And things like, that's a very innocuous example, but you then imagine applying that to like, I don't know, your Bitcoin wallet or your bank account or your company's CRM data, or actually things get bad very, very quickly.
36:00So this problem of making sure that you are still not doing incorrect or silly things, letting your agent have access to things that you under common sense wouldn't let it have access to. I think that's still the, I don't want to overpromise people and I don't want people to take dependencies on us thinking, ah, you'll keep my bank account safe for my agent ever doing anything. Like, well, we can help. We can give you a very bounded way. But at the end of the day, you still need to make good choices for the totality of the threat model around your data and your stuff. So we are working these things.
36:30I think we have a very ambitious roadmap for the next couple of years, I think, for minimum, but it's not there now. I don't think anybody's there now. No. And I mean, shared responsibility model is pretty common here. And I always like that term. It's just a fancy way of basically saying, don't do really stupid stuff, user. We can't be responsible for the really, really stupid stuff. So yeah, we can't solve everything. So yeah, exactly. I think that's exactly the right term. And no matter what we do, there will always be a shared responsibility model. We can do more, but I think that will universally be true.
36:58Yeah, absolutely. One thing you touched on a bit earlier, and does sort of come into this, I'm just curious to get a little bit more detail on is the secrets and credentials piece. Again, how does that work? I think you've mentioned proxies being part of the solution here, but how does sort of knowing that credentials aren't somehow being stored somewhere or repurposed or anything like that? It's hard without, if we had whiteboards and pictures and stuff, it'd be a lot easier. But you'd imagine, think of the things as what goes in the box and what stays out of the box. At the end of the day, what stays out of the box are the secrets.
37:30So what goes in the box is a placeholder that looks and acts to the thing inside like it's a secret, but it's not a real one. So the threat model there, it's a little bit blurry and it's nuanced. The threat model there is, well, it's impossible to leak your secret. So at least per this contract, the secret is never presented in the box. Nothing in the box can ever get it. However, that placeholder secret still has, by definition, access to things. What we literally do is we run a proxy on the outside that intercepts traffic and we'll swap out your, for example, if you put a secret in as like your Anthropic API key, what does it literally do?
38:04It looks like a bearer token in the authorization header. And every time the cloud agent makes a call to anthropic.com, it sets authorization header colon and then stacks in that key that's like SK dash blah, blah, blah, blah, blah, blah. So we put a fake one in and the call will go out and then we will intercept it and we'll say, ah, you put that key in the box. I'm going to take that key out. I'm going to put the real key in and I'm going to send it off to Anthropic. And when the call comes back, I will conversely strip and sanitize the things on the backside and now give it back inside the box.
38:33The box is none the wiser. It's a good thing. We're literally man in the middle in a healthy way. Now the threat, so the good thing is your secret can never be leaked. Nobody, you can never take that secret and send it out. You can't take your AWS keys. You can't take your bank account keys. You can't, the agent could never take them and go post them to Payspin or something like that. That can't happen. However, the agent does have access to the same service behind the scenes because by definition, we're still letting the call through. We're just keeping the secret out. So the agent could still go, for example, well, if it has access to your bank account, it doesn't have your bank account keys, but it could still transfer money for you.
39:07By definition, we are not able to stop that, at least not today. So that's kind of the nuance on the threat model there. So, you know, the agent will still have access to the things it has access to. and then mostly you bound when you revoke a credential you will then be able to immediately contain the damage but the damage will still have been done so yeah one thing i haven't i guess asked or touched on is i guess all of sbx is that all open source as well is that something it's like all right it is a lot of open source it is not 100 open source so the core vm today we've chosen not to for a couple of reasons one there's competitive and commercial things but putting that aside, actually, we're not right now trying to go get a community of people to go contribute to it.
39:47It's very hardcore systems technology. And honestly, we actually, I'm curious if you've had others on the podcast with this, like we might be hypothetically speaking, drowning in AI pull requests from everywhere else across the industry on the LDO. Yeah. Yeah. We are hearing a lot of that. Yeah. Yeah. And so we're trying to go fast. We're trying to solve a lot of problems. I'm not terribly interested right now in dealing with it all. We're already dealing with it in a bunch of the places where, I mean, Docker clearly is, to be clear, I want to say this, Docker is a very open source friendly company.
40:14We're our entire ethosist company. Our intent is to be open source. We are trying to make pragmatic choices as we go around it. But the answer to your question is today, it is not a hundred percent open source. It is a lot of open source, but there are parts of it that are not and various reasons why. It is all free, however. So. Yeah, no, that's amazing. Let's, let's talk about this dogfitting thing. Like it was, I think it was a blog post by one of your colleagues that sort of went through this a little bit and it's called the fleet but which sort of explains but maybe you could sort of semi-summarize like what is this how i guess is docker dogfooding that term of like you know trying your own products basically when in this case like i think stress testing it how has that actually worked well one i philosophically deeply believe every company should dog food their things if they don't dog food their things what are you doing so the funny thing the nuance for us is not a hundred percent of the company can dog food because we have a lot of people that work on this core thing itself and putting nested virtualization and gets actually very hard.
41:08So we have a funny dichotomy of who can dog food and who cannot on a daily basis. Cause by the, maybe the definition of dog food to me is I want most of the company living their lives in our products. And like they do exactly what our customers do. We have constraints where it's literally not possible because they're building it. But the things that have been incredibly helpful for us, cause we find a lot of the friction and the same things, we have a lot of things we need to work on that make it better and more usable and all the rest. But typically speaking, we will find a lot of things for our customers do.
41:33But the better question besides me to ask this question to is actually our CISO, who I think is like elated that we have moved our organization as much as possible to where all things run inside the sandbox for exactly the same reasons that our customers are calling us for. So today we dog food SBX. Most people that are running a coding agent need to do so inside of SBX. Again, there are exceptions. As I mentioned, we have an MCP gateway as well that goes hand in hand with this. We stand up an internal one for ourselves in enterprise. We run all the things we, as a company, use as an MCP through that.
42:05And then I've mentioned this a little bit. We have a teaser of a cloud and some higher level products coming. We've got a whole lot of our company actually building our agent orchestration around that. So if you think about the way the SDLC works is, okay, the human sitting there interacting with their cloud code terminal does their work day in, day out. We're probably like everybody else where most of the engineers like cloud code the best. Some people love open code. Some people love codecs. We let people do what they want on their things of choice. But then you have a lot of things that can run in the background and can actually be scaled out.
42:35So we do have a lot of autonomous agents that run either by because they're tagged in a pull request or they literally run and watch. They're just set up with GitHub triggers or Git triggers and watch the repo, look for security scans, look for test coverage patterns, like all the common stuff you'd expect. And then they'll go either generate new issues and give things to people or they'll actually generate all the way to the code and like reissue a new pull request. Today, we're not all the way to the point where we're letting a lot of things through without a human review because mostly we have a security product where the contract is determinism.
43:08So we kind of philosophically don't do that. We do back and forth on this all the time, like on a weekly basis. But that contract looks like then what most of the engineering organization and product organization works with. And then you also have everybody that has automated their life where they have daily reports and this, that and the other thing. Now on the one we're working on actually kind of aggressively is actually getting our go-to-market and our GNA teams working with the same stack and dogfooding. The product as is, is very much, I mean, Docker is a developer company. We're built for developers.
43:38The product is for developers to go do stuff. But it turns out a whole lot of people are learning to code or learning to, like, it's this funny thing now, well, everybody's a developer now, your definition of developer here. And so what we're actually working through right now is, okay, how do we effectively dog food our own products and our non-engineering and product organizations? That's kind of the last untapped frontier for us in terms of how that's going. Yeah. So that's a very holistic approach to dog fooding from what you've just described. So that's great to hear. As we sort of start to vaguely sort of cruise to the end of the episode, I'm curious about the future of this.
44:12And you have touched on it a little bit, like the sort of the hosted piece to this, because that was going to be one of my first questions. Like, is this only going to be local or what does it maybe look like in the future? You have touched on a hosted idea without maybe revealing anything you can't today. But like, what would be the reasons, I guess, for offering a fully hosted version of this? I'm not terribly secretive with our roadmap, because I think it's honestly, I always kind of joke like, well, if you can't look at a company, write down the roadmap in about 20 minutes of thinking, they probably have the wrong roadmap.
44:38So for us doing a hosted version and a cloud version is, I think, incredibly important. Because again, like there's many reasons why I need a cloud. So many times I want to work on my laptop. I need access to Excel files. I need access to local things. I want access to local resources. And there's many times, well, I want to close my laptop and I want to go to sleep or I want to be on an airplane and I want stuff to keep going, in which case I need to be able to go back and forth. Our term for this is Remocal. So our cut on this is, well, this SBX contract that you have will be able to move portably back and forth from the laptop to the cloud.
45:09That's kind of the big thing we're working on, along with the, again, traditional sort of scale out cloud resources for the cases that you'd expect. So getting that working, I think, unlocks just a whole lot of things for people, whether it's background agents or agent swarms or whatever it is. That's like a big one. Really working on that sort of, now you're back to like why on the laptop. I think many things actually need access to the laptop that are like, oh, say the microphone or the camera or again, Excel files or whatever else it is. And so really working on how to safely constrain access to host resources above and beyond what we've done.
45:45Like we already have a bunch of special things in for GPUs and things like that. but sort of general purpose acts to how agents can do kind of do what's called computer use is I think one of the bigger things we're working on. We are, I think I've long been the like top evangelist for open models. And so I think one of the big investments for us will be just making sure this keeps working with allow you to plug in and swap in open models. I think the year of open models is upon us already and it's coming faster. Definitely. I think that's incredibly important. And then again, a huge focus for us, it really is on this team and commercial governance product.
46:17And so just doing a whole lot more there around the richness of policy that can be expressed, how exactly agent identity works and like delegated ways and stepped on ways. It's a very rich and complex domain. And so most of the customer asks and the usage asks we see in practice beyond, hey, I'm trying to work in a little world. Why didn't this agent work? Really very quickly turn into that, which is, OK, well, I'm trying to roll this out to 1 ,000, 10 ,000 people. And I have this very complicated thing and it needs to talk to that SSO system and that access control system. But I need the agents to do this then, now.
46:49Like that's, I think what we're, I think we're going to be busy for a while working on that. But I think the nice thing is, as we build this for the comp, as an ethos, making the hard things possible and solving the complex problems in a way that is easy and digestible allows us to, I think, have a platform that will work for everybody. Because for us, part of the goal is great. We have Docker. We have probably a billion people running Docker somewhere, somehow across the... I don't even know. I can't count it all. I think we have a responsibility to make sure that the ecosystem and community of Docker users is able to come forward into the world of AI and making sure that everything we do provides them the same sort of safety uplift and the same sort of convenience uplift is actually important to us.
47:31It's just an ethos. We have a ton of work to go do to solve problems for the companies that are having real pain points on, I mean, everybody's scrambling, right? It's like, well, I was told I need to roll out Claude, but either because the people in my company are screaming at me that they need it or because, you know, the board is screaming at me, everybody's screaming at me, I need to roll it out. And it's hard. How do I roll this out with safety? How do I not like risk the business and so on? Our whole lot of a roadmap focuses on enabling that with more and more and more expressive controls and then making sure it's accessible to everybody.
48:03So that's the digest radio safe version of the roadmap. So. Yeah, that's awesome. Something I was just thinking there was, is there a world, obviously, I don't expect this to be on the roadmap, even in like maybe the next year, but is there a world where you could see, for example, Docker sandboxes running on mobile devices, as opposed to a laptop. Developers love showing off that they've been doing something on their phone, which is currently SSHed into some other box somewhere. But what do you see on this front? isn't in the next quarter. No, we've talked about something. It comes up. We're Docker.
48:33It's a ubiquitous runtime that runs everywhere. We get asked for everything. Like it has to be put on thermostats. You get asked to put on cell phones. Like I don't think it's out of the question. In particular, I think the mobile case, obviously the sandboxing primitives, the operating system primitives, everything about that is totally different, but having some type of an experience, like where again, where I can give you the, it's actually less about the sandbox and more about the safety. So I think it is in bounds for us to say, well, I'm giving you this portable run, runs, run anywhere contract that you can run with safety and speed and productivity.
49:04Having an answer for that on the mobile phone, I think is actually probably inevitable for us. I think it'll probably require us doing a different shape and it won't look like, oh, I brought SBX to the phone as is. I think it'll require us being, okay, understand the constraints of the environment, understand the problem we're trying to solve, stick to the ethos, make it as coherent as we possibly can across this but i would not rule that out that's for sure yeah that makes sense not ruled out but equally i love that phrase i think it was well i don't know who said it originally but i remember hearing toby luke say it's like when you say yes to something you're effectively saying no to everything else so what you say yes to you know you've got to be very very clear on that one that is a roadmap effectively yeah yeah there's a great steve jobs quote is like focus is not saying yes focus is saying no and so it's not a no forever but it is a no for right now so yeah well just to recap yeah where's the sort of best place for a developer to go get up and running with spx i mean easy one as well you should go to docker.com and there's docker.com slash sandboxes but if you're on mac you can do brew install spx and win get on windows and so your linux distribution of choice for your installers so but i would start there and then it is usually pretty quick to get to hello world and get to your first agent being up and running And then it's a world of threat models and virtual machines and agent controls from there.
50:21Amazing. Well, Mark, thank you so much. Really appreciate you making the time. And yeah, I think this will definitely be one where we're following along pretty closely and no doubt catch up in a year or two. That'd be great. Thanks for having me. Thank you.
50:47VE IN THE HANDborough 1 rho �go이제 2 uou 1313
51:03u 13 13
From the publisher
The most useful coding agents can mutate their environments by downloading packages, writing files, and connecting to services across the network. However, that freedom also presents dangers, and promises to usher in a new wave of security threats.
Docker recently announced Docker Sandboxes, which give each agent its own isolated micro VM while preserving the familiar ergonomics of a container. A standard container shares the host’s kernel, but a micro VM emulates hardware and runs its own kernel, giving a stronger security boundary around code that cannot be trusted.
Mark Cavage is the President and COO of Docker, and he previously worked at companies including Stripe, AWS and Oracle. In this episode, Mark joins Gregor Vand for a wide-ranging conversation that includes why agents break the immutability assumptions containers were built on, how micro VMs differ from both containers and traditional VMs, and the still-unsolved challenge of giving agents scoped, trustworthy access to sensitive services and data.
Sponsorship inquiries: sponsor@softwareengineeringdaily.com
The post Docker and Sandboxing AI Agents appeared first on Software Engineering Daily.
