In short
The episode explains how NanoClaw enables long-running “personal AI agents” while addressing security problems seen in earlier systems like OpenClaw.
Guest backgrounds
Gabrielle Cohen is the founder of NanoClaw. He studied physics and computer science at Tel Aviv University, worked ~7 years at Wix (front-end/fullstack; led a dev team), then returned to building around Cloud Code and started NanoClaw. Kevin Ball (Kate Ball) is VP Engineering at Mento, an independent engineering coach, co-founded/served as CTO of two companies, founded the San Diego JavaScript Meetup, and organizes “AI in Action” via Latent Space.
Key claims
Agents can’t be trusted; NanoClaw uses zero-trust orchestration. Each agent runs in its own Docker container; credentials stay outside the agent and are injected via a proxy “vault.” Sensitive actions require human-in-the-loop approval. Communication is message-based via host-managed inbox/outbox SQLite DBs.
Notable examples
Cohen describes using OpenClaw for sales-pipeline automation, then discovering it downloaded/stored all WhatsApp messages in plain text. NanoClaw instead stores only messages from the connected group by default. He also gives an “Andy” WhatsApp agent that schedules price checks and sends sale alerts, and a multi-agent PR QA workflow (review → test plan → VM testing → human-approved GitHub CLI command).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to AI Agents
0:00 to 0:45
Learn about the potential and challenges of AI agents as digital assistants.
“AI agents have shown remarkable potential to function as persistent digital assistants that are capable of monitoring data, managing communications, and taking action autonomes.”
Overview of OpenClaw
0:45 to 1:55
Explore OpenClaw's role in the evolution of AI agents and its security shortcomings.
“own Docker container, keeps credentials entirely outside the agent's environment, and enforces human-in-the-loop approval for sensitive actions.”
Introduction to NanoClaw
1:55 to 2:52
Discover how NanoClaw addresses the security issues found in OpenClaw.
“This is a topic I was very excited to get the opportunity to discuss.”
Gabrielle Cohen's Background
2:52 to 4:26
Learn about Gabrielle Cohen's journey and how he came to build NanoClaw.
“I think a lot more folks have heard of OpenClaw at this point, but what is NanoClaw?”
NanoClaw's Security Approach
4:26 to 5:50
Understand NanoClaw's zero-trust security model and its implications for AI agents.
“And I think we've seen some of that played out.”
Technical Details of NanoClaw
5:50 to 7:18
Dive into the technical architecture of NanoClaw and how it functions.
“No credentials or secrets in the agent's environment.”
Agent Isolation and Control
7:18 to 8:23
Learn about how NanoClaw isolates agents to enhance security and control.
“LLMs on their own can only sort of be trusted.”
Interaction with Messaging Platforms
8:23 to 9:37
Discover how NanoClaw manages interactions between agents and messaging platforms.
“but let the agent run wild within its environment and then enforce things on the boundary of that environment.”
Efficiency and Task Management
9:37 to 14:00
Explore strategies for optimizing agent performance and managing tasks efficiently.
“Agent does its thing, uses tools, bash commands, whatever it wants to do going wild.”
Efficient Data Handling by Personal Agents
14:00 to 14:33
Explore how agents manage data more efficiently through polling and scripting.
“The agent just gets whatever data gets pushed in.”
Show all 30 chapters
Improving Token Usage in Email Processing
14:33 to 15:42
Learn about strategies to reduce token consumption when processing emails.
“If I'm having a polling every five minutes or 10 minutes, that's going to burn through tokens.”
Integrating Agents into Daily Life
15:42 to 16:43
Hear a story about using an agent as a personal assistant and its capabilities.
“So these scripts that the agent runs or is able to schedule, I presume those run inside of the container so that they're in the free-for-all security model.”
Task Scheduling and Agent Communication
16:43 to 18:49
Discover how agents schedule tasks and communicate with the host process.
“And then gave her all these explanations about scheduling jobs and stuff.”
Agent Lifecycle Management
18:49 to 22:44
Understand how agents are managed in terms of persistence and lifecycle.
“But for one-off tasks, it calls an MCP that go out to the host process, the host process then writes a message to the inbox.”
Agent Lifecycle Management
23:52 to 25:36
Understand how agents are managed in terms of persistence and lifecycle.
“And when you use our link, you're supporting our show, notion.com.”
Context Management and Compaction Strategies
25:40 to 28:00
Delve into how agents manage context and handle compaction effectively.
“Let's talk a little bit about lifecycle, right?”
Session Management and Context Handling
28:00 to 29:30
Learn how NanoClaw manages session data and context for AI agents.
“So we set it a little bit lower and set the compaction threshold at 165k.”
Multi-Agent Communication in NanoClaw
29:30 to 31:40
Explore how different agents interact and communicate with each other.
“Let's talk a little bit about interaction models.”
Approval Gates and Security in Agent Interactions
31:40 to 35:13
Understand the importance of approval gates in agent communication for security.
“outside of the channels really useful is an approval gate.”
Agent Factories and Automation in Code Review
35:13 to 38:17
Discover how agents facilitate code review and testing in software development.
“First thing I think with agents is I'm like, okay, I want it to build things for me.”
Product Philosophy and Minimalism in Open Source
38:17 to 42:00
Learn about the core philosophy driving the design and functionality of NanoClaw.
“It's not the agents just automatically merging pull requests.”
Keeping NanoClaw Minimal and Customizable
42:00 to 44:39
Learn about the philosophy of maintaining a minimal codebase and supporting customization.
“So keeping it small, keeping it minimal, that's a core part of the philosophy of the project.”
Challenges of Version Control in Open Source
44:40 to 46:17
Explore the issues surrounding documentation and version control for forks in an open source project.
“modular way of coding, which is in a way kind of just coding best practices.”
Funding and Future of NanoClaw
46:18 to 47:57
Discover how funding impacts the sustainability and development of the NanoClaw project.
“But I think there's a huge advantage to this approach.”
Innovations in Agent Customization
47:58 to 49:24
Learn about the innovative skill-based approach to customizing agents in NanoClaw.
“and we can have a big community around the open source project.”
Integrating External APIs with NanoClaw
49:25 to 51:46
Understand how to integrate external services like Spotify into NanoClaw through skills.
“So we have to be able to pull in changes.”
Future of AI Agents and Customization
51:47 to 56:00
Explore the evolving landscape of AI agents and their customization capabilities.
“Because my nanoclaw is different than your nanoclaw.”
Exploring NanoClaw and Custom Agents
56:00 to 1:00:44
Learn about the customization of AI agents using NanoClaw and the intricacies of model compatibility.
“If you allow OpenCode and you allow Pi, they can handle the multiplexing to whatever model you want.”
Building Efficient AI Agents
1:00:44 to 1:02:28
Discover insights on building efficient AI agents and the pitfalls to avoid in the process.
“So we're getting close to the end of our time.”
Optimizing AI Use Cases
1:02:28 to 1:03:34
Understand the importance of focusing on value creation before cost optimization in AI projects.
“And then in terms of, I think a lot of people get stuck on which model, which agent don't get stuck there.”
Transcript
Automatic transcript. May contain errors.0:00AI agents have shown remarkable potential to function as persistent digital assistants that are capable of monitoring data, managing communications, and taking action autonomes. over long periods. OpenClaw was one of the first serious attempts to fulfill that vision, connecting frontier coding agents to messaging platforms like Slack and WhatsApp and letting them run continuously in the background. However, OpenClaw largely set aside questions of security to pursue that vision, leaving credentials exposed in the agent's environment and giving agents broad access to data and services far beyond what any given task required.
0:38Nanoclaw is an open-source project that takes a zero-trust approach to agent orchestration. Rather than relying on instructions to constrain agent behavior, it isolates each agent in its own Docker container, keeps credentials entirely outside the agent's environment, and enforces human-in-the-loop approval for sensitive actions. Gabrielle Cohen is the founder of Nanoclaw, and he joins Kevin Ball to discuss the security architecture behind Nanoclaw, how the agent sandbox and proxy model work in practice, how agents communicate with each other and with the host orchestration process, how the project approaches context window management and long-lived agent sessions, and more.
1:18Kevin Ball, or Kate Ball, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through Latent Space. Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.
1:52Gabriel, welcome to the show. Thank you. Great to be here. Yeah, I'm excited to dig in with you. This is a topic I was very excited to get the opportunity to discuss. But before we do, let's start with you a little bit. Can you share just a quick overview of your background and how you came to build Nanoclaw? Absolutely. I studied physics and computer science at Tel Aviv University. And actually, during university, I started to work at Wix.com as a student. I worked there as a software developer, started in front, then moved over to Fullstack. and then led a team of developers there. Spent a total of about seven years.
2:29Actually stopped programming for a few years and took a break and worked in public relations and promoted tech companies. And then jumped back in about a year and a half ago or a year and two months ago, right around when Cloud Code came out. I felt like there was this huge opportunity I had to dive back in, started building things. And that's when I got to NanoCloud. So for those who may not be familiar, I think a lot more folks have heard of OpenClaw at this point, but what is NanoClaw? Let's just give the quick TLDR. So NanoClaw is an open source project, and it is a safe, secure alternative to OpenClaw.
3:08So OpenClaw is this wild experiment, the way I see it. It's this wild experiment of if I put aside all of the security, safety, and even software quality concerns. I just put them aside and try to see, can I get a lot of value if I just take frontier coding agents and connect them to everything? And the answer is you can get a ton of value. I used OpenClaw. I was building an AI native marketing agency and I set up OpenClaw in a chat with me and my co-founder and told it, you've got the sales data from our company mounted in your environment. Look at that and you're managing our sales pipeline.
3:51You're now our sales manager. And it very quickly started to do the work of an employee. And what that means is giving us a morning briefing. Here are the deals. Here are the things you need to follow up on, setting reminders. We're just dumping raw data at it and it's making sense of it and updating files. But then very quickly, I saw these massive security issues. So that's why I sat down to build NanoClaw so I can get the value that I saw in OpenClaw, but without compromising on safety and security. I think this is something I've definitely had on my eye. Since OpenClaw came out, I was like, oh, this is a disaster waiting to happen.
4:25It's like fascinating. And I think we've seen some of that played out. But let's talk about the approach you take in NanoClaw. Like how do you make it secure? What is the security model that makes this work better? Yeah, OpenClaw is definitely this disaster unfolding that you can't look away. But for NanoClaw, the mindset is, first of all, agents can't be trusted. AI in general, it's non-deterministic. No matter how many exclamation points you put after the big, bold instruction saying, don't delete my production database. I've joked that if you're digging through a code base and you see in the claw.nb file at the top in big bold letters, it says never, never run, drop production DB, then you know that that agent has deleted a DB before and that it's going to do it again because it still can if the instructions are there.
5:20So the idea is don't trust agents. I actually approach it with a zero trust mindset of saying as long as your agent is able to interact with any kind of unsanitized data, if it's looking at emails, if it's looking at pull requests, that's unsanitized data. People can stick in their prompt injections. And it means that your agent can be prompt injected and it can act against you. It can end up turning out to be malicious. So you need to treat the agent as if it's malicious. So with NanoClaw, we isolate the agents. Each agent is running in its own sandbox, isolated from the host machine, isolated from other agents and we're very intentional about what we put in the agent's environment so first step is isolated second step don't put any credentials secrets api keys tokens into the agent's environment now you do want to give the agent the ability to interact with credentialed services and apis so the credentials tokens api keys sit outside of the agent's environment All requests that are leaving the agent's environment get proxied through a small vault that sits next to the agent and credentials get injected outside of the agent's environment.
6:33So that's the second part. No credentials or secrets in the agent's environment. And then the third part is enforcing policies and access control at that proxy gateway point and including human and loop approval. so I can connect my agent to my email for example the agent doesn't hold the actual credentials that allow it to access the email but its requests get proxy credentials get added and I can set a policy saying it can read my emails freely but if it tries to send an email or delete an email that requires my approval and then an approval request gets sent to me and I need to approve each action this sandboxing and proxy approach is something I feel like I've seen bubbling up in a number of places as people grapple with, yeah, this core problem that agents cannot be trusted.
7:22LLMs on their own can only sort of be trusted. But then once you introduce any sort of outside influence, it's like dealing with user generated content again, right? Like, you can't trust it, you've got to sanitize in some form. And here, sanitization looks like it's isolated. Can we dive into the details? How do you isolate it? You said it's in its own container. But like, what does that end up looking like? What is the interface to the proxy? Like, How do all those pieces work? So we have the agents running in Docker containers. There's also an option to run it in Docker sandboxes, which is a new thing that Docker has been working on specifically for running agents.
7:58And that has some extra hardening. And the agent runs in that environment. It gets a folder mounted, and that's its file system. So it can create files. It can create folders. It can organize them in all kinds of ways. it can manipulate its environment in any way. The approach that we took with Nanoclaw is not trying to inspect every action and every request, but let the agent run wild within its environment and then enforce things on the boundary of that environment. So you mount a file system. One Nanoclaw instance can run on a machine in a VM, say, and have many agents running in that one instance.
8:40There is a host process that's responsible for orchestrating the different agents and containers and routing the messages to the right container with the right agent with the right session. And then the way it's implemented, there are two files that are two SQLite DBs, one for the agent's inbox and one for the agent's outbox. So the host process, when it gets a message and determines that it's supposed to be routed to an agent, agent A, it writes the message into the DB that's mounted for that agent. And the agent is pulling the DB for new messages every time. It's actually not the agent, to be more precise.
9:20It's a loop. There's a little process running inside the sandbox. It's pulling the inbox. Every time a new message arrives, it just pushes it to the agent. And that's by default, Agent SDK, Claude's Anthropics Agent SDK can also run codecs or open code. Agent does its thing, uses tools, bash commands, whatever it wants to do going wild. It can do all kinds of wild stuff in its environment, but it can't access the host machine or other agents environments. And then when the agent finishes and outputs a response, the little process running in the sandbox with the agent writes that to the agent's outbox.
10:02And the host process is pulling the outbox, gets the response and routes it to the correct chat. That could be a Slack channel, a Slack DM, could be a WhatsApp group, wherever it's supposed to go to. Interesting. So those input and output SQLites, they're only used for input messages in both directions? Yes, and we made a decision that everything is messages. So there are a couple environment variables that are injected into the container. Besides that, everything going into the agent's environment, everything coming out is set as messages. So for example, if I want to do a scheduled task, have the agent run at a certain time every day and send me a briefing of my sales data, my sales pipeline.
10:50So that's a message with a process after time. And if I want it to be a recurring task, then it's a message with process after and a certain recurrence set to it. So everything is messages. Approval requests are responses to approval requests. all of that is just messages or system messages or regular messages but everything is messages fascinating and why have the external process managing that rather than the agent directly sending messages you want to be able to enforce what the agent can message and where it can get messages from so you've got for example a whatsapp or telegram whatsapp is maybe the most clear example of that, you have a WhatsApp bridge, WhatsApp connector.
11:36If you're connecting with normal WhatsApp, not a business WhatsApp, so most people are using a package called Baileys, it kind of reverse engineered the API that WhatsApp desktop uses, and that connects using that same API. So it's a non-sanctioned, let's say, community connection to WhatsApp. That connection connects to your WhatsApp account, your WhatsApp identity, and that same connection can connect to every WhatsApp group that you're in, every contact, every message. It just downloads everything. When I started to use OpenClaw, one of the things that I noticed that made me go, okay, wait a second, I can't actually use this for a real use case for a business, for a production use case, is I started to dig through the logs, debugging something.
12:20And I saw that OpenClaw was actually downloading and storing every single message from every single WhatsApp group and every contact in plain text in a file on my computer. Not just the one group that I had added it to, but every single group. Now with NanoClaw by default, when I connected to a group, it only stores messages from that one specific group. Any message that comes in, it just ignores. So now if the agent is directly connecting with WhatsApp and sending messages directly to WhatsApp, it can message any group and it can change any kinds of configurations to get messages from any proof.
12:57Totally. So it makes 100 % sense to me why this would be outside in the host outside of the container, right? You don't want to give the agent direct access to WhatsApp. The thing that I was curious about is once you're inside of the container, you've passed it through SQLite, you've already got, you're past your like proxy boundary, essentially, you're past your moment of control. But you mentioned there's still a separate process that is like pulling the input and passing it to the agent or vice versa, translating from the agent into the output. And I was curious about that inside of the proxy boundary separation.
13:27Yeah, that's not part of the security model. That's part of the assumption is anything within the agent's environment is the agent will be able to get access to it no matter how hard I try to prevent it. That's for better quality, better quality output from the agent and to save tokens. It's all about the tokens. So for the agent to pull the database itself constantly and then have whatever comes into that database just pipe straight into the agent, you can do that. You can implement an agent loop like that. And it works. You can do that. And I've implemented these really minimal versions of clause that just do that.
14:02The agent just gets whatever data gets pushed in. But that's not very efficient. So we could take it a step further. for the schedule of tasks, if I set a scheduled job to run, say I want an agent to check my email and process email, send me a message on WhatsApp if an email came in from a specific thing that I'm waiting for. So the agent is kind of polling. It's got a tool it can use to read my email. It's polling my email inbox and checking for new messages coming in. That burns through tokens like crazy. If I'm having a polling every five minutes or 10 minutes, that's going to burn through tokens.
14:37And that's actually the kind of stuff that Anthropic was banning accounts for or blocking accounts from using their subscriptions for. So we went and implemented, I think it's about a month and a half or two months ago, the ability for the agent to set a script that runs before a scheduled task. So instead of the agent waking up with a prompt that says, go check the inbox and see if there's new messages and send a WhatsApp message if there are. Instead, there's a little script that runs. The agent writes the script. When I tell the agent, hey, do this for me, it's got instructions saying, if this is really frequent, write a script first, have the script run, and the script runs, checks something, and programmatically decides whether to wait the agent or not.
15:20And that's both for tokens and also for context windows. So we're not standing the agent with all kinds of inputs. Messages that come in to the agent's environment, we want to wrap them in a specific way. Each message gets wrapped saying who this came from, what's the timestamp, and gets presented nicely to the agent in a way that packs it immediately and nicely into this context window. Okay, so I want to dig more on this. This is great. I'm loving this. So these scripts that the agent runs or is able to schedule, I presume those run inside of the container so that they're in the free-for-all security model.
15:54I don't know. Do you have names for these two security models, like the outside and inside? Yeah, I call them the host, not the host process or orchestration. And the agent, it's the agent runtime. Okay, so I'll use those. I think the free-for-all is a little more fun as a term. So these are running in agent runtime. Agent can write scripts. Now, those get triggered based on messages or you have, it sounds like a cron or something like that? Or is the cron managed outside in the host and triggers an inbox message of some sort? Yeah. So the agent can schedule tasks for itself, schedule recurring tasks.
16:31So I gave an agent to my wife and she just put her in a WhatsApp group with an agent said, Hey, this is Andy. He's your personal assistant. Good luck. And I left. And then she's, you know, in the beginning, it's just like questions like you would ask chat, but then she started to go like, wait, Andy, what can you actually do? Like what capabilities do you have? And then gave her all these explanations about scheduling jobs and stuff. And then she said, okay, can you, tell me when a certain item goes on sale and and he was like yeah and then she came to me and was like hey and he's gonna update me when this pair of jeans is on sale and and i was like a bit skeptical and i was like is he really though or is your agent just hallucinating a lot of times you can tell an agent something be like remember this and i'll be like great note it and then you come back to it and you're like wait what did you do with that information and it's just like i'm just remembering it just noting it mental note and you're like no you gotta write that somewhere put it in your memory or something.
17:24Right. So I was like, is he really going to, or is he just like, yeah, I'll do that. And then the next turn, whatever. So then I asked Andy, how are you going to notify Julia jeans on sale? And he said, yeah, I have a list of websites that sell this product and list of product pages. And I've set for myself a recurring job every morning, 9am to visit each of the websites using the browser that I have in my environment. And I checked the price and then I compared to the previous prices. and then if it's on sale, I'm going to update you by sending you a message in WhatsApp. So the agent schedules for itself, it has these tools of scheduling tasks, and it uses them at its own discretion to accomplish a task that the user asks it to do.
18:07So the agent does that by setting, it can set a cron schedule, and it can set a one-off route task by just setting the time when that task should run. And then the agent actually has an MCP tool to schedule a task that the agent pulls the tool and then that process that runs inside the agent's environment sends it as a system message to the host process. The host process writes it into the agent's inbox. The reason why the agent can't just write it to its own inbox is because it's a SQLite little database. And because we've got the host process and the agent, they're running in different processes, the SQLite DV doesn't actually sync between them.
18:52necessarily and you get conflicts so you have to use a specific mode to read and then you get conflicts and you start getting all kinds of errors if they're both reading and writing from the same db so that's why you've got an inbox that only the host writes to and only the agent reads from and then you've got the outbox that only the agent writes to only the host reads from so that's why it kind of has to do this whole loop and the agent can't write if it's on inbox so to make sure I understand for recurring tasks, the agent is essentially just scheduling a cron in its own agent runtime. But for one-off tasks, it calls an MCP that go out to the host process, the host process then writes a message to the inbox.
19:35So the agent doesn't actually schedule a cron in its own environment. The reason is because the agent's environment isn't always up or isn't always running. Okay. This is another good question. So yeah. Okay. So they're both going through the same flow, which actually I'm guessing you then use in other types of context, where essentially you have a tool that wraps around, here is a deterministic way that I want to send a particular message to my host, system message. Sends that, the host schedules either on a cron or if it's a one-off via a message with a timestamp, then it triggers the agent by sending it to its inbox.
20:12And then at whatever time the polling process inside the agent runtime will say, oh, it's ready and hand it to the agent. that's very close so there's another detail there though that the aging containers spin up and spin down based on when the agent needs to be active so when a message comes in the i can have in one nanoclaw instance running on a machine and then actually we stress tested this last week or two weeks ago one machine with 16 cpus 64 gigabytes ram can run about 200 containers aging containers like containers running agents, running cloud SDK in parallel at the same moment concurrently.
20:52But on a smaller machine, if I only have two CPUs, eight gig RAM, I'm only going to be able to run eight or 10 agents, containers in parallel. So if I have an instance and I have a whole bunch of different agents, because I've got different agents and different groups and different people, I don't want to have those containers running and up all the time. Very quickly, I run out of CPUs, run out of RAM, run out of space. So the containers spin up and spin down. The persistency is in the mounted folders and in the container images. So the agent can't write something in the future, for example, future cron job in its own environment that has to be written to the host.
21:38Now, the way it's implemented, just to keep everything modular and have each agent, each session managing its own things and keep the host process relatively simple. The agent calls a tool, like you said, gets, it's a message that's sent to the host process saying this agent wants to schedule a task for this time. The host process immediately writes it to the agent's inbox, but with a process after timestamp. and then the container will spin down after some period of idle. By default, 30 minutes, no messages coming out from the agent, no activity, the container spins down. And then the host process is doing this sweep over all the agent outboxes and inboxes looking for messages that need to be processed.
22:23And when it sees a message that has its process after timestamp set for the current time, it spins up that container. and then the agent, the loop running inside the container, by default, immediately when it starts up, just starts pulling the inbox, sees the message there, pushes it into the agent SDK. So that's the process there. Awesome. Agents are getting smarter every day, but even the smartest agents get stuck without the right context and the right tools. That's where Notion comes in. With the recent launch of custom agents, Notion became the collaborative AI workspace where teams and agents work side by side.
22:57And now their new developer platform is turning that workspace into infrastructure developers can build on. What stands out to me about Notion's developer platform is how much they've built underneath the surface. The platform ships with a set of new primitives that change what you can actually do. Workers are Notion-hosted sandboxes where you run custom code, database syncs, agent tools, webhook triggers, without spinning up your own infrastructure. The CLI lets developers and coding agents read from Notion, take actions, and deploy workers through one interface. And with the external agent API, you can bring agents like Claude or Codex into Notion as workspace participants with real permissions and triggers, not just a sidecar tab you're copying from.
23:33Pulling in research sources tracked externally, syncing them into Notion, and having an agent help triage what is worth covering all in the same place the team already works. Learn more about Notion's developer platform today at notion.com.s-e-d. That's all lowercase letters, notion.com.s-e-d, to try Notion's developer platform today. And when you use our link, you're supporting our show, notion.com. Your customer lives in the real world and it's messy. Intermittent connections, mid-onboarding drop-offs, edge cases on devices you've never tested. Mobile apps reflect reality in a way no other surface does.
24:10Yet, from an engineering perspective, they're the hardest to understand. It's common for mobile engineers to see green backend dashboards, normal error rates, no crashes, but inevitably, somewhere, there's a frustrated user watching your app spin. After a few seconds, they'll lose patience, close the app, and turn their attention somewhere else. They may never come back. The worst part? Most observability tools never see any of it. BitRift, on the other hand, captures 100 % of mobile data, unsampled and in real time, so it's immediately queryable by engineers and AI agents. It's mobile observability built for the real world.
24:44Try BitRift today. bitrift.io slash signup. Think about your mobile app source code. Once it hits the app store, it's out in the wild. And without the right protection, decompiling is easy for malicious actors looking to steal your IP or tamper with your software. That's where GuardSquare comes in. GuardSquare provides the highest level of mobile app security for Android and iOS applications and SDKs. Their advanced tools integrate seamlessly into your CICD pipeline. We're talking polymorphic, multi-layered code hardening techniques and automated runtime application self-protection, paired with mobile application security testing and real-time threat monitoring to deliver the highest level of mobile app security without compromise.
25:31Don't leave your hard work exposed. Secure your mobile applications today. Go to guardsquare.com to learn more. Let's talk a little bit about lifecycle, right? So you've already talked about how you'll spin up, spin down, you're managing persistence. It's in the mounted folder and in the inbox outbox. So presumably these things could be as long lived as you want. They can be quite durable. How do you deal with things like context window compaction, those types of things for one of these agents? Yeah, the compaction and how we do that is we take a very simplistic approach and try to not reinvent the wheel and try to not build in a way where code is going to be redundant with the next update of agent SDK, the next model release, the next version of cloud code.
26:19So we're actually just using cloud code default compaction with just some instructions, minimal instructions telling the agent what's important to retain in the context. And what's really important to retain is the sequence of messages going in and out from the user and the agent Tool calls you can draw. And actually the content of the message can be summarized. And the instructions tell the agent to summarize the content of the message if it's long. But the sequence of messages saying the sequence, the order, who the messages are from, the timestamps, that needs to be preserved. And not necessarily all of them.
26:56You can summarize and say there's this whole beginning of the conversation that happened. And then here's the last 15 messages. And then one or two of those might be really long. summarize the actual content of the messages, but keep the wrapper saying, here was a message from the user at this timestamp. And here's a response from me at this timestamp. So compaction, we let the model, Opus, Claude, Ord Codex, GP55, we let them handle that. And in terms of the session, it's actually just one session that keeps going by default. So the session just goes on and on, and it auto compacts when it hits the limit.
Read the full transcript
27:32We set the limit a little bit below the default. So default for agent SDK is 200K at the moment. And it's kind of maybe a little bit solved with the last couple models, but the models used to get really, I think they get some kind of like context anxiety when they would get close to their context window threshold or compaction threshold. So they start to get kind of dumb, try to finish things really quick and take shortcuts. So we set it a little bit lower and set the compaction threshold at 165k. And then by default, it just keeps going, same session, endlessly. That works quite well. When it compacts, there's a pre-compact hook that takes the full session and puts it in the agent's environment.
28:17And there are instructions in claw.md saying you've got the full context of conversations from before compactions in the conversations folder. So the agent can look through previous sessions if it needs to pull in previous context. And the agent has instructions saying, save things, kind of the LLM wiki model of just save things in your environment, create files, create folders, create indexes that specify what you have in your environment, what the different files and folders have in them, and update your Cloud.md to reflect what are in those files and folders. and even create systems for tracking things and systems for your own memory and then just designate what that system is.
29:00And that works quite well. The issue is that at some point, the actual cloud code session file, the JSON-L file, just starts to get really, really big. It can get to be 50, 100, 200 megabytes. And at some point, it starts to break down a bit. So that's something that we need to add to at some point, to rotate that session and pull in some context to the new session. But actually just compacting three sessions for that chat assistant use case, it works really well. Let's talk a little bit about interaction models. So you mentioned you can spin up a whole bunch of different agents. Can they talk to each other?
29:39Like, what does that end up looking like? I know a lot of folks who are way down in this are like, oh, yeah, I have my product manager and my sales manager and my this and my that. So how does the multi-agent story work within NanoClaw? Yes, your agents can talk to each other and your agents can actually create other agents. So your agent has an MCP tool, create agent, and it can give a definition, name, instructions, what that agent is supposed to look like. And that creates a second agent in its own container with its own file system, its own environment, own memory. and then when your agent creates an agent it now has in addition to destination of the slack channel you're talking to it from and maybe your whatsapp group that you've connected the same agent it now has another destination of agent b and agent a can send messages to agent b and that just means it can write a message that it wraps in a little xml tag saying to agent b and so just like it can send you a message in Slack and WhatsApp, they can send a message to Agent B.
30:41That's again, just a message coming out of the agent's environment. It gets put into its outbox. The host process picks that up and routes it instead of to a messaging app, it just routes it to the other agent and puts it in its inbox. And that's agent to agent communication within one nanopah instance. It's just messages in the outbox going into the other agent's inbox. And in order for me to allow agents to speak one to each other, I just need to configure them so that they have the other agent in their list of known and approved destinations. So that's both for permissions and for discovery.
31:19So the agent knows who it can talk to. The other mode is just putting them in the same channel and then just having them speak to each other through whatever messaging app. So if you want to have visibility into what they're saying and you want to have a conversation where you're part of the conversation, then you just add two different agents to the same Slack channel. They talk to each other the same way you do. One thing that is missing in order to make the agent to agent direct communication outside of the channels really useful is an approval gate. So the same way I can put an approval gate on the agent sending a message, sending an email, I really need to have the ability to have an approval gate between agents talking to each other.
32:00And a really powerful use case for agent to agent is having two different agents that belong to two different people. So I've got my agent, you've got your agent, and I can ask my agent a question, ask it to do something. As part of doing that task, it can reach out to your agent and ask it a question, ask it to grab some information. If we trust each other with all of our information, then we can just let that happen, let our agents chatter behind our backs and maybe that's okay. But if you're talking to your agent about some private things that you wouldn't want to be shared with other people.
32:31So then you want to have an approval gate before your agent sends a response. And you want to see what that response is and verify that it's not sharing something that you wouldn't want being sent out. When you have that type of multi-person, multi-agent, and this is fascinating because we're getting into a question that everybody's grappling with now, which is multiplayer AI. Does that need to be through an external channel like a Slack or something like that? Do you have a peer-to-peer model in this? Like, how is that working? So I think if multiple people are going to be talking to the same agent, it does make sense to have that happening in like a Slack channel.
33:04I think there's a false sense of, and this is just from like the psychology of the end user. If I'm talking to an agent in a DM, that feels like a private conversation. And if somebody else is speaking to that same agent in a DM, then that gives me a false sense of privacy because that agent could just be sharing everything I'm telling it with another person in a different VM. So I think the right way to do it is if we're all speaking to the same agent, probably have that in the group and have that visible and immediately apparent to all the users. So everybody's aligned that this is a public forum and anything I share with this agent is public information.
33:45Where it gets interesting is I'm chatting with my agent. It's got all this information, very useful information that's really useful to people on my team. And the same way people on my team can come and ask me questions and I can share with them some of that information, but not share other pieces of information that I've got in my head. You do want to have your agents as they're building up this LLM wiki, this knowledge graph of everything they know, a Wikipedia of you and everything you know and everything you're working on. You want them to be able to share some of that information with other people and not share other parts.
34:17And that's where I think approval gates are interesting and are useful. I think the first step is an approval gate where every single message going out, I see and I have to approve. And then you can start playing around with maybe classifiers. So maybe a model, a smart model, it doesn't have to necessarily be an agent. Maybe it's just an LLM call that looks at what's being shared and looks at a set of instructions and makes a call on whether this is blatantly safe and innocent information, or if it should be routed to me to get my approval. And that's obviously going to be trading off a little bit of safety for convenience.
34:56So you got to make sure you shouldn't be doing that with highly sensitive information and anonymous users from the public engaging with it. But if it's for someone on my team, and I know that some of the information maybe isn't ideal for other people to have, but it isn't going to cause a real crisis. So maybe I'll allow an LLM to look at some of the responses and auto-improve some of them. Interesting. Okay, another slightly different topic. So I'm an engineer. First thing I think with agents is I'm like, okay, I want it to build things for me. Maybe I want to use multiple agents, one that builds and one is a code review agent or things like that.
35:32But the only inputs and outputs here are messages. How do I navigate the creation of artifacts using these agents or getting those in and out of those kind of local container agent runtimes? So we've built internally and will likely open source this soon, maybe mid-June or maybe recently. I'm not sure. But we created our own agent factory. And I know everybody's doing that. Every company that's worked or its way that's got their own agent factory, our agent factory reviews pull requests to the NanoCode open source project. It reviews, but the most useful part of what it does isn't actually the review.
36:12The review is somewhat useful. So maybe we'll catch some edge cases. But I think the core part of the review in an open source project is really about, is this aligned with the project philosophy approach where we're going? And that is a lot of taste and judgment. And agents aren't very good at that, even with instructions. So they're doing the security review, code quality review. But then the agent creates a testing plan. And then the testing plan goes to a testing agent. The testing agent gets a VM that's been spun up for it with the pull request branch checked out in that VM. The testing agent SSHs into the VM, runs the NanoClaw instance of this PR with the PR's code changes, and then actually runs it.
36:58It runs some automated tests, but it actually runs it and then starts sending messages via Telegram and reading out the responses. So doing real kind of QA, you know, testing, you would do manual testing just before you deploy something. And it does that in a real environment, checks the results, sees, first of all, I sent in a message in Telegram. Did I get a response? What is the response? And then can actually also look through the logs and look through the DB and see that everything makes sense and is as it should be. And it can improvise. So if the environment doesn't have the necessary dependencies for whatever reason, it can install them.
37:32It can add test cases and then it writes out testing results. And then everything gets posted, the whole process of triage, review, test plan, test results, decision at the end. It all gets written out to a channel. And then at the end, the first agent gets the test results and then has a tool where it can write a GitHub CLI command. It doesn't have credentials in its environment. So it just writes the command as sort of a draft, and that command gets presented to us in the Slack channel, and then we can approve or reject. And then if we approve the command, it gets run with our credentials. So if I approve it, it gets run with my credentials.
38:12If Daniel approves, it gets run with his credentials. So that's us taking the action. It's not the agents just automatically merging pull requests. It's the agents going through this process. We look at the work they've done, ask some follow-up questions, do some tests ourselves, and decide to merge or press changes or close or whatever. Those artifacts, so the agent is running its environment. It writes a test plan markdown file to a specific location and then sends a message to the host saying, I've done the test plan. This is ready for testing. the host process orchestration, non-AI process, purely deterministic, gets the agent's message and then grabs the test file and then we'll put it in the, say the test plan.
38:59It grabs it from the agent, the review agent that writes the plan, puts it into a location in the environment of the agent that runs the tests and then sends that agent a message going, hey, you've got a test plan in your environment, do something with it. So if I'm hearing you properly, essentially the agent is writing artifacts, whether it's a markdown file or code or whatever, it's writing that locally in a location. It sends a message to the host saying, hey, I've got this artifact here. I'd like this other agent potentially to take a look at it. And the host then either copies it or does something to pull it out of that environment and puts it over in the other agent's environment and sends it a input message saying, hey, you have this new stuff, go and look at it.
39:40Is that correct? Yeah, that's exactly correct. We, rather than having the agent write arbitrary messages to the host, we have it, we give it a tool that it calls saying send to testing, and then it calls a send to testing tool. And then the send to testing tool writes a, you know, deterministically writes a message out to the outbox that the host gets. And the agent has to put the message in a specific location. It can't just arbitrarily write locations for the host process to read because there's all kinds of security issues with that. So there's a specific location. The agent has to put the message in its environment, the file, the artifact.
40:18The host process only grabs it from that specific location. If it's not in that location, the host returns an error message to the agent. Okay. Let's shift direction a little bit. I want to ask you about some of the sort of product philosophy. You mentioned judgment for where the open source project is going. And I think you've taken a few very deliberate choices in how you architect. And the forking model is very interesting. So maybe give me the big overview or the quick TLDR, and then we can dive into details of what is your product philosophy? What belongs in this open source project? And what does not?
40:55How do you approach it? So the product philosophy is about minimalism. There are prerequisites. It has to be secure. It's got to follow the security model. But then a key part is minimalism, keeping it small, keeping it simple. I think this whole world, the Wild West of running autonomous agents that can run bash commands and write arbitrary code, it's something that everybody should be a bit nervous about. That shouldn't, for developers, for security experts, that should give you a bit of pause. So I don't think, when I sat down to write this, I said, okay, this is inherently something that is worrying in order for me to be confident that when i'm adding new capabilities new features the project is evolving it's staying secure in order for other people to be confident running it i don't think they should just assume that it's okay because it's got a bunch of stars or you know some other people looked at it and thought it was okay i think people should be looking at it themselves and verifying if it's okay so in order for it to be auditable it has to be small.
42:03So keeping it small, keeping it minimal, that's a core part of the philosophy of the project. Small and minimal though means that it doesn't have everything that everybody needs. So my sort of rule of thumb is I'm only going to merge something or add something if I think it's relevant for 80 to 90 % of users. So if something's a bit niche, it's a very specific environment. It's a very specific use case, even if it makes sense for that use case and it's clean, I'm likely not going to merge it. Maybe if it's really minimal, if it's only a couple lines of code and it adds something great, I'll merge it.
42:40But if it's adding hundreds of lines of code and it's something that's a bit niche, even if everything is fine, it makes sense, it adds value for certain users, I would merge that. I'm very intentionally trying to keep the code base small. So then if it's minimal, if it's small, it's got 80, 90 % of what people need. And it means that there might be a small subset, 20 % of users that it's got everything they need and they don't need to customize it. But actually for most users, it only gets them 90 % of the way there. And everybody's got to add a couple of things to make it work for their use case.
43:14And I think though, that that's true for any tool you use. Every tool that you use, you wish there was just that one thing that worked a little bit differently, or it just had that extra capability. So the approach that we took with Nanoclaw is customization on the code level. Changing the code is supported. We're trying to support that as a first class capability. So everybody can get exactly what they want and only what you want. You can get rid of the things you don't want. So if you want it to run on Telegram, on Signal and not have WhatsApp, then your fork only has Signal. It only has a code related to signal, it doesn't have any code related to WhatsApp.
43:55It doesn't have dependencies from WhatsApp. And then if you want yours to format the messages in a different way, if you're using it purely as a coding agent and you don't want it to have messages wrapped with timestamps or some other stuff, you can change the formatter and change how messages are piped into the agent. And all of those things kind of work together. If it's really minimal and really small, then people can customize it. And we're not going to be constantly pushing out new updates and new versions that are going to break their customizations. And it can stay maintainable. What it does mean, though, in order to be able to customize it in a way that upstream changes are not going to break you and you'll be able to pull them in, you have to follow a sort of modular way of coding, which is in a way kind of just coding best practices.
44:45But you do have to code in a clean way and add things in a clean way. What that means is essentially add as many files as you want. But if you're changing existing files and you will need to, you're going to have to integrate somewhere, keep your integration points really minimal. So have one two line integration, you import something, you call it in one place, and then all of your other customizations should be in new files. And if you follow that, you're going to be pretty safe. There are going to be minimal conflicts when you're pulling in upstream changes. One thing that we haven't done a good enough job with is communicating that in the documentation, in the instructions for coding agents, communicating that to the community, to the people who are using it.
45:32There's definitely been some frustration and some challenges with people updating their versions, pulling in upstream changes when they've actually forked the repository and diverged from upstream. that's actually something that we've been working on in the last few days of just defining what is our contract with the open source community with fork maintainers people who have their own forks what's our responsibility towards them and what do they have to abide by if they want to ensure that we're not going to break their fork with every every new update and what are the things what's the surface area of the project that we can commit to and if we do change it we'll consider it a breaking change and will provide a clear, well-defined migration path.
46:17So that's something that we're defining now. But I think there's a huge advantage to this approach. Ultimately, we've got NanoClaw open source, and then we've got NanoCo, which is a company. It's our company. We've raised capital, raised$12 million. That positions us really well to continue to build and maintain NanoClaw as a free open source MIT license project, make sure that that's not going to kind of drop off and be abandoned. We're going to be able to continue to maintain it and build it as infrastructure that people can build on as a free project people can use for their own personal use as something to use in your company, in your business, for your team, you can build services, products, platforms on top of it.
46:59As a commercial entity, we're focused on bringing agents to companies. So companies that are not building their own internal agent platform on top of mouth off, but they just say, I don't want to build it myself. I don't know how to build it myself. I do know how, but I don't want to dedicate the engineering hours to building our own internal agent platform for the marketing team, for the sales team. So for those companies, we're going in and helping them get set up with agents. And it means we're doing some services, FBs, creating a custom deployment for each company, and then managing and maintaining agents over time.
47:32So that means that we have a lot of forks ourselves. And you're incented to make that easy. Exactly. We will be the largest maintainer of nano club forks going forward at some point. I mean, we are already and we will be in the future. So that means we have to figure that out. We have to nail it. We have to figure out how to maintain a bunch of forks and not break those forks. And we're dedicated. And that's a big focus right now, figuring out how to do that well so that we can do it for ourselves. We can do it for the community. and we can have a big community around the open source project. And I think it's doable, but it is a different approach.
48:12I don't know of another project that's taken this approach before. It's interesting to me, like it is very like AI forward. It's like taking this model that we have now of software is this much more moldable, everybody can change it type of thing. I saw this change play out a little bit in the front end with the rise of things like Shad CN, which is like the old model was, hey, you've got this library and maybe you extend it in some ways. And the new model is now just copy this component and then make it yours. Do whatever you want. And you're kind of taking that at the whole project level here of saying, hey, copy this and then make it yours.
48:46Yeah, I love that you said ChatCN because that to me is an analogy that I have in my mind and something that I've thought about. And in some ways is one of the inspirations for it. And I haven't heard anybody else talk about that. I mean, it's obviously super popular in front end, but it's, I don't know how well known it is more broadly. But yeah, it's sort of the Shadzian model, but there's a bit of a difference there because with your component library, you create your button. There isn't that much innovation happening on the HTML level. You copy once and you're done. You're not going to need to do the update.
49:19So you have another layer on top of that. Because we're building on coding agents and the coding agents are constantly changing multiple times a day. So we have to be able to pull in changes. We have to be able to expand and improve capabilities. Even if we said Nanoclaw is perfect, it's complete, which it isn't. There's a lot to add, even if coding agents are static. But the coding agents aren't static. They're constantly evolving and changing. So we have to be constantly updating. It means our users have to be able to be pulling in updates. So yeah, it's a bit complicated. But the way we're doing it, and this was the thing.
49:50So Andre Carpathy tweeted this big thread about Claws, and he called out specifically Nanoclaw, and said it slightly blew his mind, which I think is the height of compliments from Project Propathia. Yeah, you've made it now. Yeah. And the thing that he specifically found mind-blowing was this approach of using skills as the way to customize Nanocl, to customize your fork. So if you add integration with Spotify, you go and modify your fork so that it connects to Spotify. So you make all these code changes, then you want to contribute it back to the project so other people can have their Nanocl agents talking to their Spotify.
50:27The way to contribute that back, Spotify is not the kind of thing that's going to be merged into the core code base because it isn't a capability that 80 and 90 % of users need, but it's a great capability. It's really cool. It's definitely the kind of thing you want to share and other people would find useful. So the way you do that is you create a skill. And this skill is not a sort of runtime skill that the agent uses to do some work for you. It's a skill that's telling the agent, here's how you can modify Nanoclaw to connect to Spotify. Oh, fascinating. Yeah, you think of that as a skill. You're wrapping up essentially the code changes necessary as a prompt and maybe a set of scripts or whatever, anything that can go in a skill.
51:07And you're saying this is a modify Nanoclaw to connect to Spotify skill, and it will change whatever's needed in the host process and communicate whatever's needed to the agent runtime and make it so your Nanoclaw can talk to Spotify. Exactly. So the skill standard includes with it the concept of having scripts, but it can actually have arbitrary files. So putting TypeScript files in the skill folder and saying this is a TypeScript file that you've got to add. Oh, just copy this over here. Fascinating. Yeah, yeah, yeah. That's really interesting. It's almost deterministic. I just need to copy it.
51:44But then it's not quite deterministic because I do need to integrate it in some place. And that place is going to change. It's not an exact line number. Right, right. Because my nanoclaw is different than your nanoclaw. So it's going to be, yeah, yeah, yeah. No, this makes a ton of sense. Yeah, but we're communicating to the coding agent the intent of here's what we're trying to do. Here's how it's supposed to work. Go to this file. You'll need to import it there and you'll need to add it somewhere within the function that handles the routing of the messages, but not an exact line number. And even the edit is not an exact edit.
52:17It might need to change it a little bit and add some parameters. And the agent will be able to figure it out and do that almost deterministically. Some skills are deterministic. all I need to do. So if there's a type of skill that is needed often enough, for example, adding messaging apps. So when I need to integrate with a new messaging app, those skills, for the most part, can be applied fully deterministically. And to do that, we just need to create a kind of barrel file. It's a registry, you add an import, you've got a little registry thing, anything that's imported into that file gets registered and is available in runtime.
52:54time. So for skills like that, we add the extra overhead, the extra level of abstraction, because it's a hotspot and it's needed. And there you've got a few files that you add, and then you've got appends. So append only, appending an import to a file. And then you've got maybe something that you need to add to a command you run. So npm install, I need to install some dependencies. There's no conflicts there, and that can be done basically deterministically. And then, by the way, most skills also have a certain amount of user instructions. So the coding agent runs all these things, adds all the code, integrates it, runs some tests to verify that the integration works.
53:35So integration tests together with the unit tests. And then the agent is prompted to walk the user through authenticating with Spotify, getting the token, and it can guide you through. Yeah, yeah, yeah. That's fascinating. One more direction I want to go down. So you referenced Claude Code a lot. You also threw in a line about being able to use GPT-55. And you've talked about the way that the agent underlying tool landscape is changing very quickly. So one of the things that I immediately wonder is, can I plug in whatever agent harness I want? Can I use Codex? Or recently, I've been going deep on using Pi for things.
54:12Can I swap in Pi as the core agent harness? Or how does that layer of customization work? Yes, we have a agent provider abstraction. And that's another place where we have that registry model where you can plug in another coding agent without having to kind of mess with existing files, just add files, append to existing files and add dependencies. we have supported with skills codecs and open code because the project is really small because you've got reference implementations there of how other coding agents were added if you point pi at the project and say add an integration for pi here it will do it in like five minutes and you know 50 whatever a couple hundred k tokens so it's a trivial kind of addition it's very greenfield coding.
55:06And that's, I think, one of the ideas behind it, that any capability you want to add, it doesn't matter if it's the 10th capability or the 1 ,000th capability, it's always greenfield because the project stays really small, really minimal. So you can add that. If you add it, you can contribute AdPy as a skill to the project. And there's a pull request open, and maybe we'll add it soon to add ACP, Aging Client Protocol, which once we merge that, it will just support basically every coding agent but we do integrate on the level of coding agents so we don't integrate with like llm and route api calls from cloud codes to different llms i think that isn't really working anymore probably hasn't worked well for a while but in the as we move forward i think it's going to work it's going to be worse and worse because the agents the models are being trained in the harness.
56:01They're being trained in Cloud Code. Yeah, and Codex. If you allow OpenCode and you allow Pi, they can handle the multiplexing to whatever model you want. Like that's not a big deal. Yeah, if you want to use all kinds of models, use Pi, use OpenCode. But also, I mean, the instructions don't really transfer over. There's a whole family of models that are like Cloud family of models that are distillations or whatever, Cloud. And they probably, there's some compatibility between them. But when you go between Claude and Codex and Gemini, they're very different models, very different instructions, skills, tool definitions.
56:35Things don't just transfer over as well as people would like them to. Well, so in that domain, right? So I'm playing a lot. And there's a few things that Nanoclaw has done that I've stolen in different plays for some of my own play. Like you have a lot of really interesting stuff in here. But one of the things I've been thinking about is how do you create small custom agents with self-tuned, eviled prompts for particular purposes? And sometimes you want to do different models. So the thing I'm kind of wondering about is how much of the prompt of the different things, of the skills, et cetera, is kind of baked in to NanoClaw?
57:14How much can you customize per agent where you say, for example, let's pick one of your tools. You had a messaging tool, right? That is how you send messages. Can I have a custom messaging tool per agent? Because I happen to know this one's running GPT-5 and this one's running Opus, and they need slightly different prompts to be able to use the tool properly? Not by default, but you can fork the project and make that kind of change. So changing the instructions, the tool definitions of the default tools. We have the ability, you can set for each agent instructions, skills, models npm packages and mcp servers and some other things but the ability to tweak the tool definitions per agent don't have that maybe it's something we should add you may not need to right it may be that everybody's running clod models and or clod family models and they stay as the front runner but yeah thinking about this world of who knows which model is going to be best in a year and and as you say highlight they behave very differently we are moving towards CLI tools as well.
58:17So there's a whole set of tools for scheduling tasks, schedule, list, update, delete. And we're moving towards having all of that as a CLI tool. As you start to add more and more, it just gets to a ridiculous amount of tools. And it's a big debate, but I think CLI is pretty safe. And then it starts to become less and less critical, I think, because it isn't always there in context. Even with the MCP tools, now that you're using tool search, So a lot of the tool definitions aren't actually in context by default. And you do have the ability to adjust your own instructions per agent. There's certain pieces of the Claude MD that are like system definitions about how to, you know, the environment that it's in and how to schedule things and instructions like that.
59:06Those are the kind of things you have to fork to adjust. In the end of the day, though, the harness, the orchestration, NanoClaw as orchestration, it's trying to not do too much. So it isn't injecting a lot of messages. It isn't changing things. It formats the messages going in. It provides some instructions about how to use the tools. But we're building very intentionally in a way of assuming that models are going to get better, agents are going to get better. Let's leave as much as possible to the models and to the agents. so that we benefit from the next update. And so the next update doesn't just make weeks of engineering work become redundant.
59:48So memory system, for example, I have no doubt that Anthropic has dozens, maybe more people on making the agent memory better. That's probably at every level from tuning the instructions to pre-training, to post-training, to everything. So we're not trying to build really elaborate memory systems that are kind of deterministic and all this kind of clever engineering. It's just a set of instructions that tells the model, here's an idea, some guidelines about how you could save things in files and folders and move things from files to folders and reorganize them and keep an index of what files and folders you have.
1:00:27And you can change this definition of the memory system in this actual file to make it more or less complicated. But it's your memory system to manage. as you see fit. And that guarantees us that as the models get better, they're going to manage memory better. So we're getting close to the end of our time. Is there anything we have not talked about that you think would be important to leave listeners with? I think one of the things I would say is last year I spent a lot of time building agents out of LLM calls in an agent framework. The one I was using was OpenAI's framework, which is also agents SDK.
1:01:11But that's where you're building agents by, you got an agent loop and you're defining everything and you got to deal with caching and you got to deal with the tool definitions and compaction and everything. That's really hard, right? Thinking about caching is hard. Thinking about compaction is hard. Session management. There's months and months and months of work just to get yourself a basic agent that has the core capabilities that Claude Code has or that Codex has or that OpenCode has. So if some people out there are still in that place where they're going, maybe I'll use LangeChain or LangeGraph and I'm going to build an agent, my advice to you is don't do that.
1:01:50Save yourself the work. There's so much value that can be unlocked from using Cloud Code, from using Codex, using their SDKs to build all kinds of wild agents on top of that with all kinds of orchestration and workflows, using skills, using instruction, agents using, you know, regular good old engineering, building products and platforms on top of the existing agents, focus on unlocking the value over there rather than re-implementing the basic element, basic building block of an agent. Look at a coding agent as a building block and then build a product or a platform or something on top of that.
1:02:29And then in terms of, I think a lot of people get stuck on which model, which agent don't get stuck there. any of the frontier models are good enough, but use a frontier model when you're building. Don't try to optimize for costs. If the thing you're building for doesn't make sense, unless it's super optimized for costs, it might not be a high enough value use case where it makes sense to dedicate your time building towards that. Find a use case where if you nail this, it's going to be worth spending tokens on a frontier model. And there's definitely optimization to do, but first focus on just nailing the thing, creating a lot of value, doing something that's really valuable using coding agent, using frontier model.
1:03:10And once you figure that out, then maybe think about the economics of it and how do you optimize low-women cost. Seems like a good place to wrap.
1:03:34Thank you.
From the publisher
AI agents have shown remarkable potential to function as persistent digital assistants that are capable of monitoring data, managing communications, and taking action autonomously over long periods. OpenClaw was one of the first serious attempts to fulfill that vision, connecting frontier coding agents to messaging platforms like Slack and WhatsApp and letting them run continuously in the background. However, OpenClaw largely set aside questions of security to pursue that vision, leaving credentials exposed in the agent’s environment and giving agents broad access to data and services far beyond what any given task required.
NanoClaw is an open source project that takes a zero trust approach to agent orchestration. Rather than relying on instructions to constrain agent behavior, it isolates each agent in its own Docker container, keeps credentials entirely outside the agent’s environment, and enforces human-in-the-loop approval for sensitive actions.
Gavriel Cohen is the founder of NanoClaw and he joins Kevin Ball to discuss the security architecture behind NanoClaw, how the agent sandbox and proxy model work in practice, how agents communicate with each other and with the host orchestration process, how the project approaches context window management and long-lived agent sessions, and more.
Kevin Ball or KBall, is the vice president of engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript meetup, and organizes the AI inaction discussion group through Latent Space.
Please click here to see the transcript of this episode.
Sponsorship inquiries: sponsor@softwareengineeringdaily.com
The post NanoClaw and the Rise of Personal AI Agents appeared first on Software Engineering Daily.
