In short
Cloudflare’s work on agents and Model Context Protocol (MCP), focusing on “Code Mode” (models writing TypeScript to call tools) and how Cloudflare exposes its ~2,500 API endpoints through an MCP server using only ~1,000 context tokens. Also covers why many people misconfigure MCP, and how to safely execute model-generated code.
Guest backgrounds
- Matt Carey: Works at Cloudflare on the Agents SDK (open source) and MCP support. Joined Cloudflare in October; previously worked on MCP and agents. Releases include Code Mode and a server-side Code Mode running inside an MCP server.
- Nicky Pike: Field CTO at Coder.com; describes Coder as DevRel for the C-suite and focuses on bridging customer voice to product teams.
Key claims
- MCP isn’t “broken”; Cloudflare uses it to its best capabilities.
- Tool-calling can bloat context windows; Code Mode reduces context by letting the model write code against a small set of MCP tools.
- Server-side Code Mode executes safely on Cloudflare dynamic workers (sandboxed, restricted outbound fetch).
- Most MCP servers map one API endpoint to one tool; Cloudflare instead uses “search” and “execute” tools so the model can discover and call endpoints on demand.
Notable examples
- Cloudflare MCP server exposes the full Cloudflare API behind two tools: a code-based search over the OpenAPI spec and an execute path using cloudflare.request.
- Demo workflow: model writes code to deploy a Next.js site to Cloudflare Workers (no code saved locally).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding the Role of a Field CTO
1:04 to 2:10
Nicky Pike explains the responsibilities and importance of a Field CTO.
“So I get that question a lot and it's, you know, half the people understand it, half the people don't.”
Challenges of Local Development
2:10 to 3:16
Discussion on the issues with local development environments and security concerns.
“And, you know, the time to first commit is a metric that almost everybody knows.”
Introduction to Matt Carey and Code Mode
4:11 to 6:00
Matt Carey shares insights about his podcast and the concept of Code Mode.
“It's called You've Been a Bad Agent and we just chat absolute rubbish about agents.”
Navigating MCP and AI Deployment
6:00 to 8:13
Discussion of the challenges and developments in AI deployment and MCP at Cloudflare.
“Maybe some you can help, some you can't help.”
Explaining Model Context Protocol (MCP)
8:13 to 11:34
Matt explains the Model Context Protocol and its significance in agent development.
“There was a number, but the number keeps changing.”
Revolutionizing Agent Functionality with TypeScript
11:34 to 14:06
Exploration of how TypeScript can enhance agent functionality in MCP.
“with just writing TypeScript versus all the context-filling, tool calling, I guess you'd call it, right?”
Understanding Context Windows in AI Models
14:06 to 18:11
Learn about the limitations of context windows in AI models and the implications for automation.
“And so you really don't want to be filling the context window too much.”
Innovations in Cloudflare's Code Mode
18:11 to 23:05
Discover innovations in Cloudflare's Code Mode and how it enhances API interactions.
“The proper thing to do is like you make the agent do code execution and then you become like a code agent.”
The Impact of Friday Demos at Cloudflare
23:05 to 24:28
Explore the significance of Friday demos at Cloudflare and their role in innovation.
“and then it can just call the endpoints workers script deployment and it'll just deploy it to the cloud super super fun And you can have these like insane demos where none of the code ever gets saved on your machine.”
The Evolution of Developer Workflows
24:28 to 27:43
Understand how developer workflows have evolved, particularly with AI tools like Claude.
“Look at this thing I built and then 50 % of the time it breaks, but then 50 % of the time it's awesome.”
Show all 46 chapters
Interacting with AI for Future Innovations
27:43 to 28:07
Learn how to effectively interact with AI tools for ongoing innovation and development.
Evolving Workflow with AI Models
28:07 to 29:11
Learn how the speaker's coding workflow has changed with AI tools like Claude.
“It's changed a huge amount over the past.”
Custom Git Solutions for Safety
29:11 to 30:50
Discover the importance of custom Git solutions for protecting code integrity.
“now I'm just going to have an IDE open when I want to actually visually look at code When I want to review, I'll open an ID.”
Effective Communication with AI Models
30:50 to 32:31
Understand how to better engage with AI models for optimal results.
“Like even chatting with the model, it's just feeding a dopamine rush that is not real.”
Organizing Code Projects Efficiently
32:31 to 34:36
Explore strategies for organizing code projects to enhance productivity.
“I don't think I'd do anything special with Claude.”
Navigating Open Source Contributions
34:36 to 37:05
Learn about the dynamics of contributing to open source projects.
“So therefore it's cool to just open that one directory.”
Balancing Speed and Quality in Development
39:48 to 42:00
Discuss the challenges of maintaining code quality while using automated tools.
“in my home lab, in my dev lab without Tailscale connectivity.”
Navigating Open Source Challenges
42:00 to 43:20
Learn about the changing landscape of open source collaboration and team dynamics.
“proud of, not the way that Claude told us to do it the first time around.”
Team Structure and Product Focus
43:20 to 44:44
Explore how a small team manages multiple products and maintains autonomy.
“How does your wind blow when it comes to that?”
Balancing Specialization and Cohesion
44:44 to 46:30
Discover the importance of maintaining a cohesive team despite specialization.
“I mean, that is the worry, the way we do our team.”
Keeping Up with Rapid Changes
46:30 to 48:22
Understand the pressures of staying current in a fast-paced development environment.
“Yeah, I'd be a little worried about the two, especially when things move so fastly.”
Work-Life Balance and Mental Load
48:22 to 50:29
Examine how constant problem-solving affects personal life and mental health.
“best support them how can i support the developers building on our on our on our platform yeah um Yeah, I think we're all stressed about falling behind.”
Evolving Work Processes
50:29 to 52:34
Learn about changes in work habits and coding practices over time.
“I, since before agents, like I started my career writing code by hand, as most people listening, this probably did.”
Managing Multiple Projects Efficiently
52:34 to 56:00
Discover strategies for managing multiple projects without losing focus.
“domain expertise and what i was building i could just like smash it out whereas now because i sit at things, I feel like the coding agents sit in between me and the code now.”
The Challenges of One-on-One Planning
56:00 to 56:58
Discusses the difficulties of managing AI-generated plans in one-on-one settings.
“if I'm doing one-on-one, I actually often find myself sabotaging the model because I'm thinking faster than it is writing.”
Trusting the Model: Iterative Planning Process
56:58 to 59:08
Explores the iterative process of planning and reviewing with AI models.
“And like, do I need to iterate on this plan?”
Prompting Techniques in AI
59:08 to 1:01:45
Covers various prompting techniques and their effectiveness in AI interactions.
“and say well we're missing this here and that's not right there what do you think the next thing is that I ask it after it presents all these challenges from high to low what do you think I So fix the plan.”
Voice to Text: A New Approach
1:01:45 to 1:03:34
Discusses the benefits and challenges of using voice-to-text technology.
“Because it feels so not smart on my part.”
Memory and AI: New Perspectives
1:03:34 to 1:06:08
Introduces the concept of memory in AI and its implications for development.
“By the time I finally got a prompt that I was happy with, I'd actually sent it six times and then canceled it and brought it back.”
The Intersection of AI and Healthcare
1:06:08 to 1:10:04
Explores the lack of AI fluency among professionals like doctors and its impact.
“I mean, I do want to talk about memory and what you're working on there, because I'm curious about how, I've never played with the memory side of things at all, and so I'm super curious.”
Understanding AI in Medicine
1:10:04 to 1:10:51
Explore how AI tools are being adopted by professionals like doctors.
“She's not designed to be a business person.”
Delaying AI: A Concern for Professionals
1:10:56 to 1:11:51
Discuss the reluctance of professionals to embrace AI technology.
“We've got folks that are super intelligent, like doctors, that are not really fluent in AI.”
Transformative Tools: Granola
1:11:52 to 1:12:46
Learn how Granola is revolutionizing work for professionals like doctors.
“Unless you have anything to say about that.”
AI’s Impact on Professional Workflows
1:12:47 to 1:14:00
Examine how AI can streamline workflows for various professions.
“My wife introduced somebody to who's also a doctor and she sits down with folks and she would spend three hours of literally every evening cramming all of her notes.”
Personal Use of Granola for Improved Writing
1:14:01 to 1:15:06
Discuss personal experiences using Granola for writing improvements.
“But I really struggled with putting my words to paper in a way that like flows and makes sense and is cohesive and has a start, a middle and an end and all of the good stuff that you need for writing.”
Memory and AI Agents: Concepts and Challenges
1:18:56 to 1:24:00
Delve into the challenges of implementing memory in AI agents.
“So memory, it's like such a loaded term.”
Building Flexible Agents with Cloudflare
1:24:00 to 1:26:00
Explore the challenges and considerations in developing flexible agents using Cloudflare products.
“So it's multifaceted and I'm still trying to work out in my head what I want to focus on because I don't think I can get all of these things right in the first time.”
Innovations in Data Handling with DuckDB and Parquet
1:26:00 to 1:33:20
Learn about the advantages of using DuckDB and Parquet for efficient data handling and querying.
“I've become a real big fan of EPUB books.”
Home Lab and Automation Insights
1:33:20 to 1:36:20
Discuss the benefits of home lab setups and automation tools for personal projects.
“I have a great friend and he's one of my colleagues now, actually, since I joined Clubfler.”
Exploring Local Models and AI Usage
1:36:20 to 1:38:01
Examine the current landscape of local AI models and the speaker's approach to using them.
“because all of my interests, like nothing has to be private to that point.”
Exploring Open Source Models
1:38:01 to 1:38:56
Discussing the excitement of experimenting with new open source models in AI.
“Every three weeks, a new version came out of the model.”
Innovating DNS Solutions
1:38:57 to 1:44:45
Introducing a custom DNS server built in Rust and its AI-driven enhancements.
“and you might enjoy this as a fellow Homelabber, up and coming maybe.”
Leveraging Code Mode
1:44:46 to 1:46:58
Advice on using code mode effectively in day-to-day development.
“Happy to rant about my DNS hole, which is super cool.”
Enhancing CLI for Agents
1:46:59 to 1:52:01
Discussing improvements to command-line interfaces making them more agent-friendly.
“then have a look at some of those options.”
The Power of Agent-Aware CLIs
1:52:01 to 1:52:48
Learn about the advantages of agent-aware command-line interfaces in software development.
“But in most cases, I'm just like telling the agent to do that stuff for me because why would I do that anymore when I don't have to?”
Show Notes & Community Involvement
1:52:49 to 1:53:19
Discover how listeners can access show notes and contribute to enhancing the content.
“We went deeper than I thought on some cool stuff, though, but I enjoyed it though.”
Transcript
Automatic transcript. May contain errors.0:01What's up friends, Adam here. This is the changelog. I got an awesome show for you. Matt Carey from Cloudflare working on agents and MCP. Lots of fun things obviously happening at Cloudflare. Tons of releases, tons of momentum. And even Matt will tell you here in this podcast, it's moving fast and he's trying to keep up. If you've been curious about Cloudflare's platform, how they enabled all their APIs via their MCP without destroying your token count. Well, this shows for you a massive thank you to our friends and our partners at fly.io. That is the home of changelog.com. Learn more at fly.io.
0:43OK, let's do this.
0:52Well, friends, this episode is brought to you by our friends at Coder.com, secure environments where developers and agents work in parallel. And I'm joined by Nicky Pike, field CTO for Coder. Nicky, what is a field CTO? So I get that question a lot and it's, you know, half the people understand it, half the people don't. So a field CTO, I describe it very simply as we're DevRel for the C-suite. So we provide a bridge between the customer voice, between the C-suite and the managers and the leadership teams of our customers back into our product. And then we go through and we help enable our teams to have the same message to make sure that the message is correct and that we're building on something that people actually want, not just something that we think they want.
1:32OK, so we're taking the laptop away from the developer. Not really, though. We're putting them in a cloud development environment, a secure environment where they can work with their agents in parallel. These are blessed environments. What's wrong with the laptop? The laptop is the trap here. And not only because the fact that it could be stolen, you could lose it, it breaks, and you're out of work while you're waiting for a new one. But there's also just the consistency that you got there. We all know developers. Developers are going to be looking for some of the latest and greatest. And if you're not really controlling how they get out there, that's where you get this.
2:01It works on my machine. It doesn't work in production. It doesn't work anywhere else because you don't have that consistency. You don't have that ability to really standardize what that environment looks like. And this is a problem not only for new people coming in, you know, the onboarding statement is average, I think, is like four to five weeks for a new employee to really get their local laptops set up and ready to start doing their first time of code. And, you know, the time to first commit is a metric that almost everybody knows. And the reason they can't do that is because there's a lot of tribal knowledge out there.
2:28They got to go talk to other developers. What are we using? Where do we get our dependencies? Are we getting them from public? Are we getting them from private repositories? But there's also the security and the supply chain aspect of this. When you have local machines out there, look at like the Shai Halud, you know, that virus that went out not long ago. This was a compromise of the NPM public repositories. They went and downloaded things. NPM did what it did. Next thing you know, you're compromised. But when you use something like what we're doing with cloud development environments, then you can mandate and you can put restrictions on there to say, hey, you can only go get your packages from our private repo.
3:03Those packages are expected to have been thoroughly vetted. We know that they're clean. Now, does this stop everything like Shai Halud? No. If that compromised package gets into your private repo, you can still have that. But it really reduces the surface area of the attack. And it also reduces the blast area of the compromise should it happen. Because if your laptop gets compromised and you have to kill the laptop for whatever reason, that's weeks out of work while you're either fixing that or you're getting a new laptop in. the cloud development environments allows you to kill that start back up fresh and you're back and running in five minutes.
3:35You don't have to wait all that time. Well, friends, the first step is to go to coder.com, install Coder, self-hosted environments for your teams to enjoy, to standardize around, and it's open source. So you can try it out today. Once again, coder.com.
4:11well matt carey good to see you on the pod thanks for taking my invite i saw code mode out there and i was like you know what let's talk about code mode so what do you think yeah well thanks for having me you podcast often yeah so i actually have one with a friend of mine but we don't do it super often a couple of couple of times a month a couple times a month what do you talk about what's the show called? It's called You've Been a Bad Agent and we just chat absolute rubbish about agents. That sounds like fun. How long is the show? Well, it starts and we just start rolling and I don't speak to him that much because he's in San Francisco.
4:46So I'm in Europe and so we just use it. We started having chats every couple of weeks just because I like to catch up and then we're like, we should record this. And that's where we started recording as a podcast. So it's literally just a chat between the two of us. Sometimes it's 20 minutes. Sometimes it's an hour and a half. We don't really edit it. And it just gets dumped online. So I'm probably going to get sued one day for something I say on that. Don't do that, man. That'll get you sued. We don't want to get sued. You know, one thing I'm really curious about, really, and why I want to talk to you, because, you know, obviously Codemote is cool and there's a misconfiguration, I would probably say, and you could probably agree with that, with how folks are thinking about MCP.
5:29And I think I've been in that camp too. I think we're all sort of just navigating this new world and trying to figure out how these tools work. And there's a race, obviously. Cloudflare is involved in that race. You've got blood out there in the water on X between you and Vercel. I mean, you've got stuff happening. And I just think about the sheer size and weight of Cloudflare. Maybe you can or cannot speak to just how you personally feel about these outages. Maybe some you can help, some you can't help. But I just think about the state of AI and how it's being deployed, accepted and deployed.
6:09And so the acceptance is one thing, but the deployment of it is another in a large organization like yours. And you're in charge of agents. You're in charge of MCP. You can share with the audience what you really do there. That's kind of what I want to cover is that bigger landscape of the deployment, acceptance of AI, the misconfiguration of MCP, and that kind of stuff. What do you think? Yeah, definitely. Let's go. I can chat a little bit about me, for instance. I work on Agents, specifically the Agents SDK at Cloudflare. So I work on the open source stuff. You'll have seen a bunch of my colleagues on X, Twitter, if you're on Twitter.
6:49I also work on some of the open source stuff for MCP and how we can support MCP at Cloudflare, Model Context Protocol. It's kind of why I joined Cloudflare. I was working a bunch on MCP and I thought they had a really good avenue there to build the best agents with durable objects and give them the best tools via MCP. And I was like, that looks super cool. So it's kind of why I really wanted to join this team. and yeah I've been been here since October we released code mode in the summer of last year to like do programmatic tool execution basically you write code over your tools rather than calling tools and then Anthropic followed up with a bunch of cool stuff after that and then just a few weeks ago well a week or so ago we released a server side code mode so running code mode inside an MCP server.
7:42So the model doesn't need to, or the agent doesn't need to call tools. The agent can just write code that acts upon the tools that we have on our, like in our backend. And then all of that code is executed super safely and securely on dynamic workers on our server side. And so your agent that calls the tool can just write code and yeah, it all gets executed on the server. It meant that we could put our whole Cloudflare API, all 2 ,500 odd endpoints. There was a number, but the number keeps changing. So every time I remember the number, it's out of date. But yeah, so around 2 ,500 endpoints, we could put all of that behind one MCP server that actually works, that fills 1 ,000 tokens of context.
8:29And there was a lot going around about, like, oh, fixing MCP. It's like, I don't think we're fixing MCP. We're just using MCP to the best of its capabilities. And it was a really well-designed protocol, I believe. And I think it continues to be well-iterated on. And I think it was maybe not used as well as it could have been initially. Yeah. Could you explain MCP a little bit for us? I mean, I know I've got a good dip in my toes in the water of MCP, of course. But for the uninitiated or less initiated, what exactly is the model context protocol MCP? What exactly is that? And maybe what are the myths about it that are incorrect?
9:12And what are the things you like most about it? Yeah, so it came out in November of 2024 by two guys from Anthropic, David and Justin. And the whole idea was how can we let Claude, in this case, Claude Desktop actually, how can we let Claude Desktop do things on my computer? How can I let it access my Apple Notes? How can I let it access maybe my web browser, maybe Figma, maybe Blender, maybe like whatever was on my computer? How can I let it do that? How can I let it read my code directly? That would be pretty cool. Yeah. Now we have code code and all of that stuff, but it wasn't around yet. So they came up with a protocol that consisted of tools, prompts, and resources.
10:01And tools are the things everyone talks about. Tools are like the functions, like function calling. The prompts are instructions that the server can hold that the client might want to use at some point and can request. I think of them as instructions. They're kind of like directions. They're almost like skills. There is a debate about whether skills are a prompt or a resource at the moment. So we can get into that. But the resources are like documents that might be held on the server that the client might want to use. and like tools have far and away the most amount of usage but initially it was like how can we define some tools that can be used by an agent that i don't own and vice versa how can an agent give access to tools that it doesn't own when it's being built and for that you need some sort of like standardized protocol and when these guys made it it was just local first communicated via standard IO.
11:02And then like once it got some usage and they introduced a remote protocol and then that remote protocol has changed a couple of times. Now, like pretty much every big SaaS company publishes an MCP server. I saw Datadog publish theirs last week, which is pretty cool. And yeah, I think I'm pretty bullish on MCP's like ability to be the protocol that agents use to access services in the future. Help me understand this breakthrough that you made with just writing TypeScript versus all the context-filling, tool calling, I guess you'd call it, right? That a lot of folks are kind of getting, I guess it's kind of wrong, but it's kind of how it's designed, but maybe you're applying the application incorrectly.
11:51Talk about the way that you've remodeled it to write TypeScript versus tool calling and fill in the context window. Yeah, so when LLMs first came out, they just produced text, right? And then at some point, I can't quite remember exactly what it was, but I'm pretty sure it was after the big chat GPT moment, like function calling became popular. And I remember the first model that could do function calling really well was, I think, GPT-4. And GPT-4, you could ask it, like, what's the weather in London? And it would reply, call weather function, param location equals London. and you could take that, that structured piece of information and you could plug that into some JavaScript, some Python, some code and you could call a third-party API, a weather API with London as the argument, the city argument, and you could get a result.
12:45And that result, whatever it was, you would pass back into the model as the tool result or as the function core result and then the model would continue generating and make it all nice and pretty for you. And that was like how LLMs performed actions in the outside world. And we, I guess like rightly or wrongly assumed that each individual function would be like a piece, like would be like a hand that the model could use in the outside world. And then they were renamed to tools after a while. And that makes more sense. Like each function was a tool that the model could use in the outside world.
13:22but there is like a problem where as you try to get these agents to do more and more things you add more and more tools and then at some point you start filling the initial context window of the model so for instance like the github mcp server is always one that's used and they've done loads of work on it so like i don't not not throwing any shade but initially when it came out it was like um 15 000 tokens or something and now i think it's a little bit less and they they do some stuff to dynamically add tools or not but like if you're filling the context window with sort of 20 000 tokens initially before you've even given the model your task like the the models of yesteryear like dbt4 they had much smaller context windows and so you were filling them very quickly And even now, the foundational models, the best ones, even though they have maybe have a million tokens context window or 200k to a million for the normal ones, they do start losing power around the 50k mark, like all of them.
14:26And this is quite well documented. And so you really don't want to be filling the context window too much. And you see in code code, you'll have like compaction step. That's triggered because you've filled the context window. So say you've added like 20 tools or whatever. You've like got a really chunky context window. But now that's only 20 things the model can do or the agent can do. Imagine you want to have like proper personal AI or something that can actually try and automate your job. Like imagine how many individual functions you do in your job. It's way more than 20. It's probably more than 100.
15:04It's probably near a few hundred. so like you can't really automate anything unless you have the ability to add all of those functions into a model so uh in the in the summer uh Kenton and Sunil who uh my colleagues of mine worked out that you could or they were a really great blog post um the code mode blog post the original one amazing name from Kenton like really really stunning uh it's like I think it caught on a It's like you could just use code mode. And the idea was falling back to an idea that had been around for a while. Like I'm pretty sure Hugging Face did a research paper on Code Act a while ago.
15:42But the idea was that models should just write code. Like AI has been trained on so much code, we should just write code. And if we write code, the code can interact with the functions that we want to use. And so code mode generated a TypeScript API or TypeScript SDK, really, for the underlying functions, the molecule call. And then the model just wrote code to compose those SDK calls. And I guess the good innovation for this, the reason why it's, I think, slowly taking off and we're going to see much more adoption this year is the advent of something called a dynamic worker loader, which is a Cloudflare primitive, but other people have started building similar things, if not the same.
16:26um but like this is a primitive that allows you to execute a sandbox worker as a string like so from a string from some code that's a string you can just be like eval this new function this except it runs on a separate host in a fully sandboxed environment in a va isolate and what does this mean like traditionally people got very very scared when you say give me code and i'm going execute it on my machine because there are so many different ways that you can mess someone up by doing that you can like out of memory them you can access m variables you can do loads of stuff that's really hard to protect against and this sandbox is very particular sandbox that's not a full vm it's just it's just a va isolate allows you to spin up like billions of these little scripts almost instantaneously if you wanted to like at cloudflare scale at global scale and just run these pieces of code, like very, very safely and securely.
17:25You can even like restrict the outgoing fetch. It's called a global outbound. And you can say, I only want the outgoing fetch to be able to access example.com or mysass.com or I don't want to access anything. Just run code that's entirely constrained in this host. And this is really cool because it allows this like code act idea, this code mode idea to like really take off because the model can just write code. It doesn't matter if it's prompt injected or if it's like trying to be, I don't know, trying to be adversarial. Like it's running, the code is running in a super safe environment and that's all good.
18:02And that meant that the original code mode blog post had this code execution happening in the agent. And that was like the smart thing to do, right? The proper thing to do is like you make the agent do code execution and then you become like a code agent. and that's what happens. But this relies on every, or on the agent that wants to do it, actually shipping code execution in the agent. And like, it turns out that that is also quite tough. And although for the last like six months, we've been shouting like, if you'll have problems with context window, just get the model to write code. Like not as many people did it as we thought would do it.
18:43And so I just took what we'd done previously and moved it into our MCP server and basically just said, what does this enable now? We have that massively reduced context window allowance and we can use it to enable our MCP server to access the whole of the Cloudflare API. Do all of the possible endpoints that you wanted to call in the Cloudflare API, you can now be accessible via MCP. And before, you could come to this, I guess what I'm trying to say is you could come to the same conclusion in the summer by putting code mode inside the coding agent. You could have connected as many MCP servers as you wanted, but put code mode inside the coding agent.
19:27But fundamentally, that's quite hard to do. So how can we show the value of it? Well, we put it inside the MCP server and that's what we did. And it built like basically a one of a kind MCP server. where no one had really seen this type of thing before, where I can have one coding agent that uses a thousand tokens of context access our whole Cloudflare platform. I think that was pretty cool. Sorry, that was a bit rambling, but that was the point. It was a good deep dive. I got some questions about that. So what is fundamentally different about the MCP server that's uniquely different than every other to enable this many APIs and that reduction in the context window?
20:08Yeah, so most MCP servers would map like one API, one endpoint to one tool. And they'd be like, oh, we want to get all issues or post an issue or like delete an issue. Yeah, like one-to-one mapping. And sometimes you might do like a one-to-many mapping a little bit. If you had like a particular workflow that was very common on your website, you might like create a tool for that workflow. There was a very common workflow that customers do. But you're kind of restricted to around 10 to 20 tools, but maybe up to 25. I know Cursor has a 40 tool limit max for supplied MCP servers. So theoretically you could fill that, but like it's getting pretty large.
20:59It's getting pretty large. And you're still not covering anywhere near. You're still cherry picking like for big platforms like the Cloudflare platform, like GitHub, for instance. I use GitHub and Cloudflare, but like you can imagine any large API, you're not covering anywhere near the full amount. full amount of the breadth of the API. But when you put code mode in front of it, you're now just asking the model to write code over that API. And so the Cloudflare MCP server exports just two tools, a search tool and an execute tool. So this is combining a similar idea to tool search, which is present inside Cloud Code, as also present inside Cursor, I think, where they will search for the right tool on demand, and then they'll load the right tool, depending on the user intent.
21:49So we do that, but we do that on the server side. So there's search and execute. But the critical thing that no one else does is for search, we let the model write code to search over the Cloudflare OpenAPI spec. Code. There's no search function or anything. The model just goes like spec.paths and then filters like super naively, and it works. and then for execute we say model here or agent here is a cloudflare um fetch client call cloudflare.request to make a request the cloudflare api and and we just let the model go for it uh and so you end up with like something that's super flexible if you say something like build me a worker that's uh hosts a next js website that can do this this this this and this you would hope that the agent can just write the code maybe using vNext a new fun thing that we just built to write Next.js and deploy it on Cloudflare they write the code and then the model can look search through workers scripts deployment APIs it finds it pretty quickly normally and then it can just call the endpoints workers script deployment and it'll just deploy it to the cloud super super fun And you can have these like insane demos where none of the code ever gets saved on your machine.
23:17It only exists in the context of your chat with the coding agent and in the cloud. And that was the first demo I ever did of this. And I think it went pretty well. And it went so well, they were like, we have to ship this. Really? Wow. What was that like, that demo? Can you take us into that day? And, you know, what was, who was there? Where was it? Was it online only? Yeah, yeah. So we do Friday demos. The Friday demos are pretty legendary. Cloudflare, like we, so I'm part of a team called like developer platform. So I work on workers. I work on, yeah, on anything built on top of the developer platform.
23:59And yeah, we have like our own demo sections on Fridays. And yeah, like the whole, most of this org like turns up and we probably have five to seven, awesome demos. Some of the coolest stuff you've seen come out of Cloudflare was probably demoed at one of these platform sessions. Often only a short time before it was released to the public, which I think is pretty cool. It was good crack. You get a couple of minutes. Look at this thing I built and then 50 % of the time it breaks, but then 50 % of the time it's awesome. A lot of these things happen, I would imagine like maybe late nights, maybe just late in your ability to keep thinking, where you're sort of knee-deep in innovation.
24:49You're iterating over something. Can you take us into how you stumbled upon this, write some TypeScript against an SDK? How did you stumble into this? Was it a thought? Were you in the shower? Were you on a run? How did this iteration come to be? well i've been putting it off for ages uh we knew it was something that we wanted to do uh so i don't i don't know if the like the idea had been around for ages like sunil my uh my tech lead he i don't know if it's infamously now but he built the agents SDK which is the thing i work on the first iteration he built over a weekend and then he went to his uh product leads and was like guys can we ship this and they were like maybe let's like hold off a week and then he shipped it a week later and the first version was literally just like um durable object export um export agent as durable object and it was like clean it was like one-liner um and that was the first version of the agent's SDK but that was like his that was like a really cool innovation there um to use agents as durable objects and I guess for this uh and that was a weekend um I guess for this like we knew it was something that we wanted to do something that had been badgering me for ages to like get on and do it and work it out and work out the tool thing why can't we just have loads of tools like surely this is possible um and yeah i put it off for a couple of weeks and then i think it was a tuesday just sat down and was like all right i'm gonna do it had a had a chat with claude i reckon about how we might want to do it worked out very quickly that it would be i was been working a lot on mcps that worked out very quickly that the best use case for this would be how can we enable every mcp server to host an unlimited number of tools and that this seemed quite possible um with the search and execute paradigm and that yeah like the model i previously i don't know i've been working on this for quite a long time like in my in previous jobs and stuff i'd always had problems with search functions so whenever you give the model a search function now you need an eval like 100 you need an eval because you need to work out that if you change the parameters of your search like does the search get better or worse for your task and i was just pretty like i was very drawn to the idea of code mode that i would never have to do like an eval in that way um again because the model can write code and as the models get better at writing code my my thing would get better like i keep everything in distribution and so i was really really really drawn to that idea like i knew these models they're just going to get better at writing code let's lean to their strengths and just yeah let the model get yeah let the keep the model in the distribution it was trained in rather than like some hacky search function it's never going to be in that distribution so i was like right let's go and that sort of bore it all together like the two pieces do you have an unlimited context window like do you have my context window when you're working with claude give me an example of oh do i have something special that's kind of a tongue-in-cheek request or ask i guess but uh what i'm trying to get to is is less that real response i'm happy to take it but more so be more clear with what you're like this back and forth are you dropping files are you sort of you know sort of micro-contexting where you sort of pull some out of the context you take it to a file like what does the actual interaction look like to to go back and forth with claude to innovate for the future like this Yeah, so my workflow with Claw has actually changed quite a lot over the past.
28:26I'm sure everyone's has. It's changed a huge amount over the past. It's constantly changing, yeah. Yeah, I would say. A year and a half ago, I was big into Cursor. I really liked it. And I was using it mostly for the tab model, I'd say, about a year and a half ago. And then January last year, I was like, no, the agent model is the future. I need to like work out how I can just prompt like how can I just prompt things into existence like what guardrails do I have to put on things how can I make the feedback loop how can I have tests that have good patterns and all this sort of good stuff to like build basically my code bases for agents and this was about January last year and then when Claude Code came out I was like right now this is definitely the future now I'm just going to have an IDE open when I want to actually visually look at code When I want to review, I'll open an ID.
29:17Otherwise, just straight in the terminal. Let's just chat. Let's build something. And I tend to run everything on dangerously skipped permissions, like 100 % of the time. I have some, like, sandboxing on my laptop. It's a custom thing that I'm not, like, don't super want to talk about. Okay. But it's quite good. Yeah. Yeah. I'm not going to talk about QC. I have. Okay. So I have my own version of Git that runs on my, that runs on my machine and it's just aliased. And it just stops like the model doing like stuff that I don't want it to do. So it stops it force, force pushing to branches, overwriting stuff, even because I have admin permissions on a lot of repos.
30:04So to me, this was like the base level of thing that I didn't want to happen. And I didn't want anyone to be, I didn't want any agent running on my laptop to be able to overwrite a remote repo. That's like base level. My own laptop OS, like I don't mind breaking stuff. Like I don't mind breaking anything locally, but I don't want to break anything externally. And so I have a few like aliases like that where I've just like completely overwritten the Git internally. And my like Git wrapper is called Zaggy. It's public on GitHub. I built it in Zig actually. super fun to build using libgit2 um and it just like has a bunch of these protections out the box so i really don't mind running stuff in dangerously skip permissions i tend to do it all the time i chat with the chat with them all about things i want to do but i tend to come to sit down at my laptop with a preconceived notion of what i want i think a lot of the time when i sit down at my laptop without any idea of what i want i almost feel like i'm scrolling on instagram or Twitter or something.
31:08Like even chatting with the model, it's just feeding a dopamine rush that is not real. Like when I know what I want, I feel like I can evaluate stuff really well. So it took that Tuesday for me to sit down and be like, right, I know what I want. Let's just do it. And in that case, I think I can be hugely effective. I think most people can. Just speaking straight English to the model and nothing fancy. Yeah, just go for it. So a lot like a chat, like a real chat. You're not doing some act as Cloudflare master worker slash developer. You know, like these sort of like hacky things. That used to be a thing.
31:45That used to be a thing, yeah. I'm not doing any of that. I'm not acting as like every once in a while I might put it in a role, but I mean just it's like one out of a hundred, if that. It's usually just here's the problem. Here's where I'm trying to go. Here's where I'm at. Here's what's in between us. Let's just riff kind of thing. You know, what's here, what's there. And I sort of trust the model in those senses. So I'm not trying to wield it and force it into a mode. I kind of just kind of give it the trust it needs to do its great job. anytime you're fighting i feel like personally for me anytime i'm not getting good results or i'm not or i'm fighting it i'm trying to push it into an area where it's just not so much not good at it we're just in uncharted territories or just something like that and i feel like the more i just talk like i would a normal engineer next to me or a colleague that's where i get my best results yeah you want to always keep it in distribution like um if you're doing something too wacky then you're probably going to have less you're probably going to have erratic results because like if the agent never saw or the model never saw anything like that in pre-training or in its post-training like RL stuff was that if you're like working on a common programming language's code base you're speaking to it like maybe like you'd see in a GitHub issue with language that is like intelligible then I think you're pretty good.
Read the full transcript
33:12I don't think I'd do anything special with Claude. I would say that I tend to only use Opus now since Opus 4.6 came out or 4.5 or whichever one it was before Christmas. I think that was a big step change in being able to do things more autonomously. And so what I tend to do now actually, because I work on a lot of different repos, I just open Cloud Code or actually Open Code we use a lot, Cloudflare. I just open that in my code folder on my laptop. I basically always just open it in my code folder and then I direct to the repo that we're working on. Sometimes I'm like, right, make a new work tree, but everything lives in the top level code folder because I work a lot on libraries and on products and the products use the libraries.
34:03And so it's nice if it's all, I'm working constantly at that top level and then I can move between stuff more fluidly. that's actually one of the reasons why i don't use cursor as much or like any of the ides as much now because just like having the ability to like open an agent in that top level folder super nice yeah i guess if your context across projects or libraries are connected it makes a lot of sense but if you have disparate projects where like that is totally its own thing or this is its own thing you kind of want to have a directory of a silo is kind of how your code directory is it's like this is a silo of all Cloudflare work.
34:41So therefore it's cool to just open that one directory. Is that what you're saying? Yeah, like I would even have, like even like personal work. I don't think it really, yeah. I don't ever see the model like going into other directories that it's not meant to. And I do actually watch quite a lot. Like I don't, maybe, maybe this is a mega security thing, but I think it's pretty good. I have a working directory, like I have a directory of working code that I'm currently working on on my machine. And I just open it in that. And I tend to reuse patterns a lot. So I mostly work on open source work, right?
35:24So for instance, how the development worked with the Cloudflare MCP server was I built it as a POC. I published it on my personal GitHub. It was published. and then once I got enough buy-in, once people thought it was good, once the quality was there, then we moved it over to a Cloudflare legit one and we did a big release post. And I do that with quite a lot of stuff. Like there is procedure and things to making a new repo on the Cloudflare org and like it has to meet a certain quality bar. So for just POCs and testing and stuff, I still want version control. So I just use my personal GitHub and it's fine.
36:02Well, I use my, yeah, I just use my own org. Well, that's clearly a lie to do that without having any real issues. I know that – I mean it's so sensitive whenever you're – I mean whenever you're in your position as a brand, as a company, you do have to have locks on the doors. You know what I mean? And that's not a lock on the door, but that's cool that you have that kind of autonomy to, one, explore, and two, not get any backlash for publishing to your personal GitHub. where it's like you could be seen as like, I'm trying to take, and you're not obviously, if I'm trying to take the Cloudflare Thunder, no effect, I'm trying to innovate, and I'm just trying not to bother our main org and our brand integrity with my little toy here until it becomes not a toy.
36:48And like we have a private internal version control as well that we also use, but for things where I want to share it and I want to even see other people if they're interested in it and things like, it just makes sense. You want to get it out there. I think I'm in a very special situation in the company where I work predominantly on open source. And so there is a more freedom allowed there because anything that I share is public by its very nature and is going to become public. If I'm working on the agent's SDK, we have to be really, it's hard to have even a proper release because everyone sees what you're doing as you're doing it.
37:28And so like to even like do a bit of experimentation is like quite, it's quite tough not to get found out. And like you still want to be able to do a proper release even as an open source library.
37:44Well, friends, you know, I'm a big fan of Tailscale. And you know what? I could not do anything. I'm serious. Anything without my Tailnet. I'm here with my good friend, Alex Kretschmar from Tailscale. Alex how do you describe tail scale versus a vpn how do you describe tail scale to someone who's not in the know well the biggest difference between tail scale and a traditional vpn is how the traffic flows when you look at a traditional vpn the traffic flows through a central hub and then out to your client devices on the back end with tail scale every device makes a connection directly to every other device that means effectively you're cutting out the middleman and you get much better performance as a consequence And so that mesh network that you've built, you've got to have a way to control how the data flows between different devices because you don't have that central choke point anymore.
38:34We have a thing called access policies, which allow you to granularly define using ACLs and grant policies, which nodes are allowed to talk specifically to which other nodes on which protocols, on which ports and which users are allowed to even connect to different things all over the tailscale encrypted tunnels, which underneath use the WireGuard technology. Yeah, it's not my LAN, it's my TAN, my Tailscale Area Network. But your word for it is Tailnet, right? The Tailnet is the word that we invented to call the logical grouping of devices that form your Tailscale network. Much like you might have a LAN of devices or something like that at home, effectively the Tailnet, we call it something different because those devices can transcend physical locations.
39:17so you can have a server in the cloud talking to your phone on the bus talking to your server in i don't know the basement of your mom's house across the other side of the ocean and that tail net is a flat network that only you can connect to and access so that's why we call it a different name from anything else is because it's it's location independent and you can connect to it anywhere well friends check out tailscale at tailscale.com totally free for your home lab and of course paid for your teams pro and enterprise. But literally I could not do anything I'm doing in my home lab, in my dev lab without Tailscale connectivity.
39:55I'm out and about, I'm here, I'm there, I'm everywhere. And I've got to access my home lab, my dev lab resources and Tailscale is how I do it personally. And you should too. Once again, check it out, tailscale.com.
40:16The reason why I ask you about how you actually work with Claude is because I think that's the curiosity of everybody. A lot of us are to some degree working in silos, even if we're working together, because there's even speculation of how large of a team can you actually work on in this new era because of how much you can get done in one slip versus as a team where you'd have to collaborate a lot more on a major feature. I'm not sure you may have some inbound conversations in Slack or that kind of thing, maybe a pull request review or something like that, or in your case, a POC as an actual repository.
40:54But I feel like in this world, what I'm hearing a lot of is it's actually kind of hard to work at this level of ability and collaborate at the same time. I think it depends how you like to work. I'm not saying my way is the way and that I'm like six months in front of all you guys. you should all work like me um because i don't think so uh i thought dax from sst anomaly open code shared a really interesting post on twitter last couple of days where he was talking about how everyone sounds like they've got it all put together um but really he doesn't think so and he knows that they don't have it all worked out like they're still working stuff out they think they're faster with coding agents than without, but not entirely sure.
41:43And it was really like a push to be like, can we just leave everything better than we found it? Like that whole thing of like coding agents. Yeah, sure. They let you work very quickly in the short term, but let's go through our code base. Let's build everything the right way, the way that we're proud of, not the way that Claude told us to do it the first time around. I thought it was really pretty uh i think everyone should remember that yeah sure we're there to do a job but we're also there to ensure that when the next person comes to have a look at the job we did they can actually have a clue what's going on and like it does work properly and it is tested and there's a lot of slop being thrown around a lot of slop prs i think on the agents stk we actually closed like prs from external collaborators for the time being.
42:33Not to say we won't open them again, but it was just getting too much. The way that open source works has to change. But us as a team is kind of interesting. We support essentially three products on our team. And our team was five until very recently. We support the agents SDK that I've talked quite a lot about, how we build agents on Cloudflare. We support MCP. So how people build MCP servers, MCP clients, how we build those on Cloudflare and also the Cloudflare supported MCV service. So the new one we just built. And also we support all the ones that Cloudflare published last year as well, the external ones.
43:11So we support those two avenues and we also support sandboxes. So the whole sandbox product in Cloudflare comes from my team as well. And so there was five of us working across these like very three distinct parts of the, of, of building agents. and so there was there's a lot of surface areas so i think we're as a team we're we're pretty well versed in like in like having our own domain and like building out what we think should be built out in our own domain it's very hard to get under each other's feet because there's so much uh so much space are you six now are you four i think we might be six now seven very soon and eight very soon after it's good it's going good how do you are you autonomous in terms of like which product you focus on at any given time i'm sure there's missions of course and there's directives but whenever you think about okay like you said that tuesday when you sat down and innovated in this way to to uh to give us these 2500 plus apis in a thousand tokens or less kind of thing um how do you sit down or how do you even think about your work when you when you when you're split across your products do you shiny objected or is it pre-directed or you totally autonomous?
44:30How does your wind blow when it comes to that? Our team is, it is kind of special in Cloudflare and it's changing a lot. So if we have this conversation in six months, it might be a very different situation. But our team is very new. So we, I think we were launched as a team under a year ago and I joined in October and people have been joining every couple of months basically for the last year so i focus on mcp i also do a bunch on the agents sdk to help support mcp to help support people building agents and i'm focusing on memory a lot at the moment and how we can build out a story for that and just support developers building on on cloudflare there other people on the team have different specialties we basically all contribute to the agents SDK and then like nourish on our team he focuses very much on sandboxes like sandboxes is his baby and he's built it from the ground up and now he's getting some support on sandboxes but we were always all contributing to agents SDK even if we were doing our other stuff because it all ties back in like like we need to have like this cohesive story and be one cohesive team and we're all when I build an agent in my spare time I use all of our products like all of the SDKs we produce I use so there is like I think the main worry for our team is how do we how do we not end up with like domain specialists too much and we have a nice tracker about who who's submitted PRs to different to different repos because I think there is a worry there that like I haven't committed to sandboxes I have no idea what's going on there like how can I answer a question when someone comes up to me at an event or something and talks about sandboxes or when I like I'm developing on it myself and I find a bug, like it'd be really nice if I could fix it.
46:19Like just very basic stuff like that. I mean, that is the worry, the way we do our team. But I think everyone is just so interested in building agents and all of these are critical parts of it that we float across it, each other quite, we float across it all, all the surface quite well. Yeah, I'd be a little worried about the two, especially when things move so fastly. I mean, like it didn't move this fast before. and I guess a year-ish ago, it was a little easier to have that disposition where you say, you know what, I'm focused on agents and MCP, but if I don't contribute to sandboxes quite that often, it's okay because it's not moving at the speed of agents, which was the case beforehand.
46:56But now it does. And so I would personally, if I were in your position or on that team, I would feel a little anxious. I'm not keeping up. And I don't know. I guess this is how I feel about most things, really. but especially if I had my particular sliver that I'm focused on totally, like agents and MCP. And I'm really curious what you're talking about with memory, what you're doing there. But I would have some anxiety about, my gosh, how do I even maintain any version of context around sandboxes when if I step away for a week or I don't pay attention to some of the side chatter, how far back do I go when it comes to progress?
47:37progress yeah i mean we've got to be we've got to be as forward-looking as possible i think for sure our team attracts a lot of dreamers i would say um yeah the guy that started our team so is like an absolute dreamer like he's thinking so far in advance like i i really respect like how how he can think like that and um yeah learn from that as much as possible the aim for us is to be ahead of the org what the org wants we should already have like ready for them to use um and i would say a year ago we were quite a long time ahead of the org and now everything is going faster like there are some people building insanely cool agents at cloud flare and yeah which like that's where the memory thing is coming out of it it's like how how can i best support them how can i support the developers building on our on our on our platform yeah um Yeah, I think we're all stressed about falling behind.
48:35That's why you'll find that a lot of Cloudflare people are permanently online, maybe a little bit too much. If you want to throw shade at us for anything, it won't be because we're not receptive to feedback online. Just curious on that note, and you can blur the line if you want to or not give the exact number, but how many hours do you think you work a day? And don't just say in front of the terminal because when you're making your coffee or you're on your back patio or you're walking your dog and you're thinking about work, that's still kind of work in a way. How much time do you truly separate from the problem set that you're dealing with or working on?
49:16And how much of that turns into Matt's life? I don't know. I think I'm thinking about this stuff all the time. Like 23-7? Yeah, something like that. I like my sleep, you know, like I like I like eight hours minimum. Well, I'll admit the moment I wake up, I'm going to sleep thinking about a problem. I'm waking up thinking about that problem. It's a sign of a good problem. Zero am I even throwing shade at you. And I think the reason why I ask this question is more of a reality check to our listening audience, because I know there's a lot of folks feeling like either they're not dipping their toe in and they're abrasive to the situation.
49:57and they're kind of late in a way, but still early, which is kind of funny to think about. Or they're just like you and I and others where they're like, I mean, the race is on. I just can't stop thinking about the things I want to change or do, and there isn't enough time in the day. Now, I'm not eking into my personal life where I can't live my life by any means, but I'm definitely thinking about the problems I'm trying to solve far more than I ever have before agents entered my life as a reality check. Yeah. I don't know if it changed. I, since before agents, like I started my career writing code by hand, as most people listening, this probably did.
50:40Um, I love how you said that. That's so awesome. Yeah. I mean, you've got to, I wrote it by hand, man. You've got to preface this. it's all organic well organic code you know organic written by me that's right the og code yeah yeah definitely worse than some of the code that called spits out definitely um i think i could i could get always get very engrossed by a problem like my girlfriend she gets so mad at me sometimes like i'm just like i get super sidetracked by stuff like incredibly attached to a problem and a solution um well more of the problem than the solution but so I don't think that has changed at all I think what has changed is how I work so I spend much more time dreaming about like a future world and like a future things that I'd like to build and or like thinking who might be best to build them and when code is cheap like you can build more stuff but you also still have a limited amount of time you can't build everything and like the hard things are the things that but the cool things are the hard things and those are the things that take time and there are like there aren't that many quick wins that you need to put in the hours every day to like work out what you want to do and how you want to do it and I guess now I'm spending less time coding like manually coding I mean I don't actually do that that much anymore I'm spending much more time thinking about like what what i'd like to build but i'm always thinking about the problems i guess i now much more scatterbrained so previously in the past i could sit down for eight hours and just code for eight hours and like that was great i was stayed in a terminal or i stayed in a ide i like never left it i like knew had enough knowledge about the domain expertise and what i was building i could just like smash it out whereas now because i sit at things, I feel like the coding agents sit in between me and the code now.
52:45So I am much more in the backseat, or at least like in the bird's eye view, over multiple different things, not normally just one, because I can, right? But it does, there is a compromise there that you do feel much more scatterbrained. You're like, here, you're there, you're like, you have to dive into this, you have to dive into this. And traditionally, not super good at that, I'm not gonna lie. hugely bad at multitasking for me um so like getting that compromise right i envy people who feel like they can productively prompt like six uh versions of cloud code or open code or whatever six coding agents at once i like i just don't see how that is humanly possible uh for me i reckon i I got three in me max because I think I can only do three problems in my head at once and still like have a meaningful output to each of them.
53:42And definitely over two by like, well, over three, 100%, but like my capability, my capability to like do something hard, massively reduces. I feel like for me, three to six, a couple of times a week is where I'll catch myself there. not like i intentionally go there yeah but i rather enjoy a one-to-one problem except for when i'm waiting for to to like i was gonna say i find it slow i find it slow i can't do it so i kind of have to do one thing but multiple things on that one thing i suppose is the way to describe it where i guess that's still kind of three but it kind of depends right How you – like it's the traditional it depends scenario there because when you're waiting, what are you doing?
54:33Like maybe even your own spaghetti and your brain is getting unraveled where you think you have the context. You're sort of planning things. So I kind of feel like my zone is like two to three because one is too slow, and it's not too slow for me. It's too slow because it's doing its thing, and it's doing dramatic stuff. it's doing a week's worth of things and that 30 minutes I'm waiting or whatever or that 3 minutes or 4 minutes I'm waiting so I find like I have to be in the 3 zone almost always but then even multiple projects that are uniquely different but similar I find that a couple times a week and if I do that more than a couple times a week I can get in that zone for hours 3-4 hours really where I'm working on like three or four different projects and like three or four things per project.
55:28That's wild. That's absolutely wild. It is kind of wild to do that kind of stuff. Really, it is. And I haven't sat back and said, how well are you doing? But what I can see is the Git commits. I can see the progress. I can see the improvements. And I see the real thing deployed and usable, not just this fake thing that maybe, you know, this agent psychosis kind of scenario where you're like, I think I'm making progress. You know what I mean? I see that. I actually like on that note, if I'm doing one-on-one, I actually often find myself sabotaging the model because I'm thinking faster than it is writing.
56:09And so I start writing stuff to correct the trajectory. And I think that's really bad. Yeah. I think it meant because, because during the planning phase, I got bored. um it meant that during the execution i'm just constantly fighting like it's trajectory and so and i do use plan quite a lot i also get plan and then get reviewed by another model i have a skill for that it's really really good um i think i nicked it from someone on twitter i honestly amazing would recommend but when i do two or three then building the plan is better because i can set off one to build a plan and then get and get reviews and that might take like 10 minutes and then I can set off another one to do it.
56:51And then the third one. And then by the time I'm like done three, I'm like, can take a breather. And I can be like, right, let's go into the first one and see what it's come up with. And like, do I need to iterate on this plan? And that's so much better than being like just one-on-one. Oh, wait five minutes. And now sabotage. Like, cause I want to, I just want to, I just wanted to implement now. I think giving it the time is nice. So if this cycle repeats itself for you, this is my cycle. I don't use plan mode a lot, but I do a different version of planning. So I wrote a GoCLI and this flow I created called AgentFlow.
57:33and it's a lot of I guess context dumping in a way but I'm making plans and those plans are called peps that is stolen from the python world where it's not a python improvement proposal it's a project improvement I guess it's a it's p what's the e stand for again enhancement that's right I'm like improvement enhancement project enhancement proposal versus Python enhancement proposal. And so what I find is, is I will either, you know, make a true spec based on RFC 2119, or sorry, 2119's protocol for like must, should, things like that. And those are for bigger things, you know, like the way an API should function or what kind of error codes we should respond with and things like that, like what the API surface is.
58:23So I'm speccing an API or different things, not literally every possible thing is getting a spec. But here's what I'm trying to get to is what I often do, and maybe this is how it works in plan with for you, is I just trust the model. And I say after they present the plan to me, I ask it to review that plan for lack of clarity and blind spots. Just that one prompt response back to it. Like nothing else. Not here's what I think is wrong with it. Like I told it what I wanted to do. I'm telling it where I'm at, what the gap is, and where we're trying to go. and so the problem is there and I'm trusting the model to kind of get us there and the plan is the iterative process and so once it presents this plan to me in my case as a pep I just say review that pep for lack of clarity and blind spots and it will go and it will review it and it comes back and say well we're missing this here and that's not right there what do you think the next thing is that I ask it after it presents all these challenges from high to low what do you think I So fix the plan.
59:25No, no, I don't. No? What did you tell? Kind of, yes, but no. I give it one more little nudge because I want to trust the model. I say, what are your suggestions for each? That's literally all I say. What are your suggestions for each? It goes and it iterates through each suggestion it gave back to me of all the problems. It's like, here's how I'd solve it. Here's how I'd solve it. And I'm like, what do you think I respond back with after that? What do you think my next prompt is? Fix the plan. Do it. literally the words do it okay so let's make the plan present the plan what clarity and blind spots are missing from this thing present it back to me a big old list what do you suggest for each it goes and does this thing presents a plan back to me do it that is literally what i do this is changing on repeat this is essentially essentially if you think about what you're doing in terms of like 20 23 prompting techniques it's like you're doing reflection by asking the model to look back at itself and see whether it's done anything silly.
1:00:25And then by asking for suggestions, you're doing chain of thought prompting because you're getting a new train of thought to go back on the original one. Yeah, so it's this reflection plus chain of thought. It's like, yeah, it's just really funny how all of these prompting techniques come back around. And what's even cooler, I think, is that the likelihood of those prompting techniques being reflected in the underlying training data is i think super high so for opus 4.6 yeah for opus 4.6 i um i find it often it like it'll do like wait at the end here are suggestions for each so like i do find it often does that that that step for you like it doesn't it doesn't need to be told and i kind of feel bad about asking it for more but all it did it presented a bunch and then it kind of gave me three so it may have given me a list of let's just say six to twelve issues in the plan right and it comes down with like three or it always gives me some version suggestion but that's not the real suggestions man i mean go back to the list you know what are your suggestions for each is a more you know four i loop through all the thing you know what i mean like that's all i'm really fascinated to do and i get such great results with that and then i kind of feel bad with my final prompt being like, do it.
1:01:45Yeah, do the thing. Because it feels so not smart on my part. Do it. Yeah, none of this is that smart on our part. So I think we have to accept that. Like, I think the smart thing is knowing when you sit down to the computer, what do you want to make? Yes, the intent. Where are we going? What should we do? Yeah, and I think the suggestion side of things, like I actually review those quite heavily on each plan iteration because I do want to make sure we're following a trajectory. Maybe that's my own nervousness around the model, but I do want to make sure we're following the right trajectory that I have in mind.
1:02:23Do you ever use voice to get longer prompts? Just recently started to do it. Actually, there's a cool thing called Handy. Handy.computer just mentioned that this week in ChangeLog News. it is an open source voice to text it's all done on your machine it's free and open source so I mean a lot of safety there in terms of what you're putting out there I've tried it a few times I like it it goes in any text box you give it but it's kind of hard to always default to that because some things are technical and you can't speak a command very well or syntax or a file path or things like that. So I find that I've just learned to type faster and more clear and it keeps, it keeps my brain in it more than I think it out loud.
1:03:16Cause if I talk out loud, I will talk a lot more to my podcast. Whereas if I type, I'm more terse and more clear. Whereas if I speak, I'm more ambiguous and thought provoking and meandering so to speak. You know, like I just would say the word, uh, and it's like, what are you talking about here? Whereas if I'm, I don't ever type the word, uh, as I'm trying to speak, because that's not what happens when you, when you write like that. That's an interesting avenue. I, I know the AMP team have some thoughts around this where they specifically made, um, enter just make a new line on an AMP originally, like in the sidebar version of AMP, uh, rather, and then command enter or control enter or shift enter or whatever it was actually executed the prompt.
1:04:02and the thought process behind that was like we want you we want to encourage users to make longer prompts to make larger expressions of intent to like fully scope the problem at hand and if we make enter a new line then they might have some inspiration to write more stuff i think with i finally got around to writing longer prompts and i'm very excited by uh cord code just added voice support where you can hold spacebar and have a speech to text model so I can like speak for five minutes or for 30 seconds or however long it is then I can dump I can do like I have a clipboard so I just like command v command v command v all of my context in below and then I think I end up with quite a nice quite a nice prompt by doing that I'm very excited by that flow and I trust Opus 4.6 way more to like execute for a longer period of time I think in the past using cursor and using cursor with I don't know what model it would have been at the time but like probably Sonnet 3 or Sonnet up to Sonnet 3.5 using cursor with those models I would like send something and then I'd be like oh no I meant this thing I need to add this more information and then I would cancel the original prompt and then compress it.
1:05:24By the time I finally got a prompt that I was happy with, I'd actually sent it six times and then canceled it and brought it back. All of that flow I thought was awful. I'm really consciously trying to make that problem more well-scoped. Yeah. So we got there by talking about the things you're working on, how you focus on agents, the open source stuff you're doing there, MCP and you mentioned that you're starting to think about memory yeah can you take me into what you mean by that what are your thoughts on that what's attracting you to that how far in are you do you feel in over your head what wisdom do you have do you have any wisdom at all where you at I have felt in over my head for the past oh my whole career I'd say good for you it was crazy crazy world we live in i think before when you were saying about feeling left behind i think so many people feel slightly left behind or a lot left behind with this version of like i think if you went and spoke to my friend a lot of my friends um in london i actually recently moved to portugal but a lot of my friends in london the software engineers uh that i knew that i lived with i went to university with um the amount of them still not using ai at all it's like wow and then and then you realize that you're in this like tiny little microcosm of people who are just obsessed with this like slot machine in a terminal um it's freaking wild uh i yeah i don't know what what was the question again see there you go there you go i'll bring you back don't you worry uh memory it was about memory really no that's i I think it's cool that you're, I mean, I'm fine to even step back in there a little bit.
1:07:19I mean, I do want to talk about memory and what you're working on there, because I'm curious about how, I've never played with the memory side of things at all, and so I'm super curious. Yeah, memory, I can talk very briefly about memory. I even have friends, too, that are zero. Like, I just, here's an interesting, somewhat of a tangent in a way, but I think it may play into what you're talking about, because it's totally right of developer. I was visiting with my newest doctor, and I live in a small town outside of Austin called Dripping Springs. And the doctor I go to, oddly enough, I'm fortunate enough to live in a town where we have a concierge doctor.
1:07:59And so I don't go there with insurance. I go there as a concierge. I pay out of pocket. I won't explain it all. I can use my HSA against it. But the point is they're a concierge-style doctor where you can be a part of a subscription, and you can go there as often as you want to, and they're all about your health. And it's not about giving you a medicine or a pill. It's about root cause issue in your life from therapy to exercise to meals to bowel movements, oddly enough, even. And I'm sitting down with this person, and she's a well-trained physician, well-trained doctor, and she's got this new practice.
1:08:36This is – now that I'm telling the story, I'm realizing how much of a tangent this is, but follow me. and I'm sitting down there with her and I'm talking to her about her business because I just naturally am an entrepreneur and I think business and I think in code and all the things. I'm very right brain business. I'm very left brain developer, which is a fantastic place to be in life. I think right now, especially now, and I'm sitting down with her and we're going through this data that I have and it says PDF and it's on her screen. And I'm like, how will I get this later. And she's like, yeah, you'll get the PDF later.
1:09:09And I'm like, but you're a concierge doctor. Don't you think you should have like a, this is my brain. Don't you think you should have like a formalized patient of record in your business, you know, and this kind of thing. And like, here's this, here's this woman who's just really well off and doing well, but she's not thinking about the data problem that people like you and I think about. And I'm thinking, gosh, I mean, the thing that scanned me earlier probably has an API. You could probably pull that data into my record, and then you could do that for everyone in your practice, and you could truly live up to your concierge doctor.
1:09:44And I guess the reason why I tell you that is that you've got these people out there, these folks out there who are super intelligent, but they're not thinking about AI at all. And she was telling me how she's really good at what she does, but she feels a little overwhelmed about the business side of her business because she's not really a business person. She's not designed to be a business person. And my response to her was, just use Claude. Do you know what she said, Matt? What did she say? Back to me. What do you think she said? I have no idea. What is Claude? Ah, yeah, nice. You know what I'm trying to say?
1:10:18Like, gosh. And so I had a brain dump on her. I'm like, okay, there's an API behind this thing here. Here's how you can pull your data over there. You need a Postgres database here. I was like, okay, Adam, you're going too nerd. And I explained what an API was. She's like, mega nerd, just don't know. And she's like, and when I got to explain to her, she's like, whatever that is, I need that. Can you do that for me? I'm like, yeah, I could probably help you with that. So now I have another job, by the way. That's wild. Helping my doctor formalize her practice on the future of AI. And so all this to say is that you've got your friends who are developers that are not using AI.
1:10:56We've got folks that are super intelligent, like doctors, that are not really fluent in AI. And it's 2026. It's March, 2026. And I'm a little nerve wracked by these folks just like being so delayed. You know, even developers, you know, there's going to be some people who listen to this thinking like, Adam, stop drinking the AI coolant. I'm an AI maximalist. It's not going away. The more you lean in, the better off you are. And you can probably attest to that, Matt, with what you're doing. But I feel like the folks that are just delaying it or feeling behind, I don't want them to feel behind. but at the same time, like it's not going to go away and leverage it.
1:11:35I said, Hey, if you don't know how to run your business or you need more help with your business, put all your problems in the cloud and it will help you at least make a system to solve them. Not actually give you the solution, but help you get to a solution. And no one's getting it. It's like cheating on your homework. It is a cheat. It is a cheat. All right. Unless you have anything to say about that. Let's end that tangent and go back to memory and stuff like that. What do you think? Yeah. Yeah, slightly on this point, I recently got my dad using Granola. And he's a doctor. And he's kind of fed up with it.
1:12:11He writes so many notes. And Granola has completely changed his whole workflow. Oh, I bet. He sees people on Zoom all the time. And he sees people in person as well. Now he just starts Granola, writes up all of his meeting notes. and it's like it's just it's been it's been like transformational for him uh just that like basic summarization like granola is a great product don't get me wrong stunning product but like it is not a complicated workflow and it's like completely changed his like quality like how long he spends doing his consultations um so yeah i guess like shout out there like there are there are some small things you can try there are some way i'm a fan i'm actually i'm a paying user of granola so yeah yeah oh wow they should pay me come on granola pay me yeah i love granola it's amazing and i'm with you on that too i think i've even dm'd uh the designer i can't remember his name in the moment but um uh i think his name sam if i recall correctly yeah sam sam's one of the founders yeah uh so answer your dms but i'm a big fan of granola i think that's revolutionary Same thing.
1:13:21My wife introduced somebody to who's also a doctor and she sits down with folks and she would spend three hours of literally every evening cramming all of her notes. Yeah. Well, in this new world, you don't have to do that. Now, you do have things like HIPAA compliance here in the United States. You got different health care concerns where you have privacy and stuff. I totally get that. You should abide by all those things. And if we don't have systems that support that, we should. But imagine the unlock in your life where you're a teacher or a doctor or someone like that where now you don't have to like arduously plan and think about your note process.
1:13:59Now you can sort of have a lot of it formalized for you, and you don't have to do all of that work to even report back to folks or summarize this 45-minute session with a patient or a friend or a colleague or whatever. You can have it do it for you. That should just be the way. anyways man i think we could probably go on that front for sure yeah i actually use um granola like on my like personal stuff for um if i have a really fun idea and i want to write something up about it because i'm actually i'm horrific at writing like i i would call myself a critique of writing or a critic of writing rather than a writer like i i love reading and I've read a lot since I was very young.
1:14:42But I really struggled with putting my words to paper in a way that like flows and makes sense and is cohesive and has a start, a middle and an end and all of the good stuff that you need for writing. So something was like, dude, just start Granola and chat to it. Go for a walk and chat to it and then and then come back. Great hack. And then yeah. And then make a really good prompt. It's like this is what I want to achieve from this. and that was the first iteration of the code mode blog post was really something yeah it was something similar to that um i thought it was really really good because like i got the points i wanted to get in because i just spouted to the ai the ai listened to me the ai didn't quite summarize but picked out the key bits of information because i really hate summaries i think they're rubbish that one person's summary is another person's i don't know mud or something it's like really really really hard to get something that summarizes something well while maintaining the full information but like things like granola you can export like a nice blog post if you know exactly what you want and i tend to know exactly what i want and i think ai is like an unlock for very opinionated people because you don't have to do the thing you just have to be very good at critiquing the thing absolutely absolutely i actually like that idea a lot i'm glad you mentioned that uh the personal use of granola because i have not considered granola in that way where it's my personal note taker because it's great at that hey friends i'm here with dan mangus co-founder and ceo of rwx dan what makes rwx and the way you're doing ci so different and interesting to our audience.
1:16:29You know, obviously we're talking to you because we want to promote what we're doing. We want more engineers to become aware of what we're doing at RWX. But I think the thing that's interesting to me is that RWX is really kind of the first major evolution in CI and the approach for CI. And this is just highly relevant with agentic-driven coding. You know, CI has largely been the same since the advent of the practice. But these platforms were created when being able to run code in the cloud was really valuable. The fact that you could spin up virtual machines that would run some automation on a git push was really impactful for engineering teams trying to build good developer processes and tools.
1:17:06But that's kind of the extent. What we've done at RWX is we've taken state-of-the-art techniques used in build systems at organizations like Google and Meta. Google has their internal build system Blaze inspired the open source Bazel tool. But every engineering team I've talked to that wants to adopt Bazel has just found it extraordinarily difficult to use and configure. You have to have a dedicated engineering team to build and maintain the rules. It's hard to extend it to work with different types of languages and frameworks that engineering teams are looking to adopt. So it's been too prohibitive to actually adopt those technologies.
1:17:41But the ideas behind Bazel are really impactful. They're similar to a lot of the ideas behind Nix. I would say Nix is kind of very similar in the difficulty to adopt. And effectively, what we've done in RWX is we've taken those techniques, and we've made it very easy for engineers or agents to actually adopt and utilize those, which namely are the automatic content-based caching and the graph-based task execution, which means that RWX eliminates all redundancy. You know, whereas other platforms are having to run the same setup steps on the same jobs in every, you know, virtual machine that's spinning up, RWX can run the setup once on on one machine and then fan out accordingly based on just your dependency graph.
1:18:24So effectively with RWX, you never have to think about parallelization at all. You know, on other platforms, it's always like, well, do I add this onto the existing job? Do I make a new job for it? But then I have to duplicate all that setup. With RWX, you just define the tasks that you want to run and the dependencies between it. And we will run it with maximum parallelization based on your dependency graph. Well, friends, a good next step is to go to rwx.com, learn more check out ci in a whole new way once again rwx.com
1:18:56let's talk about memory so uh we're going back into the deets uh we're off of our personal uh soapboxes about how ai has changed our life and how it's taken some away how we can't stop thinking about it and how we prompt etc etc but take me into the world of i guess next few agents, MCP, where does memory fit in? Yeah. So memory, it's like such a loaded term. It's such a loaded term. So it is quite hard to know where to start. Essentially, I want a way for my agents to remember a conversation that we're having right now and be able to refer back to context that I gave previously in the chat, but also to like remember conversations over time um and also to be like very programmable so i work on sdks like developers are going to program with my sdks it's like how do we how do we build something that's mega customizable to like the next new um the next new trend like for instance skills like skills are just a markdown file that's loaded into a context on demand by an agent like How can we support that in a memory system that can also support compaction of sessions, can also support continual learning, can support like the migration of a session to long term storage.
1:20:17So an agent can like search over it over time. I guess I'm just trying to work out the shape of those APIs right now. There's some really good examples, like maybe not examples, but there's some really good inspiration on the in the TypeScript world at the moment. uh like letter is very very cool letter ai master just to name like a couple they all have some some cool memory stuff and i know there are there are some there are some really cool memory startups that are actually like doing managed memory like super memory like i just like shout those guys out it's like really good inspiration for what we're trying to do what we're trying to do is not trying to replace anything like that uh but it's like how like cloudflare has some really cool storage primitives how can we let developers best use those storage primitives in the function of making a better agent.
1:21:07And I realize they're all questions rather than answers, and I don't have a huge amount of answers, so I'll probably keep my powder dry on that one. While you were sharing your ideas there, I was jetting down an idea I had. Now, this may be totally wrong, but this is how I'm currently thinking about it if I was in your shoes. Go on. All chat captured to mark down or just plain text in some way, shape, or form. So all your before compresses and goes away, it's captured. And you could probably use an AI gateway for that. And then you send them an analysis across all that history. Then you vectorize that into a database.
1:21:44Then you SDK in front of that with two calls, search. And what was it? Execute? Was that what you did? That's what you do. That's what you do right there. And you just treat your vector database on the sentiment analysis that you've been capturing as plain text just like you do your APIs. That's how you do it. So – Is that wrong or is that not even close to right? No, I think you're close. I think you're close. So there's a few things that I can't do with that that maybe consumers of my SDK would want to do. So I can't be that opinionated on where the data is stored. Like some people might want to store it in a durable object in SQLite.
1:22:27Some people might want to store it in planet scale. Some people might want to store it like that. Their data is going to live somewhere and people are normally very opinionated about that. So I can't be like, here is a vector store you must use. Although I have to have the ability for people to use like vectorize if they want to. So I need to go with more of a provider based model, I think, in terms of API design. And then the next thing about search and execute being a thing, yes, yes, definitely it's a thing. You've already made it, right? I mean, that's the model. Just leverage it. For longer-term memory, I think, and for things that need to be loaded on demand, yes.
1:23:05But there are cases where you would want to programmatically load context into a session. so um the the easiest one is like if you think of like a system prompt with some direction like in in open claw i think they call it sold or md like what is the agent like yeah who does it respond to like like what is its personality all of this little stuff this would need to be loaded on demand on the start of every session so this is like slightly different um and then the next one maybe is uh like a to-do list is some sort of working context you know like claude code had had a to-do list i don't even know if it still does anymore um but that it keeps it kept the agent on track for a while maybe they rled this out but at some point people wanted a to-do list that the agent could fill and modify over time like this that enables you to do really cool stuff like create ralph loops as well which maybe we can talk about some other time but But I need a way to be able to store all this context in a way that's super flexible and also have that ability to do continual learning and extraction of facts and also have the ability for the agent to be able to pull in stuff like skills.
1:24:24So it's multifaceted and I'm still trying to work out in my head what I want to focus on because I don't think I can get all of these things right in the first time. I just need to make something flexible enough that when the new things do come, we can add them in without breaking changes. And when you're speaking of memory, you're speaking of it as part of one of the Cloudflare products you work on. Not so much. I mean, I'm sure you have personal curiosities and I can leverage it personally, but you're talking about how you can bake it into agents, for example. Yeah, I think at some point this might be a separate SDK, but yeah, like agents SDK will be where it lives initially.
1:25:01yeah yeah so people building yeah people building agents on cloud fledgeable objects but like theoretically there is nothing to say like if you're building something on a ecs somewhere like a like a container somewhere if you're building on like a lambda function somewhere on um like you have your your like next js routes on the cell like it should be it should be pretty cross compatible for all of these things. Like there shouldn't be anything runtime specific. I think the provider model there will help because yeah, sure we can use the durable object SQLite, but also if someone wants to use NEO on a planet scale, they should be able to do that as well.
1:25:42Yeah, for sure. Would not want to dictate where you can store it at maybe even one to many stores. I don't know how hard that would be, but you know, that's where I would start to, I mean, that's what this is, right? It's all exploratory. It's like, Like that's the basis of how I would initially approach it. And I might hit two brick walls and hurt real bad and learn something new and read a book. I've become a real big fan of EPUB books. I've got an ETL that takes a book from EPUB to really good markdown and then sentiment analysis on that and then vectorizing things across it and just searching it with DuckDB and Parquet.
1:26:18Oh, wow. So reading a book now is way different than it was before. so thankful for open format EPUBs out there because that's the way to do it and like between DuckDB and Parquet and this I mean that's super fast those few things there would really lean into what you're talking about with memory and that lookup process it's super fast definitely I was chatting to some of the more data engineering people in Cloudflare and they were like yeah so how can I use ClickHouse how can I use ClickHouse and I was like ah ah shit sorry um you beat that one but like it's okay like how how yeah i don't know i don't know um what's special like what's special about click house that you don't think you can get from postgres and then he kind of rolled his eyes at me um and so that was how the conversation went but yeah more rows so much faster but really hard to set up you could do it on your own you can on prem it yourself but it's definitely a ceremony i mean it's a lot to run i mean but with at cloud flare scale you got all of that right i mean i would run click house if i was on your team definitely definitely but duck db and parquet you can run right on your mac i mean you can just run it right there and it's super fast and you can have a ton of usage just in one context but as a product you may think about it differently but duck db and parquet files is like it's the way to go I have some telemetry for an open source code review project I did a few years ago that just dumps everything in DuckDB.
1:27:53It's quite good, actually. I really like it. Yeah. It's clear. I mean, it's really interesting, too, because the agent knows how to talk to it really well. And so rather than you having to learn how to retype queries into it, the agent can query it for you. And I'm like, make me a just file command for that. And so when we sort of like centralize on a query or on a style of query, just turn that into a just file command. And I throw it a few parameters and it's like a just in time CLI in a way on a large data set. That's super fast. I mean, that's awesome to query that database any other way is just stupid.
1:28:33Like, why would you do it the hard way? That's the easy way. You know, that is the way. Maybe I'll do that with my with my claw. That sounds pretty fun. Yeah, I've been thinking about, there are some things that don't work in the situation I'm in, and there are some things where I can really take inspiration from home labs. So I've been building my claw and playing with, I really like Pi and Pi Agent from... Oh yeah, I heard about that. I haven't played with it, but I heard about it. You should play with that. There are many like it, but this one's mine. Is that one of the tagline? There are many like it, but this one's mine.
1:29:07Yeah, yeah, exactly. Pi.dev is what you're talking about? Yeah, it's really well built. Like some of the best TypeScript I've seen. It's really nice. Such a cool domain name too, pi.dev. It says there are many coding agents, but this one is mine. This one's cool. It's cool. I haven't played with it, but I saw it. I was like, yeah, that's a good nod right there. Okay. So the provider model that they have and the lower level primitive, so not necessarily the agent I don't tend to use the agent but when I'm if I'm building an agent then their primitives are pretty cool and I think I'm still in the specking phase of like working out how exactly I want to run um like my like personal AI uh yeah just like finding nice product avenues from different products that I like I really like poke from interaction I don't know was this like did we talk about this no we didn't talk about this yet that's my last call yeah poke from interactions really really nice like how they how they do like the stateful workflows in the background um i take a lot of inspiration from other products like that yeah you got to be a consumer i mean consume everything everyone's creating around ai all the new innovations even if they seem silly and toy like there's some little thing that's going on there that is inspiration elsewhere i mean i've been a home labor for a very long time now.
1:30:38I would just say I feel like I feel so thankful to be this knee deep in Linux than I ever was in my life because it's a superpower right now. So to the right of me, I have a Proxmox box with just way too much RAM and storage and CPU available. And so I essentially have my own cloud here. So I can just unleash my agents. I can build something, deploy it to that, and battle test it in almost real time on my own hardware. And I have to send it to the cloud and deal with keys and deal with payments and just whatever comes with that. I can, like, skunk works whatever I want right here, and it's too easy.
1:31:20And shout out to my buddy, as a matter of fact. On the pod recently, his name is Adam Jacob. If you know Adam Jacob from Chef. but swamp.club has changed my life y 'all okay matt you guys should you gotta check this out swamp.club okay waste waste a whole day it's not a waste spend a whole day on swamp.club and learn what you can automate especially i have if you have like a little actual raspberry pi it is software automation like you've never seen before i'm just telling you that much man it's insane okay i'll look it up i'm serious i'm enamored by this stuff i love adam he's a good friend of mine he's a super big and open source system initiative uh you know automating infrastructure etc but it's amazing so i've been doing that with my proxmox so proxmox if you're not familiar is a hypervisor so you can host vms uh lxc containers on there and so it's like a mini cloud basically for you and so standing up a new vm on proxmox is a lot of clicking in a gui old days right who's doing that well with swamp you just tell swamp hey this is the ip of my proxmox server automate all the things and i'm compressing all that down to that one phrase it's not exactly that but it feels like it and so i had this go cli that i was writing that did everything that Swamp did for me in minutes.
1:32:52And I wrote that with AI too. But Swamp automated so much stuff in my Proxmux server. Spending up a new VM, hardening that thing to be a DNS server, adding tail scale to it with my off key, with my secrets, standing up one password on there for my secrets distribution. I mean, amazing stuff. It automates so quickly. Much like code mode, it actually writes code to it doesn't it creates it via writing typescript workflows and modules and sorry models and workflows and it's just so wild that what he's done with there so it's a lot of like what you're doing with with code mode where you're like rather than calling all these tools you write the code that calls the tools kind of same thing in a way but check it out yeah no definitely definitely it looks like it no it's super cool if you're not home labbing though is you got to be home lab and what i mean by home lab is like literally standing up your own vm literally standing up your own linux ubuntu fedora pick your distro debian go wherever you want and just play don't don't drink the cloud for kool-aid too long man get your own vm get your own linux play with your own keys with your own rules with your own sudo and uh feel the metal man feel the metal i have a couple of Raspberry Pi is looking at me from the corner of my room that I need to do something.
1:34:15Plug a man, man. Ethernet those things. Yeah, let's go. Let's go. Get him in there, man. I have a great friend and he's one of my colleagues now, actually, since I joined Clubfler. And he's been telling me for ages that like, it's K3s, right? He's running K3s on his Raspberry Pis. He has like a full cluster of them. He keeps on adding another one every now and again. And he's got his agent
1:34:42deploying apps, running apps on different pods. It's kind of wild. It is wild, man. Yeah. I've got a bunch to learn about Kubernetes. I mean, even the stuff you're talking about here too. I mean, now you do have the Cloudflare account. And so you have the world's oyster in front of you, so to speak, in terms of compute and power. So I mean, I'm not saying you shouldn't use that, But there's something that changes when you go on-prem, home lab, feel the true metal of the actual physical hardware, install an actual operating system onto it, whether it's Debian or Proxmox, which is actually built on top of Debian.
1:35:22You can actually install Debian and then install Proxmox on top if you wanted to, or you could just use the Proxmox installer and just isolate from the stop, from a bootable USB. point being is like literal metal choosing your ram choosing your cpu choosing your disk storage and vme of course like there's something to that where you take parts and you make it and then you put the thing on it which is linux of course and then you build on top of that like just just something about that in this world of ai that especially now right like you may you may feel a little lost or inadequate with Linux.
1:35:58Maybe, maybe not. Well, Claude is not. I have a question to ask. Claude is not. So are you running any local models? No. And the reason why is because I'm not enough time and too lazy, I suppose. When the world's best models are available to me with the credit card swipe, I have more of that ability than time to, I even have a GPU and I'm just not even using it. because all of my interests, like nothing has to be private to that point. So I'm just like, why would I do that? Cloud's right here. Codex is right here. So I'm primarily lately a Codex GPT, GPT-5, I guess, 5.4. Usually on high, not medium, because medium is not cool.
1:36:43High is cool. Extra high is super cool, of course, but no, I don't do a lot with Google models. It takes about a year on extra high. It does, but you get some really deep thoughts, you know, For the good stuff, I'll go there, but not for most things. I'm just hanging out in high. No, not a lot with models because I just find that all my problems don't require local models, and I'm not trying to be private about any of this stuff in the way that I feel fearful to be private. It's not like I'm talking about this goiter I've gotten. It's a medical problem. I mean I don't know, but I'm not talking about anything that's embarrassing I suppose.
1:37:19I'm not doing anything nefarious. So a local model is not needed for me right now. Do I plan to? 100%. Matt, I would love to. I would love to have more time to play with local models. I just don't. So I had a good experience at the last startup I was at where we were building basically a glorified PDF parsing pipeline. And I got to play with some local models there, which was really good fun because we ended up hosting our own on H100s 100s because there was no yeah there was no need to go to the like the top of the range like gpt5 in the in that moment um it would have been way too expensive and so we needed to cut some costs and these didn't have a huge amount of usage so it was like it was really good to use h100s and play with it and like do a little bit of tweaking about like which model like have have a couple of evals oh my god this model couldn't do it this model couldn't do it this model couldn't it oh my god this model managed it right can we do right use this can we use this size but can we go or can we go a little bit smaller with this with this brand of um of model this version can we go a little bit smaller a little bit more quantized does it still manage our evals like that was really fun that was a lot of tweaking it was a lot of fun but um so i i have some like pull to want to like play with some of the new open source models i mean that if you're feeling about left behind those open source models, they make you feel left behind every three weeks.
1:38:46Every three weeks, a new version came out of the model. New versions, something new. A new leapfrog. It's a tough game, that. A tough game. You know, the one thing I will say this here, and you might enjoy this as a fellow Homelabber, up and coming maybe. Definitely. Is I've written a DNS server in Rust. It's called DNS Hole. And I've been teasing my audience about this for a while And so I'm sorry about that, but I am getting really close to releasing it. I just did some really cool code review on it. It was super dope, but I'm just nervous, I suppose, about releasing it to the world. But I'm using it.
1:39:24Right now it's my DNS server, as we speak, right here, right now. And it's a replacement for PiHole. So since you have a Raspberry Pi, you may hear about one of the first things you tend to do with a Raspberry Pi is install a PiHole or stand up a PiHole in your home lab. on your LAN. And so I've written this DNS server, but I have this idea for kind of like I want an AI that constantly sniffs my traffic. So rather than me build my block list based upon nefarious actors, I want the agent, the AI, I suppose, to sit on my network and pay attention to all the real-time, the hot path traffic that my DNS is resolving.
1:40:09And I wanted to, I suppose, with intelligence, with AI, add to my block list because it knows. I don't want to have to manage my block list, Matt. And so I want to have an add-on that calls the API and pays it into the traffic, and it's got that hot path. And it wouldn't be the primary DNS server because that would be stupid. Let's put it to the sidecar of that. But one place where I wanted to play with the local model was in that. I wanted to have the hot path of the DNS being resolved and then an AI that's localized right there, but a very small parameter, like a 1.5 or a 3 billion parameter kind of thing.
1:40:49Just enough to be intelligent about that kind of traffic. If it's something that it doesn't really know about, it just sort of files it as like, this needs deeper investigation. For the most part, it can classify most traffic as good or good or not good. But it's going to manage my block list for me rather than – and the cool thing about that is you may see people say, go get this block list or that block list. Well, that block list is not based on my traffic. And so it's this massive list that is contextually not really true to my network. And so my idea is like, let's let's add an agent in the loop there and let's make a local model and let's let that thing determine my block list based on the actual traffic coming into my network.
1:41:36And so the plan I have in place, which I don't have time to build yet, is around 5 to 10 seconds after the DNS gets called, the first resolution of it, this agent will be able to infer it, check it, and add it to the block list within 10 seconds of it entering my network. Now, I don't know how you feel about that with security, but that's about as close as an instance you can get, right? It's not days later. It's not somebody else's block list later that I'm once a day sinking. It's literally based on my traffic, almost in real time. And the moment it's seen, it's evaluated and added to the block list.
1:42:15And how would it know? My house is safe. What is a key indicator of nefarious traffic? I'm glad you asked this. I mean, subdomains is a big one. So a lot of weird characters. Man, I wish I had my notes in front of me. There is a really cool – let me see if I can get my notes in front of me. It's essentially – I'll figure out the name of it, but I'll paraphrase what it is because I can tell you that part, but I can't tell you the name of it in the moment. But it's essentially like saying, okay, you have the name Matt, right? M-A-T-T. Check. No problem with Matt. But now if you do MZA1TT, that's a weird characterization.
1:43:01So it essentially watches. It's a name for a thing that knows what the proper sequential in English or any language should be. And so when the characters are off, it flags it. And that happens a lot in nefarious actions. So like Google.com is a pretty – it's an easy way to spell out a domain or even pi.dev to pull it back to our friends at pi.dev, right? That's normal. That passes the test. So just based on the domain alone, which is DNS, it's like, well, is this a really weird subdomain with weird characters? Flag that immediately. And so when you look down all these block lists, it's a lot of that.
1:43:41And so just on that alone, you can – at the DNS level, which is like network lookups, that's the most secure you can be when it comes to stopping something in a network. Just based on this one algorithm alone, you can stop 99.9 % of bad traffic that should not be on your network. So just that alone. And that's not even intelligence. That's before the AI. So just based on that algorithm, I will check all those and block based on that or flag it for the AI to go and do deeper analysis. And the AI will take care of the 0.5 % or 0.1 % that that can't catch or that doesn't catch, that is truly nefarious and needs a little bit more sniffing.
1:44:22That's what I'm building. Cool. It's dope, man. It's dope. It's dope, man. It's super dope. Ask Claude. That's why you need a home lab, man, because then you have these kinds of ideas, man. You start worrying about your DNS, and you start worrying about how you can block the nefarious actors. Because all those block lists out there, they don't do you any justice when it's not your network and not your traffic. You do it in real time, 10 seconds later, after the first lookup. Yeah, that'd be cool. Kind of off track. Happy to rant about my DNS hole, which is super cool. But I do want to bring it home for one more thing before we tail off.
1:44:59is I would like for you to give the audience a takeaway in some way, shape, or form. If folks are like, you know what, man, this code mode is so cool, how do you use code mode? What's the first step? Give us a first step to using code mode, and how do you actually build day-to-day with code mode? Yeah, of course. So I guess the first thing would be to go and have a look at the blog post or dump it in your coding agent. so it's like blog.cloudflare.com I think forward slash code dash mode dash mcp and we'll hopefully pop it in the show notes I think that's like if you dump that in your coding agent or or just like have a read I think it's a decent read give you a lot of insight but the main thing is try not to let your AI do determinist do like individual discrete actions try to write try to write an SDK, write a CLI.
1:45:55Like a CLI is also code mode in some way because a model is writing, if it's writing bash, as far as I'm concerned, it's writing code. You know, like I prefer to write TypeScript or Python, but if it has to write bash, then go to town. So like let the model write code and get out of its way. Just like let it roll. It'll be fine. Let it roll. Yeah. And then if you want it to be like secure and you want to deploy it properly, then have a look at deployment options for building or using like a sandboxed interpreter or some type of sandbox, whether it's like a VM, whether it's, I know, Pydantic released something pretty cool called Monty, which is like a Python interpreter built entirely for code mode.
1:46:44That's pretty cool. Or dynamic worker loaders, like the Cloudflare option for running JavaScript. but like really like, yeah, sure. It's fun. You can run code locally, just eval it. But if you want it to, if you want it to be safe, secure, and I don't know, deployed properly, then have a look at some of those options. But really, yeah, the main takeaway is just let the model write code, dude. Let the model write code. Yeah. I assume you prefer TypeScript over Bash. I don't mind Bash. Bash is the limi franca of agents these days. TypeScript is the second best, I think, for the agents, but they're really fun with bash.
1:47:20I just let it rip. I like bash. I just think there'll be, there's an easier permissions model if you're generating a TypeScript SDK. Because the first thing, the disclosure of features is easier. It's much harder to generate, you can't generate types for a CLI. So the model has to go through each individual, you have to have a skill or something to tell the model to call the CLI to begin with. And then the model has to go through and look at each of the options, call help on each of the options to find the one it wants. If you can generate types up front and you give some more information, some more concise information.
1:48:03So I like that. And secondly, the permissions model, I think is better. If you run JavaScript or Python, you're running it in some type of sandbox interpreter. You can take the fetch requests that are trying to leave that sandbox. In terms of us, it's our dynamic worker loader. You can take the fetch requests that are trying to leave the isolate and you can be like, inspect them, have a look at them. What's the model trying to call? Does it meet your permissions set? Does it need special permissions? Do you need to add authentication tokens? Do you need to add, like, what's it trying to do? Basically, you can have this anti-corruption layer, this ACL that sits around the code execution.
1:48:55And I think that's a better permissions layer than just YOLOing into a terminal. I like that. I'm going to give you a nugget. I think you might like this since you're speaking like that. Go on, hit me. I have this thing I've been building to all my CLIs lately. Well, I'd say lately is like the last three months, maybe more. is dash dash agent. Nice. You got to add this flag. So I mean, we humans, we love dash dash help, right? But agents don't have that. And our help is not their help because they parse things differently. They like markdown. So my dash dash agent essentially is tell the agent what this thing is and how to use it in markdown.
1:49:33And that's what it does. It responds with a markdown standard out, you know, to the prompt. And so that's my gift to you and everyone else. I've been doing this and getting great results, but you can throw it on any command dash dash agent. So, you know, Cloudflare D login dash dash agent. Like, what is this login kind of thing? And explains it to the agent. That's a pretty easy one, though. But something maybe more difficult might be like Cloudflare D tunnel new or something like that dash dash agent. And it explains it in Markdown to the agent how to use it. And so when we have code mode, I suppose, in a CLI, you can give it the same next best tool.
1:50:14So it can parse all the commands and figure it out. But you can give it one more easier nudge by doing dash dash agent versus dash dash. That's cool. It's cool. I like the nugget. I like the nugget. Have you not thought about just doing dash dash help but detecting whether it's in an interactive environment? And if it's not interactive, then printing markdown? I suppose you could do both, really. I mean, I would alias it then at that point. I didn't think about that. That's a good point. My thought was just really like I want something special just for the agent. And that sounds cool with dash dash help.
1:50:46But you can certainly alias it if you were the first class citizen is the dash dash agent. And then if you're doing dash dash help and you didn't tell it to do that and it's discovering it, determining if it's in an interactive terminal, just give it the same thing and just alias it. That's a good point. I think there's a lot of trying to make CLIs, like some CLIs, everyone's been trying about how cool CLIs are, but some of them are like not natively useful for agents. Like some of them rely on interactive process. The one that really bugs me, and we need to finish sometime soon, but the one that really bugs me is, have you ever used change sets?
1:51:26No. Well, change sets CLIs is entirely interactive. It's like a package version manager thing. It deploys new versions when you do a change sets and then you put your change notes in there and it collates them all together when you do a release and stuff. It's like a management for open source package or for packages. It's really good. I would recommend. It's really nice, but they don't have a non-interactive version of their CLI. So if a model tries to do NPX change sets, it just like freaks out because nothing works so it's so annoying all I want is NPX change sets the package name and then the change log and I might make a PR for this because it would save my life yeah I think the more we can go the agent's way with interactivity around that I don't mind a CLI being designed for a human but also agent aware or agent native because I still use CLIs myself, or at least I want to in some cases.
1:52:31But in most cases, I'm just like telling the agent to do that stuff for me because why would I do that anymore when I don't have to? I can just like let that one thing over there spin and do this thing here and here and here. I mean, that's the better world in dramatic cases, really. So there you go. Yeah, we have gone long, and I appreciate it. We went deeper than I thought on some cool stuff, though, but I enjoyed it though. Very much so, Matt. I'll link up obviously both your blog post as well as the original CodeMode blog post that we talked about in the show. I have my robots treasure trove, this entire transcript for all the cool stuff and the bits and bobs.
1:53:12It's all in there in the show notes. So it'll all be in there as best it can be. And if not, our show notes are open source on GitHub. So if you missed something or you want to add something that is contextually true, then send a PR, I guess, or have your agents in a PR or I don't know, something like that. Awesome. Matt, thank you so much for all you do, man. It's fun talking to you. Thank you. Lovely to meet you.
1:53:40Well, friends, this show is done. Thank you for tuning in. I hope you enjoyed this conversation I had with Matt Carey from Cloudflare. Wow. I mean, like seriously, some cool stuff happening in and around this agent space, MCP, APIs, what Cloudflare is doing. They're doing some really incredible stuff. I got to tell you during the podcast, maybe you could tell I got a little FOMO. I kind of wanted to work at Cloudflare about midway through this podcast. You know, I feel like I can make a dent there. I don't know about you, but I feel like there's just so much to do, so much we can build in this very moment.
1:54:15I'm having some fun building my own stuff. But hey, that's it. This show's done. Thank you for tuning in. We'll see you again so soon.
1:54:48Man!
From the publisher
This week I'm talking with Matt Carey about Code Mode and how most of us have been thinking about MCP all wrong. Matt works on the Agents SDK and MCP at Cloudflare — we discuss how server-side Code Mode lets one MCP server expose all ~2,500 Cloudflare API endpoints in about 1,000 tokens of context, the dynamic Worker loader that runs model-written code safely in a V8 isolate, Matt's own workflow with Claude, where memory fits into the future of agents, and his Zaggy git wrapper that keeps agents from force-pushing his repos.

