The harness is the showdown, the humans are the tool calls, and have you seen my Claude Code buddy?

24 Apr 2026 · 37 min · 15 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI pricing “harness” chaos, how model choice and token limits change behavior; Claude Code/Anthropic pricing experiments; leaked Claude Code source patterns for production coding agents; and a Vercel security breach tied to third-party OAuth access.

Guests/backgrounds

No named guests; two hosts, Ben Lloyd Pearson and Andrew Ziegler, discuss as practitioners building agentic workflows (e.g., CloudCoder/Claude Code, VPS workflows, in-terminal memory like Beads).

Key claims

Pricing plans incentivize users to consume maximum tokens until limits; seat-based plans may remove user model selection; Anthropic’s Claude Code “max” tier testing suggests future cost increases; agentic systems require efficiency, observability, and strict permissioning; Vercel breach highlights supply-chain risk and OAuth token hygiene.

Notable examples

GitHub Copilot paused new Pro/Pro+ and student signups; Opus 4.5/4.6 removed while 4.7 remains; Anthropic managed agents at $0.08/hour; Claude Code leak: 12 reusable design patterns (persistent instructions, RPI explore/plan/act, tiered permissions, externalized task memory); Vercel breach via employee-granted Context AI access to Google Workspace OAuth, plus token refresh concerns.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Mourning Trixel: The Terminal Buddy

0:00 to 1:18

Exploration of the emotional attachment to a lost terminal buddy feature.

“Oh, so you saw my message about Trixel and sadly what happened.”

Changes in GitHub Copilot Pricing

1:32 to 3:24

Insights into how GitHub Copilot is adjusting its pricing and subscription models.

“You want to just start right at the top, Andrew, with this AI pricing chaos?”

Consumer Behavior in AI Usage

3:24 to 4:21

Analysis of consumer tendencies to use more expensive models due to pricing structures.

“But, you know, I, I kind of think in some ways this is a pretty simple problem to solve if you really think about it.”

Local vs. External AI Processing

4:21 to 4:59

Discussion on the advantages of local AI processing compared to relying on external services.

Model Selection and Efficiency

4:59 to 5:35

Exploration of model selection strategies and their impact on efficiency in AI tasks.

“You make a lot of great observations here.”

Predictions on Future AI Costs

5:35 to 8:09

Future outlook on AI tool pricing and potential increases, focusing on Claude Code.

“And also like understanding when you do this, you start to realize that not everything needs to get sent to the most expensive API call that you can make now that you're making just raw API calls.”

Anthropic's Constant Evolution

8:09 to 12:21

Insights into Anthropic's iterative updates and user experiences with Claude.

“You're describing like my exact trajectory.”

The Dynamics of AI Tool Pricing

14:07 to 16:12

Explore the rapidly changing landscape of AI tool pricing models and usage efficiency.

“You know, it's not just like chatting with an AI tool of your choice.”

Emerging AI Harnesses and Strategies

16:12 to 19:34

Discuss the diverging strategies of major AI providers in harnessing AI agents.

“And that infrastructure layer that manages all of that execution and memories and tools has a staggering price tag involved.”

Claude Code Insights and Design Patterns

19:34 to 24:18

Delve into reusable design patterns from Claude Code that enhance AI agent production.

“Yeah, and things around too about how you operate these tools.”
Show all 15 chapters

Vercel Security Incident Overview

24:18 to 27:27

Analyze a recent security breach at Vercel and its implications for web safety.

“That way their work doesn't influence each other.”

The Complexity of Recent Cyberattacks

28:02 to 29:12

Learn about recent cyberattacks and the complexities involved in modern infiltrations.

“Thankfully, I don't think the surface area of this attack was very big, and Eversell did everything right in this scenario with notifying everybody and rapidly responding to the problem.”

OAuth Token Risks in Supply Chains

29:15 to 31:28

Discover the importance of managing OAuth tokens and the risks associated with them.

“This is a employee that allowed access to a third party tool, which then it used as Gmail to move sideways into Vercel systems.”

Agentic Software Development and Security Risks

31:29 to 33:38

Understand the risks of agentic software development and how to mitigate them.

“If they're sitting around out there and they still have access to these people's accounts, even long after they've they've been deprecated.”

Human-AI Collaboration in Software Development

33:39 to 35:39

Explore the evolving dynamics between human developers and AI systems in software engineering.

“So, you know, I hope we can all practice a bit of a blameless culture and not point too many fingers about why this happened yet.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Ben:Oh, so you saw my message about Trixel and sadly what happened. You'll remember a few weeks ago when Claude rolled out a whole bunch of features that got, you know, quote unquote leaked. And one of them was a fun buddy system where you can like hatch a buddy. It had like a name, a description and like some fun traits and whatnot.

0:22Andrew:It would literally hang out with you in the terminal.

0:25Ben:But sadly, and what none of us knew, is that when you upgrade your clod version, which I mean, God, come on, we're all doing pretty regularly now, it gets rid of your buddy. So session by session, as I was finishing up what I was doing and closing it out and then whatnot, I lost all my sessions where Trixel existed except for one. so now Trixel is just in this one little pain on this one cloud code version in my terminal that I never use because he's been fully upgraded and it got rid of him everywhere else

0:58Andrew:it's almost like a friend that leaves without ever saying goodbye you just didn't get that chance to say goodbye I know, I'm so sad, it's like I didn't even mean to

1:06Ben:kill you Trixel, I just was trying to upgrade my cloud code version but I guess that's you know, shame on me for getting to attach to my little terminal buddy, but there'll be more.

1:17Andrew:Yeah, well, we're some friends here that aren't saying goodbye today. Welcome to the Friday Deploy brought to you by Linear B. I'm your host, Ben Lloyd Pearson. And I'm your host, Andrew Ziegler. And this week, we are covering AI pricing chaos, cloud code patterns that have been exposed, and Vercel's third-party AI breach. You want to just start right at the top, Andrew, with this AI pricing chaos?

1:43Ben:Sure. So let's talk about all the different providers right now adjusting the pricing around their harnesses and our access to their models. This is a really a pricing revolution that's happening among the biggest names in the field. One of them that's like a big story that comes to mind first is at GitHub Copilot, they made changes to their individual plans. In fact, what they did was they even paused signups for new subscriptions to GitHub Copilot Pro, ProPress, and also their student plan. And that's a pretty shocking kind of like discovery to find that they're going to stop all growth to address what is clearly a pricing problem within the model usage.

2:23Ben:So what are you seeing when you see this kind of thing happen in the news?

2:28Andrew:Yeah. Yeah. And they also mentioned, you know, they're tightening like the usage limits for individual plans. So, you know, you're potentially going to start getting less usage out of it. And they're also deprecating some Opus models. Like, you know, they did say that Opus 4.7 will remain, but 4.5 and 4.6 are going to be removed, at least from the Pro Plus tier. Which I actually thought that last bit was pretty interesting because, you know, supporting 4.7, but while dropping the other two, it makes me think that it must be to do with promotional pricing for the 4.7 launch. So it kind of feels like it's a temporary state.

3:05Andrew:Like I can't imagine that they're really getting like long-term cost savings out of that. But yeah, I mean, suspending signups, I think is the most telling aspect of this, you know, like that really says that they're concerned that they aren't going to be able to fulfill the promises that they're making to their new customers. But, you know, I, I kind of think in some ways this is a pretty simple problem to solve if you really think about it. But, you know, like most of these tools that we're using right now, they just encourage you to default to the most expensive model for everything. You know, just why not pick the best model at all times?

3:44Andrew:Because, you know, I'm being charged a fixed limit and a fixed cost. And as long as I'm not hitting my session limits, you know, that pricing model encourages me to consume as many tokens as possible until I hit my limit. So that's like that's like what I'll do. Like yesterday at the end of the, at the end of my day, I was like sitting at like 89 % clawed to session use it, uh, like out of my session limit. And I was like, I should push this to a hundred before I leave for the today. Right. Like it just wants you, makes you want to consume more. Um, and you know, a lot, and these companies really aren't thinking about efficiency yet.

4:20Andrew:Like we have covered some stories like Shopify. we recently covered how they're switching or they've been using quen to sort of use sub-agents to do things locally and i think that you know there's a there's a lot of space for companies to come in with systems that get better at delegating work to less expensive or even local models so that i can do as much as possible like on my local laptop or you know in your case you'd like to do everything on a vps like why not just do your ai processing there rather than having to rely on an external service. So yeah, if there's anyone out there that works at GitHub listening to this, I think that's your answer.

5:01Andrew:So you're welcome, I guess.

5:03Ben:Yeah, you heard it here first. You make a lot of great observations here. One of them being around having a model router that understands what level of intelligence needs to serve the requests that you're asking and not defaulting to using the most default and intensive model that you can. I mean, a lot of these providers, they do shove the strongest and the best in front of you because that's their best shot at guaranteeing you're going to have that great user experience and they're going to keep you. And so it makes sense from like a growth play of why they do it. But ultimately, like in the world I've been in of taking things that I typically did, like in a chat in CloudCoder, I would typically like or hop over to Cloud or OpenAI and use it in ChatGPT, like turning those into workflows and agents and systems with determinism and that live somewhere else, like in another machine, is that you have to understand all of like the seams around those requests.

6:01Ben:And also like understanding when you do this, you start to realize that not everything needs to get sent to the most expensive API call that you can make now that you're making just raw API calls. You start to think like, oh, this stage of the process is just a little bit of pre-processing or we're picking something out or we're cleaning something up. And like, let's throw that the haiku. Let's do something on machine, which I've done now multiple times. And understanding that model choice is a big gradient. And we have to, as consumers, get smart about choosing where on that gradient we need to be for what we're doing.

6:38Ben:But in the meantime, I feel like we're all going to just be in this environment where we're just constantly encouraged to use more of the tokens. Like even when Opus 4.7 came out and I upgraded and got access to that, immediately it's on the extra high effort, high intensity mode to the point where like I think if I ask it to ultra think it's going to dumb itself down. Like I've kind of like it's trying to operate on such a high level. Maybe I don't always need it that way. So a lot of really good observations about how it kind of like feeds the way that we use the tool. Yeah.

7:11Andrew:I mean, it seems like all these companies that are building like wrappers around the foundation models are just in a really tough place because I think what's really going to probably happen or what will be one of the many results from this trend that we're seeing, because we'll get into another story here that is very similar, is that when you're paying for like a seat based unlimited usage plan, you're probably going to like I feel like being able to select your model is a luxury that we have right now. that when you're on a seat-based model will probably get taken away from you at some point.

7:44Andrew:The provider will decide which models are going to be used for you. And that's where then you go to an API where it's more usage-based, you know, because we have some workflows, like my personal workflows are all on my seat-based consumption model. But when we deploy a workflow, it gets moved over to an API consumption-based model. Then we actually do get really conscious about which model we're picking because we have to pay for it. Exactly.

8:09Ben:You're describing like my exact trajectory. Like my cursor, when you use cursor, cursor is a model router. It defaults you to using auto mode, which then routes you to a lot of different models, including their own, to do a lot of different requests, which they know they can serve at various rates of cost. And that's how cursor has been able to just scale to such a magnitude. So you're right that these providers, they become the router just as much as they're the harness.

8:33Andrew:Yeah. All right. Well, let's talk about one of these foundation model companies. is also playing around with their pricing. So Claude Code. Andrew, do you think this is going to cost$100 a month moving forward? What do you think?

8:45Ben:A Claude Code? Yeah, I think it's probably going to cost more eventually. I'm actually very much of the mind that the costs for all of these tools are going to start ratcheting up and up and up. And you're going to start seeing equivalents that are more closer to salaries just because they're going to be able to measure the value of the output that way. and I feel like a lot of those tools might slowly get out of a lot of people's grass

9:11Andrew:yeah yeah and of course I bring it up because Anthropic got called out this week for quietly testing moving Claude code to their pro or excuse me to their max account it's currently offered at their entry-level pro account so you know a cost difference of five to ten x depending on which max account you sign up for social media was abuzz with all of the news about this and people wondering like, oh, are we going to lose access to Cloud Code? You know, all these people that are on these$20 a month accounts. There's a representative from Anthropic that, you know, said it was just a small test on like 2 % of their new signups.

9:48Andrew:However, the pricing change was visible to all users, which, you know, just added a little bit of confusion to the market. It seems like Anthropic's doing some pricing experiments on the back end, trying to see if people are willing to pay more for something that they're currently getting a lot of value out of at a very low cost. So yeah, I don't know, Andrew, what do you think about this story? Like, do you think it's something, a sign of an imminent change? I mean, you seem pretty convinced that it's coming either way.

10:14Ben:Oh yeah. My Cassandra complex is on fire with this one. Absolutely. I just, as, as, as someone who's paying for just one of the max tier, the highest level max tier accounts just one i have to caveat that because there are folks who are adjacent to this podcast who use multiple and there are a lot of folks who bought them and use them and use them heavily for their open claw subscriptions which you'll remember a few weeks ago we talked about anthropic clamped down on for the same exact reason right of people not utilizing the hardest in the tool and the allotments correctly and really messing with their pricing so as someone who's on the highest level max tier it's like i know it's not getting it taken away from me yet but i do expect the cost for my tier to go up even when i got my vps immediately after getting the vps it was i was told that like you know when it comes time for renewal it's going to cost a lot more just because like the door closed behind me and it shows how high the demand is and when we talk here and our listeners here and like we're all listening it's like we're a little bit ahead of that curve because we're all early adopters and early users.

11:25Ben:But there's a huge surging wave of demand behind us that is hard for us to comprehend. And this pricing experiment, you know, companies of this size are going to do pricing experiments all the time. I don't think that is necessarily the biggest, scariest thing in the world. If this freaked you out or if this was an existential crisis for you, then I invite you to step back and think about the places where you You can source your inference and not be so dependent upon one provider. That might mean distributing your workload across multiple tools. It might mean starting to explore some local language models or self-hosting.

12:00Ben:You might find that GPU hosting and fine-tuning costs are just much, much, much lower than this. And if your biggest concern is the harness, just remember that a coding agent is a very small loop with about three tools. And there's a lot of them available. and that's not going to get taken away anytime soon. Yeah.

12:21Andrew:Yeah. You know, the pricing experiment, I actually want to touch on that a little bit because I think it is, I think it is a little telling about the culture and the way that Anthropic works. You know, I'm, I'm using Claude, like basically every hour of the working day at this point, the platform almost feels like this almost amorphous blob that just sort of like constantly shifts in real time. Like things are always changing, adding new features, like I'll open a new thread and Claude will respond in a way that I'd never seen before and will do something that surprises me.

12:50Ben:And I feel like this is just a constant state of change.

12:53Andrew:And I don't know, like Andrew, if you've seen this, but like recently I just see random errors popping up that is like Claude couldn't do this or something. And it's like this obscure error message, but then it just keeps chugging along and solves my problem anyways. Like it just looks to me like they're constantly iterating on the platform. And I think this actually parallels like what's happening to a lot of organizations, particularly those that are starting to operate more in this AI forward or agentic way. AI tooling, we always have to remember, it's part of this stochastic system. So if you ask AI to build something that has relative complexity 100 different times, it's going to do it 100 different ways.

13:31Andrew:And not to mention, sometimes it's going to succeed, sometimes it will fail. So you have to have checks in place to capture all of these things. And there are clearly an emerging group of companies who are operating in this highly agentic manner. And I'm not just referring to engineering, like we see it a lot in engineering, but there are places where it's happening across the entire company. And they're operating in what I would call like an agent first manner. So when I say that, it's like what I mean is when you have a challenge or a problem that's presented to you, your default approach is to go into an agentic system to solve that problem.

14:07Andrew:You know, it's not just like chatting with an AI tool of your choice. It's like having an agent that you communicate with, convey meaning to it and have it go out and solve that problem completely for you all the way to production. and you can see how quickly Anthropic is iterating on their core platform it's like one day you have this cool little buddy and then the next day it's gone forever but you know I'm at these these tests like to your point these tests are probably going on non-stop on an Anthropic like on an hourly basis you know I think some companies that maybe get to a point where they can do it maybe on a daily basis or a weekly basis but It could even be happening faster with a company like Anthropic.

14:53Andrew:And I suspect that they have a good sense of how different companies use their platform, potentially even all the way down to the individual level, like which individuals are using our platform in a certain way. And based on that data, they can serve as capabilities to them to see whether or not the capability works. So I'm not really reading a whole lot into this specific case. you know you might be right that these capabilities are going to get more expensive either that or you know maybe maybe like co-work becomes the$20 a month version and code becomes the power user you know 100 or$200 a month version something like that you know and co-work maybe they can do what i was saying where they implement more control that takes the decisions out of your hands so you don't get to pick the models and and whatnot but anthropic is also struggling with the same problem that GitHub has.

15:44Andrew:Their current pricing model encourages me to consume as much as possible, as frequently as possible, as long as I don't exceed my limit because that stops me from being able to work. But yeah, it's just further proof. We need more efficiency in this space. Whoever can solve this problem of giving us these powerful systems, but in a much more efficient manner, that in my books is going to be the next winner in this space.

16:10Ben:Yeah, really well said. And I think that leads us pretty well to our next story as well, where we talk about how these major AI providers, like we've been talking about, Anthropic, but also OpenAI, Google, and Microsoft are converging around the idea that, you know, AI agents are harnesses. And that infrastructure layer that manages all of that execution and memories and tools has a staggering price tag involved. Understanding its consumption is a huge task. And they're all taking radically different pricing approaches to try to understand what the winner is going to be. So some of these I want to call out is like, we've been talking about Anthropic.

16:48Ben:They also launched this managed agents system. And this operates at$0.08 an hour. And Ben, this might be a little bit of an answer to what you were asking for of you put it in their hands and you let them decide the model and how it should run. That's kind of how managed agents work and allow people to kind of push that up into their Anthropic system and let Anthropic handle all of the nitty gritty details. It's probably easier for them and more predictable for them to price that kind of work, especially because it's more batch based. Whereas like OpenAI then immediately countered with an open source SDK because that's OpenAI's play is put the harness and its tools for building in everyone's hands.

17:27Ben:And so they provided this open infrastructure representing kind of like a fundamental split. You know, you see one go fully open source, you see one make it this closed proprietary system. I think it shows a lot about like the diverging ideas. What do you think about some of the differences in which these organizations are pricing their harnesses?

17:48Andrew:Yeah, I mean, you know, cloud service pricing has always been somewhat opaque, but you can generally estimate what your costs are going to be if you know the level of resources that you need to implement. Like you can decide that a certain service is big enough for what we need. So it has an ongoing fixed cost or, you know, maybe a cost that scales based on our estimated usage. But I feel like AI tooling pricing is still just all over the place. We've been hitting on it a lot here. And I think the competition is great, but I'm really hoping that we can get more consistency and expectations around how these things are priced.

18:26Andrew:Like it's kind of obvious that that most many of these companies are operating in unsustainable pricing models today. So it would be very, I guess, comforting to know where the industry as a whole is going to standardize over, you know, over the next year or two so that we could make better decisions about like where we need to be investing our AI usage. But yeah, and then the article specifically calls out that it's really frameworks like LangChain, CrewAI, VoltAgent, like these are the most likely that are likely to be disrupted by all of these agent harnesses that are emerging. Like that seems to be like the thing that everyone wants to productize right now is a harness on your agents.

19:12Andrew:And yeah, this article that we'll link to, it does a really wonderful job at just illustrating the current state of what I think are the two biggest trends happening right now. And that is harnesses and orchestration. You know, everyone on the leading edge is thinking about those two terms. And there's a lot of products now that are starting to emerge in that space.

19:34Ben:Yeah, and things around too about how you operate these tools. Like we've been talking with a lot of folks on this show that have been pioneering a lot of that stuff. Like we've talked about the RALF loop and research plan implement RPI. We've talked with both Jeffrey Huntley and Dex Horsey behind those ideas. From this, it's like another thing we've really learned, for example, is when we talked with OpenAI's Codex team, We had Thibaut Sotio on the show to talk about exactly this, about their play, to make it open, to take a different stance in this conversation, kind of like what you're calling out, Ben, of like, you want them to come together.

Read the full transcript

20:11Ben:It's like they are firmly divergent in their theory and their strategy on what will be prevailing. So for us as users in the middle, it's like we're in these places where we need to adapt and understand our own usage because we can't necessarily rely on our providers that provide that good, safe, infinitely scalable sandbox for us. And one thing I'll call out there, going back to what we said earlier about model routers and about like you start building agents and agentic systems, you start making those model choices yourself and you start choosing cheaper ones for different parts of the process.

20:44Ben:You know, that same thing is happening here. If you're trying to get a really strong grasp on your pricing and where to take it, that observability and understanding how much each run of your agent costs and where the costs are sunk is really valuable because that's how you become less dependent on the thrash of these pricing models. You start getting cheaper layers in between. Yeah.

21:09Andrew:And continuing on this topic of agentic harnesses, let's talk about the gifts that just keeps on giving the leaked source code from Claude Code. This next article that we'll link to in the show notes features 12 reusable design patterns that have been extracted from the Claude Code base for production coding agents. You know, things for like memory and context, workflow orchestration, tools and permissions, automation. there's just a lot of like really interesting flow charts and patterns that are in this article. The author of this article mentions that these really do represent a lot of the fundamental architectural patterns that Anthropic is leveraging, many of which may even be relevant even as these technologies continue to evolve for some time.

21:58Andrew:So yeah, this is a rare look at the the inside of a production level agent system that is used by like hundreds of thousands of people, you know, around the world right now. It's such a good analysis that like I was looking at each one of these flow charts and just being like, yeah, that makes sense. Wow. That's really cool. And just feeling like I was in the head of the anthropic engineering team. Like what did you think, Andrew?

22:22Ben:Oh, I loved this article because it perfectly captured the things from the leak or like the source code being available that we should be paying the most attention to. And that is how is Claude Code primed to do its best work? And you find traces of those ideas and theories all over the code base. And then this article collects them in the one spot. And what you get is almost like, it's like this is the canon of like how to effectively work with agents and orchestrate them at scale. It takes actually a lot of the ideas from Gastown that we've been exploring all year. and breaks them down into smaller, more fundamental parts that sure don't have colorful characters and things involved with them.

23:06Ben:But it actually finally gets at what I was really hoping to get maybe sooner, but now it's here, of a more academic and empirical language and terms that we can use to talk about these experiences we have. I've been exploring this a lot, coming from the AI hackathon I did earlier this year, and how I used my planning and preparation method there to actually execute that. And that's been a fundamental part of how I work. And so when I studied this article, what I did is I actually provided the article to my orchestrator. I asked it, you know, hey, like, check out this article of these different times of coding hard-dispractices.

23:44Ben:Like, what do you think I do? What are our opportunities to get better? And in addition to learning a lot of ways that I could improve my own flow and system, my agent itself called out some things from this roundup that I do that I really resonate with. And one of them is having persistent instruction files, obviously starting everything from having a Claude MD in its root, but specifically calling out that using a global one to set global level practices and using local ones to do local project level practices lets you scale and copy and paste and move things around a lot easier. Also, it calls out the whole explore, plan, act methodology, which is RPI, which is like what we've been all about here in separating the concerns of work into different sub-agents.

24:29Ben:That way their work doesn't influence each other. Also, things like using hooks and tiered permissions. I offload a lot of cognitive burden from the agent into a linter for every language it works in. that handles all of the cleanup and formatting and best practices without the agent having to spin cycles on it. And lastly, the biggest one that stood out was externalized task memory, which has been, I know you know I've been talking about this all year, Beads, using Beads to do all of my agentic task management. And Beads is an in-terminal memory system. There's one from Steve Yege that's very popular with Gastown.

25:06Ben:There's a much simpler one in Rust by Jeffrey Emanuel. That's the one I use. But this externalized task memory is one of the most critical parts to controlling your context window and turning it into a durable store. So that's a huge unlock. If any of those are new to you or if you haven't quite unlocked how to explore them, I challenge you to just feed this article to your agent and ask how you can get started. Yeah.

25:31Andrew:And, and, you know, to, to one of the points you made, like there wasn't a lot in here that surprised me because it really did feel like it was just validating experiences that I have with Claude every day. It was like, oh, now I understand why Claude works this way. Um, and, and I think really that's kind of the beauty of it. It's just the simplicity of this architecture behind the systems in this, you know, you mentioned the persistent instruction, uh, file system. Like it's a very simple way of getting some high level consistency. You know, I was also really fascinated by the compaction patterns that they that was shown off because, you know, I've actually been very curious about that recently just to understand like what it's doing when it does that.

26:15Andrew:because there are times that, especially when I'm using like co-work, I've run into this quite a bit now, where it will compact the conversation and I don't want to stop it because I'm at a point where I'm not ready to like end that thread yet. So it would be, it is helpful to just understand like what's the risk of me not stopping that thread moving forward. And yeah, I absolutely love this article, like truly. I feel like you could just take it and feed it into an agentic system and you would have the high-level architecture you need to build most of your core capabilities.

26:50Ben:Like, if you just ask Claude to build this for you. Yeah, this is a canon for sure. As someone who's been doing a lot of stuff on this list, I'm like, wow, I wish I would have had this list a few months ago. This list is cool.

27:02Andrew:Yeah, but I mean, this stuff is changing super rapidly. So, like, yeah, there's a lot of fundamental stuff here that's probably going to be persistent for a while. But at the same time, I imagine this is going to continue to rapidly evolve. So, you know, even though we've gotten a snapshot of what and of the way Anthropic works, you know, we've we've just mentioned how Anthropic moves very fast. They could already have additional layers on top of this that that do far more complex things. All right, Andrew, let's wrap it up with this Vercel security incident. We don't normally cover this type of stuff, but this one was particularly interesting.

27:35Andrew:So why are we covering this one?

27:36Ben:This one was pretty, pretty bad to read about. It was Vercel suffering a security breach earlier this week. I think a lot of folks definitely got emails and notifications around this. Vercel is one of the largest hosting platforms in the world. It's used by a lot of tech companies as well. This type of attack happened through an employee allowing use of a third-party tool through their company account. And then the infiltrators were able to move sideways through Google Access to maybe compromise some Vercel systems. systems. Thankfully, I don't think the surface area of this attack was very big, and Eversell did everything right in this scenario with notifying everybody and rapidly responding to the problem.

28:20Ben:I actually think that this is just another strong signal that just how dangerous it is out there right now on the web. I think that the danger of the web can't be understated. It's at an all-time high in terms of supply chain attacks and infiltrations on systems and machines because there's an inequality between the powers that agents give hackers to the defenders. And that's because it's easy to spin up and parallelize a lot of hostile, you know, infiltrating or otherwise antagonistic activity, right? But it's not so simple to use that same power to proactively and parallelize your defense. And so you're seeing a lot of systems that typically were so hard and that you never would have thought about any kind of breach like this, just falling into these scenarios where they get ensnared through really complex infiltrations.

29:15Ben:This is a employee that allowed access to a third party tool, which then it used as Gmail to move sideways into Vercel systems. Like that's pretty complex in terms of the handoffs and the visibility. And that's the kind of thing that you only get when you have a antagonistic entity out there who can have a hundred or a thousand agents monitoring every single packet that your company sends. So definitely a sign at the scary times. And we've talked about this recently when we had Dan Loring of Chain Guard here on the show. We talked extensively about the supply chain crisis for software. What do you think, Ben?

29:53Andrew:Yeah, well, fortunately, Vertel did indicate that there were no risks to the supply chain that they control. But yeah, this situation is one of my And I was really curious to dig more into this beyond what Vercel said about it. And I went over to Context AI, the company that was sort of at the center of this hack, and found they had a statement as well on this. And they mentioned how last year in June, they released this new AI office suite, which was a new self-service consumer-targeted workspace. This is a company that typically does B2B work. So it was a new type of product for them. And they deprecated this service last month.

30:33Andrew:But as a matter of fact, Vercel was actually, I don't, it sounds like they were never actually a customer of Context. It just appears that one employee went in and enabled a permission for Context's AI agents and allow all permission into the Google workspace. So basically allowing, giving Context an OAuth token that grants complete access to that person's Google workspace account. Yeah. And then Context, I guess, found out they had unauthorized access to their AWS environment last year or last month. And OAuth tokens were included in the things that were accessed as a part of that. But this is actually where I have some deeper questions about this story because context did say they're notifying customers that were impacted by this or users that were impacted.

31:24Andrew:But, you know, this tells me that they're not that these tokens aren't being refreshed on a regular basis. Right. If they're sitting around out there and they still have access to these people's accounts, even long after they've they've been deprecated. You know, there's always a risk that all tokens get out into the wild. So, you know, it's really important that, you know, it's never been more important that standard security practices are followed, like refreshing all of your tokens on a regular basis. But yeah, we really need to solve like this fundamental problem of a lack of sufficient permissions for agentic workflows that just exist across the board, no matter what tool you're using out there.

32:06Andrew:And we really need to solve this before agents are going to be able to fully take over our lives. And in the meantime, we all just need to be just extremely conscious about the permissions we're granting to our AI systems. I'm terrified that I'm going to fall victim to over-granting permissions all the time.

32:26Ben:Seriously. Yeah.

32:28Andrew:And I just attended this really great talk from a friend of mine. So shout out to Justin. He's a frequent listener on the show. and this talk was about the current state of agentic software development and he had this really wonderful reminder that comes from Simon Willison about framing AI risk. In fact I think my friend Justin said that you should tattoo this somewhere visible on your body so you don't forget it. But there's effectively three ingredients to agentic risk. First is that the agent has access to untrusted content. So this could be the internet, an email inbox, anything where an outsider can inject text.

33:06Andrew:Two, the agent has access to private data. So this could be your internal Slack, could be your Google workspace, customer database. And then third, the agent can externally communicate. So he can send email or access APIs or post something to the web. Your goal should be to only ever have one of those at a time, if possible. If you get two of them, It's a risk that can be managed. But if you have three, it's eventually going to be catastrophic at some point. Like it's pretty much guaranteed that it will collapse at some point.

33:36Ben:I completely agree with that. That is a very smart observation by Justin.

33:41Andrew:Yeah. So, you know, I hope we can all practice a bit of a blameless culture and not point too many fingers about why this happened yet. But let's all learn from it as well and understand that, you know, these are real security risks that are emerging. All right, Andrew, what are your agents up to right now? my agents uh well let's see they're actually running marathons this morning because like

34:02Ben:what you said i always feel so like i gotta use the tokens available to me and my session had restarted right before this also i was using beads earlier uh i made a comment earlier this week about how my beads system which is my agentic task management system um has recently kind of flipped from being where the agent puts the stuff it needs its sub agents to do to like the agents put human labeled beads for me to do. So sometimes the agents get blocked for something that they just don't have an ability or access to do because like Justin, I have a lot of separation of concerns. So you get one agent that has a certain capability just trying to ask to do something somewhere else.

34:39Ben:And eventually they can maybe collaborate on this. But in the meantime, I'm at the nexus. So I'll pop in and see what human beads have popped up for me while we've been here chatting. What about you?

34:50Andrew:I knew this day would come, Andrew. April 24th, 2020. the agents start calling on us rather than us calling on them.

34:59Ben:I am the human tool call. Exactly. It's actually really funny because only just two weeks ago when I was on stage at HumanX, were we talking about this exact thing in the panel? It was Angela McNeil of Thread AI who was like, you know, we built our system so the agent can make calls out to the human. And I'm sitting there on the stage thinking, oh, that's so smart. I wish that my agents would do that. And then now here two weeks later and they're doing it. Yeah. Careful what you wish for. Right.

35:28Andrew:But yeah, my agents, you know, the 34th volume of the ThoughtWorks technology radar is out. It's too much for us to cover here on the show, even though we would love to. So don't take our word for it. Go read it yourself. It's a really great guide. You know, the TLDR on this one, AI is forcing engineers to rethink the foundation of their craft. And how do we secure the permissions of hungry agents? the topic that I want to talk about so much all the time. And of course, there's a great shout out to the friend of show, Brigida Buckler, about the concept of harness engineering. So just tying it all together.

36:02Andrew:I love it. So yeah, my agents are going to be connecting to that and letting me have a conversation with it and think about what I can do with the knowledge that I gain from it.

36:11Ben:Well, my agents will be looking for a report from your agents. Yeah, all right.

36:17Andrew:They'll let you know. All right. Well, thanks everyone for joining us for the Friday Deploy presented by Linear B. We'll catch you next week. See you next time.

36:32Andrew:AI is everywhere in software engineering, but most teams still can't prove its impact. That's where the Apex framework comes in. Apex is a new operating model for engineering productivity designed to measure AI where it actually matters at the pull request level. It connects AI activity to delivery outcomes, not just tool usage. Apex is built on four pillars with AI leverage, predictability, efficiency, and developer experience. Apex helps you increase throughput without sacrificing delivery confidence or burning out your team. Because speed without predictability creates chaos and faster coding often shifts bottlenecks downstream.

37:10Andrew:If you want to operationalize AI the right way, Linear D and Apex gives you the system and the cadence to do it. Download the guide and start measuring what matters.

From the publisher

Is the era of cheap, unlimited AI tokens officially over? This week on the Friday Deploy, Andrew and Ben walk through the sudden wave of AI pricing chaos—from GitHub Copilot’s panic-paused signups to Anthropic's confusing pricing tests—and break down the terrifying Vercel security breach caused by a single over-permissioned AI tool. They also examine 12 game-changing architecture patterns exposed in the Claude Code leak to help you safely orchestrate your own agentic workflows. Finally, they discuss how to avoid the lethal trifecta of agentic security risks before mourning the tragic deletion of their Claude Code buddies.

Read the guide: The APEX Framework

Follow the show:

Follow the hosts:

Follow today's stories:

OFFERS

  • Start Free Trial: Get started with LinearB's AI productivity platform for free.
  • Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.

LEARN ABOUT LINEARB

  • AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
  • AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
  • AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
  • MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.

More from Dev Interrupted

All 208 episodes
The harness is the showdown, the humans are the tool calls, and have you seen my Claude Code buddy?Dev Interrupted · 37 min
Listen in VO