In short
The episode covers “life beyond token maxing,” arguing that once engineers have enough agent tokens to create value, companies should distribute those gains across the organization. Uber is highlighted: it’s sending “rearward deployed engineers” from engineering to other teams (legal, marketing, sales) to help them capture productivity improvements, rather than keeping agent benefits siloed.
Key claims
guardrails matter as agents get more autonomy; Anthropic’s Claude CloudCode now defaults to “auto mode,” routing tool calls through a classifier that blocks irreversible/destructive/out-of-environment actions, reducing prompt-injection risk; manual review is largely habitual (~97% auto-approved). Meta’s “Muse Glimmer” is an open, 30B parameter on-device agent model (Apache 2), trained via distillation, with long-context benchmarks and speculative decoding. It also discusses why open source matters (modular ecosystem) and “context engineering” for SDLC (agents as “citizens,” not tourists). Research: personalized coding-agent skills help only marginally; generalized skills pooled across developers perform best.
Guests
none named in this transcript (hosts Ben Lloyd Pearson and Andrew Ziegler only).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VORearward Deployed Engineers at Uber
0:45 to 3:30
Discussion on Uber's innovative use of engineers in various departments.
“within Uber, out to legal, into marketing, into sales.”
Organizational Gains from Engineering
3:30 to 4:40
Exploring how distributing engineering productivity can enhance different departments.
“But anyways, welcome to the Friday Deploy brought to you by The Near B.”
CloudCode's Auto Mode Update by Anthropic
4:40 to 7:00
Introduction of CloudCode's auto mode as a default feature to enhance safety.
“a machine is a really important guardrail and restriction for working with the tool.”
Permissions and Safeguards in AI
7:00 to 8:10
Discussion on the importance of permissions and safeguards in AI models.
“And when it actually executes those capabilities, what is it going to be doing it securely?”
Hugging Face Hack's Impact
8:10 to 9:18
Exploration of the implications of the Hugging Face hack on AI safeguards.
“There are times where I see it saying it's going to make an API request and I only skim what it's doing and don't actually really investigate and look closely at what it's doing.”
Meta's Muse Glimmer Announcement
9:18 to 9:51
Introduction of Meta's Muse Glimmer, an open-source agentic model.
“Like it actually looked like they could be potential bad actors themselves.”
Technical Innovations in Muse Glimmer
9:51 to 11:48
Discussion on the technological advancements and potential applications of Muse Glimmer.
“This is a pretty cool announcement, I feel like.”
The Future of Open Source in AI
11:48 to 14:00
Exploration of the significance of open source in the evolving AI landscape.
“I'm excited to see what people build with it.”
Technological Innovations in Local Models
14:00 to 14:55
Discover the incremental improvements, like speculative decoding, in local AI models.
“And then those chunks are sort of validated after the fact by a smarter or bigger model to confirm their validity.”
The Importance of Open Source in AI
14:55 to 17:28
Learn why open source models are crucial to prevent monopolization by a few firms.
“Like, can we own the servers and the client side and just try to monopolize all of it?”
Show all 18 chapters
Modularity and Composability in AI
17:28 to 19:41
Explore the need for modular AI systems that promote innovation and collaboration.
“So what did you think about this article, Andrew?”
Context Engineering in Software Development
19:41 to 22:58
Understand how context layers enhance decision-making in the software development lifecycle.
“preventing all of us from fracturing and just writing our own versions with these models.”
Agentic Systems and Context in SDLC
22:58 to 26:55
Learn about the role of agentic systems in software development and the need for context.
“kind of, it's a very broad article, it covers a lot of topics.”
Personalized Skills for Coding Agents
26:55 to 28:00
Examine research on the effectiveness of personalized versus general skills for coding agents.
“But then it also gives those folks and their agents the real-time relational context about what the heck is moving around them.”
Exploring Productivity Gains in Development
28:00 to 29:20
Learn about the impact of personalized versus generic skills on productivity in development teams.
“And there's actually kind of some surprising findings.”
Distributing Skills for Organizational Gains
29:20 to 31:28
Discover the importance of distributing skills across teams for greater operational efficiency.
“And frankly, I think a lot of like agentic operators do.”
Current Trends in AI Utilization
31:28 to 32:49
Understand how AI tools are being utilized and the importance of token management.
“Well, Angie, what are your agents up to this week?”
Building New Skills in AI Development
32:49 to 33:35
Learn about the opportunities for building new skills with emerging AI models.
“It's like there's so many opportunities to just be like, hey, that takes too much time.”
Transcript
Automatic transcript. May contain errors.0:04Andrew:So, Andrew, we've got a new life beyond token maxing story yet again for this week. I think that's two weeks in a row now. But it's a different kind of life beyond token maxing story than what we're typically used to. We're used to people revolting against token maxing and then putting all these restrictions on it. But here we have Uber, who's back in the news. who's known as one of the organizations that token maxed earlier this year. There's CTOs out there sharing some new ideas on how they're deploying their agents, their engineers tokens now. And it's actually a really interesting idea. Like they're now sending their engineers out into other teams within the company, within Uber, out to legal, into marketing, into sales.
0:50Andrew:And they're calling them rearward or rear deployed engineers rather than like a forward deployed engineer. It's kind of interesting. We used to call it just like inner source. But yeah, it's a pretty interesting idea of like using your engineers to, you know, you got all these productivity gains coming into your engineering organization. Why not have them to just help people on other sides of the company capture those gains? I don't know. What do you think when you read this story, Andrew?
1:15Ben:Well, I just thought that we really needed to workshop rearward deployed engineer maybe before going forward with that one or going backwards with that one in this case. So this was an interesting turn of events. You're right that we've had been a little bit of a streak lately with week after week. There's been some major company we've talked about or covered here indulging in the token maxing phenomenon and then dramatically shifting course, either reversing it and going for like a very minimizing or cost restrictive approach or in this case, realizing the more substantial opportunity. and that is that if you have engineers that are able to access and work with that amount of tokens and deliver that much amount of value maybe you're still figuring out what that value is but if it's clearly there in some form let's distribute it let's figure out how to bring these gains into other departments that aren't as enabled that don't have these like technical thinkers and frankly this is the same idea of like hiring an agency that come in and do like your ai transformation except in this case you're enlisting your smartest and most like natively and familiar folks for your ecosystem.
2:22Ben:So it's a really smart play. I think this is how you actually kind of get the organizational gains. And that's going to be a theme, I think, actually, across all the stuff we talk about in today's episode is how do you distribute the gains to work on an organizational level? And what does that mean? I think that's the big challenge. So it's exciting to see a leader like Uber really take to the charge on it.
2:43Andrew:Yeah, and speaking of this theme, I mean, we've heard similar stories from LinearBee customers where they've been this agentic leader within their organization. They've enabled the entire engineering organization to leverage agents who move faster than ever before. And then the first question they get is like, hey, can you go help other teams learn how to do that too outside of engineering? So I kind of think that this probably won't be common practice for too long. I think it's mostly gonna be like larger enterprises where you see this sort of behavior pop up or companies that are just really, really far ahead on the agentic curve.
3:17Andrew:because at the end of the day, eventually people are going to build tools for those other teams to solve their work agentically. But like I said, this is a different kind of life after token maxing that I'm totally here for. And just shameless plug, if you haven't listened to it yet, we had Andrew and I hosted a workshop with Linear B a couple of months back where we talked about this concept of token maxing and what it means to get past that and what life looks like once you've sort of moved beyond just looking at raw adoption and AI usage and start to think about where's that impact actually hitting your organization.
3:51Andrew:So yeah, cool little story. But anyways, welcome to the Friday Deploy brought to you by The Near B. I'm your host, Ben Lloyd Pearson. And I'm your host, Andrew Ziegler. And this week, in addition to token maxing, we are covering model updates from Claude and Meta, why open source still matters, context engineering in your SDLC, and some research that answers the question, do personalized skills help coding agents? So let's just start right at the top with this news from Anthropic. What's all this about CloudCode now being defaulted to auto mode?
4:27Ben:All right. So CloudCode, they're talking about auto mode here. Auto mode is something that Anthropic has very famously been tinkering with since really CloudCode hit the scene. The idea that the model could check its own permissions on commands that it runs on your a machine is a really important guardrail and restriction for working with the tool. And there's been all different layers in this discussion about where should that guardrail live, who should be in charge of it. But Anthropic has really taken the mantle here in making these very safe evaluations a premier front part experience of using Cloud Code.
4:59Ben:So auto mode has been upgraded to being a default permission setting. Up until now, it was a experimental setting that you had to turn on. And the idea is that it catches pretty much any kind of red flag command that would typically be a potentially destructive or harmful one. I think these are kinds of rewards they've earned out of all of the work they've had to do in controlling and maintaining the mythos and fable rollouts around their extreme cybersecurity abilities. They just now have so much intelligence around how to construct and create these boundaries that it makes sense they're going to make this a default.
5:37Ben:I mean, folks, like anybody who's already approving their prompts manually in Cloud Code, they approve 97 % of them just automatically anyways. This is actually a better guardrail in many of those cases because it's paying a little bit of a closer attention. They also worked with some third-party red teams to try and do prompt injection on it, did a whole bunch of testing. and was able to stamp out harmful commands that had popped up in previous generations. So really promising frontier research coming from Anthropic about how to protect a model from harming your system. Definitely a really critical part of the model hosting infrastructure, especially as you move that stuff onto your own systems.
6:20Andrew:Yeah, you know, I think overall this change is probably a net benefit. Like, you know, I expect there's probably a lot of like pessimists out there that might be looking at the potential downsides to this, where, you know, Anthropic could use this as a way to just route all of your prompts to cheaper models to save them on costs, which is probably true. That's probably going to happen to some degree. But I think there's a real upside here too. And, you know, Anthropic can also choose when to do things like implement longer thinking horizons or to use reasoning or to maybe use a better model instead of a cheaper model.
6:52Andrew:You know, and when I'm using Claude, You know, I don't always want to have to be forced to think about whether or not I need all of these different capabilities. And when it actually executes those capabilities, what is it going to be doing it securely? So, yeah, there's a couple of quotes that I wanted to call out specifically from this article. The first is about how this tool or this ability routes each tool call through a classifier that targets blocking actions that are irreversible, destructive or aimed outside your environment. They have these new safeguards in place that, you know, one of the biggest things that I'm always paranoid about when I'm when I try to start giving Claude a little more freedom to take action on my tools is am I comfortable with the permissions and the changes that it's making to the things that I'm connecting to it.
7:40Andrew:And as you mentioned, the other thing that stood out, you know, they have data that suggests that manual reviews are just habitual for most users. You mentioned the number 97%. So that means 97 % of these tool call permissions that CloudCode asks for just get approved, which seems to indicate that people are just clicking through reflexively rather than reviewing every command. And I try to be conscious about what permissions I'm giving it. But I'm not going to lie. There are times where I see it saying it's going to make an API request and I only skim what it's doing and don't actually really investigate and look closely at what it's doing.
8:20Andrew:But I've been preaching on this show for a while now that these foundational AI tools need better protections and safeguards like this that help protect us from our AI going rogue. So I think it's really great to see that, you know, Anthropic is continuing to think about this challenge and is building tooling. You know, I will say probably my biggest concern and where, you know, I do align a little bit with some of the pessimists out there is that we really do need the ability to turn things like this off when we need to and be able to manually configure as much as possible. And I keep thinking about this Hugging Face open AI hack that just happened.
8:57Andrew:And it really kind of is the perfect example of how AI safeguards can actually backfire on you. In this example, Hugging Face was unable to use some of the Mythos class models for responding to the security breach in real time because those models couldn't determine if the people trying to protect themselves were being malicious or not. Like it actually looked like they could be potential bad actors themselves. So it would refuse to do things for them. And then the team had to switch to other models that just don't have those safeguards so that they could respond to this security incident.
9:31Ben:But yeah, there's a lot of benefits that Anthropica is claiming here.
9:34Andrew:It supposedly reduces the risk of prompt injection, harmful actions. So yeah, I think this is just sort of like the next step in what is likely to be a constant iteration of better safeguards around AI. Yeah. All right, let's talk about the latest news from Meta. This is a pretty cool announcement, I feel like. Muse Glimmer. It's an open agentic model that runs on your devices. What is this, Andrew?
10:00Ben:Yeah, this is, like you described, it's a long-running agent coming from the Meta Lab. So this is an open source, open weight model. If you're familiar with the models that they release, they do so fully open source because they want a fully collaborative ecosystem. And this latest one is a 30 billion parameter open weight model under an Apache 2 license, which we've talked about since around April this year. That's becoming the really popular trend ever since we're on Gemma 4 of this type of a license, which means that you can fine-tune, train this, make your own custom private model and sell services off of it.
10:34Ben:And there's no cloud dependency required either. The idea is that it can run on local hardware. It can even be on consumer-grade GPUs. And I myself, I haven't had a chance to tinker with it yet, but I'm definitely very curious to give it a try on some of the machines I have around. And I do think that owning your inference and having this long-running agent is a really powerful and useful tool for folks, especially because this one is more focused on doing tasks and is not like a coding agent, right? So this is a really great candidate for if you have hardware and you want to run a long-running agentic assistant, especially one that lives on device or works with sensitive or private data.
11:13Ben:This becomes a really great candidate for that kind of world, of course. if you're using a model like this, you have to bring your own everything, including like your harness and the environment it's going to work in. And this comes back to what we just talked about a moment ago of things like guardrails. You know, if you're going to use Muse and you're going to ultimately have it working or operating on tasks, you have to think about the guardrails you have to bring to the system to make sure it operates safely within your environment. Of course, this is all just baseline stuff for working with any model, but particularly important when you get long running ones that live on your own device.
11:47Ben:Really cool development. I'm excited to see what people build with it. What do you think of the latest developments?
11:54Andrew:Yeah, well, first of all, I'll love to hear what you think after you get it into your lab and dissect it and benchmark it and see what your agents think about it. But, you know, also I would love to hear, you know, friend of show, Brigida Boeckler. I would love to hear her opinion on this too, because I know we just recently covered some research she's been doing around the viability of local models. And at the time, her conclusion was that they still needed some time to develop. Like they weren't quite it yet. But there was definitely potential that seems like it's on the immediate horizon. So that's pretty cool.
12:24Andrew:I'm really looking forward to, I know she's probably out there already thinking about this.
12:29Ben:A cool note for this one too, is that this one was mostly trained via distillation from a larger teacher model. That's a really important note I want to call out for our listeners because that's the trend for all of these open source models is that you get these loops where they're trained or created from synthetic data that's constructed by a smarter model or by like a more frontier model. So right now, a lot of the benefits we get in the open source world are just coming off of like the comet trails, right, of these foundation models in a sense.
13:00Andrew:Yeah, yeah. And my impression that this model specifically is sort of part of Meta's goal to attach AI to your desktop work environment. So we've been hearing all these stories about how Meta's tracking their employees' computer usage and using it to train some new models. So I imagine that that's really what has been used to sort of get this model to where it is today or one of the many things that it used. So yeah, I'll point out some things that really stood out to me on the technical front. So the first thing I noticed was that in the benchmarks, scored exceptionally well at this AA LCR benchmark, which measures a model's ability to reason about and to synthesize information from long form documents.
13:46Andrew:So we're talking like 10 to 100 ,000 tokens, which is, you know, that's a pretty notable achievement because that is still a thing that a lot of models struggle with, particularly other local models. But then really what's notable is that you know they it only requires 20 gigabytes of memory to run you know it seems like it's a that's like exceptionally low but then there were some new developments in this that were also just interesting as like technological incremental improvements like they have this concept called speculative decoding which is where they have this like super lightweight model that generates the output tokens in larger chunks rather than sequentially which is how a lot of models do it today.
14:27Andrew:And then those chunks are sort of validated after the fact by a smarter or bigger model to confirm their validity. So yeah, lots of just really cool, like incremental innovations out of this that are really making the local model space really seem like it's starting to heat up. And I think the coming months are going to be, like a lot of attention is going to be not just on this, but on all of the developments happening to local models. I agree. Yeah. And speaking of open source, let's talk a bit about why open source matters for AI. So we got an article here from the O 'Reilly Substack from Tim O 'Reilly bringing some just really great sage advice for the AI era from somebody who has been around for a lot of major technical or technological developments.
15:16Andrew:And this article argues how the open source models, the weightings, the harnesses, the context layers, the data that goes into this, it's really important that we do have effective open source competitors in this space because there's a lot of risk in relying too heavily on a small number of proprietary firms to provide these types of services. So there's a really great allegory in here to, you know, the early days of the web when you had Netscape and Microsoft who were sort of duking it out and trying to figure out, like, how do we own the entire tech stack of the web? Like, can we own the servers and the client side and just try to monopolize all of it?
15:59Andrew:And, you know, for a while, it seemed like that may actually play out. But then you had something like Apache hit the scene that, you know, quickly followed with like the LAMP stack becoming the norm. And suddenly everything is open source and nobody fully owns the tech stack of the web. And, you know, some of the risks that O 'Reilly highlights in this article that I think are really worth, you know, just paying attention to is that, you know, if we're increasingly relying on like one or two or three frontier labs to define like things like personality traits. and the guardrails that go into these, there's a risk of everything sort of like going towards the lowest common denominator.
16:39Andrew:You know, we all start to become the same with the same outputs and the same approaches to solving problems. And there's a real risk that that sort of stifles a lot of creative innovation, for lack of a better phrase. When you think about how like engineering teams are tinkering with a lot of AI today, like if you're using the frontier stuff, you're really doing more around customizing your harnesses rather than customizing like the weights of the model, for example. And you may actually like both of those may be important things to focus on. So yeah, I just really like this as, you know, somebody who's been in this industry for a very long time and has been successful.
17:21Andrew:And, you know, and he's a big friend. O 'Reilly is a huge friend of open source and long has been. It's just really great to hear his perspective on how AI is shaping things. So what did you think about this article, Andrew?
17:32Ben:I thought it was really smart how the article calls out that where the seams of open source technology is used for models. Up until now, there's been a lot of controversy around what open source even means for a model. Like, oh, you give us the model. That's great. But did you give us the training code? Did you give us the weights? Did you give us the corpus that was used to train it. All of those things have their checkboxes and people use them to grade if something's open source. But in this article from O 'Reilly, he's really focusing on how the thing we need to focus on making open source is the modularity of the stack that all of the AI stuff is operating on.
18:14Ben:The inference platforms, all of the tooling that we use to serve and store the data for them, making them as composable and modular and open as possible is actually the keys for letting all of the rest of that thrive because to your exact point um like you you get in this situation where different large players own really critical parts of just like the baseline experience of using the model like think of what we've covered so far we talked about anthropic really becoming a really having this you know a grip on the on the classifying the dangerous commands and putting guardrails on prompting and stuff like if you move into an open source world you don't have that anymore.
18:55Ben:So we have to think about what are the parts that give us the equivalents. And that's where this really cool AI potluck initiative that he calls out comes from. The idea of how do you build this very rich ecosystem? Think Linux foundation level rich ecosystem of all of the parts you need to run a cloud and making it open source. We need that equivalent for AI. And that's what the potluck is. That's a pretty cool initiative. And we'll link it for folks to check out. I think that this is a really important part that we get right. And so far, I think we are. I think of the major tooling and the parts that I use to run my open source models or my harnesses.
19:30Ben:And, you know, I feel like I have the parts I need to at least start assembling together. But that's where I think a lot of the challenges will live is supporting that ecosystem and preventing all of us from fracturing and just writing our own versions with these models.
19:46Andrew:There's a lot in this article that I really strongly agree with. It's worth giving a read. But the tone in it kind of made me feel almost like, and we talked about this before recording, it was almost like a very cautionary and sobering take, you know, almost cold, like a viewpoint on this perspective. I don't want to say pessimism, but it was almost feeling like it was bordering on having a pessimistic take about it. I really have sort of a very different perspective on it. I think that, yeah, open source is being very heavily disrupted today because of AI. And in some ways, it is sort of following the lead of these proprietary companies when it comes to the frontier of AI rather than being what's leading it.
20:26Andrew:But I also think that we're sort of primed for a bit of like an open source renaissance, actually, because, you know, you mentioned distillation earlier with Meta's new model. You know, AI has made it easier than ever to replicate and iterate on other people's ideas. And, you know, I think open source is just going to continue very closely tracing the capabilities of the frontier model companies, particularly considering, again, how easy it is to distill value out of stuff that exists. You know, we saw Claude, some of their code and architecture get released or leaked earlier this year. And immediately everyone's out there with like their own versions of how they've, you know, they've distilled it into something else.
21:05Andrew:So, you know, it's a very, very exciting time. But again, a great take from O 'Reilly on the state of things.
21:11Ben:Your SDLC looks more like a software factory every day.
21:15Andrew:How do you get ahead of that transformation? And how do you prove what it cost and what it delivered? On August 27th, DevInterrupted hosts a live roundtable on this very topic, proving AI ROI from software factories. To learn, we've invited two industry experts and past DevInterrupted guests. It's Dex Horthy of HumanLayer, who ran a fully automated factory and then shut it down. And Zach Lloyd of Warp, who publishes frequently about how he measures what his factory pays for itself. Linear B co-founder Dan Lyons will join them to discuss the power of the context layer that will make all of this possible.
Read the full transcript
21:51Andrew:Save your seat on Luma. All right, let's talk now about your STLC and context engineering. So here we have an article from our friends over at LeadDev where it really just breaks down how there's so much context scattered across all of your STLC that really needs to be pulled into your agentic systems for them to make good decisions. So there's a lot of context that exists in the way that work is specified and how it's reviewed, where and how it's tested, and, you know, the realities of it being shipped. And you really need to be accounting for all of that when you're trying to build an agentic system to contribute code into production.
22:33Andrew:So you need to think about stuff like documenting all the states of your life cycle. You know, what's the current state of your project? What's the future plans for your project? What's permanent and should never change without like a great deal of focus? And really understanding all of those things and giving them to your agents when they need it within the SD DLC. And this article, you know, kind of, it's a very broad article, it covers a lot of topics. It's kind of a little bit too much that we can cover in length here. But you know, I wanted to bring it up because this idea of a context layer for your SDLC is a topic that we're continuing to see come up like over and over again right now.
23:15Andrew:And when you really think about it, you know, early agents really just had access to like your code base, maybe they had some basic metadata, like your project management tasks or PR descriptions. But I think most of us learned like very early on that that wasn't enough context for most engineering decisions. There were some places where it was enough, but many places where it was not. And, you know, there's a lot more that goes into making good engineering decisions. So, you know, you have to think about like how well planned is your spec? You know, what learnings and decisions did you make along the way that modifies your final outputs?
23:54Andrew:What parts of the code base have risky components that need particular care or review when modifying? You know, which components get bogged down in reviews the most? Like these are all questions that you have to wait when you're initiating a project or task because they all impact like the feasibility of of completing them and you know humans you know we accumulate answers to these through experience but your agent only has as much experience as you feed into it like we covered this concept a while back of where ai agents are like tourists they it's like the first time they've ever every time they show up it's the first time they've ever been there and they only have as much context as they deliver, as you deliver to them as they go along their journey.
24:37Andrew:And then once they're done, they leave forever and go back home. But that's, you know, that's really like, you know, and this is something that, you know, linear B, like we've really have started to ingrain this into the, to what we're building for our customers. You know, we're really trying to help understand, like, how do you, how do you get that context layer into all of your, your agentic decisions, you know, and that's whether it's humans or agents, you know, if a human is making decisions, they should have a format of all of this context that works for them as well. So yeah, we've been doing this a lot with customers recently.
25:09Andrew:And I just think it's a really fascinating concept that really every engineering leader needs to be solving right now. So Andrew, what did you think about it?
25:17Ben:You had a really great coverage on what this article gave us and definitely dove into different parts of what matters for a team working on a code base together. I think that's a real big focus is a lot of times in agentic coding you get these two different groups that are talking to each other you have like the vibe coder or someone with a project that they are the only person touching anything on it and they can go really fast and they don't have these kinds of guardrails and you folks that are using it as a team and trying to build a product together and that's like a team sport and so like they need to have a lot more context sharing and a lot of the techniques that give one velocity would just totally you know wreck another one and so that nuance is really important.
25:58Ben:And this one gives us that secondary path, the idea of like, what does this look like on a team level? So the article gives us some really great tactics here. I think your point about, you know, the agents are almost like tourists and they come in and they do their work and their leave. A lot of this is around how do you turn those agents into more like citizens of this code base? Like they are native to it, fluent in it. They understand the parts that are needed to operate just as part of their operations themselves. And that's what's achieved by having these kinds of layers of context, as he describes here.
26:33Ben:I will say I wanted to go, you know, like on a fridge, like the word magnets, you can rearrange. I really wanted to do that with the title because I'm like, your SDLC is your context engineering is not saying what this article says to me. And it's really more about context engineering for your SDLC or, you know, how to put a harness on your SDLC. Because that's really what this is unlocking is it gives you as like an operational or a leader person the vantage into this code world that's shared by a bunch of folks. But then it also gives those folks and their agents the real-time relational context about what the heck is moving around them.
27:13Ben:And I mean, gosh, we talk about that all day here on Dev Interrupted because that's our story at Linear B. So really resonated with this article. Yeah.
27:21Andrew:And speaking of context for your agents, let's talk about personalized skills and some new research that came across our desk on whether or not they help coding agents. So this is a new study we read around the concept of personalized skills. and these are preferences that i'll learn from an individual developers past interactions with ai coding agents and the research wanted to find out if those were better for um coding agents than more generalized skills so they ran this test uh across 13 developers with over 200 real world developer agent sessions so it's a little bit of a limited sample size but enough to at least get some ideas about what might work.
28:06Andrew:And there's actually kind of some surprising findings. While the personalized gains did, or the personalized skills did provide some productivity gains, it was actually like kind of marginal, particularly when you looked at the impact of more generic skills that are pooled across many developers. Those are the ones that have, that produced the largest and the most consistent improvements. And I feel like there's been an ebb and a flow over the last two years or so where people, you know, think that maybe AI needs to, you know, be heavily customized to the individual situation or person or, or workflow, um, or does it need to be more broadly, you know, constrained to, to have like wider practices.
28:53Andrew:And, you know, this research seems to indicate that particularly when you're, you're trying to look at like, what's going to have the biggest impact on an organization, it's probably better to have skills that that help everyone a little bit rather than to have one person help one person a lot in just a few situations so yeah it's really great research to read for anyone that you know is out there buying or thinking about ai coding tools or building them as a way to sort of prioritize like where should you be investing your time for your team so what did you think about this
29:25Ben:research andrew this was a really great for me a revisit back to when we had karthik ramka Paul, the distinguished engineer from LinkedIn on the show, talking about how to distribute gains across the entire engineering department. We've had a lot of big leaders on the show, established enterprises, not small engineering teams, manage to get these kinds of operational gains and then also get those gains across adjacent departments, marketing and finance and HR. And how it all comes down to is having this distribution system, a place where the experiments can live together, a place where learnings like memories and skills can live together, be version controlled and distributed amongst others.
30:11Ben:And so that's the really big takeaway from this is that, you know, I felt like this article was coming for me a little bit because I have lots of personalized skills based on my own tastes for working with all sorts of stuff. And frankly, I think a lot of like agentic operators do. You just kind of accrue them over time. And this article really calls out that maybe the benefits of those are marginal or not as much as you think. I argue maybe there's a compounding effect of using a lot of those together to create, do a very domain-specific task. But maybe that argues that there's more for simplification to be done, which I agree with.
30:43Ben:But the biggest thing here is that those gains are marginal. If you distribute those same kinds of learnings, but on a general level, and to everybody, the gains and productivity across the board are just substantial. substantially larger than they can be for the individual. This really speaks to the power of us pulling together what works, especially within an organization. So, and it's really promising data that like, this is a real trend now captured in this research. So really great dive. I recommend our folks check it out when we include it. But if you haven't thought about how you're distributing skills within your team, whether it's a engineering team or product team or otherwise, like I think that is your immediate next opportunity.
31:29Ben:Absolutely.
31:30Andrew:Absolutely.
31:32Ben:Well, Angie, what are your agents up to this week? Well, they're being cautious because I'm basically out of tokens. And they have this concept of throttling themselves when that happens. So I'm getting caveman talk again, unfortunately. We're back in caveman days. But thankfully, I'm only a few hours away from a refresh. So then we'll be back. I guess they'll be talking like Shakespeare or Homer or whatever they feel like. What about your agents? What are they up to?
31:57Andrew:I feel like you need to get a little bit of sub-agent delegation and model routing into your life from the sounds.
32:02Ben:Well, I just need to take some time to try out the new Glimmer and kick some things over to that. I will say I burned my tokens doing a lot of really cool stuff this week, like working on a lot of news reading stuff. I love to read lots of stuff, but as you know, the world's moving way too fast. So we talked about an article in here recently about like there's too much information to consume it all. Like how are we reading? How are we writing? and it talked about this like layer of of like understanding what's going around in the world and the curating what matters for you so i've been trying to explore that and build that you
32:34Andrew:know using a lot of tokens in the process yeah what about you we just talked about generalized skills i i've i you know with these mythos class models now all out you know we got we got all of them sonnet opus fable i can kind of pick and choose now i've decided to kind of go back and like look at some of the stuff that that we've built in the past and how it's you know a lot of it hasn't has it it needs to be it needed to be iterated on to keep up with uh where the models are today so uh yeah definitely some uh like generalized skills happening but then also just like it's a good time to build as you mentioned it's a great time to build new skills because these mythos class uh models are you know the opus 5 fable 5 they are really great at constructing these things so um yeah mostly just trying to get toil out of the way you know that's really where we're at right now.
33:24Andrew:It's like there's so many opportunities to just be like, hey, that takes too much time. I'm going to go have AI do it for me now.
33:31Ben:Every week, it gets a little easier too. Yeah.
33:34Andrew:Well, speaking of every week, thanks again to our listeners for joining us again this week for the Friday Deploy presented to you by Linear B. It's always a pleasure to get to share our opinions and our learnings on what's happening in the space of AI and agentic development. So thanks for joining us again today. Engage with us wherever you find us out on social media, We're out on LinkedIn. We're on Substack. You can see us on YouTube as well. Leave a comment, a like, a thumbs up, whatever you can do to help us. It really does just help spread the word of the show. So thanks for sticking around to the end and we'll see you next week.
34:06Andrew:See you next time.
From the publisher
This week on the Friday Deploy, Ben and Andrew explore Uber's strategy of "rearward deploying" engineers to spread agentic AI workflows into departments like legal and marketing. They also dive into Anthropic making Claude Code's Auto Mode the default, Meta's new on-device Muse Glimmer model, and Tim O'Reilly's case for an open source AI ecosystem. Finally, they break down context engineering for the SDLC and examine new research showing why generalized agent skills outperform personalized ones.
Register: Dev Interrupted Presents: The Software Factory Roundtable
Follow the show:
- Subscribe to our Substack
- Follow us on LinkedIn
- Subscribe to our YouTube Channel
Follow the hosts:
Follow today's stories:
- After starting the tokenmaxxing panic, Uber's CTO is back with a very different AI story
- Auto mode is now the default in Claude Code for Pro, Max, and Team plans
- Introducing Muse Glimmer: An Open Agentic Model That Runs on Your Device
- Why Open Source Matters for AI
- Your SDLC is your context engineering
- Do personalized skills help coding agents? An empirical study of developer interaction histories
OFFERS
- Start Free Trial: Get started with LinearB's AI productivity platform for free.
- Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.
LEARN ABOUT LINEARB
- AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
- AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
- AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
- MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.
