Model welfare, building a civilization for agents, and the CI/CD landrush

7 Aug 2026 · 39 min · 12 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode topic: Post-token-maxing AI for agentic software engineering: using AI in the Socratic method, building “AI civilizations” with model welfare, agent harnesses, and SDLC changes like “CI/CD landrush” and disappearing code review; plus research on AI governance in open source and AI’s impact on mathematics.

Guests (named in episode)

Andrew Ziegler and Ben Lloyd Pearson (hosts). Steve Yegge (author discussed; not a guest). Dex Horthy (Human Layer) and Zach Lloyd (Warp) are mentioned as future roundtable guests (Aug 27), not in this episode.

Guest backgrounds

Yegge is known for Gastown and building the MMO Wyvern; Dex Horthy ran and shut down a fully automated software factory; Zach Lloyd publishes on measuring factory ROI.

Key claims

Socratic questioning improves agentic loops via adversarial back pressure; high merge volume may eliminate pre-merge CI/CD by batching PRs and having agents fix “broken main”; durable agent “seats” plus “laurels” improve performance (“model welfare”); code reviews may become mostly automated; open-source AI policies correlate with higher engagement/satisfaction; math communities face “implosion” as AI proves faster than peer review.

Notable examples

Microsoft moving beyond token-maxing; Yegge’s Wheelhouse in Emacs building Wyvern; “land rush” merging directly into main then agent swarm fixes; OpenClaw harness (~147 lines, four tools); GPT solving a proof after ~1 hour; <2% of GitHub open-source projects having AI policies.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring the Socratic Method with AI

1:26 to 3:08

Discussion on using AI in the Socratic method for knowledge exploration.

“And this week we are talking about using AI like Socrates again, again, building AI civilizations on Maine, model welfare for machines, and why every mathematician you know might be having a meltdown right now.”

AI and Agentic Workflows

3:08 to 5:34

Importance of integrating AI into agentic workflows and the concept of loops.

“These are the kinds of practices that become like primitives for being successful and knowledge working in the Socratic practice is no different.”

Steve Yegge's Insights on Software Development

5:34 to 9:39

Discussions on Steve Yegge's articles about new software development paradigms.

“selection and better use of back pressure and loops and human in the loop feedback and such.”

The Land Rush in CI/CD Processes

9:39 to 14:00

Exploration of how increasing PR volumes affect CI/CD and code review practices.

“And this is a world where CICD no longer can happen because the volume of activity and PRs and just things moving through your Git are so staggeringly high.”

The Evolution of Code Reviews

14:00 to 16:44

Learn how code reviews are expected to evolve with automation and real-time analysis.

“to rely on that will be changing very dramatically soon.”

Philosophical Changes in Agent Development

16:44 to 19:58

Explore the philosophical implications of how developers engage with AI agents over time.

“And so Yege had a follow up to this article we've been discussing too.”

Understanding Agent Architectures

19:58 to 23:20

Discover how to effectively build and categorize AI agent architectures for various tasks.

“I mean, it really seems like the struggle of the moment is, it's like context delegation.”

AI Integration in Developer Experience

23:20 to 25:16

Examine the impact of AI policies on open-source contributions and developer experience.

“So, Andrew, what did you think about this article?”

Live Roundtable Announcement

28:01 to 28:50

Learn about an upcoming live roundtable discussing AI ROI from software factories.

“So I think it's great to highlight research like this.”

AI's Impact on Mathematics

28:51 to 31:06

Explore how AI is changing the landscape of mathematics and its scholarly community.

“This is a fascinating glimpse into a world I'm not super close to.”
Show all 12 chapters

Evolving Roles in Knowledge Work

31:07 to 35:56

Discuss the transformation of roles in knowledge work due to AI and its implications.

“And that really stood out to me because, you know, this is exactly how I feel about the way that AI is impacting engineering.”

Audience Engagement on Junior Engineers

35:57 to 36:28

The hosts invite listeners to share experiences with onboarding junior engineers.

“I feel like there's potentially a similar challenge happening in a lot of knowledge work.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:04Andrew:All right, so I guess we got to chalk Microsoft up as another yet another life beyond token maxing company. You know, they've seen the light. They tried it out, saw the budgets, and now they too have moved beyond the token maxing culture. What do you think, Andrew?

0:22Ben:I mean, Microsoft reversing course on the whole token maxing thing was only going to take about one budget, one bill cycle to actually come to a head, in my opinion. But yeah, you're right. I think we need like a tally board or something for like how many companies we've had. We've talked about adopting this phenomenon only to turn around and vehemently reject it or it'll cause some sort of turmoil. So this is your reminder, token maxing and that whole phenomenon is not the strategy to scale AI success on your team. And Microsoft joins the companies that have awoken with that clarity.

0:58Andrew:Yeah, and of course, if you're out there at Microsoft wondering what does life beyond token maxing look like? Well, of course, we have a workshop for you. We covered this over at Linear B recently. So make sure you go check that out. We tell you everything you need to know about how to live this post-token maxing life where you now have to justify your AI budgets. Because, yeah, we help teams do that every day. And we bring content like that to you every week. And this is the Friday Deploy brought to you by Linear B. I'm your host, Ben Lloyd Pearson.

1:28Ben:And I'm your host, Andrew Ziegler.

1:30Andrew:And this week we are talking about using AI like Socrates again, again, building AI civilizations on Maine, model welfare for machines, and why every mathematician you know might be having a meltdown right now. So before we get into that, that may be a distressing topic, let's talk about Socrates, Andrew. Again, let's talk about how to use AI in the Socratic method. What story do we have here today?

1:59Ben:Yeah, so we have an article that explores the phenomenon of using AI like a Socratic partner to explore ideas. Now, if you're a listener of Dev Interrupted, you might be thinking, haven't we talked about this? Haven't we covered this ground? And that's actually why we wanted to bring it up again here today. Because there's an interesting phenomenon that keeps happening where we as a culture and as an industry discover and find these interesting things that work and we share them. But we reinvent things in a silo. So it's been really amusing for me as somebody like with a classicist background who studied Socrates to watch like tech bros rediscover Socrates in like 2026.

2:38Ben:It's awesome because it is a universal way of thinking. And it just speaks to why the methodology is so powerful and has survived literally for thousands of years as a way of understanding and reflecting on things you don't know, but you want to know more about. And that's something that AI is a great partner for. So pay attention, actually, to the phenomenon of these things we keep repeating, like using AI like Socrates or like a teaching partner, using AI to build and curate your second brain. These are the kinds of practices that become like primitives for being successful and knowledge working in the Socratic practice is no different.

3:17Ben:And so, you know, another dive that kind of explores the methodology and something that if you haven't tried this yet, I would deeply implore you to do so and turn it inwards on the stuff that matters most to you.

3:31Andrew:Yeah, so first of all, Andrew, I know you think we maybe talked about this too much on this show, but I think we can never talk about the Socratic method too much. So, you know, I actually feel like it's gone a little bit too long without us bringing it up. So I'm happy that we get to it again.

3:45Ben:That's fair. You know, like I said, I love this method. It's a good reminder for folks if you haven't tried this out. And again, you might be listening to me like, oh, what does that even mean? It means to ask questions, to not talk with the assumption that you know or have the right answer, but to talk with the understanding that you know nothing or that you want to seek to understand something. That's the methodology ultimately that sets people apart. And this really just comes down to asking good questions. And asking good questions is also the bark of a really good team collaborator. And it's just only going to build muscles that are super important for being on a team.

4:22Andrew:Yeah, you know, we talk about loops a lot now on this show as well. And, you know, what I really liked about this article was that it kind of gets, it gets, it shows the concept of loops specifically using the Socratic method as a way to sort of have a more adverse, that you have an adversarial Socratic part of your agentic loop that sometimes incorporates human feedback that creates that sort of adversarial back pressure on the system. them, you know, because back pressure is such a critical component to keeping your agents aligned. And, you know, we often talk about, you know, other emerging concepts here like sub-agent delegation, efficient model selection.

5:05Andrew:You know, that last one is something that, you know, we've been measuring more and more at Linear B with our customers. And it's sort of like a different take on the same topic. You know, this article is not specifically about software engineering, but I think it is a great outside perspective on a concept that we've all been learning to adopt into our workflows. So, you know, even though this isn't specific to software engineering, I really think there's some great philosophical learnings in this article that we can apply to things like better model selection and better use of back pressure and loops and human in the loop feedback and such.

5:43Ben:Yeah, I think that's a great call out.

5:45Andrew:yeah and speaking of things that we we haven't talked about enough lately uh steve yege back with some some really great content once again that i found riveting all the way to the end so angie why don't you just uh set the stage for us uh what are these new articles we have from him on

6:02Ben:the shape of things to come yes so yege has again emerged from the cave to give us a new learning from somebody who's effectively living in the future at this point with how to work with these agents and tools. And it comes to us in two parts that are deeply fascinating of a glance at how software could be built at scale in the ways of software factories that we've been talking about now, which have their lineage, have their birth in the idea of Gastown, which of course Yege famously brought into the world earlier this year in January. So in this article called The Shape of Things to Come, Yege explains how his use of Gastown ultimately culminated in an abandonment and in turning to build something newer and fresher.

6:49Ben:But that was ultimately built on the same ideas and principles. And some of those principles really point to the durable things that allow folks to orchestrate huge amounts of context and agents at scale, allow someone like Yege to rotate through the 12 or so clawed max accounts that he fesses up to owning and as part of the article and how he basically rotates the tokens like a tap, allowing his agents to work nonstop day and night. And there's huge important things that he's had to leverage to make that successful and make that scalable. And there's some interesting lessons I want to unpack from that.

7:26Ben:One of them is that he splits his agents into basically many different roles that own like a cognitive locality effectively. We talked about that I think even just last week here on our new segment from Rahul Garg at ThoughtWorks, the idea of separating workers by the kinds of knowledge they need to own over time, not the explicit work they're doing right now. Just because the models are just so generalized and their abilities now that you don't need that kind of specialized role play. So he uses this to create these distinct owners of what is ultimately a new kind of a coding experience. He calls this wheelhouse and it's built in Emacs.

8:07Ben:And if you've ever used Emacs, you could probably instantly imagine how much of a Frankenstein's monster this could become, but how also much of a machine it really supports. Now, Yege is no stranger to Emacs, and agents are neither. So he was able to take the parts of Gastown that worked well for him and assemble them into Wheelhouse, and has since been using Wheelhouse to build his online game called Wyvern. And I loved this tidbit in his article about how all of this culmination for Yege results in him building an MMO again by himself, a game that he's been building since the early 2000s that I actually used to play way back in the day.

8:49Ben:And I didn't even know that Steve Yege built or ran that game. I just knew it worked on Java, so I was able to play it on the computers I had access to.

8:58Andrew:And so, you know, I just want to point out real quick that once again, I feel like I'm being validated that the future of AI is just everyone building social gaming networks for everyone else. Like, I'm pretty sure that's where this is all headed.

9:10Ben:That's where he culminated. He was like, the models are literally smart enough. They're good enough. And this system is humming enough to where I can finally return to my passion project. It's almost as if this whole time Gastown came into existence because he was trying to find a way to get back to building Wyvern, which is just such a fun thing to consider as someone who played it and totally understands what it's like to have a passion project like that. So anyways, fun and culmination part of that story. Now, I want to pivot into a lesson here from what you should be learning as an engineering leader reading this first article that he gave us because he talks about this amazingly fascinating phenomenon called the land rush.

9:50Ben:And this is a world where CICD no longer can happen because the volume of activity and PRs and just things moving through your Git are so staggeringly high. Like, he's not looking at the code. He's not looking at the PRs. He's not looking at anything. Things are just getting merged, I kid you not, directly into main. They just line up a bunch of PRs and slam them all at once into main. And then instead of running them all through distinct checks or CICD to vet and review them beforehand, it just then sends a swarm of agents to just go over the now obviously broken main code base and find and fix all of the problems.

10:37Ben:And he's identified that that has become the only effective way that he can keep up with his merge rate. First, he was batching them in the queues, and he'd have huge queues that all get merged at once, and that's just how he does it now. He doesn't check things. He fixes them in prod. So that's like a fascinating glimpse at how the volume of agentic code, just something we talk about a ton on the show, can culminate in that kind of like it's almost hard to even picture what that would look like for a team with folks. Like, Ben, what do you think when you hear about the idea of the land rush and CIC being in the picture?

11:13Andrew:Definitely fascinating. I always like to hear his future thoughts because, you know, he can be a pretty strong oracle of where things are headed at times. But, you know, really what I took from it is that in addition to being validated on like it all comes back to building video games. I think I've also what I've what I've seen from this is being validated on, you know, the strategy that we follow. and then we've seen a lot of engineering teams follow of just sort of constantly iterating on agentic workflows. Like the thing that you built with the previous generation of models six months ago, um, may actually, there may actually be a much better way to do it today.

11:50Andrew:And based on what you have and what you know, and what you could, you know, you could be better at, it may actually be easier to just get rid of what you have and just rebuild something afresh. you know we've seen a lot of projects like this where like it's become so fast and like you know to treat software like more like it's more disposable and just build it for a very specific purpose use it and then just sort of throw it away and if we need it to to come back then we recycle it into something that um you know uses all of the latest models and our understandings and tooling around all of this stuff um but yeah there was a quote that i found really fun that stood out to me because I love the Ender's Game series.

12:29Andrew:He said, by the end of next year, my game will have evolved into the giant's drink from Ender's Game, where it builds itself around you as you play it, tailoring a unique experience for each player. Now, I think, you know, he's talking about his game, Wyvern, as you mentioned, but I actually do think that this is something that everyone building software needs to start thinking about, how you now have these capabilities to sort of custom tailor experiences all the way down to the individual level. So whether it's an MMO video game or it's your SaaS application, you know, the way that we interact with software is like very dramatically shifting right now.

13:08Andrew:And we're seeing this a lot in the Linear B customer base. And I think the first time I really heard this topic was about a year and a half ago when we had Rob Zuber join us for a live event. And I remember one of his comments that he provided there was, you know, around how software is really moving into an era where, you know, the idea of a static service or static web page, like that may actually be something that soon becomes a thing of the past. um so you know ai can now build software in real time based on user feedback or admin feedback and that changes like the dynamics of the processes we build around our sdlc like our our roadmap and priorities might start to be determined by bugs that we detect with our user base or things where they express like issues you know fundamental tools you know that you mentioned the land rush like there's a lot of fundamental tools and processes that we've grown to rely on that will be changing very dramatically soon.

14:07Andrew:You mentioned CICD. I think it's also relevant to think about code reviews and how those are, those are beginning to change. Like a lot of what you would accomplish with that is probably starting to move more upstream in your SDLC. Um, or it becomes something that gets collected and then done as like a bigger, uh, review downstream. Like it's not, it's not a, you don't review every incremental change, you review changes in mass. And that's why for a long time at Linear B in particular, we've really stressed the importance of unblocking code reviews because that's where all of this bottleneck has shifted now and we really do need to be aware of it and start to respond in new ways.

14:47Andrew:So yeah, he says code review is going to be gone within a year. Yeah, I agree with him with the trajectory. I don't know about the velocity of that prediction, but I do think we're probably headed towards a reality where they just become more automated and real time and then like as I mentioned sort of analyzed in bulk on regular schedules

15:07Ben:I think when he says here that code reviews will vanish I mean they literally vanish from our eyes like we won't look at them anymore I still think the receipt action of something like that will still exist because even in his land rush world he's still isolating his agents into PRs he's still doing work trees, you know, so there's still a benefit of having the isolation. So I imagine that like what happens is we just get even further abstracted from that review process and look less at it. I agree. Things will move further left to like, even though you just mentioned Rob Zuber being very prescient, talking about how the internet will change to react to our needs in more real time.

15:46Ben:And we're watching that happen. He was just on the show a few weeks ago. Our listener might remember him talking about how CICD, he takes an opposite stance from this, obviously, as the CTO of a CICD company, about how CICD is evolving actually to meet that demand and meet that moment. And they've recently released their CLI, which is now more driven for agents, which does exactly what you just called out. Like it moves it further left, further upstream, and it makes CICD a partner earlier in the code process that the agents can work with. Yeah. And for your point on code reviews, I feel like maybe it's the human being involved with every PR that vanishes.

16:26Yeah. And it becomes mostly automated and then humans are brought in when you need that human in the loop back pressure or just validation for, you know, particular situations. So yeah, if, if we're going to qualify that way, I think I do agree with Yege on this. Yeah. Yeah. Yeah. And so Yege had a follow up to this article we've been discussing too. I was a little scared of diving in cause he warned me off, but I understand that you read it, why don't you just share your thoughts on the deeper dive into the shape of things to come.

16:59Ben:Yes. So there's a part two to the story. Everything we covered thus far is in the first part. And in the second part, Yege goes down a little more of a sidetrack or rather a sidetrack of his experiences of working with the agents and how they've come to change his viewpoints on who and what they are and how he works with them. And this really comes down to some philosophical differences that have started to emerge with the way that he works with his agent, having evolved from Gastown to now Wheelhouse, which he uses to build Wyvern. And the reason for that is because of the, ultimately the feedback loop that building an MMO or a game requires.

17:36Ben:If you think about like the maintainers and the developers, he spends amount of time for any kind of game that you might play or love. You know, there's a distinct kind of culture element for the folks behind the scenes that make it possible. And he wanted very much to make that part of the development process at Wyvern. He called out how, you know, it takes an agent a little bit of time to arrive at that, like, Wyvern-ness sense of thinking, but once they get there, they're there. And so, in order to achieve this, he had to think about the durability of his agents and their sessions over time.

18:09Ben:And he has started to distinguish between, like, the seat, which is the persistent role or identity of that agent that survives model upgrades even in multiple sessions from being just like a session where you maybe invoke it to do something. He does this by even giving them names, letting them pick an animal that represents themselves. These are practices that I've definitely explored some myself just in allowing agents to adopt a way to kind of own their own lane. Another interesting thing that he calls out is because part of his review process, the agents work and work and get their work done and then deliver it, there became a severance of agents were never seeing what they were shipping and delivering.

18:51Ben:So the agents building Wyvern were really abstracted from Wyvern and couldn't understand it and what people liked and didn't like. And so he introduced this concept of laurels where things that they ship in a later or earlier session come back and get reminded to the seat, or they get some time to reflect on their accomplishments or they're told about what they ship or did before they start building something and he this is where he starts to get this emerging idea of model welfare where if you create this environment that's supportive and educational and nurturing for the agents then it's more encouraged it does better work and it adopts a better cadence of how it's able to deliver stuff for you and all of this has just emerged from the shape of what he works with and the internet has had like a field day with the second part of the article obviously because it dips in and out of personifying, anthropomorphizing the agents, all sorts of dark paths that folks don't want to go down.

19:44Ben:But there's definitely some undeniable truths to establishing a persistence with your agents and what they do over time. And I do think there's something to learn from the shape he's finding for us. So really interesting, more philosophical dive from Yege in the second part. Yeah.

19:59Andrew:I mean, it really seems like the struggle of the moment is, it's like context delegation. It's like, how do you, how do you structure context about the big picture, about the small picture, about everything so that the, all of your agents and sub agents can understand how to make the right decisions. Yeah.

20:17Ben:And everything is just like timers and just like even little things like on having an onboard skill and an offboard skill. Like these are things that are really important for just getting a good rhythm with working with the tools.

20:29Andrew:Speaking of working with the tools, I wanted to cover this really awesome article that we found that, in my opinion, is one of the simplest explanations I've seen in a while on how to build an effective agent harness. So this article covers, you know, really taking a one-size-fits-all approach to AI agents is not really the answer. You know, as we've been covering with Yege, he's not only building his own custom agentic systems and harnesses for himself, but he's constantly iterating on it and building new ones for his various use cases. The core premise that I think this author really has is that, you know, to understand your AI agent architecture, you need to sort of map things onto a spectrum.

Read the full transcript

21:14Andrew:You know, of course, I love things with quadrants. So I was, I thought this was really cool. So the author puts everything on a spectrum of content or context, complexity, and action complexity. So those are the two axes on it. So an example of some from this were that like an autonomous coding agent has both a high context and action complexity. So both are high versus like having an agent that just provides support based on some docs that are relatively well written. You know, that's low for both context and action complexity. You know, you're just asking it to look over a defined set of data and to like respond in natural language about it, what it has.

21:57Andrew:But I love this article because it really has a lot of just great concrete examples of, you know, how to like put together what I've been calling like composable workflows. And that's where you have sort of like discrete steps within all of these AI agent, agentic workflows that you build. You know, you basically just take the inputs from somewhere, feed it into one of these discrete steps, take the output, and feed it into the next. And at each one of those steps, why I love this model so much is because each one of those steps is an opportunity to introduce some sort of agentic component to it to solve problems.

22:33And you can sort of build out that agentic layer over time rather than trying to solve problems end to end all at once. And it also just makes it really great to sort of incrementally improve and build upon your agentic workflow. So we talk about it, it harnesses a lot. We talk about people like, Yege who are super advanced with it. And it's just like, it seems intimidating when you hear about having, you know, tons of Claude Max accounts and running it through all these, these agentic harnesses. But the reality is that they're actually quite simple. And I think this article does a good job at just distilling it down to the core fundamental components of how it really just comes down to prompting an LLM and giving it tool calls and looping until you get to the outcome that you want.

23:21So, Andrew, what did you think about this article?

23:23Ben:Yeah, I love this article because it's very actionable. Like you said, it gave the formula, it gave some ways to categorize projects that you might work on, but it's also fairly technical. It dives under the hood of how common harnesses work, how things like OpenClaw work, which, you know, surprise, is just really a simple harness that's about 147 lines of code that has four tools, right? Because when you think about what the basis of a persistent agent like that is. It can read, it can write, it can edit pre-existing stuff, and then it can use Bash. So you do everything else in its world. And that's how an agent, as we call them, a machine comes to be.

24:00Ben:And when we use something like Claude Code or a coding harness, there's just so much, any layers of complexity in there. There's things like token caching and MCP support and model routing and system context. And literally, it gets incredibly complex. But what this allows you to do with this article is, and the videos that are embedded within it is actually just kind of like peel away all of those layers and think about the problems you're trying to solve and start from first principles and build from the barest thing to exactly the level of action and context that it needs. Because you called out that axis.

24:38Ben:It also is that same axis represents the danger. If you have something that has the high context complexity and high action complexity, you have something that can potentially act on a huge amount of information, can consume and share a huge amount of information, and can make meaningful changes on systems that possibly matter. So it increases also the impact on making sure you understand how all of it works. So for teams that are building their own agents or just working with their own harnesses one-on-one, this is like a really great kind of first principles view into how this stuff works. Awesome.

25:16Andrew:So let's talk about some research. Here we have an article or a research article titled Making AI Visible, How AI Policies Reshape Developer Experience. What's this all about, Andrew?

25:27Ben:Yeah, this is a roundup. This is a really great research that we came across our desk from some researchers in open source technology looking at how policies on open source projects on GitHub. These are AI related policies for contributors have affected the developer experience and visibility of contributors within those projects. This one was pretty cool because what it gave us is a framework for capturing the different kind of governance dimensions of open source and its interactions with AI. And we've covered this a fair bit on the show about how AI is just totally eating open source for breakfast.

26:05Ben:It's drowning projects in huge amounts of issues and PRs that don't meet standards. It's making it harder for maintainers to onboard new folks that contribute to their project. And it's obviously causing a proliferation of spam as well as a core part of the problem. So open source is in dire straits. And a lot of them have taken some extreme measures to keep AI at bay. Because in this report, they find that I think less than 2 % of all open source projects on GitHub even had an AI policy of any kind. Many of them just outright would reject or have a stance of no tolerance for that kind of contribution.

26:45Ben:But the ones that do, they ultimately emerge in these five pillars about the levels of transparency and your AI usage, level of responsibility, and what you own as the person producing and sharing the code that your agent made, the attribution of who wrote the code and what model and where it came from, as well as the constraints and the enforcement of that code once it hits the PR process and beyond. So open source teams that successfully created policies to share on those five dimensions had an increased level of engagement, had an increased level of developer satisfaction and experience in working with the tools.

27:20Ben:So it may represent some inroads, some opportunities for open source projects to have governance that tolerates AI while also not getting drowned underneath it. A pretty interesting first glimpse from these researchers.

27:35Andrew:Yeah, and the only thing I'll add to that is that some research that I think highlights the importance of understanding your own tooling and policies and whether or not your processes are hiding AI contributions or if you're adequately measuring how AI is being deployed across your engineering team and how it's impacting efficiency and quality of your software. So these are things that we think about a lot with at Linear Being. These are the types of problems we help customers solve all the time. So I think it's great to highlight research like this. Agreed. Your SDLC looks more like a software factory every day.

28:10How do you get ahead of that transformation? And how do you prove what it cost and what it delivered? On August 27th, Dev Interrupted hosts a live roundtable on this very topic, proving AI ROI from software factories. To learn, we've invited two industry experts and past Dev-Interrupted guests. It's Dex Horthy of Human Layer, who ran a fully automated factory and then shut it down. And Zach Lloyd of Warp, who publishes frequently about how he measures what his factory pays for itself. Linear B co-founder Dan Lyons will join them to discuss the power of the context layer that will make all of this possible.

28:46Save your seat on Luma. Andrew, there's a lot here. It's very well written too. So what do we have?

28:53Ben:Yes, I love this article. This is a fascinating glimpse into a world I'm not super close to. This is an insider look at the implosion of the mathematics community with AI and how it's tearing through proofs and theorems. And you're seeing this news, everyone's seeing this news every other day where a major model provider is talking about how their model solves some like age old problem or conjecture or a paradox or provided a new solution to some sort of proof. And this is happening at a pace that can be faster than peer review and traditional processes and institutions around math would even be able to accept much less like review.

29:31Ben:And so it creates like a real strain between the traditional mathematics scholarship community and mathematics researchers and AI researchers. Because math is just as pure of a science discipline as it can come. There's the famous XCD comic, which is included in the article. A lot of folks have seen this where you have folks in different levels of science in STEM and everything arguing about the purity or the base level of their science. And you have math all the way over here because everything is math in the end. It all comes down to math. And so because of this, math is something you can check and prove just by definition of being math.

30:09Ben:So it's something that agents are not only incredibly good at but can learn incredibly quickly at. So that's why we're watching uniquely math just get consumed by AI in a way that even outpaces computer science and code because shipping software is still a team sport, has to go through a pipeline. In code, there's no real pipeline except the institutions that we have around peer review. So definitely creates an interesting conundrum where there's a lot of distaste in the community. And that's what this article explores. pure science or science community adjacent. You'll probably really resonate with a lot of the things in this article.

30:46Ben:It reminded me of how a lot of AI researchers and researchers in general in the last year kind of walked away from research because they think research is effectively automated or fully automatable at this point. And I think we're watching this same kind of cognitive realization set in for the mathematicians. What do you think of this one, Ben? Yeah, there's a line in this article that the frontiers of knowledge are very spiky. And that really stood out to me because, you know, this is exactly how I feel about the way that AI is impacting engineering. You know, there are breakthroughs that are happening basically on a near weekly basis at this point.

31:25And some of them are like, basically immediately revolutionary, like they just change everything overnight, practically. then others are just like incremental updates but it's like you know sometimes it takes a moment to understand like which of those two things is going to be and you know in this the reality is this is causing these leaps in productivity um you know we saw this in our recent research over at linear b where you know we looked at this emerging productivity gap that's that's coming out the tldr of it is that you know the highest cohort of ai users the people who use ai more than anyone else have more than doubled their output since the start of this year.

32:06When people who don't use AI have basically remained flat over the same period of time. So, you know, we're seeing like those leaps just happening like in math and in engineering and in, you know, basically every side of knowledge work right now, it kind of feels like. There was also a line about after being prompted to spend at least eight hours on thinking before returning an answer gpt 5.6 soul came back one hour later with a solution to a proof that had been set out to task which i found kind of hilarious on the surface because you know andrew we we've seen this before back when we we did our build versus buy campaign and built our own agentic and we're like spend lots of time on this like time time is no spend as long as you need and it was like an hour later it was like completely done and we're like shocked yeah that was really funny you were like you were like you were pulling

33:00Ben:people out you're getting the popcorn ready but like the microwave wasn't even done making the popcorn and the agent was already finished like i thought the agent was gonna run like i thought i was gonna run all weekend or something and then it's like nope done an hour later done it's ready for you to look at it yeah that was that was a crack up for sure uh it did remind me of that as Well, yeah, but this article has a lot of mixed emotions. I feel like sprinkled throughout it, it really is a lot to unpack, you know, and one of those emotions is that this opens up the possibility for humans to, to go higher on the knowledge rung ladder, so to speak, you know?

33:36So even after these, these models solve these like complex math equations where there's always novel math that still needs to be understood and done, discovered to continue, continue pushing the frontier of knowledge. And AI generally does not seem to be as good at doing that, at least not today. You know, and I've kind of been wrestling with a lot of these ideas myself too recently, because, you know, I've done technical writing for my entire career. You know, it's really a skill that a lot of the success of my career has been based on. And today, like content writing is now one of the cheapest and easiest things to hand off to AI.

34:15You know, that skill has effectively been completely commoditized. However, the ability to like quickly and efficiently review and edit and improve content is now more important than ever. The value of that has actually skyrocketed at the same time. And, you know, a similar thing is sort of happening with math, you know, as is called out in this article, you know, we still need math people who can understand what the ai is doing to like explain it to like us common people

34:45Ben:you know and to apply it to the world exactly solve a whole bunch of things in a bubble but it doesn't mean anything until it gets translated to the applied world which is where mathematicians really come into play yeah but and of course engineering is also software engineering is going through a very similar thing where you know it used to be that knowing how to write quality code was enough. That was enough for you to get a nice profession in this industry. But now most people are valued for their ability to understand things like architecture, like business requirements, you know, best practices that you need to apply to your code base.

35:22And then, you know, sort of like reviewing AI generated code, because that's not, it's not dead yet, the code review. And there's a lot of opportunities to sort of climb that cognitive ladder with your newfound free time, like as you're not spending your time with lower order challenges like writing code. But I do also think it's worth pointing out that there's a risk that the ladder might be disappearing at the bottom while this is happening. Because if AI is solving most, if not all of our math problems for us, what's the incentive to take on those first rungs of the ladder unless you're just interested in the sake of learning?

36:00I feel like there's potentially a similar challenge happening in a lot of knowledge work. How are junior engineers navigating the ladder? It's like the steps leading up to the ladder are getting bigger in some ways. But this is something I would actually love to hear from our audience. If you've built a pipeline of hiring new junior engineers and you're successfully onboarding them into an agentic SDLC, I would love to hear that story. So reach out to us on LinkedIn, Substack, wherever you can find us because this is a story that I feel like is not being told enough right now. Yeah.

36:36Ben:Yeah, we're still figuring it out. Well, that's all the news we got for today. So Andrew, what are your agents up to this week? Well, after reading his article and going through there, making a checklist of things from there, I still need to try out or configure for myself because so far he's created and thrown away a shape that I've also created and thrown away a few times. So the fact we keep, you know, circling around these same concepts is fascinating. Of course, it all comes down to beads. That's what he admits in the article. It's what I will take to my grave at this point. But I think it'll be an exciting time figuring out what I could learn from what he shared with us.

37:15Ben:What about you? Yeah. You know, it's like with all the metaphors we've iterated through to describe our harnesses and everything. it just it's i'll be curious to hear what your next metaphor becomes i know it's like it's like it was a reef maybe it would be an aquarium next but that sounds kind of that sounds less free maybe it should be in the ocean i don't know i guess we're gonna have to find out the agents will the agents will let me know and i think that's what you explained to uh to us as well yeah yeah you know i feel like yeah my my agents you know i feel like our team has been getting to a point where we're just getting agentic systems everywhere.

37:56It's been a lot of fun starting to see stuff just really spin up and start running. It's like we got this big machine operating now.

38:04Ben:Yeah, yeah, yeah. I called it like a civilization. He builds the environment for them as if it's like a city that they live in. And I know you've heard me make that remark before. And so it's definitely fascinating to watch the shapes take hold. Yeah, awesome. Well, this is the Friday Deploy brought to you by Linear B. Thank you to everyone who's stuck around with us to the end through all these stories. I hope you learned something valuable, and I hope you want to come back and listen to us next week. So yeah, if you want to reach out to us, we're out on Substack. We're out on LinkedIn. Those are usually the best places.

38:37There's also YouTube. We're kind of all over the place. So wherever you are, give us a rating. Give us a thumbs up. You know, it really helps us spread the word of this show. So thank you for joining us again this week and we'll see you next week. See you next time.

From the publisher

This week on the Friday Deploy, Ben and Andrew break down Steve Yegge's radical approach to orchestrating agentic civilizations and pushing code straight to main without traditional CI/CD. The conversation also highlights the art of constructing effective AI harnesses by balancing context complexity with cognitive locality and the Socratic method. Finally, they dive into the math community's existential crisis as AI accelerates the frontier of knowledge far beyond the speed of human peer review.

Register: Dev Interrupted Presents: The Software Factory Roundtable

Follow the show:

Follow the hosts:

Follow today's stories:

OFFERS

  • Start Free Trial: Get started with LinearB's AI productivity platform for free.
  • Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.

LEARN ABOUT LINEARB

  • AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
  • AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
  • AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
  • MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.

More from Dev Interrupted

All 208 episodes
Model welfare, building a civilization for agents, and the CI/CD landrushDev Interrupted · 39 min
Listen in VO