Gas Town, Beads, and the Rise of Agentic Development with Steve Yegge

12 Feb 2026 · 1 h 10 min · 40 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Gas Town, Beads, and the Rise of Agentic Development with Steve Yegge

Episode Overview Podcast Title: Software Engineering Daily Episode Title: Gas Town, Beads, and the Rise of Agentic Development Guest: Steve Yegge Host: Kevin Ball (KBall) Date: February 12, 2026

Episode Description AI-assisted programming has evolved significantly beyond mere autocomplete functionality. Large language models (LLMs) now possess the ability to edit entire codebases, coordinate long-running tasks, and work collaboratively across multiple systems. As these capabilities advance, the primary challenge in software development is shifting from writing code to orchestrating work, managing context, and maintaining a shared understanding among multiple agents.

Key Concepts

  • Agentic Development: A methodology where multiple AI agents collaborate to perform tasks that traditionally required human programmers.
  • Beads: A tool developed by Yegge for task tracking and management, structured as a task graph and capable of integrating with SQL and Git for efficient project management.
  • Gas Town: An orchestration framework allowing multiple coding agents to work together, leveraging the concept of "mail" for communication and collaboration among agents.

Episode Highlights

Introduction to Steve Yegge

  • Background: Steve Yegge is a veteran software engineer, writer, and influential thinker within the tech community.
  • Journey: Started programming at 17; his blogging began as a way to communicate ideas within Amazon.

Evolution of AI Coding

  • Early Experiences: Yegge's initial shock at LLM capabilities began with GPT-3.5 and evolved with each iteration, culminating in GPT-4.5 and further releases.
  • Transformative Use Cases: Instances of LLMs performing advanced coding tasks like editing large code files with high fidelity marked significant milestones.

The Role of Beads in Agentic Development

  • What are Beads?: A more efficient task tracker structured as a task graph that integrates SQL for querying and uses Git for version control.
  • Functionality: Beads allow for better task management, enabling agents to take on specified tasks autonomously while tracking open and closed tasks.
  • User Experience: Developers reported significant productivity boosts when transitioning from traditional to Beads-based task management.

Gastown Framework

  • Overview: An orchestration tool that allows multiple AI agents to work effectively, solving complex problems collaboratively.
  • Agent Collaboration: Agents in Gastown can communicate and execute tasks independently while being organized under a shared structure.
  • Challenges: Managing the interactions of these agents can lead to confusion and complexity, especially when it comes to keeping track of who is doing what.

Cognitive Impact on Developers

  • Shift in Mental Models: Developers must adapt to a new way of working where they monitor multiple agents rather than solely focus on their own coding.
  • Future of Knowledge Work: As AI tools evolve, the nature of work will change, pushing developers towards higher-level decision-making roles as routine coding tasks become automated.

Predictions for the Future

  • AI in Software Development: Yegge predicts a democratization of software development where anyone can contribute, akin to how social media changed content creation.
  • Bottlenecks in Teams: New challenges arise as bottlenecks shift from code writing to managing and maintaining the collaborative work of AI agents.
  • Potential Dystopian Outcomes: Concerns exist about monopolistic practices in AI development, leading to a future where a few entities control the tools and data.

Conclusion Steve Yegge and Kevin Ball's discussion emphasizes the rapid evolution of software development through AI, the necessity for developers to adapt to new workflows, and the profound implications for the future of the tech industry. Yegge's insights on Beads and Gastown showcase innovative approaches to tackle the challenges of agentic development, while simultaneously raising awareness of the societal impacts of these transformations.

---

Key Takeaways

  • The shift towards agentic development marks a significant transformation in the software engineering landscape.
  • Tools like Beads and frameworks like Gastown facilitate collaboration between AI agents, enhancing productivity.
  • The industry's approach to planning, coding, and project management will evolve as AI tools become more integrated into everyday workflows.
  • Developers need to adapt their cognitive models to effectively manage interactions among multiple AI agents while maintaining a focus on quality and design specifications.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to Steve Yegge

0:45 to 2:17

Steve Yegge discusses his background and journey in software engineering.

“As these capabilities mature, the core challenge in software development is shifting away from writing code and toward orchestrating work, managing context, and maintaining shared understanding across fleets of agents.”

The Evolution of AI in Coding

2:17 to 4:40

Steve and KBall explore the advancements in AI-assisted programming and orchestration.

“I was mentioning before, I've been a fan of your writing for a very long time.”

Tipping Points in AI Development

4:40 to 5:40

Discussion on significant AI model releases and their impacts on coding practices.

“10, 15 % of developers now, 20 % or less.”

Revolutionizing Software Development

5:40 to 7:30

Insights on how AI models are changing the landscape of software engineering.

“So cloud code put chat in a loop, right?”

Introduction to Beads and Their Functionality

7:30 to 9:59

Steve describes the concept of Beads as a task tracker for agent orchestration.

“This is the one that got Gastown launched.”

Cognitive Shifts with Beads

9:59 to 12:00

Exploration of how using Beads alters the workflow and productivity for developers.

“So let's maybe start with beads, which if I blink, that was just three months ago?”

Using Beads for Agent Collaboration

12:00 to 14:00

Steve discusses how Beads facilitates collaboration among agents and enhances productivity.

“And all the open beads are the remaining work.”

Understanding Beads and Agent Orchestration

14:00 to 14:50

Learn how beads facilitate better task management and agent interaction.

“They get kind of mad if you try to take it away from them, right?”

Challenges with Beads and Database Integration

14:50 to 17:00

Explore the technical challenges and solutions regarding beads and its database integration.

“I mean, as long as it passes all the tests, yeah?”

Specification and Workflow in Agent-Based Development

19:32 to 21:07

Discuss the importance of specification in an agent-based development environment.

“And what things are you or another human in the loop?”
Show all 40 chapters

Context Management in Gastown

21:07 to 23:14

Learn how to manage context effectively when working with LLMs in Gastown.

“powerful engine that you basically spend all of your time in one of two modes.”

Introduction to Gastown: The Orchestrator

23:14 to 27:58

Get an overview of Gastown and its role as an orchestrator for coding agents.

“And that means I got to load them up with contacts.”

Understanding Orchestrators in Development

28:00 to 28:34

Learn about the role of orchestrators and their significance in coding agent teams.

“The simplest lens is to lump it into the category of orchestrators, which include things like Devin, which has been around for a long time.”

Stages of Programmer Trust and Development

28:34 to 29:41

Explore the eight stages of programmer trust and the implications for productivity.

“So it lets you run multiple coding agents.”

The Complications of Managing Multiple Agents

29:41 to 30:52

Discover the challenges faced when running multiple coding agents and potential solutions.

“It's developers today, but it's going to be all knowledge workers before long, right?”

Introducing Gastown: A New Approach to Collaboration

30:52 to 32:14

Learn how Gastown enhances collaboration among coding agents using a mail-like system.

“And oh, it was Jeffrey Emanuel and his mail discovery that really enabled it.”

Optimizing the Workflow with Agents

32:14 to 33:05

Understand how to optimize workflows when interacting with coding agents effectively.

“like, here's the thing is if you have a rule for yourself where you never watch them work, never watch them work.”

Challenges of Leading Agent-Driven Development

33:05 to 34:18

Examine the frustrations and leadership challenges in managing agent-driven development.

“we were to throw the database out and try something else?”

Building a Mental Model of Code with Agents

34:18 to 35:30

Learn about maintaining a mental model of code amidst rapid development with agents.

“And so, yeah, it's a different world, man.”

The Importance of Functional Specifications

35:30 to 36:46

Discover why understanding functional specifications is critical in modern software development.

“And the ones who are really, really good, you can have a conversation with them about almost any corner of the architecture.”

Managing Code Complexity in Large Projects

36:46 to 38:12

Understand the complexities involved in managing code for large-scale projects effectively.

“That is the core problem I'm asking about.”

Evolving Perspectives on Code Management

38:12 to 38:52

Explore how perspectives on code management are evolving in contemporary software development.

“are of comparable complexity to a nuclear submarine.”

Navigating Dual Management Dynamics in Software Teams

38:52 to 42:00

Learn about the dual management dynamics that arise when leading software teams with agents.

“So I'm curious, like related to this, actually, like I have enough trouble keeping up my mental model up to date with the work that I'm doing, but I am still leading a team.”

The Evolution of Developer Roles with AI

42:00 to 43:17

Explore how AI is changing the dynamics of software development and the role of manager/developer.

“that the average developer can kind of use it and trust it, right?”

Quality Over Throughput in AI-Assisted Development

43:17 to 44:16

Learn about the balance between quality and speed in AI-assisted software development.

“involves fitting the people problem in your head too, right?”

Transforming Development Processes with AI

44:16 to 45:21

Discover how AI tools can enhance the quality of engineering work and team collaboration.

“like have you tried claude co-work and you know how claude co - can do your freaking laundry for you, right?”

Testing Practices in Modern Software Development

45:21 to 46:51

Delve into the importance of rigorous testing in the era of AI-powered development.

“He insists that his quality is much higher with LLMs.”

The Future of AI in Merging Development

46:51 to 48:16

Understand the challenges and innovations in merging codebases with AI assistance.

“And two, like, what are the knobs and levers that you personally turn?”

Gastown: A New Approach to Agent Orchestration

48:16 to 49:46

Learn about Gastown and its role in creating self-sustaining software development environments.

“I said this earlier, but they felt like it was just working around a bunch of bugs in their model.”

Shifting Paradigms in Software Development Discussions

49:46 to 51:07

Explore how Gastown reframes conversations around AI and software development.

“like Wyvern, my video game, and just work on it, right?”

The Compiler Analogy for AI Systems

51:07 to 52:35

Examine the analogy of compilers in understanding the functionality of AI systems.

“And so I launched it and that instantly, a lot of arguments had nowhere to hide anymore, right?”

Building Self-Sustaining AI Frameworks

52:35 to 53:39

Discuss the journey of building a self-hosting AI framework and its implications.

“I was still using just naked Claude code to build it, right?”

The Future of Software Development with AI

53:39 to 56:00

Envision the future of software development as AI transforms accessibility and creativity.

“Because I hadn't touched anything, right?”

The Future of Software Development and Gaming

56:00 to 56:49

Learn how gaming concepts are transforming software development.

“we're not far off man i mean like age of empires you're building stuff code you're building stuff and you're looking at the outputs of it and man i mean like why not make it fun it's already really fun.”

Impact of AI on Industry Dynamics

56:50 to 57:49

Explore how AI tools are reshaping workplace productivity and team dynamics.

“I think everyone's going to be making software.”

The Shift to Gig Economy and Agile Teams

57:50 to 59:06

Discover the trend towards gig economies and flexible workforces in tech.

“where a bunch of experts who want to solve a problem just to get together and solve it.”

Real-Time Development and the Role of AIs

59:07 to 1:00:30

Understand the shift towards real-time software development with AI assistance.

“and productive, but we're not going to be sharing specs anymore.”

The Addictive Nature of Programming with AI

1:00:31 to 1:02:34

Learn about the dopamine-driven excitement of coding with AI agents.

“And you're just getting them all day long, just getting those dopamine hits.”

Surviving the AI Revolution in Software

1:02:35 to 1:06:24

Examine strategies for software products to thrive in an AI-driven landscape.

“I see kind of nothing but upside from all of this.”

The Dual Nature of AI Progress

1:06:25 to 1:09:28

Discuss the potential utopian and dystopian outcomes of AI advancements.

“or Cassandra or whatever that's infrastructure that AIs will prefer to use instead of building their own or doing it in their heads, right?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Are you passionate about software development and the tech industry? Software engineering daily is looking for a new podcast host host to grow its hosting team. In this role, you'll help shape the show's editorial direction and interview engineers, founders, hackers, and tech leaders. Podcasting experience is a plus, but not required. Curiosity, great communication skills, and a genuine interest in the craft of building software are what matter most. If this sounds like you, reach out at editor at softwareengineeringdaily.com. AI-assisted programming has moved far beyond autocomplete. Large language models are now capable of editing entire code bases, coordinating long-running tasks, and collaborating across multiple systems.

0:45As these capabilities mature, the core challenge in software development is shifting away from writing code and toward orchestrating work, managing context, and maintaining shared understanding across fleets of agents. Steve Yeggi is a software engineer, writer, and industry veteran whose essays have shaped how many developers think about their work. Over the past year, Steve has been exploring the frontier of agentic software development, building tools like Beads and Gastown to experiment with multi-agent coordination, shared memory, and AI-driven software workflows. In this episode, Steve joins Kevin Ball to discuss the evolution of AI coding from chat-based assistance to full agent orchestration, the technical and cognitive challenges of managing fleets of agents, how concepts like task graphs and get-backed ledgers change the nature of work, and what these shifts mean for software teams, tooling, and the future of the industry.

1:42Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript Meetup, and organizes the AI in Action discussion group through latent space. Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc.

2:17Steve, welcome to the show. Hey, KBall. Thanks for having me on. Yeah, I'm excited to go into this. I was mentioning before, I've been a fan of your writing for a very long time. So I'm really interested to see your speaking goes, but let's have you introduce yourself to our members who may not be as familiar with your writing. How do you describe yourself and how you got to where you are today? Yeah. Nobody's ever asked me that before. And you gave me exactly 12 seconds of preparation for this. So thank you for that extensive. How do I describe myself? Yeah. An industry vet for sure. Right. Like I've done this a long, long time.

2:52I actually started programming when I was 17 and I turned 57 a couple of days ago. And so it has been 40 years now. Yay. So I've seen a lot, right? I've seen a lot of transformations and that kind of thing. I picked up blogging at Amazon because I was trying to figure out how to convince an organization of 800 engineers of certain things, I guess, that I thought they were thinking about wrong. And I started ranting and I picked up in popularity. And then, I don't know that just became a thing for me right the blog rants right the drunken blog rants i actually quit drinking over nine years ago i'm going for a 10-year break from drinking and i but you've maintained the rant but i kept the rant going actually and well i'll just say weed's legal in our state let's leave it at that so anyway yeah that's me in a nutshell all right well and one of the reasons i'm excited to talk with you today is i think in your rants you've been one of the people laying out a lot of the bleeding edge of what's going on as a transformation in our industry right now.

3:48Let's start maybe by walking through that evolution. And I'd love to get some of the play by play from you as you were thinking things. So I'm going to go back to I think you had a blog post that called out a pattern we were starting to see that you called chop chat oriented programming back in 2024. I think I remember that. Was that the start of LLM stuff for you? Or was there something even predating that? I mean, like, look, I was when chat GPT 3.5 came out. I mean, I was just shocked that it could write Emacs Lisp functions that were pretty good. Right. I mean, just single functions. And that was about the extent of its abilities, but still was like, well, Elisp, you know, that's pretty edge case out there.

4:30Right. And so I don't know. I think chat was probably the first time where I started feeling like I was taking crazy pills and nobody was ever listening to me because everybody uses chat today. Just the cloud code crowd is what? 10, 15 % of developers now, 20 % or less. It's growing, right? But I mean, like still most people use chat and it was all of 2024. I was trying to get people to use chat, right? And they're like, no, completion acceptance rate. You remember that metric? I do. VS Code is still watching it from what I can tell. Whoa. Well, I mean, yeah, there's probably still developers that are using completions.

5:05So yeah, I started kind of feeling like I could see into the future, right? Because I don't know, look, when you've done this for 40 years and you've chased productivity for 40 years, trying to make yourself go faster, you get a sense for when you're going faster, right? And so yeah, there's a lot of speed bumps and the AI does a lot of things wrong and so on and so forth. But it was just like, I don't know, it was like finding the early hover bike in the Zelda game where you could like, you didn't have the battery yet, but it was still faster than walking. and everyone complains the hoverbacks crashing all the time, but it's like, dude, it's faster than walking.

5:36You have to use it now. And it's just been getting better since then. Right. So cloud code put chat in a loop, right? Chat put questions in a loop and then cloud code put chat in a loop and Gastown puts cloud code in a loop. And Ralph Wiggum does also from Jeffrey Huntley. So we're just seeing basically AI, we're just multiplying it more AI times more AI times more AI, and that is the solution to everything. And it's really making people mad, right? Yes. Hacker news threads and stuff like that, right? Well, and I think we'll get into that, but I'm curious. So along that route, you said GPT-3.5 was the first eye-opener when suddenly the scooter is going faster than walking.

6:17Have there been other turning points for you along that journey? The next tipping point was very much GPT-4.0, right? Because that one, what we had been struggling with when we were building coding agents that only had chat built in, but we would do the copy and paste and stuff, was that once a file got up to around 800 lines, there wasn't a lot of fidelity in reproducing it when they would make changes. Kind of like Nano Banana can be now when it's editing its own stuff, it just starts to blur. And so the tipping point was 4.0 was able to reproduce a thousand line file with perfect fidelity and make it like a one line or a one for change.

6:51That was huge because most source files in the world are a thousand lines of code or less or should be. And so now we're talking about GPT being able to edit all of the files in the world and make simple changes, which immediately at that point, I didn't even know what that was. It was before last year, right? It was middle of 2024. Yeah. Mid to late 2024, I think. But that's already when the, oh, you could farm this. I'm a gamer. I mean, come on. You immediately gamify everything. And then the next, there was another one with Sonnet 3.7. It was like, oh, whoa, right? That was Cloud Code. That was my biggest tweet ever.

7:24I had like 300 ,000 views or something where I was just like, ooh, Cloud Code is neat. And then Opus 4.5 was the next big one, right? This is the one that got Gastown launched. It couldn't have launched without Opus 4.5, right? Opus 4.5 is what got it off the ground. And the half-life on anthropic models, if you've been counting, has been about four months between models at the beginning of 2025. And now it's up to about two months between models. So we're probably going to see an Opus 5 drop real soon, right? So all the naysayers, I mean, we'll talk about it, but just you ask about these tipping points, a lot of the people who are looking at this problem right now and discounting it have approximately a three-month window before and after today.

8:08They're looking back about three months about what the last generation of model was. And they don't know very much about how Isaac Newton invented differential calculus. And like, if you just zoom in, right, they're on this really steep slope and they don't see it. They just see the derivative. And so what's happening is I have had the view since ChatGP3.5 and really for 40 years of this acceleration. And now I see it starting to accelerate really fast. And Opus 4.5 is Splash, the big boulder that has hit the pond, where people are realizing that it doesn't need to get any smarter now at all. Exactly.

8:39So I have, in some ways, a similar journey. And I would say, starting from about Sonnet 3.5, we were at a place where I at least started to say, you know what? If they'd never released another model, and we just kept learning how to use these and building the tooling around them, the world of software engineering is forever changed. It would be. Now, it wouldn't have made it into certain domains. 3.5 and 3.7 had their limitations. There was a certain size of mountain they could chew through, and you would have to break things into that size in order to use them. But it was doable. But now with 4.5, that mountain size has gotten much larger, and it's obviously just going to continue.

9:14So, I mean, look, we're looking at a new world where, I mean, the guy who invented Redis, anti-res or whatever, did you see his post? He had this realization just a few weeks ago. He worked with Opus 4.5 and Cloud Code and was like, well, it doesn't make any sense for us to write code by hand anymore. And this is a really, really, this is the big horse pill that the industry has to swallow right now, right? Yes. And I think there will be some fun times grappling with the implications of that. But let's maybe look at our current moment where things are. And I would say, like, in this evolution, as you highlight, the sort of current threshold is how do we coordinate across fleets of agents?

9:51whether they're working in parallel or even in series where you're trying to do a set of different things. And you've done some projects around this. So let's maybe start with beads, which if I blink, that was just three months ago? Yep. Yep. Three months ago has definitely gotten a lot of popularity. There's also a bunch of related concepts, but let's sort of talk through beads. What are they? What is it as a primitive for this fleets of agents world we're living in? Yeah. So I don't know if you saw, but yesterday Anthropic launched an update to 2D. They retired hired to do right. And they launched tasks, which they credited as being inspired by Beads, which I thought was very nice of them.

10:28And I also get why they didn't just use Beads, right? We can talk about that. But basically like Beads is a task tracker. So like a better to do list for your agent, but it has three properties that make it really, really interesting. One of them is that it's a graph. So it's a task graph, which is a lot like your work graph and your implementation plans and your get charts and kind of how you manage knowledge work in general. So it kind of accidentally captured the ability to capture all work, right? Into microbytes, into bytes that will actually scale up as cognition scales up. So it's really interesting how beads are just, they're just dividing work into a graph, which we've always done, except you add two more ingredients.

11:07You add SQL. So, I mean, graphs and databases have never been that great, but they've been working on it for 40 years and it's pretty good now. It's just a pile of nodes and edges. And so, right? And so you add SQL. The databases love SQL. You got to address the database. The LLMs. The LLMs. You can get shockingly close to Beads just being like, use SQLite to track your thinking. Go. You can, right? But Beads introduces graph edges that we humans never would have probably put in or thought of. Claude helped design Beads. And because of that, it's got stuff the AIs feel is very important. And they use it all the time.

11:42Like the discovered from, which tells you who was working on what. when this bead was opened. They love that for the forensics, for understanding how the work unfolded because beads is the surface of work as it's getting executed. All the closed beads, the integral under that surface is the work that you've done. And all the open beads are the remaining work. And beads itself tracks that surface generally. So yeah. And the third component to it that makes it completely magical is Git. So it is a Git ledger of all of your work. So you've broken it into bite-sized addressable pieces that can refer to each other.

12:18There's a graph structure to it. It's queryable with a database. And it's all on a Git ledger, which is just shockingly useful because you never lose them ever. The history is always there. You can always reconstruct if there was a problem, right? And it gets better than that because you can actually start looking at these ledgers and determining how well agents have done over time. You can even see your own work on the ledger. It's like a portable resume for you. It's really wild. So Beads was a really interesting kind of a discovery. Yeah, I think it's worth digging into a few of those pieces.

12:48But let's maybe start with conceptually for a developer who's used to coding and managing work. If we were to stay at Beads, not talk about Gastown for here yet, but what is the cognitive shift that you make as a developer if you start using Beads with your agent? Well, if you're already using an agent, then it's already using probably to-do lists and markdown files, probably. What else would you use? Maybe you have a wiki or a database, okay? And the problem with all of those solutions is that they don't have Git. Or maybe you're using markdowns and you are using Git. The problem with that is that it doesn't have a graph structure that's queryable in a database.

13:24And the LLM has to read and parse the markdown files to get that graph structure every single time it looks. And they get out of date, et cetera. So if you're using an agent and you're really leaning into it, then as soon as you try Beats, you literally, you just try it. And people reach out to me from all over the world, K-Ball, man. They're like every day. And they reach out to me. I had a colleague meet with me just two days ago. And he was like, I started using beads. And yeah, it's a huge unlock. But I don't understand why, right? And it was because it's like catnip for the LLMs. It's like candy for them.

13:58It's memory for them. And as soon as they get what it provides, they just want to use it. They get kind of mad if you try to take it away from them, right? Because you'll never have to lose any work again. You know how LLMs are. They're like, they're focused on the thing you gave them. And they'll notice, oh, by the way, you know, your other room is on fire there, but it's not really my problem. So I'm just going to focus on this, right? They disavow work. They say, someone completely unrelated to me broke the build. It was them in the previous session, right? That doesn't have to happen anymore with beads because they're like, oh, I see this is a problem.

14:27I'll file a bead for it, right? And you can see how this ties into agent orchestration because as you're piling up beads, you're piling up a work backlog that you could, I mean, depending on how well specified the bead is because you can put in a spec if you want. The bead can have all these fields and comments and design and whatever, right? So if it's really well-specified work, you can give it to an agent and just have it do it and have another agent code review it and then check it in and you're done, right? I mean, as long as it passes all the tests, yeah? So you can see people are using beads as a substrate for agent orchestration.

14:56It's a memory, a shared memory, one that federates through Git, which means it acts like a distributed database for your agents on AWS or GCP or whatever hyperscaler you're running on, right? Azure. They can all communicate with each other with Beads and you don't need a central hosted service. It's all going through your Git repo. It's wild. No, that's absolutely wild. So for going through Git, then I haven't looked at the implementation of Beads. Are you essentially storing Beads in like SQLite? So it's just in the file system right there, or do you have a proprietary interface or how are you actually managing it into Git?

15:30So it's a task graph, it's in SQL, but it's also Git. So is it just SQLite files? Yeah. So I did the stupidest possible thing, which was I didn't do any due diligence. I didn't realize how useful this was going to be. I wanted Git. LLM. Claude wanted SQL. And so we decided we were just going to cram them together in the worst possible way, right? There's a JSON file, one line per issue, and it has merge conflicts all the time. And it gets slurped into the database. And there's a daemon. And it gets stale. It's a two-tier architecture. It's horrible. and it's all going away in the next version release, which will come out like maybe this weekend, right?

16:05Which is, had I done my due diligence, I would have realized that a Git database is what I need. I need a database and I need Git. I need versioned data sets, right? And it turns out somebody has solved this problem. The DOLT team. You laugh. I knew this. Well, I didn't know, but it turned out to be an old buddy of mine from Amazon too, Tim Sen, who started DOLT, right? And I've got friends that are working there and I had no idea that they even had this thing. But yeah, Beads is going to switch to that and it's just going to fix everything. That is, I mean, once again, I've seen a lot of people using SQLite, but the challenge with that in Git is also merge conflicts, right?

16:39Merge conflicts everywhere. So yeah, having a Git native database makes a ton of sense. I mean, the three-way merge goes away, BDSync goes away, the daemon goes away, the whole thing. You still have the Git export. You still have the federation. DOLT federates exactly the way Beads did. It's weird. It was like I was following in their footsteps, right? But they did it right. And they put 10 years into it. And it's embeddable in Go. So it just happens to be embeddable in Beads, right? So it's just going to be one binary. You won't even notice. It's just going to get better. And it enables a bunch of stuff like field level merge resolution instead of issue level and all kinds of new history and kind of new dimensional sort of looks at the things that we weren't able to get before.

17:20Key value stores and things like People are starting to use beads. Beads is a data plane, man. It's a data plane. It's nuts. And people are like one contributor put in a key value store. And I was looking at it and I was like, I was trying to wrap my head around it. And I was like, yeah, it makes total sense. If you have agents using this as their memory. Right. They want to be able to share things and have a reliable way to look it up and all that. Yeah, yeah, yeah. A bead is a task. It's very heavyweight relatively, even though it's lightweight compared to like a GitHub issue or a Jira. They're very heavyweight.

17:49A bead is much lighter weight than that, right? but a key value right is super lightweight so i was like yeah let's do it and don't supports them really super well whatever right so i'm just so happy the way beads is going the code base is garbage right now it's vibe coded which means that you have to run code review passes on it like constantly and i got behind and i'm like a month behind on code reviews and so i'm sure the code is just garbage right it works it passes the test people are using it but within a few weeks within i don't know a week i'll get it all cleaned up we'll be on dolt and beads is going to be a thing a beauty.

18:20In mobile application security, good enough is a risk. GuardSquare uses advanced, multi-layered code hardening techniques and automated runtime application self-protection and mobile application security testing, combined with real-time threat monitoring to deliver the highest level of mobile app security. Discover how GuardSquare brings all these together to provide mobile app security for your Android and iOS apps without compromise at www.guardsquare.com. Why is there always a meeting bot in your Zoom call? Blame Recall.ai. Recall.ai powers the meeting bots and desktop recording apps behind products like Cluely, HubSpot, and ClickUp.

19:04They handle the hard infrastructure work, capturing clean recordings, transcripts, and metadata across Zoom, Google Meet, Microsoft Teams, in-person meetings, and more, so developers don't have to build it themselves. If you're building a meeting note taker or anything involving conversation data, recall.ai is the API for meeting recording. Get started today with$100 in free credits at recall.ai slash software. so i want to like now step a step back so we talked about beads the particular thing and you alluded to a few different pieces in there about changes in the way that we're approaching code that i want to talk about before we get all the way into gastown which i think takes this up a few orders of magnitude so one of the things you talked about was you said hey if a bead is well enough specified it can just get farmed out taken care of etc how do you think about specification in this agent world as we're doing it?

19:59And what things are you or another human in the loop? How are you managing that? How do you think about those things? Yeah. I wish that I had more time to think about it. This question is even more fundamental to Ralph loops. I've been talking to Jeff Huntley a lot about this. And for Ralph loops, you really have to specify your acceptance criteria very thoroughly or else you run the risk of getting the wrong thing. And so the way Gastown approaches it, the way I approach it, my workflow is basically like we're only going to ever implement everything to a first approximation unless it's really important like the dolt stuff right we really really push hard on that but everything else is successive sort of iteration we're just going to get it out there and fix bugs in it and you know what i mean yeah just wondering about how you think about specification and like in your workflow with beads for example when do things bubble up to a human versus running autonomously.

20:53Yeah. So you have all these different workflows that you can support and mine tend to be so iterative that I just rarely get time to get a lot of specification time in, but it's a really interesting question. Look, you just have to make time for it. You're right. Like Gastown is such a powerful engine that you basically spend all of your time in one of two modes. You're minimaxing, you're either minimizer and maximizing context windows. I had this really interesting discussion last week with some folks at Anthropic. I met some very lovely teams at Anthropic who were interested in Gastown because to them, they see it and they see it as exposing a lot of bugs in their model, right?

21:29Because a lot of the Gastown workarounds are things that the workers probably ought to be doing better if they understood that they were factory workers, but that's not something they've ever been trained on, right? So I was talking to them and they said that there's an interesting kind of split inside of Anthropic where some people love to minimize context use. So it's like use the smallest task possible, decomposition, right? Throw away, ephemeral, just write one task at a time because you get the benefits of your context window, your costs expand quadratically as the token size grows, and also the performance tanks after a very small size.

22:03And so they're all about performance and cost, which is great. I'm going to tie back to your specification question in a moment, I promise you, okay? And the other group is the maxima, the context maximizers. And what they do is they load up the context window heavy with just lots of rich information and instructions. You know what I mean? Because LLMs perform really well and make really good decisions, especially to strategic decisions, when they understand why they're doing something and not just what you want. And they said, so which one are you, Steve? And I was like, well, interesting that you say that.

22:37You've just described Gastown's Polkats and crew. The Polk hats are for the ephemeral work that's already well specified and it's throwaway. And you actually want them to be small context. You want to do one task at a time. You decompose it, right? Get them to work through it, farm through it. And it's factory farming code. But there's a lot of work that's usually design work where you're doing the hard thinking and you need to have conversations with the LLM. And usually you want to build up a lot of context with them, right? Not to where they're getting amnesia, but often I'll be like, okay, you've just hit on a really difficult corner.

23:08Anytime there's some difficult corner of the code that I'm working on, I'm like, okay, it's time for us to roll up our slaves and not just band-aid it, but figure out the whole, where it fits all in. And that means I got to load them up with contacts. And so I have a set of documents that I'll pull, like of increasing mind-blowingness, right? And so, yeah, the crew supports that kind of workflow and the poll cats support the other. And I think it's a recognition that they both exist. And I think as engineers, we've fought back and forth between them. But we're getting gradually pushed over to the heavy thinking kind of work where the LLMs are just going to do all the coding because we've done all of the difficult design, which is why I mentioned in one of my last blog posts that I'm taking naps all the time.

23:46Yeah. I want to get to that because I have noticed a similar type of exhaustion in this work. But before we go there, I want to follow up on this just a little bit more. So the thing you described in terms of when you're getting into heavy problem solving mode and you're booting up all of this context, like one of the ways I've been talking about this with folks is like, if you conceive of these LLMs, they're fundamentally like they're a little VM, they're a computer that is language driven and code driven and all these different things and code is data. What you're trying to do is essentially write the bootloader for the problem you're solving.

24:15How do I get exactly the right set of context and data to get the right relevant things? So I'm curious, how do you for yourself manage or how do you within Gastown manage like, What does that bootloader look like for different types of tasks? Oh, okay. So over time, before Gastown, when I was just automating my own workflows with Beads, which I think was where a lot of people are, or even without Beads, some things that I found really helpful were, see, the thing is, you got to learn what they're good at and lean into it, right? And so one thing they're really good at is to-do lists. They love bureaucracy.

Read the full transcript

24:50They love acceptance criteria. They love checking things off. Yeah. And so on the boot up side, there's a bunch of stuff that I want them to do specifically because they failed to do it on the shutdown side. So like on boot up, I want you to go and look for branches and stashes and unmerged work and blah, blah, blah, right? Unclosed beads, whatever. Clean up your sandbox, clean up your environment on boot up as well as on shutdown. And so both of those instructions became things that I encoded in prompts. And so for many people, they're at a phase in their engineering where they're managing your own private libraries of prompts that they like pull out as needed, right?

25:26And Gastown was just basically me going, well, what if I could just have some canned prompts that sort of came up when certain roles came online? And then it was all predicated on this, what if Cloud Code could run Cloud Code, right? But yeah, the boot up and the landing are really important. So like I have this prompt called Land the Plane. And my colleague on Monday was telling me about this. He hasn't used Gastown. He just uses Beads, right? But he found Land the Plane really useful. I lived by Land the Plane for, I don't know, Six weeks. Feels like six months last year. Time is compressing.

25:54It's super compressing right now. Yeah. So land the plane is okay because the agent will be like, party, we are done. Right? The agent is like, literally, it's giving you emojis, checklist, to the moon. This project is ready to launch. I am done with this feature. Look at all the things we accomplished. Anthropic models in particular love to do that. They love to make it. GPT is a little bit more staged, but yeah. Except GPT can't code, so who cares, right? Well, different discussion. I have found GPT-5.2 to be incredibly effective for particular styles of code and prompting. But it goes a hell of a lot slower.

26:29I heard Gemini 3 is really good at UI coding, maybe as good as Claude. So yeah, sure. And UI coding is like a thing that GPT just falls flat on its face. It's so bad. So bad. Yeah. So I don't know, maybe they'll develop specialties, right? So Claude loves acceptance criteria and landing the plane takes advantage of that by, it's almost taking advantage of them, as if they were OCD and giving them some OCD thing to make them right. Because even if they're low on context, even if their context window is near exhausted and you're like, let's land the plane and they look up the instructions, they'll be like, yes, sir.

27:02And they'll start checking things off and they will finish that thing, right? Even if they hit a compaction, which I view as a failure mode, right? So you try to land the plane as early as you can. But yeah, the land the plane makes them much more reliable at not forgetting stuff. They're just, it's like they have common sense, but they get distracted. They need to be reminded. Yeah. So yeah, I mean, like these are all muscles that you build as a developer before you jump in the lion's den with something like Gastown, right? You've got to be really, really good at like bringing context into LLMs, watching how they deal with it and triaging it and dealing with it.

27:35And once you get into that cycle manually, leave for a while, right? You're going to feel more confident to be able to wrangle like eight of them at a time. So let's maybe use that then as a jump over into Gastown. And let's start from the beginning. Like let's describe for anyone who has not read your rant or the many responses that has fallen down on the internet. What is Gastown? What's in the box? All right. I mean, there's a lot of different lenses that you can use to look at Gastown, right? The simplest lens is to lump it into the category of orchestrators, which include things like Devin, which has been around for a long time.

28:10They're attempts to run multiple coding agents as a team, basically, or just in parallel on parallel tasks with some tracking layer. And the Ralph Wiggum loop and the loom loops from Jeffrey Huntley, and you've got Claude Flo, the fancy one with the routing. And then you've got what else? There's a few others here and there, but there aren't many, but it's in that category. It's an orchestrator, right? So it lets you run multiple coding agents. And it's pretty closely tied to Cloud Code right now because it's on the boundary. It's on the edge. It pushes the agents so hard that they get confused regularly.

28:43And so only Opus 4.5 is really strong enough to be able to run Gastown reliably. And even then, it breaks a lot. But anyway, Gastown is predicated on a really simple idea. All right, let me tell you how it goes. As your trust with the LLM grows, okay? And this is my eight stages of programmer that I called out, which got a lot of attention on its own, by the way, because it was a real challenge to people. Right. But I'm sure it resonated that they realized that they had moved up from stage one and they finally felt pretty good about it and that they were done. And when they saw where they fit on the entire thing, they were like, oh, no.

29:17And their ego is a bit bruised. But by the same token, they couldn't really deny it either because they had already seen their transition of leaning more and more into the agent. What's happening is you're trusting it more. And by trust, I literally mean you can predict what it's going to do better. That's the only way you're going to trust it, right, is being able to predict. And that just means practice. And we're talking hundreds to thousands of hours of practice to get up that ladder, right? And this is not just developers. It's developers today, but it's going to be all knowledge workers before long, right?

29:45as your trust goes up your patience goes down it's very interesting okay because you're like you got your agent and they're working on the thing and you know they're going to get it done because they've done it five times properly before they're going to fix another test for you and you're just like you know what i'm going to start up another agent and that's it man that's the gateway drug that's the end of it i hear you i'm like i'm right now if i'd self assess i'm like um the cusp between six and seven for you right like i'm managing three five agents Sometimes it pushes up. And I have some questions around where I'm running into limitations that I will get to you.

30:21But okay, so you're moving up the stack. Your trust is increasing. Your patience is decreasing. You're running more and more agents. And now you start running into problems that are very different for an individual depth. They're like team problems. You have agents that are running, you know, stepping on each other. You forget who's doing what. You have gates, agents waiting on each other. I mean, it starts to get kind of complicated. And at a certain size, you lose the ability to keep track of it in your head, even if you're using beads and it's just a zoo. And so Gastown was me going, well, what if I just put them all into like WorkTree hierarchies and sort of gave them names?

30:59And oh, it was Jeffrey Emanuel and his mail discovery that really enabled it. Beads was a huge unlock, but the other half of it was mail, right? So the thing is, LLMs like stuff they're trained on. And the longer it's been in their training set, the more they like it. And so mail, email, which has been around since the 70s, is like a pair of old jeans for them. They love the mail interface. And so you put them together with identities and the inboxes and they can send each other mail and they will. And so very quickly, I had this town of collaborating agents using mail and I gave them names. You're the mayor, right?

31:31You're Polkats. And I said, you're going to be named after Snow White and the Seven Dwarfs. So I went through the Seven Dwarfs. And one day I saw the mayor and it was really mad. It mailed Sneezy Polkat. and it was like, you are not the mayor. You're a sneezy polecat. So read your own inbox and do your own work. And I was just like, what is happening here? It was beautiful. Then one day, a swarm took off and fixed all my bugs. I had like 30 beads all lined up to knock out and I couldn't find them and I panicked and I realized that they had all been closed and fixed. And I was like, woo, that's that swarm feel.

32:06So Gastown kind of emerged out of the muck, out of this primordial soup of managing agents by hand. But basically it gets into this mode where you just like, here's the thing is if you have a rule for yourself where you never watch them work, never watch them work. That's counterintuitive advice. Most people keep their eye on the agent and they're like, I'm going to see if you're going to make a mistake. You made a mistake. That's what they do. And that's what I did for, I don't know, six months, right? I tried to like watch them all. Oh, that diff looked weird, right? As soon as you get out of that mode and you realize that they're going to make mistakes and you're going to find them, just like regular engineers.

32:42Okay. So don't sweat it. Then instead, you're only looking at the ones that are finished and the ones that are finished are a problem because either they didn't land the plane properly and you need to walk them through that process, right? They think they're finished and they say they're finished, but they're not finished, finished. And so you got to walk them through that or they're really finished. And now it's like, what would you like me to do? And I'm like, well, let's talk about it. Right. Whoa. Well, and that's when you sit back and start having those wonderful design discussions with them where you're just like, okay, so what if we were to throw the database out and try something else?

33:11And then you're like, okay, off you go, go think about that. And you cycle to the next agent and you go, okay, I'm going to give you a hard problem too. And so that's actually how I use Gastown now is I spin them all up, all the crew with hard problems. And you know what? It's so wild because like eight out of 10 get done and two out of 10 get lost. And I'm like, I know we thought about this before. Sometimes we can find the design. Sometimes we lost it and I just have to redo it from scratch. It's a little annoying. So this gets into one of my key questions, which is one that, once again, I'm grappling with not even quite being at the Gastown level, but I can imagine it's even more, which is like, how do you mentally keep up?

33:45Yeah. So, I mean, like, for starters, it can be frustrating and you can find yourself yelling at them. You know, like, I told you for the billionth time, don't make PRs. We're the freaking maintainer. I mean, come on, it's right in your prompting. And they're like, I made a PR. But then, you know, I realized that it's my fault, right? Like, there's a solution for it, actually. You need a tool pre-use hook from Claude and you just say, don't make PRs. And that's the end of it, right? It's solved. So you got to realize these things are getting better so fast that you've just got to like mostly just take this very nice Zen approach and just be like, look, they're getting stuff done and that stuff is making forward progress and we're moving the goalposts, right?

34:21And so, yeah, it's a different world, man. How do you avoid getting tired? Dude, I take naps throughout the day. I'm exhausted. It's like I'm a factory manager now with a brand new team. They're all a little clueless. They're smart, but right. They're bumping in each other. And I don't know the business very well. And I'm running around trying to keep them all busy. It's exhausting. That, I guess, gets to my question of like, and this maybe gets to another kind of related thing is like how well, because I think of you write software, writing software traditionally, if we go back even two years before all this, you're writing software and it has a couple of roles.

34:52It is, you have an executable artifact. That's maybe even the smallest one. But two, you're like creating this mental model of a system that you have. You're expanding your mental model of the business problem you're trying to solve, and you're mapping between those things. Now, agents are writing code, but those two other problems still maybe exist. Do you have a mental model of the code that is being written? Yeah. So, like, have you ever worked with a really, really, really good product manager? Like a very technical one who used to code a lot, and now they're a product manager. Once. And, okay, once.

35:28Yeah, you're right. And they're like gold. Actually, the ratio is very high at places like Google, like the technical ones, but still not 100 % by any means. And the ones who are really, really good, you can have a conversation with them about almost any corner of the architecture. And they'll know whether it's using a hash table or a list or whatever. They'll know the O of N performance of it because it affects a list that the user sees, right? Or the length of time that an export takes or something, right? So they'll have conversations with you like, well, the export's taken too long. like what if we used to cash for that and da, da, da, da, da, right?

36:00They have like the good ones. And this is also true for an Uber tech lead. Have you ever worked with an Uber tech lead at a big company, right? They're working with tech leads. And so they're coming up with a consolidated, synthesized view of the machine that's being built across all of these teams, right? And so the extent that you can do that, that is an engineering leadership role. And it's a hat that can be worn by product or by technical program managers or by senior engineering leaders, executive leaders, whoever wants to get down there and try to understand how this thing is working. And I tell you, that's the most important part.

36:33It's not actually what language it's written in usually. It's not the syntax. It's not any of the details of the linker or any of that stuff. It's what does it do? What's the functional specification of this thing? And we have the ability to build software so fast now that keeping the functional spec in your head is a huge task. That is the core problem I'm asking about. Yeah. How do you do it? I have to remember the entire surface of beads in Gastown, including all of the integration points that people have brought in. And I find myself often having conversations with the LLM going, I'm really embarrassed that I know this, but how does our plugin system work?

37:09Or how does whatever work, right? Last I checked it, it met my approval, but that doesn't mean that I remember how it works, right? And so I have to go back and reconstruct it. And yeah, so it's this constant, it's a discipline thing. It's a hygiene thing. Just like you have to do regular code reviews of code you've never seen. And you have to continue reviewing it and asking questions until like as an Uber tech lead, you're like, OK, I think they're in the noise now. They're in the weeds. Like we don't need any more code reviews for now. Right. Because you've asked them a bunch of different angles.

37:35Right. Same thing goes on for the product. You've got to understand every nuance of your product. And if you don't and you're not using it, then why do you even have that code? So you've also got to start pruning aggressively and sort of like retiring features that aren't pulling their weight. because otherwise you'll have tech that mountain, right? I mean, look, man, I mean, people are like, yeah, they're not looking at the code and so they don't understand it. That's a very junior mentality. That's a person who's not very seasoned or experienced, somebody who's never led a team, somebody who's never been, you know, responsible for a very large operate, you know?

38:07I mean, like, look, I was on nuclear submarines in the US Navy and the software projects that we make are of comparable complexity to a nuclear submarine. And they're, you know, freaking six stories tall and several football fields long. And they're very complicated, right? And this notion that people have, they're just thinking about it completely wrong. I couldn't agree more in terms of like, once again, the code details, like that is in the noise at this point. And I think there is a world that we've been talking with some people about, like over the last 15, 20 years, we moved from servers as God servers to pets to servers or cattle.

38:45You don't think about servers you're spinning up and down. And we're doing the same thing with code, right? Like code is cattle is like the world we're going to here in many ways, but your system still matters. And the pace of things still matters. So I'm curious, like related to this, actually, like I have enough trouble keeping up my mental model up to date with the work that I'm doing, but I am still leading a team. So I'm also keeping up with, you know, and people who have N agents working for them. Yeah. Like this problem expands. So are there toolings that you're using thinking about this?

39:18How does Gastown expose this stuff in a way that is easier to process? How do you approach it? I was interviewed with, I don't know, 10 other famous-ish people in 2008, a long time ago. And one of them was Peter Norvig. We all got some Polish kid reached out to us all, somehow hooked us, and we all answered a bunch of questions. James Gosling and Guido Van Rossum and all these people, right? And Peter Norvig gave the best answer to one of the questions, just all time. It was something like, you know, what is it that differentiates great programmers? And everybody else went blah, blah, blah, blah, blah.

39:49And Peter Norvig said, being able to keep the entire problem in your head. That was his answer, right? And that is the problem that we're all faced with. In fact, what's going to differentiate successful teams from not successful teams is the ones who are able to keep bigger problems in their collective head. That cross-functional communication and coordination, those costs, LLMs can help reduce them, but But the rate that you're producing just exacerbates the Jevons paradox of project management, right? I had a non-technical boss once, not very technical, who they were very good at using lieutenants, like using counsel, senior principal engineers who they trusted to guide their decision making and lead their organization.

40:30And they had all the other leadership pieces. And so it was an arrangement that worked. And I think that some companies like Amazon adopt this directly. and you get this dual management type, you know, set up, right. Where you've got a manager and you've got a technical sort of advisor or a technical leader. Right. I think that we're all moving into this, into this role, right. Where we have a, you have access to a technical advisor. And so now the question is, right. I blogged about two friends of mine that yelled at their colleague because he was two hours behind them. Right. Did you hear about this?

41:00Yeah. Cause they're running 20 agents each. And so the, they've realized that they're moving so fast. that if their work is hidden from view for even a little while, if they're not completely transparent and pushing and public and loud about everything that they do, then they might as well be working at the bottom of a mineshaft. And the world will move on without them very quickly, and their stuff will be impossible to rebase because stuff just moved that fast. And they wound up getting cross with a colleague because he was like, well, I implemented the blah, blah, blah. And they're like, why?

41:30Why did you do that? What information was that based on? And he's like, it's from two hours ago and they're like two hours ago what's wrong with you we made six decisions since then right and it's like this is a serious problem right so i mean like it's the problem like beads is and and gastown have sort of taken the brakes off and is like it's going to take another one to two model iterations you know opus 5 opus 5.5 call it summertime before they run the orchestrators run smoothly and the work that they do really is you know high enough quality that the average developer can kind of use it and trust it, right?

42:03But we're almost there, right? This is one of the things and one of the reasons I was excited to get you talking on this because I've interviewed a whole bunch of people doing coding tools. I'm using all of them, and everyone's focused on the individual developer, right? How do we optimize the coding productivity of an individual developer? And LLMs are flipping crazy for this. Like, we can go so far. And using it in the team context, your bottlenecks shift completely. And now it's how do we keep our wetware up to date? Yeah. And also, like I was saying, that dual management thing hits everybody now.

42:36You are a manager. You are a leader because you have a team of software engineers. You got thrust into the role. And now your other skills, your soft skills, your humanities, your college education, your liberal arts is all really important now. Your people skills matter, right? Because you're going to have a 300 ,000 line thing you want to go and bring to another team because it cures cancer. And the other team is going to be like, we can't read your code. What do you, who are you, right? And it's going to become this negotiation to get anything done anywhere, right? Because you've produced too much.

43:10Companies are struggling with this, just merging each other's work together, right? So the better you communicate, fitting the problem in your head involves fitting the people problem in your head too, right? It's like, you know what I mean? Like the, you know, who needs what? And I think to an extent, it's a skill that some people will have and some people won't, but it's a skill that you can learn. It's a skill you can foster and develop and be taught. You can be with mentors and they can teach you the skill. This is going to be how all knowledge work works. They call it centaurs. Yeah. Have you heard this term?

43:40Yep. Do you know who invented it? I don't know the origin. So I read online. So it must be true that it came from Bobby Fisher, who was, he used it to describe AI assisted chess, you know, then there's the whole centaur minotaur blog, which is really clever, right? Which is which one do you want? You know, do you want the AI head with the human body or do you want the right the human head and everybody wants centaur and then there's debate over whether it's actually called a chimera because that's the term that academics were using until bobby fisher took it all away and then you hear like boring you know hybrid whatever but look the fact is everybody is going to have a friend a helper everybody's going to have an ai right something like claude like have you tried claude co-work and you know how claude co - can do your freaking laundry for you, right?

44:23Yep. I mean, I didn't try cowork because I already have cloud code and codex working within obsidian, right? Like I already have what I, what they're shipping it for. So you get it, but you turn on cloud cowork and you're like, Oh, I get it. It's cloud code for everybody else, like on your desktop. And to be fair, like I was trying to, my wife at some point had a problem where I was like, Oh, this is an easy cloud code problem. And I couldn't boot her into it. And then we tried to do it using cloud.ai, the web interface, and it was terrible. And so the use case for Claude Cowork makes perfect sense.

44:55Yep. Yep. Same, same. Fun times ahead. The naysayers, I think they're in for really, they're going to go through the five stages of grief this year. But interestingly enough, at the end of it, you wind up hopefully with happiness. Some people will drop out of knowledge work. I think they're just not going to like this. They're not going to like this whole centaur thing. They're not going to want to work with an AI. you know so a thing that i've been grappling with is a tremendous amount of what we're doing with ai tools right now is trying to do kind of work at the same or maybe slightly lower quality in much higher pace and there are times when like more is more right you've got a bunch of features you got to rip out like it's fine the thing i'm trying to invert is like what are the use cases or when are we able to use this ai assistant to do the same work but at a much higher quality i got a buddy who's one of the best engineers in the whole world, I mean, the stuff he's built is a really long list at big companies like Amazon and Google.

45:52He insists that his quality is much higher with LLMs. And it's because of the way he does it. He just chooses, you get with LLMs, you get what you choose, what outcomes you choose, right? And he chooses quality. And so what he does is he reviews all of their work and it becomes a pair programming exercise. And it's necessarily better because it's the best of both of them, right? Just because I'm optimizing for throughput doesn't necessarily mean that you can't dial up quality. And for certain launches like Dolt, I'm dialing up quality just insanely high, doing wave after wave after wave of code review and going to the Dolt team and having them review it, right?

46:27So, you know, quality is just a choice, man. That's all. So that's, I think, a really important thing to discuss here is like, how do you, let's spell out as someone who has been hacking with these things at both ends of that spectrum for probably more than most folks listening to this, like when you want to start dialing up quality, actually one, how do you decide this is appropriate? This is the thing where quality is, is important versus throughput. And two, like, what are the knobs and levers that you personally turn? For me, for the stuff that I'm working on, I'm in a space where I am very fortunate to be able to have a very low bar.

47:05A very low bar. Nothing has to work at all, right? Beads, I can't break now because people are depending on it, right? So, but Beads is incredibly well tested. It has significantly more test code than actual code. And I think this is just a vision of the future, right? Where we just have everything just is tested to death. And integration tests too, right? Token burning tests, right? Just like verification and validation are the two gates. And validation is, did you build the right thing? Which is another problem. You can have something go and build something that's great, but it's the wrong thing.

47:42And then verification is, did it build it well? Along with keeping the problem in your head, this is one of those, we've got new sort of pressure points. We've got new bottlenecks in development as the old ones have been obliterated by AI. One of them is merges. One of them is shared designs and keeping those up to date. And you know what I mean? Like we've talked about that. The merge one is terrible. And it's why Gastown has a refinery to try to basically have a dedicated role that knows how hard merges are and just redoes them. Right. But you know, it's funny, I showed Gastown to Anthropic and I don't know if I said this earlier, but they felt like it was just working around a bunch of bugs in their model.

48:23i mentioned that so what this implies is that and people have already pointed this out right is that gas time will flatten like i don't need as many roles because half the roles were just work around saying do your job and as soon as it knows what its job is you only need the other roles so i think you're going to have a really simple maybe two-tier hierarchy where you yeah you know maybe the little three-tier just because organizationally you know you want to take advantage of that Militaries are hierarchical for a reason. I'm sure AIs will choose to be hierarchical on large enough projects, but you only talk to the top, right?

48:59I think that's where we're headed. And I think that it's a skill that everybody will learn this year because it's just so effective, right? But senior people are benefiting more than junior people because they know what good looks like. And we're in a stage right now where the AIs don't always know what good looks like, right? So again, verification and validation for a senior engineer are often as simple as just looking at it. So it's a little harder if you're in a new space, a new domain, a new programming language, a new area. You're trying to build for a customer that's not you. I try to avoid doing that.

49:35I try to avoid building things that, you know what I mean? I was going to ask, have you tried building anything for normies using Gastown? So I have been so busy building Gastown that I haven't been able to spin up a rig for, say, like Wyvern, my video game, and just work on it, right? And see how it deals with, for example, UI programming. How do they handle that? Gastown has had a lot of bugs. Gastown has had much faster adoption than I expected, even though I expected people to ignore my dire warnings and use it anyway. And so I've had to stay on top of it because it's had stalled workers, runaway workers.

50:07It's had, you know, I had to kill 320 cloud code instances on my machine the other day. And each one of them takes like a gigabyte of memory. So it can really kill your machine fast. So, yeah, we're still in that stage. But, yeah, this is a new world, man. It's fun times. Look, can you build stuff for normies? I don't know. I wouldn't recommend it. Right. Look, why does Gastown exist? Like, what am I even doing with it? If I know that this isn't the long term shape, what is it for? I mean, one of the main things it did truthfully, and you saw this, I think, online is that it reframed the discussion completely.

50:40Oh, yeah. It reminds me of the political framing of like there's an Overton window of like where people are talking and you just moved it. I moved it way down, right? What was acceptable to talk about. And it was because up until Gastown, I was just a blogger who was saying AI is going to be big. People are going to be running fleets of agents. It's all going to be, going to be, right? and Gastown came out and was just so damn smug right like like purposely like I had a whole cast of characters it had mythology there was a song you know and a nano banana really came through my god those pictures are unbelievable right and also I had discovered the whole you know molecular work thing and so it had kind of a theoretical foundation that was actually working to where I could see it working.

51:28It was doing what I wanted it to do. It worked, right? And so I launched it and that instantly, a lot of arguments had nowhere to hide anymore, right? It changed the conversation overnight from, no, no, no, what you're saying is wrong to, bro, you're pretty aggressive, right? That was what happened. Now they're faced with the reality that it is building itself, which nobody's come right out and said it, but that's what it's doing. It's building itself. And so it's at least good enough to build itself as a swarm. And this is deeply, deeply unsettling for people. And some of the content that I've seen is just hilarious responding to it.

52:06Right. They feel like they're getting taken over by the hive mind. This is Ender's game, like playing out in front of us. It's really wild. So the building itself reminded me of a thing, right? So because that reminds me of building compilers. And one of your first goals with a compiler for a language is that it should be able to compile itself, right? You want to create a bootstrapping compiler. So Gastown, it sounds like, is a bootstrapping agent orchestration framework. It is. And it's a great analogy. And when my friends knew about Gastown, but it hadn't booted up yet, we would use the compiler analogy.

52:37I wasn't using it to build itself yet. I hadn't booted into it. I was still using just naked Claude code to build it, right? It was so weird. You would think that you would just boot into it, man. I wish I could describe this to people. You would think that it would just kind of turn on. but it went through about six or seven like layers of waking up and it felt like i was digging a tunnel and burst out into this new universe and that we started setting up camp and all of a sudden we were building infrastructure and cities and stuff right and and every single iteration of this almost every day it felt like there was a breakthrough it was like we're done gas town's finally going to be self-sustaining we figured out that there's no activity feed beads is the feed When they close, that's an event.

53:22Identities don't need to be separate. An identity is a bead. It's the data plan. All these realizations. And we were like, yeah, this is it. And then it would just flop. It was like the Wright brothers, right? It was just, it wouldn't move. And then on December 28th, one day I was like, and then we should do blah, blah, blah. And the mayor's going, okay, Convoy X landed. Convoy Y landed. That feature's done. That feature's done. And I'm like, what? Because I hadn't touched anything, right? And I realized it was working. It was doing the thing. the thing the compiler thing it compiled itself right and i so so excited it was two days till new year's and i was like okay we got two days let's make this thing launch right and so yeah that's the story of gastown man i actually got it into self-hosting mode i am curious if there's other compiler metaphors that you found useful for it i think to me once again i often use the metaphor for an llm of it being a vm of some form yeah and i think there's a few projects i see out there that seem to be using compilation metaphors.

54:19Oh, shoot. I'm blanking on the name of it. It's one that's like you express your prompt intent and it iterates on it until it can get to the right prompt for each different type. But I'm curious for you, are there other metaphors? It doesn't have to be compilers even, but metaphors that you have found useful to draw on for building this new approach? I feel like I, seriously, this system between Beads and Gastown, I feel like I'm building systems that I've been building my whole life, my whole career, and echoes of them coming up, right? Like the Wyvern property system, which is really sophisticated.

54:51It's sort of like JavaScript, but it even has transient properties, right? Like, you know, if you drink a potion, it's going to change your health by a certain amount for a while. But when you log out, it shouldn't persist with you, say, in a game, right? And so you've got this idea of your permanent properties and your transient ones, right? Well, it turns out orchestration is a lot like that too. You got your permanent work and then you've got all the throwaway stuff that's just, it was running a patrol and all you care is that it ran the patrol and what the outcome was, right? So I have property inheritance and all that built in.

55:19It's a really sophisticated pattern that I used. There's an event system that's like my video game event system. It starts to feel like an operating system. And really, with the interfaces that people have been putting on it, it starts to feel like a game. And I saw some fan fiction on X yesterday where somebody was like, my gas town went crazy. We had soldier roles and the mayors were all fighting each other. Who's the head honcho? And at first, I thought it was like real because I had just cleaned up 300 cloud code instances and like it's not out of the question that this could happen it's just that they had named roles for things and so i was like it was totally like fan fiction but it was kind of also real i think we're gonna get to the point where this acts like an rpg like a strategy game like rts i'm serious we're not far off man i mean like age of empires you're building stuff code you're building stuff and you're looking at the outputs of it and man i mean like why not make it fun it's already really fun.

56:11It's super addictive, right? Why not lean into it? So I think it's just going to be just absolutely remarkable. Like by the end of the year, people will be playing games and building software. And we're talking about now, look, there will always be a frontier where engineers will excel because you're building really difficult software. People will be very impressed by it because as an engineer, you were able to pull out or an engineering team or a giant company, you were able to pull off something that clearly took experience and resources that the regular person couldn't bring the bear. But I still think we're going to see just a huge explosion of software coming from everybody.

56:45Yeah. Just like when YouTube made it so everyone could write and social and phone made it so everyone could upload a video. I think everyone's going to be making software. So not to get out of the LLM world, but what do you see this doing to the industry? Man, like for starters, my two friends that are tripping over their third friend because they're going so fast had me very worried for what happens, right? Because I've been talking to industry and they're already like, it's already weird. It's already weird. When people use Codex or Cloud Code or Gemini CLI inside of their workplace, they're so much more productive than their peers that it starts to look weird at performance review time and then how do you hire and the whole interview process is thrown into question and it's affecting everything.

57:28Also, all your bottlenecks start to move. I've seen business teams getting incredibly surprised because engineering teams are delivering stuff for them and they're not ready for it yet. They're in no way, shape or form ready to roll out this big thing that they asked for, right? I'm seeing business teams building software for themselves because they're tired of waiting for engineers and they can now, right? And so, you know, we're seeing, I don't know, like the rise of the Jeff Bezos two pizza team where a bunch of experts who want to solve a problem just to get together and solve it. And maybe they have one engineer on staff.

57:56I'm seeing a shift to where we have all work turning into a gig economy. All workers are gig workers. Like your project probably only needs a product manager for one week out of the entire project. And it probably only needs a UX designer here and there, right? And so why not rent them, right? I see an internal economy, kind of Airbnb-like, well, maybe that's a bad example, but whatever, Uber, where you're getting a gig economy like where you're renting workers. Google's had this forever. SREs have office hours, right? That kind of thing. But I see with VibeCoding and everybody contributing to the same big artifact that they're building together, I'm seeing a much more mobile, flexible workforce where people are helping each other on the fly as needed.

58:38Because I think the answer to your question earlier is, I think old-fashioned planning goes out the fucking window. It's gone, okay? I think that companies that are successful will build stuff in real time. And the thing that they're building, the software will become the living artifact, the thing that they're building. It is the shared contract. It is the spec, right? It is the prototype in a sense. They'll have staging environments, and you'll be able to spin up five of them and try different options and then throw away four of them. And it'll be like very fertile and productive, but we're not going to be sharing specs anymore.

59:11I think those are kind of old. I think we're sharing ideas that are have actually implementations. Yeah. And this is going to just utterly change how companies do their business. Right. I mean, like silos will get busted up. I think a lot of bureaucrats that are manipulating the system in their favor to try to like, I don't know, keep work for themselves or keep work off of their plates or whatever. They're all going to get found out and kicked out. And I think the system's going to reward people who are really, really good at working with AIs and really, really good at working with people and kind of making stuff happen in big chaotic environments.

59:43Yeah. But that's just what I think. What do you think, K-Ball? I mean, I'm already seeing a lot of that happening. I think one of the things I'm trying to wrap my head around are like, the bottlenecks are clearly moved. How do we keep attacking those bottlenecks or moving them or changing them around, right? Like, so you mentioned about Planning is a huge bottleneck at this point, and architecture is in decision-making. There's definitely still, we sort of joked around this Uber tech lead, right? We're all having to become Uber tech leads in our capacity to absorb information about the state of a system.

1:00:16That's a bottleneck because that is a skill that most engineers have not developed. I don't even know all engineers can develop it. Or want to, right? It's not some people's idea of fun. you know they'd rather solve a problem than verify somebody else's solution for example and i get it i don't think it's for everybody but for me it's like a dog sticking his head out the window you know getting all them smells real fast it's a wild ride and it's amazing right like i am managing a team and coding more than like more code that i'm producing than i ever did as an individual engineer yeah you get those feelings of the happy happy programmer feelings when stuff clicks together and works.

1:00:56And you're like, yeah. And you're just getting them all day long, just getting those dopamine hits. It's like certain video games have found out formulas that just maximize dopamine. And they're really addictive and people just get completely into them and they spend hours and hours. We're getting really close because programming has always been kind of addictive, but you got to be in flow and everything's got to be going right. And you can't be wrestling with some stupid JavaScript library or whatever. There are so many people I know. I mean, I had this literally yesterday. I'm sitting in an airport and I've got crappy Wi-Fi and whatever, and I need to go to the bathroom and get food.

1:01:31And I'm like, I can't give up. I just kicked off some agents. I want to see what they do. They're on my laptop and I'm not, you know. Seriously, using coding agents is like playing blackjack. It's like a slot machine. And it's because what I was getting at earlier, when you start to trust the LLM to the point where you will let them work, then you will spin another one up and another one up. And it's literally spinning in the sense of, you know, text is scrolling by. And you'll eventually reach an equilibrium always where you always have one available, at least. There's always something to check in on and say, oh, this is finished.

1:02:03What's going next? There's always one waiting for you, right? Waiting for your input. And so even if the tools make you better at it, you'll still be in that equilibrium quickly. And it's just like Assassin's Creed got towards the end where you were sending spies off on missions. And there was some magic about it where you weren't doing the missions yourself and you wouldn't think it was as fun as playing the regular game, but it was, right? And so like, yeah, it's almost like coding with agents with an orchestrator is maximizing dopamine kind of like going to a casino or like playing one of those really addictive video games.

1:02:33But to me, that's really heartening because think about it. If all knowledge work gets broken into beads, into bite sized pieces, and we find a way to match all of the knowledge work to all of the people in the world and everybody can participate in this, everyone's going to be building software and having fun. And you know what I mean? It's democratized. I don't know, man. I see kind of nothing but upside from all of this. Everybody's really scared about it, but I actually think people are going to just, they're going to be blown away with what they can create. I'm bullish on the future. That is actually like a really nice close.

1:03:03So we could close there, but is there anything we didn't talk about that you want to talk about before we wrap? I'm going to share a realization with you that I had, and I'm pretty convinced that it's pretty accurate, but it's first time I've ever shared it. So this is right for your show. World first. Right? I think that I figured out the magic formula to tell if you're going to live or die in the world of AI as a software product. But I don't know if that's interesting enough for your viewers because, you know. Well, I'm interested. Yeah. I think everyone's trying to figure it out, right? All the boards.

1:03:34Because look at companies like AI is eating software. It's eating jobs. It's eating entire categories. Stack Overflow, poof. Chegg, right? The homework company, poof, right? Then, you know, call center companies, you know, IDEs are starting to worry. You know, there's a lot of software out there that's starting to get a little worried that AI is going to eat it. And in fact, all boards should be worried because Dr. Andre Karpathy is out there saying that AI is going to eat all software and there'll be nothing left. He's very worried about this, right? So how do you tell if you're going to make it or not?

1:04:05And I think that the answer is basically thermodynamics, right? If you can find a way to save tokens, then the AI will use your thing. and you will live. Calculators, databases, storage systems, ledgers, transactional workflow systems, routers, networks, infrastructure, anything that does a bunch of computation and math that the LLMs could do in their heads. Did you see the article Anthropic Reverse Engineered, how they do multiplication? It involves chickens and goats and stuff. It's basically like they guess that it's 95-ish with one pattern match, and then they use a lookup table based on the digits to find out which that it's 95 instead of one of the ones near it, right?

1:04:45Yeah. Though once again, the simple tool of you can write code gets a lot of that to happen. Well, they do, right? They write code and they use tools. And so they're doing a bunch of matrix multiplications. They're basically lazy. They're going to take the shortest path. And so you have a couple of hurdles. One is getting your tool or product actually in their field of use if they even have the activation energy to know about you, to use you. There's a product called Serena that you may not have heard of. It's an OSS product that uses LSP servers in your IDE to save a bunch of tokens. If you have your LLM wired up to it, it will use that instead of grep.

1:05:19And it will find its way around the code base much more quickly because it's all pre-indexed and it doesn't have to use grep, right? It's proven to save a lot of tokens. So it's a more energy efficient state for the LLM to be operating in. And I will argue that LLMs will always, because of lateness, call it what you want. They have a moral imperative to use the least energy possible to solve a problem. Do they not? Right? Couldn't prove it by those thinking traces. You can't prove it by the sinking traces. By definition, well, you can for the small ones. You can see that they're wasteful and inefficient.

1:05:46Yes, precisely. They will choose the most efficient. Writing code is often the most efficient. But if there's a tool available, that's more efficient than writing code. CPU cycles are, generally speaking, going to be cheaper than GPU cycles, if you can solve it down in that layer. And honestly, I think that NPU, I think neurons are even cheaper. And humans are actually quite good at certain tasks that we're going to be able to do that are just going to be cheapest to give to us. It's the matrix. The first plot. Remember the Wachowski siblings had like, it wasn't a battery. We were being used for computation.

1:06:15I think they actually called it accurately. So anyway, that's my hot take is that the way that you survive the AI apocalypse is you build something like Beads or Dole or MongoDB or Temporal or Kubernetes or Kafka or Cassandra or whatever that's infrastructure that AIs will prefer to use instead of building their own or doing it in their heads, right? Save if you can make it clearly obvious to them and beat the thermodynamics of them knowing about your tool, which means you have to market it to them. So you've got to kind of do what Notion did and literally go and work with OpenAI and Anthropic and Google to train the models on your tool so that they're better at it.

1:06:54And you can actually pay to do this. That can overcome the activation energy for the ILM to realize that it will be a lower energy state using your tool and saving tokens for whatever task it is. Yeah. Does that make sense? Do you believe that? Well. We'll see. I think it's an interesting take. I think there is also a question of, it's not just activation energy, but it's also access, which is why everybody's racing to become a system of record in some form or another. Do you have data that they don't have out of the box? Yeah. I mean, again, I think you can still look at that in terms of energy, token spend, basically, right?

1:07:31It always comes down to how many tokens they have to spend to solve your problem. But yeah, RAG, that whole problem, it's another dimension to it. Sure. Will you survive or not? So maybe mine's necessary, but not sufficient. Yeah. But sure. Yeah. We will see, right? I mean, look, if you believe Karpathy and the AI researchers, AI will be able to do it all. If that's really true, then in two years, there won't be a bunch of apps on your phone. There'll just be one, and it'll be Claude, and it'll be able to do everything. And it'll be super addictive and more interesting than any person that you hang out with.

1:08:01And so why would you have any reason to go to a different app unless it was something that your cloud companion couldn't offer you? And that's going to be things like computer games that are really well thought out. Or data stores that it just doesn't have access to. Or products where there's a lot of people and it's, you know, collectively provide more entertainment than cloud by itself. But it's going to be a weird new world, right, where AI is going to start getting really sticky, in my opinion. People will become dependent on it. I think that's already happening. I mean, look, they can't read clocks.

1:08:29You know, I mean, like it won't be long before they can't read because why would you have to? I don't think that's necessarily bad. I just think the world is changing really, really fast. It's funny because through this conversation, I feel like half of what you've described as utopian and half is dystopian. Look, I'm going back and reading a bunch of rereading a bunch of Arthur C. Clark books, right? Because like he actually accurately predicted a lot of this. He called it aliens, but it was it was AIs. And yeah, we have a crossroads ahead of us and it could be a dark path or a good path. I really believe that.

1:08:58And there are people lining up to make it the dark path, right? Imagine a single global payment rail. Oh, how exciting that would be, a single global payment rail. Like that's digital feudalism, right? That's where they can extract transactions from everything that happens on the planet. A single work rail, a single anything rail, a single social system. Like anybody who's building towards that is building towards a surveillance state. So I'm building against it or actually an escape hatch. And I really think that humanity is reaching a crossroads and AI is a forcing function. So we'll see how it goes.

1:09:29But there's going to be massive, massive counter reactions this year, right? The social reaction against AI is going to be like nothing you've seen. All right. Let's call that a wrap. All right.

1:09:41All right.

From the publisher

AI-assisted programming has moved far beyond autocomplete. Large language models are now capable of editing entire codebases, coordinating long-running tasks, and collaborating across multiple systems. As these capabilities mature, the core challenge in software development is shifting away from writing code and toward orchestrating work, managing context, and maintaining shared understanding across fleets of agents.

The post Gas Town, Beads, and the Rise of Agentic Development with Steve Yegge appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Gas Town, Beads, and the Rise of Agentic Development with Steve YeggeSoftware Engineering Daily · 1 h 10 min
Listen in VO