Context engineering with Dex Horthy

15 Jul 2026 · 1 h 32 min · 32 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Context engineering for LLM coding agents, including “context window physics,” the “dump zone,” and loop engineering (e.g., slow/iterated loops). Also covers software factories, spec-driven development drift, and “harness engineering” (engineering around an agent’s tool/integration layer). Dex shares why agentic systems can break in production and how to verify changes with CI and debugging tools.

Guest backgrounds

Dex Horthy is founder of HumanLayer (coined “context engineering” before it became popular). He previously worked in platform engineering (Sprout Social, Aspiration), built a pre-Kubernetes container orchestrator (Replicator), and did customer-facing platform work that closed ~12 deals in 3 months. He later co-founded HumanLayer/Metalytics (data engineering pivot) and has spent ~2 years interviewing hundreds of AI engineers.

Key claims

Quality is limited by how you “push the right context” (token-level inputs), not just model choice. Longer contexts don’t equal better results; attention/instruction budgets degrade performance. Loops work when verification/backpressure exists; but “dark factories” that ship without human reading code degrade for 3–6 months.

Notable examples

Lightsoft “lights-off software factory” shut down after ~4 months; etcd bug diagnosis via Antisys casualty analysis; CI volume scaling needs parallel, volume-aware orchestration (Buildkite).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Dex Horthy's Journey in Tech

1:05 to 1:16

Dex shares his unconventional path from physics to software engineering.

“If you work with agents, your job is no longer just writing code, it's specifying and testing it.”

Dex Horthy's Journey in Tech

1:21 to 2:38

Dex shares his unconventional path from physics to software engineering.

“Today we're talking about pushing the right context into models so that they write better code.”

Dex Horthy's Journey in Tech

2:42 to 10:00

Dex shares his unconventional path from physics to software engineering.

“30-day all-access trial, no credit card, and an actual human engineer on standby.”

The Evolution of Software Factories and AI

10:00 to 14:00

Discussion on the rise of software factories and integrating AI into development.

“You know, I think a lot of engineers are afraid that if they go do a customer facing thing, they lose all their credibility.”

12-Factor Agents Manifesto Overview

14:00 to 18:56

Explore the origins and principles of the 12-factor agents manifesto.

“And one of them was this now famous 12-factor agents manifesto.”

Introduction to Context Engineering

18:56 to 20:48

Learn about context engineering and its impact on AI application development.

“toby looky from shopify he says i really like this idea of like context engineering and i'm like i I wrote about this two months ago.”

The Importance of Context in AI

20:48 to 24:42

Understand why managing context is essential for effective AI applications.

“And understanding that is a lot more powerful than trying to learn memory and trying to pick some agent framework off the shelf and some memory framework off the shelf.”

Harness Engineering and Its Definition

24:42 to 28:00

Uncover the concept of harness engineering and its significance in AI.

Understanding Context Engineering in AI

28:00 to 37:47

Explore the significance of context windows and instruction budgets in AI engineering.

“Yeah, I mean, so the longer context windows are good.”

Understanding Context Engineering in AI

37:51 to 39:11

Explore the significance of context windows and instruction budgets in AI engineering.

“the open source key value store used by Kubernetes.”
Show all 32 chapters

The Challenges of Loop Engineering

40:14 to 42:00

Discuss the implications and challenges of using loop engineering in software development.

“I think that with today's models, today's programming languages, today's infrastructure, you might get away with not reading the code.”

The Challenges of Code Maintainability

42:00 to 45:56

Explore the complexities of maintaining software and how code quality impacts long-term development.

“like make your code base more maintainable over time and keep you from falling into this trap of like, OK, well, now if I change something over here, I broke something over here.”

Evolution of the Software Factory

45:56 to 48:21

Understand the historical context and evolution of the software factory concept before AI advancements.

“Do you know the first definition of software factory the first time it was used?”

Agentic Software Factories and Their Implications

48:21 to 51:26

Discuss the implications of replacing human roles with AI agents in the software development process.

“without a drawing, but I'll try to draw it out.”

Future of Software Development with AI

51:26 to 56:00

Investigate the future potential and challenges of integrating AI into software development workflows.

“we're going to treat the whole system as a black box.”

Evaluating AI in Software Development

56:00 to 57:10

Discussing the evaluation of AI's capabilities in software feature development.

“to evaluate the quality of code that's written.”

The Evolving Software Factory

57:10 to 58:50

Exploring how software development processes are changing with AI integration.

“Some actually have the agents already one-shotting bugs.”

Loop Engineering and Human-AI Collaboration

58:50 to 1:00:20

Strategies for improving software loops through human and AI collaboration.

“Instead of not touching it, just literally saying every user reported issue becomes a PR through the loop.”

The Research Plan Implement Framework

1:00:20 to 1:02:50

Understanding the Research Plan Implement (RPI) framework and its impact.

“It was this technique that worked really well for hard problems and complex code bases.”

Spec-Driven Development vs. RPI

1:02:50 to 1:05:20

Comparing spec-driven development and the original RPI framework.

“And then it started to and you could edit it as well.”

Intentional Compaction in Context Engineering

1:05:20 to 1:07:30

Discussing the method and importance of intentional compaction in context engineering.

Navigating the Smart Zone in AI Models

1:07:30 to 1:10:03

Strategies for optimizing AI model performance by managing context.

“If you ask a model, like build a plan of steps to go build this app, it's like, cool, we're going to do the database and then we're going to do the services layer.”

Understanding Model Behavior and Context

1:10:03 to 1:13:02

Learn about the complexities of AI model behavior and context management.

“Like if you said something where you were angry or frustrated or just wanted to point out that it's done something wrong, it would respond with you're absolutely right.”

From Token Harder to Token Smarter

1:13:03 to 1:14:48

Explore the evolution of coding strategies from efficiency-driven to quality-focused.

“Let's talk about some observations on how software engineering is changing.”

Dark Factories and Software Automation

1:14:49 to 1:17:42

Discover the concept of dark factories and its implications for software development.

“But the full dark factory where you don't read any code, yeah, it's a good way to maximize your token utilization.”

AI's Impact on Code Quality and Development Process

1:17:43 to 1:20:25

Discuss how AI can enhance coding practices and maintain quality outputs.

“And like just today, it doesn't feel like like basically you need humans in the loop to be able to do that.”

Introducing Human Layer: The Future of Coding

1:20:26 to 1:24:09

Learn about Human Layer, an AI IDE designed to enhance collaborative coding.

“leads you to success, let's talk about your company that's, you've, you've just come out of stealth, human layer.”

Revolutionizing Code Collaboration

1:24:09 to 1:25:57

Learn how new workflows in coding can surpass traditional pull request models.

“I mean, what it reminds me is like what GitHub did to software team.”

The Importance of Location for AI Startups

1:25:57 to 1:27:43

Explore the advantages and community dynamics of being in Silicon Valley for AI startups.

“And I was like, damn, I learned a useless thing just for somebody's ego.”

Hiring Trends in AI Development

1:27:43 to 1:28:51

Understand the skills and traits sought after in standout AI engineers.

Challenges in Distributed Systems

1:28:51 to 1:30:15

Discuss the complexities of building collaborative platforms and distributed systems.

Recommended Reading for Software Engineers

1:30:15 to 1:30:39

Discover essential books that can improve your approach to software design.

“Yeah, and then you got Kubernetes a decade later.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Dex Horthy:What happens when you let AI agents ship code for months and no developer reads a single line? Today's guest tried exactly that. He built a lights-off software factory and four months later he had no choice but to shut it down as things just stopped working. Dex Horthy is the founder of HumanLayer and the person who coined the term context engineering days before Andrzej Karpathy and Tobi Luzka made it famous. He spent the last two years talking to hundreds of AI engineers about what actually works when you build with LLMs and is testing the most extreme ideas with his own team. In today's conversation, we discuss context engineering, what it is and the physics of context windows, including what the dump zone is.

0:34Dex Horthy:Loop engineering, from the Ralph Wiggum technique to the slow loop that Dex's team runs every night to wake up to code cleanup PRs. The rise of software factories, from a NATO conference in 1968 through DevOps to today's agentic factories. Spec-driven development and why specs always drift from the code itself. And many more. If you want to understand increasingly important concepts like concept engineering and harness engineering, or want to know how far you can push the let agents build everything idea from someone who pushed it further than almost anyone, then this episode is for you. This episode is presented by Antisys.

1:06Dex Horthy:If you work with agents, your job is no longer just writing code, it's specifying and testing it. And Antisys is the most effective method of verifying agendic code today. Today's episode is brought to you by BuildKite, the CI orchestration platform trusted by OpenAI, Entropic, Cursor, NVIDIA, Uber, Canva, and more. Today we're talking about pushing the right context into models so that they write better code. Right after that starts working, your agents will write more code. A lot more. Trusting that code avalanche is where many teams face a challenge today. Every change that an agent makes still has to be built, tested, and proven safe before it ships.

1:39Dex Horthy:Worked on my machine is not enough, so you obviously need CI. But when agents are pushing 5, 10, or 50 times the commit volume to your pipelines, faster CI runners won't save you. Shaving 30 seconds off a single build is meaningless when a queue has 100 plus jobs deep. What you really want is a CI system that gets faster as the volume grows, and CI that offers instant parallelization to give you unlimited concurrency and to intelligently route changes at runtime. This is what BuildKite does, and why global software leaders continue to rely on it. The same architecture that observed the scale of Shopify and Uber a decade ago now runs about 1.4 billion job minutes a week across Cursor, Meta, Reddit, and Snowflake.

2:18Dex Horthy:While the rest of the CI world are cracking under the weight of re-architecting their platform, Buildkite continues to reliably grow. Agents running on your infrastructure are Buildkite. Any cloud, any chip, your secrets, your skill. Every artifact and log is captured, so when something fails, either you or your agents have immediate insight for why. As you're entering the context you'll give to your agents, think about how you'll verify what they hand back. If your system is buckling under the increased volume, head to buildkite.com slash pragmatic. 30-day all-access trial, no credit card, and an actual human engineer on standby.

2:46Dex Horthy:His name's Ola, and he's very helpful. So Dex, welcome to the podcast. Super stoked to be here, dude. Before we get into some of the context engineering and some of the more spicy stuff as well, how did you get into tech? How did you fall in love with computers? Oh, man. So I was doing undergrad as a physics major, and I realized that I didn't like academia. And there's basically two or three paths out of physics. It's basically you go get a PhD, or you go into finance, or you go do programming. at that time this was you know 2012 2011 when it was like in the middle of undergrad and deciding what to do and i had done an internship when i was in high school i was working with nasa researchers to a jet propulsion lab in california they had just gotten this really high fidelity like the most uh you know fine-grained data set of altitudes like the heights of very very like top top top topographical map of the south pole of the moon and the south pole of the moon is really interesting because some of the craters there are so deep because of the angle it has it got hit by meteor storms like no other part of the moon so there's very deep craters that have never seen sunlight and so there's frozen liquid water in there from the formation of the moon and so scientists were really interested in getting down there and exploring and uh so we had this really fine grained map and it's like okay cool let's build software so that i i have point a to point b i know the limitations of my rover can you know max incline up is this max incline down is that find a path from point a to point b that doesn't like break those rules of incline so i was you know 17 i had never cracked a cs textbook so i wrote i basically like wrote a really naive bad version of dykstra's algorithm for pathfinding uh so i was in college i was like i don't know if i want to do the academics thing but i really enjoyed programming back in the day and so uh so i decided to go i got like half of a cs minor and then started working on a api platform team at a software company in Chicago and Sprout Social, right?

4:45Yes. And basically never went back. Yeah. And then where did you go from there? Where did you pick up like the parts of the trade? Because very early on, your first job, that's not really common. You were doing platform engineering back in, you know, more than a decade ago. From that point, it took me about two or three months to notice that like the most valuable work that was being done in the company was being done by like of course it's obviously like the first couple engineers who know everything and understand where everything was and like you spend a day on a support ticket from a customer and they solve it in five minutes but like you have to solve it so you learn and whatever and i realized like the most valuable people in the company were the people that were building the developer platform cicd sandbox environments preview stuff and so i kind of like that was my first step into the journey and i've basically been obsessed with software factories since that like three or six months into my first job.

5:38We talk about software factories now, but you're talking about software factories back then. So like you were starting to already think that this is how we can produce better software inside the pre-AI world, right? Well, and I'm always surprised. Like there's a huge class of developers that say, I don't want to work on CICD. I hate CICD. I'm like, really? Because building the thing that builds the thing and building the thing that builds the thing that builds the thing is like, as software engineers, we're lazy. We want to do the most high leverage thing that makes our job easier. So how do we, if we can build a thing that helps us build a thing that helps has moved faster, then that's the best use of my time as a lazy engineer.

6:11And then you went to another startup, Aspiration. Aspiration, yeah. Aspiration, also platform engineering. Yeah, I was brought in and then like three months into the job, the VP of engineering who hired me quit or got fired. I don't know, there was some drama about it. I probably shouldn't talk about it. And then I was there for about a year and was kind of like acting CTO for a while, like hired a couple of people, helped hire the new VP of engineering, but I was out of there. I don't think I'll ever do consumer again. I think I'm actually a B2B guy. Good to know. And then you went to Replicator where you spent like a good like solid like four years and went from engineer for the deployed engineer to product manager.

6:47Yeah, I did core engineering for like two years that we were building a container orchestrator like before Kubernetes, before Docker Swarm was really a thing. We built our own orchestrator. The founders had this vision that like, oh, Docker is going to make it much easier to ship on-prem software. And when I say on-prem, I don't mean literally like a rack in a colo. It's more like, hey, look, bring the app to where the data is rather than sending the data up to some cloud vendor. And Docker makes it much easier to package up apps and move them around. And so they had this thesis that basically you could build a platform that the experience that you get when you use GitHub Enterprise, which is like you install it and it has this admin panel, but then you just get GitHub running in your data center and your code never has to leave your data center.

7:28suddenly you could build a generic SaaS where everybody could have that. So I did two years in engineer there. And then our head of sales, we parted ways with our head of sales. And honestly, I was having a lot of arguments about the software factory with our CTO. And it's kind of like almost like a too many cooks in the kitchen kind of thing. I'm sure many listeners have had this experience of like, well, yeah, I know I have these tickets to build, but like CI sucks. I got to fix CI because it's too slow. Or it's like, there's too many different bills. It's always breaking. I'm like, I want to fix that.

7:56And then I'm going to do the end is just like, Dex, I need you to stop fixing the build pipeline and like do the tickets I gave you. I'm sure you've had this experience perhaps. Yeah. And then was this what led you to either forward deploy engineering? Yeah. So I like, I really loved our customers. Our customers, our customers are HashiCorp, DataStax, Puppet, all these really cool engineering brands, Travis CI, CircleCI. I was like, yeah, I actually love working with our customers. Our customers are awesome. And it was a great way to like get in the trenches. A lot of really good engineers who were solving the hardest problem at the company, which is like, how do we take this three to five year old SaaS platform and package it all up so that someone who knows nothing about our architecture can run it reliably in their own AWS VPC, in their own on-prem data center, whatever it was.

8:39And so I spent, I was our first kind of customer facing engineer. And it was about three months. I, we closed, I met with like every company customer that was like kind of in the pipeline but wasn't moving sales wise and we closed like 12 deals in three months and the ceo was like holy crap dex like the the investors are taking my calls again like i don't i know you want to get back to coding but like i need you to go hire three people and like build this team out because i think you might have been like born for this wow yeah so i did that for about four years built that org to like 25 people and then zerp happened and uh it got a lot smaller and we kind of realized like hey we have a product that's like pretty good uh and we've been solving what lots of early startups do is like okay there's some usability issues we'll throw we'll get a bunch of smart people throw them in the trenches with our customers great for sales great for retention all this stuff and it was like oh we actually like the margins on that aren't aren't good enough and so we basically were like cool we actually just need to make the product way more usable do a more plg shaped thing plg meaning product-led growth product-led growth make a little more self-service so you don't need an expert to teach you how to use it i was like cool if that's the most important thing then i want to go be a product manager because i have tons of opinions.

9:46I've now spent four years in the trenches with our customers. I have a laundry list of roadmap things that I think would make the product way easier to use and adopt and implement and deploy. And now you went to the full argument towards a dark side. Exactly. Yeah, I did. I was like, this is going to kill my street cred, isn't it? But I was really glad. You know, I think a lot of engineers are afraid that if they go do a customer facing thing, they lose all their credibility. And like, yes, I wasn't coding for 10 hours a day. I was coding for like three or four hours on a Saturday for fun. But I mean, we were helping people build yaml we were building clis we owned a lot of the tooling that customers use but it was like the last mile delivery side of it not the core platform and like on a more personal note i had spent the last like most of my 20s feeling like okay a little bit introverted a little bit like socially awkward what i what a lot of engineers i'm sure experience and uh i had talked to my uncle's a music producer so he used to work with like randy newman and a bunch of like really famous musicians.

10:40Oh, wow. Yeah. This guy, Mitchell Froome. And he, he, I was sitting with dinner with him at some point. And when I was, I think it was when I was still in undergrad, but he gave me this lecture. He was basically like, if you want to be really good at something, you have to make it the only thing you do. The guy playing guitar nights and weekends, trying to get his band off, off the ground will probably never achieve greatness. The people who become great are the people who basically make it like, if I don't play guitar, I don't eat. And you go and you sit on the street all day and you play for 14 hours a day or whatever it is that's the only way to become great so i said okay instead of trying to like read self-help books about how to be less introverted and less socially awkward like what if i just made it my freaking job to just talk to people and make friends and like help people and solve their problems and uh i think it worked out i recommend it i think everyone should spend a year or two at least doing something really like customer facing did you do this because you felt that it was holding you back being introverted or like what what what and i i know you got the motivation from the whole musician motivation i get it on one part but what was it that you said like is a customer facing thing that i'm gonna be doing it because clearly you were pretty great at like writing code by that point you could argue you were doing it night and day so where did you find that like i actually i think like customer facing or like getting this introvert off of me did you feel that i was holding you back or you just wanted to be good at it it was just kind of a thing that was like interfering with my like general life satisfaction and it was also like I'm not a very type A person.

12:05I'm very disorganized. I don't know if people call it like, okay, I'm like ADHD now. That's why I can run 30 clods in parallel or whatever it is. But it was like, I was really bad at email and calendars and spreadsheets. It was like, didn't care about these, didn't understand them. And so like another side effect of this was like, it just forced me to be organized and keep a lot of things going. And so like, I don't know, there's like weird benefits you get from like stepping outside your comfort zone and learning like industrial disciplines that are separate from what you've been doing. And so the opportunity presented itself and I was oh, I like working.

12:32I'll try this for a little bit. Started going really well. I'm like, cool, let's keep, let's see, let's see how far this thread goes. And then afterwards, you're now in your second startup, you became a founder, and you also got involved in AI pretty early, as it was even before it was so obvious that it would change how it would change how we develop software, right? Well, I would say I was, I was later than I could have been, because we started the company, me and a buddy in Chicago started a company in the data engineering space in about 2020, November, 2020, we decided in like August of 2020.

13:03This is metalytics. Metalytics. Um, technically still the same company as human layer. We just like pivoted the, the, the mission, but, uh, yeah, the, the, the advice I got from every angel investor that, you know, people who just knew CTOs I'd worked for before and stuff, they were just like, look, hitting a lot of heads wins. And I don't know if you know, like the whole DBT data engineering, five trend, that whole arc where it was like this huge party and tons of investor money going into all these different companies and then within by like 2021 2022 there was kind of the zerp thing and just this general realization that the tam for those sorts of tools is not as big as everyone thought it was market yes the total addressable market for those sort of tools was was not as quite as big as uh as we all thought it was um so it was a hard place to raise money it was a hard place to get customers yeah and then i i met you at while you were at human layer nsf at an event we you actually talked and we chatted afterwards.

13:54But by that, this was about a year ago, you were already, you started to have some really strong opinions on using AI. And one of them was this now famous 12-factor agents manifesto. Are we calling it a manifesto now? I'm calling it a manifesto. It's a manifesto, I'm calling it. Let's talk about this. This was 12 engineering principles to build reliable production-ready apps. How did you come up with this? And maybe we can also talk about some of them. Yeah, so I'll kind of like go to like around August, the co-founder I was working with, they kind of burned out and left. And it was very, we're on good terms.

14:29It was very mutual. And I decided to just start messing with AI stuff. And I was building AI agents. And what was really in vogue right then was like the Lang chain, the crew AI, these like agent frameworks. And it seemed like there was a ton of, you go in the crew AI discord, there's 10 ,000 people. It's like, okay, this feels like the right shape. And there's clearly this eco. You go in every single one of those projects. They have a ChromaDB plugin. They have like a Composio plugin. There's like, clearly like this is the shared interface that everybody is building for. I said, okay, what's missing from all of this?

14:58The agents can call tools, but it's really hard to like control which tools they call. And if it's a chat bot, obviously you can show a prove deny in the UI of your application. But I kind of was obsessed with what I would call like outer loop agents or proactive agents that would run in the background, get triggered by events. I mean, OpenClaw is basically like the biggest manifestation of this of like you have a heartbeat it wakes up it sees if there's any work to do it tries to do stuff and my thought was like i'm not going to trust that agent to do anything meaningful if i can't get like a slack message or an i message or something when it wants to do something and kind of guarantee deterministically that i can approve or deny that or deny it with feedback and say actually no do it like this so we played in that space for a while and talked to a lot of founders and founding engineers and builders we came and did yc in the fall of 2024 with this idea.

15:49We're building out this API platform and it was sort of like pager duty, but like it wasn't who's on call to fix the servers. It was like, who's on call to this like routing mechanism for like who needs to approve this agent and can they like escalate it or delegate it or defer it, all this stuff. And we built it for this ecosystem. Great. I like chain five. There's so many grip tape. There was so many in that in that time. And then I talked to tons of AI engineers who were actually building really interesting things and like actually making money doing six figure contracts shipping ai to the enterprise and all of them had tried that stuff for like a month or two and then they had thrown it out and they were just writing all the api calls by hand and they were building more things that look more like pipelines and workflows than these sort of like hands-off call tools in a loop kind of thing and so i talked to a hundred people and i spent a lot of time a lot a lot of time hanging out with one of my best friends uh vibov from boundary so they built a programming they built like there's like proto buffs for ai thing and And I think they're about to launch their full-fat programming language Turing complete thing.

16:49But he had this way of thinking about agents and building with models and building with inference where it was a lot more about understanding what structured output really is under the hood. And every single step in your AI workflow is just tokens in, tokens out. And your job as an engineer is figure out, okay, what tokens do I need to put in to maximize the chance that the tokens out are going to be good? And kind of distilled all these ideas into about 12 principles and wrote about it on GitHub, posted just like this like 12 page GitHub repo, threw it on Hacker News, got like 500, it was on the front page for like two days.

17:21And I think it really resonated with a lot of people. Yeah, so I'll just quickly read the 12 principles and then let's talk about like one or two that resonate. So the 12 are natural language of tool calls, own your prompts, own your context window, tools are just structured outputs, unify execution state and business state, launch, pause, resume with simple APIs, contact humans with tool calls, own your control flow compact hours into context window small focused agents trigger from anywhere meet user where they are make your agent a stateless reducer haha the stateless yeah the stateless reducer one was a little actually someone hit me up on twitter and uh corrected me it's actually it's actually a transducer because there's technically multiple steps in the workflow but there we go but but uh but of this one this this was a year ago so like which is like forever and uh and and how the tooling is is evolving which ones still stick with you or if you're like all right these were good that that still seemed to hold off yeah i think i spent most of march writing it published this in april uh and then swix hit me up from ai dot engineer and he said hey can you come you want to come talk about this so i gave this talk 12 factor agents in like june 6th i think and uh small room maybe like it was packed but it was like maybe a hundred people that was the year at ai engineer where like the lower physically like on the on the second basement floor was all the super corporate stuff and you go up a level it's a little bit more and then like on the top floor is all the like weird cutting edge like startup stuff that like you probably shouldn't care about yet kind of thing so we were up there on the top of this like weird way of thinking about agents uh and then about a week later two weeks later uh toby looky from shopify he says i really like this idea of like context engineering and i'm like i I wrote about this two months ago.

19:04This is great. Toby gets it. And then a week later, Andre Karpathy is like, well, I really like, I think what we should think about is not prompt engineering, but context engineering. I was like, yes, that's my, anyways, I don't know. If you ask Gemini, depends what day it is. They will tell you either me or Toby or Andre came up with context engineering. You can't really own a word. Like, I don't, no one remembers who invented the word prompt engineering, but of all the factors, factor three of own your context window. And basically the only way you can, whether it's agentic or a single step at a pipeline, the only way you can impact the quality of your output from AI is by caring a lot about what the inputs and crafting them.

19:38So let's talk about context engineering, which I am going to credit you that you coined it. I did some research and like, I think you were earlier, but a few days. So there we go. You coined it. We're adding to the, we're adding to SEO juice. We'll have it in a transcript. Dex coined context engineering. Well, and like asterisk on that is basically like I learned about context engineering from talking to these hundred engineers and founders. I just kind of like what was the same about what they were all doing? And I put a name on it. So like I didn't invent doing it. I was just like, I think I think there's this thing and like vocabulary and names are really important and having like clean ways to talk about the problem, especially when like a lot of the content about AI right now is so much hype and jargon that is like meaningless.

20:18I was like, okay, I think there's a word here that is useful to builders that explains how they should be thinking about building their software. So what is context engineering? It's kind of like de-abstracting a lot of the abstractions that have been layered on top. So you have a rag, you have memory, you have agentic history, you have structured output. You have all these things that are like different ideas in the frame of agentic programming. And at the end of the day, they're all like different ways to pass tokens into a model and ask it to produce usually some structured output. And understanding that is a lot more powerful than trying to learn memory and trying to pick some agent framework off the shelf and some memory framework off the shelf.

20:59I mean, those are these things are all really good. If you want to get to like 80 percent, you want to get a really good demo. But when you have to go from 80 % to 95 % or 99%, you need to go down a level and think about what's everything we're putting into the context window. What order is it going in, depending on which model we're doing. And all of this stuff matters. You have all of these levers that you can pull. And it just felt like the right abstraction for thinking about how do I get AI to do the thing I want as accurately as possible. Why is context engineering started to become more talked about?

21:31It was about a year ago. Did it have to do with the context window that we could pass onto LLMs pretty much? Did it start to expand or did we just start to realize that we can do a lot more by passing on from, you know, the easiest one is, of course, system prompts. But of course, whenever you build an LLM behind the scenes, you will pass additional context as well. Not just to prompt the user, you will add a bunch of subdesk, I guess, a dirty secret of any LLM. But why do you think the focus is moving on to like, all right, context is important? I think it always was important. I think what had to happen is a ton of smart people, again, like all these builders I talked to, a ton of smart people had to like focus really hard on producing.

22:12Like I want to make software that I can sell. I want to make something that's accurate enough that I'm proud of and I can sell to an enterprise and they're going to be happy with it. And there's just like the easiest way to get to really high quality applications is by thinking at that token level, thinking about a string of different LLM calls, like rather than just tools in a loop and it's kind of open ended and very flexible, but not that reliable. thinking of agents as as workflows as pipelines as some mix between maybe a couple tools in a loop versus just hey i have my tools and i have my model and i have my system prompt and these are the only levers i have and it's actually no you have way more levers it's going to take more work and you're going to have to like understand the lm with a deeper intuition but it was a thing that we always needed and it just took time for people to build with this technology to figure out that like this is the layer of abstraction that allows you to break through the quality ceiling and how are cost and context engineering connected yeah um i don't know i was i was talking about this with uh someone this morning um about like when you're working with lms one of the things i like to say is kind of like make it run make it right make it fast see if the world's best lm at the time i think we did a podcast episode that at the time it was like oh three see if oh three can solve your problem and then give it to people and see if they want that and then if people want it and you use it a lot then go do a bunch of context engineering because your engineering time is always the bottleneck like humans trying to figure out and solve problems and build evals and improve and try different dimensions or set up jepper or whatever it is is always going to be more expensive than just using a smarter model until you have millions of requests a day and then it's like okay we're gonna do a bunch of context engineering break this up into three calls and get it to work on gpt4o and then we're gonna take two of those and make those two work on gpt4o and i'm using old model names.

24:01But the point is like for a certain task in your workflow, can you get GPT OSS 120B, which is like one one thousandth of the cost of Opus? Can you get it to solve parts of the problem so that the tokens and the things you're using the smartest frontier models for are just the things that you really need that level of intelligence? But you shouldn't go build all of that and overengineer it until you've proved that you need it, that it's valuable, that it's like, OK, this is now I'm going to get to Eli Goldratt. And like, what is the the he had this book the goal right it was about how to model your factory and i'm sure we'll get to that when we talk about software factories it was like what is the bottleneck in your system and one day it will be latency and cost but it's probably not that when you first start out and context engineering is how you move from the you you add human effort to the equation to improve the efficiency the speed the price the cost efficiency of your system interesting and then one thing that came up more recently and a lot later uh as soon as recently is harness engineering what is harness engineering so i made a post in like october i think about or maybe november of like hey there's this new thing that i see is like i'm calling it harness engineering my definition that i had at the time is not what actually this guy viv who's at lang chain now does a lot of really good writing on agents and how to think about harness and he had written something called harness engineering like a couple weeks before me but i hadn't read it at that point and my take was basically like okay when you build an agent you use heart use context engineering when you use an agent because we gave this talk in august of 2025 about like how to apply context engineering to how you use coding agents and that kind of evolved into this idea of like how do you take a harness like cloud code like codex how do you engineer against the integration points of that harness so commands mcps skills how you organize your code base how do you kind of optimize the environment that the coding agent runs in to like get the best results the same way with context engineer how do you optimize the inputs to every single prompt well harness engineering just is like how do i raise the floor so that every single turn of this thing the results are as good as possible and the term got super blurry and some people think harness engineering means building a harness and some people think harness engineering means building around a harness i actually like what martin fowler came up with uh as usual he's very good at naming things and he kind of defined the you have the lm and then you have the inner harness which is like the thing that the tool definitions and the integration points that like say like a cloud code or a codex or a amp actually exposes that's your inner harness and then you have the outer harness which is the stuff that you the human do to customize that for your specific needs your code base your languages etc that's the best definition i think we have for harness engineering it's interesting how naming is still so important isn't it well it's like as soon as you name anything people or most people i'm actually surprised the context engineering still means the same thing to most people that it did a year ago and that it's even still relevant like that's honestly the craziest thing to me is like you wrote how many things were written about ai 15 months ago still matter or still interesting or are still like have good advice baked into them stuff changes a lot i think context engineering has been so long lived because it's it's grounded in the fundamentals of how transformer attention works and until we have post transformer models or linear attention or whatever it is, which who knows when that's going to happen, context engineering will be interesting and important to anyone building on AI.

27:25Can we talk about the physics of context? You had a tweet, this one, the context reality check. This is a graph of as you get to 1 million contexts, just the quality just drops, it goes down. What do we need to know about the context? Because again, we now have models that do have a 1 million context window, maybe we'll have even longer ones. But when you start to just put in more stuff into the context, it starts to become less efficient. Like what do we know so far in terms of from the practical perspective of like someone who is using the context window to add on a bunch of stuff? May that be MCP?

28:02May that be tools? May that be skills? May that be all of these things? Yeah, I mean, so the longer context windows are good. You can talk to it for longer. Like they're doing a good job. But at the end of the day, like especially when you had like Opus, it was like opus 4.5 and then opus 4.51 mil or 4.6 and 4.61 mil you're not actually getting a like smarter model like the intelligence of the model is is what drives its ability to attend to all of the tokens in the context window to figure out on the next turn which parts of this 100k or 200k context window are the most relevant to making the decision of like what is the next tool we call and doing that over and over again in a loop so i don't know there was some study that came out in 2025 which found that and again these are old models so like inflate your numbers but it was like frontier lms can follow about 150 to 250 instructions before it starts to drop off their ability to follow all the instructions just like drops off pretty quickly and i think laurie voss at arise i haven't actually looked at the data but they did a study with like the next generation models a year later and it looks like it's like much better the number of instructions you can get in but in any case you have like i split context engineering into like two categories you have like the the most people think about like the information budget of like okay i can do rag and i can pull out chunks of this document rather than putting the entire book into my context window i can just go grab the pages that matter but it's also your instruction budget is like if you give the model too many instructions and especially too many conflicting instructions and that's in your initial prompt and also like if you have a conversation you start going down a path and then you change your mind and you start going down and you're like, actually, I don't want to do any of that.

29:42I want to do this. It's like it's a lot of computation the model has to do to notice that it has to ignore that whole thing. And when both of those things are kind of far back enough in the context window that they're only half getting attended to, your likelihood that it's like actually going to like remember the exact instructions you gave it 100 ,000 tokens ago is like it goes down quite significantly. This is all very interesting because as engineers, we are expected when we're AI engineers, which now a lot of software engineers are, meaning you just use LLMs to build software. Underneath, there's an LLM layer somewhere.

30:16You're an AI engineer. Congratulations. But it sounds like the expectation is to be a good software engineer pre-AI, you need to understand how to write good code. And it helps when you understand a little bit of the underlying. We didn't need to do that that much over time, but it never hurts. But it sounds like right now we're in this phase that to be an engineer who can write an efficient AI system that use LLMs, you need to understand the dynamics of the context. You need to understand why stuffing your context one way or the other can introduce latency. And all of these, it sounds like it's kind of more of an intuition.

30:56And of course, there's some understanding. But from talking to you, you're like, well, it does this computation. like i know you know because you've tried it out right like i'm not i'm not a phd in machine learning like i couldn't actually go like draw a mathematical proof of how this works but we know attention is quadratic and the more stuff you put in the more it has to spread this attention out over everything this just feels like an absolute new area and like a little bit very different to like what we're used to like software engineering which is like pretty kind of like black and white right it compiles or doesn't compile that's true i mean there's a different kind of intuition i was talking about this earlier as well is like there's a different kind of intuition that you that you develop over years as a software engineer and uh there's many categories of it but the one i'll call attention to that is like a thing that you cannot teach you cannot do you cannot learn in a textbook the only way to learn it is like i know bad patterns in software because i have debugged them at three in the morning this is my buddy jake from netflix said this in his talk at ai engineer code it's just like there's no better way to learn what is good and what is bad and what works and what doesn't than suffering through the thing that doesn't work.

31:59Well, speaking of suffering through the things that doesn't work, a new paradigm that is spreading up is loops. Loop engineering, the idea that instead of writing prompts, just write loops, set up your loops. And this all started with the Ralph Wiggum technique where it will just, well, I guess that's an early version of loops that were just loops around. And now we're hearing what some of the biggest labs talking about that they're actually just doing loop engineering. what is your take on have you done some loop engineering yourself have you set up some loops and what do you think is good about it and what do you think is bad about it yeah so i think of loops as i mean this could i could ramble on this for 10 minutes this is an entire talk but i'll try to i'll try to lay out some high level stuff and then we can dig in wherever you think is most interesting we had ralph wiggum as actually a year and four days ago was the first time i saw the Ralph Wiggum demo and like Jeff Hundley was just like visiting SF and he just like came through and like dropped everybody's jaws with his like yeah I just ran Sonnet around the clock and spent six grand in six weeks and like I built an entire Gen Z programming language look at it compiles and it has a stage two compiler where the compiler for the language is written in the language itself and all that insane and the core lesson from all of that I think was the idea of back pressure which is basically and i think a lot of people were doing this for a very have been doing this for a long time which is how do i let the model check its own work how do i automate the process of getting feedback into the model and there's lots and lots of different flavors of this you can have deterministic linters you can have unit tests like part of what made the programming language easy to build with ralph is a programming language can be infinitely verified you write you write the code in the language you compile it if the compiler fails you go fix the compiler you run the program if the program fails you go fix the compiler like it's like it's very very verifiable and i think the lesson in loops engineering is like if you can make a problem very verifiable you can kind of like treat it like a black box and then have it loop because it will keep improving itself because of the verification loop is already there exactly and so like you can do this with cicd is like i do this every time i do it really so i'm like i'm higher the CI CD is slow cool go research the code base make a change make a pull request run the test see if it's faster try again run the pull run run the test push push to the branch check again see if it's faster and so it's like if it can verify its own work in a loop instead of design instead of saying let's try this approach or let's try that approach or success and being really back and forth you just say like my goal is to make CI faster and you tell the model here's the steps here's the five here's the five steps you're going to write some code you're going to commit it you're going to push it you're going to launch a subagent to watch the job until it's finished it's going to tell you what happened then you're going to decide what to do next and so that's that's like the very simplest example i have of like designing loops and you just said the goal which is cloud code and and i think codecs have both chip slash goal which is you just set the goal and it iterates until it reaches it or as long as it makes progress towards it exactly and so it's like if it's verifiable if you can measure this is auto research too auto research is like hey go make this model twice as fast and like it's just a prompt that tells the model like go to it over and over again and try things until it actually has good results so that's what i think of loops engineering i don't know we we do a very interesting kind of loops engineering where like the the challenge is like i think it's very easy to get very excited about building the thing that builds the thing or building the thing that builds the thing that builds the thing we talked about uh and so people say oh we need to like redo everything as this big like agentic first factory, maybe even a dark factory, and they're like redesigning their entire thing to be their infrastructure for the next five years.

35:43And I'm sure one thing we know of in engineering, and especially pragmatic engineering, is how can you make this more incremental? How can you make it more continuous? And a lot of people don't have the option to just, hey, I ran a Ralph loop for three days and it fixed every linter error at our code base. Here's a 60 ,000 line PR. Who wants to review it and who wants to sign off on merging and deploying it and that there's not going to be any bugs nobody so i think the the thing i'm most excited is actually like what we call like iterated loops or like slow loops where we basically have a cron job we have the loop the the structure of the loop is really easy it's like run this linter fix one thing commit and push and then we run that every night in our github actions and we wake up every morning to one pr that makes the code base a little bit better i like the slow loops yeah and it has two dimensions so you can add now we have a blueprint for it and actually kyle just shipped a skill so that you can build these yourself you can add more like feedback mechanisms so we have react doctor for the front end we have another anti-pattern that has no deterministic tooling but kyle's just like here's what good looks like here's what bad looks like go fix one thing and bring it as like prop narrowing basically we have a bunch of optional props and most of them don't need to be optional it's like here's how to make the prop not optional so that you know that the code just is like cleaner and easier to reason about and so you can add more conditions more things of like fix one thing i want to wake up to a pr so now we wake up to like four prs because there's four separate things and then the other dimension you can do here is as you gain confidence you can increase the scope instead of fixing one thing fix four things and so these are like other ways to think about loops where it's like something that's not a human triggers it to start whether it's you know an alert from sentry whether it's a user feedback like support ticket whether it's pm writes a ticket whether it's a test is failing any of or it's a cron it runs on a schedule but it's like the trigger should be something that you don't have to like press a button on and there's a defined workflow and it makes everything a little bit better.

37:38Dex Horthy:That's just described letting agents fix things without a human pressing a button. But what if a bug is too difficult not just for an agent but also for human to reproduce let alone fix? This is where presenting sponsor Antisysys comes in. I was recently pairing with the Antisysys team where we did a walkthrough of how they helped fix a nasty bug in etcd, the open source key value store used by Kubernetes. This is a bug that actually happened in etcd. The team noticed that the linearization validation assertion failed during the regular antithesis runs. This is not good because the linearization guarantees strong consistency, so this needs to be fixed.

38:11Dex Horthy:So what the etcd team did was run a casualty analysis inside antithesis. This generates this graph, which is a bug probability graph. Here, the x-axis is virtual time, and the y-axis is probability. Now we see that something happened just before virtual time 24 that caused a huge jump in the probability that the bug would occur. Going deeper, we can look at the entire set of timelines. Vertical lines going down represent events branching off from the same state. And the purple dots are where the bug happens. If we look closely enough, we see that all of the failures come from one parent branch.

Read the full transcript

38:46Dex Horthy:Gotcha! This is such a useful debugging tool. In the end, the team was able to figure out that process pauses were causing the bug using all these Antisys debugging tools. This non-deterministic bug was diagnosed in a deterministic way. How cool is that? Oh, and this is an actual bug that then got fixed in etcd. You can see the bug and the fix in etcd's GitHub repo. Honestly, the tools that Antisys built for debugging feel pretty darn futuristic, but they are also really powerful. Head over to antisys.com slash pragmatic to learn more. I'd also like to talk about our season sponsor, Sentry. Sentry is a tool I use for application monitoring on all of my projects, including their pragmatic engineering backend.

39:22Dex Horthy:I've used it for 10 years now, starting with when I worked at Uber. A neat Sentry feature I'm liking is their Sear AI agent, which helps investigate production errors. For example, here's an actual error I had in my application. I can just ask Sear what might be the root cause, and it brings context, and it can also make a plan to fix it right from the web interface. And the nice thing is how Sear also works great from Slack as well, not just from the web. One place I find even more handy to use Sentry is from codex or clock code using Sentry MCP. Also, you can set up neat automations like when a resolve Sentry issue resurfaces, you can kick off a cursor agent or GitHub Copilot agent to investigate the regression, read the relevant code, and open a PR with a suggested fix.

40:00Dex Horthy:I'm not a fan for using AI tools just for the sake of it, but I really like the practical integrations where I can fix errors faster and with more context. Check out Sentry at sentry.io slash pragmatic and start monitoring and fixing regressions today. And with this, let's get back to Dex and to agentic loops that trigger themselves. Now, you said we can get more ambitious and we can add more things to it, but I'm going to quote you with one of your tweets, which says, this may surprise you that this is coming from me, but I think we're in for a one to three year period where stuff might break at 3am and you're relying on loops to fix it and nobody understands what's under the hood and you're looking at it as an existential threat to your company.

40:40Yes. Yeah, that one was great. That one did a lot of numbers. It resonated. Here's the other side of it. I think that with today's models, today's programming languages, today's infrastructure, you might get away with not reading the code. The problem with loops is at a certain point, you're going to generate so much code that you can't read it anymore. This is the StrongDM Dark Factory. This is Ryan LaPopolo's Harness Engineering. Just spend as many tokens as possible. We tried this. We built a Lightsoft software factory in July of 2025. And by November, we had shut it down. I think it takes about three to six months of you shipping all the time with nobody reading the code before you realize like, wow, this is getting way worse.

41:21And it's easier to start over than it is to fix it. Like the models have made the code base so bad that it is actually going to be easier to just like rethink this from scratch. And maybe that's OK because we have AI and it's easier to rebuild things from nothing. And like usually when engineers say like, oh, we can't fix this. We have to rebuild it. the feedback is like no just refactor in place just constantly keep the code base getting better you mentioned what i said you'll notice what i said was not use loops to ship the features that users want we use loops to actually improve the code base quality and we read all the code because we care about how it's architected and we care not just about the system architecture but what i would call the program design which i think is something people are going to where are the interfaces where are the seams how are we doing dependency injection all of these things that like make your code base more maintainable over time and keep you from falling into this trap of like, OK, well, now if I change something over here, I broke something over here.

42:10This is the classic problem software engineering that like software engineering was invented in the 1970s because we realized we needed techniques for avoiding that problem of like this giant ball of spaghetti. And I don't think the models are smart enough. And I don't think we actually have the training and the benchmarking and the eval techniques to get models to write code that is more maintainable over time versus they're all trained on sweet bench and sweet bench looking things right all of the benchmarks are basically like here's a commit in django here's an issue that was filed around that time see if you can create the fix that the human created and it's django and it's apache and it's there's 100 repos in go and c++ and typescript and java and all these different languages but they're all it's like the problem with training models on maintainability is like the cost function of bad architecture and bad program design can't be evaluated by running the unit test because it hits you three to six months later when you're like holy crap like no one can make it's the software has become so hard to change is this not similar to how senior software engineers why it took years for someone to become a senior because typically and in some environments you became you can become a senior you're faster, typically fast moving where there's a bunch of issues and you have to keep fixing it.

43:25Sometimes, you know, some people are working in the same place for 10 years and they're still not that level. The point was, it just takes time for you to understand the small mistake that you make right now that snowballs into like something disastrous later and you get hit by it and you realize like, okay, things like, you know, like testing matters, architecture matters, tech depth can actually be a killer. You know, we don't talk about it anymore, but we used to talk about how tagged up kills or slows down companies so badly pre-ai that their competitors can overtake them or they're just like stuck with a two-year refactor not shipping any new features and the competition you know ships a bunch bunch of other stuff and now they're ahead and i will say like it is possible that gpt7 will fix this but if you are turning the lights off in your software factory and you're saying like hey you know what like we're not going to read the code it's fine the models are smart enough if we give it the right feedback and just throw enough tokens at the problem it will keep getting better this is what led to this tweet like that might work but if nobody read the code in three months and you replace all of your all of your like code review with loops of like hey if a user complains we give it to an agent if something crashes we give it to an agent if a if a pm writes a ticket we give it to an agent if a ceo writes an obnoxious essay about what we should be building in slack we give it to an agent yeah and then you stop reading the code because that's going to produce way too much like no one can read it and like the the PR reviews become the bottleneck.

44:44So you replace that with agentic testing and agentic, agentic code review. But none of these things have intuition for software architecture because we haven't trained it in yet. And so you're going to wake up one day and you're going to have an issue with this happened to us. And like we got through it and at the time of like, it was still worth it. It was like spent three weeks onboarding back into the code base that we had stopped reading three months ago, because no matter how much sophisticated expert prompting, we could not get Opus. I think it was Opus 4.1 at the time. we could not get opus 4.1 to actually find the root cause we had to go spend several days digging through the code and figuring out like oh there's just actually a primary key that's being routed through this whole thing that needs to be changed to a different type of object and it needs its own table this actually happened to you this happened to us yeah and when it happened i was like you know what that sucked that was terrible but we did it we solved it and uh it's still worth it's still worth not reading the code for most of the time at the cost of every once in a while i'm have to spend two weeks fixing an issue by hand.

45:39And I don't believe that anymore because I think the amount of code we're able to write now is actually like 10x or 100x. And I think the problem is just getting worse. So let's talk about software factories. Yeah. In your mind, because I feel it's an overloaded word, but what do you think of a software factory before AI and now post AI? Do you know the first definition of software factory the first time it was used? No. It was a NATO conference in 1968. Oh, Grady Boosh would know about this. Yeah, exactly. Yeah, great you should ask grady about it they talked about the idea of like okay you actually need to build a system of steps and like just like a factory floor you have like the coding part and the testing part and the validation part and the integration part we had no cicd we barely had version control like but you needed a factory and then it was adopted by like um toshiba and a bunch of companies and then the the next moment was like devops and you have like this idea of like okay we're gonna do cicd we're gonna automate we're gonna use chef and ansible puppet whatever all these technologies is like instead of having dudes running around data centers like resizing disks and stuff or clicking around the aws console yeah exactly it was like cool we build loops the server hits 90 disk space that sends an alert to nagios nagios triggers a chef run chefs makes the the disk bigger feedback loops right this has been around for a while and in 2018 i want to say this guy nick chalain who was uh he was like the cto or chief software officer of the air force he wrote this 100 page essay of hey the dod needs a software factory the department of defense yeah the department of defense and the air force and he called it dev sec ops factory and he said we need all the things that all of the good enterprises are using we need jenkins we need like code quality scanning we need security scanning we need cicd we need to be able to ship we're shipping once every three months or once a year we need to be able to ship every day like all these other companies and the way we do that is we actually embrace all these automations and technologies so that engineers are are the 90 of the issues are caught by automations instead of people actually like manually checking it or manually reading the code or manually integrating modules together wow talk about forward thinking in in the government i know no as i was surprised like oh nice like this is i mean and that was part of it is like hey look we're falling behind and like you know i don't know exactly all the reason but i imagine also about like attracting really good talent is like, hey, look, if we have like the modern software stack and we're building things fast and we care about efficiency and we care about people using people's time well, we care about them spending time on the hard parts of the job, not manually looking for SQL injections, like you could automate that.

48:11So this was software factories pre-AI. Pre-AI. Now I've heard the term a lot more because of AI. Yeah. Is it the same? Is it different? So this is really hard to say without a drawing, but I'll try to draw it out. At the core of a software factory, you have like a source of work most you can imagine a linear a jira a source of truth your object whether it's a spreadsheet or whatever it is you have like what stages is the work in yep and pre-ai you would take you know you would maybe do some architecture review planning you would maybe do some sprint planning and then people would take tickets off the queue and they would go build them and then you would make a pull request and people would review it and you would run ci checks and then you would send it to prod and then it would make contact with your users and your users would complain about stuff and that would go to your support team and back into your work tracker and it would crash and you would have issues and that would go into your monitoring stack and that would go into your tracker and that was your loop and then people would take stuff off the tracker based on priorities product managers engineering managers engineers prioritizing work and they would go and do that and the first change is like this long wind lots of phases and this is also why when like a developer shifts a bug but by the time it comes back to you it might be two or three months or even longer and by the time it gets fixed it might be a year or two and you know this is why when you're using a piece of software.

49:24It's like that annoying bug and you talk with customer support, but it's just a very like long latencies at each part of the factory, if you will. Yeah, and the step where someone pulls a work item off the queue and starts working on it is, you know, a couple hours to a couple days before it actually gets integrated into everything else and touches user. And that's in a great world, right? Sometimes you go build it and then you merge it and then it actually gets released three months later. But we're going to assume we're in a fairly modern, And like we're somewhere like the Netflix or a meta where engineers are capable of shipping a hundred times a day or a thousand times a day.

49:57But it still takes two, three hours to do the work. And now with an agentic factory, what you do is you take out that person building the thing and you replace it with an agent building the thing. And so you have orchestration to trigger things. You have a sandbox. You have an LLM. You have an inner harness. You have an outer harness, which is like the dev environment you build for the agent. And maybe you give it a browser. You give it a video recorder. if you use like things like cursor background agents that kind of built this outer harness around the inner harness that is the coding agent.

50:22And then you make PRs with that. Problem there is that like, okay, now it takes 10 minutes to do a build instead of two hours or two days. And so now the bottleneck is code review. So, okay, let's throw a bunch of AI agents at code review and let's do agentic testing so that like we can basically catch a lot of the easy stuff and humans are only focused on the most like important critical core parts of the code base. And then the next level up of your agentic factory as you do the top. It's like, okay, and then it gets deployed. It goes to prod and a user complains. You just hook your support queue right up to the agent.

50:53Someone complains about something, agent tries to fix it. And instead of looking at a ticket and then saying, okay, go send, you just close that loop. And instead, every time something goes wrong, you just get a PR. And then every time something crashes in Sentry or Datadog or whatever, it goes into the tracker. It gets picked up by an agent and you get a PR. This is the ramp inspect thing. This is the, the only difference is like, then you have so much code to review and people say, well, let's try turning the lights on. Let's just take all the human testing and review steps out and we'll say, okay, cool.

51:20If users complain, then it's broken. And if users don't complain, then it's working. And we're not going to read the code. We're going to use, we're going to treat the whole system as a black box. So you said you tried this out when it was like Opus 4.0 and you built the software factory it was running beautifully until it just blew up on your faces. How do you think of this model? Because I can see an ideal world where it works, but clearly we're not in an ideal world. Where do you think we are? Like, and could some of this actually work at some point? Or, you know, like, what progress are you seeing right now?

51:50And what is the today, the situation? Like, how much of this do you believe we can automate or should we automate? Yep. So if you know me, you follow my stuff, you know, I stand for three things. Number one is like cutting through the hype and the jargon and going, trying things and talking to people who are using things and figuring out which parts of this actually work and are valuable. Number two, we talked about words. I try to find and protect useful bits of language because I think it helps us all move forward. And when you take a useful word like agents or you take a useful word like software factory and then you semantically diffuse it, this is another Martin Fowler word, you make it mean everybody likes the word and it all becomes hype and everyone starts.

52:27Agents means nothing anymore. Agents could be a chatbot. It could be a Slackbot. It could be a coding agent. It could be tools in a loop, whatever it is. So I like to protect important, useful words and like help help us all like elevate the conversation out of that hype and jargon. And then I care a lot about going one level down beneath where I'm generally working. I think there's always this is the same thing with context engineering is like I was rarely actually going and like building LLMs or understanding or training LLMs. But knowing how they're trained and how Transformers works informs how you build at one layer up.

53:00and for the software factory my version of that is i spent the last couple weeks going really deep on uh reinforcement learning with uh verifiable rewards rlvr which is like this very productionized like it's not like rlh rlhf is still like fairly academic and pure rlvr is this like it's a machine in these labs of how we train these models and i'm studying like the benchmarks for coding agents and the techniques for training them and how we like give it a small problem have it solve it, delete the test changes it made, revert them, apply a test patch, see if it passed. And then even the frontier this year, we have like, we can get into this later, but like frontier code and SWE marathon, these new benchmarks that are supposed to be like better at evaluating models ability to maintain a code base over time and write maintainable code.

53:47Um, and they are better, but I don't think they're sufficient. But it was basically this idea that like the only thing that made cloud code good was reinforcement learning. And the dimension along which it got good was like we made a model, we trained the model and the harness together. And so the model got really good at calling the specific tools in that harness, really good at reading files, writing files, searching for files, all this stuff through doing these problems. And that was what made it feel so much better than all the other CLI coding agents that came before it. And so people are like, okay, that was so much better.

54:17And they're just going to keep getting better. But it's like, it got really good in one dimension and the dimension that they're not getting better in because it's hard, expensive maybe we need to like get a lot more creative with how we design these these verifiers and benchmarks is in how do i make code that in three months is gonna like improve the productivity of humans and agents mostly agents but humans and agents in the code base instead of making it worse over time and so you think that part is just missing we haven't seen too much improvement i haven't seen obviously no one knows what the labs are doing internally because it's all very secret But I think if we looking at where the bench, the benchmarks tend to reflect where the labs are, right?

54:58If there is no benchmark that can convey to me, did this model write code that is going to make my code base better or worse? The best we have is I think frontier code from the cognition team is really interesting. They have like did the test pass and then they have like two layers of model review. So they have a judge model that checks, OK, is the patch the model made similar to the patch that is like the golden answer set? So even if the model didn't write the exact code that the benchmark was expecting, was it functionally equivalent? And the next one is like a code quality review from another judge model.

55:31And that's better, but it's not sufficient. And this is why I also think agentic code review is like, yes, it will catch things and it will raise your floor. But I don't believe the model writing the code is the same model reading the code. And if you ask a model, hey, is this code good? It's going to be like, oh, yeah, it's great, comprehensive. It's got unit tests. You've tried this, I'm sure. and you say, okay, review this PR that my coworker wrote and tell me everything that's wrong with it. And I was like, oh, it has this problem and this problem and this is sycophantic and they want to tell you what you want to hear.

55:59And so like, it's really hard for me to trust a model to evaluate the quality of code that's written. And so I have some ideas on like, okay, can you build a benchmark where the model builds 20 features in a row and maintains the code base the whole time? And it doesn't know what features are coming. You treat it like a real product team where you don't know what you're going to build next week until you get there and you find out what's most important. and then can we try to evaluate, like, can we build a problem like that that's hard enough that most frontier models fail by issue six or seven?

56:29Is it fair to say that, you know, like we've had the software factory, like before AI, it was just like lots of loop. It was like the PM giving a ticket to the dev, the dev building and deploying to production, user customers using it, customer support getting tickets, and then creating PM triaging, and it just kind of goes around in this loop. Is it fair to say that the software factory of how a company, a team builds and maintains software that is changing because now everyone's replacing some parts of it? Maybe the least advanced teams will just be devs are starting to use cloud code or codecs to write faster.

57:07They're not spending as much time on there. Some others are also having the deployment, the feedback. Some actually have the agents already one-shotting bugs. Is it fair to say that the software factory is just changing everywhere? maybe at different speeds, but everyone, I think every team who is building production software, they're frantically experimenting, trying, and everyone's at a different pace. You'll have the AI native starters where most of this will have agents in them, and you'll have the laggers who are more cautious ones. They have agents in a few places, but not in the others.

57:39And I think that's the key. It's like, if you want to do loops engineering, you should build one loop at a time and you should keep them small and contained. Basically, I think everything except stop reading the code is really good advice. Take support tickets and turn them into tickets in your system and then maybe turn those into PRs. Great. The advice that I have and like what we kind of like are chasing at human layer is like, how can I add another checkpoint in that factory? So instead of having one human viewpoint where you're reviewing PRs and sometimes there are a hundred lines and sometimes there are a thousand lines, but it's quite a lot of effort for, especially if it's bad, especially if it needs a rework, it's quite a lot of effort for a human to be like okay this is wrong go change it in this way and then you loop back to the agent and then you come with another one and like doing a lot of loops on there once once the direction has been committed to it's really hard to steer off like you're better off just kind of restarting from scratch how do you build like controls and mechanisms around that and then my take is like if you do a little bit of human agent planning and like discussion before you hand it to the implementer whether it's i mean planning and specs whatever you want to call it again this is Spectre-driven development is another word that has become kind of very like muddled as far as what it means.

58:49But basically, how can we spend an hour before we start building so that the PR when we read it only takes 20 minutes because the code is perfect? Instead of not touching it, just literally saying every user reported issue becomes a PR through the loop. And then we read that PR and it takes six hours because it's back and forth with making changes and things. I'm all about like, let's find leverage. And so you basically, you have three options in the software factory world. If you're going to go all in on Ingentic software factories, you can turn the lights off and just let everything flow and pray that you don't create too much slop and pray that the next generation of models comes fast enough before you create a giant pile of ash.

59:29you can slow way down and read every PR and read every line of code. And then you're only going to really get modest benefits from AI because that becomes, I think you should expect maybe 30 to 50 % lift in productivity is kind of what I see when we go into teams. Or you can find the right leverage points where humans can actually, an hour spent over here in planning can save you four hours in implementation in terms of fixing and going back and getting the design right. And that's what I call like seeking leverage. You can find the right leverage points for the agents to guide the work. Then you can actually move like two to three times faster while maintaining a like 99 % like accuracy to like, if the humans were carefully writing this code by hand, how would it come out now jumping a little bit back to ideas i will come back to this this was earlier maybe it was last year but you had the research plan implement can we talk about the original research plan implement framework and then also what you've learned about this approach what what you got wrong about it yeah sure yeah so um i mean the first time we talked about rpi was in august of 2025 um and it was basically like the research was this thing of like hey before you go build anything go read lots and lots of code use a bunch of subagent sub-agents in parallel understand all the code.

1:00:44It was this technique that worked really well for hard problems and complex code bases. You just ask Claude to do a thing that would read three files and make a change. It would have no context. So you start the research. You don't even tell it what you're working on. You just tell it, hey, can you tell me how this system works and this system and how they connect together? And then you get a markdown doc out. And this is the context engineering part is like that would take 100 ,000 tokens of context, but you would get a 10K token doc out of it that summarized it. Then you would start a new context window and you would do planning.

1:01:14And the planning would be, and I actually realized like the plans that we were building last summer were actually terrible, but it would basically be this long, you would say, okay, now here's what we're building. Here's the research doc, build a plan to implement it. And in retrospect, now that we see like everyone is obsessed with how do I get agents to work for longer? I think the reason why in like May, June, July, August of 2025, that a lot of people became am really interested in planning was it was a very powerful lever to get agents to work for longer. If you said, build me a B2B SaaS for burrito delivery, you'd get like a homepage and that's it.

1:01:50But if you said, build me a plan, it would build out this big plan. And then in the next context window, you'd say, hey, here's the plan. Here's all the changes we're going to make. Go implement. It would actually keep going until the plan was done. So the plan was a really good way to anchor an agent and remind it that like, hey, you're not done until this is all finished. So that was the original RPI. And the plan doc, what was bad about it is it didn't give you leverage. The plan was every single line of code that was going to change, like in diff blocks and like all the new stuff to write.

1:02:15And so like people would review these plans. We recommended this. We told people to read the plans. We read all our plans. And then eventually I found myself like, I just kind of skimmed the plans. And so you're not really using it as a way to re-steer the agent. It's just kind of there. And then you go write the code and there's a crap. Some people would review the plans and the code. And it's like, okay, well, the plan was took you 20 minutes to read. And then the pull request takes you 20 minutes to read. And they're different. And so you actually doubled the amount of time you're spending reading code instead of like doing less of it.

1:02:43You've been anti-leverage. And hang on, was spec different development not related to this? The one that Amazon Kiro, for example, and GitHub workflows again a year ago did, which was it also it first generated a plan and it had the human review it. And then it started to and you could edit it as well. and then it went off and implemented this part and it looked beautifully on the surface. It should have worked great, but it's tossed into the garbage outside of some maintenance projects, I think. It just didn't work. All the feedback I got, people just stopped using it because it just didn't really work that well.

1:03:15It just rhymes to the RPI framework a little bit, the original one, right? Well, so our thing too, the biggest difference between RPI and spec-driven development, and some people refer to RPI as spec-driven dev because for some people, sdd all it means is i use a bunch of markdown files while i'm coding and forget what's in them i just spec driven those are my specs and i'm using them to drive development there was this open ai researcher who talked about spec driven dev and like hey stop reading the code just write the specs and treat like the coding part is compiling specs into code that part never really materialized maybe with gpt7 you know um but the challenge i'm on a github issue in spec kit uh that has been open for a year and every couple of weeks I get it.

1:03:58There's a new email on the thread of people complaining about this problem. Like, okay, I edit my specs and then I edit the code and the code drifts and the specs. How do I keep the specs up to date as the code is changing? And it's basically like, you now have two sources of truth and it's, it stops being useful. And so like, that's why when RPI, the idea of the docs is they're all for a while we kept them around, but after two or three months, we're like, oh, these are actually like tactical execution docs. I do the research, I do the plan, I do the implementation, I throw the docs out. And the next time I need research, I just do it from scratch because tokens are cheap and my time is expensive.

1:04:30And the amount of time I might waste if I reuse a research that is no longer in sync with the real state of the code base. So we just create it live every time. This is why it's like context engineering still matters. Creating artifacts that compress the state of the code base and compress the intent of the builder into small things that can be reused in the future for the scope of a task. is like a very powerful, like tactical approach, but it's not a thing. Like I have very few opinions on like what sorts of docs that you should leave lying around your code base that are like evergreen. I've seen people try to maintain parity between documentation or specs and the code itself.

1:05:07And I don't think anyone actually like found it very useful. Like you can do it and it works, but it's like the ratio of the effort it takes to keep them up to date. And trivially, you could do this with AI probably, but I've never known anyone who was like, yeah this is great and we're glad we have it like you could do it and it might help but i i don't think anyone found it useful enough to like maintain a system to keep the specs in the code in sync versus just using the code as the source of truth always now you mentioned something interesting which is with context engineering you need to sometimes compact and you've previously talked about intentional compaction that when context is noisy deliberately compress the useful part into a clear like markdown artifact verify it and then start a fresh conversation can we talk about this kind of compaction and why it's important and and it sounds like it's going to be a building block where it already is for context engineering right yeah you know frequent intentional compaction is the building block it is it is completely comes from context engineering is context engineering is like how do we get the most out of today's models how do we change what we're putting into the model into the context window into the agentic chat how do we control that in such a way that we get the best results possible which means doing as much work as possible in the smart zone the you know first hundred thousand tokens of the context window and uh this intentional frequent intentional compaction is basically like okay the research step we're going to go read a bunch of code and turn it into a doc that's our compaction we take that forward in the next session we're going to read we're going to read the ticket and the intent and turn that into a design document that we call it it's like okay here's the high level spec of what we want to do here's the high level like current state desired end state and then a bunch of design questions the model has kind of like a very thorough maybe even over-engineered like plan mode and then you take the research and the design and you do a new session new context when you're like cool you've compressed the intent and you've compressed the state of the code base so that you can then do your planning of like okay we know what the end state looks like we know where we're going now let's break down how we're going to get there all of these different steps of the process exist because models have shortcomings in each of these phases so the research is pretty hands-off i don't read the research docs it's just like go read a bunch of code and then like make a doc out of it models are pretty damn good at that if you ask it to find a bug and have opinions about the code base that's different but if you just ask it what is the intent and how does this stuff fit together uh that's usually pretty straightforward but designing the end state of the of the software the architecture and the program design models are not great at they make a lot of like they make decisions and sometimes they're right and sometimes they're wrong so we want to have a human in the loop there and then the steps to get there i we talked about this before, but models love making what I call like horizontal plans.

1:07:42If you ask a model, like build a plan of steps to go build this app, it's like, cool, we're going to do the database and then we're going to do the services layer. Then we're going to do the API and then we're going to do the front end. And it's like, well, that actually kind of sucks because we're going to be on the other side of 2000 lines of code. And let's imagine this is an existing code base, right? We're going to make changes to all these different parts of the system. I can't test it till the end. And so what I would do is like, okay, how would I have built this if I were building my hand?

1:08:03Well, okay, I would probably create a mock API endpoint with fake data and then I would go kind of get the front end kind of how I want it to look and then I would actually go like build a services layer and actually wire the data through and then I would make a database migration and make my new table and then I would actually add a lot of business logic and then I would add a bunch of error handling and it's completely orthogonal to how model like models will write the database layer and all the error handling without ever like anyone's ever touched or seen the code or whatever it is and so this is another place where we like to have humans involved because humans have really good taste and judgment and like I would rather read five separate little mini diffs of like things that i can manually verify and explore then read 2 000 lines of code and be like well it's not working i don't know where you don't know where because you wrote the code you were supposed to get it right we talk about compaction context engineers like how can you stay in the smart zone of the context window which is again the dumb zone i will say disclaimer it's really good training wheels if you don't have intuition about this so let's just define these things what is the smart zone and what is the dumb zone so it's it's it's a little bit blurrier than like i would like i would like it to be i think in november we we talked and said oh that's about the first 40 of the context window but then we had million smart zone yeah then we had million token context windows so then i changed it to like the first hundred thousand tokens if it's a really like 4.8 i usually will go up to like 200k but basically the thing jeff only had and ralph wiggum was like the less context window you use the better outcomes you'll get and basically the smart the smart zone mean meaning if you have context in that first part it should work a lot better and then like the dumb zone is like once you have stuff there it's kind of forget about it like it'll be confused it's not going to do much like it'll degrade yeah and there are times and this is an intuition thing like i will often go up to three four hundred k token four is rare but i will go up to 250 300k tokens for certain types of work where my intuition tells me that i can keep working without without degrading the performance but if you don't have good llm intuition like 100k for smaller models 200k for these like really beefy like codex and opus 4.8 models is usually a good like training wheel guideline of like if you pass there your quality of results may be degrading the biggest tell i see for this is often the uh models trying to get the test to pass and your 200k token well let me try this okay let me try that and it's like trying a bunch of stuff and it's getting more and more extreme and it's like thing oh let me delete your dot end file and try again like this is where things get really really weird and so it's like if you start to see certain types if i'm like oh we're at 300k tokens and i need to like fix the unit test i'm like cool write everything we did to a file or even i'll just do like a built-in compaction depending on the model and then i'm starting a new session at 30k or 50k tokens i'm like cool we're gonna do a hard thing which is you're gonna get this freaking test to pass and you're not gonna be stupid about it by the way one thing that you said like about the the model being dumb is you said that if the model ever tells you you are absolutely right you should start over and we've all had that when it tells me like oh you know you didn't you're absolutely right and i'm like we just get annoyed but why should we start over what's happening there in your observations yeah that's great yeah and the new the new you're absolutely right i think is uh you're right to push back on that right yes that's opus right yeah opus is like you didn't run the test did you you're right to push back on that i totally did it but no for me You're absolutely right was always what the model would respond.

1:11:25If you were like, that's totally wrong. You did it. Like if you said something where you were angry or frustrated or just wanted to point out that it's done something wrong, it would respond with you're absolutely right. And most of us have had the experience of it says that and then it continues to do the wrong thing. So it's like once it starts doing dumb things, because there's four things in your context window that matter. There's like the size of it. How many tokens? There's like the quality of the information is like, is there any incorrect information? Like if the model had some thinking trace where it decided the wrong thing was true, is there missing information?

1:11:57Does this is like have context missing that it should have? And then there's the trajectory. And the trajectory is very subtle, but you may have had sessions. The trajectory meaning you're prompting? The actual history of everything. I call it trajectory is like the actual history of like what the agent has done in the past. Yep. And so if I say, hey, make this change and the agent makes the change and then it runs the test and then they're broken and then it fixes the test. I have very high confidence the next change I asked it to make it's going to follow that path again because it's like okay here's a conversation and the last time the user asked me to do a thing I made the change I ran a test test broke and fixed the test and then I told the user but if I say make a change and it makes the change it doesn't run the tests then I'm on a different trajectory and if I say okay make another change it's like basically the they're auto-regressive so they're predicting the next what's the next message in this conversation and so the example we talked about in uh no vibes allowed was of course they're like, hey, the model makes a mistake and then you yelled at it and then it made another mistake and then you yelled at it.

1:12:50And then it's like, cool, what's the next message in this conversation? Well, look, if I read the history, I should probably make another mistake so the human can yell at me. So I was like, OK, that's a great that is a great example of like time to start over. Let's talk about some observations on how software engineering is changing. One thing you talked about recently on the evolution of the coding meta is going from token harder to token smarter can we talk about what you mean by token harder and token smarter yeah so token harder is i mean i'm in a i'm in a group chat called hyper engineering and it's all like people trying to max out their cloud subs oh wow okay it's just like okay how do i keep my yeah that sounds like a fun is a fun place it's a fun place but it's like all token harder it's like look at all the side projects i built it's look at everything that uh i i've gotten my cloud token i've got six six cloud code accounts i've gotten all of them maxed out every five hour period i've timed it out so i always use all the tokens and it starts up immediately when the limit resets and so it's like ed mean getting into eli goldrad and the goal is like optimizing for utilization and efficiency of one node in your factory rather than the end-to-end goal of like how do we ship value and things that people like that are stable and like will last a long time but that's my idea of token harder and it's the same thing with the dark factory thing is like hey, if you remove humans from code review, you can push more tokens through the system.

1:14:09So we talk about software factories, but what is the dark factory? Ah, so the dark factory is this, comes from this idea of like, there are factories where everything is automated by robotics. So you can imagine like a car factory where it's all robots building the cars and they don't have lights because there's no humans. Oh, so that's where it comes from. Dark factory. Yeah, you walk in there, there's no lights. There's not even light switches. So it will be the fully automated software factory where it will be like no human input, basically. No human input. Raw materials go in, cars come out.

1:14:40Yep. And I think in a micro, like you can have many loops that are dark in your thing of like, hey, if the code review agent comes back with a problem, you loop that back to the builder agent, it fixes it and comes back. And that's dark. You don't need a human loop for that. But the full dark factory where you don't read any code, yeah, it's a good way to maximize your token utilization. It's like if your belief is like, my job is to extract as much intelligence out of the machine god as I can, because that's how I get the most value and the most leverage on my time, then token harder. And my take is basically what we talked about before.

1:15:12Token smarter is like, okay, how do I move faster? How do I get as much value out of AI as I can without having to turn the lights off while still maintaining control and taste and judgment and understanding the system architecture and having a lot of like applying my hard-won opinions through 10 years of software engineering to the design of the program so that I can feel confident that the code's going to get better and more maintainable over time. It's the same thing of like, you look at like the SRE team inside Google. They brought out this book, SRE, Site Reliability Engineering. And the whole take was like, hey, we're going to go from one data center to five data centers.

1:15:48And we need the same six person team to be able to manage five data centers. And we need the same six person team to be able to manage 50 data centers next year. And it's basically how do we apply software to this problems so that instead of scaling linearly of like okay every data center needs five devops people so we need to scale the people with these things how do we continually automate the parts that we don't need so a little bit orthogonal and maybe even like contradictory to what i just said but this idea of like how do you find leverage and the way the way well i think what you were saying there is like from google that never seek to remove those sres from the process at all they just said like look can we think ahead and scale yourselves and they actually grew the team it wasn't actually six people it was more like i think google specifically said okay we have five data centers next year we'll have 50 you're six of you we do not want to have 60 people we don't want and and then management leader and all that is like how can we do it with like 12 or like or like 10 and then one will have 500 and now actually their sre has grown but but of course yeah but but they never you know i think as engineers like we feel pretty threatened when someone says like all right we just want to have zero engineers like i mean that's not a fun place to work out but it sounds like it's not a possible place to work at if they have zero engineers neither of us could work there right but don't understand the token smarter is like let's keep humans in the loop let's keep adding value and figure out what are the parts which are not as relevant boring where we don't need it and so like one developer can probably do more than before but you are built to like be part of this whole thing and the lights are on in the factory yeah and it's like basically i think i think what i'm trying to get to is like the connection here is like sre built It's a thing where like headcount scales at like a square root function or a logarithmic function, whereas their output scales like linearly.

1:17:32And you want to say that the way you do that is with good architecture and good program design. And so in order to like avoid this problem where you have to throw more people or more tokens at the problem, if you design good software in such a way that it gets more maintainable and more scalable over time. And like just today, it doesn't feel like like basically you need humans in the loop to be able to do that. Let's talk about. uh ai slop at one point you wrote yeah ai can write your code but it can also write your specs and prds but the same the same rule is always slop in slop out if you outsource your thinking you're gonna get garbage yep um so yeah that's basically the idea is like the way we think about like getting high quality outputs is like yeah you could write the code by hand or you could sit with a model and work back and forth and go maybe a little bit faster and you have control and every time it makes a change you go read the change and if it's bad you tell it nope we want it like this and you kind of incrementally slowly this is like kind of the stage two or stage three version of working with agents where like the agent's writing all your code but you're kind of very much in the loop and this will make you go faster but it won't make you go that much faster it won't make you go anywhere near there's like there's like that level and then there's like the maximum speed you can go while still caring about the code and then there's like the maximum speed you can go if you turn the lights off and so we always think about it as like in terms of leverage is like okay let me take everything starts with like a sentence or a voice note ramble.

1:18:54Like I want to build this thing and it's going to work like this, whatever it is. Let's say like on average, like two sentences, I got to fix this thing. Or there's a support ticket. I got to fix this thing. If you can turn that with AI into a one pager and then turn that one page and make sure that's correct. And then turn that one pager into a three pager and make sure that's correct. And then turn that three pager into a 10 page, like detailed outline. Then you can write a hundred pages worth of code. And it's maybe not perfect. You shouldn't like sweat over these documents and make sure they're perfect but you're increasing the chance that like you're decreasing the uncertainty of the outputs it's like you can think of like you have like a line of like where it's going and then you have like the probabilities of where like it might go in that range if you are kind of reviewing along the way as you get more and more detailed into how what you're building and how you want it to be built you kind of collapse the uncertainty and the set of end states that you could land in that's me doing the physics thing of like you got to superimpose all these probabilities and like i don't know i have this thing that like i think people who really like playing real-time strategy games uh are probably going to be really good with ai because you kind of have to like i don't know matt pocock was just talking about fog of war and like things that are at the frontier of like there's stuff we don't know about this problem yet how can we find that out and how can i make the best decision now knowing what i have seen there's a i've seen a couple pieces of information and so there's a 30 chance it's this and there's a 40 chance that it's this how could I get more information so in my head I can like recalculate those probabilities and decide what's the most likely path that's going to lead us to success.

1:20:24Speaking of the most likely path that leads you to success, let's talk about your company that's, you've, you've just come out of stealth, human layer. What is human layer and what is the probability that you're setting up for success? That's a good question. A hundred percent, a hundred percent probability, uh, Maybe 110. But no. So Human Layer is an AI IDE. It's a collaboration platform. And it is building blocks for your software factory. And the basic pitch is like engineers solving hard problems and complex code bases. Basically, there's two categories of builders. There's like vibe coders building side projects.

1:20:59And then there's people building production software where the stakes are high. And if something breaks, we're going to get fined millions of dollars. Or, you know, we're going to lose millions of dollars of money for the company. And there's a whole spectrum in between there. But it's like, if you're kind of in the left half of that spectrum, you're building software that matters and it has to last and be around for a while, then we are helping people like that solve problems two to three times faster without descending into slop. It's like, how do you maintain that near human level of quality and move two to three times faster?

1:21:27And what were the ideas that you built and that you came with? One idea that we're really excited about right now, I mean, it all comes from this RPI and this like using specs to like, I mean, I've kind of been hinting at it this whole time, right? Of like, okay, cool. Like start really high level and zoom in layer by layer and re-steer and like find that leverage that helps you move faster and increase the chance that your agent's going to build exactly what you want or something that's really high quality. The other thing I think that's really interesting that where I just posted yesterday, I said, hey, chat, should we kill the pull request?

1:21:58And that's something I can't talk too much about, but basically the idea is like the IDE of the future needs to be rethought from the ground up for agents. And it might not even be a like, I don't know, a lot of editors kind of started with the text field and bolted on an agents tab. And then eventually you've seen like cursor three. I can't even find the text editor. I know it exists. People have told me you can get to a text view of files, but it's also very agent first. And so we started from the ground up of like, what is an IDE design for helping a developer interact with and manage the work of agents?

1:22:29And then we zoomed out and said, how do we make this collaborative and build in a sync engine and durable streams and all of these like pieces of tech that enable me to get human input and feedback on what I'm doing with agents in real time rather than waiting for the pull request time. And great engineering teams have been doing this for decades of like, hey, we're going to have a design review where we're going to talk about how we're going to build the thing as like a two page Google doc or whatever, 10 page, whatever, however. PRD, ERD. Yeah. AR architecture requirements document. And then you go to sprint planning and you break it down into little tickets and you decide who's going to do what.

1:23:03It's like AI can help with all of this. You should if you're just using AI to write the code, you're missing out on a lot of the benefits that AI can bring to your SDLC. and a lot of people say like well we don't need any of those meetings anymore because we have the loop we have the dark factory things just fly around the loop but it's like okay but if you want to actually move faster and maintain quality then like you should have these checkpoints before you go to actually write the code and you should use ai to help with that so we built this like cloud platform that's kind of has like a google doc style component where you can comment and the agent can surface like mock-ups and mermaid diagrams and html and all these things so basically how do we make agents like bigma style every everything's in the cloud everything's collaborative I see all my co-worker sessions.

1:23:40They see all of mine. It's almost like the benefit that Slack had over email was that you didn't have to be in every conversation to know what was happening. You could maintain, you could see all these channels light up. You could check on them. Okay, I don't care about any of that. But if you saw a conversation that you cared about, you could jump in on that. And it's like, how do we do that for engineering work versus like we really had these like very strict, even when we called it agile, it was very waterfally like PRD, ARD, tickets. everyone goes and builds for a day and then you get the pr back and then one person reviews it how do you create this more just like soup and like what is the data model for that world where you have like agentic traces you have documents you have tasks and projects that group these things you have actual git diffs being streamed everywhere where it's like why would i review all the code at once when i can just always every everybody's work lives in a shared environment that anyone can go interact with.

1:24:36I mean, what it reminds me is like what GitHub did to software team. Before GitHub and its competitors, you might have a tracker somewhere, but most teams were just kind of like inside the company. You didn't know what one team was. I remember pre-GitHub, like, you know, you had individual teams. Some of them had like a board with stickers, but no one else in the company knew what they were doing. They were all working in isolation. And now when you have GitHub or even the internal version of GitHub, inside a company you can always see when you go to a team you see the pull request flying you can join in you have history it's all it is all kind of connected and it came together and now it's like you know for a very long time i was like you duh you're gonna use github or or people will copy it so do i sense that you're trying to build something like this this workflow for like when you have the software factories which are like dark factories and loops at a bunch of places how can we have this this new way of working which which will feel natural but like coming up with it like is hard work and it's counterintuitive how can we do something that accomplishes what github did but like 10x better like more can specifically like more continuous and more real time and more collaborative than like these discrete units of work that is like the pull request well i now i'm starting to understand why you're saying maybe we should kill the pull request because pull request was invented by github right like it's it is not part of git but they do it as a way for you to do a code review merge before it goes in and be able to modify it or like just reject it or etc and it's probably a lot better than whatever we had before which i guess was like emailing your git patch to linus and ask him to merge it into the kernel or whatever they still do it it works but it works for them that's the point but it only works for them yeah i don't know anybody else who does that i mean i'm sure even before get up for you you guys had what like cvs or cvs so if you have a lot of money for microsoft they made us use subversion at in undergrad because the guy who invented subversion uh was a u chicago guy the year after I graduated, they switched everybody to Git.

1:26:34And I was like, damn, I learned a useless thing just for somebody's ego. Specifically for AI startups or startups building on top of AI or building AI products, how important do you think location and network is, especially you are based in the valley. We see research that AI startups are more frequently funded from here than normal startups as well. Do you see this advantage? And also, So do you see some disadvantages of being a specific, may that be Silicon Valley or elsewhere? I don't have really strong opinions on this. Actually, like Paul Graham gave a talk in Sweden about why SF is cool.

1:27:08Rather than just regurgitate that, I will forward people onto that one. We can put it in the show notes or whatever. But he talks about all of the dynamics of Silicon Valley and the paid forward culture and the like people take you way more seriously just because you're based here. I lived in Chicago for a long time. I have a lot of really good friends from high school, from college, from going up in L.A. and never before have I felt like so locked in with like my people have more never have I felt more seen more connected like there's just so many people here again talking about the founder thing people who care deeply who are incredibly competent who like we all have all the same types of problems we love all the same types of things like I don't do land parties where we play video games but all my buddies will come over and we'll sit in the office till 11 we'll just do co-working and like hack on cool fun fun projects and stuff and like you can't do that anywhere else there's not enough like a critical mass for that to just happen organically everywhere you go and and i absolutely love it i wouldn't trade it for anything yeah i think the critical mass is it nails it on on the head when it comes to hiring what types of folks are you hiring for specifically because i'm interested in how hiring changes and and what what a standout engineer means for you and how you are trying to you know confirm that those traits exist in general we we are looking um for people who are have really strong software fundamentals so understand distributed systems understand like the core fundamentals of cs and operating systems and these kind of things i mean you don't have to be a phd in freaking kernel design or whatever but it's a lot easier we can we can teach we can teach somebody i think to be a really good ai developer in a few months you can build enough intuition where you are you know accelerated off the ground and you can go like keep growing there it's really hard to teach someone a cs undergrad program and in three months and what's the problem space that you're excited about and in software engineering or even product engineering or building products that you think in the next few years is going to be one of the interesting things that you're going to be attacking my co-founder could talk more about this but like there's a lot of interesting things happening in in real time and cloud and sandboxes in sync and kind of like using these new building blocks that have gotten really solid in the last couple years we're big fans of the electric sequel team we're users of durable streams it's like how can you build systems that kind of are a lot more spread out and distributed and almost like decentralized this is really interesting for coding because you want to be able to run coding agents anywhere you want to be able to run them for a short time for a long time on demand on a schedule all these things and have them all be part of this kind of like brain so i don't know parts of what we're doing are really boring like all our data is in postgres and then parts of what we're doing is really interesting um but there's a lot of distributed systems problems there's a lot of infrastructure problems like we are building tools for ai but there's a lot of problems in building collaboration platforms that are really really hard and there's a lot of new tech that makes it easier and more interesting but it's still uh by far from an easy problem it sounds like what you're saying is like the infralayers it's some extent a new infralayers being built and it'll take some time and but it'll be like just new new blocks and it will eventually become the primitives like for cloud we have from it already but it took freaking decades to get those together or more.

1:30:15Yeah, you had AWS in what, like 2008? 2006. Yeah, and then you got Kubernetes a decade later. And as closing, what's a book or reading that you would recommend? Something that you personally enjoyed? Nowadays, we talk a lot about Refactoring by Martin Fowler. Classic. I think it's because we spent a lot of time improving the design of existing code and trying to figure out how to get models to build code that is easy to maintain and easy to read and easy to understand and easy to build on. I feel like I probably have a better answer than that, but that's what's top of mind these days. We're reading a lot of classics of software engineering.

1:30:51Refactoring, clean code, the pragmatic programmer, all that stuff, I think is more relevant now than it has ever been. Love it. Well, Dex, thanks so much. This was fun. This was a blast, dude. Thanks for having me on. This was great. I had a lot of fun.

1:31:03Dex Horthy:I don't know about you, but I really enjoyed this conversation. Dex is such a big believer in gender coding, yet he's the one warning us that if you stop reading the code, you have about three to six months before your codebase becomes easier to rewrite than to fix. And this comes from Furt has experience. His team built a light soft software factory, ran it, and then had to shut it down. I also like the idea of the slow loop. Loop engineering feels like a somewhat meaningless term to me. What Dex's team does is actually pretty boring. A cron job runs every night, fixes one issue or one anti-pattern, and opens one small pull request.

1:31:38Dex Horthy:The team wakes up to a code base that's a little bit better every morning, and devs still need to review and prove it. This is a practice that honestly any engineering team could just adopt today. Finally, I really enjoyed the history lesson. The term software factory comes from a NATO conference in 1968. The idea of software used to build software with analogies to a factory is more than 60 years old, and every generation of our industry has tried to automate more of the loop of building software. AI agents are just yet one more attempt, although probably the most successful one. Do check out show notes below for the related The Pragmatic Engineering Deep Dives that go even deeper into AI engineering and other related topics.

1:32:14Dex Horthy:If you've enjoyed this podcast, please do subscribe on your favorite podcast platform and on YouTube. A special thank you if you also leave a rating on the show. Thanks and see you in the next one.

From the publisher

Brought to You By:

• Antithesis – verify your system’s correctness without human review or traditional integration tests – and avoid bugs or outages.

• Buildkite – CI software built to absorb whatever your coding agents throw at the build queue.

• Sentry – application monitoring software considered “not bad” by millions of developers.

—

Knowing how LLM contexts work and how to work around context limitations – aka “context engineering” – is becoming more important for software engineers working with LLMs. Let’s look into what works and what doesn’t, today.

In this episode of The Pragmatic Engineer podcast, I sit down with the CEO and cofounder of HumanLayer, Dex Horthy, who coined the term “context engineering”. We discuss the ideas behind this context engineering, harness engineering, loop engineering, software factories, why his approach to AI-assisted software development has evolved, and how HumanLayer is helping engineering teams automate more of the software development lifecycle without sacrificing code quality.

—

Timestamps

00:00 Intro

01:33 Dex’s path into tech

03:34 Early work in platform engineering

05:28 Replicated

11:24 Metalytics

12:36 12-factor agents

18:27 Context engineering

23:38 Harness engineering

26:11 Context overload

30:45 Loop engineering

44:34 Software factories before and after AI

50:33 Automation limits

55:18 Three options for automating

59:00 RPI framework

1:04:16 Intentional compaction

1:11:48 Token harder vs. token smarter

1:16:44 AI slop

1:19:15 HumanLayer

1:29:09 Book recommendation

—

The Pragmatic Engineer deepdives relevant for this episode:

• How Uber uses AI for development: inside look

• Are AI agents actually slowing us down?

• AI Tooling for Software Engineers in 2026

• Vibe Coding as a software engineer

• How Claude Code is built

• AI Engineering in the real world

• The AI Engineering Stack

• How AI-assisted coding will change software engineering: hard truths

• The creator of OpenClaw: "I ship code I don't read"

—

Production and marketing by ⁠⁠⁠⁠⁠⁠⁠⁠https://penname.co/⁠⁠⁠⁠⁠⁠⁠⁠. For inquiries about sponsoring the podcast, email podcast@pragmaticengineer.com.



Get full access to The Pragmatic Engineer at newsletter.pragmaticengineer.com/subscribe

More from The Pragmatic Engineer

All 45 episodes
Context engineering with Dex HorthyThe Pragmatic Engineer · 1 h 32 min
Listen in VO